Session library

Can AI Predict if you will go Viral

Paul Greenberg · May 22, 2024

video marketingvideo productionai powered marketing

Join David Berkowitz of AI Marketers Guild for a fascinating conversation with Paul Greenberg, the innovative mind behind Butter Works, a digital video firm that harnesses AI to predict social video success. In this episode, Paul delves into his journey from leading major digital and media operations, like CollegeHumor and Nylon, to pioneering predictive AI technologies in video production. 

Discover how Paul’s venture utilized machine learning, computer vision, and natural language processing to empower creators with data-driven insights for content strategy. 

This engaging discussion not only explores the practical applications of AI in digital media but also addresses the broader impacts of AI on creativity and production efficiencies. 

Whether you're a content creator, a marketer, or just curious about the intersection of AI and media, this episode offers a wealth of insights on leveraging technology to predict content success and streamline creative processes.

[0:30] How Does Predictive AI Forecast Social Video Success and Engagement?

Answer / Description: Predictive AI forecasts social video success by analyzing historical performance data alongside computer vision, natural language processing (NLP), and topic clustering to identify unsaturated, high-engagement content niches. By mapping video metrics across axes of total views, audience engagement, and topic saturation, creators can mathematically determine which concepts to produce before investing resources in production.

Paul Greenberg, former CEO of CollegeHumor and founder of the digital video firm Butter Works, demonstrates this approach using an analysis conducted for Animal Planet. Instead of guessing what would resonate with audiences, the team ingested tens of thousands of animal videos across social networks, transcribing audio and analyzing visual assets. This data was clustered into subtopics (such as dogs, cats, animal babies, and animal births) and mapped onto a matrix: views on the X-axis, engagement on the Y-axis, and volume of existing content represented by bubble size.

The sweet spot for creators consists of high-view, high-engagement topics with small bubble sizes, indicating low saturation. For example, while dog and cat videos perform exceptionally well, their high saturation makes it difficult to break through the noise. Conversely, "animal babies" and "animals giving birth" represented high-potential, low-saturation opportunities where creators could ride the wave of an emerging trend ahead of competitors.

Keywords: predictive AI for video, social video success, content saturation analysis, AWS computer vision, topic clustering, content strategy matrix, Animal Planet case study


[4:13] Why Do Video Tags Fail Compared to AI Computer Vision and NLP in Content Analysis?

Answer / Description: Traditional tags and metadata fail because they are primarily designed by creators to influence platform algorithms rather than provide accurate, objective descriptions of video content. In contrast, AI computer vision and natural language processing (NLP) analyze the actual visual frames and spoken dialogue, eliminating creator bias and misleading tag noise.

Greenberg highlights how simple metadata analysis is insufficient for understanding modern social video dynamics. Creators often populate tag fields with trending keywords that have little to do with the actual subject matter of their videos, creating a high level of data noise. By shifting to Amazon Web Services (AWS) AI tools, the firm could execute frame-by-frame analysis to identify exact visual elements, while using NLP to interpret transcribed audio files.

This deep-data approach revealed counterintuitive insights that standard tags would miss. For example, while "cats" and "kittens" are conceptually similar, the AI analysis proved that cat videos are highly successful, whereas kitten videos underperform. The computer vision analysis explained this discrepancy: kittens move very little, making them better suited for static images, while active cats doing dynamic actions (such as riding skateboards) generate the motion necessary to drive video engagement.

Keywords: metadata limitations, AI computer vision, video transcription NLP, automated video analysis, content tagging noise, social video performance data


[7:59] How Can Predictive AI Optimize Video Length and Formatting for Fuzzy Content Categories?

Answer / Description: For complex or "fuzzy" content categories with less defined boundaries, predictive AI overlays dimensions like optimal video length alongside topic clusters to map out precise production blueprints. By analyzing historical viewer retention, AI can specify the exact duration required to maximize engagement for specific subgenres.

Using a case study of a travel and luggage brand, Greenberg explains how predictive models handle categories that do not have distinct divisions like animal species. The team mapped topics like "packing hacks" and "what to pack" but introduced a third dimension: video length, visualized through color-coded density in bubble charts.

This multidimensional analysis allowed the brand to understand not just what topics to cover, but how long each specific video format needed to be to succeed. The data revealed that a detailed "what to pack" video requires a longer runtime to satisfy viewer intent, whereas "packing hacks" perform better when delivered in a shorter, more rapid format. Combining topic clustering with length optimization provides creative teams with precise, data-driven guardrails for production.

Keywords: video length optimization, multi-dimensional AI clustering, travel video marketing, viewer retention data, topic modeling, video formatting guardrails


[9:50] Which Facial Expressions in the First 3 Seconds Drive the Most Engagement on TikTok?

Answer / Description: TikTok videos that begin with confused, angry, or disgusted facial expressions in the first three seconds drive significantly higher engagement than those starting with happy or smiling faces. Happy faces correlate with lower viewership and engagement, while negative or intense emotions create an immediate narrative hook and air of mystery.

Butter Works utilized computer vision to analyze the emotional states displayed in the critical first three seconds of TikTok videos. Contrary to traditional marketing logic which favors positive, smiling faces, the data demonstrated that starting a video with a happy expression statistically dampens performance. Instead, starting with a confused, angry, or disgusted face engages viewers because it triggers psychological curiosity; users feel compelled to stay to discover the cause of the reaction and its eventual resolution.

The research also coupled these emotional starting points with specific video durations to determine the perfect formula. For instance, the data indicated that a video starting with a disgusted face is most successful when kept to 26 seconds, while a confused face performs best at 31 seconds. While these metrics function as directional guardrails rather than rigid rules, they provide critical structure for structuring short-form video hooks.

Keywords: TikTok hook optimization, 3-second hook facial expressions, computer vision emotion detection, social media engagement analysis, viewer curiosity hooks, short-form video formulas


[15:20] How Can Brands Use AI Object Detection to Choose the Best Filming Locations?

Answer / Description: Brands can use AI computer vision to detect objects in top-performing videos within their industry, indicating which settings and visual backgrounds naturally appeal to their target audience. By identifying frequently occurring background objects, brands can reverse-engineer the ideal environment for their commercial shoots.

Greenberg shares an example of a beverage brand seeking to design a highly successful video campaign. The predictive AI ingested and analyzed successful beverage content, running object-detection models to catalog every visible item. The AI identified recurring items like high-end interior design elements, appliances, wood textures, people seated with laptops, and specific furniture.

Synthesizing this object-detection data made the ideal location choice obvious: a coffee house setting natively contained all of these high-performing visual signals. Rather than setting up an abstract studio shoot, the brand filmed on-location in a coffee shop, maintaining creative control over the casting, script, and music while operating within the mathematically validated visual parameters provided by the AI.

Keywords: computer vision object detection, visual background optimization, brand video locations, reverse-engineer video success, video metadata analysis, retail video marketing


[18:07] How Can Generative AI Workflows Reduce Video Production Time and Costs?

Answer / Description: Generative AI workflows can reduce video production time by up to 85% and costs by 40% when used for rapid prototyping, script ideation, and asset generation. By integrating tools like ChatGPT, Runway, and Eleven Labs, teams can automate highly repetitive pre-production and asset creation stages without sacrificing output quality.

At the entertainment company Meet Cute, Greenberg implemented a generative AI workflow that transformed how they developed short-form promotional videos. The team polled their Instagram audience on favorite romantic comedy tropes (e.g., "enemies to lovers" vs. "second chance at love"). They fed the winning poll results directly into ChatGPT to generate a fake movie trailer script, getting them 80% to 90% of the way to a final draft within minutes.

To produce the trailer, they leveraged Runway for stock-style video generation and Eleven Labs for synthesized audio and voiceovers. This stack slashed a traditional six-to-eight-hour production cycle down to an hour and a half, while reducing hard costs by 40%. Instead of using these savings to downsize staff, the company kept its headcount stable and reinvested the efficiency gains to double or triple their overall content output, scaling from one video per day to three.

Keywords: generative AI video workflow, ChatGPT script writing, Eleven Labs synthetic audio, Runway AI video generator, production efficiency gains, AI-assisted content scaling


[21:30] Does AI Video Automation Threaten the Core Value Proposition of Creative Agencies?

Answer / Description: AI video automation does not eliminate the need for creative agencies, but it shifts their value proposition away from manual labor toward high-level curation, strategic consulting, and creative direction. AI serves as a "co-pilot" that automates repetitive preparation and analysis, but human oversight remains critical to prevent low-quality outputs.

Address-engine and agency leaders express concern that as AI automates media buying and content creation, the traditional agency model is under threat. Greenberg counters this by arguing that AI tools are fundamentally no different from previous technological advancements like green screens or digital editing software. While AI can process hundreds of reference videos or generate initial script drafts, it cannot guarantee high-quality creative output without skilled humans driving the narrative.

If an agency produces poor creative work, backing it with AI insights or automated production will not make it successful. The primary value of the modern agency lies in establishing directional guardrails, editing "janky" AI-generated drafts, and applying taste and strategy to raw assets. Agencies that leverage AI to handle the tedious aspects of brainstorming, pitch-deck creation, and rough drafting can deliver superior results in less time.

Keywords: creative agency value proposition, AI as a co-pilot, human-in-the-loop creative, automated content production, agency efficiency, content curation


[24:00] What Are the Top AI Tools for Video Generation, Optimization, and Social Repurposing?

Answer / Description: Leading tools for AI video generation and repurposing include Runway and Moon Valley for text-to-video generation, Auggie for advanced AI editing, and Opus Clip for automated social media resizing and speaker tracking. While text-to-video tools offer creative control, template and stock-based tools like Fliki and Vista offer highly realistic outputs for faster turnarounds.

The discussion highlights a bifurcated landscape in AI video software. True generative engines like Runway and Moon Valley allow users to build surrealistic scenes entirely from text prompts. These are ideal for stylized concepts, though they can require significant post-production work to stitch together. For social media and business-focused creators, new entries like Christa Wolf's "playday" focus heavily on social formatting and quick-turn creative needs.

For social repurposing, Opus Clip (referred to as "Opus") is highlighted as an exceptionally reliable tool for turning long-form podcasts or webcasts into platform-ready vertical clips. Unlike alternative platforms that struggle with formatting, Opus successfully tracks active speakers, automatically reframes the video to keep faces centered, and overlays highly accurate captions, solving the complex visual challenges of multi-camera editing with a single click.

Keywords: Runway text to video, Opus Clip social media, Moon Valley AI, Auggie video editor, automated video editing, video repurposing software


Answer / Description: Brands should mitigate intellectual property (IP) and copyright risks by utilizing generative AI primarily for internal ideation, rapid prototyping, and non-monetized organic social media marketing rather than final commercial assets. Because pure AI-generated assets cannot be copyrighted, maintaining a "human-in-the-loop" workflow to modify and finish the work is essential for legal ownership.

Intellectual property remains a massive point of friction, with many corporate legal departments actively disabling built-in AI features (such as Adobe's Firefly/AI integrations) due to liability concerns. Under current legal frameworks, pure outputs from engines like Midjourney lack human authorship and cannot be protected by copyright. This presents a double-sided risk: not only could a brand face plagiarism claims, but their own competitors could duplicate their AI-generated assets without legal consequence.

To navigate this landscape, consultants recommend a risk-managed approach. Using AI for rapid prototyping—such as pitch-deck imagery or preliminary script generation—saves enormous upfront capital while keeping the final commercial output hand-made or heavily modified by human artists. For external marketing, keeping AI-generated content restricted to organic, non-monetized social media reduces the risk profile while allowing the brand to benefit from the speed of generative workflows.

Keywords: AI copyright infringement, Adobe Firefly legal concerns, generative AI rapid prototyping, intellectual property guardrails, non-monetized AI content, legal risk mitigation


[37:30] Does the Widespread Use of AI Content Generators Lead to Ideation Commoditization?

Answer / Description: Yes, the widespread use of shared generative AI models threatens to commoditize ideas, as algorithms pulling from the same data pools inevitably produce homogeneous, repetitive content. To counteract this "shrinkage of idea diversity," organizations must retain senior creative talent to introduce unique human perspectives, emotional nuances, and strategic differentiation.

Thomas and other industry panel members point out a hidden pitfall of the AI revolution: a notable compression in the diversity of creative concepts. Because models like ChatGPT and Midjourney are trained on static, historical datasets, they operate by predicting the most statistically probable next step. If every marketer and creator uses the same underlying models for brainstorming, the resulting content naturally converges on a generic, highly standardized average.

This homogenization emphasizes the critical importance of human talent. While companies may be tempted to lay off creative staff to cut costs, doing so leaves them without the intellectual database required to evaluate whether an AI-generated idea is actually good, fresh, or different. Retaining senior copywriters and directors is necessary to spot algorithmic patterns, break conventional molds, and inject genuine emotional hooks that a machine cannot calculate.

Keywords: creative commoditization, algorithmic homogeneity, idea diversity shrinkage, senior creative talent, AI content homogenization, human creative differentiation