← All 10 studies
Field guide · Audience intelligenceViral Nation2025 to 2026

Criticism Into Creative Direction

How sentiment, LLM topic analysis, and human review turned scattered criticism into a simple “do this, not that” brief, then checked the lesson against a later campaign with 86.8% positive sentiment.

Methods and toolsPythonXLM-RoBERTaLLM topic analysisHuman reviewCampaign measurement

Select creator posts were filling with negative comments. The client could see that humans were upset, but not what to do next. A red sentiment score did not say whether viewers disliked the product, distrusted AI, felt tired of ads, objected to the creator, or were arguing about something else. Every explanation led to a different response. Without that distinction, the safest choice was also the least useful one: watch the number and guess.

This is the gap between social listening and audience intelligence. Listening counts what humans said. Intelligence explains which reactions matter, what caused them, and which decision should change. I led the work to close that gap. The method combined machine sentiment, LLM-assisted topic analysis, stable labels, and human review. The end product was not another chart. It was a short creative rule: do more of this, avoid that, and watch these risks.

The method now operates at much larger scale. The reporting system contains 956,729 analyzed audience comments across 20,466 posts. That figure shows the reach of the reusable system, not the size of this one campaign study. The later campaign check in this article used 1,664 captured comments. Keeping those two levels separate matters because scale and proof are not the same thing.

Sentiment is a warning light, not a diagnosis.

Sentiment sorts language into broad feelings such as positive, negative, neutral, or unclear. That is useful for finding a change. It is weak at explaining the change. Two posts can both be 30% negative for completely different reasons. One may trigger a real product concern. The other may attract criticism aimed at the creator. Treating those cases as equal would send the client in the wrong direction.

The target of the feeling matters as much as the feeling itself. A comment saying “mind your own business” sounds hostile, but it may be defending the creator against another commenter. A polite question about photo access may signal a serious privacy concern. A strong system therefore needs stance: what position the person takes, and what that position is aimed at.

The same negative label can require very different action
What the comment targetsWhat it may meanPossible response
Product or privacyThe feature feels risky or unclearShow choice, access, and control
AI usefulnessThe tool feels forced into an easy taskStart with a real need; make AI the helper
AdvertisingThe post feels scripted or over-soldUse a more natural creator voice
Creator or harassmentThe reaction is not about the campaignDo not rewrite product strategy around it

A total can hide the decision.

An overall score blends every creator, platform, post, and conversation into one number. A large creator community can dominate it. A viral argument can make one post look like a product failure. The useful unit is often smaller: a theme within a creator-platform group, connected to the exact post and message that caused it. The goal is not to make the number more complex. It is to keep enough context to make the number honest.

Build the analysis as a chain of checks.

Comment data arrives messy. The same text can appear more than once. Replies depend on a parent comment. Some rows contain only tags, stickers, or emojis. Platform fields do not always match. Before asking a model what humans mean, the pipeline has to preserve the post, creator, platform, reply link, likes, and original text. Those fields are the evidence around the sentence.

The workflow separates repeatable machine work from judgment. Code cleans and joins the records. A sentiment model supplies a first label. An LLM finds and codes the themes that need language understanding. Fixed rules handle empty or low-signal text. Human review checks the places where an error would change the recommendation. Each stage writes an output that can be inspected instead of hiding the whole process inside one prompt.

Raw commentsClean contextSentiment and stanceFrozen topicsHuman QACreative action
Each stage makes one kind of decision. Keeping the stages separate makes errors easier to find and the final advice easier to defend.

Do not ask one model to do everything.

One large prompt can label sentiment, invent topics, summarize the result, and recommend action. It can also change its labels halfway through, merge unrelated ideas, and produce totals that cannot be reproduced. Small stages are slower to design but safer to run. A failure has a location. The team can rerun topic assignment without changing sentiment, or correct a sentiment label without rebuilding the raw data.

Let the model discover language, then freeze the rules.

A fixed topic list created before reading the comments can miss the thing humans actually care about. A fully open LLM can find new ideas, but it may call the same concern “privacy,” “data access,” and “photo scanning” in three different batches. The method needs both freedom and control.

The discovery pass keeps useful, unique comments and asks the LLM for repeated topic phrases across the campaign. A second pass merges close phrases into a short taxonomy. That taxonomy is then frozen. During final assignment, the model must select one approved topic. It cannot create, rename, or skip a label. Low-signal rows such as empty comments, tags, and emoji-only reactions use simple local buckets instead of spending model time on text with little meaning.

Open discovery
Read comments across the campaignFind repeated ideasMerge overlapping language
then
Closed assignment
Freeze the reviewed taxonomyAssign one approved topicReject unknown labels
Discovery is open because the audience may raise an issue nobody expected. Measurement is closed because a stable total needs stable labels.

Topics explain the “why.”

Sentiment can say that a post is receiving criticism. Topics reveal whether the criticism is about AI being unnecessary, data access feeling intrusive, sponsored content feeling fake, or the creator being disliked for an unrelated reason. The first three can shape the brief. The last one should usually be reported as context, not turned into product advice.

Review the mistakes that could change the advice.

A random accuracy sample is useful, but it is not enough for a high-risk read. Rare negative comments can drive the client decision even when the overall campaign is positive. Replies, sarcasm, mixed statements, and criticism quoted by a supporter are also easy to misread. Quality checks should therefore follow decision risk, not only statistical neatness.

In the later campaign, all 78 negative-coded comments were read by a person. They separated into 21 AI, privacy, or product concerns; 15 advertising or authenticity concerns; 17 creator or personal criticisms; 10 cases of hate or harassment; and 15 ambiguous or unrelated negatives. Only 36 comments were clear campaign, feature, or execution complaints. That distinction changed the story from “78 humans rejected the campaign” to a much narrower and more useful set of concerns.

The positive side also received a targeted check. Review covered 202 unique positive comments, including every positive label containing risk words, the 50 most-liked positives, samples across creator-platform groups, and comments marked as questions or criticism. No clear campaign-negative comments were found inside that group. This does not prove every positive label was perfect. It tests the failure that mattered most: hiding campaign criticism inside the positive total.

Review every rare risk

When a small class can change the recommendation, read the full class instead of trusting an average.

Sample the large class

Use targeted checks to look for serious errors hidden inside the much larger positive group.

Keep the context

Read replies with their parent comments and connect labels back to the post, creator, and platform.

Preserve corrections

Store manual overrides as history so the team can see what changed and improve the next run.

Translate every theme into a choice.

A theme table is still not a strategy. The last step asks what the creative team can change. If humans question why AI is needed, begin with a problem worth solving. If the technology feels cold, begin with the person, memory, or relationship. If photo access causes concern, show that use is clear and optional. If sponsorship feels fake, keep the creator's normal voice and make the disclosure natural.

The main insight was simple: humans were more open to AI when it helped with a meaningful or difficult task. They pushed back when AI appeared to replace taste, craft, or an action that already felt easy. The new creative direction therefore put the human reason first and the tool second. AI became the supporting mechanism, not the star of the story.

From audience language to a usable creative rule
Audience signalAvoidDo instead
“Why does this need AI?”Lead with the featureLead with a real human need
The result feels less heartfeltMake AI the heroKeep the person and memory at the center
Photo access feels unclearAssume trustShow choice and control plainly
The post feels scriptedForce a product speechUse the creator's own story and voice

A good recommendation can be used before the next post ships.

“Sentiment was negative” describes the past. “Open with why the memory matters, then show how the feature removes work” changes the future. That is the test for an insight. A producer should be able to use it when choosing a hook, writing a brief, or reviewing a draft. If the statement cannot change a choice, it is still analysis, not guidance.

Check the lesson in later work.

A useful idea should create a prediction. Here, the prediction was that human-first creative would draw a stronger response than creative that opened with the mechanics of making content. A later related campaign gave the team a chance to check that pattern across 15 concepts, 55 posts, and 1,664 captured comments.

Concepts that opened with a person, memory, or milestone averaged a 4.52% engagement rate. Concepts that opened with sorting, editing, speed, or unshared photos averaged 1.36%. The second group did buy cheaper reach: $0.09 per view compared with $0.14. That tradeoff matters. Human stories were better at earning response; problem-first openings were better at buying inexpensive views. The best choice depends on the campaign job.

The later comment set was 86.8% positive: 1,444 of 1,664 captured comments. Only 36 comments, or 2.2%, were clear campaign, feature, or execution complaints after manual review. These results fit the earlier creative lesson, but they do not prove that the lesson caused them. The concepts were not randomly assigned, the campaigns differed, Facebook comment coverage was incomplete, and one creator community supplied 43.7% of the captured comments.

Audience reactionExplain the themesChange the briefMeasure later workRefine the rule
The loop ends only when later work is measured. A result that fits the prediction strengthens the lesson; it does not turn an observational comparison into a controlled experiment.

What the evidence supports

The reusable comment system has analyzed 956,729 non-creator audience comments across 20,466 posts. That proves operating scale. The later campaign study contained 1,664 captured comments and 15 manually coded concepts. It supports the measured sentiment, complaint mix, and directional difference between human-first and creation-problem openings.

The evidence does not support a causal claim that one brief produced the later result. It also does not mean that 86.8% of humans directly approved of the company or AI. Many positive comments were about creators, families, relationships, and memories. The result describes the captured creator-community response. It should be read with its coverage and concentration limits.

86.8%

positive sentiment in the later campaign

Measured
1,664

later comments coded

Measured
84.0% vs 62.6%

human-story vs problem-first sentiment

Measured
4.52% vs 1.36%

human-story vs problem-first engagement

Measured

The most valuable output was not 86.8%. It was the path from “humans are upset” to a decision the client could use: lead with a meaningful human need, let the product remove real work, keep AI in a supporting role, show control clearly, and make the sponsorship feel natural. The models made a large conversation readable. Human judgment made the answer safe. The later campaign gave the team evidence that the direction was worth keeping.