When To Apply AI Moderation to UXR (And When to Keep It Human)

Light

post-banner

By Phil Heuring, SVP, Qualitative Insights at Material

Introducing AI moderation into UX research is often touted for its speed, scalability and efficiency. But often these discussions don’t take into account whether human moderation is more appropriate in each use case. Evidence today suggests that the differences in data between the two moderation approaches are real and impact the outcomes.
For UXR initiatives to tap the potential benefits of AI moderation while maintaining their integrity, brands must know when to integrate technology and when to stay human-centered.

 

AI Moderation and Human Participants

Research and practical experience with human-AI interaction are early and even contradictory at times, but there have been some revelations that can help UX researchers decide when to introduce AI moderation.
Some studies suggest people react more dynamically to humans, showing an increase in positive emotions, openness and extraversion — but also more neuroticism. Others show more restricted vocabulary use when interacting with a conversational AI, though it may also be a judgment-free zone, encouraging greater candor. One specific study shows brain scan activity differs even when people can’t consciously separate an AI voice from a human’s. In this particular case, human speech triggered empathy and memory centers, while AI put internal error alarms on high alert. If this were moderation, it’s a small leap to expect that behaviors and verbal responses would also differ.
This doesn’t necessarily guarantee human moderators offer truer outcomes. Still, so long as differences persist in how humans interact with bots, we must expect that the data we receive is also different.
Despite varying results across studies, one shared finding is that differences between human-to-human and human-to-AI interactions are real. When participants speak to an AI bot, they likely code-switch rather than answering how they would normally. Since human habits won’t change as quickly as language model advancements, this difference will likely continue to be a factor that researchers must take into account.

 

A Case Study Comparing an AI Moderator to a Human Moderator

Bringing this all to a practical level for UX research, Material analyzed a recent study examining concepts for a financial service product. We compared transcripts from participants who were interviewed by both a human and an AI moderator. The same objectives, stimuli and discussion guide were applied for each.
This is one case study and not a perfect A/B study, as in-depth interviews evaluating concepts are generally not identical by design: probing is adapted based on context and real-time moderator judgment. Additionally, participant fatigue sets in sooner with AI moderators, so our sessions with humans could run slightly longer. To help adjust for this, we included twice the participants in the AI moderation transcript sample (N8 vs. N4). Since AI generally achieves a higher sample for the same cost, this places both human and AI techniques in a comparable and practical real-world scenario.
We used two language models to examine the transcript results side-by-side and cross-check findings. Often, the goal of concept evaluation in in-depth interviews is data “richness,” so we prompted these models to define that in four distinct areas: novel ideas, figurative speech, tone/sentiment and vocabulary/sentence structure. The results are summarized in the chart below.
Material+

These analyses show that content and tone vary distinctly between moderation modes.  The AI moderator is judged to solicit clear, comparative and functional expressions, while the human elicits more layered complexity. Both language models generally agree that the human-moderated transcripts offer more texture, expression and personal revelation. As noted earlier, this is one specific example. Whether these differences hold broadly and which moderation mode will best serve your research purposes depends on the objectives and constraints of each specific project.

 

Choosing When to Lean into AI or Human Moderation

Choosing between AI and human moderation ultimately comes down to your project’s objectives and parameters. In concept testing, AI moderation can elicit rapid reactions from a broad audience to quickly de-risk a “go/no-go” business decision. But if the goal is to optimize a concept, unpacking complex user mental models to guide the next iteration is often better suited to human-led moderation.
Below is a helpful tool to help you determine where to lean into AI or humans for moderation. Thinking across your project, judge which characteristics overlap the most with your overall situation.
Material+

Consider “Yes, and” for Your UXR Approach

To borrow a lesson from improvisational theater, the best moderation approach is often not either-or but rather “Yes, and.”
Some of Material’s most impactful projects combine wide-scale AI moderation with a small human-moderated sample. This allows scale and speed with added context and depth, as the human can build on learnings and hypotheses generated from the AI work. This combination of approaches can capture different expressions of truth to inform decision-making.
Across AI and human-led modes, Material designs research engagements to deliver nuanced, actionable insights at the speed your business demands. Reach out to our UXR team to learn what our approach could mean for your next project.