All case studies
Nova Dynamic Media · in progress·Sole owner of the AI platform

Grouping Live Audience Questions With an Agent

Clustering thousands of event questions on Amazon Bedrock AgentCore

Bedrock AgentCoreLangGraphAmazon Nova 2DynamoDBTerraformPython
7,857 → 69
questions collapsed to distinct meanings, one dev session
~100x
dedup saving on embedding the largest session
0.985 AUC
Nova 2 on a hand-labeled paraphrase benchmark
~90s
warm run on a 69-text session (down from ~16 min cold, unpooled)
Read-only
the agent never writes the platform's data

Options on the table

Keyword / exact-match grouping
Misses paraphrases and cross-language duplicates. 'When does it go on sale?' and its Cantonese phrasing never meet.
Embedding + agglomerative clustering
Group by meaning, not words. Complete linkage makes the similarity floor a hard promise every pair in a group clears.
Titan v2 embeddings
At its shipped floor, recall on genuine paraphrases was zero; it matched near-identical strings, never rewordings.
Amazon Nova 2 embeddings
Embeds across languages directly, so a Chinese and English version of one question land close without a translation crutch.

Architecture at a glance

Load
scoped, read-only
Normalize
strip + filter
Pivot
translate, cached
Embed
Nova 2, cached
Cluster
complete linkage
Report
medoid + demand

Background

Nova's live events already push real-time questions to moderators. The missing piece was sense-making: when thousands of questions arrive, which concerns actually matter, and how many people share each one. This agent is the answer, and it is the piece I am building now on Amazon Bedrock AgentCore.

The hard part is not clustering in the abstract. It is doing it reproducibly, across languages, on a corpus that is mostly repetition, without ever dropping a genuine attendee question because of a schema quirk or a model hiccup.

Why translate before embedding

The pivot to a common language is load-bearing for two reasons. It collapses Traditional and Simplified and English phrasings of one question onto a single string, so they share one cache entry and one embedding, and it gives the moderator a representative they can actually read. Newer multilingual embeddings narrow the gap on raw cross-language matching, but the pivot still earns its place on cost and on display.

Choosing the similarity floor

The floor is a property of the embedder, not of the problem, so it was measured, not guessed: a hand-labeled set of attendee questions with deliberate near-miss traps (an attendee limit against a poll limit, a free trial against enterprise pricing). The earlier embedder looked fine on a load-test full of literal edit variants but had near-zero recall on genuine rewordings, which is exactly the case that matters. The current floor sits at the benchmark's F1 peak, chosen because under-merging is invisible to a moderator while over-merging is not.

Read-only and fail-open

Two rules keep the feature safe to ship. The agent reads the platform's data and never writes it; grouping is a derived view stored in the agent's own table, so dropping those records changes nothing else. And every model-driven stage fails open: a text that cannot be translated is grouped in its original language, a text that cannot be embedded is left out, and neither fails the run. A worse grouping always beats returning nothing.

What I took away

Keep the deterministic core deterministic and let only the ends be model-driven, both failing open. A moderator can see when nothing was narrowed; they cannot see questions that were silently dropped. So when a stage fails, degrade to a slightly worse grouping, never to nothing.
Want the deeper architecture behind this? Let's talk.
Get in touch