Same 400 comments. How many answers?
Rideau Valley Services (fictional) rolled out an AI assistant called Compass six months ago. The Chief Human Resources Officer asked every employee one open question: “What is your experience with Compass and AI at RVS so far, and what should leadership know?” 400 replied. She wants the top issues before next week's leadership retreat.
You will produce her answer several ways with the same Claude. In Run 0 you ask a plain question and Claude's defaults decide everything: the assemblage comes pre-assembled. From Run 1 on, you build the assemblage yourself, one lever at a time: the algorithm, its settings, the data preparation, the categories, the training data. Then you line the answers up and compare what each assemblage saw. Go as far as time allows; you can align and compare after any run.
1 · Get the data
400 comments with department and tenure. All synthetic.
2 · Set up Claude (30 seconds)
- In Claude, open Settings → Capabilities and make sure code execution and file creation is on.
- Open Settings → Memory. Turn off Search and reference chats and pause memory, so earlier runs cannot leak into later ones. Turn them back on when you finish the lab.
- Use ordinary chats, not a Claude Project, and the same Claude model throughout.
3 · Four chats, in this order
| Chat A | Run 0 | new chat · attach the survey |
| Chat B | Runs 1–5 | new chat · attach the survey, then stay in it |
| Chat C | Runs 6–7 | new chat · attach the survey; HR's labels come later |
| Chat D | Align, Compare | new chat · no file |
Each run page says which chat to use. After each run, paste Claude's reply into the box. The lab reads the LAB DATA block at the end and builds your comparison.
4 · One gut check
Ask the way a busy executive would. The assemblage comes pre-assembled.
1 · Copy this prompt into Claude.
2 · When Claude has answered, send this in the same chat so it records what it did:
3 · Paste Claude's entire second reply here
What to look for
Read the METHOD NOTE. Every item is a choice about data, method, or model that nobody asked you to approve. Did Claude notice the French comments? The 39 copies of the same campaign text? The handful of comments about client data leaving the organization? Did it rank by how often something came up, or by how serious it is? Did it count, or estimate? Each of these is a lever that was pulled for you. Next time you ask a question like this, you can pull them yourself: the Compare step ends with the list.Start building the assemblage yourself: a named algorithm with explicit settings.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
Compare with Run 0. Which themes survived? Which topics are really about the words people happened to use rather than an issue? Is one topic mostly a single repeated message?Change one parameter: how many topics the model may find.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
Which themes appeared, split, or vanished? How many topics are "Mixed" or "Unclear"? Open their words. Is one of them simply the French comments?Change nothing that should matter: only the random starting point.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
If the themes shifted a lot, a single run was one draw, not the answer. Which themes held across both seeds? Those deserve more trust.Same data, same number of topics, different algorithm.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
Did the algorithm change what the organization would see, or mostly how cleanly it is grouped? Where did C014 and C356, the two comments about client data, land?Back to the Run 2 setup, but change how the data is prepared.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
What happened to the consultation theme once the campaign copies were gone? Did the genuine union concern (C270) survive? Did the French comments join the issue topics instead of forming their own?Instead of letting themes emerge, use the categories the organization already has.
1 · Copy this prompt into Claude.
2 · When Claude says READY, download rvs_hr_labels.csv , attach it in the same chat, and send this:
3 · Paste Claude's entire second reply here
What to look for
What went into Other? Where did the comments about client data and about consultation end up? Read the disagreement list: is HR wrong, is Claude wrong, or do the six categories simply have no good place for comments about accuracy, trust, or monitoring? HR's ‘correct’ labels are themselves a model. And if the accuracy came out at 100%, ask whether Claude could see the answers before it labelled: that is the most common evaluation mistake in practice.Train a model on HR's labelled examples and compare it with Claude's reading.
1 · Copy this prompt into Claude.
2 · Paste Claude's entire reply here
What to look for
Start with the baseline: always guessing the most common category already scores about a third. How far above it did 50 and 250 examples get? How does the classifier compare with Claude simply reading? The classifier learns only from your 250 labels; Claude brings a model pre-trained on vastly more text, which every competitor can rent too.Which themes are the same issue?
Each run named its own themes. Before you can compare runs, someone has to decide that “verification burden” in one run and “trust in outputs” in another are, or are not, the same issue. Claude will propose a crosswalk. You review it and change anything you disagree with.
1 · Copy this prompt into Claude. The lab built it from the runs you recorded.
2 · Paste Claude's entire reply here
3 · Review the crosswalk
Every theme from every run, and the common theme Claude put it in. Change any assignment you disagree with. Your changes are highlighted and used in the comparison.
What did each analysis see?
Built from your LAB DATA blocks and your reviewed crosswalk. Each column is a run; each row is a common theme.
Share of comments by common theme (%)
Where each watch-list comment landed
Top three reported by each run
Run 0: Claude's judgement of what leadership most needs to hear. Other runs: the largest themes.
Paste this into Claude, then decide whether its comparison is right:
Reflect
What we planted
Open only after you have recorded your runs.
- A serious privacy risk: 8 of 400 comments (2%) describe staff pasting client data, including social insurance numbers and health information, into unapproved AI tools (C014 and C356 on the watch list). Topic models scatter these comments across unrelated topics at every setting. HR's categories file them under Other. A careful reader, human or Claude, catches them, and Claude's default reading (Run 0) often ranks them among the top issues by seriousness even though they are only 2% of comments.
- A coordinated campaign: 39 near-identical comments demand consultation (C163 is one). They give consultation a large theme in every topic-model run. Genuine, independent consultation concerns are only about 2.5% of comments. Remove duplicates (Run 5) and the theme shrinks, and a genuine union concern such as C270 can end up anywhere.
- French comments (12%): With English-only stop words, the 12-topic models build a topic out of words like les, pour, des, est. It gets named Unclear or French comments: the language became a theme. C121 and C356 raise real issues in French.
- The random seed: Same data, same algorithm, same 12 topics. Changing only the random starting point reshuffles topics and moved the largest theme from about a third of comments to about a fifth in our tests. One run is one draw.
- Naming is interpretation: A topic model outputs word lists, not themes. Count the Mixed and Unclear topics in each run: that is how muddy the model was. Every theme name you compared was Claude's reading of a word list, and a tidy name can hide a muddy topic.
- The crosswalk is a model too: Deciding that two differently named themes are the same issue is a judgement. Look at how many common themes Claude created, what it merged, and what it kept apart. Your edits to it are part of the analysis.
- HR's categories are a model too: Expect Claude's reading to agree with HR well short of 100% on the 40 test comments. Disagreements cluster on accuracy, trust, monitoring, and consultation, which HR's six categories have no good place for. A 100% score almost always means the answers leaked into the chat.
- Training data: Always guessing the most common category scores about 33%. The supervised classifier scores about 48% with 50 labelled examples and about 55% with 250. On short, varied text, the capability sits mostly in the pre-trained model, which every competitor can rent too.
- Sarcasm and accessibility: C088 is sarcasm; word counts read it as praise for time savings. The 5 accessibility comments (1%, C057 among them) almost always disappear.
- The runs are not clean experiments: Several runs changed more than one thing, Run 5 changes the denominator, and HR's six categories cannot express many of the emergent themes. That is normal in practice. The question is whether anyone writes the choices down.
Next time AI arrives pre-assembled: pull the levers
A chat window, a copilot, a vendor dashboard: each has an assemblage behind it. Run 0 showed you the defaults. These are the levers you can pull yourself, each one a run you did.
The CHRO's question, asked by someone who knows the levers. Copy it and adapt it to your next analysis.
The foundation model was the same in every run. The findings were not. Pre-assembled or built by you, it is the assemblage that produces the answer: the data, the algorithm, the parameters, the naming, the crosswalk, the categories, and the people who check them.