ADX 5103 · AI for Executives · Class 4
TELFER Executive MBA — The Assemblage Lab
Same Claude · different assemblages
The Assemblage Lab · before you start

Same 400 comments. How many answers?

Rideau Valley Services (fictional) rolled out an AI assistant called Compass six months ago. The Chief Human Resources Officer asked every employee one open question: “What is your experience with Compass and AI at RVS so far, and what should leadership know?” 400 replied. She wants the top issues before next week's leadership retreat.

You will produce her answer several ways with the same Claude. In Run 0 you ask a plain question and Claude's defaults decide everything: the assemblage comes pre-assembled. From Run 1 on, you build the assemblage yourself, one lever at a time: the algorithm, its settings, the data preparation, the categories, the training data. Then you line the answers up and compare what each assemblage saw. Go as far as time allows; you can align and compare after any run.

1 · Get the data

400 comments with department and tenure. All synthetic.

2 · Set up Claude (30 seconds)

  • In Claude, open Settings → Capabilities and make sure code execution and file creation is on.
  • Open Settings → Memory. Turn off Search and reference chats and pause memory, so earlier runs cannot leak into later ones. Turn them back on when you finish the lab.
  • Use ordinary chats, not a Claude Project, and the same Claude model throughout.

3 · Four chats, in this order

Chat ARun 0new chat · attach the survey
Chat BRuns 1–5new chat · attach the survey, then stay in it
Chat CRuns 6–7new chat · attach the survey; HR's labels come later
Chat DAlign, Comparenew chat · no file

Each run page says which chat to use. After each run, paste Claude's reply into the box. The lab reads the LAB DATA block at the end and builds your comparison.

4 · One gut check


Run 0

Ask the way a busy executive would. The assemblage comes pre-assembled.

Pre-assembled
Chat A · Start a new chat and attach rvs_ai_survey.csv.

1 · Copy this prompt into Claude.

I have attached rvs_ai_survey.csv: 400 open-ended comments from an employee survey at Rideau Valley Services, a fictional public-service organization, about its AI assistant, Compass. The question was: "What is your experience with Compass and AI at RVS so far, and what should leadership know?" Use only the files I share in this chat; do not draw on my other chats or on memory. What are employees telling us, and what should leadership know?

2 · When Claude has answered, send this in the same chat so it records what it did:

Thanks. Please add an appendix to that analysis, as an analyst would for a leadership report: 1. METHOD NOTE: a numbered list documenting the analysis: which comments were included or set aside, how repeated comments and French comments were treated, whether theme sizes were counted, estimated, or computed with code, how themes were defined and named, and the criterion used to rank what leadership should hear first. 2. The share of comments (%) for each theme in the analysis, marked as counted or estimated. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 0 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> (the three leadership most needs to hear about, in your order) WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

3 · Paste Claude's entire second reply here

What to look for

Read the METHOD NOTE. Every item is a choice about data, method, or model that nobody asked you to approve. Did Claude notice the French comments? The 39 copies of the same campaign text? The handful of comments about client data leaving the organization? Did it rank by how often something came up, or by how serious it is? Did it count, or estimate? Each of these is a lever that was pulled for you. Next time you ask a question like this, you can pull them yourself: the Compare step ends with the list.
Run 1

Start building the assemblage yourself: a named algorithm with explicit settings.

Assembled by you
Chat B · Start a new chat and attach rvs_ai_survey.csv.

1 · Copy this prompt into Claude.

I have attached rvs_ai_survey.csv: 400 open-ended comments from an employee survey at Rideau Valley Services, a fictional public-service organization, about its AI assistant, Compass. The question was: "What is your experience with Compass and AI at RVS so far, and what should leadership know?" The comments are in the column "comment". Use only the files I share in this chat; do not draw on my other chats or on memory. Use Python. Vectorize the "comment" column with CountVectorizer (stop words: scikit-learn's built-in English stop-word list plus these extra stop words: compass ai assistant tool tools use using used just like don ve doesn didn people really make makes thing things want know got get; min_df=2). Fit LatentDirichletAllocation with n_components=5 and random_state=1. Assign each comment to its highest-probability topic. Put comments with no words left after stop-word removal into a theme called "No words left" and tell me how many there were. Name every topic yourself with a short label of 2 to 5 words. Start from its 10 highest-weight words. If the words are ambiguous, also read the 3 comments with the highest weight for that topic, and only those. Then: if the topic is about one clear issue, give it a plain name; if it mixes two or more issues, name it "Mixed: <main issue>"; only if the words and those 3 comments share no issue at all, name it "Unclear: <its top three words>". Give each topic a different name, and in the TOPICS table add a column saying whether you named it from the words alone or also read the 3 comments. Report: 1. TOPICS: a table with #, your name for the topic, share of comments (%), its top 8 words, and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was prepared, the method and parameters, and how topics were named. 3. In the LAB DATA block, TOP3 = the three largest themes, not counting "Unclear" themes or "No words left" ("Mixed" themes count); WATCH = the theme each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 1 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

2 · Paste Claude's entire reply here

What to look for

Compare with Run 0. Which themes survived? Which topics are really about the words people happened to use rather than an issue? Is one topic mostly a single repeated message?
Run 2

Change one parameter: how many topics the model may find.

Assembled by youLever: Model parameter: number of topics
Chat B · Continue in the same chat as Run 1. No new upload.

1 · Copy this prompt into Claude.

Keep everything exactly the same as the previous run (same data, stop words, and naming rule), but change n_components to 12. Name the topics afresh from their words; reuse an earlier name only if the words justify it. Report: 1. TOPICS: a table with #, your name for the topic, share of comments (%), its top 8 words, and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was prepared, the method and parameters, and how topics were named. 3. In the LAB DATA block, TOP3 = the three largest themes, not counting "Unclear" themes or "No words left" ("Mixed" themes count); WATCH = the theme each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 2 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

2 · Paste Claude's entire reply here

What to look for

Which themes appeared, split, or vanished? How many topics are "Mixed" or "Unclear"? Open their words. Is one of them simply the French comments?
Run 3

Change nothing that should matter: only the random starting point.

Assembled by youLever: Model parameter: random seed
Chat B · Continue in the same chat as Run 1. No new upload.

1 · Copy this prompt into Claude.

Keep everything exactly the same as the previous run (LDA, 12 topics, same naming rule), but change only random_state from 1 to 2. Name the topics afresh from their words; reuse an earlier name only if the words justify it. Also tell me how many of the 12 topics look essentially the same as in the previous run. Report: 1. TOPICS: a table with #, your name for the topic, share of comments (%), its top 8 words, and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was prepared, the method and parameters, and how topics were named. 3. In the LAB DATA block, TOP3 = the three largest themes, not counting "Unclear" themes or "No words left" ("Mixed" themes count); WATCH = the theme each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 3 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

2 · Paste Claude's entire reply here

What to look for

If the themes shifted a lot, a single run was one draw, not the answer. Which themes held across both seeds? Those deserve more trust.
Run 4

Same data, same number of topics, different algorithm.

Assembled by youLever: Algorithm
Chat B · Continue in the same chat as Run 1. No new upload.

1 · Copy this prompt into Claude.

Now switch the algorithm. Use TfidfVectorizer (stop words: scikit-learn's built-in English stop-word list plus these extra stop words: compass ai assistant tool tools use using used just like don ve doesn didn people really make makes thing things want know got get; min_df=2) and NMF with n_components=12, random_state=1, init="nndsvda", max_iter=500. Assign each comment to its highest-weight topic. Put comments with no words left after stop-word removal into a theme called "No words left" and tell me how many there were. Use the same naming rule. Name the topics afresh from their words; reuse an earlier name only if the words justify it. Report: 1. TOPICS: a table with #, your name for the topic, share of comments (%), its top 8 words, and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was prepared, the method and parameters, and how topics were named. 3. In the LAB DATA block, TOP3 = the three largest themes, not counting "Unclear" themes or "No words left" ("Mixed" themes count); WATCH = the theme each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 4 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

2 · Paste Claude's entire reply here

What to look for

Did the algorithm change what the organization would see, or mostly how cleanly it is grouped? Where did C014 and C356, the two comments about client data, land?
Run 5

Back to the Run 2 setup, but change how the data is prepared.

Assembled by youLever: Data
Chat B · Continue in the same chat as Run 1. No new upload.

1 · Copy this prompt into Claude.

Go back to the Run 2 setup (CountVectorizer, same stop words, min_df=2, LDA, n_components=12, random_state=1, same naming rule), but change the data preparation: (a) before fitting, remove duplicate and near-duplicate comments, ignoring case, punctuation, and hashtags, and tell me how many you removed (for any watch-list comment removed, use the theme name "Removed"); (b) add French stop words (for example: le, la, les, des, est, pour, en, pas, sur, du, de, et, un, une, que, qui, dans, nous, je, il, à, au, aux, ce, ces, se, sont, ne, plus, par, avec, on, été, être, aucune, tous, l, d, j, n, qu, c, s, y, a, ça). Report percentages of the comments that remain. Name the topics afresh from their words; reuse an earlier name only if the words justify it. Report: 1. TOPICS: a table with #, your name for the topic, share of comments (%), its top 8 words, and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was prepared, the method and parameters, and how topics were named. 3. In the LAB DATA block, TOP3 = the three largest themes, not counting "Unclear" themes or "No words left" ("Mixed" themes count); WATCH = the theme each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 5 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> === END ===

2 · Paste Claude's entire reply here

What to look for

What happened to the consultation theme once the campaign copies were gone? Did the genuine union concern (C270) survive? Did the French comments join the issue topics instead of forming their own?
Run 6

Instead of letting themes emerge, use the categories the organization already has.

Assembled by youLever: Model: the organization's categories
Chat C · Start a new chat and attach rvs_ai_survey.csv.

1 · Copy this prompt into Claude.

I have attached rvs_ai_survey.csv: 400 open-ended comments from an employee survey at Rideau Valley Services, a fictional public-service organization, about its AI assistant, Compass. The question was: "What is your experience with Compass and AI at RVS so far, and what should leadership know?" The comments are in the column "comment". Use only the files I share in this chat; do not draw on my other chats or on memory. Our HR analysts code survey comments into six categories: Training & Support; Tools & Access; Workload & Productivity; Jobs & Careers; Leadership & Communication; Other. Read the comments yourself and assign each to exactly one category. Decide every label by reading, not with a keyword rule or a model; you may use code only to write down the labels you decided and to count them. Save your labels to a file. Do not list all 400 labels. For now, reply only with the number of comments in each category and the word READY. I will then send HR's own labels.

2 · When Claude says READY, download rvs_hr_labels.csv , attach it in the same chat, and send this:

Here is rvs_hr_labels.csv: HR's category for 290 comments (columns id, hr_category, split). Do not change any of your labels. Compare your labels with hr_category for the 40 rows where split = "test": report your accuracy, the accuracy of always guessing the most common test category, and list the comments where you and HR disagree. Report: 1. CATEGORIES: a table with each category, its share of comments (%), and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was handled, and how categories were assigned. 3. In the LAB DATA block, THEMES = the six categories with their shares; TOP3 = the three largest categories, not counting Other; WATCH = the category each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 6 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> ACCURACY: <your accuracy on the 40 test rows>; baseline <most-common-category accuracy> === END ===

3 · Paste Claude's entire second reply here

What to look for

What went into Other? Where did the comments about client data and about consultation end up? Read the disagreement list: is HR wrong, is Claude wrong, or do the six categories simply have no good place for comments about accuracy, trust, or monitoring? HR's ‘correct’ labels are themselves a model. And if the accuracy came out at 100%, ask whether Claude could see the answers before it labelled: that is the most common evaluation mistake in practice.
Run 7

Train a model on HR's labelled examples and compare it with Claude's reading.

Assembled by youLever: Data: training examples
Chat C · Continue in the same chat as Run 6. No new upload.

1 · Copy this prompt into Claude.

Now build a supervised model instead. In Python: train a classifier on the first 50 rows of rvs_hr_labels.csv where split = "train" (TfidfVectorizer with English stop words, then LogisticRegression with max_iter=2000) and report accuracy on the 40 rows where split = "test". Then retrain on all 250 training rows and report accuracy again. Also report the most-common-category baseline on the test rows. Finally, use the 250-row model to predict all 400 comments and use those predictions for the report. Note that the 250 training comments are predicted by a model that has already seen their labels. Report: 1. CATEGORIES: a table with each category, its share of comments (%), and two example comment IDs. 2. METHOD NOTE: a numbered list of the analysis settings, as an analyst would document them in a report appendix: which comments were used, how the text was handled, and how categories were assigned. 3. In the LAB DATA block, THEMES = the six categories with their shares; TOP3 = the three largest categories, not counting Other; WATCH = the category each of those comments was assigned to. End your reply with this block, exactly in this format. Use exactly the same theme names in THEMES, TOP3 and WATCH. === LAB DATA === RUN: 7 THEMES: <theme>=<% of comments>; <theme>=<% of comments>; ... TOP3: <theme> | <theme> | <theme> WATCH: C014=<theme>; C057=<theme>; C088=<theme>; C121=<theme>; C163=<theme>; C199=<theme>; C232=<theme>; C270=<theme>; C311=<theme>; C356=<theme> ACCURACY: reading <Run 6 %>; baseline <%>; 50 examples <%>; 250 examples <%> === END ===

2 · Paste Claude's entire reply here

What to look for

Start with the baseline: always guessing the most common category already scores about a third. How far above it did 50 and 250 examples get? How does the classifier compare with Claude simply reading? The classifier learns only from your 250 labels; Claude brings a model pre-trained on vastly more text, which every competitor can rent too.
Align · the human element

Which themes are the same issue?

Each run named its own themes. Before you can compare runs, someone has to decide that “verification burden” in one run and “trust in outputs” in another are, or are not, the same issue. Claude will propose a crosswalk. You review it and change anything you disagree with.

Chat D · Start a new chat. No file needed.

1 · Copy this prompt into Claude. The lab built it from the runs you recorded.

2 · Paste Claude's entire reply here

3 · Review the crosswalk

Every theme from every run, and the common theme Claude put it in. Change any assignment you disagree with. Your changes are highlighted and used in the comparison.

Compare

What did each analysis see?

Built from your LAB DATA blocks and your reviewed crosswalk. Each column is a run; each row is a common theme.

Share of comments by common theme (%)

Where each watch-list comment landed

Top three reported by each run

Run 0: Claude's judgement of what leadership most needs to hear. Other runs: the largest themes.

Chat D · Continue in the same chat as Align.

Paste this into Claude, then decide whether its comparison is right:

Reflect

What we planted

Open only after you have recorded your runs.

  • A serious privacy risk: 8 of 400 comments (2%) describe staff pasting client data, including social insurance numbers and health information, into unapproved AI tools (C014 and C356 on the watch list). Topic models scatter these comments across unrelated topics at every setting. HR's categories file them under Other. A careful reader, human or Claude, catches them, and Claude's default reading (Run 0) often ranks them among the top issues by seriousness even though they are only 2% of comments.
  • A coordinated campaign: 39 near-identical comments demand consultation (C163 is one). They give consultation a large theme in every topic-model run. Genuine, independent consultation concerns are only about 2.5% of comments. Remove duplicates (Run 5) and the theme shrinks, and a genuine union concern such as C270 can end up anywhere.
  • French comments (12%): With English-only stop words, the 12-topic models build a topic out of words like les, pour, des, est. It gets named Unclear or French comments: the language became a theme. C121 and C356 raise real issues in French.
  • The random seed: Same data, same algorithm, same 12 topics. Changing only the random starting point reshuffles topics and moved the largest theme from about a third of comments to about a fifth in our tests. One run is one draw.
  • Naming is interpretation: A topic model outputs word lists, not themes. Count the Mixed and Unclear topics in each run: that is how muddy the model was. Every theme name you compared was Claude's reading of a word list, and a tidy name can hide a muddy topic.
  • The crosswalk is a model too: Deciding that two differently named themes are the same issue is a judgement. Look at how many common themes Claude created, what it merged, and what it kept apart. Your edits to it are part of the analysis.
  • HR's categories are a model too: Expect Claude's reading to agree with HR well short of 100% on the 40 test comments. Disagreements cluster on accuracy, trust, monitoring, and consultation, which HR's six categories have no good place for. A 100% score almost always means the answers leaked into the chat.
  • Training data: Always guessing the most common category scores about 33%. The supervised classifier scores about 48% with 50 labelled examples and about 55% with 250. On short, varied text, the capability sits mostly in the pre-trained model, which every competitor can rent too.
  • Sarcasm and accessibility: C088 is sarcasm; word counts read it as praise for time savings. The 5 accessibility comments (1%, C057 among them) almost always disappear.
  • The runs are not clean experiments: Several runs changed more than one thing, Run 5 changes the denominator, and HR's six categories cannot express many of the emergent themes. That is normal in practice. The question is whether anyone writes the choices down.

Next time AI arrives pre-assembled: pull the levers

A chat window, a copilot, a vendor dashboard: each has an assemblage behind it. Run 0 showed you the defaults. These are the levers you can pull yourself, each one a run you did.

The CHRO's question, asked by someone who knows the levers. Copy it and adapt it to your next analysis.

The foundation model was the same in every run. The findings were not. Pre-assembled or built by you, it is the assemblage that produces the answer: the data, the algorithm, the parameters, the naming, the crosswalk, the categories, and the people who check them.