Skip to content

Bundle Tools — MCP Tool Reference

A Bundle binds one Project, one Questionnaire, and a non-empty selected subset of that Questionnaire's Campaigns to exactly one linear Bronze → Silver → Gold dataset chain. These tools manage the Bundle itself and its quality story; see Dataset Tools for the Bundle-scoped stage operations (extract, code + weight, refine, export).

create_bundle

Create a pipeline Bundle.

Parameter Type Required Default Description
name string Yes — Bundle name
project_id string Yes — Owning project (write access required)
questionnaire_id string Yes — Questionnaire whose completed surveys feed Bronze
campaign_ids string[] Yes — Campaign subset to measure. Must name at least one: a Bundle's measurement base is the surveys of the campaigns it selects

Returns: { bundle_id, name }


list_bundles

List pipeline Bundles, optionally filtered to one project.

Parameter Type Required Default Description
project_id string No — Restrict to Bundles owned by this project

Returns:

{
  "items": [
    {
      "id": "...",
      "name": "Wave 1",
      "project_id": "...",
      "questionnaire_id": "...",
      "campaign_ids": ["..."],
      "chain": {
        "bronze": {"dataset_id": "...", "status": "ready"},
        "silver": {"dataset_id": "...", "status": "processing"}
      }
    }
  ],
  "count": 1
}

chain keys are the stages that have a dataset (bronze/silver/gold); a stage with no dataset yet is simply absent — treat it as missing.


clone_bundle

Deep-copy a Bundle's ready datasets and recipes into an independent variant.

Parameter Type Required Default Description
bundle_id string Yes — Source Bundle
name string No "{source} (copy)" Name for the clone

Returns: { bundle_id, cloned_stages } — cloned_stages lists which stages (bronze/silver/gold) had a ready dataset to copy.

The clone starts from a byte-identical copy of each ready stage and carries the same stage recipes (coding selection and its dimension codebook, weighting targets, refine field specs), so it can be re-derived independently — e.g. with different weighting targets — without touching the original Bundle. This is how you compare variants (different weighting, different refinement) of the same raw extraction.


delete_bundle

Delete a Bundle and its dataset chain.

Parameter Type Required Default Description
bundle_id string Yes — Bundle to delete

Returns: { success: true }

Deletion is soft — the Bundle and its chain datasets are marked deleted immediately; their underlying data files are cleaned up by a background maintenance process shortly after.


Quality Tools

Quality is measured at the Bundle, not on individual datasets. Three of the four representativeness measures grade against the Bundle's re-assignable current Strategy (current_strategy_id) — assign one before expecting them to be measured. The Calibration Targets are the weighting instruction, not a benchmark: grading the weighted result against the spec the weighting aimed at would make a successful raking run score well by construction.

get_bundle_quality

The Bundle's Representativeness story in one call — one sample traced from design intent to deliverable:

  1. Strategy → Pool (selection) — did sampling achieve the design? Per Pool the Bundle's Campaigns drew on, aggregated only when they agree on one Strategy. Graded from each member's values recorded when the Pool was drawn; a Pool without a complete record reads not_measurable (pool_has_no_complete_snapshot)
  2. Strategy → Actual (fielded) — how far off was the realized base before weighting? With a per-campaign breakdown; no_strategy until a Strategy is assigned
  3. Strategy → Weighted (weighted) — did weighting recover the design? Graded against the same Strategy as (2), so the two differ by exactly what weighting recovered
  4. Pool → Actual (fielding_shift) — what did fielding itself contribute? Pool-referenced, over pool-borne Surveys only, so it answers a different question from (1)–(3) and is not a fourth point on their scale
  5. Response quality (response) — entropy, straightlining, Cronbach's alpha, acquiescence, non-response (see the Quality Metrics Reference)

Alongside them, excluded_profile describes what the per-Bundle completeness threshold removed, against what it kept — compared on the Strategy's own factors, because two row counts cannot say whether a cut introduced bias. An empty excluded set reports empty rather than a divergence of zero.

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID

Returns: { quality: {...} } — the five measures plus excluded_profile and chain/strategy context. Each carries a state (measured, not_measurable, no_strategy, no_weighted_factors, empty, not_derived, processing, invalidated, error) with a reason, so one unmeasurable or failing measure never hides the others and a zero is never mistaken for a grade. Measured payloads include overall score, composite error, and per-factor deviation analysis; the weighted payload also carries weighting_diagnostics (design effect, effective sample size) — the variance cost of the recovery.

A Bronze extracted under an older extraction schema version is not graded: fielded, fielding_shift, response, excluded_profile and weighted each read not_measurable with reason bronze_schema_stale until the Bronze is re-extracted (create_bronze_dataset) and the Silver derived again. selection reads no dataset and is still graded.


compare_bundle_quality

Bronze-vs-Silver representativeness comparison for the Bundle's chain, graded against the Bundle's current Strategy. Its recovery field is the figure to read: a plain difference of weighted-mean error over the Strategy factors measured on both sides, so equal error reductions at different absolute levels give equal figures and two Bundles are comparable. When the two measures share no measured factor it is null with a recovery_reason rather than 0.0 — nothing was comparable, which is not the same as weighting having achieved nothing.

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID

Returns: { comparison: {...} } carrying recovery (the overall figure described above) and recovery_per_factor — the same subtraction per factor, in error units. Both are error differences, not percentages: a positive figure means the weighting reduced error by that much. A per-factor percentage cannot be stated meaningfully here, because its denominator would be the Bronze error, which is the quantity weighting exists to drive toward zero. Both the Bronze and Silver must be ready first; a chain that isn't ready yet fails with a conflict error (409 with {"error": "conflict"} on REST) rather than a validation error — finish or re-run the derive, then compare.

The comparison's own state is measured when both sides were graded. It is no_strategy when the Bundle has no current Strategy, and not_measurable with reason bronze_schema_stale when the Bronze was extracted under an older extraction schema version; in both, bronze, silver and recovery are null and the call succeeds. Clear the stale state by re-extracting (create_bronze_dataset) and deriving Silver again.


get_bundle_coding

The Bundle's open-end coding state: what may be coded, which questions the analyst chose, and — per chosen question — the dimensions proposed for it, their categories, which are selected, and whether the degraded no-credential path produced them. Read it before selecting or re-deriving.

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID

Returns: { candidates, kinds, selected, discovered, units, awaiting_review, coding_outcome, bronze_status, silver_status, pending_review }.

  • candidates — every open-text unit this Bundle's Bronze offers, each {unit_key, title, control_type, roster_block_id, columns}. Only Textarea qualifies, and appearing here implies nothing about coding: what the column holds is the analyst's classification (kinds), and the unit must be classified before it can be selected. Editbox is a numeric input in QML (it requires a min and a max), so its answers are counts and amounts rather than text a codebook can describe. A Roster is one candidate whose columns name the Bronze columns it pools; a QuestionGroup sub-question and a Matrix cell each stay their own candidate.
  • kinds — the analyst's answer kind per candidate, passed through undefaulted: descriptive (prose, coded by meaning), nominal (a fixed vocabulary, coded by exact value) or identifier (names a person or record, never coded). A candidate with no entry is unclassified and cannot be selected; null means nothing has been classified. Read-only here: the analyst classifies in Balansor and no tool writes a kind.
  • selected — the persisted unit choice, passed through undefaulted: null means "not yet chosen", [] means "chosen: nothing". Only the second is a decision someone made.
  • discovered — maps each coded unit to {column: dimension label} — the record of which columns the last derive actually emitted.
  • units — the per-unit dimension state, keyed by unit key. Each carries reviewed (has anybody decided which dimensions become columns?), selected (the chosen dimension ids, null while unreviewed and [] for "reviewed, chose nothing"), degraded (the round came from the no-credential path — it is what tells a genuine one-dimension result apart from that path's single dimension), answered / missing, and dimensions. Each dimension carries its slot id, label, the frozen column name, its origin (proposed / authored / degraded, or verbatim for a nominal unit's one exact-value dimension), whether it is selected, its categories as {id, label} ({id} only for a verbatim dimension, whose category -1 is the reserved other), the per-column assignment counts, and overlap — how many of that dimension's own categories an average answer also cleared, which reads on the codebook rather than on the respondents.
  • awaiting_review — the selected units whose dimensions nobody has reviewed. Those units produce no coded column while the derive still succeeds, which is why this is named rather than inferred.
  • coding_outcome — what the last derive attempted and got no column from: {attempted, coded, lost}, each lost entry {unit_key, kind} with kind one of not_codable, unclassified or no_column. A unit failing does not fail the derive — the other units' columns are real work — so without this the loss has no signal at all: the dataset reads ready, and discovered above still names the unit from an earlier derive, because that map is deliberately kept across runs. Check it before reporting a Silver as complete. null means the derive recorded nothing, which is not the same as losing nothing: every Silver derived before this field existed is in that state, and an empty lost list is the value that says a run lost nothing. The reason a loss happened is deliberately reduced to kind — the underlying message can be arbitrary text, and this door never emits anything that might carry a respondent's answer.
  • bronze_status sits beside silver_status because an empty candidates means one of two different things — no Bronze has been extracted yet, or this questionnaire asks no open-text question at all — and the list alone cannot tell them apart.
  • pending_review — true while either decision is outstanding: a ready Bronze with candidates and no unit selection, or a selected unit awaiting dimension review.

No verbatim respondent text ever comes out of this door. A category is projected as its id and label only, and the per-dimension sample answers are not projected at all — both exist for the researcher's own review surface in Balansor. A nominal unit's category labels are respondent answers, so its categories are projected as ids with counts and no label.

Unit selection is writable over MCP/REST via set_coding_selection below, because it is a judgement about the question, made before any model has run. Classifying a column, selecting dimensions, renaming them and editing their categories are not, and will not be — those are judgements about the researcher's own data and about generated content, and there is deliberately no tool for them on any transport. An agent driving the pipeline should surface pending_review: true and the awaiting_review list to its user rather than re-deriving past them. See Open-End Coding for the full workflow.


set_coding_selection

Set which open-text units this Bundle codes into categories. The choice is a per-Bundle analysis decision on the Silver recipe, not a property of the questionnaire, and it can only name columns the analyst has already classified as descriptive or nominal (see kinds on get_bundle_coding).

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID
selected string[] Yes — Full replacement set of coding unit keys, in the vocabulary get_bundle_coding's candidates offers

Returns: { bundle_id, selected }.

Requires manager permissions and an identified user — service-token-only callers are rejected, so an unattended run cannot select on nobody's behalf. Present the candidates and their kinds to the researcher and write the answer they give you rather than choosing for them. A unit that is unclassified or classified identifier is refused with the reason and the tool writes nothing; an agent cannot classify, so ask the researcher to do it in the Silver page's coding picker.

The list is a full replacement set, not a delta, and an empty list is a decision ("code nothing"), not an omission. A key that is not a candidate on this Bundle's Bronze is refused — including a single Roster iteration's column name, which looks like a real column but would pool nothing; re-read the candidate list and select the pooled unit key instead. A selection change is a Silver input: it invalidates a ready Silver and everything downstream, and applies on the next derive (a derive already in flight keeps the selection it started with).

Deselecting a unit that backs a Calibration Target is not refused here — that conflict surfaces at the derive, which refuses with weighting_factor_deselected so the recipe can be fixed in either order.


assign_bundle_strategy

Assign (or clear) the Bundle's current Sampling Strategy — the benchmark selection, fielded and weighted all grade against. Requires manager permissions and an identified user (service-token-only callers are rejected); assigning a strategy also requires access to the strategy's own project. Reassigning it recomputes all three of those measures; it does not touch the Calibration Targets, which are the weighting instruction.

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID
strategy_id string No null Strategy to assign; null clears the assignment

Returns: { bundle_id, strategy_id, strategy_name }

To advance a strategy with reality-grounded outcome factors (e.g. political preference from the survey responses plus an external benchmark), use advance_sampling_strategy — it clones the strategy (the original is never edited) and the clone can then be assigned here. See the Sampling Tools Reference.


set_calibration_targets

Edit the Bundle's Calibration Targets — the editable spec derive_silver/code_open_ends actually weight to. It is the weighting instruction, not the benchmark: the weighted measure is graded against the current Strategy, so a Calibration Targets set that misses the design shows up as distance from that design rather than as a perfect score against its own aim. Seeded once from the current Strategy's factors the first time a researcher opens the Silver stage in Balansor; from then on it is independent of the current Strategy. Requires manager permissions and an identified user.

Parameter Type Required Default Description
bundle_id string Yes — Bundle UUID
weighting object No — Full replacement block {use_strategy_factors, factors}
factors array No — Shorthand for weighting={factors, use_strategy_factors: false} — supplying a bare factor list is itself an opt-out of live-Strategy tracking

Each factor is {column, target_distribution, name?, factor_type?, bucket_ranges?} — column names a bare column already present in the coded frame (a Bronze demographic/outcome column, a coded open-end dimension, or a multi-select option indicator), and target_distribution maps each value to its target proportion (summing to 1.0).

A target key is the column's stored value, not its caption

For a coded answer the stored value is the option's numeric code, so a five-point scale is targeted as {"1": 0.1, "2": 0.2, ...} — not {"Strongly agree": 0.1, ...}. If none of a factor's declared categories occurs in its column, the derive fails with an error naming the column, how many distinct values it holds, and the categories you declared. It does not quietly return weights of 1.0. A multi-select option indicator is 0/1. A coded open-end dimension is one factor across its whole category set — one target per category, 0 (unclassified) included, since its categories partition the respondents; a distribution that omits one is refused rather than silently normalized.

A target may also name a respondent custom attribute. Its column is _custom_<key>, where <key> is the attribute's key on the respondent record (_custom_region for custom_attributes.region). A custom attribute is extracted into a dataset only when a Sampling Strategy factor or a Calibration Target of the Bundle names it, so a target is how a new attribute enters the dataset: it does not have to be in the Bronze yet. Read which attributes the Bronze holds from custom_attribute_keys on get_dataset_schema.

Returns: { bundle_id, weighting, warnings }. warnings is a list of notes on a save that succeeded, empty when there is nothing to say. It carries one entry when the saved targets weight on a custom attribute the Bundle's Bronze was not extracted with: re-extract Bronze (create_bronze_dataset) before deriving Silver, or the derive is refused with weighting_attributes_not_extracted. If no respondent of the Bundle carries an attribute by that name, the name is wrong; correct it instead. With no Bronze yet there is no warning, because the first extraction writes every named attribute. The text is the same sentence a researcher sees when saving the targets in Balansor.

Two conditions are refused with a conflict error before anything is written — the same preflight the Silver derive itself runs:

  • a target naming a coding category that was rejected or has since vanished from the latest coding run;
  • a target naming a raw multi-select answer column. That column stores a combination of the selected options, so every observed combination would become its own weighting cell. Target the per-option indicator columns instead — one target per option.

Stage Operations

Extraction, coding + weighting, refinement, and export are all invoked with a bundle_id on the Dataset Tools. A typical flow:

  1. create_bundle → get bundle_id
  2. create_bronze_dataset(bundle_id) → extract Bronze
  3. code_open_ends(bundle_id) (or its alias derive_silver) → codes open-ends, then weights against the Calibration Targets, producing Silver
  4. create_gold_dataset(bundle_id, operations=[...]) → refine Silver into Gold using the operation catalog
  5. export_dataset(dataset_id, format) → export any ready stage
  6. get_bundle_quality(bundle_id) → the Representativeness story at any point

run_bundle_pipeline

Steps 2 through 4 as a single call. Each step is the same tool listed above, so the same datasets are produced and the same audit events land, and any refusal a raw stage would give you is given here unchanged — plus stopped_at_step, naming which stage a stop came from.

Parameter Type Required Default Description
bundle_id string Yes — The Bundle whose chain to run
include_demographics boolean No true Enrich the extracted Bronze rows with respondent demographics. false is refused with named_attributes_require_enrichment when the Bundle names a custom attribute (see create_bronze_dataset)
gold_name string No — Name for the Gold dataset. Defaults to "{bundle_name} — Gold"; the Gold recipe itself still comes from the Bundle's own stored recipe

Returns: { task_id, status, step, bundle_id, bronze_dataset_id, silver_dataset_id }. Poll task_id with get_task_status until the run reaches completed, partial, or failed; the poll reports the stage the run is on and, once finished, each stage's dataset id (bronze_dataset_id, silver_dataset_id, gold_dataset_id) and gold_skipped: the entries of the stored Gold recipe the build set aside because the Silver no longer has their column, in the shape create_gold_dataset returns as skipped. It is empty when nothing was set aside.

Only the Bronze extraction and the Silver derive's enqueue happen while you wait. Gold cannot start until the Silver is ready, which is minutes of coding and weighting away, so the Gold step runs on its own after the derive finishes and the run stays pollable throughout.

Two things to know:

  • A stop leaves what was built. A Bronze is the Bundle's Bronze whether or not a Gold followed it, so nothing is undone. Re-running resumes over the same rows rather than duplicating a stage.
  • The whole chain's entitlements are checked before anything is built. The chain crosses both the gateway tier and the AI-assistant tier (the coding derive needs the latter), so a subscription covering only the first is refused up front rather than handed a Bronze it cannot derive from.