AI Functions & Model Options¶
"AI" in Askalot is not one model doing everything. It is a set of distinct functions, each chosen for a specific job, each with its own requirement and its own degree of choice. Some are platform-fixed because a per-user choice would break correctness. Others are yours to select, because they spend your budget and touch your data.
This page is the complete map: what each function does, what it needs to work, what you can change, and what happens when something is missing.
At a glance¶
| Function | What it does | Runs on | Your choice? |
|---|---|---|---|
| Embedding | Turns text into vectors for semantic search and clustering | Askalot's own inference server (BGE-M3) | No — platform-fixed |
| Layout detection | Finds headings, paragraphs, tables and reading order in a PDF | Askalot's own GPU (Docling) | No — infrastructure |
| OCR | Reads text out of scanned pages and images | Not enabled yet | Planned |
| Knowledge-graph extraction | Pulls entities and relations out of your documents | Claude (your Direct LLM slot) | Yes |
| Re-ranking | Re-scores retrieved passages for relevance | Askalot's own inference server (BGE reranker) | On/off per query |
| Query keyword extraction | Works out what to look for in the knowledge graph | Claude (your Direct LLM slot) | Yes |
| Open-end coding | Proposes the dimensions open-ended answers vary on, and names the fallback's groups | Claude (your Direct LLM slot), or keyword statistics when no slot is set | Yes — provider and model |
| Reasoning (the agents) | Designer, Manager, Analyst, Respondent | Claude (your Reasoning slot) | Yes |
| Document summaries | The index of what your uploaded documents cover | Claude (your Reasoning slot, cheap tier) | Follows Reasoning |
| Respondent simulation | Answers surveys as a persona (llm mode) |
Claude Haiku, via your Reasoning provider | Provider only |
Everything marked "Claude" runs through your own Anthropic API key, Claude subscription, or AWS Bedrock account — see Choosing your models.
Retrieval functions¶
These four turn an uploaded document into something an agent can search. They run in sequence: parse the page, split it into chunks, vectorize the chunks, build a knowledge graph over them.
Embedding — platform-fixed¶
What it does. Converts every document chunk and every search query into a 1024-dimensional vector, so "how do I calculate a design effect" can find a passage that never uses those words.
| Model | BGE-M3 (multilingual, 8K context) |
| Where it runs | Askalot's own GPU inference server — your text is not sent to a third party for this step |
| Requirement | The inference server must be reachable. It sleeps when idle and wakes on first use (a few extra seconds). |
| Options | None. The slot exists in Profile Settings but is fixed and read-only. |
Why you cannot change the embedding model
Search only works if your documents were indexed with the same model the query uses. Two different models produce vectors that are not comparable — switching would not degrade search, it would silently return nothing useful from everything indexed before the switch. So the choice is made once, at the platform level, and kept.
This also keeps your documents' vectors on Askalot's own hardware rather than on a third-party embedding API.
Layout detection — platform infrastructure¶
What it does. Before text can be chunked it has to be read correctly: headings identified, columns un-interleaved, tables recovered as tables rather than as scrambled runs of numbers.
| Model | Docling layout analysis + TableFormer (table structure recovery) |
| Where it runs | Askalot's own GPU |
| Requirement | Applies to PDFs. Plain text, Markdown and CSV skip this step entirely. |
| Options | None — it is infrastructure, not a model preference |
Why this is not a setting
CPU and GPU produce byte-identical output here — the GPU is roughly four times faster, and that is the only difference. Parsed documents are also cached and shared, so a document one person uploads is not re-parsed for the next person. A per-user preference would be quietly ignored whenever someone else's parse populated the cache first, which is worse than having no setting at all.
OCR — planned¶
What it does (when it lands). Reads text out of scanned pages, photographed documents, and images embedded in a PDF.
| Status | Not enabled. OCR is switched off in the current document pipeline. |
| What this means today | A PDF with a real text layer indexes normally. A scanned PDF — pages that are images of text — yields little or no text, so the document indexes as effectively empty rather than failing with an error. |
| How to work around it | Run the scan through OCR yourself before uploading, or upload a text-layer version. |
| Planned options | Whether OCR runs at all, and the language hint, are the two knobs under consideration. No dates yet. |
Check scanned uploads
If you upload a scan and the Designer cannot find anything in it, this is almost certainly why — not a retrieval failure. Open the document's indexing status: a scanned page set produces a very small chunk count.
Re-ranking — on/off per query¶
What it does. After the graph and vector search return candidate passages, a cross-encoder re-reads each candidate against your actual question and re-scores it. It is more accurate than vector similarity because it looks at the query and the passage together rather than comparing two independent summaries of them.
| Model | BGE-reranker-v2-m3 (cross-encoder) |
| Where it runs | Askalot's own GPU inference server |
| Default | Off |
| Options | On or off, per query |
Why it defaults to off
Quality is not free here. Scoring 40 full-length candidates costs about 8 seconds per query — real quality, real latency. For most retrieval the graph + vector ranking is already good enough, so the extra wait is not worth paying on every question. Turn it on when precision matters more than speed.
Generative functions¶
These call a Claude model. They spend your credits and send the relevant text to whichever provider you selected, so each one is tied to a slot you control.
Knowledge-graph extraction — your Direct LLM slot¶
What it does. Reads each chunk of your document and pulls out entities and the relations between them, building the graph that makes retrieval conceptual rather than merely keyword-based. This is the most token-hungry AI step in the platform — a long PDF is many chunks, and each one is a model call.
| Runs on | Claude, via your Direct LLM slot (Anthropic API or AWS Bedrock) |
| Requirement | A configured Direct LLM slot with a working credential |
| Options | Provider and model |
| If unconfigured | Indexing fails with a clear error telling you to finish setup |
Why extraction is billed to you
Extraction is the single largest AI cost in document indexing. Attributing it to the person who uploaded the document is what makes that cost visible and attributable, instead of pooling it invisibly into platform overhead. The same principle runs through every generative function on this page: the account you chose is the account that pays, and you can always see which one that is.
Query keyword extraction — your Direct LLM slot¶
What it does. When you ask a question, a fast, cheap model call works out which entities and concepts to look for in the knowledge graph before retrieval runs.
| Runs on | Claude, via your Direct LLM slot |
| Requirement | Same slot as extraction above |
| Options | Provider and model — a small, fast model is the right choice here |
| If unconfigured | Depends on the search: keyword lookup in your own documents keeps working (it matches on meaning alone and needs no model), while the broader knowledge-graph searches refuse with a clear error — those read your documents with a model before they can answer, so without one there is nothing to run |
Open-end coding — dimension proposal, and the fallback's labels¶
What it does. Open-ended answers are coded into labelled categorical variables (see Open-End Coding). The AI part is the proposal round: a model reads each question plus a sample of its answers and names the dimensions those answers vary on, with starter categories. Matching answers to those categories, deriving the columns, weighting and exporting all run on the platform's own embedding service and need no credential of yours.
| Runs on | Claude, via your Direct LLM slot |
| Requirement | A configured Direct LLM slot with a working credential |
| If unconfigured | Coding does not refuse. The round groups the answers themselves and proposes one dimension instead, labelled Degraded everywhere it appears so a genuine one-axis result is never mistaken for it |
| If the credential is rejected | The round fails and says so — it does not fall back to the degraded path, because a silent downgrade would be indistinguishable from a genuine result |
The degraded round names its own groups from keyword statistics — the top distinguishing words of each group, e.g. "delivery, late, refund". That runs on Askalot with no network call and no cost, which is the whole point: it is the path for a researcher who has no credential to spend. Those names are yours to change — the review screen lets you rename every dimension and every category before you keep any of them.
Why a rejected credential fails the round
The failure mode this prevents is the quiet one. If a refused credential dropped silently to the degraded path, you would get a single dimension flagged Degraded — which is exactly what a legitimately credential-less round produces — and no way to tell that a working credential would have read your question instead. An expired AWS session mid-run is exactly the case that would trigger it. So the round stops and tells you, and you re-run it after fixing the credential.
Reasoning — the four agents¶
What it does. Everything the Designer, Manager, Analyst and Respondent do: reading your research paper, planning a questionnaire, writing and validating QML, running a campaign, interpreting quality metrics.
| Runs on | Claude, via your Reasoning slot |
| Options | Anthropic API key, personal Claude subscription, or AWS Bedrock |
| Tiers | High / medium / low — each stage of work uses the tier that suits it, all on the provider you chose |
| If unconfigured | The agent stops and asks you to finish setup — it never silently switches provider |
Local models are deliberately excluded here: the agent framework requires Claude-compatible models. That is a capability constraint, not a judgement about local model quality.
Document summaries — follows Reasoning¶
What it does. Produces the short index of what each uploaded document covers, so the Designer knows what material exists and what to go looking for. One cheap, single-turn call per document plus one to combine them.
| Runs on | Your Reasoning slot, cheapest tier |
| Options | Follows your Reasoning choice — nothing separate to configure |
Respondent simulation — provider only¶
What it does. In llm distribution mode, answers each survey question as a
demographic persona, carrying the conversation history forward so a respondent who
said "unemployed" does not describe a workplace three questions later.
| Runs on | Claude Haiku, through your Reasoning slot's provider |
| Options | Provider only — the model is fixed to the fast, cheap tier because the job is high-volume and short-form |
| Alternative | The three rule-based modes (realistic, random, stratified) need no AI at all |
Choosing your models¶
Three settings, under Profile Settings → AI Models. They are independent — each covers a different way of invoking a model, and choosing one does not constrain the others.
| Slot | Covers | You can choose |
|---|---|---|
| Embedding | Semantic indexing and search | Fixed to the local model — shown but not editable |
| Direct LLM | Knowledge-graph extraction, query keyword extraction, open-end dimension proposal | Anthropic API, or AWS Bedrock |
| Reasoning | The four agents, document summaries, respondent simulation | Anthropic API, Claude subscription, or AWS Bedrock |
Credentials are entered separately under Profile Settings → AI Providers. An organization administrator can set organization-wide credentials; your personal credential takes precedence when both exist.
No shared key, no silent fallback
Askalot never runs your AI work on a platform-owned credential. A model slot resolves your personal credential first, then your organization's — and stops there. If neither is set, the feature tells you to configure it rather than quietly running on someone else's account.
The same rule covers expiry. If your AWS session has lapsed, you get a re-authenticate prompt, not a silent switch to a different provider or a degraded result that looks like a real one.
Where your data goes
Embedding, layout detection and re-ranking run entirely on Askalot's own hardware — that content is not sent to a model provider. The generative functions send the relevant text to the provider you selected: Anthropic directly, or AWS Bedrock inside your own AWS account if you need IAM, VPC or data-residency control.
When something is missing¶
Askalot fails loudly and specifically, by design. The table below is the map from symptom to cause.
| Symptom | Cause | Fix |
|---|---|---|
| Indexing fails with a credential error | Direct LLM slot has no working credential | Add a key under Profile Settings → AI Providers, or re-authenticate AWS |
| A knowledge-graph search refuses to run (plain keyword lookup still works) | Direct LLM slot has no configured provider or credential | Configure a provider and model under Profile Settings → AI Models → Direct LLM |
| Dimension proposal fails part-way with a credential error | Direct LLM credential rejected — an expired AWS session is the usual cause | Re-authenticate, or replace the key, then propose again. The derive that follows needs no credential of yours |
| Agent asks you to finish setup | Reasoning slot has no working credential | Choose a provider and add its credential |
| Indexing succeeds but the document seems empty | Scanned PDF with no text layer — OCR is not enabled yet | Upload a text-layer version, or OCR it first |
| First action after a quiet period is slow | Local inference models were asleep and are waking | Nothing — the action completes; the navigation bar's Local models indicator shows the state |
| Retrieval feels imprecise | Re-ranking is off by default | Turn it on for that query, accepting the added latency |
Related Documentation¶
- AI-Assisted Features Overview — what the four agents do
- Open-End Coding — how open-ended answers become labelled categorical variables
- Questionnaire Generation — document indexing in context
- Claude Integration — using Askalot from Claude