Configuration & providers
The Assistant is configured entirely through the root Assistant configuration section
(bound to AssistantOptions) plus two host calls on the backend and one module flag on the
frontend. This page covers the providers, the wiring, and the runtime behaviour you can tune.
For what the feature does, see the AI Assistant overview. For DB-enforced tenant isolation, see Tenant isolation & SQL hardening.
Models
Section titled “Models”A chat picks the model that answers it, from the list under Genie:Assistant:Providers:
"Genie": { "Assistant": { "Providers": [ { "Name": "DeepSeek R1", "Provider": "DeepSeek", "Endpoint": "https://api.deepseek.com/v1", "Model": "deepseek-reasoner", "MaxTokens": 10000, "IsDefault": true }, { "Name": "Local Llama", "Provider": "Ollama", "Endpoint": "http://localhost:11434", "Model": "codellama:7b" } ] }}An array, so the order is the order the model picker lists — and the index an out-of-band
API key targets. Name is free text: it is what the picker
shows and what a chat remembers, so it can read like a name rather than a model id. It must be unique;
IsDefault marks the one a new chat starts on; Enabled: false takes a model out of service without
deleting the entry that documents it.
Why per-chat rather than per-deployment: the useful choice differs by question. A reasoning model is worth its latency for a report someone is about to publish, and wasted on “how many suppliers do we have”.
Every model is reached through Microsoft.Extensions.AI’s IChatClient. The built-in kinds are
all served by the OpenAI client pointed at the kind’s endpoint — DeepSeek, Ollama and Gemini speak
OpenAI Chat Completions — and a host can bring any other provider as a ChatClient entry
(below). The kinds, and what goes with each:
Provider |
Default Endpoint (optional) |
Example Model |
API key |
|---|---|---|---|
| OpenAI | https://api.openai.com/v1 |
gpt-4o-mini, gpt-5-mini |
required |
| DeepSeek | https://api.deepseek.com/v1 |
deepseek-chat, deepseek-reasoner |
required |
| Gemini | https://generativelanguage.googleapis.com/v1beta/openai |
gemini-2.5-flash |
required |
| Ollama | http://localhost:11434 (served under /v1) |
qwen3, llama3.1:8b |
not required (local) |
| OpenAICompatible | none — required | whatever the server serves | optional |
| ChatClient | not used | optional (sent as ModelId) |
the host’s client handles it |
Endpoint is the API root. Leave it out and a hosted kind uses its own service. A trailing
/chat/completions is accepted and stripped, so entries written with the full completions URL keep
working; Ollama’s root gets /v1 added, and Gemini’s native root (…/v1beta) is pointed at its
OpenAI-compatible surface (…/v1beta/openai). Gemini is called through that surface on purpose: its
native API took the key in the query string, where proxies and HTTP logs record it — the compatible
surface takes it as a bearer header.
Every entry is validated at startup — a missing or duplicate Name, or a Provider that is not
one of the kinds above, fails there rather than on the first question someone asks.
Per-entry settings beyond those: MaxTokens, Temperature, TopP, TimeoutSeconds,
StreamIdleTimeoutSeconds, EnableThinking, ReasoningEffort, OrganizationId.
What is the same for every kind, because it is decided once in the provider layer rather than per provider:
- Retries. A
408/429/5xxor a transport failure is retried twice, with backoff, before the turn fails. A rejected request (400/401/404) is never retried — it is refused identically every time — and is reported with the status, the URL dialled and the server’s own explanation (with a key-length hint on a401, never the key). - Running out of tokens. An answer cut off at
MaxTokenstriggers the assistant’s compaction retry on every kind. (Before, only DeepSeek and OpenAI’s reasoning path noticed.) - Reasoning never becomes the reply (see the thinking notes below).
- A streamed answer that goes silent is abandoned after
StreamIdleTimeoutSeconds. - Keys stay out of errors and logs: any echo of the key in a provider’s error body is redacted before it reaches an exception, the turn trace or the log.
- OpenTelemetry: each model call is an
Activityon theGenie.Assistantsource (a no-op unless a listener is attached), recorded without message content — the turn trace remains the only place chat text is logged.
Running a local model (llama.cpp, Ollama, vLLM, LM Studio)
Section titled “Running a local model (llama.cpp, Ollama, vLLM, LM Studio)”Ollama has its own kind; any other server that exposes the OpenAI API (llama-server, vLLM,
LM Studio, a proxy) is an OpenAICompatible entry with its API root as Endpoint:
{ "Name": "Local Qwen", "Provider": "OpenAICompatible", "Endpoint": "http://my-host:18437/v1", "Model": "qwen2.5-coder-7b" }No ApiKey is needed for either. Beyond the endpoint:
-
Size the options to the server’s context window, not to a hosted model’s. For a
--ctx-size 8192server, the 10000-tokenMaxTokensdefault is larger than the entire window; llama.cpp will reject or clamp it. The prompt and the reply share that budget:Prompt component Approx. tokens Built-in rules and formatting instructions ~1,400 Tool definitions (5 information tools) ~600 Schema context at SchemaTableLimit: 5~600 User context, app-docs index ~400 That leaves roughly 5,000 for history, the question and the answer. Lower
SchemaTableLimitandConversationHistoryLimitto buy headroom; the tables you drop stay reachable viaget_table_schema. -
CPU inference is minutes, not seconds. Token generation is memory-bandwidth bound, so a 7B Q4 model runs at single-digit tokens/second whatever the core count. With
EnableThinking: falsethe call does not stream, so raiseTimeoutSecondsto cover a whole response body, and raiseRequestTimeoutSecondseither way. Turn the trace on to see where the time actually goes. -
Thinking models stream their reasoning. With
EnableThinkingon (the default), Ollama andOpenAICompatibleentries stream, and a reasoning model’sreasoning/reasoning_contentfield is shown in the Thinking panel — never as the reply.
Bringing your own IChatClient
Section titled “Bringing your own IChatClient”Any provider with a Microsoft.Extensions.AI implementation — Azure OpenAI / Foundry, Amazon
Bedrock, Anthropic’s C# SDK, Google’s Google.GenAI, OllamaSharp, or your own — can back a model in
the picker without engine code. Declare the entry with Provider: "ChatClient", and register the
client under the entry’s Name after AddGenieAssistant / AddGenieApp:
{ "Name": "Company Claude", "Provider": "ChatClient", "Model": "claude-sonnet-5", "MaxTokens": 8000 }// Program.cs — after AddGenieAssistant / AddGenieApp:builder.Services.AddKeyedChatClient("Company Claude", sp => /* your IChatClient */);The entry keeps everything configuration gives the built-in kinds — its place in the picker,
IsDefault, Enabled, MaxTokens / Temperature / TopP (sent as ChatOptions), Model (sent as
ModelId when set), ReasoningEffort (when it names a ReasoningEffort level) and the idle deadline —
and the same shared rules: truncation, reasoning kept out of the reply, and the turn trace. With
EnableThinking on it is called streaming, and any TextReasoningContent it yields goes to the
Thinking panel. The host owns the client’s lifetime (and its own timeouts and retries); the engine
never disposes it. An entry whose client is not registered fails the turn with a message naming the
AddKeyedChatClient call to make.
Backend wiring
Section titled “Backend wiring”The umbrella AddGenieApp already registers the Assistant module and maps the hub. A host that
composes the engine directly adds two calls:
services.AddGenieAssistant(builder.Configuration); // provider + MCP tools + services// …app.MapAssistantHubs(); // the /hubs/assistant-chat SignalR hubBoth entry points also accept code-first overrides (config-first + in-code override, code wins —
including registration-time values like Provider and ChatWidgetEnabled):
// Direct composition:services.AddGenieAssistant(builder.Configuration, o =>{ o.Provider = "OpenAI"; o.ChatWidgetEnabled = true; o.WelcomeChips = ["Which products are low on stock?"];});
// Umbrella AddGenieApp — via the fluent GenieBuilder:builder.Services.AddGenieApp(genie => genie .LoadFromConfiguration(builder.Configuration) .ConfigureAssistant(o => o.Model = "deepseek-chat"));AddGenieAssistant (ServiceCollectionExtensions) binds AssistantOptions, registers the one HTTP
client every built-in kind sends through (GenieAssistant), the core services (SchemaContextBuilder, NaturalLanguageQueryService,
ConversationService, ChatHistoryService, AppDocsService, AssistantQueryExecutor, …), the
four MCP tools, the McpOrchestratorService, the selected provider, and the AssistantChat
authorization policy that gates the hub. The policy authenticates with Cookie or JWT Bearer
(the same scheme pair as the notification hub’s ApiPolicy), so the SignalR ?access_token=
authenticates the connection and an unauthenticated negotiate gets a 401 — not a cookie login
redirect.
Configuration reference
Section titled “Configuration reference”Configuration lives in Genie:Assistant:
"Genie": { "Assistant": { "Providers": [ // the models the picker offers — see Models above { "Name": "GPT-4o mini", "Provider": "OpenAI", "Endpoint": "https://api.openai.com/v1", "Model": "gpt-4o-mini", "IsDefault": true } ], "ConversationMode": true, // keep chat context across turns "AutoGenerateTitle": true, // AI-summarise the first message into a chat title "DocsPath": "AssistantDocs", // optional: a dir of *.md help docs (merged over baseline) "ChatWidgetEnabled": true, // gates the AssistantChat auth policy / hub "ChatWidgetEnabledForRoles": "*", // "*" or a list/CSV of role names, e.g. ["Admin","Analyst"] "RequestTimeoutSeconds": 120, // overall deadline for one turn (MCP loop + streaming) "MaxIterations": 5 // tool-calling passes before the assistant must answer }}Other notable AssistantOptions knobs (all optional, sensible defaults):
- Per-model (on a
Providersentry, not here):MaxTokens(10000),Temperature(0.1),TopP(0.9), andTimeoutSeconds(60) — the timeout for a single provider HTTP request (connect + headers, and the full body of a non-streaming completion), applied per attempt. StreamIdleTimeoutSeconds(90, per-model) — how long a streamed answer may go without delivering a token before the call is abandoned.TimeoutSecondscannot cover this: it ends at the response headers, and an SSE body arrives after them. Observed against DeepSeek on 2026-09-15 —200 OK, then a: keep-alivecomment every 12 seconds and no tokens at all, on bothdeepseek-reasoneranddeepseek-chat, with a funded account. A keep-alive deliberately does not reset the clock: it proves the connection is alive, which is the one thing not in doubt. The caller is told the provider accepted the request and then sent nothing, and which model did it. Set to0to disable.ReasoningEffort(per-model; unset by default) —minimal,low,mediumorhigh, sent asreasoning_effort. Worth setting on the GPT-5 and o-series families, whose default ismediumand which pay for it on every pass of the tool-calling loop rather than once per answer: measured against GPT-5 Mini, passes ran 20–76 seconds each and an eight-pass turn hit the request deadline with no answer. The trade is smaller than it sounds, because most passes are “read the schema block, emit one tool call” — a formatting decision, not a reasoning problem. Unset sends nothing, and a server that rejects the parameter has it dropped after one round trip, so it is safe on an entry pointed at a proxy or a local model. (Adeepseek-reasonerentry still sendshighwhen this is unset, as it always has.)MaxIterations(5) — how many tool-calling passes one turn may take before the assistant is made to answer with what it has. Raise it for genuinely multi-step work: editing a report can spend a pass reading the draft, one per source it checks, one previewing and one applying, and running out mid-edit leaves the user with an explanation instead of a change.RequestTimeoutSeconds(120) — the overall deadline for one assistant turn: the whole multi-iteration MCP tool-calling loop plus the provider’s streamed reasoning read. A streamed SSE body isn’t covered byTimeoutSecondsonce the response headers arrive, so without this a stalled provider stream would leave the chat stuck on “Thinking” forever; when it elapses the caller is told “The assistant timed out.” Set to0to disable.ConversationHistoryLimit(3) — how many recent Q&A pairs (beyond the system prompt) are sent each turn, to control token usage.UserMemoryMaxEntries(0 — off) — how many of the user’s recent questions, taken from their other chats, are injected into every conversation’s system prompt. A chat’s own history is always sent in full and is unaffected; this is the cross-thread bleed, and it is off because a new chat that already knows what the last one was about is the opposite of what opening one is for. Set a small positive number to turn it on. (Before this it could not be turned off at all:0selected the old default of 8 instead of disabling it.)SchemaTableLimit(15) — how many tables are described inline in the system prompt; the rest remain reachable viaget_table_schema.ToolResultMaxChars(8000) — cap on a single tool result before truncation. A tool can opt out withIAssistantTool.TruncateResult => false, and the report tools do: half a report is not a cheaper report, it is one with no closing tag, and the model cannot tell that what it was handed was incomplete. Keep the default for anything row-shaped, where the first N rows still answer the question.WelcomeMessage/WelcomeChips— the greeting and quick-start suggestion chips for an empty chat, served to the React UI viaGET api/assistant/config. Per-field precedence: the host UI’sassistant.*config > these server values > built-in defaults — and there are no default chips: when neither the host UI nor the server provides any, no chips render.
API keys — keep them out of source
Section titled “API keys — keep them out of source”ASP.NET Core layers the environment-variable provider after the JSON files, and nested keys map
via the __ separator. Providers is an array, so a key targets an entry by its index:
# the FIRST entry in Genie:Assistant:Providerssetx Genie__Assistant__Providers__0__ApiKey "sk-…"# then restart the shell/IDE so the process inherits itThe index is positional: reorder the array and an override follows the position, not the model that
used to be there. A source’s connection string works the same way
(Genie__ReportingSources__1__ConnectionString). Ollama is local and needs no key.
A key addressed by name rather than index — …Providers__DeepSeekR1__ApiKey — fails at startup
rather than being ignored. Configuration is a flat key-value store, so that path binds as an entry
with a key and no name; refusing to start is the right outcome, because a secret that is silently
dropped looks identical to one that worked until the provider rejects the request.
Tool categories
Section titled “Tool categories”Every active tool’s name, description and parameter schema is written into the system prompt on
every provider call. A flat registry therefore grows the prompt with each capability added — and
prompt size is what decides whether a smaller local model still emits a well-formed TOOL_CALL:. So
tools declare a category, and only the categories relevant to what the user is doing are offered:
| Category | Tools | Active |
|---|---|---|
information |
execute_query, get_table_schema, get_app_doc, request_clarification, chart_result |
always |
reports |
report_get_draft, report_preview_data, report_apply_change |
only in the report designer |
wizards |
wizard_get_definition, wizard_apply_change |
only in the wizard designer (gated on Update of WizardDefinitions) |
workflows |
workflow_get_definition, workflow_apply_change |
only in the workflow designer (gated on Update of Wf_DefinitionsView) |
Five tools on an ordinary page, eight in the report designer, seven in the wizard or workflow designer — never the full set everywhere.
The React shell derives the mode from the current route and sends it with each message; nothing needs
configuring. A host tool keeps working unchanged, because IAssistantTool.Category is a default
interface member returning information:
public sealed class MyTool : IAssistantTool{ public string Name => "my_tool"; // Category not declared -> "information", offered on every turn.}Declare AssistantToolCategories.Reports, Wizards or Workflows to be offered only alongside that designer.
A mode also changes the instructions, not just the tool list. The bulk of the system prompt is
written for the data path — query the database, answer in business language — and on the designer
none of it applies. So report-designer mode appends its own block at the very end of the prompt,
naming the open report and stating the rules that override what came before: read the report with
report_get_draft before saying anything about it, never offer a change as XML in the reply (the
user has no way to apply it), and put every change through report_apply_change as a complete
document. Without
that block a smaller model treats a designer request like any other question and simply answers it.
report_get_draft returns the report XML contract after the document — the elements and attributes
a report may use, condensed from Report authoring. It is carried on the tool
result rather than in the prompt because it is reference material, and because it is then fetched once,
on the call the prompt already requires before any change. Without it the assistant’s only evidence for
what an element accepts is whatever the open report happens to contain, so anything the report does not
already use is a guess drawn from ordinary charting-library conventions — Type for Kind, a Series
attribute, a SplitBy that does not exist — and each wrong guess costs a whole round trip, re-sending
the document and waiting on another completion. Three of them is the turn.
Which databases a chat reads
Section titled “Which databases a chat reads”A chat reads one or more of the sources configured under
Genie:ReportingSources,
picked from the checkbox list under the chat’s title bar. Default is the application’s own database,
described from Genie’s modeled entities and filtered to the tables the caller may reach through a
permitted view; every other source is described by a host-written schema (see below).
The picker only offers sources the caller may actually query — the source name is the RBAC resource, so “may I see it in the list” and “may I query it” are the same question. Every name the browser sends is re-validated server-side against what exists, what is enabled, what the caller may view, and how many sources one chat may read at once (four). A name that fails fails the turn rather than being dropped: answering from a narrower set than the user believes they ticked would look identical to there being no data.
The selection is fixed for the duration of a turn and free to change between them. Fixed within a turn because a source decides both the connection opened and the tenant column rows are filtered on, so a mid-turn switch is a cross-tenant hazard rather than an inconvenience.
A chat remembers its selection, so reopening a thread does not silently answer the next question from a different database than the ones above it in the transcript.
Report designer
Section titled “Report designer”Opening the assistant on a report pre-ticks every source that report’s datasets already read. This is what lets it edit a report at all: an assistant that can only see one database cannot write a dataset for another, and before this it would produce a correct query for the wrong source and be rejected at the last step.
With more than one source in context, the assistant is told to ask which source a new dataset should use rather than choose — “total products” against a live table and against a replica are different numbers, and only the user knows which they meant.
Describing a source
Section titled “Describing a source”A source other than Default has no modeled entities, so nothing can generate a schema from it — a
host writes one. That is an ISourceSchemaProvider, bound to a source by its own Source
property:
public sealed class AnalyticsSchemaProvider : ISourceSchemaProvider{ public string Source => "Analytics";
public SourceSchema GetSchema() => new( [new SourceTableSchema("wh_Inventory_Products", "Product master.
| Column | … |")], Rules: "ALWAYS read with FINAL and add IsDeleted = 0.");}services.AddSingleton<ISourceSchemaProvider, AnalyticsSchemaProvider>(); // after AddGenieAssistantThe binding is a property the compiler sees, not a type name in configuration — so there is nothing to activate by reflection and no path-versus-classname ambiguity. Providers are applied in registration order: a later one overrides an earlier one’s table of the same name and appends its rules, so a small correction can layer over a large description. One that throws is logged and skipped, never allowed to take the assistant down.
Resolved once per source, as a singleton. Do not put per-user or per-request logic here; that is
what IAssistantContextContributor is for.
From a file
Section titled “From a file”SourceSchemaFile.Load(path) reads the same record from disk, for when the description is
documentation a data engineer should edit without a rebuild:
public SourceSchema GetSchema() => SourceSchemaFile.Load("Assistant/analytics.schema.xml");It dispatches on what the path is — .xml, .md, or a directory of either (plus a reserved
_rules.md taken whole as source-wide rules). Both formats produce the same record, so nothing
downstream can tell which was used.
The XML format:
<SourceSchema Source="Analytics"> <Table Name="wh_Inventory_Products" Summary="Product master — where most questions start."> <![CDATA[ | Column | Type | Notes | |---|---|---| | UnitCost | Decimal(18,2) | use this for stock VALUE |
There is no on-hand column — on-hand is the SUM of wh_Inventory_StockMovements.Quantity. ]]> </Table> <Rules> <Rule>ALWAYS read with `FROM <table> FINAL` and add `AND IsDeleted = 0`.</Rule> </Rules></SourceSchema>| Element | Meaning |
|---|---|
<SourceSchema Source="…"> |
Checked against the source the provider serves, so a file copied and re-pointed cannot quietly describe the wrong database. |
<Table Name="…" Summary="…"> |
Name is what the assistant writes in SQL. Summary is the one line in the prompt’s index; omit it and the first sentence of the body is used. |
| the table body | Opaque — handed to the model verbatim. There is deliberately no <Column> element: the half of a description worth having is the part a column grid cannot hold, like the note above about on-hand being a SUM. |
<Rules><Rule> |
Rules for the source as a whole, rendered as a list beneath the table index. |
Markdown is the same information with less ceremony — ## TableName declares a table and everything
under it is the body.
What XML earns over it is load-time validation, naming the file: a duplicate table name, a
<Table> with no name, an empty <Rule>, or a root Source that disagrees with the provider all
throw at load rather than surfacing later as an assistant that cannot find a table.
The catalog is the assistant’s only knowledge of a source: nothing probes the database, so a table nobody describes is a table it cannot name. A source with no provider at all is reported as undescribed rather than silently answered from the application database.
Query rules
Section titled “Query rules”SourceSchema.Rules is where the requirements that make a query correct rather than merely valid
belong — and for a Genie-warehoused ClickHouse target there is one that matters enormously:
- Every table here is a ReplacingMergeTree holding every version of a row, and a soft-deleted record keeps its last values rather than leaving. ALWAYS read with `FROM <table> FINAL`, and ALWAYS add `AND IsDeleted = 0` on a table that has the column.- Bind parameters the ClickHouse way — `{Name:Type}`, not `@Name`.Without FINAL a query sees every historical version of every row; without IsDeleted = 0 it sees
rows someone deleted. Neither fails — they just inflate the answer.
Declaring a source in code
Section titled “Declaring a source in code”Genie:ReportingSources entries can equally be declared or adjusted through the builder:
services.AddGenie(genie => genie .LoadFromConfiguration(configuration) .UseSource("Analytics", source => { source.Dialect = SourceDialect.ClickHouse; source.Connection = "Analytics"; // a ConnectionStrings key, or a literal source.TenantColumn = "CompanyId"; }));This overload configures rather than replaces, so anything LoadFromConfiguration bound survives
unless you set it. That is deliberate: an overload that rebuilt the entry would silently drop
TenantColumn, and a source with no tenant column refuses every non-System caller — safe, but
baffling if all you meant to change was a password. The three-argument
UseSource(name, dialect, connection) delegates to it and behaves the same way.
Access and tenant scoping
Section titled “Access and tenant scoping”The two protections the primary database provides do not reach a warehouse, so they move into the engine:
-
RBAC — the source is its own resource. A caller without
Viewon a resource named after the source is refused. (Warehouse tables sit behind no view, so the permission filtering that scopes the modeled schema has nothing to bite on.) -
Tenant — on the primary database the caller’s company is pushed into a DB session context and row-level-security filters rows whatever SQL the model wrote. ClickHouse has no such policy, so Genie wraps the query instead:
SELECT * FROM ( <the model's query> ) AS genie_scopeWHERE genie_scope."CompanyId" = {genie_tenant:Int64}An outer predicate constrains every row the inner query could produce, through any join, union or grouping. The cost is that the query must project the tenant column — the schema context says so up front, and if it is missing the database says so and the model corrects itself. Failing in that direction is the point: a filter that silently did not apply would be worse than none.
System callers cross tenants and are not filtered. Admin callers are — they administer one
company, and RLS scopes them on the primary path too.
What else follows the source
Section titled “What else follows the source”- The dialect — the prompt’s SQL rules, the wrong-dialect backstop and the row cap all switch to
the source’s engine, and
execute_query’s own description names it. - Report datasets — a change the assistant proposes puts every
<DataSet>on the active source, andreport_apply_changerejects a document that mixes sources. A hand-authored report may still mix them; the restriction is on what the assistant writes, because it has read one schema.
Tracing a turn
Section titled “Tracing a turn”One user message fans out across the hub, the conversation service, the MCP iteration loop and every tool it calls. Reconstructing “what happened to this message” from ordinary log lines means stitching a dozen events together by conversation id and hoping none interleaved with another user’s turn.
Genie:Assistant:Trace fixes that from both ends: it collects the whole turn into one tagged
block, and it stamps a short key on every other line the turn writes, so the whole message is one
grep.
"Genie": { "Assistant": { "Trace": { "Enabled": true, // default false "Verbosity": "Io", // Steps | Io | Full "Live": false, // also echo each step as it happens "MaxDetailChars": 20000 // per-payload cap; 0 = unlimited } }}The block
Section titled “The block”The event is written under the Genie.Assistant.Trace source context — route it to its own
Serilog sink, or filter it out — carrying Key, TurnId, Outcome, ConversationId, UserId,
CompanyId and TotalMs as properties, with the block itself in the message. At Verbosity: "Io":
[Assistant] A7C3F1 Answer in 43258.9ms — 42:7 (qwen2.5-coder-7b) [IN] [A7C3F1]-42:7 User Input — 28 chars, mode=report-designer key=PurchasingPerformance | show me low stock products [CTX] [A7C3F1]-42:7 Schema:CacheHit 1589 chars [ITER] [A7C3F1]-42:7 Iteration 1 (41074 ms) [LLM-REQ] [A7C3F1]-42:7 2 messages, 5400 chars (40915 ms) [THINK] [A7C3F1]-42:7 1204 chars | The user wants products below reorder level… [LLM-RES] [A7C3F1]-42:7 142 chars | TOOL_CALL: execute_query | PARAMS: {"sql":"SELECT …"} [TOOL-CALL][A7C3F1]-42:7 execute_query | {"sql":"SELECT …"} [TOOL-RES][A7C3F1]-42:7 execute_query — 0 rows, 128 chars | TOOL_RESULT: 0 rows [OUT] [A7C3F1]-42:7 Answer — 142 chars, 0 block(s) | No products are currently below their reorder level.| Tag | What it marks |
|---|---|
[IN] |
The user’s message, verbatim |
[CTX] |
Context assembly — schema, history, memory, the active tool set |
[SYS] |
The system prompt handed to the model |
[ITER] |
One pass of the MCP loop; everything below it nests under the pass that caused it |
[LLM-REQ] |
What was sent to the provider this pass |
[THINK] |
The model’s chain-of-thought, when it emits one |
[LLM-RES] |
The provider’s raw response, before any sanitising |
[TOOL-CALL] |
A tool invocation and its parameters |
[TOOL-RES] |
What the tool returned, including an ERROR: verdict |
[BLOCK] |
A rich block (chart, action card) emitted toward the UI |
[OUT] |
How the turn ended, and the answer text |
[ERR] |
A provider exception, a malformed payload, an unhandled throw |
Lines beginning | are verbatim captured text belonging to the step above them, indented so a
multi-line prompt cannot be mistaken for a run of sibling steps.
It flushes on every exit — an answer, a clarification, a pipeline error, the overall-deadline timeout, a user cancel, a dropped connection, an unhandled exception — because the turns worth reading are usually the ones that failed.
The key
Section titled “The key”A7C3F1 identifies the chat session; 42:7 is the chat id and the user message’s sequence
number within it. So grep A7C3F1 gives you the whole conversation and grep 'A7C3F1-42:7' gives
you exactly one message.
The key is derived from the conversation id rather than generated, so it is the same key every time that conversation is resumed and after a restart — a key copied out of yesterday’s log still finds the live chat.
The same A7C3F1-42:7 is pushed onto Serilog’s LogContext for the duration of the turn, so every
line the provider, the schema builder and each tool writes carries it too:
[23:25:17 INF] [A7C3F1-42:7] [Assistant:AppDocs] Loaded 3 app doc(s)[23:25:18 INF] [A7C3F1-42:7] Sending request to OpenAI — Model: qwen2.5-coder-7b, MessageCount: 2[23:26:01 INF] [A7C3F1-42:7] [Assistant] A7C3F1 Answer in 43258.9ms — 42:7[23:26:02 INF] [------] Hangfire server heartbeatThree lines are emitted before the turn’s message row exists and so cannot carry the key: the hub’s
receive and conversation-open lines (both Debug, and reproduced as the block’s [IN] and [CTX]
steps) and Created assistant thread ….
How much to keep
Section titled “How much to keep”| Verbosity | Keeps | Size |
|---|---|---|
Steps |
The outline and timings — which pass called which tool, and where the time went | negligible |
Io |
+ the question, the thinking, every provider response, every tool payload, the answer | ~5 KB a turn |
Full |
+ the system prompt, the schema, the docs index, and the whole message array resent each pass | 50–100 KB a turn |
Io is the day-to-day setting: it answers “what did the model actually say?” without dragging in the
prompt and schema that dominate Full and barely change between turns. Reach for Full when the
question is about the prompt itself.
MaxDetailChars caps any single captured payload and marks the cut explicitly
(… [truncated 3100 chars]), so a clipped 13 KB report can never be mistaken for a short one.
Live: true additionally echoes each step as a Debug line the moment it happens, on top of the
block. Use it when a turn is slow or hanging and the question is “where is it right now?” — a
40-second provider call otherwise looks identical to a deadlock until the turn ends.
When disabled the accumulator records nothing and no event is written, so the pipeline pays nothing. This replaces the former per-iteration Debug transcript, which wrote one event per provider call.
Extending the assistant (host seams)
Section titled “Extending the assistant (host seams)”The assistant is designed to be extended and overridden from the caller project, entirely through DI — no forking:
-
Business context in the prompt —
IAssistantContextContributor. Register any number of implementations; each renders as a titled=== {Title} ===section in the system prompt (after the user-context block, before the app-docs index), ordered byPriority(lower first). A contributor that throws is logged and skipped — it never fails the turn. Contributors run on every assistant turn, so keep sections short and cache expensive queries:public sealed class StockSummaryContextContributor(MyContext db, IMemoryCache cache): IAssistantContextContributor{public string Title => "INVENTORY SNAPSHOT";public async Task<string?> BuildContextAsync(AssistantUserContext user, CancellationToken ct) =>await cache.GetOrCreateAsync("stock-summary", async e =>{e.AbsoluteExpirationRelativeToNow = TimeSpan.FromMinutes(5);return $"Current totals: {await db.Products.CountAsync(ct)} products.";});}// Program.cs — after AddGenieAssistant / AddGenieApp:services.AddScoped<IAssistantContextContributor, StockSummaryContextContributor>(); -
Custom / replacement tools —
IAssistantTool.services.AddScoped<IAssistantTool, MyTool>()adds a tool to the MCP loop (itsName/Description/ParameterSchemaare injected into the prompt automatically). Tool names are case-insensitive and the last registration wins, so a host tool namedexecute_queryreplaces the built-in. -
Replacing the provider —
IAssistantProvider. Register your own implementation afterAddGenieAssistantand single-service resolution takes the last registration:services.AddScoped<IAssistantProvider, MySemanticKernelProvider>();— the whole pipeline (orchestrator, titles, SQL generation) flows through it.
The Inventory sample ships a working contributor
(StockSummaryContextContributor) plus server-configured welcome chips.
Frontend createGenieApp module
Section titled “Frontend createGenieApp module”Mount the assistant by enabling the module. It renders app-wide over every route — an AI icon in
the header (left of the notification bell) opening a themed bottom-right panel. An optional
top-level assistant block overrides the panel content:
createGenieApp({ apiBase: "/api/v1", modules: { assistant: true }, assistant: { title: "Inventory Copilot", // header title subtitle: "Online", // status line welcome: "Hi! Ask me anything.", // empty-state greeting suggestions: ["Low stock?"], // quick-start chips ([] renders none, beating server chips) placeholder: "Ask Copilot…", // composer placeholder },});welcome and suggestions fall back to the server’s Genie:Assistant:WelcomeMessage /
Genie:Assistant:WelcomeChips (fetched once from GET api/assistant/config when the panel first opens),
then to the built-in greeting — with no built-in chips.
The panel is multi-conversation: a conversations list (open / rename / delete / new, plus
multi-select bulk delete) and the conversation view — user/assistant bubbles as sanitised Markdown
(GFM tables, lists, code) via marked + DOMPurify, a live Thinking disclosure, a collapsible
SQL block and result-table preview for data answers, suggestion chips, a Shift+Enter
composer, and a stop control while a reply is in flight. It is accent-/theme-aware.
The UI talks to the backend over AssistantChatHub when connected and falls back to REST
(api/assistant: GET /config for the welcome content, POST /query, thread CRUD
GET/POST /chats, GET/PATCH/DELETE /chats/{id}, POST /chats/delete for bulk delete)
otherwise — clarification pills and inline charts work on both transports. GenieAssistant and
GenieAssistantButton are also exported for hosts that want to place or configure them manually
(custom title, subtitle, suggestions, welcome). See
Frontend configuration.
Persistence, titles and error masking
Section titled “Persistence, titles and error masking”- Persistence. Conversations live in the database (schema
Genie, tablesAssistantChat+AssistantChatMessage), owned per user (UserId+CompanyId). Non-admins only see their own threads. A thread is saved only once its first reply succeeds; a failed first turn is rolled back. Host apps must add the migration for the two tables (dotnet ef migrations add AddAssistantChat -c YourContext), and again after an engine upgrade that adds a column to them — the nullableAssistantChatMessage.ProviderName, which records the model that produced each reply so the UI can label it, arrived that way. - Titles. New threads show a truncated placeholder immediately. With
AutoGenerateTitleon (default) the provider summarises the first message into a concise title after the first reply, pushed live to the list (ConversationRenamed); on failure the placeholder is kept silently. A manual rename (PATCH /chats/{id}) always overrides a generated title. - Error masking. Internal error detail (a missing/invalid key, a raw provider failure) is
surfaced only to users holding the
Systemrole (AssistantErrors.ForUser). Everyone else sees a generic “The assistant is temporarily unavailable…” message. Generated SQL is likewise returned to admins only.
Where to go next
Section titled “Where to go next”- Tenant isolation & SQL hardening — the read-only login + RLS setup that makes tenant scope DB-enforced.
- AI Assistant overview — the tool-orchestration model and prompt context.