Skip to content

Configuration & providers

The Assistant is configured entirely through the root Assistant configuration section (bound to AssistantOptions) plus two host calls on the backend and one module flag on the frontend. This page covers the providers, the wiring, and the runtime behaviour you can tune.

For what the feature does, see the AI Assistant overview. For DB-enforced tenant isolation, see Tenant isolation & SQL hardening.

A chat picks the model that answers it, from the list under Genie:Assistant:Providers:

"Genie": {
"Assistant": {
"Providers": [
{ "Name": "DeepSeek R1", "Provider": "DeepSeek", "Endpoint": "https://api.deepseek.com/v1",
"Model": "deepseek-reasoner", "MaxTokens": 10000, "IsDefault": true },
{ "Name": "Local Llama", "Provider": "Ollama", "Endpoint": "http://localhost:11434",
"Model": "codellama:7b" }
]
}
}

An array, so the order is the order the model picker lists — and the index an out-of-band API key targets. Name is free text: it is what the picker shows and what a chat remembers, so it can read like a name rather than a model id. It must be unique; IsDefault marks the one a new chat starts on; Enabled: false takes a model out of service without deleting the entry that documents it.

Why per-chat rather than per-deployment: the useful choice differs by question. A reasoning model is worth its latency for a report someone is about to publish, and wasted on “how many suppliers do we have”.

Every model is reached through Microsoft.Extensions.AI’s IChatClient. The built-in kinds are all served by the OpenAI client pointed at the kind’s endpoint — DeepSeek, Ollama and Gemini speak OpenAI Chat Completions — and a host can bring any other provider as a ChatClient entry (below). The kinds, and what goes with each:

Provider Default Endpoint (optional) Example Model API key
OpenAI https://api.openai.com/v1 gpt-4o-mini, gpt-5-mini required
DeepSeek https://api.deepseek.com/v1 deepseek-chat, deepseek-reasoner required
Gemini https://generativelanguage.googleapis.com/v1beta/openai gemini-2.5-flash required
Ollama http://localhost:11434 (served under /v1) qwen3, llama3.1:8b not required (local)
OpenAICompatible none — required whatever the server serves optional
ChatClient not used optional (sent as ModelId) the host’s client handles it

Endpoint is the API root. Leave it out and a hosted kind uses its own service. A trailing /chat/completions is accepted and stripped, so entries written with the full completions URL keep working; Ollama’s root gets /v1 added, and Gemini’s native root (…/v1beta) is pointed at its OpenAI-compatible surface (…/v1beta/openai). Gemini is called through that surface on purpose: its native API took the key in the query string, where proxies and HTTP logs record it — the compatible surface takes it as a bearer header.

Every entry is validated at startup — a missing or duplicate Name, or a Provider that is not one of the kinds above, fails there rather than on the first question someone asks.

Per-entry settings beyond those: MaxTokens, Temperature, TopP, TimeoutSeconds, StreamIdleTimeoutSeconds, EnableThinking, ReasoningEffort, OrganizationId.

What is the same for every kind, because it is decided once in the provider layer rather than per provider:

  • Retries. A 408/429/5xx or a transport failure is retried twice, with backoff, before the turn fails. A rejected request (400/401/404) is never retried — it is refused identically every time — and is reported with the status, the URL dialled and the server’s own explanation (with a key-length hint on a 401, never the key).
  • Running out of tokens. An answer cut off at MaxTokens triggers the assistant’s compaction retry on every kind. (Before, only DeepSeek and OpenAI’s reasoning path noticed.)
  • Reasoning never becomes the reply (see the thinking notes below).
  • A streamed answer that goes silent is abandoned after StreamIdleTimeoutSeconds.
  • Keys stay out of errors and logs: any echo of the key in a provider’s error body is redacted before it reaches an exception, the turn trace or the log.
  • OpenTelemetry: each model call is an Activity on the Genie.Assistant source (a no-op unless a listener is attached), recorded without message content — the turn trace remains the only place chat text is logged.

Running a local model (llama.cpp, Ollama, vLLM, LM Studio)

Section titled “Running a local model (llama.cpp, Ollama, vLLM, LM Studio)”

Ollama has its own kind; any other server that exposes the OpenAI API (llama-server, vLLM, LM Studio, a proxy) is an OpenAICompatible entry with its API root as Endpoint:

{ "Name": "Local Qwen", "Provider": "OpenAICompatible",
"Endpoint": "http://my-host:18437/v1", "Model": "qwen2.5-coder-7b" }

No ApiKey is needed for either. Beyond the endpoint:

  • Size the options to the server’s context window, not to a hosted model’s. For a --ctx-size 8192 server, the 10000-token MaxTokens default is larger than the entire window; llama.cpp will reject or clamp it. The prompt and the reply share that budget:

    Prompt component Approx. tokens
    Built-in rules and formatting instructions ~1,400
    Tool definitions (5 information tools) ~600
    Schema context at SchemaTableLimit: 5 ~600
    User context, app-docs index ~400

    That leaves roughly 5,000 for history, the question and the answer. Lower SchemaTableLimit and ConversationHistoryLimit to buy headroom; the tables you drop stay reachable via get_table_schema.

  • CPU inference is minutes, not seconds. Token generation is memory-bandwidth bound, so a 7B Q4 model runs at single-digit tokens/second whatever the core count. With EnableThinking: false the call does not stream, so raise TimeoutSeconds to cover a whole response body, and raise RequestTimeoutSeconds either way. Turn the trace on to see where the time actually goes.

  • Thinking models stream their reasoning. With EnableThinking on (the default), Ollama and OpenAICompatible entries stream, and a reasoning model’s reasoning / reasoning_content field is shown in the Thinking panel — never as the reply.

Any provider with a Microsoft.Extensions.AI implementation — Azure OpenAI / Foundry, Amazon Bedrock, Anthropic’s C# SDK, Google’s Google.GenAI, OllamaSharp, or your own — can back a model in the picker without engine code. Declare the entry with Provider: "ChatClient", and register the client under the entry’s Name after AddGenieAssistant / AddGenieApp:

{ "Name": "Company Claude", "Provider": "ChatClient", "Model": "claude-sonnet-5", "MaxTokens": 8000 }
// Program.cs — after AddGenieAssistant / AddGenieApp:
builder.Services.AddKeyedChatClient("Company Claude", sp => /* your IChatClient */);

The entry keeps everything configuration gives the built-in kinds — its place in the picker, IsDefault, Enabled, MaxTokens / Temperature / TopP (sent as ChatOptions), Model (sent as ModelId when set), ReasoningEffort (when it names a ReasoningEffort level) and the idle deadline — and the same shared rules: truncation, reasoning kept out of the reply, and the turn trace. With EnableThinking on it is called streaming, and any TextReasoningContent it yields goes to the Thinking panel. The host owns the client’s lifetime (and its own timeouts and retries); the engine never disposes it. An entry whose client is not registered fails the turn with a message naming the AddKeyedChatClient call to make.

The umbrella AddGenieApp already registers the Assistant module and maps the hub. A host that composes the engine directly adds two calls:

Program.cs
services.AddGenieAssistant(builder.Configuration); // provider + MCP tools + services
// …
app.MapAssistantHubs(); // the /hubs/assistant-chat SignalR hub

Both entry points also accept code-first overrides (config-first + in-code override, code wins — including registration-time values like Provider and ChatWidgetEnabled):

// Direct composition:
services.AddGenieAssistant(builder.Configuration, o =>
{
o.Provider = "OpenAI";
o.ChatWidgetEnabled = true;
o.WelcomeChips = ["Which products are low on stock?"];
});
// Umbrella AddGenieApp — via the fluent GenieBuilder:
builder.Services.AddGenieApp(genie => genie
.LoadFromConfiguration(builder.Configuration)
.ConfigureAssistant(o => o.Model = "deepseek-chat"));

AddGenieAssistant (ServiceCollectionExtensions) binds AssistantOptions, registers the one HTTP client every built-in kind sends through (GenieAssistant), the core services (SchemaContextBuilder, NaturalLanguageQueryService, ConversationService, ChatHistoryService, AppDocsService, AssistantQueryExecutor, …), the four MCP tools, the McpOrchestratorService, the selected provider, and the AssistantChat authorization policy that gates the hub. The policy authenticates with Cookie or JWT Bearer (the same scheme pair as the notification hub’s ApiPolicy), so the SignalR ?access_token= authenticates the connection and an unauthenticated negotiate gets a 401 — not a cookie login redirect.

Configuration lives in Genie:Assistant:

appsettings.json
"Genie": {
"Assistant": {
"Providers": [ // the models the picker offers — see Models above
{ "Name": "GPT-4o mini", "Provider": "OpenAI",
"Endpoint": "https://api.openai.com/v1", "Model": "gpt-4o-mini", "IsDefault": true }
],
"ConversationMode": true, // keep chat context across turns
"AutoGenerateTitle": true, // AI-summarise the first message into a chat title
"DocsPath": "AssistantDocs", // optional: a dir of *.md help docs (merged over baseline)
"ChatWidgetEnabled": true, // gates the AssistantChat auth policy / hub
"ChatWidgetEnabledForRoles": "*", // "*" or a list/CSV of role names, e.g. ["Admin","Analyst"]
"RequestTimeoutSeconds": 120, // overall deadline for one turn (MCP loop + streaming)
"MaxIterations": 5 // tool-calling passes before the assistant must answer
}
}

Other notable AssistantOptions knobs (all optional, sensible defaults):

  • Per-model (on a Providers entry, not here): MaxTokens (10000), Temperature (0.1), TopP (0.9), and TimeoutSeconds (60) — the timeout for a single provider HTTP request (connect + headers, and the full body of a non-streaming completion), applied per attempt.
  • StreamIdleTimeoutSeconds (90, per-model) — how long a streamed answer may go without delivering a token before the call is abandoned. TimeoutSeconds cannot cover this: it ends at the response headers, and an SSE body arrives after them. Observed against DeepSeek on 2026-09-15 — 200 OK, then a : keep-alive comment every 12 seconds and no tokens at all, on both deepseek-reasoner and deepseek-chat, with a funded account. A keep-alive deliberately does not reset the clock: it proves the connection is alive, which is the one thing not in doubt. The caller is told the provider accepted the request and then sent nothing, and which model did it. Set to 0 to disable.
  • ReasoningEffort (per-model; unset by default) — minimal, low, medium or high, sent as reasoning_effort. Worth setting on the GPT-5 and o-series families, whose default is medium and which pay for it on every pass of the tool-calling loop rather than once per answer: measured against GPT-5 Mini, passes ran 20–76 seconds each and an eight-pass turn hit the request deadline with no answer. The trade is smaller than it sounds, because most passes are “read the schema block, emit one tool call” — a formatting decision, not a reasoning problem. Unset sends nothing, and a server that rejects the parameter has it dropped after one round trip, so it is safe on an entry pointed at a proxy or a local model. (A deepseek-reasoner entry still sends high when this is unset, as it always has.)
  • MaxIterations (5) — how many tool-calling passes one turn may take before the assistant is made to answer with what it has. Raise it for genuinely multi-step work: editing a report can spend a pass reading the draft, one per source it checks, one previewing and one applying, and running out mid-edit leaves the user with an explanation instead of a change.
  • RequestTimeoutSeconds (120) — the overall deadline for one assistant turn: the whole multi-iteration MCP tool-calling loop plus the provider’s streamed reasoning read. A streamed SSE body isn’t covered by TimeoutSeconds once the response headers arrive, so without this a stalled provider stream would leave the chat stuck on “Thinking” forever; when it elapses the caller is told “The assistant timed out.” Set to 0 to disable.
  • ConversationHistoryLimit (3) — how many recent Q&A pairs (beyond the system prompt) are sent each turn, to control token usage.
  • UserMemoryMaxEntries (0 — off) — how many of the user’s recent questions, taken from their other chats, are injected into every conversation’s system prompt. A chat’s own history is always sent in full and is unaffected; this is the cross-thread bleed, and it is off because a new chat that already knows what the last one was about is the opposite of what opening one is for. Set a small positive number to turn it on. (Before this it could not be turned off at all: 0 selected the old default of 8 instead of disabling it.)
  • SchemaTableLimit (15) — how many tables are described inline in the system prompt; the rest remain reachable via get_table_schema.
  • ToolResultMaxChars (8000) — cap on a single tool result before truncation. A tool can opt out with IAssistantTool.TruncateResult => false, and the report tools do: half a report is not a cheaper report, it is one with no closing tag, and the model cannot tell that what it was handed was incomplete. Keep the default for anything row-shaped, where the first N rows still answer the question.
  • WelcomeMessage / WelcomeChips — the greeting and quick-start suggestion chips for an empty chat, served to the React UI via GET api/assistant/config. Per-field precedence: the host UI’s assistant.* config > these server values > built-in defaults — and there are no default chips: when neither the host UI nor the server provides any, no chips render.

ASP.NET Core layers the environment-variable provider after the JSON files, and nested keys map via the __ separator. Providers is an array, so a key targets an entry by its index:

Terminal window
# the FIRST entry in Genie:Assistant:Providers
setx Genie__Assistant__Providers__0__ApiKey "sk-…"
# then restart the shell/IDE so the process inherits it

The index is positional: reorder the array and an override follows the position, not the model that used to be there. A source’s connection string works the same way (Genie__ReportingSources__1__ConnectionString). Ollama is local and needs no key.

A key addressed by name rather than index — …Providers__DeepSeekR1__ApiKey — fails at startup rather than being ignored. Configuration is a flat key-value store, so that path binds as an entry with a key and no name; refusing to start is the right outcome, because a secret that is silently dropped looks identical to one that worked until the provider rejects the request.

Every active tool’s name, description and parameter schema is written into the system prompt on every provider call. A flat registry therefore grows the prompt with each capability added — and prompt size is what decides whether a smaller local model still emits a well-formed TOOL_CALL:. So tools declare a category, and only the categories relevant to what the user is doing are offered:

Category Tools Active
information execute_query, get_table_schema, get_app_doc, request_clarification, chart_result always
reports report_get_draft, report_preview_data, report_apply_change only in the report designer
wizards wizard_get_definition, wizard_apply_change only in the wizard designer (gated on Update of WizardDefinitions)
workflows workflow_get_definition, workflow_apply_change only in the workflow designer (gated on Update of Wf_DefinitionsView)

Five tools on an ordinary page, eight in the report designer, seven in the wizard or workflow designer — never the full set everywhere.

The React shell derives the mode from the current route and sends it with each message; nothing needs configuring. A host tool keeps working unchanged, because IAssistantTool.Category is a default interface member returning information:

public sealed class MyTool : IAssistantTool
{
public string Name => "my_tool";
// Category not declared -> "information", offered on every turn.
}

Declare AssistantToolCategories.Reports, Wizards or Workflows to be offered only alongside that designer.

A mode also changes the instructions, not just the tool list. The bulk of the system prompt is written for the data path — query the database, answer in business language — and on the designer none of it applies. So report-designer mode appends its own block at the very end of the prompt, naming the open report and stating the rules that override what came before: read the report with report_get_draft before saying anything about it, never offer a change as XML in the reply (the user has no way to apply it), and put every change through report_apply_change as a complete document. Without that block a smaller model treats a designer request like any other question and simply answers it.

report_get_draft returns the report XML contract after the document — the elements and attributes a report may use, condensed from Report authoring. It is carried on the tool result rather than in the prompt because it is reference material, and because it is then fetched once, on the call the prompt already requires before any change. Without it the assistant’s only evidence for what an element accepts is whatever the open report happens to contain, so anything the report does not already use is a guess drawn from ordinary charting-library conventions — Type for Kind, a Series attribute, a SplitBy that does not exist — and each wrong guess costs a whole round trip, re-sending the document and waiting on another completion. Three of them is the turn.

A chat reads one or more of the sources configured under Genie:ReportingSources, picked from the checkbox list under the chat’s title bar. Default is the application’s own database, described from Genie’s modeled entities and filtered to the tables the caller may reach through a permitted view; every other source is described by a host-written schema (see below).

The picker only offers sources the caller may actually query — the source name is the RBAC resource, so “may I see it in the list” and “may I query it” are the same question. Every name the browser sends is re-validated server-side against what exists, what is enabled, what the caller may view, and how many sources one chat may read at once (four). A name that fails fails the turn rather than being dropped: answering from a narrower set than the user believes they ticked would look identical to there being no data.

The selection is fixed for the duration of a turn and free to change between them. Fixed within a turn because a source decides both the connection opened and the tenant column rows are filtered on, so a mid-turn switch is a cross-tenant hazard rather than an inconvenience.

A chat remembers its selection, so reopening a thread does not silently answer the next question from a different database than the ones above it in the transcript.

Opening the assistant on a report pre-ticks every source that report’s datasets already read. This is what lets it edit a report at all: an assistant that can only see one database cannot write a dataset for another, and before this it would produce a correct query for the wrong source and be rejected at the last step.

With more than one source in context, the assistant is told to ask which source a new dataset should use rather than choose — “total products” against a live table and against a replica are different numbers, and only the user knows which they meant.

A source other than Default has no modeled entities, so nothing can generate a schema from it — a host writes one. That is an ISourceSchemaProvider, bound to a source by its own Source property:

public sealed class AnalyticsSchemaProvider : ISourceSchemaProvider
{
public string Source => "Analytics";
public SourceSchema GetSchema() => new(
[new SourceTableSchema("wh_Inventory_Products", "Product master.
| Column | … |")],
Rules: "ALWAYS read with FINAL and add IsDeleted = 0.");
}
services.AddSingleton<ISourceSchemaProvider, AnalyticsSchemaProvider>(); // after AddGenieAssistant

The binding is a property the compiler sees, not a type name in configuration — so there is nothing to activate by reflection and no path-versus-classname ambiguity. Providers are applied in registration order: a later one overrides an earlier one’s table of the same name and appends its rules, so a small correction can layer over a large description. One that throws is logged and skipped, never allowed to take the assistant down.

Resolved once per source, as a singleton. Do not put per-user or per-request logic here; that is what IAssistantContextContributor is for.

SourceSchemaFile.Load(path) reads the same record from disk, for when the description is documentation a data engineer should edit without a rebuild:

public SourceSchema GetSchema() => SourceSchemaFile.Load("Assistant/analytics.schema.xml");

It dispatches on what the path is — .xml, .md, or a directory of either (plus a reserved _rules.md taken whole as source-wide rules). Both formats produce the same record, so nothing downstream can tell which was used.

The XML format:

<SourceSchema Source="Analytics">
<Table Name="wh_Inventory_Products" Summary="Product master — where most questions start.">
<![CDATA[
| Column | Type | Notes |
|---|---|---|
| UnitCost | Decimal(18,2) | use this for stock VALUE |
There is no on-hand column — on-hand is the SUM of wh_Inventory_StockMovements.Quantity.
]]>
</Table>
<Rules>
<Rule>ALWAYS read with `FROM &lt;table&gt; FINAL` and add `AND IsDeleted = 0`.</Rule>
</Rules>
</SourceSchema>
Element Meaning
<SourceSchema Source="…"> Checked against the source the provider serves, so a file copied and re-pointed cannot quietly describe the wrong database.
<Table Name="…" Summary="…"> Name is what the assistant writes in SQL. Summary is the one line in the prompt’s index; omit it and the first sentence of the body is used.
the table body Opaque — handed to the model verbatim. There is deliberately no <Column> element: the half of a description worth having is the part a column grid cannot hold, like the note above about on-hand being a SUM.
<Rules><Rule> Rules for the source as a whole, rendered as a list beneath the table index.

Markdown is the same information with less ceremony — ## TableName declares a table and everything under it is the body.

What XML earns over it is load-time validation, naming the file: a duplicate table name, a <Table> with no name, an empty <Rule>, or a root Source that disagrees with the provider all throw at load rather than surfacing later as an assistant that cannot find a table.

The catalog is the assistant’s only knowledge of a source: nothing probes the database, so a table nobody describes is a table it cannot name. A source with no provider at all is reported as undescribed rather than silently answered from the application database.

SourceSchema.Rules is where the requirements that make a query correct rather than merely valid belong — and for a Genie-warehoused ClickHouse target there is one that matters enormously:

- Every table here is a ReplacingMergeTree holding every version of a row, and a soft-deleted record
keeps its last values rather than leaving. ALWAYS read with `FROM <table> FINAL`, and ALWAYS add
`AND IsDeleted = 0` on a table that has the column.
- Bind parameters the ClickHouse way — `{Name:Type}`, not `@Name`.

Without FINAL a query sees every historical version of every row; without IsDeleted = 0 it sees rows someone deleted. Neither fails — they just inflate the answer.

Genie:ReportingSources entries can equally be declared or adjusted through the builder:

services.AddGenie(genie => genie
.LoadFromConfiguration(configuration)
.UseSource("Analytics", source =>
{
source.Dialect = SourceDialect.ClickHouse;
source.Connection = "Analytics"; // a ConnectionStrings key, or a literal
source.TenantColumn = "CompanyId";
}));

This overload configures rather than replaces, so anything LoadFromConfiguration bound survives unless you set it. That is deliberate: an overload that rebuilt the entry would silently drop TenantColumn, and a source with no tenant column refuses every non-System caller — safe, but baffling if all you meant to change was a password. The three-argument UseSource(name, dialect, connection) delegates to it and behaves the same way.

The two protections the primary database provides do not reach a warehouse, so they move into the engine:

  • RBAC — the source is its own resource. A caller without View on a resource named after the source is refused. (Warehouse tables sit behind no view, so the permission filtering that scopes the modeled schema has nothing to bite on.)

  • Tenant — on the primary database the caller’s company is pushed into a DB session context and row-level-security filters rows whatever SQL the model wrote. ClickHouse has no such policy, so Genie wraps the query instead:

    SELECT * FROM ( <the model's query> ) AS genie_scope
    WHERE genie_scope."CompanyId" = {genie_tenant:Int64}

    An outer predicate constrains every row the inner query could produce, through any join, union or grouping. The cost is that the query must project the tenant column — the schema context says so up front, and if it is missing the database says so and the model corrects itself. Failing in that direction is the point: a filter that silently did not apply would be worse than none.

System callers cross tenants and are not filtered. Admin callers are — they administer one company, and RLS scopes them on the primary path too.

  • The dialect — the prompt’s SQL rules, the wrong-dialect backstop and the row cap all switch to the source’s engine, and execute_query’s own description names it.
  • Report datasets — a change the assistant proposes puts every <DataSet> on the active source, and report_apply_change rejects a document that mixes sources. A hand-authored report may still mix them; the restriction is on what the assistant writes, because it has read one schema.

One user message fans out across the hub, the conversation service, the MCP iteration loop and every tool it calls. Reconstructing “what happened to this message” from ordinary log lines means stitching a dozen events together by conversation id and hoping none interleaved with another user’s turn.

Genie:Assistant:Trace fixes that from both ends: it collects the whole turn into one tagged block, and it stamps a short key on every other line the turn writes, so the whole message is one grep.

"Genie": {
"Assistant": {
"Trace": {
"Enabled": true, // default false
"Verbosity": "Io", // Steps | Io | Full
"Live": false, // also echo each step as it happens
"MaxDetailChars": 20000 // per-payload cap; 0 = unlimited
}
}
}

The event is written under the Genie.Assistant.Trace source context — route it to its own Serilog sink, or filter it out — carrying Key, TurnId, Outcome, ConversationId, UserId, CompanyId and TotalMs as properties, with the block itself in the message. At Verbosity: "Io":

[Assistant] A7C3F1 Answer in 43258.9ms — 42:7 (qwen2.5-coder-7b)
[IN] [A7C3F1]-42:7 User Input — 28 chars, mode=report-designer key=PurchasingPerformance
| show me low stock products
[CTX] [A7C3F1]-42:7 Schema:CacheHit 1589 chars
[ITER] [A7C3F1]-42:7 Iteration 1 (41074 ms)
[LLM-REQ] [A7C3F1]-42:7 2 messages, 5400 chars (40915 ms)
[THINK] [A7C3F1]-42:7 1204 chars
| The user wants products below reorder level…
[LLM-RES] [A7C3F1]-42:7 142 chars
| TOOL_CALL: execute_query
| PARAMS: {"sql":"SELECT …"}
[TOOL-CALL][A7C3F1]-42:7 execute_query
| {"sql":"SELECT …"}
[TOOL-RES][A7C3F1]-42:7 execute_query — 0 rows, 128 chars
| TOOL_RESULT: 0 rows
[OUT] [A7C3F1]-42:7 Answer — 142 chars, 0 block(s)
| No products are currently below their reorder level.
Tag What it marks
[IN] The user’s message, verbatim
[CTX] Context assembly — schema, history, memory, the active tool set
[SYS] The system prompt handed to the model
[ITER] One pass of the MCP loop; everything below it nests under the pass that caused it
[LLM-REQ] What was sent to the provider this pass
[THINK] The model’s chain-of-thought, when it emits one
[LLM-RES] The provider’s raw response, before any sanitising
[TOOL-CALL] A tool invocation and its parameters
[TOOL-RES] What the tool returned, including an ERROR: verdict
[BLOCK] A rich block (chart, action card) emitted toward the UI
[OUT] How the turn ended, and the answer text
[ERR] A provider exception, a malformed payload, an unhandled throw

Lines beginning | are verbatim captured text belonging to the step above them, indented so a multi-line prompt cannot be mistaken for a run of sibling steps.

It flushes on every exit — an answer, a clarification, a pipeline error, the overall-deadline timeout, a user cancel, a dropped connection, an unhandled exception — because the turns worth reading are usually the ones that failed.

A7C3F1 identifies the chat session; 42:7 is the chat id and the user message’s sequence number within it. So grep A7C3F1 gives you the whole conversation and grep 'A7C3F1-42:7' gives you exactly one message.

The key is derived from the conversation id rather than generated, so it is the same key every time that conversation is resumed and after a restart — a key copied out of yesterday’s log still finds the live chat.

The same A7C3F1-42:7 is pushed onto Serilog’s LogContext for the duration of the turn, so every line the provider, the schema builder and each tool writes carries it too:

[23:25:17 INF] [A7C3F1-42:7] [Assistant:AppDocs] Loaded 3 app doc(s)
[23:25:18 INF] [A7C3F1-42:7] Sending request to OpenAI — Model: qwen2.5-coder-7b, MessageCount: 2
[23:26:01 INF] [A7C3F1-42:7] [Assistant] A7C3F1 Answer in 43258.9ms — 42:7
[23:26:02 INF] [------] Hangfire server heartbeat

Three lines are emitted before the turn’s message row exists and so cannot carry the key: the hub’s receive and conversation-open lines (both Debug, and reproduced as the block’s [IN] and [CTX] steps) and Created assistant thread ….

Verbosity Keeps Size
Steps The outline and timings — which pass called which tool, and where the time went negligible
Io + the question, the thinking, every provider response, every tool payload, the answer ~5 KB a turn
Full + the system prompt, the schema, the docs index, and the whole message array resent each pass 50–100 KB a turn

Io is the day-to-day setting: it answers “what did the model actually say?” without dragging in the prompt and schema that dominate Full and barely change between turns. Reach for Full when the question is about the prompt itself.

MaxDetailChars caps any single captured payload and marks the cut explicitly (… [truncated 3100 chars]), so a clipped 13 KB report can never be mistaken for a short one.

Live: true additionally echoes each step as a Debug line the moment it happens, on top of the block. Use it when a turn is slow or hanging and the question is “where is it right now?” — a 40-second provider call otherwise looks identical to a deadlock until the turn ends.

When disabled the accumulator records nothing and no event is written, so the pipeline pays nothing. This replaces the former per-iteration Debug transcript, which wrote one event per provider call.

The assistant is designed to be extended and overridden from the caller project, entirely through DI — no forking:

  • Business context in the prompt — IAssistantContextContributor. Register any number of implementations; each renders as a titled === {Title} === section in the system prompt (after the user-context block, before the app-docs index), ordered by Priority (lower first). A contributor that throws is logged and skipped — it never fails the turn. Contributors run on every assistant turn, so keep sections short and cache expensive queries:

    public sealed class StockSummaryContextContributor(MyContext db, IMemoryCache cache)
    : IAssistantContextContributor
    {
    public string Title => "INVENTORY SNAPSHOT";
    public async Task<string?> BuildContextAsync(AssistantUserContext user, CancellationToken ct) =>
    await cache.GetOrCreateAsync("stock-summary", async e =>
    {
    e.AbsoluteExpirationRelativeToNow = TimeSpan.FromMinutes(5);
    return $"Current totals: {await db.Products.CountAsync(ct)} products.";
    });
    }
    // Program.cs — after AddGenieAssistant / AddGenieApp:
    services.AddScoped<IAssistantContextContributor, StockSummaryContextContributor>();
  • Custom / replacement tools — IAssistantTool. services.AddScoped<IAssistantTool, MyTool>() adds a tool to the MCP loop (its Name/Description/ParameterSchema are injected into the prompt automatically). Tool names are case-insensitive and the last registration wins, so a host tool named execute_query replaces the built-in.

  • Replacing the provider — IAssistantProvider. Register your own implementation after AddGenieAssistant and single-service resolution takes the last registration: services.AddScoped<IAssistantProvider, MySemanticKernelProvider>(); — the whole pipeline (orchestrator, titles, SQL generation) flows through it.

The Inventory sample ships a working contributor (StockSummaryContextContributor) plus server-configured welcome chips.

Mount the assistant by enabling the module. It renders app-wide over every route — an AI icon in the header (left of the notification bell) opening a themed bottom-right panel. An optional top-level assistant block overrides the panel content:

createGenieApp({
apiBase: "/api/v1",
modules: { assistant: true },
assistant: {
title: "Inventory Copilot", // header title
subtitle: "Online", // status line
welcome: "Hi! Ask me anything.", // empty-state greeting
suggestions: ["Low stock?"], // quick-start chips ([] renders none, beating server chips)
placeholder: "Ask Copilot…", // composer placeholder
},
});

welcome and suggestions fall back to the server’s Genie:Assistant:WelcomeMessage / Genie:Assistant:WelcomeChips (fetched once from GET api/assistant/config when the panel first opens), then to the built-in greeting — with no built-in chips.

The panel is multi-conversation: a conversations list (open / rename / delete / new, plus multi-select bulk delete) and the conversation view — user/assistant bubbles as sanitised Markdown (GFM tables, lists, code) via marked + DOMPurify, a live Thinking disclosure, a collapsible SQL block and result-table preview for data answers, suggestion chips, a Shift+Enter composer, and a stop control while a reply is in flight. It is accent-/theme-aware.

The UI talks to the backend over AssistantChatHub when connected and falls back to REST (api/assistant: GET /config for the welcome content, POST /query, thread CRUD GET/POST /chats, GET/PATCH/DELETE /chats/{id}, POST /chats/delete for bulk delete) otherwise — clarification pills and inline charts work on both transports. GenieAssistant and GenieAssistantButton are also exported for hosts that want to place or configure them manually (custom title, subtitle, suggestions, welcome). See Frontend configuration.

  • Persistence. Conversations live in the database (schema Genie, tables AssistantChat + AssistantChatMessage), owned per user (UserId + CompanyId). Non-admins only see their own threads. A thread is saved only once its first reply succeeds; a failed first turn is rolled back. Host apps must add the migration for the two tables (dotnet ef migrations add AddAssistantChat -c YourContext), and again after an engine upgrade that adds a column to them — the nullable AssistantChatMessage.ProviderName, which records the model that produced each reply so the UI can label it, arrived that way.
  • Titles. New threads show a truncated placeholder immediately. With AutoGenerateTitle on (default) the provider summarises the first message into a concise title after the first reply, pushed live to the list (ConversationRenamed); on failure the placeholder is kept silently. A manual rename (PATCH /chats/{id}) always overrides a generated title.
  • Error masking. Internal error detail (a missing/invalid key, a raw provider failure) is surfaced only to users holding the System role (AssistantErrors.ForUser). Everyone else sees a generic “The assistant is temporarily unavailable…” message. Generated SQL is likewise returned to admins only.