Skip to content

AI Assistant overview

The Genie Assistant is a natural-language chat that answers questions about your application’s data and how to use the app. A user asks in plain language (“how many tickets are open this week?”); the assistant generates SQL, runs it against the database, and replies with the answer as sanitised Markdown — figures, bullet lists, or a result table.

It is named Assistant everywhere in code (Genie.Engine/Features/Assistant). Under the hood it is a provider-agnostic, MCP-style tool orchestrator streamed over SignalR and scoped to the caller’s RBAC permissions — so it can only see and describe what the user is already allowed to see.

  • Natural language → SQL → answer. The model turns a question into a read-only SELECT, the engine validates and runs it, and the model phrases the rows back in business language (never leaking table or column names).
  • App how-to questions. Beyond data, it answers “how do I add/edit X in the app?” from Markdown help docs (baseline engine docs + host docs).
  • “What can I access?” It is given the caller’s roles and full permission map, so it can explain what the user may do and avoid suggesting actions they aren’t permitted to perform.
  • Charts, in the conversation. Ask it to show, plot or compare something and the answer comes with a chart under it (see below).
  • Report authoring. Open the report designer and it can read the report you are editing, preview what a dataset returns, and propose changes you apply with one click — see Assistant-assisted authoring.

Most replies are Markdown. Some carry a block rendered underneath the prose:

  • A chart. Produced by the chart_result tool, which runs the SELECT itself rather than asking the model to re-emit the rows — cheaper, and it removes the most likely thing for a model to get wrong. It goes through the same safety chain as execute_query: SELECT-only validation, the row cap, and the tenant-scoped read-only executor. Charting is a way to display data, never a second way to query it. Rendering reuses the report page’s own chart component, so there is no charting dependency in the package and the colours follow your theme and accent. The chart is drawn at the panel’s real width — labels stay the size the stylesheet asks for however narrow it is — and a bar chart of long or numerous category names is turned on its side, where each name gets its own row.
  • An Apply card. A proposed change with Apply and Dismiss. Applying is entirely client-side: the card names a handler that the owning page registered, and that page writes to its own in-memory state. The server never applies anything. If the owning page isn’t open the card says so and stays put.

Because a chart needs room, the panel header has an expand toggle that widens it, and the panel’s top edge can be dragged to change its height (double-click the handle to reset).

The panel can also be moved: drag it by its title bar and it stays where you drop it, clamped so it can never end up off-screen. Once it has been moved, a reset-position button appears in the header to send it back to its default corner. Position, width and height are per-session choices — they are not persisted, so a reload starts from the default again.

Above the composer, a + button offers whatever the current page can contribute to the conversation. On the report designer that is Current report, on the wizard designer Current wizard, and on the workflow designer Current workflow (not while previewing a running instance); other pages offer nothing and the button is hidden.

Attaching is what tells the server which tools the turn may use, so the chip is not decoration — it is the switch. Nothing is attached automatically. Arriving on the report designer changes nothing about the conversation: until you attach the document, the turn gets the information tools and the assistant answers questions. Attach it and the authoring tools appear; remove the chip and they go again.

That is deliberate. Attaching is what turns the assistant from something you ask questions of into something that rewrites your document, and it also puts a multi-kilobyte document within the model’s reach — neither should follow from the route you happen to be on. It is one click, in the + menu, and the chip shows you it happened.

One context is attached at a time; picking from the menu replaces what is there, and navigating to a page that no longer offers it clears it rather than carrying it along.

Answering in the chat while a document is attached

Section titled “Answering in the chat while a document is attached”

The first way to get a chat answer is simply not to attach the report — see above. But while you are working on one you will want it attached, and a question will still come up mid-edit. So: add “reply in chat” to the message — or tell it not to change the report, wizard or workflow — and that turn answers in the conversation and proposes nothing, attachment or no attachment: no Apply card, no XML.

This is enforced, not requested. When the message carries that instruction the change-proposing tool is not offered to the model at all, so there is nothing for it to call however the rest of the prompt reads. The reading tools stay, because the question may well be about the open document (“summarise this report, just reply in chat”), and so does chart_result — a chart is drawn in the conversation and changes nothing, so “plot the last 24 hours, reply in chat” is answered with a picture and a sentence.

Phrasings that switch it on:

You write What happens
“…, reply in chat” / “answer in the chat” / “chat only” Answered in the conversation; nothing proposed
“don’t change the report, just give me the figures” Same
“show me the trend without changing the report” Same
“don’t change the report title, add a chart instead” Ordinary authoring — this asks to change something specific

That last row is the line the rule is drawn on: the instruction has to end the clause. A request to change one particular thing reads, word for word, like a request to change nothing, so only where the phrase stops decides which it is. And whichever way it goes, the failure is cheap in one direction only — an answer in the chat when you wanted an edit costs one message; an edit you asked nobody to make costs an undo.

The assistant may ask a follow-up question when a request is genuinely ambiguous — but only one, and answering it ends the matter: request_clarification is withheld from the turn that carries your answer, so the next thing that happens is a query.

That is enforced rather than asked for, because the rule as a prompt line had a hole in it. “At most one clarifying question per user message” is true of every message individually, and an answer to a clarification is a new message — so a model could narrow indefinitely and still be following the rule. It did: one request to plot an availability trend was met with which series, then which technology set, then how many sites, without a row being read, and the turn that finally had everything it needed ran out of time.

Anything still open after one question is assumed rather than asked about, and the answer says what was assumed. Scope, row counts and the choice between near-identical metrics are explicitly named in the prompt as things to decide rather than ask about.

Rather than a single prompt round-trip, the assistant runs a multi-turn tool-calling loop (McpOrchestratorService). The provider protocol is embedded in the system prompt: to call a tool the model replies with exactly

TOOL_CALL: <tool_name>
PARAMS: {"key": "value"}

A reply may carry up to three of these when the calls do not depend on each other — they run in the order written and the results come back together, which is what a model expects when it batches two lookups. Anything past three is not run and the model is told which ones, by name.

The orchestrator parses those directives, dispatches each named tool, appends the results as tool messages, and loops — until the model returns a plain-text answer, requests clarification, or the iteration cap (5) is reached. This keeps the design provider-agnostic: any provider that can emit text can drive the tools; no provider-specific function-calling API is required.

┌─ McpOrchestratorService · loop, max 5 turns ──────┐
│ system prompt: schema + app-docs + RBAC + rules │
│ │
┌──────────┐ │ ┌─────────┐ ── TOOL_CALL ▶ ┌────────┐ │ ┌──────────────┐
│ User │──┼─▶ │ model │ │ tool │ ├─▶ │ Markdown │
│ question │ │ │ │ ◀─── result ── │ │ │ │ answer │
└──────────┘ │ └─────────┘ └────────┘ │ └──────────────┘
└───────────────────────────────────────────────────┘
Prefer a picture?

Each tool implements IAssistantTool (name, description, parameter schema, ExecuteAsync):

Tool What it does
execute_query (ExecuteQueryTool) Validates a generated SELECT (SELECT-only, no dangerous keywords, row-capped) and runs it through IAssistantQueryExecutor, which enforces tenant scope at the DB layer. Returns rows as JSON.
get_table_schema (GetTableSchemaTool) Returns the column definitions for a specific table when they weren’t already in the compact schema context.
get_app_doc (GetAppDocTool) Returns the full Markdown body of an application help document by key (from the app-docs index in the system prompt) — for “how do I use X” questions.
request_clarification (RequestClarificationTool) Asks the user one follow-up question when the request is ambiguous, with 2–4 suggested answers the UI renders as one-tap pills (tap to answer) plus an “Other…” pill that focuses the composer for a free-text reply. Only the latest clarification is answerable — history renders the pills inert. Returns a sentinel the orchestrator relays back to the user. Not offered on the turn that answers one: see below.

The assistant never sees more than the user does. Two layers enforce that:

  • Permission-filtered schema context. The schema library the model receives is built from Genie’s modeled metadata (EntityStore), not the raw DB catalog, and is filtered to tables the caller can reach through a view they hold View/List on. Engine and identity internals are never modeled as domain entities, so they are never described to the model. If the caller can access no tables — or asks about data outside their access — the model is instructed to answer plainly that they don’t have permission, rather than guess table names.
  • DB-enforced tenant isolation. For real isolation the generated SQL runs through a dedicated read-only, least-privilege DB login, and the caller’s CompanyId/UserId/IsSystem is pushed into a DB session context so Row-Level Security filters rows regardless of what SQL the model wrote. See Tenant isolation & SQL hardening.

This mirrors the engine-wide rule: prompt text is a hint, not an authorization boundary — the database re-enforces access. See the Security model.

The chat runs over a dedicated SignalR hub, AssistantChatHub (/hubs/assistant-chat), separate from the notification hub. As the model works, the hub streams intermediate output to the caller:

  • ReceiveReasoningChunk — chain-of-thought reasoning streamed live while the model thinks (shown in an expandable Thinking disclosure that collapses once the answer lands).
  • ReceiveMessage — the final assistant reply (Markdown, plus generated SQL for admins and a result table for data answers).
  • ConversationStarted / ConversationRenamed — a brand-new thread is announced only once its first reply exists, then its AI-generated title is pushed live.
  • ReceiveError / MessageCancelled — role-masked errors and cancellation of an in-flight reply (CancelMessage).

If the hub is unavailable the UI falls back to REST (POST /api/assistant/query). The fallback carries the reply’s rich blocks (charts) just as ReceiveMessage does, and it is gated by the same AssistantChat policy — so it is a fallback for a dropped connection, never a way around ChatWidgetEnabled.

Conversations are persisted in the database (schema Genie, tables AssistantChat and AssistantChatMessage) — not Redis — so threads survive restarts and each user can keep multiple chats. Key behaviours:

  • Each thread is owned by its user (UserId + CompanyId); non-admins only ever see their own.
  • A thread is persisted only once its first reply succeeds — a failed first turn is rolled back, so there are no empty ghost threads.
  • Titles start as the first message truncated, then (when AutoGenerateTitle is on) the provider summarises the first message into a concise title that updates live; a manual rename always wins.
  • A reply belongs to the thread that asked for it. Leaving a conversation while it is still being answered — to the thread list, to another chat, or to a new one — no longer carries the “thinking” state along with you, and the answer lands in its own thread rather than whichever one is on screen when it arrives. The thread list marks the conversation still being answered, and opening it shows the pending question with the reply still coming.
  • One reply at a time. The hub runs one turn per connection, so while any thread is being answered the composer in the others will not send: doing so would cancel the running reply server-side.
  • Each reply names the model that wrote it (beside its timestamp), stored per message. The model picker stays live for the life of a thread, so a transcript can hold answers from more than one model and each is labelled with its own.

Every request injects four things into the system prompt, assembled by the orchestrator:

  1. A schema library from EntityStore (dual-database safe, modeled types/enums/FK targets, permission-filtered as above; falls back to a raw catalog read only when no entities are modeled).
  2. An app-documentation index — a compact list of Markdown help docs; the model pulls a full doc on demand via get_app_doc. The engine ships a baseline set; a host adds its own via Genie:Assistant:DocsPath (host wins on collision).
  3. Host-contributed business context — any IAssistantContextContributor implementations the caller project registers in DI render as titled sections (live business signals, KPIs, domain glossaries…). See Extending the assistant.
  4. The caller’s roles and full permission map (from IPermissionManager) plus mandatory security rules (SELECT-only, soft-delete filtering, and — for non-admins — a hard CompanyId filter requirement).