Workshop 05 · LLM stack & data retention

SwissDnaCode — LLM stack summary

A shared record of the fifth workshop — the recommended two-tier Gemini stack and layered architecture, Swiss hosting and data residency, cost as the central risk, the data-retention and consent documentation, the cancer-diagnostics restriction, and a further prototype pass across admin, doctor and client. Firmer cost figures and the final model choice are refined next week.

SwissDnaCode · Mendelio LLM stack · data retention · prototype pass Completed

Executive summary

The recommended stack. A two-tier Gemini setup: Gemini 3.1 for synthesis and the medical path, and Gemini Flash Lite for light / lifestyle questions and intent routing — Flash Lite is adequate for lifestyle at much lower token cost, while the medical path uses the stronger model. Built on Google Vertex AI (Agent Builder / Agent Engine) in the Zürich region, all-in-one and the cheapest of Google / Amazon / Azure, with a contract preventing training use of the data. The layers are largely business logic and software surface — consent / release workflows, audit trails, storage, the bot's memory — not just "the AI." To be finalised next week.

Cost is make-or-break. ~60,000 interactions / month ≈ $1,000–1,400 / month before optimisation; optimisation phase 1 is ~5–8 weeks; firmer figures land next week. Model-switching / migration cost is a key concern (avoiding lock-in): much can be stored and migration can be step-by-step. No ready-made medical chatbot exists — MedLM is discontinued, MedGemma is research-only; Vertex AI Search for Healthcare is the candidate. TOON (a JSON alternative, up to ~80% fewer tokens) and Claude "skills" are being evaluated for savings.

Legal & product. The informed-consent record is stored with timestamp and name, exportable as a PDF via a name / date-of-birth search — minimal but legally sufficient. Cancer / genetic-cancer diagnostics are doctor-only: the patient must not see them and cannot ask about them in the client bot. Confirmed: an 18+ age gate, mandatory height / weight in registration, the "Academy" section with uploadable onboarding videos, critical hints auto-surfacing the major DNA findings, and dictation kept in the prototype but removed later if Swiss-German dialect performance is poor. The prototype link goes live Monday; the next call is Thursday 16 July.

Recommended LLM stack — two-tier Gemini

The main topic. The recommendation is a two-tier setup that routes each question to the right-sized model: the strong model carries synthesis and anything medical, the light model handles everyday lifestyle chat and decides where a question should go. This keeps quality where it matters while keeping the token bill down where it doesn't. Still to be refined with next week's update.

Gemini 3.1 · synthesis
Medical path & synthesis. The stronger model handles the doctor path and any answer that combines DNA, findings and history into guidance — where accuracy and grounding matter most.
Flash Lite · routing
Lifestyle & intent routing. The light model answers everyday lifestyle questions and classifies intent (simple / complex / medical) so only the questions that need the strong model reach it — currently adequate at much lower token cost.

Two-tier Gemini recommended. Gemini 3.1 for synthesis / medical, Flash Lite for light questions + routing — to be finalised next week.

Right-size by intent. Routing sends only the questions that need it to the expensive model, protecting both quality and cost.

Final model choice refined next week with more precise cost figures

Layered architecture — mostly software, not "the AI"

Chris clarified that the layers are largely the business logic and software surface — consent / release workflows, audit trails, data storage, the bot's summaries and memory — not just the model. The client's stance: the "how" is the agency's domain; they mainly need a working interface and to decide costs jointly.

L1 — model & hosting

  • The LLM in the background
  • Hosting & data access

L3 — orchestration

  • Routing of questions
  • Retrieval of patient data
  • Release / consent workflows — fastest via Vertex Agent Builder / Agent Engine

L4 — prompts, audit & learning

  • Prompt management
  • Revision-proof audit trails
  • Continuous learning

L5 — healthcare building blocks

  • Medically-tuned reusable components
  • Layered on top of the general model

Model landscape & medical building blocks

Models evaluated: Gemini, Claude, GPT, Llama 4, DeepSeek, Qwen. Open-source models can be self-hosted (own server or cloud) but carry higher operating cost and integration effort — not a native click-install. Orchestration effort: ~1–3 weeks via Google vs. ~2–4 weeks self-hosted. Crucially, no ready-made medical chatbot was found.

Managed vs. self-hosted

  • Evaluated: Gemini, Claude, GPT, Llama 4, DeepSeek, Qwen
  • Open-source is self-hostable but higher opex + integration effort
  • Orchestration ~1–3 wks on Google vs. ~2–4 wks self-hosted

No ready-made medical bot

  • MedLM — discontinued
  • MedGemma — research only (free building block for CT / MR / EHR understanding)
  • Vertex AI Search for Healthcare — the candidate
Recommendation may narrow next week — the conversational engine is a general Gemini model + our own orchestration

Swiss hosting & data residency

Gemini is Google, which has a Zürich server location; the cloud can be set up in Zürich and routed to keep data in Switzerland, and a contract with Google can prevent the data being used for training. The client's priority is legal certainty toward the authorities — being able to state, in good conscience, that hosting is in Switzerland. Google is recommended: all-in-one and the cheapest of Google / Amazon / Azure — always with the cost caveat.

Vertex AI · Zürich region. Cloud set up in Zürich and routed to keep patient data in Switzerland.

No-training contract. A contract with Google prevents the data being used to train models.

Why it matters: the priority is legal certainty toward authorities — SwissDnaCode must be able to state in good conscience that hosting is in Switzerland. Google is recommended as all-in-one and cheapest of the three big clouds, subject to the cost review.

Cost — the central risk

Cost is repeatedly stressed as make-or-break: if the model cost blows up the cost structure, the product can't be priced or sold — especially with the DNA-analysis cost on top. Model-switching cost is the other key concern: how hard it is to migrate the accumulated knowledge from Gemini to open source later.

~$1,000–1,400 / month before optimisation at ~60k interactions / month. Optimisation phase 1 ≈ 5–8 weeks; firmer figures next week.

Avoid lock-in. Much of the knowledge can be stored; migration could be step-by-step. The doctor / client might even choose between models (Gemini / Claude / DeepSeek) — though a different model may answer differently.

Cost drives everything downstream

The DNA-analysis cost sits on top of the model cost, so the model bill cannot be allowed to blow up the price structure — otherwise the product can't be priced or sold. Next week should provide a first basis to discuss the calculation and budget.

Migration to be clarified internally — both the technical view (what can be stored and moved) and the business / data view (whether letting users pick a model is worth the inconsistency it introduces).

Data retention, consent & documentation

A 10-year retention obligation for medical data was raised; whether it applies is debated — it is the customer's data, but the practice has access and stores findings, and Chris sees it applying to the findings. To be clarified. The consent record must be legally protective without being heavy: keep it minimal but sufficient.

Consent record with timestamp + name. The online / informed-consent record is stored with a timestamp and name for legal protection; the audit trail captures this.

Export on demand. A report / PDF can be generated in any format; a search by name / date-of-birth can export everything held on a patient.

DNA-sampling consent stays personal. Informed consent for DNA sampling is done via a personal video call — it cannot be replaced by a pre-recorded video; a clip is recorded.

Manual upload first. Storing full videos on the existing practice system isn't feasible; the first path is a manual upload. A Zoom-API integration is Vision.

Clarify whether the 10-year retention obligation applies and how the consent record is stored on the practice side

Cancer / genetic-diagnostics restriction

Nadine can upload cancer / genetic-cancer diagnostics as a PDF in the doctor segment, but the patient must not see it and must not be able to ask about it in the patient bot. This matches the previously agreed restriction and is considered legally important.

Doctor
Doctor bot — full access. The doctor bot always has full access, including pathologies. Cancer / genetic-cancer findings are uploaded as a PDF in the doctor segment.
Patient
Patient bot — blocked. The patient bot blocks all pathology / cancer questions and refers to Dr. Farkas. The document is never visible to the patient nor retrievable by the patient bot.

Data-visibility rule: this is not just an answer guardrail — cancer / pathology documents never cross into the client's view or the client bot's retrievable context. Confirmed and legally important.

Knowledge-base training & rollout

Training the bot with validated authors / books: prefer condensed summary PDFs (e.g. ~30 pages instead of 500, generated via other AIs) to reduce tokens; some books could be uploaded to the knowledge base. The first online model should already be good enough — broad external video / social scraping for training is Vision.

What goes in

  • Condensed summary PDFs (~30 pp vs. 500) to cut tokens
  • Validated authors / books uploaded to the knowledge base
  • First online model already "good enough"; scraping is Vision

Rollout options

  • (a) Log into the Google interface to train — structured, clustered by topic, compared against the previous version
  • (b) A section in the software where a version is approved and "rolled out to all doctors / patients"
The knowledge-push interface and version-rollout mechanism to be discussed with the programmer

Token-optimisation techniques

Two levers are being evaluated to reduce token cost on suitable use cases — directly relevant given cost is the central risk.

Claude "skills"

  • A memory network of already-asked items
  • The agency's staff use skills to cut costs

TOON format

  • An alternative to JSON
  • Can save up to ~80% of tokens
  • Being evaluated for suitable use cases

Regulatory & company location

Sensitive data falls under GUMG / GUMV. Medical-device classification: if the system supports therapist / doctor decisions it tends to fall into a higher regulated class, and under the EU AI Act it would be high-risk AI; Switzerland's stance is less clear. Currently only the two founders use it as doctors — no resellers / therapists yet.

Frameworks in play

  • GUMG / GUMV — Swiss genetic-testing law for sensitive data
  • Decision-supporting software → higher regulated class (EU MDR)
  • EU AI Act → high-risk AI; Switzerland's stance less clear

Roles & domicile

  • Only the two founders act as doctors today
  • Therapist resellers would not receive the doctor bot — anything medical refers to Dr. Farkas
  • Company domicile outside EU / CH brings different regulation — still to be addressed
Clarify medical-device / high-risk-AI classification, GUMG/GUMV implications and company-location topics

DNA data & SNP selection

Nadine is finalising with the lab partner which genes / SNPs to include in the panel; the raw data is analysed and may differ from the booklet. For Switzerland this is a different product than the Swiss partner offers. The result can be delivered as a PDF or another format; the team can convert any data format via the interface.

Panel & format

Nadine will get a sample of what the partner sends and discuss the best format with Chris. Any format is convertible via the interface, so the panel decision drives the content, not the plumbing.

Nadine: finalise SNP / gene selection with the lab partner · obtain a sample report · agree the format with Chris

Prototype walkthrough — updates

Another pass across the three surfaces. Admin gains the Academy; the doctor file gets critical hints and a same-page bot; the client registration is reordered around mandatory fields with an age gate and a collapsible optional section.

Admin

  • User roles
  • "Academy" (term kept) — add / categorise videos
  • Uploadable onboarding videos, cloud-hosted

Doctor

  • Doctor-bot quick access (visual placeholder ok)
  • Critical hints full-width at the top — auto-surface major DNA findings (scope TBD)
  • Notes with "add" at the top; findings, anamnesis, SNPs, medications
  • Bot hides other panels and expands on the same page; can go full size

Client

  • Mandatory fields first; height / weight mandatory (DNA needs them)
  • Birth time optional, kept under date of birth
  • Optional questions collapse behind a button with a persistent hint
  • Works on mobile / tablet / desktop (desktop like the doctor view)

Scope the critical hints: Nadine wants them to auto-surface the major findings from the DNA report (e.g. strong abnormalities, allergies) — but the scope must be defined so it doesn't dump everything. Test by uploading a report during LLM testing.

Age gate & medical mandatory questions

Age gate: 18+ via date of birth plus explicit confirmation of being at least 18. Mandatory Stammdaten: name, date of birth, email, phone, address, language, height, weight.

Medical mandatory questions (main goal, biggest current health topic, onset of complaints, known diagnoses, operations / hospitalisations, medications / hormones, supplements, allergies, family history, diet, sleep & stress) are structured with free-text and multi-select; consent is shown in registration and pre-activated where applicable.

Dictation / voice input

Both now lean against dictation because of dialect — Swiss High German is near-dialect. Writing makes users more deliberate and concise and reduces cost; power users can paste text prepared elsewhere. Observations: ChatGPT dictates best; Claude's dictation is weaker.

Writing preferred. More deliberate, more concise, and cheaper than long dictated input in dialect.

Keep it in the prototype. Remove it from V1 if dialect performance proves poor in practice.

Testing plan

The ETH-affiliated, IT-savvy customer will stress-test the client view only (not doctor / admin), and only once the real version starts — not the first prototype variant — possibly joining a call for IT-to-IT feedback. The founders test the prototype (link live from Monday) and may start familiarising themselves with the two Google LLMs.

Founders test from Monday. The prototype link goes live Monday; both founders click through and think it through. Chris sends the exact access details (likely a special URL / registration).

ETH customer — client view only. Stress-tests only the client view, only once the real version starts, possibly on an IT-to-IT call.

Explore the two Google LLMs. Founders may begin familiarising themselves with Gemini 3.1 and Gemini Flash Lite once access details arrive.

Decisions at a glance

The decisions taken or confirmed in this session — several remain subject to next week's cost review.

Two-tier Gemini (3.1 synthesis / medical + Flash Lite light / routing), to be finalised next week.

Google (Vertex / Agent Builder), Zürich, all-in-one — subject to the cost review — with a no-training contract.

Doctor bot full access (incl. cancer / pathology); patient bot blocks pathology / cancer and refers to Dr. Farkas.

18+ age gate (date of birth + explicit confirmation).

Registration: mandatory fields first; optional questions behind an expandable button; birth time optional; height / weight mandatory.

"Academy" kept; onboarding videos included in the MVP.

Doctor file: critical hints full-width top (auto-populated, scope TBD), notes add-at-top, bot quick access + single-page expansion.

Consent record via the audit trail with timestamp / name, exportable as a PDF via name / date-of-birth search — minimal but legally sufficient.

Dictation kept in the prototype, removed later if dialect performance is poor.

Zoom-like video consent = Vision; first path is a manual upload. The ETH customer tests only the client view, once the real version starts.

Deliverables, to-dos & open points

Monday is an internal Mendelio day: finalise the LLM comparison, refine the prototype, and fix the FTP upload so the link goes live. The LLM analysis is updated next week with firmer costs to reach a first budget basis.

Workshop 5 summary — decisions & open points (this document)
Recommended two-tier Gemini stack + layered architecture (presented)
Prototype link live Monday — FTP upload fixed
LLM analysis updated next week — more precise cost figures + first budget basis
Critical-hints auto-population, age gate, consent documentation / PDF export, optional-question grouping
Evaluate the TOON format and skills for token savings

SwissDnaCode — to-dos

  • Nadine: finalise the SNP / gene selection with the lab partner; get a sample report and agree the format with Chris
  • Nadine: define what belongs in the "critical hints" (which DNA findings auto-surface)
  • Both: test the prototype (link live from Monday) and give feedback
  • Both: optionally begin exploring Gemini 3.1 and Flash Lite (access details Monday)
  • Clarify the 10-year retention obligation and how the consent record is stored on the practice side
  • Later: prepare condensed knowledge-base summaries (validated books / authors)

Mendelio — to-dos

  • Monday internal: finalise the LLM comparison, refine the prototype, fix the FTP upload
  • Update the LLM analysis next week with more precise costs → first budget / calculation basis
  • Send the exact Gemini access details on Monday
  • Clarify internally the model-switch / migration path (technical + business / data)
  • Implement critical-hints auto-population, the age gate, consent documentation / PDF export, optional-question grouping
  • Evaluate the TOON format and skills for token savings
  • Clarify medical-device / high-risk-AI classification, GUMG/GUMV implications and company-location topics

Next steps & process

The prototype link goes live Monday; the founders test it and think it through. Chris runs the LLM comparison internally on Monday, refines the prototype, and updates the analysis with firmer costs next week — a first basis to discuss the calculation and budget.

Monday · internal + go-live. Chris finalises the LLM comparison with the team, refines the prototype, fixes the FTP upload so the link goes live, and sends the exact Gemini access details.

Founders test. Both founders test the prototype and may start familiarising themselves with the two Google LLMs.

Next week · firmer costs. The LLM analysis is updated with more precise figures to reach a first budget / calculation basis.

Next call · Thursday 16 July, 08:00. Moved from the usual slot (Nadine has a Friday clinic day in Adorf and family visiting Saturday). Chris sends the invitation with the link.

Abbreviations

AI ActEU Artificial Intelligence Act
APIApplication Programming Interface
CHFSwiss Franc (currency)
EHRElectronic Health Record
FTPFile Transfer Protocol (used to publish the prototype)
GUMGSwiss Federal Act on Human Genetic Testing
GUMVOrdinance to the Swiss Human Genetic Testing Act
LLMLarge Language Model
MDREU Medical Device Regulation
MVPMinimum Viable Product (V1)
RAGRetrieval-Augmented Generation
SNPSingle Nucleotide Polymorphism (a DNA variant)
TBDTo Be Decided
TOONToken-Oriented Object Notation (a compact JSON alternative)