Workshop 05 · LLM stack & data retention
A shared record of the fifth workshop — the recommended two-tier Gemini stack and layered architecture, Swiss hosting and data residency, cost as the central risk, the data-retention and consent documentation, the cancer-diagnostics restriction, and a further prototype pass across admin, doctor and client. Firmer cost figures and the final model choice are refined next week.
The recommended stack. A two-tier Gemini setup: Gemini 3.1 for synthesis and the medical path, and Gemini Flash Lite for light / lifestyle questions and intent routing — Flash Lite is adequate for lifestyle at much lower token cost, while the medical path uses the stronger model. Built on Google Vertex AI (Agent Builder / Agent Engine) in the Zürich region, all-in-one and the cheapest of Google / Amazon / Azure, with a contract preventing training use of the data. The layers are largely business logic and software surface — consent / release workflows, audit trails, storage, the bot's memory — not just "the AI." To be finalised next week.
Cost is make-or-break. ~60,000 interactions / month ≈ $1,000–1,400 / month before optimisation; optimisation phase 1 is ~5–8 weeks; firmer figures land next week. Model-switching / migration cost is a key concern (avoiding lock-in): much can be stored and migration can be step-by-step. No ready-made medical chatbot exists — MedLM is discontinued, MedGemma is research-only; Vertex AI Search for Healthcare is the candidate. TOON (a JSON alternative, up to ~80% fewer tokens) and Claude "skills" are being evaluated for savings.
Legal & product. The informed-consent record is stored with timestamp and name, exportable as a PDF via a name / date-of-birth search — minimal but legally sufficient. Cancer / genetic-cancer diagnostics are doctor-only: the patient must not see them and cannot ask about them in the client bot. Confirmed: an 18+ age gate, mandatory height / weight in registration, the "Academy" section with uploadable onboarding videos, critical hints auto-surfacing the major DNA findings, and dictation kept in the prototype but removed later if Swiss-German dialect performance is poor. The prototype link goes live Monday; the next call is Thursday 16 July.
The main topic. The recommendation is a two-tier setup that routes each question to the right-sized model: the strong model carries synthesis and anything medical, the light model handles everyday lifestyle chat and decides where a question should go. This keeps quality where it matters while keeping the token bill down where it doesn't. Still to be refined with next week's update.
Two-tier Gemini recommended. Gemini 3.1 for synthesis / medical, Flash Lite for light questions + routing — to be finalised next week.
Right-size by intent. Routing sends only the questions that need it to the expensive model, protecting both quality and cost.
Chris clarified that the layers are largely the business logic and software surface — consent / release workflows, audit trails, data storage, the bot's summaries and memory — not just the model. The client's stance: the "how" is the agency's domain; they mainly need a working interface and to decide costs jointly.
L1 — model & hosting
L3 — orchestration
L4 — prompts, audit & learning
L5 — healthcare building blocks
Models evaluated: Gemini, Claude, GPT, Llama 4, DeepSeek, Qwen. Open-source models can be self-hosted (own server or cloud) but carry higher operating cost and integration effort — not a native click-install. Orchestration effort: ~1–3 weeks via Google vs. ~2–4 weeks self-hosted. Crucially, no ready-made medical chatbot was found.
Managed vs. self-hosted
No ready-made medical bot
Gemini is Google, which has a Zürich server location; the cloud can be set up in Zürich and routed to keep data in Switzerland, and a contract with Google can prevent the data being used for training. The client's priority is legal certainty toward the authorities — being able to state, in good conscience, that hosting is in Switzerland. Google is recommended: all-in-one and the cheapest of Google / Amazon / Azure — always with the cost caveat.
Vertex AI · Zürich region. Cloud set up in Zürich and routed to keep patient data in Switzerland.
No-training contract. A contract with Google prevents the data being used to train models.
Why it matters: the priority is legal certainty toward authorities — SwissDnaCode must be able to state in good conscience that hosting is in Switzerland. Google is recommended as all-in-one and cheapest of the three big clouds, subject to the cost review.
Cost is repeatedly stressed as make-or-break: if the model cost blows up the cost structure, the product can't be priced or sold — especially with the DNA-analysis cost on top. Model-switching cost is the other key concern: how hard it is to migrate the accumulated knowledge from Gemini to open source later.
~$1,000–1,400 / month before optimisation at ~60k interactions / month. Optimisation phase 1 ≈ 5–8 weeks; firmer figures next week.
Avoid lock-in. Much of the knowledge can be stored; migration could be step-by-step. The doctor / client might even choose between models (Gemini / Claude / DeepSeek) — though a different model may answer differently.
Cost drives everything downstream
The DNA-analysis cost sits on top of the model cost, so the model bill cannot be allowed to blow up the price structure — otherwise the product can't be priced or sold. Next week should provide a first basis to discuss the calculation and budget.
Migration to be clarified internally — both the technical view (what can be stored and moved) and the business / data view (whether letting users pick a model is worth the inconsistency it introduces).
A 10-year retention obligation for medical data was raised; whether it applies is debated — it is the customer's data, but the practice has access and stores findings, and Chris sees it applying to the findings. To be clarified. The consent record must be legally protective without being heavy: keep it minimal but sufficient.
Consent record with timestamp + name. The online / informed-consent record is stored with a timestamp and name for legal protection; the audit trail captures this.
Export on demand. A report / PDF can be generated in any format; a search by name / date-of-birth can export everything held on a patient.
DNA-sampling consent stays personal. Informed consent for DNA sampling is done via a personal video call — it cannot be replaced by a pre-recorded video; a clip is recorded.
Manual upload first. Storing full videos on the existing practice system isn't feasible; the first path is a manual upload. A Zoom-API integration is Vision.
Nadine can upload cancer / genetic-cancer diagnostics as a PDF in the doctor segment, but the patient must not see it and must not be able to ask about it in the patient bot. This matches the previously agreed restriction and is considered legally important.
Data-visibility rule: this is not just an answer guardrail — cancer / pathology documents never cross into the client's view or the client bot's retrievable context. Confirmed and legally important.
Training the bot with validated authors / books: prefer condensed summary PDFs (e.g. ~30 pages instead of 500, generated via other AIs) to reduce tokens; some books could be uploaded to the knowledge base. The first online model should already be good enough — broad external video / social scraping for training is Vision.
What goes in
Rollout options
Two levers are being evaluated to reduce token cost on suitable use cases — directly relevant given cost is the central risk.
Claude "skills"
TOON format
Sensitive data falls under GUMG / GUMV. Medical-device classification: if the system supports therapist / doctor decisions it tends to fall into a higher regulated class, and under the EU AI Act it would be high-risk AI; Switzerland's stance is less clear. Currently only the two founders use it as doctors — no resellers / therapists yet.
Frameworks in play
Roles & domicile
Nadine is finalising with the lab partner which genes / SNPs to include in the panel; the raw data is analysed and may differ from the booklet. For Switzerland this is a different product than the Swiss partner offers. The result can be delivered as a PDF or another format; the team can convert any data format via the interface.
Panel & format
Nadine will get a sample of what the partner sends and discuss the best format with Chris. Any format is convertible via the interface, so the panel decision drives the content, not the plumbing.
Another pass across the three surfaces. Admin gains the Academy; the doctor file gets critical hints and a same-page bot; the client registration is reordered around mandatory fields with an age gate and a collapsible optional section.
Admin
Doctor
Client
Scope the critical hints: Nadine wants them to auto-surface the major findings from the DNA report (e.g. strong abnormalities, allergies) — but the scope must be defined so it doesn't dump everything. Test by uploading a report during LLM testing.
Age gate & medical mandatory questions
Age gate: 18+ via date of birth plus explicit confirmation of being at least 18. Mandatory Stammdaten: name, date of birth, email, phone, address, language, height, weight.
Medical mandatory questions (main goal, biggest current health topic, onset of complaints, known diagnoses, operations / hospitalisations, medications / hormones, supplements, allergies, family history, diet, sleep & stress) are structured with free-text and multi-select; consent is shown in registration and pre-activated where applicable.
Both now lean against dictation because of dialect — Swiss High German is near-dialect. Writing makes users more deliberate and concise and reduces cost; power users can paste text prepared elsewhere. Observations: ChatGPT dictates best; Claude's dictation is weaker.
Writing preferred. More deliberate, more concise, and cheaper than long dictated input in dialect.
Keep it in the prototype. Remove it from V1 if dialect performance proves poor in practice.
The ETH-affiliated, IT-savvy customer will stress-test the client view only (not doctor / admin), and only once the real version starts — not the first prototype variant — possibly joining a call for IT-to-IT feedback. The founders test the prototype (link live from Monday) and may start familiarising themselves with the two Google LLMs.
Founders test from Monday. The prototype link goes live Monday; both founders click through and think it through. Chris sends the exact access details (likely a special URL / registration).
ETH customer — client view only. Stress-tests only the client view, only once the real version starts, possibly on an IT-to-IT call.
Explore the two Google LLMs. Founders may begin familiarising themselves with Gemini 3.1 and Gemini Flash Lite once access details arrive.
The decisions taken or confirmed in this session — several remain subject to next week's cost review.
Two-tier Gemini (3.1 synthesis / medical + Flash Lite light / routing), to be finalised next week.
Google (Vertex / Agent Builder), Zürich, all-in-one — subject to the cost review — with a no-training contract.
Doctor bot full access (incl. cancer / pathology); patient bot blocks pathology / cancer and refers to Dr. Farkas.
18+ age gate (date of birth + explicit confirmation).
Registration: mandatory fields first; optional questions behind an expandable button; birth time optional; height / weight mandatory.
"Academy" kept; onboarding videos included in the MVP.
Doctor file: critical hints full-width top (auto-populated, scope TBD), notes add-at-top, bot quick access + single-page expansion.
Consent record via the audit trail with timestamp / name, exportable as a PDF via name / date-of-birth search — minimal but legally sufficient.
Dictation kept in the prototype, removed later if dialect performance is poor.
Zoom-like video consent = Vision; first path is a manual upload. The ETH customer tests only the client view, once the real version starts.
Monday is an internal Mendelio day: finalise the LLM comparison, refine the prototype, and fix the FTP upload so the link goes live. The LLM analysis is updated next week with firmer costs to reach a first budget basis.
SwissDnaCode — to-dos
Mendelio — to-dos
Open points — still to decide
The prototype link goes live Monday; the founders test it and think it through. Chris runs the LLM comparison internally on Monday, refines the prototype, and updates the analysis with firmer costs next week — a first basis to discuss the calculation and budget.
Monday · internal + go-live. Chris finalises the LLM comparison with the team, refines the prototype, fixes the FTP upload so the link goes live, and sends the exact Gemini access details.
Founders test. Both founders test the prototype and may start familiarising themselves with the two Google LLMs.
Next week · firmer costs. The LLM analysis is updated with more precise figures to reach a first budget / calculation basis.
Next call · Thursday 16 July, 08:00. Moved from the usual slot (Nadine has a Friday clinic day in Adorf and family visiting Saturday). Chris sends the invitation with the link.