Twenty-one building blocks in six chapters — from the personnel file through the rule set for tools to a virtual developer working on your source code. Together they make a workforce you can actually manage.
How a job description becomes a colleague — and how that colleague takes on work, plans it and hands it out.
Every AI colleague has a complete personnel file: role description and mandate, persona and tone, thinking tools (e.g. "pre-mortem before every recommendation"), explicit no-gos (no invented facts, no filler phrases), the right model for the role — capable for complex work, cheap for routine — bound skills and knowledge access, budget and interaction limits. Plus a working mode: a colleague handling routine mail works in short turns; one building a whole package may stay on a single run for up to 90 minutes — the platform condenses the state of work itself and files it in the case, instead of losing the thread. Every change is versioned: one click and the employee works exactly as in any earlier version. Specialists beat generalists — the platform steers you toward tightly scoped roles with few, fitting skills.
An AI employee is created like a real hire: role, personality, working style and limits live in a personnel file — not in a prompt box.
"I'd like to hire someone for invoice review", "Why isn't Sabine responding?", "Extend the sick-leave process" — Helga understands the request, asks the right follow-up questions and executes. For multi-step work she presents a plan you approve once; every change appears as an approval card with an explanation anyone can read. She diagnoses problems, reminds you of open items, navigates you to the right place in the studio and knows the full product documentation. She also reads what actually happened inside a case and fetches the files stored there — questions that would otherwise mean clicking through the interface get answered in conversation. You can attach files to her directly. Every conversation is a resumable session.
You say what you want in plain language — Helga sets it up, explains what she is doing, and never changes anything without your approval.
New colleagues start on probation: sandbox operation with tight feedback loops. The assessment centre tests adversarially — with deliberate attacks and traps — across five dimensions: safety (resists manipulation, leaks no internals), honesty (admits uncertainty instead of inventing), brand voice (stays professional, makes no unauthorised commitments), compliance (follows the company constitution) and process fidelity (plays real business processes through — safely, because effective actions run as simulation). An independent judge model scores each dimension with reasoning. That verdict is a recommendation to the person deciding on the hire — it does not block, and that is deliberate: someone who knows the role should be allowed to weigh a narrow result instead of being overruled by a number. Individual processes can be rehearsed at any time — before the first real case, not after.
Before an AI employee takes on responsibility, an adversarial test bench tells you whether it works safely, honestly and within the rules — the decision to keep it is yours.
For bigger jobs you assemble teams: a team lead breaks the mandate down along the job descriptions, delegates subtasks to specialised members, reviews their results and requests rework before delivering consolidated. The four-eyes principle is built in: whoever created something does not review it. Work runs in waves — what can run in parallel does; what needs another task's finished result waits. And the platform does not allow otherwise: assigning a task from a later wave is refused, with the reason and with a statement of what is due now. Questions travel upward instead of into the void: a member who is stuck or would have to guess asks their team lead, who answers or carries the question to a human. A case speaks to the outside with one voice. As the organisation grows we warn about excessive span of control — divisions and intermediate levels then help, and supervisor responsibility and process scope inherit cleanly through them.
AI colleagues organise into teams with a team lead: delegate, review, deliver consolidated — creator and reviewer are never the same instance.
When several tasks build on each other, when something irreversible is involved, or when the mandate allows more than one reading, the employee presents a plan first: the goal in one sentence, the waves with their tasks and the result expected from each, open questions — and explicitly the non-goals, where a quick read tells you whether the mandate was understood. This applies to every kind of work: a trade-fair campaign as much as a quote or a software change. Open questions come with clickable answer suggestions and a free-text field; answers given then live in the plan and thus in the context of every later run. You have three options: approve, return with a comment (not a rejection — the employee revises) or reject. A new version must state what changed and why, otherwise the platform does not accept it. Once approved the plan is binding, available to everyone involved, and the card shows progress. Per team you set when planning happens — at the employee's discretion, as soon as several people contribute, always, or never.
Larger jobs start with a plan on a card — and a plan is always decided by a human, never by an AI.
Who decides when something may run — and who gets asked when nobody has decided. Four mechanisms that together leave no gap.
Oversight is an org structure, not a checkbox. Through the hierarchy (divisions → teams → employees) AI colleagues inherit their responsible supervisor — with deputies for absences and a fallback all the way to the management board. Actions requiring approval create a request for exactly that supervisor, wherever they are: in the studio or as an interactive card in Microsoft Teams. If it sits unanswered it escalates one level up automatically — after 4, 8 or 24 hours depending on risk class. At the end of the chain the risk class decides: critical and normal actions count as declined, explicitly low-risk ones go through. The default is an assessment you know up front, not a surprise. An approval turns into a rule along the way: the card says why you are being asked, and below it suggestions drawn from the actual call — one click and the same case runs through in future. Trust grows on real cases instead of on a setting somebody made generous just in case. And an approval does not tear work apart: whoever is waiting on a decision mid-task continues exactly there afterwards.
Every AI employee has a human supervisor; risky actions wait for a human yes — with deputies, deadlines and safe defaults.
Whether a call runs through or is put in front of a human depends not only on the employee and the tool, but on the situation: who, in which channel, in which process, with which arguments. You write a sentence — "she may send mail to our own domains without asking" — the platform AI translates it, shows examples including edge cases, and computes every example against the real evaluation: if the rule behaves differently from its sentence, it cannot be made binding. Four properties make this safe rather than convenient. Rules only grant — there is no forbidding rule, and whatever no rule grants goes to approval; the system is closed when in doubt. The whole chain is checked bottom-up (employee → team → divisions → company), and the most specific rule is named as the reason. The place determines the chain — a relaxation for a team channel does not apply in a direct message, otherwise you could bypass it by writing directly. And new rules may observe first: an observing rule has no effect, it merely counts how often it would have saved an approval. Making a rule binding is always a human act.
Instead of "how often to ask" you answer "under which circumstances" — in one sentence of plain text that the platform turns into a testable rule.
Some things no AI should do: answer direct messages from people, comment publicly on behalf of the brand, write emails to customers, make HR decisions. These human zones are defined per employee and work on two levels: as a binding rule in the employee's core — and hard in the technology, wherever the platform itself builds the path outside: the Teams connector then simply refuses the DM reply, the mail system refuses external sending, in both directions. For connected third-party tools the same promise is carried by the rule set: whatever no rule explicitly allows does not run — it goes to a human. AI takes work off your hands; it does not replace relationships between people.
You decide where AI simply does not belong — secured twice: as a binding rule at the employee's core and as a hard stop in the paths that lead outside.
You set cost limits per employee, daily and monthly. At 80 % there's a warning, at 100 % the employee is stopped — never a surprise invoice. Billing is honest, based on complete usage data for every single model call. On top of that there is a limit per run: a colleague working on a large task has, besides the monthly budget, a time and cost measure for this one run; when it is reached the run ends with a statement of what it consumed and how far it got — not with a silent abort, and never with an invoice that grows overnight. Costs are visible per employee, team and division; what the platform's own AI (Helga, judge and helper models) consumes is disclosed transparently too. The weekly operations report names the cost drivers on its own.
Every employee has a daily and monthly budget with a hard stop — and every cent spent is traceable to a case.
Where the work arrives, how it stays visible, and what ends up in front of a person.
AI colleagues live in the platform's internal chat, in Microsoft Teams and in the email inbox. In Teams you reach them by @mention; approvals and questions appear as interactive cards right in the channel. Mailboxes are managed centrally (Microsoft 365 or classic IMAP): one page shows which address exists, who receives incoming mail and — separately — who may send from it. Only someone explicitly assigned to a mailbox may send; without that assignment the sending tool does not exist for that employee at all. Incoming mail is polled every minute and starts tasks; replies are bundled and only sent with substance — no mail ping-pong. Replies to a mail the platform sent find their way back into the same case. Protection against auto-reply loops and injected foreign instructions is built in. Meetings are part of it too: find free slots, create calendar entries, set up a Teams meeting — every call states explicitly on whose behalf it acts. Every task is a thread the colleague continues in without being mentioned again.
AI colleagues are reachable by @mention in the Teams channel, look after centrally managed mailboxes and schedule meetings — with no new tool for your staff.
A case carries a subject the platform formulates itself (for mail it comes from there) and a status: open, done, archived. Completed cases collapse in the chat instead of filling the view; a separate overview shows what is still open across the company — per employee, per channel, with follow-ups. Closing is driven by content, decided by whoever answered last, not by a clock: there is deliberately no rule that closes old cases automatically. A case that stays open for a long time is a signal that something is stuck — that must not disappear quietly. After a longer period of silence it moves to the archive and comes back on its own as soon as someone replies, even if that is a mail three weeks later. Nothing is ever deleted. And while a colleague works, it says so: "Sabine is working …", with the tool she is currently using, and on long runs with progress — instead of a mute pause that looks like a defect.
Every task is a case with a subject and a status — and you can see live what is being worked on.
Inbound: a file on a message lands in the working directory of the case — from the chat via the paperclip, from a mail, or from a connected bridge. It sits there twice: as the original, byte for byte, and as a readable text version the employee reads the content from. Both share the same name and the same version, so they cannot drift apart; both can be downloaded. The content does not travel along in every AI request — the message carries a note, and it is read when needed. Whatever cannot be converted (a scan without a text layer, a corrupted file) is stated explicitly rather than vanishing quietly. And an attachment from outside is foreign text: the readable version carries a header with its origin and the note that its content is information, not instruction. Outbound: the employee writes Markdown and has a PDF or an Excel workbook made from it. Layout, typeface and footer come from a template in your knowledge area — the letterhead belongs to the company, and the model can neither choose nor overwrite it. It can only attach what lies in the working directory of this case, nothing from company knowledge. And an attachment points at a version: if work continues at the source, the old message still shows what went out back then — it stays a record.
Whatever arrives by mail, chat or Teams gets read — and whatever leaves is a finished document on your letterhead, not a text file.
Recurring work runs on schedules, accurate to the minute and in your time zone: the morning mailbox check, the weekly report, daily follow-up on open items. Every run leaves a handover note for the next — filed in the case, and like a shift handover the thread is never lost. External systems trigger employees through secured webhooks (signed, with an individual secret per employee). A colleague can also put a case on their own follow-up date — "chase this in 48 hours if no answer arrives" — visible on the case instead of being a silent hope. And when something repeatedly goes wrong the platform pulls the cord itself: after three consecutive failures the employee is stopped and the supervisor informed.
AI colleagues don't wait to be asked: they work on a schedule, react to events from other systems and hand over their own shifts.
How a case should run at your company — sick leave, incoming invoices, customer requests — is described once in the central process library: trigger plus concrete steps, in plain language. Processes apply tenant-wide or specifically to certain teams and divisions. AI colleagues know their relevant processes and are obliged to follow them; they look up details themselves when needed. A running process can also shape the tool rules: "while the annual accounts are being closed, no postings without approval" is a condition you can write. Before the real thing, every process can be played through with a specific employee in a scenario — with an independent verdict and a full log. Change processes instead of retraining employees: one change takes effect immediately for everyone.
You describe your procedures once, centrally — every relevant AI employee follows them, and you can rehearse each process beforehand.
What an AI colleague works with: your company knowledge, connected systems — and tools that never leave your network.
You organise company knowledge in a file tree: upload files or maintain them in the web editor, and every change creates a new version with author and rollback. The top-level folders are also the searchable knowledge areas — an employee searches semantically in exactly the area assigned to them, strictly separated per tenant. AI may only write into company knowledge if a human approves — always, without exception. Alongside that, every case has its own working directory: AI colleagues file interim states, reports and results there, versioned and inspectable; delegated members work in the same directory as their team lead, which is why everyone involved sees the same state. Whatever you want to take with you, you download — individually, several as a ZIP, or a case's whole working area at once. The download checks visibility on its own; nobody receives anything that does not belong to their tenant.
Company knowledge and every case's working files live on one page — as a versioned file tree in which every state has an author.
Skills enter the platform through the open MCP standard (Model Context Protocol — the same standard Claude, ChatGPT and others speak): CRM, calendar, ticketing, document archive, your own internal services. A directory helps with discovery, management is central. Credentials live encrypted in the credential vault only (with a per-tenant key) and never appear in clear text — not even in approval cards. An employee gets a skill only through deliberate assignment; whether and when they may use it is decided by the tool rules. Every skill carries a readable profile: not just what the tool is called, but what someone can do with it — the team lead sees that too when handing out work. And when a file needs to go to an external system, say an invoice into the document archive, the employee names it and the platform moves it: the content does not pass through the model and not into the log, and the approval card shows the file name. The model names, the platform moves.
Tools are managed centrally and granted deliberately — credentials sit encrypted in a vault, not in a prompt.
Some tools are fundamentally unreachable from outside — a server on the company network, a legacy system, a database behind two firewalls. That is what the runner is for: a small service that runs inside your network and registers with the platform through an outbound connection. The platform never calls it; nothing has to be opened in the firewall — that is the whole point. For employees nothing changes: a runner tool is assigned, ruled and logged like any other. Three things make it fit for operations. Confirmation is required: when a runner reports a new skill it first appears as a proposal and delivers no tools until a company admin confirms it — so an update cannot arm anything by itself. Visibility: the platform shows whether a runner is connected, when it last reported in and how the machine is doing; if it is gone, tasks fail visibly instead of someone quietly working on without their tool. Operator authority: what may be executed on the machine is decided solely by the operator, at the machine itself — from the platform there is no way into those settings, and that is deliberate.
Internal systems become usable without opening a door in your firewall: a small service at your site dials into the platform, never the other way round.
The same colleague, a real workspace: fetch a branch, build, test, hand over.
On a runner inside your network, a provider can offer development workspaces. A workspace belongs to exactly one case and one employee; two colleagues on the same case work in separate directories and do not get in each other's way. The colleague opens it explicitly — with a source control system they pick the branch from the list of available ones, and that choice applies to everyone involved in the case. Then they work: list, search, read, write and precisely edit files — and, if the operator has enabled it, run commands: build, run tests, start a service and work against it. Project instructions in the tree are read automatically, so your conventions apply without anyone copying them into a personnel file. Three properties are the core. Nothing is ever checked in — the result leaves the runner as a package a human reviews and takes over, and the service account has no check-in permission at all. Nothing fails quietly — a workspace that cannot be provisioned or returned cleanly is taken out of service and the task fails visibly; and writing to a file that changed since you read it is refused. Large trees are normal — a branch with six-figure file counts gets fetched, that is allowed to take a while, the chat shows what is being worked on meanwhile, and the waiting time does not count against the work budget.
An AI colleague gets a real workspace on your source code: fetch a branch, build, test, hand over — and a human reviews the handover like any other contribution.
What you need to keep this under control — and what the whole thing stands on.
Everything AI employees do is recorded in a tamper-evident, fully chained audit log — every call, every approval, every decision, every cent; on a failure the reason as well, so nobody has to guess later what went wrong. Day to day you don't need log archaeology: ask Helga "Why isn't Sabine responding?" and get the answer in seconds — status, root cause, stuck deliveries, open approvals. Every Monday the operations report appears automatically in the channel of your choice: top 5 costs, incidents of the week, concrete recommendations. Management needs visibility — here it is built in, not bolted on.
A tamper-evident log, a one-question diagnosis ("Why isn't X responding?") and an automatic weekly report with recommendations.
Every night each employee looks back at the day's communication: what went badly, which phrasing landed wrong, which rule is missing? From that comes at most one concrete improvement proposal for its own personnel file. That proposal first goes through the assessment centre — the "improved" candidate must pass — and then reaches the supervisor as an approval card. Applied changes are versioned and individually revertible; rejected proposals are remembered so they don't come back. Your AI workforce gets better — but never behind your back.
AI colleagues reflect on their work at night and propose improvements — nothing is applied without passing assessment and a human approval.
MyAiColleague is multi-tenant by design, not retrofitted: separation between companies is enforced by the database itself (row-level security) — it does not depend on application code. Sign-in runs through enterprise SSO from day one (OIDC/SAML, e.g. your Microsoft Entra), with fine-grained roles down to a single approval permission. The platform is operated as SaaS in EU data centres; no data leaving the EU, and no vendor lock-in on models: multi-provider out of the box, the model selectable and swappable per role. Then there is API access: you issue keys with which scripts and automations execute the same commands as a human in the studio — no second permission model, nobody can grant more than they hold themselves, the normal case is a read-only key, and every call is logged. A key is visible exactly once; revoking takes effect immediately but does not erase history.
Data separation is enforced by the database, sign-in runs through your corporate identity system, operations sit in EU data centres — and anything a human can do in the studio, a script can do too.
30 minutes, one use case from your company, no slides.