Deskwerk· Workshop

From the engine roomOperations

7 AI Agents in 4 Months: what I talked about at Zendesk Showcase Munich

In June I got to stand on stage at Zendesk Showcase in Munich and talk about what happens when you don’t announce AI in customer support, you just build it. Here’s the long version for everyone who wasn’t there. Plus the details that didn’t fit into 20 minutes of stage time.

The starting point: a USP loses its protection

For years our 24/7 support was a unique selling point. Then capable AI came along, and suddenly any competitor can be reachable around the clock, at a fraction of the effort. At the same time, our contact volumes kept climbing.

And our chatbot at the time? It solved a good half of the chat requests, but under the hood it was a rigid decision tree: painful to program, tedious to maintain, and it could only do chat. For phone and email there was simply nothing. Meanwhile, roughly 80% of requests are routine cases.

The strategic math was simple: without a transformation, we fall behind the market on efficiency and service quality. With one, 24/7 service becomes something that can’t be copied so easily again.

The plan: a two-track strategy

From the start, the project ran on two tracks that reinforce each other:

Technology modernization: an AI that learns, across all three main channels (chat, phone, email), around the clock, in four languages (DE, FR, IT, EN), wired into the core systems: player data, player protection, Zendesk as the ticketing heart.

Upgrading the team’s role: away from reactive ticket-clearing, toward AI supervision, escalation management, and proactive care. The goal: shift 30% of working time to higher-value work. One doesn’t work without the other. Roll out the technology alone and you end up with frustrated agents and an unsupervised bot.

Just as important as the scope was what was not in it: complex cases, VIP care, and compensation decisions stay with humans. At most five or six core systems get integrated, not everything with an API. That line saved us from scope creep more than once.

Before the first sprint began

The most underrated part of the project maybe happened before the official start. Three things were ready before the first sprint week began:

The knowledge base was structured and cleaned up: FAQs consolidated, SOP documents digitized. An AI trained on a messy Help Center gives messy answers. Then a persona and tone-of-voice concept per language: How does the bot speak? How formal in French, how direct in German? And an integration concept with API architecture, data flows, and escalation logic. Most of the integrations to the core systems were even pre-built.

Without that prep work, the four-month plan would have been wishful thinking. With it, it was aggressive but doable.

Four months in fast-forward

Sprint 1 (late January – mid-February): Build out the chat flows completely: greeting, intent detection, standard requests, escalation logic. Internal testing in all four languages, then UAT with the support team: answer quality, tone, escalations. The team validates, not just the project lead.

February 16: Go-Live chat. After three weeks, on all platforms at once, website and app. Quick win first: value you can see immediately, while the harder pieces are still in progress.

And because a bot in production has to be controllable, the agents got an app of their own: if something’s down, a few clicks add an extra message at the start of the chat, or steer individual bot use cases on purpose, across every channel and without a detour through IT. Communicate incidents proactively, before the wave of requests hits. The live demo shows what that looks like.

Sprint 2: Two weeks of intensive monitoring with daily review of every interaction, hotfixes, prompt adjustments, the top 10 missing knowledge entries filled in. In parallel, the Voice architecture took shape: Speech-to-Text and Text-to-Speech in four languages, integrated into the phone system.

Early March: Voice pilot at one location. Deliberately starting where the requests are more standardized: opening hours, directions, reservations. Then rollout to four locations.

Then came the reality check. The Voice tracks slipped by about three weeks. The new phone platform turned out to be far from intuitive, a UI change from the vendor right before the project start introduced new bugs, and the race conditions that come with Voice ate up serious project time, like misconfigured tokens that cut the connection to the bots. What saved us: the old solution kept running in parallel the whole time. Only what was ready went into production; on systemic failures, we migrated back immediately. Staggered migration with a fallback option: unglamorous, but worth its weight in gold.

April: Voice for the online business (login problems, payments, bonus rules, the more complex cases) went into 24/7 operation after the pilot. The email automation was finished and validated, but its production Go-Live was deliberately pushed to a later phase: a scoping decision in favor of stabilizing Voice, not a failure. Sprint 6 delivered the reporting dashboard with real-time KPIs across every channel, ROI tracking, and a utilization overview, plus the passed compliance audit.

The numbers: including the ones we missed

Chat: 65% autonomous resolution against a stretch target of 70 (OKR score 0.93), with 48 live use-case dialogs. Speech recognition: over 90% accuracy in all four languages, including Swiss German, which is anything but a given. Email: 96% categorization accuracy across 15 intent categories, validated on more than 200 test emails. All told: seven bots, more than 100 use-case dialogs, over 25 API integrations. Well beyond what the project brief had promised.

And Voice? 27% land-based, 20% online. Well under the 50% target, and still the right call. We deliberately gave callers the choice: bot or human, right at the start. Around 70% pick the human first. That opt-in logic structurally drags the rate down, but it puts the experience first. And it gives us real acceptance data instead of forced bot contacts.

My favorite detail from the Voice design: on the call, the bot proactively inspects the account and suggests the likely reason right away: “Is this about your play break?” No menu read out loud, no “Press 3.” That’s the difference between a voice-response system and an agent.

Compliance is architecture, not a footnote

In a regulated environment, one thing is non-negotiable: sensitive topics like gambling addiction or money laundering, the AI detects reliably and escalates to a human immediately. Every AI decision lands in the audit trail, the exclusion check runs automatically against the player-protection system, the data flows are documented. The compliance audit (GDPR, anti-money-laundering law, casino legislation) passed in the last sprint. Because it was built in from Sprint 1, not bolted on at the end.

On top of that, quality assurance in production: if an answer’s confidence drops below 80%, a human takes over. A keyword blacklist triggers independently of that. Interactions are reviewed weekly; the feedback loop between team and bot is a fixed process. Optimizing the resolution rate is routine operations, not a project.

What we learned, and what got confirmed

Voice isn’t chat: we knew it, now we have the proof. Speech recognition is the solved part; latency, caller behavior, and telephony infrastructure are their own domains of complexity, each with its own learning curve. The real value of the project sits exactly here: we now have real Voice operations experience. Acceptance data, patterns, edge cases from everyday work. No consultant can hand you that; it only comes from running it in production. And it’s the foundation the next generation of Voice builds on.

Stretch goals are a leadership tool, not a forecast. The 50% for Voice was set deliberately high to create pull, not because we thought it was a sure thing. The key is knowing that yourself: automation takes time. The rate grows in production with every optimization cycle, not in the project plan.

The big one: can doesn’t mean should. Just because a process can be automated doesn’t mean it belongs to the bot. We used the Value-Irritant-Matrix as our guide: whatever only annoys customers and the company alike gets eliminated, not automated. The repetitive work with no advisory value goes to the AI. And the conversations that create value for both sides belong, deliberately, in human hands. That’s exactly what we’re freeing up the time for.

What comes next

The project was never an end in itself. The real goal is scale: absorbing growing request volumes and the peaks around events and promotions, without adding headcount, at steady wait times and the same service quality. The infrastructure for that is in place now.

And it keeps growing. Top of the list: context exchange between the bots. Someone who first writes to the chat bot and then calls shouldn’t have to start over: the Voice bot should know what was discussed in chat. A memory across channels, that’s the next build stage. On top of that, the Voice expertise we’ve built flows straight into the next generation: Zendesk-native Voice bots, Whisper-based, with much lower latency and natural conversational interruption.

The core message: automate tasks, not people

The AI takes the repetitive work off our hands: password resets, status questions, the endless copy-paste. That gives my team time for the conversations that matter: listening, thinking along, building relationships.

The result in numbers: 75% more active customers than three years ago, operations around the clock, and not a single extra position for it. Not because we cut jobs, but because the AI absorbs the growth while people do what people do better.

What you can take away

  1. Start with the contact reasons, not the tool. Volume and top topics determine what’s worth automating. The Value-Irritant-Matrix decides what should be automated in the first place.
  2. Invest in the prep work. A cleaned-up knowledge base, tone-of-voice, integration concept: that decides the pace you can afford later.
  3. Quick win first. One channel, three weeks, a visible result. That buys you the trust for the harder pieces.
  4. Don’t expect too much early on. Automation takes time: 40–60% autonomous resolution in the first year is realistic, depending on your contact reasons. After that, the rate keeps growing in production.
  5. Design good fallbacks. The bot has to know when it’s done: a limited number of attempts, then handoff to a human, with the context so far instead of “Sorry, I didn’t catch that” on an endless loop. How gracefully a system fails decides whether people accept it.

Want to know what this would look like in your support operation? I build systems like this outside my day job too. Get in touch.

← Back to: Operations