bookmi.
Try free
← Back to blog
Tech2026-01-18· 12 min· Bookmi Tech Team

Learning from Sierra AI: Design Principles for Restaurant AI Agents

Adapting Sierra AI's τ-bench evaluation, deterministic coding, and state management for Bookmi.

Learning from Sierra AI

Last year, we were puzzled.

Bookmi's AI agent was performing excellently on industry-standard benchmark tests. It answered questions accurately, retrieved information correctly, and generated grammatically perfect responses. However, when actually deployed to restaurants, the feedback from owners was quite different from what we expected.

"The AI can answer questions, but reservation changes never seem to go smoothly."

That single comment fundamentally changed our way of thinking.

Correct Answers and Correct Outcomes Are Different Things

Traditional AI benchmarks—GSM8K, MMLU, and many others—evaluate one answer for one question. Can it solve math problems? Does it have general knowledge? These are certainly important capabilities.

But what happens on the restaurant floor is far more complex.

When a customer says "I'd like to change tomorrow's reservation from 7 PM to 8 PM," what's required of the AI isn't simply answering "Yes, we can change that." It needs to confirm the current reservation, check availability at 8 PM, execute the change, send a confirmation message, and if necessary, adjust the kitchen's preparation schedule—all within a natural conversation with the customer.

Whether the task can be completed. We realized that this is the evaluation criterion that truly matters.

Task Completion Rate: A New Metric

We revisited our evaluation methodology from scratch.

The new approach doesn't evaluate AI on "single Q&A" but on "entire tasks." Creating new reservations, changing times, modifying party sizes, handling cancellations, responding to special requests—we simulate these real scenarios and measure whether the AI can handle them correctly from start to finish.

The results were surprising. Models that scored over 90% on traditional tests sometimes achieved only around 60% in task completion rate. Conversely, models we had fine-tuned ourselves, while scoring slightly lower on traditional metrics, performed significantly better in task completion.

Evaluation Scenarios Designed Specifically for Restaurants

The evaluation scenarios we designed encompass the unique complexities of restaurant operations.

Peak-time reservation conflicts—when reservations pile up at 7 PM on Friday, can the AI appropriately suggest alternatives? Allergy handling—when a customer has multiple food allergies, can it identify safe options from the entire menu? Last-minute changes—when the party size increases significantly on the day of reservation, can it coordinate everything from seat rearrangement to additional ingredient orders?

We've prepared hundreds of such scenarios and run tests automatically every week.

A Cycle of Continuous Improvement

Evaluation isn't a one-time effort.

We analyze real conversation logs every day. Cases where the AI failed, cases where customers were dissatisfied, cases where operators had to intervene—all of these become valuable data for the next improvement.

Task completion rate has improved from 72% at initial deployment to 94% today. But we're not satisfied yet. Because within that remaining 6%, there are still moments where we failed to meet customer expectations.

Pursuing perfection, we continue to evolve our AI every single day.

Answering calls, replying to reviews, following up

AI can take over most of the work that eats your time. Try it for 30 days, no credit card.

Start 30-day free trial →Read other articles →