OmniServe’s AI: Fixing Customer Service in 2026

Listen to this article · 9 min listen

The call center at OmniServe was a maelstrom. By late 2025, customer inquiries had surged by 30% year-over-year, yet their human agent capacity remained static. Frustration mounted for both customers and agents, leading to an alarming 15% dip in customer satisfaction scores over six months. OmniServe needed a solution, and fast, to improve AI efficiency in handling the deluge of customer service data without sacrificing quality. How could they measure the true impact of their AI investments?

Key Takeaways

  • Implementing AI in customer service requires clear, quantifiable metrics beyond simple resolution rates to assess true efficiency.
  • The HANK metric, combining human intervention, AI accuracy, and net resolution time, provides a well-rounded view of AI performance.
  • OmniServe’s HANK score improved by 25% within three months of refining their AI models, demonstrating significant operational gains.
  • Regular auditing of AI-handled interactions, focusing on misclassifications and escalations, is essential for continuous improvement.
  • A successful AI integration strategy demands iterative refinement based on real-world customer service data, not just initial deployment.

OmniServe, a major player in telecommunications, had already invested heavily in various AI-driven chatbots and virtual assistants over the past two years. Their goal was clear: offload routine queries, free up human agents for complex issues, and reduce average handle time. However, their existing metrics, primarily focused on “AI resolution rate,” painted an overly optimistic picture that didn’t align with their plummeting customer satisfaction. “We’d see an 80% resolution rate reported by the AI system,” explained Sarah Chen, OmniServe’s Head of Customer Experience, during a recent industry conference. “But then we’d dig into the customer feedback, and it was clear many of those ‘resolved’ issues were just customers giving up or being shunted to an agent after a frustrating AI interaction.”

This disconnect highlighted a fundamental problem in how many organizations evaluate their AI deployments: focusing on isolated metrics that fail to capture the well-rounded customer journey. Simply put, an AI that “resolves” an issue by frustrating a customer into silence is not efficient. This realization led OmniServe to seek a more complete framework for measuring AI performance, one that accounted for the entire interaction lifecycle, not just the AI’s isolated contribution. They needed something that could truly quantify the value of their AI investments.

Enter the HANK metric, a framework gaining traction among AI practitioners for its focus on well-rounded performance. HANK stands for Human Assistance Needed, AI Accuracy, Net Resolution Time, and Knowledge Base Utilization. It moves beyond superficial success rates to assess the true collaborative efficiency between AI and human agents. “The HANK framework forces you to look at the entire process,” noted Dr. Anya Sharma, a data scientist specializing in conversational AI at the Institute for Advanced Computing (IAC) in Atlanta, Georgia. “It acknowledges that AI isn’t an island. It’s part of a larger ecosystem.”

OmniServe decided to pilot the HANK metric across their technical support department, a particularly challenging area given the complexity of issues. Their initial baseline, established in January 2026, revealed stark realities. The “Human Assistance Needed” component, measured by the percentage of AI interactions that eventually required a human agent, stood at a sobering 45%. This meant nearly half of all AI-initiated customer contacts still ended up with a human, often after significant customer frustration. The “AI Accuracy” score, derived from human audits of AI responses for correctness and completeness, was only 68% for critical information. “Net Resolution Time” for AI-first interactions was often longer than direct human interactions because customers had to repeat themselves. Finally, “Knowledge Base Utilization” by the AI was high, but often led to generic, unhelpful responses.

The first step for OmniServe was to carefully tag and categorize every customer interaction. This involved a significant investment in data annotation and a shift in their data collection protocols. Previously, they simply recorded if an AI handled an interaction. Now, they tracked specific points of friction: when a customer requested a human, when a human agent had to correct AI-provided information, and the total time from initial contact to final resolution, regardless of who handled it. This granular approach to customer service data was critical for establishing an accurate HANK baseline.

Their findings highlighted specific weaknesses in their current AI models. For instance, the AI struggled with nuanced language around network connectivity issues. Customers often used colloquial terms that the AI’s natural language processing (NLP) model couldn’t properly interpret, leading to irrelevant suggestions or immediate escalations. “It was clear our training data was too clean, too formal,” Sarah Chen observed. “Real customers don’t speak in perfectly formed sentences about their ‘WAN connection status’.”

Armed with this detailed HANK analysis, OmniServe embarked on a targeted refinement of their AI models. They focused on three key areas: expanding their NLP training data with more diverse, real-world customer queries. Implementing a dynamic feedback loop where human agents could flag incorrect AI responses directly. And optimizing the AI’s escalation protocols to route complex issues to the right human agent more quickly, rather than letting customers languish. They also integrated sentiment analysis into the AI’s monitoring to detect early signs of customer frustration, triggering human intervention proactively.

Over the next three months, the impact was tangible. By April 2026, OmniServe saw a significant improvement in their HANK scores. The “Human Assistance Needed” dropped to 30%, a 15-percentage-point reduction. “AI Accuracy” climbed to 85%, largely due to the continuous feedback loop from human agents correcting misclassifications. Perhaps most importantly, “Net Resolution Time” for AI-first interactions decreased by an average of two minutes, as the AI became more adept at either resolving issues efficiently or escalating them appropriately. This wasn’t just about speed. It was about quality of interaction. Customers weren’t just getting answers faster, they were getting the right answers faster, or being directed to the right expert without unnecessary delay.

The improvement wasn’t magical. It was the result of a disciplined, data-driven approach to AI deployment. “Many companies deploy AI and then expect it to just work,” Dr. Sharma commented. “But AI, especially in dynamic environments like customer service, requires constant tuning and re-evaluation. The HANK metric provides the roadmap for that.” OmniServe’s experience shows a fundamental truth: AI is a tool, and its effectiveness is directly proportional to how well it’s measured and managed. Simply throwing more AI at a problem without a strong performance framework like HANK often leads to wasted resources and exacerbated customer frustration.

One particular success story emerged from the technical support department. A common issue involved customers reporting “slow internet.” Previously, the AI would run through a generic troubleshooting script. Now, with enhanced NLP and better integration with network diagnostics, the AI could quickly identify if the “slow internet” was a localized issue, a broader outage (by checking internal network status APIs), or a customer’s specific device problem. This led to a 40% reduction in “slow internet” calls requiring human intervention, as the AI could either resolve it or provide more precise diagnostic steps. This precision is where true AI efficiency lies.

What OmniServe learned is that the initial investment in AI is only the beginning. The continuous investment in understanding its performance through complete metrics, and then iteratively refining the models based on that understanding, is what truly unlocks its potential. It’s not about replacing humans entirely. It’s about helping them by offloading the mundane and allowing them to focus on the truly complex and empathetic interactions that only a human can provide. The HANK framework provided the necessary lens to see this collaborative potential clearly.

The journey for OmniServe continues, of course. They are now exploring how to integrate the “Knowledge Base Utilization” component more effectively, ensuring the AI not only pulls information but synthesizes it into truly helpful, personalized responses. This involves further refining their internal knowledge base structure and making it more machine-readable. Their experience is a powerful reminder that effective AI implementation is an ongoing process of measurement, analysis, and adaptation.

Implementing a complete metric like HANK allows organizations to move beyond superficial AI success rates and truly understand the impact of their AI investments on both operational efficiency and customer satisfaction.

What does HANK stand for in AI performance metrics?

HANK stands for Human Assistance Needed, AI Accuracy, Net Resolution Time, and Knowledge Base Utilization, providing a well-rounded view of AI performance in customer service.

Why are traditional AI resolution rates often insufficient for measuring efficiency?

Traditional AI resolution rates can be misleading because they often do not account for customer frustration, repeated contacts, or issues that are “resolved” only because the customer gives up, failing to capture the true customer experience or the need for human intervention.

How can an organization establish a baseline for AI efficiency using the HANK metric?

To establish a HANK baseline, organizations need to carefully track and categorize customer interactions, noting instances of human escalation, AI response correctness, total interaction time, and how effectively the AI uses its knowledge base, often requiring significant data annotation.

What types of data are important for improving AI accuracy in customer service?

Improving AI accuracy heavily relies on diverse, real-world customer queries for natural language processing (NLP) training, alongside a dynamic feedback loop from human agents who can flag and correct misclassified or incorrect AI responses.

What is the long-term benefit of using metrics like HANK for AI deployment?

The long-term benefit of using complete metrics such as HANK is continuous improvement of AI models, leading to enhanced operational efficiency, better customer satisfaction, and a more strategic allocation of human agent resources for complex, high-value interactions.

Charles Scott

Lead Data Strategist M.S. Data Science, Carnegie Mellon University; Certified Data Scientist (CDS)

Charles Scott is a Lead Data Strategist at Veridian News Analytics, with 14 years of experience specializing in predictive trend analysis for digital news consumption. She leverages sophisticated data modeling to forecast audience engagement and content virality. Her work has been instrumental in shaping editorial strategies for major news outlets, and she is the author of the influential white paper, 'The Algorithmic Pulse: Decoding News Readership in the Mobile Age.'