Multi-Engine Speech Analytics for Contact Center Intent Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech analytics systems in contact centers primarily focus on semantic analysis, which may not accurately capture non-semantic characteristics like emotion, gender, and personality, leading to potential misinterpretation of speech intentions.

Innovation Solution

A speech analytics system that integrates both semantic and non-semantic analysis capabilities, utilizing multiple speech analytics engines to detect keywords, emotions, age, and gender, and providing event notification messages for appropriate handling of calls based on these characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single speech analytics system is used for both semantic and non-semantic analysis, then the system complexity increases and processing accuracy may deteriorate, but deploying separate systems increases cost and infrastructure complexity

Engineering Contradiction:
Improvespeech interpretation accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech analytics system is divided into multiple independent speech analytics engines, each specialized in specific analysis types (semantic, non-semantic, emotion detection, speaker identification). Each engine processes speech data independently and outputs results that are integrated by the system, allowing modular complexity management while maintaining high processing accuracy for each specialized function.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If multiple speech analytics engines are deployed to provide both semantic and non-semantic indicators, then the processing capability and accuracy improve, but the system complexity and resource consumption increase

Engineering Contradiction:
Improvespeech analysis capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system dynamically selects and activates only the speech analytics engines required for each specific analysis task. Rather than running all engines continuously, the system adjusts the active processing components based on the current analytical needs, reducing unnecessary computational resource consumption while maintaining versatile speech analysis capabilities when needed.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If semantic analysis is performed alone, then the processing speed is maintained, but the interpretation accuracy of speech intentions deteriorates due to lack of non-semantic context

Engineering Contradiction:
Improvespeech intention detection accuracyVSAvoidadditional processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Non-semantic analysis engines perform preliminary analysis of speech characteristics (emotion, speaker identity, tone) before semantic analysis is conducted. This preliminary processing extracts contextual information that enhances the subsequent semantic interpretation, improving overall speech intention detection accuracy without significantly increasing total processing time due to the parallel architecture.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9014364B1Contact center speech analytics system having multiple speech analytics engines
Publication Date: 2015.04.21 ALVARIA INC
  • US9014364B1 patent drawing
  • US9014364B1 patent drawing
  • US9014364B1 patent drawing

AI summary

Various embodiments of the invention provide methods, systems, and computer-program products for providing a plurality of speech analytics engines in a speech analytics module for detecting semantic and non-semantic speech characteristics in the audio of a call involving an agent in a contact center and a remote party. The speech analytics module generates event notification messages reporting the detected semantic and non-semantic speech characteristics and these messages are sent to an event handler module that forwards the messages to one or more application specific modules. In turn, the application specific modules provide functionality based on the semantic and non-semantic speech characteristics detected during the call such as, for example, causing information to be presented on the screen of a computer used by the agent during the call.