Phoneme Sequence Analysis for Multi-Language Phrase Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing contact center audio recordings are time-consuming and inefficient, particularly when searching for specific phrases, as they fail to account for regional pronunciation, cadence differences, and homologues, leading to false positives and difficulties in multi-language environments.

Innovation Solution

A multi-language system and method that uses an audio analysis engine to identify and mark specific phrases within spoken audio recordings by comparing phonemes, adjusting for regional accents and cadence, and allowing multiple audio representations of a phrase to be associated with a single identifier, enabling accurate capture and visualization of phrase usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If keyword recognition is used to search for phrases in audio recordings, then the analysis process is accelerated, but false positives increase due to inability to account for regional pronunciation and cadence differences

Engineering Contradiction:
Improveanalysis efficiencyVSAvoidphrase identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments phrases into individual phonemes and creates phoneme sequences for comparison. This segmentation allows the system to analyze speech at the phoneme level rather than relying on whole-word keyword matching, enabling accurate identification of phrases despite variations in pronunciation, cadence, and regional accents. The phoneme sequence comparison method breaks down the phrase recognition problem into smaller, more manageable units that can be systematically compared against the audio recording.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If manual review of entire audio recordings is performed to ensure accurate phrase identification, then measurement precision is maintained, but loss of time increases significantly

Engineering Contradiction:
Improvephrase identification accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical process of manual audio review with an automated computer-based phoneme sequence comparison system. The system automatically converts audio recordings into phoneme sequences and compares them against target phrase phoneme sequences, eliminating the need for manual listening and review. This substitution maintains high accuracy through systematic phoneme-level analysis while dramatically reducing the time required for phrase identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If single-language keyword search is used, then device complexity is minimized, but adaptability decreases when handling multi-language and multi-accent environments

Engineering Contradiction:
Improvesystem simplicityVSAvoidmulti-language support capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal phoneme-based search system that can handle multiple languages and accents through a common phoneme sequence comparison framework. Rather than creating separate search systems for each language, the system uses phoneme sequences as a universal representation that works across different languages and regional pronunciations. This universal approach allows the same core algorithm to identify phrases in various languages and accents without requiring language-specific customization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10469623B2Phrase labeling within spoken audio recordings
Publication Date: 2019.11.05 ZOOM INT AS
  • US10469623B2 patent drawing
  • US10469623B2 patent drawing
  • US10469623B2 patent drawing

AI summary

A system and method for multi-language phrase identification within spoken interaction audio capable of adjusting for regional pronunciation (accents), cadence differences, and homologs. In this system, a spoken interaction audio data store supplies spoken audio data such as contact center call recordings to be analyzed for a specific phrase or set of phrases. Phrases are entered as natural language text and converted to the phonemes representative of the phrase audio using the invention's language packs and stored in a data store. Spoken interaction and phrase audio are converted to a digital format allowing comparison using multiple characteristics. Phrase matches are stored for subsequent post analysis display and analytics generation.