Unified Audio Search Using Multi-Model Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional audio search systems in call recording databases face challenges in accurately identifying keywords due to variations in pronunciation, accents, and languages, leading to suboptimal search results as they rely on a single language model, failing to effectively process audio files with multiple languages and dialects.

Innovation Solution

The system employs multiple acoustic and language models to process audio streams, generating multiple search tracks and results, which are then combined into a unified search result using clustering and confidence scoring techniques to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single language model is used for audio processing, then the system complexity is reduced, but the search accuracy deteriorates due to inability to handle multiple languages and accents

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple language models and acoustic models into a unified search system. Multiple search tracks generated from different models are merged into a single unified search result, allowing the system to handle multiple languages and accents while maintaining manageable complexity through integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system employs a universal audio processing framework that can accommodate multiple language models and acoustic models simultaneously. This multi-functional approach enables the same system to process diverse audio content in various languages and dialects without requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple language models are employed to handle diverse languages and accents, then the search accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
Improvesearch accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the audio processing into separate tracks for different language models and acoustic models. Each track processes the audio independently, allowing parallel computation and enabling the system to handle multiple languages simultaneously without sequential processing delays.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from single-model processing to multi-dimensional processing by employing multiple language models and acoustic models concurrently. This dimensional expansion allows the system to process audio in parallel across different linguistic dimensions, improving accuracy while managing processing time through parallelization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If traditional single-model speech recognition is used, then the system is simpler to implement, but keywords are missed due to inaccurate conversion of audio files

Engineering Contradiction:
Improvekeyword identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple search tracks generated by different language models and acoustic models into a unified search result. This combination allows the system to capture keywords that might be missed by individual models, improving keyword identification accuracy through collective processing power.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS7725318B2System and method for improving the accuracy of audio searching
Publication Date: 2010.05.25 NICE SYST INC
  • US7725318B2 patent drawing
  • US7725318B2 patent drawing
  • US7725318B2 patent drawing

AI summary

A system and method for improving the accuracy of audio searching using multiple models to process an audio file or stream to obtain search tracks. The search tracks are processed to locate at least one search term and generate multiple search results. The number of search results is equivalent to the number of models used to process the audio stream. The search results are combined to generate a unified search result. The multiple models may represent different languages, dialects and accents.