Contextual Audio Frame Splitting for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems fail to accurately interpret user-specific speech patterns due to static audio slicing, which separates words related to the same context into different frames, leading to inaccuracies in speech recognition and interpretation, especially in user-specific applications, and struggle with variations in tone, emotion, and urgency, as well as complications from multiple audio channels.

Innovation Solution

A system that dynamically splits speech signals into audio frames based on context changes, adapts audio slicing windows to user interactions, and uses multiple audio processing modules to enhance context detection and transcription accuracy, particularly in conference calls.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If static audio slicing is used with fixed duration or word count, then the speech signal can be processed in uniform segments, but words related to the same context are separated into different frames leading to inaccurate speech recognition

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic audio slicing where the slicing duration is adjusted based on detected context changes in the speech signal. The system uses sentiment analysis and context detection to identify when contextual shifts occur, then adapts the slicing window accordingly. This ensures that words belonging to the same context remain within the same audio frame, improving recognition accuracy while maintaining processing efficiency through automated adaptation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the slicing parameter (duration) dynamically based on the detected context. When context changes are detected through sentiment analysis, the slicing window is adjusted to align with contextual boundaries. This parameter adaptation allows the system to maintain both processing efficiency and accuracy by optimizing frame boundaries according to the actual speech content structure.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If static audio slicing is used, then the system structure remains simple, but it fails to account for individual variations in speech patterns such as rate, pause, and intonation

Engineering Contradiction:
Improvesystem structureVSAvoiduser-specific speech pattern adaptation
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system performs self-adaptation by automatically detecting context changes and adjusting slicing parameters based on the speech signal characteristics. Through integrated sentiment analysis and context detection, the system serves itself by dynamically optimizing the audio slicing without requiring manual configuration for each user, thereby achieving user-specific adaptation while keeping the overall system structure manageable.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces dynamic adjustment capabilities that allow the system to adapt to individual speech patterns. By continuously monitoring speech rate, pauses, and intonation through sentiment analysis, the system dynamically modifies slicing parameters to match each user's unique speech characteristics, enhancing adaptability without requiring completely different system architectures for each user.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If contextually dynamic slicing is implemented, then speech recognition accuracy improves, but the system complexity increases due to multiple processing steps

Engineering Contradiction:
Improvecontext detection accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines sentiment analysis, context detection, and audio slicing functions into an integrated processing pipeline. By merging these operations and allowing them to work synergistically, the system achieves high context detection accuracy while managing complexity through functional integration rather than separate independent modules, reducing overall system complexity despite the advanced capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements a multi-functional processing module that performs sentiment analysis, context detection, and dynamic slicing within a unified framework. This universal approach allows the same core infrastructure to handle multiple tasks, reducing the need for separate specialized systems and thereby controlling complexity while maintaining high accuracy in context detection and speech recognition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250391421A1System and method for contextual analysis and metadata database generation for user-specific speech patterns
Publication Date: 2025.12.25 BANK OF AMERICA CORP
  • US20250391421A1 patent drawing
  • US20250391421A1 patent drawing
  • US20250391421A1 patent drawing

AI summary

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.