Speech Processing Data Modification for Accent Resilience

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech processing systems face challenges in handling a wide range of speaking accents and vocal irregularities, leading to performance degradation, and are prone to errors in classification tasks due to noise and errors in automated speech recognition processes.

Innovation Solution

The system configures modification pipelines to process input data based on speakers' accents and semantic similarity, generating modified data that includes a broader range of speech properties, which is then used for training, validation, and classification tasks, enhancing resilience to variations and noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional training techniques use batches of correlated training data, then the model performs accurately on similar speech properties, but the model performance degrades when input data has different speech properties such as various accents and vocal irregularities

Engineering Contradiction:
Improvemodel performanceVSAvoidresilience to speech property variations
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by generating synthetic speech data with diverse accents, vocal irregularities, and speech properties before actual model training. This preparatory data augmentation ensures the model encounters varied speech patterns during training, improving its adaptability to different speech properties while maintaining reliable performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by systematically varying speech properties in the training data, including accent types, vocal characteristics, noise levels, and speech conditions. By modifying these parameters across the training dataset, the model learns to handle diverse speech variations, resolving the contradiction between maintaining performance on familiar data and adapting to new speech properties.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If automated speech recognition is used for classification, then speech data can be converted to text data, but errors in ASR recognition cause incorrect classification

Engineering Contradiction:
Improveautomated classificationVSAvoidclassification accuracy
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system implements feedback mechanisms where classification results are evaluated and used to refine the ASR process. When classification errors are detected, the system adjusts the ASR parameters or re-processes the speech data, creating a closed-loop system that continuously improves classification accuracy while maintaining automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Before performing automated classification, the system performs preliminary actions by pre-processing speech data to enhance quality, remove noise, and optimize features for ASR. This preparatory processing reduces the likelihood of ASR errors, thereby improving classification accuracy while preserving the benefits of automation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11417317B2Determining input data for speech processing
Publication Date: 2022.08.16 CAPITAL ONE SERVICES LLC
  • US11417317B2 patent drawing
  • US11417317B2 patent drawing
  • US11417317B2 patent drawing

AI summary

Aspects described herein may relate to the determination of data that is indicative of a greater range of speech properties than input text data. The determined data may be used as input to one or more speech processing tasks, such as model training, model validation, model testing, or classification. For example, after a model is trained based on the determined data, the model's performance may exhibit more resilience to a wider range of speech properties. The determined data may include one or more modified versions of the input text data. The one or more modified versions may be associated with the one or more speakers or accents and/or may be associated with one or more levels of semantic similarity in relation to the input text data. The one or more modified versions may be determined based on one or more machine learning algorithms.