Concurrent Multi-Path Audio Demixing for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems face significant challenges in noisy environments, such as those found in vehicles, where ambient noise and interference lead to increased word error rates, making them less effective for real-life applications compared to controlled lab conditions.

Innovation Solution

The implementation of a real-time blind source separation technique for concurrent multi-path processing of audio signals, which uses a system comprising processors and a memory to demix mixed audio content into source-specific signals, improving recognition accuracy even with a small number of microphones and in the presence of interference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If more microphones are used to suppress interference and improve speech recognition robustness, then recognition accuracy improves, but device complexity and cost increase

Engineering Contradiction:
Improvespeech recognition robustnessVSAvoidnumber of microphones
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the audio signal processing into multiple independent paths: a first path for determining demixing parameters and a second path for applying these parameters to demix audio content. This segmentation allows the system to achieve robust speech recognition with fewer microphones by processing signals through specialized parallel pathways rather than requiring more microphones.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary determination of demixing parameters in the first signal processing path before applying them in the second path. This preliminary action enables the system to prepare separation parameters in advance, improving recognition robustness without requiring additional microphones during the actual speech recognition process.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If blind source separation techniques are applied to demix audio content in real-time, then speech recognition accuracy improves in noisy environments, but algorithm delay increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidalgorithm delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements continuous real-time processing through concurrent multi-path signal processing, where demixing parameters are continuously determined and applied without interrupting the audio signal flow. This continuous processing maintains speech recognition accuracy while minimizing algorithm delay by avoiding batch processing interruptions.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system dynamically adjusts demixing parameters in real-time based on changing audio conditions in noisy environments. The first signal processing path continuously updates demixing parameters, which are then immediately applied in the second path, allowing the system to adapt to dynamic noise conditions without significant processing delay.

Inventive Principle:
Principle #15Dynamics

3Productivity

If concurrent multi-path processing is implemented to reduce algorithm delay, then processing speed improves, but system complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidsignal processing structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the signal processing system into two distinct but concurrent paths: a first path for determining demixing parameters and a second path for applying them. This segmentation enables parallel processing that improves productivity while keeping each individual path relatively simple, balancing processing speed with manageable system complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12080274B2Concurrent multi-path processing of audio signals for automatic speech recognition systems
Publication Date: 2024.09.03 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US12080274B2 patent drawing
  • US12080274B2 patent drawing
  • US12080274B2 patent drawing

AI summary

A system and method for concurrent multi-path processing of audio signals for automatic speech recognition is presented. Audio information defining a set of audio signals may be obtained (502). The audio signals may convey mixed audio content produced by multiple audio sources. A set of source-specific audio signals may be determined by demixing the mixed audio content produced by the multiple audio sources. Determining the set of source-specific audio signals may comprises providing the set of audio signals to both a first signal processing path and a second signal processing path (504). The first signal processing path may determine a value of a demixing parameter for demixing the mixed audio content (506). The second signal processing path may apply the value of the demixing parameter to the individual audio signals of the set of audio signals (508) to generate the individual source-specific audio signals (510).