Vehicle Audio Signal Processing for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems in vehicles face challenges in accurately distinguishing intended voice commands from various background audio sources, leading to false positives and reduced recognition rates due to noise and echo interference.

Innovation Solution

A system and method that process audio signals captured from microphones in vehicles by identifying and reducing unwanted audio sources, using reference signals from mixed audio outputs to improve speech recognition, with separate processing for human perception and machine recognition, employing noise reduction and echo cancellation techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If audio is captured continuously from the microphone to enable speech recognition, then the system can respond to voice commands, but false positive recognition results occur due to background noise and audio sources

Engineering Contradiction:
Improvespeech recognition response capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent extracts and removes known audio sources (navigation prompts, music, chimes, text-to-speech output) from the microphone signal before processing. By identifying these specific audio sources and selectively removing them, the system prevents false positives while maintaining continuous listening capability for genuine voice commands.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing stage between audio capture and speech recognition. This intermediate processor analyzes the microphone signal, identifies known audio sources, and removes them before the signal reaches the speech recognition system, thereby improving recognition accuracy without sacrificing responsiveness.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If echo cancellation and noise reduction are applied to process the audio signal, then recognition rates improve, but the processing complexity increases

Engineering Contradiction:
Improverecognition rateVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by removing known audio sources from the microphone signal before the speech recognition processing occurs. This pre-processing step simplifies the subsequent recognition task by eliminating predictable interference sources, thereby improving recognition rates without requiring overly complex real-time processing during the recognition phase.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple audio sources are present in the vehicle environment, then the audio system provides rich functionality, but the speech recognition system experiences false positives and reduced accuracy

Engineering Contradiction:
Improveaudio system functionalityVSAvoidvoice command detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent selectively extracts and removes specific known audio sources (navigation prompts, music, chimes, text-to-speech) from the mixed microphone signal. This extraction process preserves the functionality of these audio sources while eliminating their interference with speech recognition, allowing the system to maintain rich audio functionality without sacrificing detection accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3002754B1System and method for processing an audio signal captured from a microphone
Publication Date: 2018.06.27 2236008 ONTARIO INC
  • EP3002754B1 patent drawingFigure 1
  • EP3002754B1 patent drawingFigure 2
  • EP3002754B1 patent drawingFigure 3

AI summary

A system and method for processing an audio signal captured from a microphone may reproduce a known audio signal with an audio transducer into an acoustic space. The known audio signal may include content from one or more audio sources. A microphone audio signal may be captured from the acoustic space where the microphone audio signal comprises the known audio signal and one or more unknown audio signals. Processing control information may be accessed. The known audio signal may be reduced in the microphone audio signal responsive to the processing control information where the processing control information indicates one or more characteristics of a downstream audio processor that processes the microphone audio signal.