Voiceprint Segmentation for Speech Recognition in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face challenges in accurately recognizing voice commands in noisy environments, such as vehicles, where multiple voices are present, leading to insufficient accuracy due to interference from other speakers.
Innovation Solution
A speech recognition method that segments captured voice information, extracts voiceprint information, matches it with local voiceprint data, and combines filtered voice segments to determine accurate semantic information, improving recognition accuracy by filtering out irrelevant voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech recognition is performed on captured voice information in a multi-speaker environment, then the system can process voice commands, but the recognition accuracy deteriorates due to interference from other speakers
Solution Approach 1:
The captured voice information is segmented into multiple voice segments based on voiceprint characteristics. The system divides the mixed audio signal into distinct segments corresponding to different speakers, allowing separate processing and identification of the target speaker's voice from the background noise of other speakers.
Solution Approach 2:
The system extracts voiceprint information from the captured voice data and compares it with stored voiceprint templates to identify and extract the target speaker's voice segments. This extraction process separates the useful voice information from the harmful background noise of other speakers.
2Measurement precision
If voiceprint matching is performed to filter voice segments, then the accuracy of identifying the intended voice improves, but the processing time and computational complexity increase
Solution Approach 1:
Voiceprint template information is pre-stored in the system for multiple speakers. This preliminary preparation of reference data allows for rapid comparison and matching during actual speech recognition, reducing the processing time required for voice identification without compromising accuracy.
Data Source
AI summary
A speech recognition method includes segmenting captured voice information to obtain a plurality of voice segments, and extracting voiceprint information of the voice segments; matching the voiceprint information of the voice segments with a first stored voiceprint information to determine a set of filtered voice segments having voiceprint information that successfully matches the first stored voiceprint information; combining the set of filtered voice segments to obtain combined voice information, and determining combined semantic information of the combined voice information; and using the combined semantic information as a speech recognition result when the combined semantic information satisfies a preset rule.


