Voiceprint Segmentation for Speech Recognition in Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies face challenges in accurately recognizing voice commands in noisy environments, such as vehicles, where multiple voices are present, leading to insufficient accuracy due to interference from other speakers.

Innovation Solution

A speech recognition method that segments captured voice information, extracts voiceprint information, matches it with local voiceprint data, and combines filtered voice segments to determine accurate semantic information, improving recognition accuracy by filtering out irrelevant voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech recognition is performed on captured voice information in a multi-speaker environment, then the system can process voice commands, but the recognition accuracy deteriorates due to interference from other speakers

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidnoise interference from other speakers
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The captured voice information is segmented into multiple voice segments based on voiceprint characteristics. The system divides the mixed audio signal into distinct segments corresponding to different speakers, allowing separate processing and identification of the target speaker's voice from the background noise of other speakers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts voiceprint information from the captured voice data and compares it with stored voiceprint templates to identify and extract the target speaker's voice segments. This extraction process separates the useful voice information from the harmful background noise of other speakers.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If voiceprint matching is performed to filter voice segments, then the accuracy of identifying the intended voice improves, but the processing time and computational complexity increase

Engineering Contradiction:
Improvevoice identification accuracyVSAvoidprocessing time for voiceprint matching
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Voiceprint template information is pre-stored in the system for multiple speakers. This preliminary preparation of reference data allows for rapid comparison and matching during actual speech recognition, reducing the processing time required for voice identification without compromising accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11562736B2Speech recognition method, electronic device, and computer storage medium
Publication Date: 2023.01.24 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11562736B2 patent drawing
  • US11562736B2 patent drawing
  • US11562736B2 patent drawing

AI summary

A speech recognition method includes segmenting captured voice information to obtain a plurality of voice segments, and extracting voiceprint information of the voice segments; matching the voiceprint information of the voice segments with a first stored voiceprint information to determine a set of filtered voice segments having voiceprint information that successfully matches the first stored voiceprint information; combining the set of filtered voice segments to obtain combined voice information, and determining combined semantic information of the combined voice information; and using the combined semantic information as a speech recognition result when the combined semantic information satisfies a preset rule.