Speech Recognition Anchor Feature Extraction Hybrid Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition methods in hybrid speech scenarios suffer from low accuracy and inability to trace and recognize a specific target speaker without prior estimation of the number of speakers.

Innovation Solution

The method determines an anchor extraction feature of a target speech within a hybrid speech to obtain a mask, allowing for recognition of the target speech without pre-learning the number of speakers, using a double-layer embedding space for improved feature concentration and stability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of speakers in hybrid speech is learned or estimated in advance, then speech recognition can be performed, but the system cannot trace and extract a specific target speaker's speech

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtarget speaker extraction capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an anchor extractor as an intermediary component that generates anchor extraction features from the mixed speech signal. This anchor feature serves as a mediator that guides the attention mechanism to identify and extract the target speaker's speech without needing to pre-estimate the number of speakers. The anchor extractor acts as a bridge between the mixed speech input and the target speech output, enabling both recognition and target-specific extraction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the speech recognition problem from a traditional approach that operates directly on the mixed speech signal to one that first projects the signal into an anchor feature space through the anchor extractor. This dimensional transformation allows the system to operate in a new feature space where target speaker identification and extraction can be performed more effectively, resolving the contradiction between general recognition and specific target extraction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If traditional speech recognition methods are used, then processing can be performed, but accuracy is relatively low and target speaker cannot be traced

Engineering Contradiction:
Improvespeech processing efficiencyVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces traditional mechanical speech recognition approaches with a neural network-based system that incorporates attention mechanisms. The attention mechanism dynamically weights different components of the speech signal based on their relevance to the target speaker, allowing the system to focus computational resources on the most informative parts of the signal. This substitution of traditional methods with neural attention-based methods simultaneously improves both accuracy and efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If the system processes hybrid speech without anchor extraction, then computation is simpler, but feature concentration and stability are insufficient

Engineering Contradiction:
Improveprocessing complexityVSAvoidfeature concentration and stability
Core Design Contradiction:
Device complexityVSStability of the object's composition

Solution Approach 1:

The patent applies preliminary action by introducing the anchor extractor that generates anchor extraction features before the main speech recognition processing. This pre-processing step concentrates the relevant features and stabilizes the representation of the target speaker's speech characteristics. By performing this feature concentration action in advance, the subsequent recognition processes benefit from more stable and concentrated features without requiring overly complex processing in later stages.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3767619B1Speech recognition method and apparatus
Publication Date: 2024.12.04 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3767619B1 patent drawingFigure 1~2
  • EP3767619B1 patent drawingFigure 3
  • EP3767619B1 patent drawingFigure 4

AI summary

Embodiments of the disclosure relate to the speech recognition technology of artificial intelligence. A speech recognition method, a speech recognition apparatus, and a method and an apparatus for training a speech recognition model are provided. The speech recognition method includes: recognizing a target word speech from a hybrid speech, and obtaining, as an anchor extraction feature of a target speech, an anchor extraction feature of the target word speech based on the target word speech; obtaining a mask of the target speech according to the anchor extraction feature of the target speech; and recognizing the target speech according to the mask of the target speech.