Voice Identification via Dual Decoding Paths

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice identification technologies, particularly those using joint modeling of acoustic and language models, face challenges in balancing identification accuracy and acoustic diversity, leading to compromised accuracy rates due to language constraints and limited domain adaptability.

Innovation Solution

The proposed method employs double decoding using independent acoustic models, where one model is generated by acoustic modeling alone and the other by joint acoustic and language modeling, expanding the decoding space and improving accuracy by incorporating acoustic diversity and multi-feature fusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If joint modeling of acoustic and language models is used, then voice identification accuracy is improved, but acoustic diversity is reduced

Engineering Contradiction:
Improvevoice identification accuracyVSAvoidacoustic diversity
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the voice identification process into two independent decoding paths: one using acoustic model only and another using joint acoustic-language model. This segmentation allows both acoustic diversity and identification accuracy to be preserved by processing the same input through different model configurations independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the results from two separate decoding processes (acoustic model only and joint acoustic-language model) by selecting the best identification result between them. This combination strategy leverages the strengths of both approaches to achieve superior overall performance.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If joint modeling of acoustic and language models is used, then voice identification accuracy is improved, but domain adaptability is reduced

Engineering Contradiction:
Improvevoice identification accuracyVSAvoiddomain adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the modeling approach into two parallel paths: one maintaining pure acoustic modeling for domain adaptability and another using joint acoustic-language modeling for accuracy. This allows the system to adapt to different domains while maintaining high identification accuracy through the dual-path architecture.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If language constraints are applied in acoustic modeling, then identification precision is improved, but correct paths are pre-clipped

Engineering Contradiction:
Improveidentification precisionVSAvoidcorrect path retention
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent inverts the traditional approach by running two decoders in parallel: one with language constraints and one without. The unconstrained decoder ensures that correct acoustic paths are not pre-clipped, while the constrained decoder provides precision. The final result is selected by comparing both outputs, thereby inverting the problem of path clipping.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS11145314B2Method and apparatus for voice identification, device and computer readable storage medium
Publication Date: 2021.10.12 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11145314B2 patent drawing
  • US11145314B2 patent drawing
  • US11145314B2 patent drawing

AI summary

Embodiments of the present disclosure provide a method and apparatus for voice identification, a device and a computer readable storage medium. The method may include: for an inputted voice signal, obtaining a first piece of decoded acoustic information by a first acoustic model and obtaining a second piece of decoded acoustic information by a second acoustic model, where the second acoustic model being generated by joint modeling of acoustic model and language model. The method may further include determining a first group of candidate identification results based on the first piece of decoded acoustic information, determining a second group of candidate identification results based on the second piece of decoded acoustic information, and then determining a final identification result for the voice signal based on the first group of candidate identification results and the second group of candidate identification results.