Voice Identification via Dual Decoding Paths
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice identification technologies, particularly those using joint modeling of acoustic and language models, face challenges in balancing identification accuracy and acoustic diversity, leading to compromised accuracy rates due to language constraints and limited domain adaptability.
Innovation Solution
The proposed method employs double decoding using independent acoustic models, where one model is generated by acoustic modeling alone and the other by joint acoustic and language modeling, expanding the decoding space and improving accuracy by incorporating acoustic diversity and multi-feature fusion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If joint modeling of acoustic and language models is used, then voice identification accuracy is improved, but acoustic diversity is reduced
Solution Approach 1:
The patent segments the voice identification process into two independent decoding paths: one using acoustic model only and another using joint acoustic-language model. This segmentation allows both acoustic diversity and identification accuracy to be preserved by processing the same input through different model configurations independently.
Solution Approach 2:
The patent merges the results from two separate decoding processes (acoustic model only and joint acoustic-language model) by selecting the best identification result between them. This combination strategy leverages the strengths of both approaches to achieve superior overall performance.
2Measurement precision
If joint modeling of acoustic and language models is used, then voice identification accuracy is improved, but domain adaptability is reduced
Solution Approach 1:
The patent segments the modeling approach into two parallel paths: one maintaining pure acoustic modeling for domain adaptability and another using joint acoustic-language modeling for accuracy. This allows the system to adapt to different domains while maintaining high identification accuracy through the dual-path architecture.
3Measurement precision
If language constraints are applied in acoustic modeling, then identification precision is improved, but correct paths are pre-clipped
Solution Approach 1:
The patent inverts the traditional approach by running two decoders in parallel: one with language constraints and one without. The unconstrained decoder ensures that correct acoustic paths are not pre-clipped, while the constrained decoder provides precision. The final result is selected by comparing both outputs, thereby inverting the problem of path clipping.
Data Source
AI summary
Embodiments of the present disclosure provide a method and apparatus for voice identification, a device and a computer readable storage medium. The method may include: for an inputted voice signal, obtaining a first piece of decoded acoustic information by a first acoustic model and obtaining a second piece of decoded acoustic information by a second acoustic model, where the second acoustic model being generated by joint modeling of acoustic model and language model. The method may further include determining a first group of candidate identification results based on the first piece of decoded acoustic information, determining a second group of candidate identification results based on the second piece of decoded acoustic information, and then determining a final identification result for the voice signal based on the first group of candidate identification results and the second group of candidate identification results.


