Acoustic Model Library for Voice-Adaptive Speech Translation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Translation devices often fail to distinguish between different speakers, leading to difficulties in recognizing individual voices during communication, which affects user experience and communication effectiveness.

Innovation Solution

A data processing method and device that utilize an acoustic model library with different timbre characteristics to determine and apply a target acoustic model based on voiceprint recognition, allowing for voice conversion that matches the timbre characteristics of individual users, either default, current, or preferred.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fixed timbre is used for synthesizing target language text, then the translation device can maintain consistent output quality, but different users cannot be distinguished and user experience deteriorates

Engineering Contradiction:
Improveoutput consistencyVSAvoidspeaker recognition
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements dynamic timbre selection by switching between different acoustic models based on voiceprint recognition results. The system transitions from a static fixed-timbre approach to a dynamic adaptive approach where the timbre characteristics change according to the identified speaker, resolving the contradiction between output consistency and speaker recognition.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the timbre parameter of the synthesized speech by selecting different acoustic models with distinct timbre characteristics corresponding to different users. This parameter change enables the system to maintain reliability through consistent synthesis quality while improving ease of operation by allowing speaker distinction through unique timbre profiles.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If voiceprint recognition and acoustic model selection are added, then speaker recognition is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker recognitionVSAvoidsystem structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-establishing an acoustic model library containing multiple acoustic models with different timbre characteristics before actual use. During operation, the system only needs to perform voiceprint recognition and select from pre-prepared models, which simplifies the real-time processing complexity while maintaining improved speaker recognition capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating multiple acoustic model copies with different timbre characteristics in the acoustic model library. Each model serves as a template for synthesizing speech with specific timbre properties, allowing the system to achieve speaker recognition without complex real-time timbre generation, thus managing device complexity effectively.

Inventive Principle:
Principle #26Copying

3Ease of operation

If multiple acoustic models are maintained in a library, then timbre characteristics can be preserved, but storage requirements and processing time increase

Engineering Contradiction:
Improvetimbre preservationVSAvoidmodel selection time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent applies local quality by organizing the acoustic model library with specific timbre characteristics assigned to different acoustic models. Each model in the library has localized, specialized timbre properties that can be quickly matched to voiceprint recognition results, enabling efficient model selection while preserving distinct timbre characteristics for different speakers.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11354520B2Data processing method and apparatus providing translation based on acoustic model, and storage medium
Publication Date: 2022.06.07 BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
  • US11354520B2 patent drawing
  • US11354520B2 patent drawing
  • US11354520B2 patent drawing

AI summary

In present disclosure, a data processing method, a data processing device, and an apparatus for data processing are provided. The method specifically includes: receiving a source language speech input by a target user; determining, based on the source language speech, a target acoustic model from a preset acoustic model library, the acoustic model library including at least two acoustic models corresponding to different timbre characteristics; converting, based on the target acoustic model, the source language speech into a target language speech; and outputting the target language speech. According to the embodiments of the present disclosure, the recognition degree of the speaker corresponding to the target language speech output by the translation device can be increased, and the effect of user communication can be improved.