Automatic Interpretation Apparatus Voice Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic interpretation technologies face challenges in real-time processing performance due to the need for extracting voice features from source voice files and transmission delays, and they cannot convert personalized synthesis voices into desired tones.
Innovation Solution
An automatic interpretation method and apparatus that allows the utterer terminal to transmit voice feature information and automatic translation results to the correspondent terminal, enabling real-time voice synthesis without requiring the source voice file, and allowing for the adjustment and conversion of tone in the personalized synthesis voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the correspondent terminal extracts voice features from the source voice file provided by the utterer terminal, then a personalized synthesis voice similar to the utterer's voice can be reproduced, but the processing time increases and real-time processing performance deteriorates
Solution Approach 1:
The voice feature extraction process is performed in advance by the utterer terminal before transmission to the correspondent terminal. The extracted voice features are pre-processed and prepared, eliminating the need for time-consuming extraction at the correspondent terminal during real-time interpretation.
Solution Approach 2:
The voice feature extraction function is separated from the correspondent terminal and relocated to the utterer terminal. Only the extracted voice features are transmitted, not the entire source voice file, reducing data transmission volume and processing time at the correspondent terminal.
2Measurement precision
If the source voice file is transmitted from the utterer terminal to the correspondent terminal for voice feature extraction, then personalized synthesis voice can be achieved, but transmission delay occurs and real-time processing performance deteriorates
Solution Approach 1:
Instead of transmitting the entire source voice file, only the essential voice features are extracted and transmitted. This dramatically reduces the data volume that needs to be transmitted over the communication channel, minimizing transmission delay.
Solution Approach 2:
The voice features serve as a compact representation or copy of the essential characteristics of the source voice. This compressed feature representation maintains the necessary information for synthesis while occupying minimal transmission bandwidth.
3Measurement precision
If the correspondent terminal performs voice feature extraction from the source voice file, then personalized synthesis voice similar to the utterer can be reproduced, but the device complexity and processing load increase
Solution Approach 1:
The complex voice feature extraction process is removed from the correspondent terminal and performed at the utterer terminal. The correspondent terminal only needs to receive and process the already-extracted features, significantly reducing its processing complexity.
Solution Approach 2:
The voice features act as an intermediary representation between the source voice and the synthesis process. They encapsulate the essential characteristics needed for personalized synthesis while simplifying the processing requirements at the correspondent terminal.
4Device complexity
If traditional voice synthesis is used without hidden variables and additional voice features, then the synthesis process is simpler, but the ability to convert personalized synthesis voice into desired tones is lost
Solution Approach 1:
The voice feature representation is extended from traditional dimensions to include hidden variables and additional voice features. This additional dimensional information enables tone conversion capabilities while maintaining a manageable synthesis process through structured feature organization.
Solution Approach 2:
The synthesis system utilizes adjustable parameters within the voice feature representation (hidden variables and additional features) to enable tone conversion. By modifying these parameters, the system can transform the personalized synthesis voice into desired tones without fundamentally changing the synthesis architecture.
Data Source
AI summary
An automatic interpretation method performed by a correspondent terminal communicating with an utterer terminal includes receiving, by a communication unit, voice feature information about an utterer and an automatic translation result, obtained by automatically translating a voice uttered in a source language by the utterer in a target language, from the utterer terminal and performing, by a sound synthesizer, voice synthesis on the basis of the automatic translation result and the voice feature information to output a personalized synthesis voice as an automatic interpretation result. The voice feature information about the utterer includes a hidden variable including a first additional voice result and a voice feature parameter and a second additional voice feature, which are extracted from a voice of the utterer.


