A Personalized Speech Translation Method and Device Based on Speaker Features
What is AI technical title?
AI technical title is built by Patsnap AI team. It summarizes the technical point description of the patent document.
A technology of speech translation and speaker, applied in the field of speech translation, can solve the problem of not solving the problem of applying personalized translation system to speaker characteristics, etc.
Active Publication Date: 2022-02-01
SICHUAN CHANGHONG ELECTRIC CO LTD
View PDF16 Cites 0 Cited by
Summary
Abstract
Description
Claims
Application Information
AI Technical Summary
This helps you quickly interpret patents by identifying the three key elements:
Problems solved by technology
Method used
Benefits of technology
Problems solved by technology
[0007] The present invention provides a speaker-based personalized speech translation method and device to solve the problem in the prior art that speaker features are not applied to the entire personalized translation system from speaker speech to text to speech
Method used
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more
Image
Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
Click on the blue label to locate the original text in one second.
Reading with bidirectional positioning of images and text.
Smart Image
Examples
Experimental program
Comparison scheme
Effect test
Embodiment 1
[0043] see Figure 1-3 , a method for personalized speech translation based on speaker characteristics, comprising the following steps:
[0044] Step 1, collecting the speaker's voice, extracting the speaker's voice acoustic feature, and converting it into a speaker feature vector;
[0045] The method for extracting the acoustic features of the speaker's speech is specifically to perform windowed Fourier transformation on the speaker's voice to obtain linear features, and then process the acoustic features of the speaker's speech through a Mel filter.
[0046] The speaker's speech acoustic features extracted by collecting people with different intonation features are input into the deep speech recognition model, and then trained with a deep learning network to obtain the speaker feature vector model corresponding to the speech acoustic features of different speakers.
[0047] The speaker's speech acoustic features extracted by the speaker are input into the speaker feature ve...
Embodiment 2
[0062] In this embodiment, a personalized speech translation device based on speaker features includes a speaker audio feature extraction unit, a speaker speech recognition unit, a translation unit, an encoder unit, and an end-to-end text feature-to-audio feature unit.
[0063] Speaker audio feature extraction unit, which performs windowed Fourier transformation on the speaker's voice to obtain linear features, and then obtains the speaker's voice acoustic features through Mel filter processing, and inputs the target voice acoustic features into the speaker feature vector model to get the speaker feature vector.
[0064] The speaker's speech recognition unit, which recognizes the speech as corresponding text according to the speaker's feature vector combined with the speaker's speech acoustic feature as the neural network input of the text recognition model.
[0065] The translation unit is used to translate the speaker's language into the target language. This unit translatio...
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More
PUM
Login to View More
Abstract
The invention discloses a personalized voice translation method based on speaker features, which comprises the following steps: collecting the speaker's voice, extracting the voice acoustic features of the speaker's voice, and converting it into a speaker feature vector; combining the speaker feature vector with the speaker Speech acoustic features for speaker text recognition; translate the speaker's text into the target language text; combine the text encoding of the target language generated in the previous step with the speaker feature vector generated in the first step to obtain a target with speaker characteristics Text vector; generate the target speech from the target text vector generated in the previous step through the text-to-speech model. By adding a speaker feature extraction network, the present invention can add different speakers' tone and intonation into the process of speech recognition and text-to-speech, helping to translate the speaker's meaning more accurately. The invention also discloses a personalized speech translation device based on speaker characteristics.
Description
technical field [0001] The invention relates to the technical field of speech translation, in particular to a speaker-based personalized speech translation method and device. Background technique [0002] With the development of globalization and the increase of exchanges between different countries, the importance of real-time voice translation is increasing. When the tone of the speaker changes in traditional voice translation, it may not be able to express the meaning of the speaker, and different regions have different interpretations. Certain words may have different pronunciations, which highlights the importance of personalized translations. [0003] At the same time, in the process of translation, there may be cases where the translated result is different from the actual application result due to the difference in the speaker's accent and intonation. For example, the message that the speaker wants to express is "Is there a hot dog seller nearby?" After speech recog...
Claims
the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More
Application Information
Patent Timeline
Application Date:The date an application was filed.
Publication Date:The date a patent or application was officially published.
First Publication Date:The earliest publication date of a patent with the same application number.
Issue Date:Publication date of the patent grant document.
PCT Entry Date:The Entry date of PCT National Phase.
Estimated Expiry Date:The statutory expiry date of a patent right according to the Patent Law, and it is the longest term of protection that the patent right can achieve without the termination of the patent right due to other reasons(Term extension factor has been taken into account ).
Invalid Date:Actual expiry date is based on effective date or publication date of legal transaction data of invalid patent.