Methods related to speech synthesis, training method of coarticulation model, and related devices
By training the speech flow sound change model, using pronunciation analysis and phonetic flow data adjustment, the problem of mechanical sound in existing speech synthesis is solved, achieving more natural speech synthesis and improving user experience.
Patent Information
- Application Number
- CN202011468064.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-14
AI Technical Summary
Existing speech synthesis methods are difficult to reduce the mechanical sounds of synthesized speech, affecting the user's experience of using it.
By obtaining training data, including text data, pinyin annotation data and phonetic flow data, the speech flow sound change model is trained. This model generates more natural phonetic flow data by performing pronunciation analysis and phonetic flow data adjustments on text data, thereby performing speech synthesis.
It improves the naturalness of speech synthesis, enhances the accuracy and flexibility of speech logos, and improves the user experience and usage experience.
Smart Images

Figure CN114694627B_ABST
Abstract
Claims
1. A training method for a sandhi model, characterized in that, it includes: Obtain training data, wherein the training data includes text data, pinyin annotation data of the text data, and phonetic stream data of the text data; wherein, the phonetic stream data of the text data is data annotated with phonetic symbols according to the rules of broad transcription; Generate standard phonetic stream data of the text data based on the text data, standard speech data of the text data, and pronunciation prosody of the text data; Input the text data, pinyin annotation data of the text data, and standard phonetic stream data of the text data into an initial model for model training. When the output of the initial model matches the standard phonetic stream data of the text data, obtain the sandhi model.
2. The training method for a sandhi model according to claim 1, characterized in that, the training method for the sandhi model further includes: Perform prosody analysis on the text data and obtain the word segmentation result after prosody analysis; Adjust the phonetic stream data of the corresponding text data according to the word segmentation result after prosody analysis and the phonetic stream data of the corresponding text data; Input the adjusted phonetic stream data of the text data and the corresponding text data, pinyin annotation data of the text data into the initial model for retraining to obtain the sandhi model.
3. A method for speech synthesis, characterized in that, it includes: Perform word segmentation on the text to be processed to obtain the word segmentation result of the text to be processed; Perform pinyin annotation on the word segmentation result of the text to be processed to obtain the pinyin annotation information of the text to be processed; Input the text to be processed and the pinyin annotation information of the text to be processed into a sandhi model to obtain the first phonetic stream data of the text to be processed; wherein, the sandhi model is trained by the training method for a sandhi model according to claim 1 or 2; Perform speech synthesis on the text to be processed based on the first phonetic stream data.
4. The method for speech synthesis according to claim 3, characterized in that, the method further includes: Perform prosody analysis on the first phonetic stream data and compare the result of the prosody analysis with the word segmentation result of the text to be processed corresponding to the first phonetic stream data; When the result of the prosody analysis is different from the word segmentation result of the text to be processed corresponding to the first phonetic stream data, input the result of the prosody analysis into the sandhi model again to obtain the second phonetic stream data; Perform speech synthesis on the text to be processed based on the second phonetic stream data.
5. The method for speech synthesis according to claim 4, characterized in that, the step of performing prosody analysis on the text to be processed and comparing the result of the prosody analysis with the word segmentation result of the text to be processed corresponding to the first phonetic stream data includes: Perform stress annotation on the first phonetic stream data of the text to be processed to obtain the first phonetic stream data after stress annotation; Perform prosody boundary division on the first phonetic stream data after stress annotation to obtain the second word segmentation result; Compare the second word segmentation result with the word segmentation result of the text to be processed corresponding to the first phonetic symbol stream data.
6. A human-computer interaction method, characterized in that, the human-computer interaction method includes: Receiving the user's conversation information and determining a response text for the conversation information based on the conversation information; Performing pinyin annotation on the response text to obtain pinyin annotation information of the response text; Inputting the response text and the pinyin annotation information of the response text into a speech flow phonetic change model to obtain first phonetic symbol stream data of the response text; wherein, the speech flow phonetic change model is trained by the training method of the speech flow phonetic change model described in claim 1 or 2; Performing speech synthesis on the response text based on the first phonetic symbol stream data to obtain response speech; Presenting the response speech to the user.
7. A speech synthesis device, characterized in that, the speech synthesis device includes a pinyin annotation module, a processing module, and a synthesis module, the pinyin annotation module is configured to perform pinyin annotation on the text to be processed to obtain pinyin annotation information of the text to be processed; the processing module is configured to input the text to be processed and the pinyin annotation information of the text to be processed into a speech flow phonetic change model to obtain first phonetic symbol stream data of the text to be processed; wherein, the speech flow phonetic change model is trained by the training method of the speech flow phonetic change model described in claim 1 or 2; the synthesis module is configured to perform speech synthesis on the text to be processed based on the first phonetic symbol stream data.
8. An electronic device, characterized in that, comprising a mutually coupled memory and a processor, the processor is configured to execute program instructions stored in the memory to implement the training method of the speech flow phonetic change model described in claim 1 or 2 or the speech synthesis method described in any one of claims 3-5 or the human-computer interaction method described in claim 6.
9. A computer-readable storage medium, on which program instructions are stored, characterized in that, when the program instructions are executed by a processor, the training method of the speech flow phonetic change model described in claim 1 or 2 or the speech synthesis method described in any one of claims 3-5 or the human-computer interaction method described in claim 6 is implemented.
Citation Information
Patent Citations
Voice data annotation method and device
CN113593522A