Style-controllable speech synthesis system and method based on large model

By introducing the style feature extraction and discretization module and the style adjustment module, the problem of insufficient emotional expression in existing speech synthesis systems is solved, the controllable adjustment of speech style is achieved, and the naturalness and appeal of speech are improved. It is suitable for scenarios such as intelligent voice assistants and audiobooks.

CN120673744APending Publication Date: 2025-09-19GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510984140.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing speech synthesis systems have shortcomings in the emotional expression of speech and are unable to flexibly adjust the emotional style of speech according to user needs, resulting in the synthesized speech sounding monotonous and lacking in vividness and appeal.

Method used

A style feature extraction and discretization module and a style adjustment module are introduced to extract style features from text information and user-specified emotional tags, adjust the intonation, speaking speed, timbre and emotion of the speech, and make the synthesized speech conform to the style specified by the user.

Benefits of technology

It improves the naturalness and appeal of speech, enables the synthesized speech to be adjusted in style according to user needs, enhances the vividness and personalized experience, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673744A_ABST
    Figure CN120673744A_ABST
Patent Text Reader

Abstract

The invention relates to a style-controllable speech synthesis system and method based on a large model. The system comprises a text input module for receiving text information input by a user; the voice large model module is used for converting input text information into voice signals based on a pre-trained voice large model; the style feature extraction and discretization module is used for extracting style features and speaker timbre features from input text information and emotion tags specified by a user; the style adjusting module is used for adjusting the style of the voice generating process according to the extracted style features and the voice signals; and the voice output module is used for outputting the voice signal subjected to style adjustment as a playable audio file. According to the invention, by introducing the style feature extraction and discretization module and the style adjustment module, the synthesized voice can be adjusted according to the emotion style specified by the user, so that the naturalness and infection of the voice are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a style-controllable speech synthesis system and method based on a large model. Background Art

[0002] With the development of artificial intelligence (AI), text-to-speech (TTS) technology has been widely used in intelligent voice assistants, audiobooks, and voice broadcasting. However, existing TTS systems lack emotional expression in speech and are unable to flexibly adjust the emotional style of speech according to user needs. As a result, the synthesized speech sounds monotonous and lacks liveliness and appeal.

[0003] Therefore, it is necessary to provide a style-controllable speech synthesis system and method based on a large model. By introducing a style feature extraction and discretization module and a style adjustment module, the synthesized speech can be adjusted according to the emotional style specified by the user, thereby improving the naturalness and appeal of the speech. Summary of the Invention

[0004] The purpose of the present invention is to provide a style-controllable speech synthesis system and method based on a large model. By introducing a style feature extraction and discretization module and a style adjustment module, the synthesized speech can be adjusted according to the emotional style specified by the user, thereby improving the naturalness and appeal of the speech.

[0005] In order to solve the problems existing in the prior art, the present invention provides a style-controllable speech synthesis system based on a large model, comprising:

[0006] Text input module: receives text information input by the user;

[0007] Speech model module: Based on the pre-trained speech model, it converts the input text information into speech signals;

[0008] Style feature extraction and discretization module: extracts style features and speaker timbre features from input text information and user-specified emotion tags;

[0009] Style adjustment module: adjusts the style of speech generation process based on the extracted style features and speech signals;

[0010] Speech output module: outputs the style-adjusted speech signal as a playable audio file.

[0011] Optionally, in the large model-based style-controllable speech synthesis system,

[0012] User-specified emotion labels include style reference speech and speaker reference speech.

[0013] Optionally, in the large model-based style-controllable speech synthesis system, the style feature extraction and discretization module analyzes the style tendency from the reference speech and extracts style features based on the analysis results. The style features include the emotion, rhythm and volume of the speech.

[0014] Optionally, in the large model-based style-controllable speech synthesis system, style adjustment includes adjusting the intonation, speaking speed, timbre and emotion of the speech.

[0015] The present invention also provides a style-controllable speech synthesis method based on a large model, comprising the following steps:

[0016] Receive text information input by the user;

[0017] Based on the pre-trained speech model, the input text information is converted into speech signals;

[0018] Extract style features and speaker timbre features from input text information and user-specified emotion tags;

[0019] Adjust the style of speech generation process according to the extracted style features and speech signals;

[0020] Output the style-adjusted speech signal as a playable audio file.

[0021] Optionally, in the style-controllable speech synthesis method based on a large model,

[0022] User-specified emotion labels include style reference speech and speaker reference speech.

[0023] Optionally, in the large-model-based style-controllable speech synthesis method, the style feature extraction and discretization module analyzes the style tendency from the reference speech and extracts style features based on the analysis results, and the style features include the emotion, rhythm and volume of the speech.

[0024] Optionally, in the large model-based style-controllable speech synthesis method, style adjustment includes adjusting the intonation, speaking speed, timbre and emotion of the speech.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] (1) Rich style expression: The present invention introduces a style feature extraction and discretization module and a style adjustment module, so that the synthesized speech can be adjusted according to the emotional style specified by the user, thereby improving the naturalness and appeal of the speech, and enhancing its vividness and appeal.

[0027] (2) Strong adaptability: The system of the present invention can process a variety of text content, styles and timbres, and is suitable for different application scenarios, such as intelligent voice assistants, audio books, voice broadcasts, etc.

[0028] (3) Improved user experience: Users can customize the style and timbre of the speech according to their needs, improving the personalization of the speech synthesis system and user experience.

[0029] (4) The style-controllable speech synthesis system obtained by training based on a large amount of stylized speech data in the present invention can process a variety of text contents and speech styles. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A flowchart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.

[0032] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0033] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0034] With the development of artificial intelligence (AI), text-to-speech (TTS) technology has been widely used in intelligent voice assistants, audiobooks, and voice broadcasting. However, existing TTS systems lack emotional expression in speech and are unable to flexibly adjust the emotional style of speech according to user needs. As a result, the synthesized speech sounds monotonous and lacks liveliness and appeal.

[0035] In order to solve the problems existing in the prior art, the present invention provides a style-controllable speech synthesis system based on a large model, comprising:

[0036] Text input module: receives text information input by the user; preferably, also includes inputting the expected style reference voice and the expected speaker reference voice.

[0037] Speech model module: Based on the pre-trained speech model, it converts the input text information into speech signals (i.e., speech tokens);

[0038] The style feature extraction and discretization module extracts style features and speaker timbre features from the input text and user-specified emotion tags. The user-specified emotion tags include style reference speech and speaker reference speech. The module analyzes the style tendencies of the reference speech and extracts style features based on the analysis results. Style features include emotion, prosody, and volume.

[0039] Style adjustment module: performs style adjustment on the speech generation process based on the extracted style features and speech signals; style adjustment includes adjusting the intonation, speaking speed, timbre and emotion of the speech to make the speech conform to the style specified by the user.

[0040] Voice output module: Outputs the style-adjusted voice signal as a playable audio file for users to play.

[0041] The present invention also provides a style-controllable speech synthesis method based on a large model, such as Figure 1 As shown, the following steps are included:

[0042] Receive text information input by the user; preferably, also include inputting expected style reference voice and expected speaker reference voice.

[0043] Based on the pre-trained speech model, the input text information is converted into speech signals (i.e., speech tokens);

[0044] The style features and speaker timbre features are extracted from the input text and user-specified emotion tags, where the user-specified emotion tags include style reference speech and speaker reference speech. The style feature extraction and discretization module analyzes the style tendencies of the reference speech and extracts style features based on the analysis results. The style features include emotion, prosody, and volume of the speech.

[0045] The style of speech generation is adjusted according to the extracted style features and speech signals; style adjustment includes adjusting the intonation, speaking speed, timbre and emotion of the speech to make the speech conform to the style specified by the user.

[0046] The style-adjusted speech signal is output as a playable audio file for the user to play.

[0047] Specifically, Figure 1 Middle, 1. Left part: large speech model

[0048] (1) Text: The original text data entered by the user.

[0049] (2) Text Tokenizer: Converts the input text into a series of text tokens, which are the basis of text-to-speech conversion.

[0050] (3) Style Tokenizer: Extracts style tokens from the reference speech. These tokens represent the style speech features, such as intonation, speaking speed, emotion, volume, etc.

[0051] (4) Style Encoder: Encodes style markers and generates style embeddings. These embeddings contain the style information of the speech and are used to guide style adjustment during the speech synthesis process.

[0052] (5) Text-Speech Language Model: This model receives textual tokens and style embeddings and generates a sequence of speech tokens. This model is responsible for combining textual information with style information to generate speech output that conforms to a specific style.

[0053] 2. Right side: speech synthesis module

[0054] (1) Speech Prompt: Contains the initial information for speech synthesis, which may include basic parameter settings for speech.

[0055] (2) Speaker Encoder: Extracts speaker embeddings from speech prompts. These embeddings contain speaker feature information and are used to maintain the consistency of speaker identity in speech synthesis.

[0056] (3) Conditional Flow: This module receives the speech tokens, style embeddings, and speaker embeddings output by the speech model and generates the speech mel-spectrogram. This module is responsible for adjusting the synthesis of the speech mel-spectrogram based on the style and speaker characteristics.

[0057] (4) Vocoder: Converts the Mel-spectrogram output by the conditional stream module into an audible speech waveform. The vocoder is the last step in speech synthesis and is responsible for converting the synthesized speech signal into human-audible sound.

[0058] (5) Speech: The final output speech signal can be played or used in other applications.

[0059] 3. Other Notes

[0060] (1) Style Embedding and Speaker Embedding: These two embeddings are optimized through orthogonal loss to ensure that style and speaker characteristics can independently and effectively affect the final speech output during the speech synthesis process.

[0061] 4. Overall process:

[0062] The user inputs text and reference speech. The system extracts text and style tags through a text tokenizer and a style tokenizer. Then, the system generates a speech tag sequence that matches a specific style using a style encoder and a large text-to-speech model. These tag sequences are then converted into the final speech output using a conditional stream module and a vocoder.

[0063] In summary, the present invention has the following advantages compared with the prior art:

[0064] (1) Rich style expression: The present invention introduces a style feature extraction and discretization module and a style adjustment module, so that the synthesized speech can be adjusted according to the emotional style specified by the user, thereby improving the naturalness and appeal of the speech, and enhancing its vividness and appeal.

[0065] (2) Strong adaptability: The system of the present invention can process a variety of text content, styles and timbres, and is suitable for different application scenarios, such as intelligent voice assistants, audio books, voice broadcasts, etc.

[0066] (3) Improved user experience: Users can customize the style and timbre of the speech according to their needs, improving the personalization of the speech synthesis system and user experience.

[0067] (4) The style-controllable speech synthesis system obtained by training based on a large amount of stylized speech data in the present invention can process a variety of text contents and speech styles.

[0068] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A style-controllable speech synthesis system based on a large model, characterized by: include: Text input module: receives text information input by the user; Speech model module: Based on the pre-trained speech model, it converts the input text information into speech signals; Style feature extraction and discretization module: extracts style features and speaker timbre features from input text information and user-specified emotion tags; Style adjustment module: adjusts the style of speech generation process based on the extracted style features and speech signals; Speech output module: outputs the style-adjusted speech signal as a playable audio file.

2. The style-controllable speech synthesis system based on a large model as claimed in claim 1, characterized in that: User-specified emotion labels include style reference speech and speaker reference speech.

3. The style-controllable speech synthesis system based on a large model as claimed in claim 2, characterized in that: The style feature extraction and discretization module analyzes the style tendency from the reference speech and extracts style features based on the analysis results. The style features include the emotion, rhythm and volume of the speech.

4. The style-controllable speech synthesis system based on a large model according to claim 1, characterized in that: Style adjustment includes adjusting the tone, speed, timbre and emotion of the voice.

5. A style-controllable speech synthesis method based on a large model, characterized in that: The following steps are involved: Receive text information input by the user; Based on the pre-trained speech model, the input text information is converted into speech signals; Extract style features and speaker timbre features from input text information and user-specified emotion tags; Adjust the style of speech generation process according to the extracted style features and speech signals; Output the style-adjusted speech signal as a playable audio file.

6. The style-controllable speech synthesis method based on a large model according to claim 5, characterized in that: User-specified emotion labels include style reference speech and speaker reference speech.

7. The style-controllable speech synthesis method based on a large model according to claim 5, characterized in that: The style feature extraction and discretization module analyzes the style tendency from the reference speech and extracts style features based on the analysis results. The style features include the emotion, rhythm and volume of the speech.

8. The style-controllable speech synthesis method based on a large model according to claim 5, characterized in that: Style adjustment includes adjusting the tone, speed, timbre and emotion of the voice.