Personalized speech synthesis method, device and equipment based on scene perception and natural language description
By conducting semantic analysis of the text, generating scene sound characteristics and style control parameters, combining scene perception and natural language description, dynamically adjusting the speech synthesis parameters, solving the problem of mismatch between speech synthesis and scene in the existing technology, and achieving personalized and scene-adaptive speech synthesis.
Patent Information
- Application Number
- CN202510434793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing voice synthesis technology is difficult to dynamically adjust the voice background according to different application scenarios, and it is impossible to fully understand and utilize the user's personalized voice style requirements, resulting in the mismatch of synthetic voice and actual scenarios and lack of dynamic adaptability.
By conducting semantic analysis of the text, scene sound characteristics and style control parameters are generated, scene perception and natural language description are combined, speech synthesis parameters are dynamically adjusted to generate personalized and expressive voice.
The scene adaptability and personalized needs of speech synthesis are realized, the flexibility and controllability of speech synthesis are improved, the generated speech is more in line with actual needs, and the sense of substitution and reality are enhanced.
Smart Images

Figure CN120148475A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis technology, and particularly to a personalized speech synthesis method, device, and equipment based on scene perception and natural language description. Background Art
[0002] Speech synthesis (Text-to-Speech, TTS) technology converts text information into speech signals and is widely used in fields such as intelligent assistants, navigation systems, audiobook devices, and customer service robots, providing an efficient and natural solution for human-computer interaction. In recent years, speech synthesis models based on deep learning have made significant improvements in speech quality, fluency, and timbre fidelity, basically meeting the basic speech generation requirements in general scenarios.
[0003] However, most speech synthesis methods in related technologies rely on a single speech model and cannot dynamically adjust the speech background according to different application scenarios, resulting in a mismatch between the synthesized speech and the actual application scenario; moreover, in terms of emotion recognition methods, they usually rely on label conditions to generate speech and cannot comprehensively understand and utilize the user's personalized speech style requirements, such as "a cheerful tone" or "a deep male voice", etc. Therefore, related technologies lack the ability to dynamically adapt to specific application scenarios, user personalized needs, and interaction contexts. Summary of the Invention
[0004] In view of the above problems, embodiments of this application provide a personalized speech synthesis method, device, and equipment based on scene perception and natural language description to overcome or at least partially solve the above problems.
[0005] In the first aspect of the embodiments of this application, a personalized speech synthesis method based on scene perception and natural language description is disclosed. The method includes: Performing semantic analysis on a first text to obtain a first semantic feature, where the first text includes text content and a scene description; Generating a scene classification based on the first semantic feature to obtain a scene vector, where the scene vector represents a scene category; Generating a scene sound feature based on the scene vector, where the scene sound feature includes detailed features of the application scene and audio features related to the scene; Performing semantic analysis on a second text to obtain a second semantic feature, where the second text includes at least a natural language description of the speech style; Performing emotion category prediction and style feature extraction based on the second semantic feature to obtain style control parameters, where the style control parameters include: pitch, speech rate, tone, and volume; Adjust the synthetic speech parameters of the text content according to the scene sound characteristics and the style control parameters to obtain synthetic speech.
[0006] Optionally, adjusting the synthetic speech parameters of the text content according to the scene sound characteristics and the style control parameters to obtain synthetic speech includes: Fuse the scene sound characteristics and the style control parameters to obtain a comprehensive style adjustment parameter; Dynamically adjust the synthetic speech parameters of the text content according to the comprehensive style adjustment parameter to obtain synthetic speech.
[0007] Optionally, generating scene sound characteristics according to the scene vector includes: Match the scene vector with each scene sound template in the scene template library to obtain a target scene sound template that matches the scene vector; Generate scene sound characteristics according to the target scene sound template and the parameters of the speech synthesis model.
[0008] Optionally, performing semantic analysis on the first text to obtain first semantic features includes: Perform word segmentation and annotation processing on the first text through a natural language processing model to obtain the first text after word segmentation and annotation; Perform semantic encoding on the first text after word segmentation and annotation through a text encoder to obtain first semantic features.
[0009] Optionally, performing emotion category prediction and style feature extraction according to the second semantic features to obtain style control parameters includes: Perform emotion category prediction on the second semantic features through an emotion analysis model to obtain a first speech control parameter, where the first speech control parameter includes a first pitch, a first speech rate, and a volume; Perform tone recognition on the second semantic features through a tone recognition model to obtain a second speech control parameter, where the second speech control parameter includes a tone; Perform style conversion on the second semantic features through a style conversion model to obtain a third speech control parameter, where the third speech control parameter includes a second pitch and a second speech rate; Integrate the first speech control parameter, the second speech control parameter, and the third speech control parameter to obtain a style control parameter.
[0010] Optionally, performing emotion category prediction on the second semantic features through an emotion analysis model to obtain a first speech control parameter includes: Perform emotion category prediction on the second semantic features through the emotion analysis model to obtain a target emotion category label; According to the mapping relationship between the emotional category label and the voice parameter, map the target emotional category label to the first voice parameter.
[0011] Optionally, perform style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter, including: Process the second semantic feature through the style conversion model to obtain a style vector; Match the style vector with each style template in the style library to obtain a target style template that matches the style vector; Use the voice control parameter corresponding to the target style template as the third voice control parameter.
[0012] Optionally, the second text further includes context information of the text content; perform semantic analysis on the second text to obtain a second semantic feature, including: Parse and encode the second text through a language parsing model to obtain a second semantic feature.
[0013] In a second aspect of the embodiments of the present application, a personalized speech synthesis device based on scene perception and natural language description is disclosed. The device includes: A first analysis module, configured to perform semantic analysis on a first text to obtain a first semantic feature, where the first text includes text content and a scene description; A first generation module, configured to generate a scene classification based on the first semantic feature to obtain a scene vector, where the scene vector represents a scene category; A second generation module, configured to generate a scene sound feature based on the scene vector, where the scene sound feature includes detailed features of the application scenario and audio features related to the scene; A second analysis module, configured to perform semantic analysis on a second text to obtain a second semantic feature, where the second text includes at least a natural language description of a speech style; A first extraction module, configured to perform emotional category prediction and style feature extraction based on the second semantic feature to obtain a style control parameter, where the style control parameter includes: pitch, speech rate, tone, and volume; A first adjustment module, configured to adjust the synthesized speech parameter of the text content according to the scene sound feature and the style control parameter to obtain synthesized speech.
[0014] In a third aspect of the embodiments of the present application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the personalized speech synthesis method according to the first aspect of the embodiments of the present application are implemented.
[0015] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the personalized speech synthesis method based on scene perception and natural language description described in the first aspect of the embodiments of the present application are implemented.
[0016] In the fifth aspect of the embodiments of the present application, a computer program product is disclosed, including a computer program. When the computer program is executed by a processor, the steps of the personalized speech synthesis method based on scene perception and natural language description described in the first aspect of the embodiments of the present application are implemented.
[0017] The embodiments of the present application include the following advantages: In the embodiments of the present application, the first text includes text content and a scene description. By processing the first text, a scene sound feature including detailed features of the application scene and audio features related to the scene is generated, which can recognize and adapt to multiple application scenes (for example, news, station, conversation, etc.); thus, according to the scene description of the first text, the style and atmosphere of the synthesized speech are adjusted, which not only improves the adaptability of the synthesized speech and generates more practical speech in multiple scenes, but also enhances the sense of immersion and realism of speech synthesis. The second text includes at least a natural language description of the speech style. By processing the second text, a style control parameter is obtained, which can accurately analyze and extract the natural language description of the speech style, and then adjust parameters such as the emotional category, tone feature, pitch, and speech rate of the synthesized speech; thus, the method can not only understand rich language descriptions, but also flexibly adjust the speech output to meet the personalized speech needs of users, significantly improving the flexibility and controllability of speech synthesis. In this way, by combining scene information and natural language description, the method can generate personalized, expressive and context-compliant synthesized speech. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a flowchart of the steps of a personalized speech synthesis method based on scene perception and natural language description provided by the embodiments of the present application; Figure 2 is a schematic diagram of a scene perception module provided by the embodiments of the present application; Figure 3It is a schematic diagram of a natural language description parsing module provided by an embodiment of the present application; Figure 4 It is a schematic diagram of an adaptive regulation module provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a personalized speech synthesis device based on scene perception and natural language description provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] To make the above objects, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0021] Speech synthesis models based on deep learning have achieved significant improvements in speech quality, fluency, and timbre fidelity. For example, the speech synthesis model WaveNet generates natural speech waveforms through a deep convolutional neural network; the speech synthesis models Tacotron and FastSpeech directly generate mel spectrograms from text through an end-to-end neural network architecture and combine models such as WaveNet to generate high-quality speech.
[0022] To improve the expressive ability of speech, emotion-specific text-to-speech (Emotion-Specific TTS) has emerged. Speech synthesis methods based on emotion recognition control the emotional characteristics of speech (such as joy, anger, sorrow, and happiness) by analyzing the emotional information of the text; although speech synthesis methods based on emotion recognition adjust features such as pitch, speech rate, and stress of speech to a certain extent to express emotions. However, such methods mostly rely on emotion labels or manual annotations, and in practical applications, the accuracy and diversity of emotional expression are still greatly limited.
[0023] Therefore, the speech synthesis methods in the related technologies have the following limitations: 1) Insufficient scene adaptability. Existing speech synthesis methods usually use a single speech model and are difficult to adjust the scene sound according to different application scenarios, resulting in the synthesized speech not matching the scene atmosphere and lacking a sense of immersion. 2) Lack of natural language description ability. Existing speech synthesis methods mainly rely on label-conditioned speech generation and are difficult to understand and utilize natural language description information (such as "lively tone", "low-pitched male voice", etc.) to guide the speech synthesis process, restricting the flexibility and controllability of speech synthesis.
[0024] To overcome the limitations of related technologies, embodiments of the present application provide a personalized speech synthesis method based on scene perception and natural language description. This method can generate scene sounds according to the generated content, solve the problem that the synthesized speech does not match the scene atmosphere, and can understand and utilize natural language description information to adjust the synthesized speech style. Therefore, it can generate personalized, expressive and context-compliant synthesized speech.
[0025] The following will describe the personalized speech synthesis method based on scene perception and natural language description provided by the embodiments of the present application with reference to the accompanying drawings, respectively through Section 1.1 Personalized Speech Synthesis Method, Section 1.2 Scene Perception, Section 1.3 Natural Language Description Analysis, and Section 1.4 Adaptive Adjustment.
[0026] 1.1 Personalized Speech Synthesis Method: Referring to Figure 1 as shown, Figure 1 is a flowchart of the steps of a personalized speech synthesis method based on scene perception and natural language description provided by the embodiments of the present application. As Figure 1 shown, the personalized speech synthesis method based on scene perception and natural language description may include steps S110 to S160: Step S110: Perform semantic analysis on the first text to obtain a first semantic feature, where the first text includes text content and a scene description.
[0027] Among them, the text content in the first text refers to the content that needs to be synthesized into speech, and the text content can be news, stories, conversations, etc.; the scene description in the first text refers to the description of emotional color and occasion requirements. For example, the scene description can be a news broadcast scene description, a station scene description.
[0028] The first semantic feature is used to represent the semantic information of the first text. By performing semantic analysis on the first text, the first semantic feature is obtained. In some embodiments, a deep learning model can be used to perform semantic analysis on the first text to obtain the first semantic feature corresponding to the first text.
[0029] Step S120: Generate scene classification according to the first semantic feature to obtain a scene vector, where the scene vector represents a scene category.
[0030] Specifically, according to the semantic information represented by the first semantic feature, the corresponding scene category is identified, and a scene vector is generated according to the scene category. For example, if it is identified as a news broadcast scene according to the first semantic feature, a scene vector representing the news broadcast scene is generated; if it is identified as a station scene according to the first semantic feature, a scene vector representing the station scene is generated.
[0031] In some embodiments, the first semantic feature may be input into a pre-trained scene classification model, and the pre-trained scene classification model generates a scene classification based on the input first semantic feature to obtain a scene vector.
[0032] Step S130: Generate a scene audio feature according to the scene vector, where the scene audio feature includes detail features of the application scenario and audio features related to the scene.
[0033] In the embodiments of the present application, the scene audio feature is a scene audio feature adapted to the current target scene (i.e., the scene corresponding to the scene description in the first text); among them, the detail features of the application scenario may be news broadcasts, emotional dialogue detail information; the audio features related to the scene may be speech features such as volume, timbre, and background noise.
[0034] Different scene vectors correspond to different scene audio features. For example, for a news broadcast scene, a more formal and quiet scene audio feature is generated; while for a scene such as a station, a more noisy scene audio feature is selected.
[0035] Step S140: Perform semantic analysis on the second text to obtain a second semantic feature, where the second text includes at least a natural language description of the speech style.
[0036] In the embodiments of the present application, the required speech synthesis style is characterized by a natural language description. For example, the natural language description of the speech style may be descriptions such as "a cheerful tone" or "a deep male voice". The second semantic feature is used to represent the semantic information of the second text, and the second semantic feature is obtained by performing semantic analysis on the second text.
[0037] In some embodiments, semantic analysis may be performed on the second text based on a deep learning model to obtain a second semantic feature corresponding to the second text.
[0038] Step S150: Perform emotion category prediction and style feature extraction according to the second semantic feature to obtain a style control parameter, where the style control parameter includes: pitch, speech rate, tone, and volume.
[0039] Among them, the emotion category refers to emotions such as cheerfulness, sadness, and anger; the style feature refers to features such as pitch, speech rate, and tone. The style control parameter is a style control parameter generated according to personalized speech adjustment requirements such as emotion category and tone feature (for example, cheerfulness, deepness, fast speech rate, etc.).
[0040] In some embodiments, an emotion analysis model, a tone recognition model, a style conversion model, etc. may be used to perform emotion category prediction and style feature extraction on the second semantic feature to obtain a style control parameter.
[0041] Step S160: Adjust the synthesis voice parameters of the text content according to the scene voice feature and the style control parameter to obtain a synthesized voice.
[0042] In the embodiment of the present application, the scene voice feature includes the detailed features of the application scenario and the audio features related to the scenario, and the style control parameters include parameters such as pitch, speech rate, tone, and volume. Therefore, the style and voice characteristics of speech synthesis are dynamically adjusted by using the scene voice feature and the style control parameter, so as to realize dynamically adjusting the features such as pitch, speech rate, and tone of speech synthesis according to user needs, scene changes, and emotional feedback, and ensure the personalization and context adaptability of the voice output.
[0043] Through the above implementation process, the first text includes the text content and the scene description. By processing the first text, a scene voice feature including the detailed features of the application scenario and the audio features related to the scenario is generated, which can identify and adapt to multiple application scenarios (for example, news, station, conversation, etc.). Therefore, according to the scene description of the first text, the method adjusts the style and atmosphere of the synthesized voice, which not only improves the adaptability of the synthesized voice and generates a voice that better meets the actual needs in multiple scenarios, but also enhances the immersion and realism of speech synthesis. The second text includes at least a natural language description of the speech style. By processing the second text, style control parameters are obtained, which can accurately analyze and extract the natural language description about the speech style, and then adjust parameters such as the emotional category, tone feature, pitch, and speech rate of the synthesized voice. Therefore, this method can not only understand rich language descriptions, but also flexibly adjust the voice output to meet the personalized voice needs of users, significantly improving the flexibility and controllability of speech synthesis. In this way, by combining scene information and natural language description, this method can generate a personalized, expressive, and context-compliant synthesized voice.
[0044] 1.2 Scene perception: Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a personalized speech synthesis method based on scene perception and natural language description. In this method, the "performing semantic analysis on the first text to obtain a first semantic feature" in the above step S110 specifically may include sub-steps S110-1 to step S110-2: Step S110-1: Perform word segmentation and annotation processing on the first text through a natural language processing model to obtain a first text that has been word-segmented and annotated.
[0045] Step S120-2: Perform semantic encoding on the first text that has been word-segmented and annotated through a text encoder to obtain a first semantic feature.
[0046] In the embodiments of the present application, the tokenization and annotation processing of the first text refers to: splitting the first text into the smallest semantic units (e.g., words, sub-words, or characters, etc.), and annotating each smallest semantic unit (e.g., part-of-speech tagging, named entity recognition, etc.). The natural language processing model can be models such as BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder model based on Transformer), BiLSTM-CRF (Bidirectional Long Short-Term Memory-Conditional Random Field), etc., and the text encoder refers to the BERT or Transformer model based on deep learning.
[0047] In this way, semantic analysis of the first text is achieved through the natural language processing model and the text encoder, obtaining the first semantic feature corresponding to the first text, so as to subsequently achieve scene perception based on the first semantic feature and generate a scene sound feature adapted to the scene, improving the adaptability of the synthesized speech and enhancing the sense of immersion and realism of the speech synthesis.
[0048] Further, the "generating a scene sound feature according to the scene vector" in the above step S130 may specifically include sub-steps S130-1 to step S130-2: Step S130-1: Matching the scene vector with each scene sound template in the scene template library to obtain a target scene sound template that matches the scene vector.
[0049] Step S130-2: Generating a scene sound feature according to the target scene sound template and the parameters of the speech synthesis model.
[0050] In the embodiments of the present application, the scene template library contains multiple preset scene sound templates, and each scene sound template represents the scene sound features in different scenes. Matching the scene vector with each scene sound template in the scene template library can be achieved by calculating the similarity between the scene vector and each scene sound template (e.g., calculating the cosine similarity between the scene vector and the scene sound template), and determining the scene sound template corresponding to the maximum similarity as the target scene sound template that matches the scene vector.
[0051] The parameters of the speech synthesis model include parameters such as pitch, speech rate, tone, and volume. After obtaining the target scene sound template that matches the scene vector, combining the target scene sound template with the parameters of the speech synthesis model to obtain a scene sound feature adapted to the target scene (i.e., the scene corresponding to the scene description in the first text).
[0052] In this way, the method can identify and adapt to a variety of application scenarios (such as news, stations, conversations, etc.), adjust the style and atmosphere of the synthesized speech according to the generated scene sound features, thereby improving the adaptability of the synthesized speech, generating more practical speech in a variety of scenarios, and enhancing the immersion and authenticity of speech synthesis.
[0053] Refer to Figure 2 As shown, the methods of the above steps S110 to S130 can be implemented based on a scene perception module. Figure 2 It is a schematic diagram of a scene perception module provided by an embodiment of the present application. Among them, the scene perception module includes: a natural language processing model, a text encoder, a scene classification model, and a scene template library.
[0054] Specifically, the process of generating scene sound features based on the first text is as follows: The first text is segmented and annotated by a natural language processing model to obtain the segmented and annotated first text, and the segmented and annotated first text is semantically encoded by a text encoder to obtain a first semantic feature; then, a scene classification model is used to generate a scene classification according to the first semantic feature to obtain a scene vector; finally, the scene vector is matched with each scene sound template in the scene template library to obtain a target scene sound template that matches the scene vector, and scene sound features are generated according to the target scene sound template and the parameters of the speech synthesis model.
[0055] In this way, the specific application scenario is recognized by the scene perception module, and the speech features suitable for the current scene are generated, so that the speech synthesis can adjust the style of the synthesized speech according to the scene of the input text. Therefore, the method can adjust the style and atmosphere of the synthesized speech according to the scene description of the first text, not only improving the adaptability of the synthesized speech and generating more practical speech in a variety of scenarios, but also enhancing the immersion and authenticity of speech synthesis.
[0056] 1.3 Natural language description parsing: Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a personalized speech synthesis method based on scene perception and natural language description. In this method, the second text further includes the context information of the text content; specifically, "performing semantic analysis on the second text to obtain a second semantic feature" in the above step S140 may include: parsing and encoding the second text by a language parsing model to obtain a second semantic feature.
[0057] In the embodiments of the present application, in order to better understand the requirements of personalized speech styles, the context information of the text content is further included in the second text. The language parsing model can be a model based on BERT or GPT (Generative Pre-trained Transformer), and the second text is input into the language parsing model. The language parsing model decodes the input second text and outputs the second semantic features.
[0058] In this way, by parsing and encoding the second text through the language parsing model, the understanding of the second text is realized. Subsequently, control parameters such as the pitch, speaking speed, and tone of the speech can be generated based on the second semantic features obtained through the language parsing model, so as to realize the emotional control of the synthesized speech.
[0059] Further, the step of "performing emotional category prediction and style feature extraction according to the second semantic features to obtain style control parameters" in the above step S150 may specifically include sub-steps S150-1 to step S150-4: Step S150-1: Performing emotional category prediction on the second semantic features through an emotion analysis model to obtain a first speech control parameter, where the first speech control parameter includes a first pitch, a first speaking speed, and a volume.
[0060] Among them, the emotion analysis model can be an LSTM-based model. The second semantic features contain the semantic information of the second text. The second semantic features are input into the emotion analysis model, and the emotion analysis model performs emotional category prediction according to the input second semantic features to obtain the first speech control parameter.
[0061] Specifically, performing emotional category prediction on the second semantic features through the emotion analysis model to obtain a first speech control parameter includes: performing emotional category prediction on the second semantic features through the emotion analysis model to obtain a target emotional category label; according to the mapping relationship between the emotional category label and the speech parameter, mapping the target emotional category label to the first speech parameter.
[0062] In the embodiments of the present application, different emotional category labels correspond to different speech parameters (for example, speech parameters such as pitch, speaking speed, and volume). The mapping relationship between the emotional category label and the speech parameter is determined in advance. After obtaining the target emotional category label through the emotion analysis model, based on the mapping relationship between the emotional category label and the speech parameter, the target emotional category label is mapped into the corresponding semantic parameter, that is, the first speech parameter.
[0063] Step S150-2: Performing tone recognition on the second semantic features through a tone recognition model to obtain a second speech control parameter, where the second speech control parameter includes a tone.
[0064] Among them, the tone recognition model can be a model based on CNN (Convolutional Neural Network) or Transformer. The second semantic feature contains the semantic information of the second text. The second semantic feature is input into the tone recognition model, and the tone recognition model recognizes the corresponding tone according to the input second semantic feature to generate corresponding control parameters.
[0065] Step S150-3: Perform style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter, where the third voice control parameter includes a second pitch and a second speech rate.
[0066] Among them, the style conversion model can be a model based on Transformer or GAN (Generative Adversarial Network). The second semantic feature is input into the style conversion model to convert the text description into specific acoustic features, that is, the third voice control parameter.
[0067] Specifically, performing style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter includes: processing the second semantic feature through the style conversion model to obtain a style vector; matching the style vector with each style template in the style library to obtain a target style template that matches the style vector; using the voice control parameter corresponding to the target style template as the third voice control parameter.
[0068] In the embodiment of the present application, the style library contains multiple preset style templates, and each style template represents the voice control parameters corresponding to different styles. Matching the style vector with each style template in the style library can be achieved by calculating the similarity between the style vector and each style template (for example, calculating the cosine similarity between the style vector and the style template), and determining the style template with the maximum similarity as the target style template that matches the style vector. Finally, the voice control parameter corresponding to the target style template is used as the third voice control parameter.
[0069] Step S150-4: Integrate the first voice control parameter, the second voice control parameter, and the third voice control parameter to obtain a style control parameter.
[0070] In the embodiment of the present application, the first voice control parameter corresponding to the emotion label, the second voice control parameter corresponding to the tone, and the third voice control parameter corresponding to the style are integrated to obtain a set of final style control parameters. In this way, the conversion of natural language description into corresponding style control parameters is realized, so as to guide subsequent speech synthesis through the style control parameters and ensure that the speech style highly matches the user's needs.
[0071] It can be understood that the above-mentioned sentiment analysis model, tone recognition model, and style conversion model can be separate neural network models or different sub-models in a neural network model.
[0072] Referring to Figure 3 As shown, the methods of the above steps S140 to S150 can be implemented based on a natural language description parsing module. Figure 3 It is a schematic diagram of a natural language description parsing module provided by an embodiment of the present application. Among them, the natural language description parsing module includes: a language parsing model, a sentiment analysis model, a tone recognition model, a style conversion model, and a style library.
[0073] Specifically, the process of generating the style control parameter based on the second text is as follows: the second text is parsed and encoded by the language parsing model to obtain the second semantic feature; then, the sentiment analysis model is used to predict the sentiment category of the second semantic feature to obtain the target sentiment category label, and according to the mapping relationship between the sentiment category label and the speech parameter, the target sentiment category label is mapped to the first speech parameter; the second semantic feature is recognized for its tone by the tone recognition model to obtain the second speech control parameter; the second semantic feature is processed by the style conversion model to obtain a style vector, the style vector is matched with each style template in the style library to obtain the target style template that matches the style vector, and the speech control parameter corresponding to the target style template is used as the third speech control parameter; finally, the first speech control parameter, the second speech control parameter, and the third speech control parameter are integrated to obtain the style control parameter.
[0074] In this way, based on the natural language description parsing module, it can accurately parse and extract the natural language description about the speech style, and then adjust parameters such as the sentiment category, tone feature, pitch, and speech rate of the synthesized speech. This module can not only understand rich language descriptions but also flexibly adjust the speech output to meet the personalized speech needs of users, significantly improving the flexibility and controllability of the speech synthesis system.
[0075] 1.4 Adaptive adjustment: Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a personalized speech synthesis method based on scene perception and natural language description. In this method, in the above step S160, "adjust the synthesis speech parameters of the text content according to the scene sound feature and the style control parameter to obtain the synthesized speech" specifically includes sub-steps S160-1 to step S160-2: Step S160-1: Fuse the scene sound feature and the style control parameter to obtain a comprehensive style adjustment parameter.
[0076] Step S160-2: Dynamically adjust the synthetic speech parameters of the text content according to the comprehensive style adjustment parameters to obtain synthetic speech.
[0077] In the embodiments of the present application, the scene sound features and style control parameters are fused based on the Fusion Mechanism to obtain comprehensive style adjustment parameters, that is, a control parameter including scene information and speech style in the comprehensive style adjustment parameters.
[0078] Among them, the fusion mechanism refers to the process of effectively combining features from different sources or modalities (for example, scene sound features and style control parameters) through specific strategies or mathematical models to generate unified comprehensive style adjustment parameters; the fusion mechanism can be methods such as weighted linear fusion (that is, linearly superimposing after assigning weights to scene sound features and style parameters), neural network fusion (that is, using fully connected layers, LSTM or Transformer to encode feature interactions), and attention mechanism (that is, dynamically allocating feature importance through attention weights).
[0079] Dynamically adjusting the synthetic speech parameters of the text content according to the comprehensive style adjustment parameters can be to dynamically adjust the synthetic speech parameters of the text content using a deep learning model (for example, a time series model based on LSTM or GRU). For example, when the current scene is a "news broadcast" scene, the clarity and formality of the synthetic speech will be enhanced; when the second text contains "sad" emotions, the pitch will be lowered and the speech rate will be reduced, so that the generated speech sounds more in line with the emotional tendency.
[0080] Refer to Figure 4 As shown, the method of step S160 above can be implemented based on an adaptive regulation module, Figure 4 which is a schematic diagram of an adaptive regulation module provided by the embodiments of the present application. Among them, the adaptive regulation module includes: a fusion mechanism and a parameter adjustment model.
[0081] Specifically, the process of adjusting the synthetic speech parameters according to the scene sound features and style control parameters is as follows: based on the fusion mechanism, fuse the scene sound features from the scene perception module and the style control parameters from the natural language description parsing module to obtain a comprehensive style adjustment parameter; then, through the parameter adjustment model, dynamically adjust the synthetic speech parameters of the text content to obtain synthetic speech.
[0082] In this way, by combining the outputs of the scene perception module and the natural language description parsing module through the adaptive regulation module, the style and speech characteristics of speech synthesis are intelligently adjusted; this module can dynamically adjust features such as pitch, speech rate, and tone of speech synthesis according to user needs, scene changes, and emotional feedback, ensuring the personalization and situational adaptability of speech output.
[0083] Based on the same inventive concept, an embodiment of the present application further provides a personalized speech synthesis device based on scene perception and natural language description. Refer to Figure 5 as shown in Figure 5 FIG. 5 is a schematic structural diagram of a personalized speech synthesis device based on scene perception and natural language description provided by an embodiment of the present application. The device includes: A first analysis module 510, configured to perform semantic analysis on a first text to obtain a first semantic feature, where the first text includes text content and a scene description; A first generation module 520, configured to generate a scene classification based on the first semantic feature to obtain a scene vector, where the scene vector represents a scene category; A second generation module 530, configured to generate a scene sound feature based on the scene vector, where the scene sound feature includes detailed features of the application scenario and audio features related to the scene; A second analysis module 540, configured to perform semantic analysis on a second text to obtain a second semantic feature, where the second text includes at least a natural language description of a speech style; A first extraction module 550, configured to perform emotion category prediction and style feature extraction based on the second semantic feature to obtain a style control parameter, where the style control parameter includes: pitch, speech rate, tone, and volume; A first adjustment module 560, configured to adjust the synthesis speech parameter of the text content according to the scene sound feature and the style control parameter to obtain a synthesized speech.
[0084] In an optional embodiment, the first adjustment module includes: A parameter fusion module, configured to fuse the scene sound feature and the style control parameter to obtain a comprehensive style adjustment parameter; A parameter adjustment module, configured to dynamically adjust the synthesis speech parameter of the text content according to the comprehensive style adjustment parameter to obtain a synthesized speech.
[0085] In an optional embodiment, the first generation module includes: A scene sound template matching module, configured to match the scene vector with each scene sound template in a scene template library to obtain a target scene sound template that matches the scene vector; A scene sound feature generation module, configured to generate a scene sound feature according to the target scene sound template and the parameters of a speech synthesis model.
[0086] In an optional embodiment, the first analysis module includes: A text processing module for performing word segmentation and annotation processing on the first text through a natural language processing model to obtain the first text after word segmentation and annotation; A semantic encoding module for performing semantic encoding on the first text after word segmentation and annotation through a text encoder to obtain a first semantic feature.
[0087] In an optional embodiment, the first extraction module includes: An emotion category prediction module for predicting an emotion category of the second semantic feature through an emotion analysis model to obtain a first voice control parameter, where the first voice control parameter includes a first pitch, a first speech rate, and a volume; A tone recognition module for recognizing a tone of the second semantic feature through a tone recognition model to obtain a second voice control parameter, where the second voice control parameter includes a tone; A style conversion module for performing style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter, where the third voice control parameter includes a second pitch and a second speech rate; A parameter integration module for integrating the first voice control parameter, the second voice control parameter, and the third voice control parameter to obtain a style control parameter.
[0088] In an optional embodiment, the emotion category prediction module is specifically configured to: predict an emotion category of the second semantic feature through the emotion analysis model to obtain a target emotion category label; and map the target emotion category label to the first voice parameter according to a mapping relationship between the emotion category label and the voice parameter.
[0089] In an optional embodiment, the style conversion module is specifically configured to: process the second semantic feature through the style conversion model to obtain a style vector; match the style vector with each style template in a style library to obtain a target style template that matches the style vector; and use the voice control parameter corresponding to the target style template as the third voice control parameter.
[0090] In an optional embodiment, the second text further includes context information of the text content, and the second analysis module includes: A text parsing module for parsing and encoding the second text through a language parsing model to obtain a second semantic feature.
[0091] The embodiment of the present application further provides an electronic device. Refer to Figure 6 , Figure 6 is a schematic structural diagram of an electronic device provided by the embodiment of the present application. As Figure 6As shown in the figure, the electronic device 600 includes: a memory 610 and a processor 620. The memory 610 and the processor 620 are communicatively connected via a bus. A computer program is stored in the memory 610, and the computer program can run on the processor 620, thereby implementing the steps of the personalized speech synthesis method based on scene perception and natural language description described in the embodiments of the present application.
[0092] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the personalized speech synthesis method based on scene perception and natural language description described in the embodiments of the present application are implemented.
[0093] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the personalized speech synthesis method based on scene perception and natural language description described in the embodiments of the present application are implemented.
[0094] The embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0095] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.
[0098] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0099] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0100] The above has introduced in detail a personalized speech synthesis method, device and equipment based on scene perception and natural language description provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A personalized speech synthesis method based on scene perception and natural language description, characterized in that: The method comprises: Performing semantic analysis on a first text to obtain a first semantic feature, wherein the first text includes text content and a scene description; Perform scene classification and generation according to the first semantic feature to obtain a scene vector, wherein the scene vector represents a scene category; Generate a scene sound feature according to the scene vector, wherein the scene sound feature includes detailed features of the application scene and audio features related to the scene; Performing semantic analysis on the second text to obtain a second semantic feature, wherein the second text at least includes a natural language description of the voice style; Performing emotion category prediction and style feature extraction according to the second semantic feature to obtain style control parameters, wherein the style control parameters include: pitch, speech speed, tone and volume; According to the scene sound features and the style control parameters, the synthesized speech parameters of the text content are adjusted to obtain synthesized speech.
2. The method according to claim 1, characterized in that According to the scene sound feature and the style control parameter, adjusting the synthesized speech parameter of the text content to obtain the synthesized speech includes: The scene sound feature and the style control parameter are integrated to obtain a comprehensive style adjustment parameter; According to the comprehensive style adjustment parameters, the synthesized speech parameters of the text content are dynamically adjusted to obtain synthesized speech.
3. The method according to claim 1 or 2, characterized in that: Generating a scene sound feature according to the scene vector includes: Matching the scene vector with each scene sound template in a scene template library to obtain a target scene sound template that matches the scene vector; A scene sound feature is generated according to the target scene sound template and the parameters of the speech synthesis model.
4. The method according to claim 3, characterized in that Performing semantic analysis on the first text to obtain a first semantic feature includes: Performing word segmentation and tagging processing on the first text through a natural language processing model to obtain a first text that has been word segmented and annotated; The first text that has been segmented and annotated is semantically encoded by a text encoder to obtain a first semantic feature.
5. The method according to claim 1 or 2, characterized in that: According to the second semantic feature, sentiment category prediction and style feature extraction are performed to obtain style control parameters, including: Performing emotion category prediction on the second semantic feature through an emotion analysis model to obtain first voice control parameters, where the first voice control parameters include a first pitch, a first speech speed, and a volume; Performing tone recognition on the second semantic feature through a tone recognition model to obtain a second voice control parameter, wherein the second voice control parameter includes the tone; Performing style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter, wherein the third voice control parameter includes a second pitch and a second speech rate; The first voice control parameter, the second voice control parameter and the third voice control parameter are integrated to obtain a style control parameter.
6. The method according to claim 5, characterized in that The emotion category prediction is performed on the second semantic feature by using an emotion analysis model to obtain a first voice control parameter, including: Performing sentiment category prediction on the second semantic feature by using the sentiment analysis model to obtain a target sentiment category label; According to the mapping relationship between the emotion category label and the speech parameter, the target emotion category label is mapped to the first speech parameter.
7. The method according to claim 5, characterized in that Performing style conversion on the second semantic feature through a style conversion model to obtain a third voice control parameter includes: Processing the second semantic feature through the style conversion model to obtain a style vector; Matching the style vector with each style template in the style library to obtain a target style template that matches the style vector; The voice control parameter corresponding to the target style template is used as the third voice control parameter.
8. The method according to claim 5, characterized in that The second text also includes context information of the text content; performing semantic analysis on the second text to obtain a second semantic feature includes: The second text is parsed and encoded through a language parsing model to obtain a second semantic feature.
9. A personalized speech synthesis device based on scene perception and natural language description, characterized in that: The device comprises: A first analysis module, configured to perform semantic analysis on a first text to obtain a first semantic feature, wherein the first text includes text content and a scene description; A first generating module, configured to generate a scene classification according to the first semantic feature to obtain a scene vector, wherein the scene vector represents a scene category; A second generating module, configured to generate a scene sound feature according to the scene vector, wherein the scene sound feature includes a detailed feature of the application scene and an audio feature related to the scene; A second analysis module is used to perform semantic analysis on a second text to obtain a second semantic feature, wherein the second text at least includes a natural language description of the voice style; A first extraction module, configured to perform emotion category prediction and style feature extraction according to the second semantic feature to obtain style control parameters, wherein the style control parameters include: pitch, speech speed, tone and volume; The first adjustment module is used to adjust the synthesized speech parameters of the text content according to the scene sound characteristics and the style control parameters to obtain synthesized speech.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the personalized speech synthesis method based on scene perception and natural language description described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Voice synthesis method and device, computer readable medium and electronic equipment
CN111292719A
Voice synthesis method and device, synthesis model training method and device, medium and equipment
CN112786006A
Speech synthesis method, speech synthesis device, electronic equipment and storage medium
CN116580691A
Speech synthesis method and device, electronic equipment and storage medium
CN117524190A
Voice generation method, virtual human voice generation method and voice generation system
CN118942443A
Cited By
Personalized sound cloning method and device based on dynamic adaptation, equipment and medium
CN120748363A
Audio rendering method and device based on speech synthesis, equipment and storage medium
CN121617383A