Speech synthesis scheme and device, electronic equipment, storage medium and program product

Through the combination of the LLM semantic understanding module and neural codec, deep semantic features and multimodal context semantic information are obtained, and high-quality pronunciation features and parameters are generated through acoustic modeling and emotional rhythm control modules, which solves the problem of insufficient pronunciation and emotional expression in speech synthesis and achieves high-quality and efficient pronunciation synthesis.

CN120126445APending Publication Date: 2025-06-10SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510210087.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The speech emotion expression of speech synthesis in the prior art is insufficient, which limits the further development of speech synthesis technology.

Method used

Three-layer semantic analysis (word-level, sentence-level, and chapter-level) is performed through the LLM semantic understanding module, combining the context perception mechanism and the multi-grained feature extraction mechanism to obtain deep semantic features and multimodal context semantic information. These features are then transferred to the neural codec for acoustic feature compression, and high-precision acoustic features are generated through the acoustic modeling module. Finally, in the emotional rhythm control module, high-precision acoustic features and multimodal context semantic information are combined to generate voice parameters with emotion and rhythm annotations, and efficient timbre migration is achieved through the tone migration module.

Benefits of technology

It significantly improves the naturalness and expressiveness of speech synthesis, and solves problems such as lack of semantic understanding, low encoding and decoding efficiency, limited acoustic modeling capabilities, unnatural emotional expression, and unstable tone transfer quality in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126445A_ABST
    Figure CN120126445A_ABST
Patent Text Reader

Abstract

The invention provides a voice synthesis scheme and device, electronic equipment, a storage medium and a program product, and relates to the technical field of voice processing, and the method comprises the steps: inputting a text content into an LLM semantic understanding module, and obtaining a deep semantic feature and multi-modal context semantic information corresponding to the text content; transmitting the deep semantic features to a neural codec, and outputting compressed acoustic features; inputting the compressed acoustic features into an acoustic modeling module, and outputting high-precision acoustic features; inputting the high-precision acoustic features and the multi-mode context semantic information into an emotion and rhythm control module, and inputting voice parameters with emotion and rhythm marks; and inputting the voice parameters with the emotion and rhythm marks and the reference audio into a tone migration module to obtain a synthetic voice of a tone corresponding to the reference audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a speech synthesis solution, device, electronic device, storage medium, and program product. Background Art

[0002] With the rapid development of artificial intelligence technology, speech synthesis technology has been widely used in fields such as intelligent customer service, voice assistants, radio dubbing, and audiobooks, becoming an important interface for human-computer interaction. The introduction of deep learning has significantly improved the naturalness and expressiveness of speech synthesis.

[0003] However, in related technologies, the lack of emotional expression in the synthesized speech has been restricting the further development of speech synthesis technology. Therefore, how to better perform speech synthesis has become an urgent problem in the industry. Summary of the Invention

[0004] The present invention provides a speech synthesis solution, device, electronic device, storage medium, and program product to solve the problem of how to better perform speech synthesis in the prior art.

[0005] The present invention provides a speech synthesis method, including: Inputting the text content into the LLM semantic understanding module to obtain the deep semantic features corresponding to the text content and multi-modal context semantic information; Transmitting the deep semantic features to the neural codec to output compressed acoustic features; Inputting the compressed acoustic features into the acoustic modeling module to output high-precision acoustic features; Inputting the high-precision acoustic features and the multi-modal context semantic information into the emotion prosody control module to input speech parameters with emotion and prosody annotations; Inputting the speech parameters with emotion and prosody annotations and a reference audio into the timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio.

[0006] According to the speech synthesis method provided by the present invention, the step of inputting the text content into the LLM semantic understanding module to obtain the deep semantic features corresponding to the text content and multi-modal context semantic information includes: Performing word-level semantic parsing on the text content according to the word-level semantic layer in the LLM semantic understanding module to obtain word-level semantic features; Performing tone and context information parsing on the word-level semantic features according to the sentence-level semantic layer in the LLM semantic understanding module to obtain sentence-level semantic features; According to the discourse semantic layer in the LLM semantic understanding module, parse the overall structure and logical relationship of the text for the sentence-level semantic features to obtain discourse semantic information; Based on the word-level semantic features, sentence-level semantic features, and the discourse semantic information, determine the deep semantic features corresponding to the text content; Analyze the temporal context, emotional context, and topic context of the text content through a context awareness mechanism to obtain multi-modal context semantic information.

[0007] According to a speech synthesis method provided by the present invention, transmit the deep semantic features to a neural codec to output compressed acoustic features, including: Input the deep semantic features into the harmonic layer encoding unit in the neural codec for fundamental frequency detection, harmonic analysis, and feature encoding to obtain a first feature vector containing fundamental frequency and harmonic structure information; Transmit the first feature vector to the resonance layer encoding unit in the neural codec for formant extraction, vocal tract parameter estimation, and feature compression processing to obtain a resonance feature vector; Transmit the resonance feature vector to the noise layer encoding unit in the neural codec for noise separation feature quantization and coding compression to obtain a second feature vector containing aperiodic component information; Input the second feature vector into a decoder for decoding to obtain compressed acoustic features.

[0008] According to a speech synthesis method provided by the present invention, input the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features, including: Use a multi-scale window through the local attention unit in the acoustic modeling module to capture local acoustic features; Extract global acoustic features through the global attention unit in the acoustic modeling module; Based on the position encoding unit in the acoustic modeling module, explicitly model the temporal position into the local acoustic features and global acoustic features to obtain encoded features; Determine the high-precision acoustic features according to the encoded features.

[0009] According to a speech synthesis method provided by the present invention, input the high-precision acoustic features and the multi-modal context semantic information into an emotion prosody control module to input speech parameters with emotion and prosody annotations, including: The emotion analysis module in the emotion prosody control module performs explicit emotion analysis, implicit emotion analysis, and mixed emotion analysis on the high-precision acoustic features and the multi-modal context semantic information respectively to obtain emotion annotation information; The prosody prediction module in the emotional prosody control module performs pitch prediction, phoneme duration prediction, and phoneme energy prediction on the high-precision acoustic features to obtain prosody annotation information; Based on the emotional annotation information and the prosody annotation information, speech parameters with emotional and prosody annotations are obtained.

[0010] According to a speech synthesis method provided by the present invention, the speech parameters with emotional and prosody annotations and a reference audio are input into a timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio, including: The timbre encoding unit in the timbre transfer module extracts timbre features, speaking style features, and audio personalization features from the reference audio to obtain the target timbre features of the reference audio; Through the fast adaptation module in the timbre transfer module, the target timbre features are mapped into the speech parameters with emotional and prosody annotations to obtain a synthesized speech with the timbre corresponding to the reference audio; Among them, the contrast learning module in the timbre transfer module is used to improve the accuracy and robustness of the timbre transfer of the timbre transfer module.

[0011] The present invention also provides a speech synthesis device, including: A semantic understanding module for inputting text content into an LLM semantic understanding module to obtain deep semantic features and multi-modal context semantic information corresponding to the text content; An encoding and decoding module for transmitting the deep semantic features to a neural codec and outputting compressed acoustic features; A modeling module for inputting the compressed acoustic features into an acoustic modeling module and outputting high-precision acoustic features; A prosody control module for inputting the high-precision acoustic features and the multi-modal context semantic information into an emotional prosody control module and inputting speech parameters with emotional and prosody annotations; A synthesis module for inputting the speech parameters with emotional and prosody annotations and a reference audio into a timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the speech synthesis method described in any one of the above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the speech synthesis method described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned speech synthesis methods.

[0015] The speech synthesis scheme, device, electronic device, storage medium and program product provided by the present invention can deeply mine the deep semantic features of the text through the three-layer semantic parsing structure (word level, sentence level, and chapter level) of the LLM semantic understanding module, and at the same time combine the context perception mechanism and the multi-granularity feature extraction mechanism to obtain multimodal contextual semantic information. This design not only improves the depth and breadth of semantic understanding, but also solves the problems of shallow semantic understanding and insufficient use of context information in traditional methods. In the acoustic feature compression stage, the neural codec module adopts a hierarchical coding structure (harmonic layer, resonance layer, and noise layer) to efficiently compress deep semantic features and output compressed acoustic features. Through the improved decoder structure and progressive training scheme, the reconstruction quality is further optimized, and the problems of low efficiency and poor reconstruction effect of traditional codecs are solved. Then, in the acoustic modeling stage, the adaptive acoustic modeling module combines the hybrid attention mechanism and the dynamic convolutional network to improve the feature extraction efficiency and modeling ability, and output high-precision acoustic features. This design significantly enhances the accuracy and robustness of acoustic modeling, making up for the limited acoustic modeling ability in the prior art. In the emotion and prosody control stage, the emotion and prosody control module generates speech parameters with emotion and prosody annotations through the emotion analysis network and the precise prosody prediction network driven by the large model. The emotion analysis network combines the emotion feature extraction network and the emotion dictionary enhancement mechanism to achieve accurate emotion analysis; the prosody prediction network generates natural emotion expression and prosody changes through the multi-task prosody predictor and the long-range dependency modeling mechanism. This step effectively solves the problems of unnatural emotion expression and imprecise prosody control in traditional methods. Finally, in the timbre transfer and speech synthesis stage, the efficient timbre transfer module achieves efficient timbre transfer based on the timbre encoding network driven by StyleGAN and the contrastive learning optimization strategy. The quality and stability of timbre transfer are ensured through the hierarchical style encoding structure, the adaptive style mixing mechanism and the triple contrastive learning framework. At the same time, the fast timbre adaptation network achieves fast adaptation and high-quality synthesis of timbre through the lightweight adaptation module and the progressive training strategy. It solves the problems of unstable timbre transfer quality and low adaptation efficiency in the existing technology. Therefore, this application systematically solves problems such as shallow semantic understanding, low encoding and decoding efficiency, limited acoustic modeling capabilities, unnatural emotional expression, and unstable timbre migration quality through modular design and optimization strategies, and achieves high-quality and efficient speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the speech synthesis method provided by the present invention; Figure 2 It is the overall architecture diagram provided by the present invention; Figure 3 It is a schematic structural diagram of the semantic understanding module provided by the embodiments of the present invention; Figure 4 It is a schematic diagram of the acoustic modeling module provided by the present invention; Figure 5 It is a schematic diagram of the emotion prosody control module provided by the present invention; Figure 6 It is a schematic structural diagram of the voice conversion module provided by the present invention; Figure 7 It is a schematic structural diagram of the speech synthesis device provided by the present invention; Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0018] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0019] Figure 1 It is a schematic flowchart of the speech synthesis method provided by the present invention. As Figure 1 shown, the method includes the following: Step 110, input the text content into the LLM semantic understanding module to obtain the deep semantic features corresponding to the text content and multi-modal context semantic information; In the present invention, the text content can be any form of input, such as a sentence, a paragraph or a dialogue. These text contents are sent to the LLM semantic understanding module for processing. The text content may contain complex language structures, implicit semantic information and context dependencies, all of which need to be accurately understood and extracted by the module.

[0020] The LLM semantic understanding module adopts a multi-layer semantic parsing structure, mainly including the following three levels: Word-level semantic layer: Parse each word in the text to extract its basic features such as phonemes, tones, and part of speech. This layer is implemented through a deep neural network (such as Transformer) and can capture the local semantic information of words.

[0021] Sentence-level semantic layer: Analyze the grammatical structure, mood, and context information of sentences. This layer uses the attention mechanism and context modeling to understand the semantic relationships within sentences, such as the subject-predicate-object structure and modification relationships.

[0022] Discourse-level semantic layer: Understand the overall structure and theme information of the text from a global perspective. This layer uses a long-range dependence modeling mechanism to capture the logical relationships, emotional tendencies, and context coherence in the text.

[0023] Deep semantic features refer to the implicit and high-level semantic information in the text, such as emotional tendencies, intentions, themes, etc. Through the multi-layer semantic parsing structure of the LLM semantic understanding module, the text content is parsed and abstracted layer by layer, and finally a high-dimensional semantic vector is generated to represent the deep semantic features of the text.

[0024] These deep semantic features not only contain the literal meaning of the text but also cover its implicit semantic information, providing a rich semantic basis for subsequent speech synthesis.

[0025] In the present invention, multi-modal context semantic information refers to the context information related to the text content, including dialogue history, emotional changes, scene information, etc.

[0026] The LLM semantic understanding module records and analyzes the context information of the text through a context awareness mechanism. For example, in a dialogue scenario, the module will record the historical state, emotional changes, and dialogue focus of the dialogue, thereby generating context semantic information related to the current text.

[0027] In addition, the module also extracts features from multiple granularities such as the phoneme level, word level, and sentence level through a multi-granularity feature extraction mechanism, and dynamically adjusts them in combination with the context information to ensure that the generated semantic information has high accuracy and adaptability.

[0028] More specifically, in order to further improve the accuracy and generation effect of semantic understanding, the LLM semantic understanding module introduces a hybrid retrieval enhanced generation mechanism.

[0029] This mechanism retrieves relevant semantic information from a pre-trained knowledge base through the retrieval paths of semantic similarity and prosody patterns, and fuses it with the semantic features of the current text.

[0030] This retrieval-augmented generation mechanism can not only supplement the missing semantic information in the text, but also optimize the expression of semantic features, providing richer semantic support for subsequent speech synthesis.

[0031] Step 120: Transmit the deep semantic features to a neural codec to output compressed acoustic features. In the present invention, the deep semantic features are fed into the neural codec module as input. These semantic features are high-dimensional vectors extracted from the LLM semantic understanding module and contain the deep semantic information of the text. To convert these semantic features into acoustic features suitable for speech synthesis, the neural codec module adopts a hierarchical coding structure, decomposing the input features into three levels for processing: the harmonic layer, the resonance layer, and the noise layer. The harmonic layer is responsible for extracting the fundamental frequency and harmonic structure of the speech signal, the resonance layer models the vocal tract features, and the noise layer processes the aperiodic components in the speech. This hierarchical coding structure can capture the details of the speech signal more precisely, ensuring that the compressed acoustic features have high fidelity.

[0032] During the encoding process, the neural codec module reconstructs the features through an improved decoder structure. The improved decoder adopts multiple discriminators and a progressive training scheme to gradually optimize the reconstruction quality. Specifically, the discriminators evaluate the reconstructed acoustic features through an adversarial learning mechanism to ensure that they are as close as possible to the real speech features. At the same time, the progressive training scheme gradually improves the performance of the decoder through staged training, avoiding the performance bottleneck caused by one-time training in traditional methods. This design not only improves the efficiency of encoding and decoding, but also significantly enhances the quality of the reconstructed acoustic features.

[0033] In addition, the neural codec module also introduces a dynamic adjustment mechanism to adaptively adjust the encoding and decoding processes according to the characteristics of the input semantic features. For example, for text with rich emotions, the module will enhance the encoding capabilities of the harmonic layer and the resonance layer to better capture the emotional information in the speech; for speech signals with more noise, the module will strengthen the processing ability of the noise layer to ensure that the reconstructed acoustic features are clear and natural. This dynamic adjustment mechanism further improves the adaptability and robustness of the module, enabling it to handle diverse speech synthesis tasks.

[0034] The neural codec module outputs compressed acoustic features. These features not only retain the key details of the original semantic information, but also achieve high-quality feature representation through an efficient compression and reconstruction process. The compressed acoustic features provide a solid foundation for subsequent acoustic modeling and speech synthesis, significantly improving the naturalness and clarity of speech synthesis.

[0035] Step 130: Input the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features. In the present invention, the compressed acoustic features are fed into the acoustic modeling module as input. The core structure of the acoustic modeling module is an improved Conformer encoder, which combines a hybrid attention mechanism and a dynamic convolutional network, and can efficiently capture the temporal correlation and long-range dependence in the acoustic features. The hybrid attention mechanism enhances the module's ability to capture key information in the acoustic features through self-attention and cross-attention mechanisms, while the dynamic convolutional network further improves the efficiency of feature extraction through an adaptive convolutional kernel generation network and a feature extraction optimization network.

[0036] During the feature extraction process, the acoustic modeling module adopts a phased optimization strategy to gradually improve the accuracy of the acoustic features. First, the module preliminarily processes the input features through a shallow network to extract their basic information; then, it finely models the features through a deep network to capture their high-order information. This phased optimization strategy not only improves the efficiency of feature extraction but also avoids the performance bottleneck caused by one-time processing in traditional methods. In addition, the module also introduces a multi-scale feature fusion mechanism to fuse features at different levels, ensuring that the generated high-precision acoustic features contain both local details and global consistency.

[0037] To further enhance the robustness of the acoustic modeling, the module also introduces a dynamic adjustment mechanism. This mechanism adaptively adjusts the modeling process according to the characteristics of the input acoustic features. For example, for a speech signal rich in emotion, the module will enhance the modeling ability for harmonic features and resonance features to better capture the emotional information in the speech; for a speech signal with more noise, the module will strengthen the modeling ability for noise features to ensure that the generated high-precision acoustic features are clear and natural. This dynamic adjustment mechanism further improves the adaptability and robustness of the module, enabling it to handle diverse speech synthesis tasks.

[0038] The acoustic modeling module outputs high-precision acoustic features. These features not only retain the key details of the input acoustic information but also achieve a higher-quality feature representation through an efficient modeling and optimization process. The high-precision acoustic features provide a solid foundation for subsequent emotion prosody control and speech synthesis, significantly improving the naturalness and clarity of speech synthesis.

[0039] Step 140, input the high-precision acoustic features and the multi-modal context semantic information into the emotion prosody control module, and input the speech parameters with emotion and prosody annotations; In the present invention, high-precision acoustic features and multi-modal context semantic information are input into the module as inputs. The high-precision acoustic features are obtained from the previous acoustic modeling module, which contains detailed acoustic information of the speech, such as pitch, duration, intensity, etc. The multi-modal context semantic information comes from the LLM semantic understanding module, which provides the deep meaning and context background of the text, such as whether this sentence is happy, sad, interrogative, or declarative. The combination of these two provides a solid foundation for the generation of emotion and prosody.

[0040] The present invention captures the emotion in the text through an emotion analysis network. Identify the emotional tendency from the text. For example, if the text expresses an excited emotion, the module will capture the keywords and tone, and then generate corresponding emotion annotations to make the speech sound full of vitality. If it is a sad text, the module will adjust the intonation of the speech to make it sound low and slow, conveying a sad feeling.

[0041] At the same time, the module will also generate natural rhythm and intonation for the speech through a prosody prediction network. Prosody is a very important part of speech, which determines the fluency and expressiveness of speech. For example, in an interrogative sentence, the module will automatically raise the tone at the end of the sentence to make it sound more like a natural interrogative sentence; in a declarative sentence, the module will maintain a steady intonation to make the speech sound more natural. The prosody prediction network will analyze the grammatical structure and semantic information of the text, and generate prosody annotations that match the emotion to ensure that the rhythm and intonation of the speech are just right.

[0042] To make the generation of emotion and prosody more accurate, the module will also dynamically adjust the intensity of emotion and prosody according to the characteristics of the high-precision acoustic features. For example, in a scene with strong emotion, the module will enhance the emotional expression of the speech to make it sound more vivid; while in a scene with gentle emotion, the module will reduce the intensity of emotion to ensure that the speech sounds natural and not exaggerated.

[0043] Step 150, input the speech parameters with emotion and prosody annotations, and the reference audio into the timbre transfer module to obtain the synthesized speech with the timbre corresponding to the reference audio.

[0044] In the present invention, the speech parameters with emotion and prosody annotations and the reference audio are input into the timbre transfer module as inputs. The speech parameters are output from the emotion and prosody control module, which contains the acoustic information of the speech as well as the details of emotion and prosody. The reference audio is the source of the target timbre, which may be a real speech sample of a person, or a specific sound style. The task of the module is to transfer the timbre characteristics of the reference audio to the synthesized speech while retaining the emotion and prosody information in the speech parameters.

[0045] The timbre transfer module can refer to a StyleGAN-driven timbre encoding network that can extract timbre features from the reference audio. This network is like a timbre extractor that can capture the sound style in the reference audio, such as the brightness, thickness, and sound quality of the timbre. Through a hierarchical style encoding structure and an adaptive style mixing mechanism, the module can efficiently transfer these timbre features to the synthesized speech to ensure that the timbre of the synthesized speech is consistent with the reference audio.

[0046] To make the timbre transfer more accurate, the module also introduces a contrastive learning optimization strategy. This strategy ensures the quality and stability of timbre transfer through a triple contrastive learning framework and a hard example mining mechanism. Simply put, the module will compare the timbre features of the synthesized speech and the reference audio, continuously optimize the transfer process, and ensure that the timbre of the synthesized speech is as close as possible to the reference audio. At the same time, the hard example mining mechanism will focus on those timbre features that are difficult to transfer to further improve the transfer effect.

[0047] In the present invention, through modular design and optimization strategies, problems such as shallow semantic understanding, low encoding and decoding efficiency, limited acoustic modeling ability, unnatural emotional expression, and unstable timbre transfer quality are systematically solved, achieving high-quality and high-efficiency speech synthesis.

[0048] Optionally, inputting the text content into the LLM semantic understanding module to obtain the deep semantic features and multi-modal context semantic information corresponding to the text content includes: Performing word-level semantic parsing on the text content according to the word-level semantic layer in the LLM semantic understanding module to obtain word-level semantic features; Performing tone and context information parsing on the word-level semantic features according to the sentence-level semantic layer in the LLM semantic understanding module to obtain sentence-level semantic features; Performing overall structure and logical relationship parsing of the text on the sentence-level semantic features according to the discourse semantic layer in the LLM semantic understanding module to obtain discourse semantic information; Based on the word-level semantic features, sentence-level semantic features, and the discourse semantic information, determining the deep semantic features corresponding to the text content; Analyzing the temporal context, emotional context, and topic context of the text content through a context awareness mechanism to obtain multi-modal context semantic information.

[0049] In the present invention, the text content is sent to the word-level semantic layer for parsing. The main task of this layer is to perform semantic analysis on each word in the text, extract its basic features, such as phonemes, tones, part of speech, etc. Through a deep neural network (such as Transformer), the word-level semantic layer can capture the local semantic information of words and generate word-level semantic features. These features provide the basis for subsequent semantic parsing. For example, in the "word-level semantic layer in the LLM semantic understanding module" mentioned in the document, the text content is initially parsed in this way to extract the semantic features of each word.

[0050] More specifically, the word-level semantic features are passed to the sentence-level semantic layer for further parsing. The task of the sentence-level semantic layer is to analyze the grammatical structure, mood, and context information of the sentence. Through the attention mechanism and context modeling, the sentence-level semantic layer can understand the semantic relationships within the sentence, such as the subject-predicate-object structure, modification relationships, etc., and generate sentence-level semantic features. These features not only contain the literal meaning of the sentence but also capture its implicit mood and context information. The "sentence-level semantic layer performs mood and context information parsing on the word-level semantic features" mentioned in the document is the specific implementation of this step.

[0051] On the other hand, the sentence-level semantic features are sent to the discourse semantic layer for global parsing. The task of the discourse semantic layer is to understand the structure and logical relationships of the text from an overall perspective. Through the long-range dependence modeling mechanism, the discourse semantic layer can capture the theme information, logical coherence, and emotional tendency in the text and generate discourse semantic information. These information provide a global perspective for the deep semantic understanding of the text. The "discourse semantic layer performs parsing on the overall structure and logical relationships of the text for the sentence-level semantic features" mentioned in the document is the specific description of this step.

[0052] Based on the word-level semantic features, sentence-level semantic features, and discourse semantic information, the LLM semantic understanding module further integrates these features to determine the deep semantic features corresponding to the text content. The deep semantic features are a high-dimensional vector that synthesizes the literal meaning, implicit semantics, and global logical relationships of the text, providing a rich semantic basis for subsequent speech synthesis.

[0053] Optionally, the deep semantic features are transmitted to a neural codec to output compressed acoustic features, including: Input the deep semantic features into the harmonic layer encoding unit in the neural codec for fundamental frequency detection, harmonic analysis, and feature encoding to obtain a first feature vector containing fundamental frequency and harmonic structure information; Transmit the first feature vector to the resonance layer encoding unit in the neural codec for formant extraction, vocal tract parameter estimation, and feature compression processing to obtain a resonance feature vector; Transmit the resonance feature vector to the noise layer encoding unit in the neural codec for noise separation feature quantization and encoding compression to obtain a second feature vector containing aperiodic component information; Input the second feature vector into the decoder for decoding to obtain the compressed acoustic features.

[0054] In the present invention, the deep semantic features are input into the harmonic layer encoding unit in the neural codec. The main task of the harmonic layer encoding unit is to analyze and encode the fundamental frequency and harmonic structure of the speech signal. Through fundamental frequency detection and harmonic analysis, the module can extract the periodic components in the speech signal and generate a first feature vector containing fundamental frequency and harmonic structure information. This step ensures that the harmonic components in the speech signal are accurately captured and provides a basis for subsequent encoding.

[0055] In the present invention, the first feature vector is transmitted to the resonance layer encoding unit in the neural codec. The task of the resonance layer encoding unit is to extract and compress the formants and vocal tract features of the speech signal. Through formant extraction and vocal tract parameter estimation, the module can capture the vocal tract features in the speech signal and generate a resonance feature vector. This step further compresses the features of the speech signal while retaining the key vocal tract information.

[0056] In the present invention, the resonance feature vector is transmitted to the noise layer encoding unit in the neural codec. The main task of the noise layer encoding unit is to separate and encode the aperiodic components in the speech signal. Through noise separation and feature quantization, the module can extract the noise components in the speech signal and generate a second feature vector containing aperiodic component information. This step ensures that the aperiodic components in the speech signal are accurately captured and provides comprehensive feature information for subsequent decoding.

[0057] In the present invention, the second feature vector is input into the decoder for decoding. The task of the decoder is to restore the compressed feature vector to acoustic features. Through a hierarchical decoding structure, the module can efficiently reconstruct the acoustic features while retaining the key information of the original speech signal. The decoder adopts an improved decoder structure and a progressive training scheme to gradually optimize the reconstruction quality and ensure that the generated acoustic features have high fidelity.

[0058] Optionally, inputting the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features includes: Utilize a multi-scale window through the local attention unit in the acoustic modeling module to capture local acoustic features; Perform global acoustic feature extraction through the global attention unit in the acoustic modeling module; Based on the position encoding unit in the acoustic modeling module, the temporal position is explicitly modeled into the local acoustic features and the global acoustic features to obtain encoded features; Determine the high-precision acoustic features according to the encoded features.

[0059] In the present invention, the compressed acoustic features are input into the local attention unit in the acoustic modeling module. The main task of the local attention unit is to capture the local detailed information in the acoustic features. Through the multi-scale window mechanism, the module can analyze the acoustic features at different scales and extract the local acoustic information. For example, for the short-term changes in the speech signal (such as the rapid switching of phonemes), the local attention unit can accurately capture these details to ensure that the generated acoustic features have a high-precision local performance.

[0060] The acoustic features are passed to the global attention unit in the acoustic modeling module. The task of the global attention unit is to extract the global information in the acoustic features from an overall perspective. Through the global attention mechanism, the module can capture the long-range dependencies in the acoustic features, such as the intonation changes and emotional expressions in the speech signal. The global attention unit not only focuses on the local details but also analyzes the acoustic features from a global perspective to ensure that the generated acoustic features have consistency and coherence.

[0061] In the present invention, the local acoustic features and the global acoustic features are fed into the position encoding unit in the acoustic modeling module. The main task of the position encoding unit is to explicitly model the temporal position information into the acoustic features. Through position encoding, the module can capture the sequential relationship of the acoustic features in time, such as the syllable order and intonation changes in the speech signal. The position encoding unit fuses the temporal information with the local and global acoustic features to generate encoded features. These encoded features not only contain the details of the acoustic information but also retain the sequential relationship in time, providing a basis for the subsequent generation of high-precision acoustic features.

[0062] Finally, based on the encoded features, the acoustic modeling module determines the high-precision acoustic features. Through the multi-layer neural network and the feature fusion mechanism, the module integrates the local acoustic features, the global acoustic features, and the temporal position information to generate high-precision acoustic features. These features not only retain the key details of the input acoustic information but also achieve a higher-quality feature representation through the local and global attention mechanisms and position encoding. The high-precision acoustic features provide a solid foundation for the subsequent speech synthesis, significantly improving the naturalness and clarity of the speech synthesis.

[0063] Optionally, inputting the high-precision acoustic features and the multi-modal context semantic information into the emotional prosody control module, and inputting the speech parameters with emotional and prosodic annotations, including: The sentiment analysis module in the sentiment prosody control module performs explicit sentiment analysis, implicit sentiment analysis, and hybrid sentiment analysis on the high-precision acoustic features and the multi-modal context semantic information respectively to obtain sentiment annotation information; The prosody prediction module in the sentiment prosody control module performs pitch prediction, phoneme duration prediction, and phoneme energy prediction on the high-precision acoustic features to obtain prosody annotation information; Based on the sentiment annotation information and the prosody annotation information, speech parameters with sentiment and prosody annotations are obtained.

[0064] In the present invention, the high-precision acoustic features and the multi-modal context semantic information are input into the sentiment analysis module in the sentiment prosody control module. The main task of the sentiment analysis module is to perform sentiment analysis on the input features and generate sentiment annotation information. The module achieves this goal through three analysis methods: explicit sentiment analysis, implicit sentiment analysis, and hybrid sentiment analysis. Explicit sentiment analysis directly extracts explicit sentiment words and tones from the text, such as "happy", "sad", etc.; implicit sentiment analysis captures the implicit sentiment tendency in the text through context semantic information, such as inferring sentiment through intonation, context, etc.; hybrid sentiment analysis combines the results of explicit and implicit sentiment analysis to generate comprehensive sentiment annotation information. For example, if the text expresses "I'm so happy!", explicit sentiment analysis will directly capture the sentiment of "happy", while implicit sentiment analysis will further confirm the intensity of the sentiment through the exclamation mark and tone, and finally generate accurate sentiment annotation information.

[0065] Next, the high-precision acoustic features are input into the prosody prediction module in the sentiment prosody control module. The main task of the prosody prediction module is to predict the prosody features of speech and generate prosody annotation information. The module achieves this goal through three key prediction tasks: pitch prediction, phoneme duration prediction, and phoneme energy prediction. Pitch prediction analyzes the fundamental frequency change of speech to generate intonation information; phoneme duration prediction analyzes the duration of each phoneme to generate rhythm information; phoneme energy prediction analyzes the energy distribution of speech to generate stress information. For example, in an interrogative sentence, the prosody prediction module will raise the pitch at the end of the sentence to make it sound more like a natural interrogative sentence; when emphasizing a certain word, the module will increase the energy and duration of the word to make it sound more prominent. Through these prediction tasks, the prosody prediction module generates prosody annotation information that matches the sentiment and semantics.

[0066] Finally, the sentiment annotation information generated by the sentiment analysis module and the prosody annotation information generated by the prosody prediction module are integrated to generate speech parameters with sentiment and prosody annotations. These speech parameters not only contain the acoustic information in the high-precision acoustic features, but also inject details of sentiment and prosody, making the synthesized speech sound more natural and expressive.

[0067] Optionally, input the speech parameters with emotion and prosody annotations, and the reference audio into the timbre transfer module to obtain the synthesized speech with the timbre corresponding to the reference audio, including: Extract the target timbre features of the reference audio through the timbre encoding unit in the timbre transfer module for timbre feature extraction, speaking style extraction, and audio personalization extraction of the reference audio; Through the fast adaptation module in the timbre transfer module, map the target timbre features into the speech parameters with emotion and prosody annotations to obtain the synthesized speech with the timbre corresponding to the reference audio; Among them, the contrast learning module in the timbre transfer module is used to improve the accuracy and robustness of the timbre transfer of the timbre transfer module.

[0068] In the present invention, in the process of inputting the speech parameters with emotion and prosody annotations and the reference audio into the timbre transfer module, the module generates a synthesized speech with the same timbre as the reference audio through timbre encoding, fast adaptation, and contrast learning. First, the reference audio is sent to the timbre encoding unit in the timbre transfer module, which extracts the timbre features of the reference audio through spectral analysis or a deep learning model (such as StyleGAN), such as the brightness, thickness, and sound quality of the timbre. At the same time, the timbre encoding unit also analyzes the speaking style of the reference audio, such as the speaking speed, intonation, and stress distribution, captures the personalized speaking manner, and extracts the personalized features in the reference audio, such as specific pronunciation habits or timbre details, to ensure the accuracy of timbre transfer. Finally, the timbre encoding unit generates the target timbre features, representing the timbre, style, and personalized information of the reference audio.

[0069] The target timbre features are passed to the fast adaptation module, which maps the target timbre features into the speech parameters with emotion and prosody annotations to generate a synthesized speech with the same timbre as the reference audio. The fast adaptation module fuses the target timbre features with the speech parameters through a lightweight neural network or an adapter model to ensure that the timbre of the synthesized speech is the same as that of the reference audio. At the same time, the module dynamically adjusts the timbre transfer process according to the emotion and prosody information in the speech parameters to ensure the naturalness and expressiveness of the synthesized speech. In addition, a progressive training strategy is adopted to gradually optimize the timbre transfer effect and avoid the performance bottleneck caused by one-time transfer. Finally, a synthesized speech with the same timbre as the reference audio is generated while retaining the emotion and prosody information in the speech parameters.

[0070] To improve the accuracy and robustness of timbre transfer, the timbre transfer module also introduces a contrastive learning module. The contrastive learning module optimizes the timbre transfer process by comparing the features of the synthesized speech, the reference audio, and the original speech through a triple contrastive learning framework, ensuring that the timbre of the synthesized speech is as close as possible to the reference audio. At the same time, the contrastive learning module focuses on the timbre features that are difficult to transfer through a hard example mining mechanism, and improves the robustness of timbre transfer through hard example mining and targeted optimization. In addition, the contrastive learning module reduces noise and distortion in the timbre transfer process through contrastive learning, ensuring the stability and high quality of the synthesized speech. Finally, the timbre transfer result optimized by the contrastive learning module ensures that the timbre of the synthesized speech is highly consistent with the reference audio.

[0071] In the embodiment of the present invention, the timbre transfer module realizes high-quality and high-efficiency timbre transfer, solves the problems of unstable timbre transfer quality and low adaptation efficiency in traditional methods, and provides key support for personalized speech synthesis. This process not only retains the emotional and prosodic information in the speech parameters, but also realizes high-quality synthesized speech with the same timbre as the reference audio, significantly improving the naturalness and expressiveness of speech synthesis.

[0072] Figure 2 The overall architecture diagram provided by the present invention is as Figure 2 shown, including: Text input: The system receives the text input by the user.

[0073] Preprocess the text by cleaning, tokenizing, annotating, etc., to prepare for subsequent semantic understanding.

[0074] Semantic understanding module: Word-level semantic layer, extract the phoneme and tone features in the words, and perform feature fusion through convolutional neural network and recurrent neural network.

[0075] Sentence-level semantic layer, analyze the tone and context information of the sentence, and use the attention mechanism and Transformer structure for feature integration.

[0076] Discourse-level semantic layer, understand the overall structure of the text, extract the theme information and discourse relationships, and perform semantic summarization.

[0077] Context awareness mechanism, record the dialogue history state, track emotional changes, and maintain the dialogue focus.

[0078] Multi-granularity feature extraction mechanism, extract phoneme-level, word-level, and sentence-level features, and realize dynamic adjustment of features.

[0079] Hybrid retrieval enhanced generation mechanism, retrieve through semantic similarity and prosodic patterns, and use the attention mechanism to perform weighted fusion on the results.

[0080] Neural codec module: Harmonic layer coding unit; extracts the fundamental frequency and harmonic structure of the speech signal.

[0081] Resonance layer coding unit; models vocal tract characteristics and extracts formants and vocal tract parameters.

[0082] Noise layer coding unit; processes aperiodic components, separates noise, and performs feature quantization.

[0083] Improved decoder structure; sets multiple discriminators to optimize the reconstruction quality in the time domain, frequency domain, and envelope domain, and adopts a progressive training scheme.

[0084] Adaptive acoustic modeling module: Improved Conformer encoder; includes local attention unit, global attention unit, and relative position encoding unit to enhance feature extraction efficiency and long-range dependence modeling ability.

[0085] Dynamic convolution network optimization; adaptive convolution kernel generation network, feature extraction optimization network, and network structure optimization unit to improve feature extraction and modeling ability.

[0086] Emotional prosody control module: Large model-driven emotion analysis network; emotion feature extraction network, emotion dictionary enhancement mechanism, and mixed emotion processing mechanism to achieve natural emotional expression.

[0087] Accurate prosody prediction network; multi-task prosody predictor and long-range dependence modeling mechanism to generate natural prosody features.

[0088] Efficient timbre transfer module: StyleGAN-driven timbre coding network; hierarchical style coding structure and adaptive style mixing mechanism to ensure timbre transfer quality.

[0089] Contrastive learning optimization strategy; triple contrastive learning framework and hard example mining mechanism to improve the robustness and adaptability of the model.

[0090] Fast timbre adaptation network; lightweight adaptation module and progressive training strategy to achieve efficient timbre adaptation.

[0091] The processed features are reconstructed into a speech signal through the decoder, and the generated speech is optimized and adjusted to ensure naturalness and expressiveness.

[0092] The speech synthesis device provided by the present invention will be described below. The speech synthesis device described below can be mutually referred to the speech synthesis method described above.

[0093] Figure 3 It is a schematic diagram of the semantic understanding module structure provided by the embodiment of the present invention, asFigure 3 As shown, in the semantic parsing part, there are three levels: Word level: It includes phoneme recognition and tone analysis. These information are integrated through feature fusion to capture the pronunciation and intonation features of individual words.

[0094] Sentence level: It involves mood recognition and context analysis. These information are synthesized through feature integration to understand the overall mood and context relationship of the sentence.

[0095] Discourse level: It includes topic modeling and semantic analysis. The semantic information of the entire discourse is integrated through semantic summarization to grasp the topic and overall meaning of the text macroscopically.

[0096] In the feature processing part, it includes context awareness and feature extraction: Context awareness: It is divided into temporal context, emotional context and topic context. These information are used to extract phoneme features, word-level features and sentence-level features.

[0097] Feature extraction: Convert semantic information at different levels into specific features to provide data support for subsequent processing.

[0098] The entire module gradually goes deeper in a hierarchical manner, from words to sentences and then to discourses. Finally, the semantic information is converted into processable data through feature extraction.

[0099] Figure 4 Schematic diagram of the acoustic modeling module provided by the present invention, as Figure 4 shown, includes: In the Conformer encoder, there are the following key components: Attention mechanism: It includes local attention and global attention, which are used to capture information at different scales.

[0100] Feature fusion: Combine the outputs of local and global attention to form a richer feature representation.

[0101] Position encoding: It includes hierarchical encoding, position awareness and dynamic update, which are used to retain hierarchical information and position information in the sequence.

[0102] In the dynamic convolutional network, the key steps include: Convolution kernel processing: Kernel size generation and kernel weight generation. The size and weight of the convolution kernel are dynamically generated to improve the flexibility of the model.

[0103] Dilation adjustment: Adjust the dilation rate of the convolution kernel to control the size of the receptive field and better capture features in different ranges.

[0104] Feature extraction: Extract useful features from the input data.

[0105] Feature selection: Select the most important features and remove redundant information.

[0106] Feature enhancement: Enhance the selected features to improve the performance of the model.

[0107] This system can effectively process and optimize features by combining the attention mechanism and the dynamic convolutional network, improving the performance of the model in speech recognition or other sequence modeling tasks.

[0108] Figure 5 Schematic diagram of the emotional prosody control module provided by the present invention, as Figure 5 shown, including: The emotional analysis network includes emotion recognition, emotion processing, intensity analysis, change tracking, conflict resolution, transition processing, and feature fusion. Emotion recognition detects the type of emotion, emotion processing further analyzes or adjusts the emotion, intensity analysis evaluates the intensity of the emotion, and change tracking monitors the change of emotion over time. Conflict resolution deals with the contradictions between emotions, transition processing ensures smooth emotional changes, and feature fusion integrates emotional features into a unified representation.

[0109] The prosody prediction network includes F0 contour generation, phoneme duration, energy prediction, long-range modeling, pattern memory, and optimization mechanism. F0 contour generation generates the fundamental frequency change curve, phoneme duration predicts the duration of each phoneme, and energy prediction involves the loudness or intensity of speech. Long-range modeling considers the prosody patterns in a long time range, pattern memory stores and recalls common prosody patterns, and the optimization mechanism adjusts and optimizes the prediction to ensure naturalness and accuracy.

[0110] The emotional analysis network processes emotional features, and the prosody prediction network converts these emotional features into specific speech parameters such as pitch, duration, and energy. In this way, the entire module generates corresponding emotional speech according to the input emotional information. Conflict resolution and transition processing ensure natural emotional changes, and long-range modeling and pattern memory guarantee the coherence and diversity of prosody.

[0111] Figure 6 Schematic diagram of the timbre transfer module structure provided by the present invention, as Figure 6 shown, including: Adapter module: Responsible for quickly adjusting the general speech model to the needs of a specific user or scenario.

[0112] Reference encoding: Extract or generate a reference encoding to guide the adaptation process of the model.

[0113] Mapping network: Maps the reference encoding to the parameter space of the model to adjust the model.

[0114] Synthesis decoding: Uses the adjusted model parameters to generate speech.

[0115] Below the fast adaptation module, there are also training strategies: Feature adaptation: The model automatically adjusts according to the input features.

[0116] Parameter sharing: Different parts of the model share the same parameters.

[0117] Online fine-tuning: After the model is deployed, it is fine-tuned according to new data.

[0118] Contrast framework: The model is trained by comparing different samples.

[0119] Content contrast: Compare the content information of the speech.

[0120] Style contrast: Pay attention to the style features of the speech, such as timbre and emotion.

[0121] Identity contrast: Distinguish the voice features of different speakers.

[0122] Below the contrast learning module, there is also hard example mining: Hard example identification: Identify samples that are difficult for the model to process.

[0123] Difficulty assessment: Evaluate the difficulty of the samples.

[0124] Weight assignment: Assign weights according to the difficulty of the samples.

[0125] Coding structure: Define the structure of the timbre coding network.

[0126] Timbre feature: Extract and encode the timbre features of the speech.

[0127] Style modeling: Model the style features of the speech.

[0128] Feature preservation: Ensure that key timbre and style features are retained.

[0129] Below the timbre coding network module, there is also style mixing: Feature separation: Separate different features in the speech.

[0130] Intensity control: Control the intensity of the style features.

[0131] Consistency preservation: Ensure the consistency of the speech during the style mixing process.

[0132] In the present invention, through fast adaptation, contrast learning, and the timbre coding network, efficient, flexible, and high-quality speech synthesis or processing is achieved. Fast adaptation ensures that the model quickly adapts to new data, the contrast learning module improves the robustness and generalization ability of the model, and the timbre coding network focuses on capturing and retaining the timbre and style features of the speech. These modules work together to generate natural and personalized speech output.

[0133] Figure 7 Schematic structural diagram of the speech synthesis device provided by the present invention, as Figure 7 shown, including: The semantic understanding module 710 is configured to input the text content into the LLM semantic understanding module to obtain the deep semantic features corresponding to the text content and the multi-modal context semantic information; The encoding and decoding module 720 is configured to transmit the deep semantic features to the neural codec and output the compressed acoustic features; The modeling module 730 is configured to input the compressed acoustic features into the acoustic modeling module and output high-precision acoustic features; The pitch control module 740 is configured to input the high-precision acoustic features and the multi-modal context semantic information into the emotional pitch control module and input the voice parameters with emotional and prosodic annotations; The synthesis module 750 is configured to input the voice parameters with emotional and prosodic annotations and the reference audio into the timbre transfer module to obtain the synthesized speech with the timbre corresponding to the reference audio.

[0134] In the present invention, through the three-layer semantic parsing structure (word level, sentence level, discourse level) of the LLM semantic understanding module, the deep semantic features of the text can be deeply mined. At the same time, combined with the context awareness mechanism and the multi-granularity feature extraction mechanism, multi-modal context semantic information is obtained. This design not only improves the depth and breadth of semantic understanding, but also solves the problems of insufficient semantic understanding and underutilization of context information in traditional methods. In the acoustic feature compression stage, the neural codec module adopts a hierarchical coding structure (harmonic layer, resonance layer, noise layer) to efficiently compress the deep semantic features and output the compressed acoustic features. Through the improved decoder structure and the progressive training scheme, the reconstruction quality is further optimized, and the problems of low efficiency and poor reconstruction effect of traditional codecs are solved. Then, in the acoustic modeling stage, the adaptive acoustic modeling module combines the hybrid attention mechanism and the dynamic convolutional network to improve the feature extraction efficiency and modeling ability, and outputs high-precision acoustic features. This design significantly enhances the accuracy and robustness of acoustic modeling, making up for the deficiency of limited acoustic modeling ability in the existing technology. In the emotion and prosody control stage, the emotion prosody control module generates speech parameters with emotion and prosody annotations through the emotion analysis network driven by the large model and the precise prosody prediction network. The emotion analysis network combines the emotion feature extraction network and the emotion dictionary enhancement mechanism to achieve precise emotion analysis; the prosody prediction network generates natural emotion expressions and prosody changes through the multi-task prosody predictor and the long-range dependence modeling mechanism. This step effectively solves the problems of unnatural emotion expression and inaccurate prosody control in traditional methods. Finally, in the timbre transfer and speech synthesis stage, the efficient timbre transfer module realizes efficient timbre transfer based on the timbre encoding network driven by StyleGAN and the contrast learning optimization strategy. Through the hierarchical style encoding structure, the adaptive style mixing mechanism and the triple contrast learning framework, the quality and stability of timbre transfer are ensured. At the same time, the fast timbre adaptation network realizes fast timbre adaptation and high-quality synthesis through the lightweight adaptation module and the progressive training strategy. The problems of unstable timbre transfer quality and low adaptation efficiency in the existing technology are solved. Therefore, through the modular design and optimization strategy, this application systematically solves the problems of insufficient semantic understanding, low codec efficiency, limited acoustic modeling ability, unnatural emotion expression, unstable timbre transfer quality, etc., and realizes high-quality and high-efficiency speech synthesis.

[0135] Figure 8 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 8As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute a speech synthesis method, which includes: inputting the text content into the LLM semantic understanding module to obtain the deep semantic features and multi-modal context semantic information corresponding to the text content; Transmitting the deep semantic features to a neural codec to output compressed acoustic features; Inputting the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features; Inputting the high-precision acoustic features and the multi-modal context semantic information into an emotion prosody control module to input speech parameters with emotion and prosody annotations; Inputting the speech parameters with emotion and prosody annotations, and a reference audio into a timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio.

[0136] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0137] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech synthesis method provided by the above-mentioned various methods. The method includes: inputting the text content into the LLM semantic understanding module to obtain the deep semantic features and multi-modal context semantic information corresponding to the text content; Transmitting the deep semantic features to a neural codec to output compressed acoustic features; Input the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features; Input the high-precision acoustic features and the multi-modal context semantic information into an emotion prosody control module to input speech parameters with emotion and prosody annotations; Input the speech parameters with emotion and prosody annotations, as well as a reference audio, into a timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio.

[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a speech synthesis method provided by the above various methods. The method includes: inputting text content into an LLM semantic understanding module to obtain deep semantic features corresponding to the text content and multi-modal context semantic information; Transmit the deep semantic features to a neural codec to output compressed acoustic features; Input the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features; Input the high-precision acoustic features and the multi-modal context semantic information into an emotion prosody control module to input speech parameters with emotion and prosody annotations; Input the speech parameters with emotion and prosody annotations, as well as a reference audio, into a timbre transfer module to obtain a synthesized speech with the timbre corresponding to the reference audio.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech synthesis method, characterized in that: include: Input the text content into the LLM semantic understanding module to obtain the deep semantic features and multimodal contextual semantic information corresponding to the text content; Transmitting the deep semantic features to a neural codec and outputting compressed acoustic features; Inputting the compressed acoustic features into an acoustic modeling module to output high-precision acoustic features; Inputting the high-precision acoustic features and the multimodal contextual semantic information into an emotional prosody control module, and inputting speech parameters with emotional and prosody annotations; The speech parameters with emotion and rhythm annotations and the reference audio are input into a timbre migration module to obtain a synthesized speech with a timbre corresponding to the reference audio.

2. The speech synthesis method according to claim 1, characterized in that: The step of inputting the text content into the LLM semantic understanding module to obtain the deep semantic features and multimodal contextual semantic information corresponding to the text content includes: According to the word-level semantic layer in the LLM semantic understanding module, the text content is subjected to word-level semantic analysis to obtain word-level semantic features; According to the sentence-level semantic layer in the LLM semantic understanding module, the word-level semantic features are analyzed for tone and context information to obtain sentence-level semantic features; According to the text semantic layer in the LLM semantic understanding module, the sentence-level semantic features are analyzed for the overall structure and logical relationship of the text to obtain text semantic information; Determining deep semantic features corresponding to the text content based on the word-level semantic features, sentence-level semantic features, and the text semantic information; The temporal context, emotional context and topic context of the text content are analyzed through a context-aware mechanism to obtain multimodal contextual semantic information.

3. The speech synthesis method according to claim 1, characterized in that: The deep semantic features are transmitted to the neural codec, and compressed acoustic features are output, including: Inputting the deep semantic features into the harmonic layer coding unit in the neural codec, performing fundamental frequency detection, harmonic analysis and feature coding, and obtaining a first feature vector containing fundamental frequency and harmonic structure information; The first feature vector is transmitted to the resonance layer encoding unit in the neural codec to perform resonance peak extraction, vocal tract parameter estimation and feature compression processing to obtain a resonance feature vector; Transmitting the resonance feature vector to the noise layer encoding unit in the neural codec for noise separation feature quantization and encoding compression, and obtaining a second feature vector containing non-periodic component information; The second feature vector is input into a decoder for decoding to obtain compressed acoustic features.

4. The speech synthesis method according to claim 1, characterized in that: The compressed acoustic features are input into an acoustic modeling module to output high-precision acoustic features, including: Capturing local acoustic features using a multi-scale window through a local attention unit in the acoustic modeling module; Performing global acoustic feature extraction through a global attention unit in the acoustic modeling module; Based on the position encoding unit in the acoustic modeling module, the temporal position is explicitly modeled into the local acoustic features and the global acoustic features to obtain encoding features; The high-precision acoustic feature is determined according to the coding feature.

5. The speech synthesis method according to claim 1, characterized in that: The step of inputting the high-precision acoustic features and the multimodal contextual semantic information into an emotion and prosody control module and inputting speech parameters with emotion and prosody annotations comprises: The emotion analysis module in the emotion rhythm control module performs explicit emotion analysis, implicit emotion analysis and mixed emotion analysis on the high-precision acoustic features and the multimodal contextual semantic information to obtain emotion annotation information; The rhythm prediction module in the emotional rhythm control module performs pitch prediction, phoneme duration prediction and phoneme energy prediction on the high-precision acoustic features to obtain rhythm annotation information; According to the emotion annotation information and the prosody annotation information, speech parameters with emotion and prosody annotations are obtained.

6. The speech synthesis method according to claim 1, characterized in that: Inputting the speech parameters with emotion and rhythm annotations and the reference audio into a timbre migration module to obtain a synthesized speech with a timbre corresponding to the reference audio, including: The timbre encoding unit in the timbre migration module extracts timbre features, speaking style and audio personalization from the reference audio to obtain target timbre features of the reference audio; By means of the fast adaptation module in the timbre migration module, the target timbre features are mapped to the speech parameters with emotion and rhythm annotations to obtain a synthesized speech with the timbre corresponding to the reference audio; Among them, the contrast learning module in the timbre migration module is used to improve the accuracy and robustness of the timbre migration performed by the timbre migration module.

7. A speech synthesis device, characterized in that: include: A semantic understanding module, used to input text content into the LLM semantic understanding module to obtain deep semantic features and multimodal contextual semantic information corresponding to the text content; A codec module, used for transmitting the deep semantic features to a neural codec and outputting compressed acoustic features; A modeling module, used for inputting the compressed acoustic features into an acoustic modeling module and outputting high-precision acoustic features; A rhythm control module, used for inputting the high-precision acoustic features and the multimodal contextual semantic information into an emotional rhythm control module, and inputting speech parameters with emotion and rhythm annotations; The synthesis module is used to input the speech parameters with emotion and rhythm annotations and the reference audio into the timbre migration module to obtain the synthesized speech with the timbre corresponding to the reference audio.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the speech synthesis method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Voice cloning method and device based on emotion enhancement and related medium

    CN120599998A

  • A voice cloning method, device, and related medium based on emotion enhancement

    CN120599998B

  • Speech recognition method and system based on large language model

    CN120783764A