Speech synthesis method and device, computer equipment and storage medium

By extracting features from timbre, emotion, and prosody encoders and fusing them to generate speech, the problem of single timbre and prosody in traditional speech synthesis methods is solved, and high-quality speech synthesis is achieved.

CN121789635APending Publication Date: 2026-04-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional speech synthesis methods are relatively limited in terms of timbre, emotion, and rhythm, making it difficult to meet users' needs for high-quality speech synthesis.

Method used

A timbre encoder, an emotion encoder, and a prosody encoder are used to extract timbre, emotion, and prosody feature vectors from the speech prompt text, respectively. The target synthesized speech is generated by feature fusion and then converted into speech using a vocoder.

Benefits of technology

It achieves feature extraction and effective fusion of voice prompt text from three dimensions: timbre, emotion, and rhythm, and synthesizes target speech with specific timbre, rich emotion, and natural rhythm to meet users' needs for high-quality speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789635A_ABST
    Figure CN121789635A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of audio processing, and relates to a voice synthesis method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a voice prompt text; inputting the timbre prompt text into a timbre encoder for timbre encoding to obtain a timbre feature vector; inputting the emotion prompt text into an emotion encoder for emotion encoding to obtain an emotion feature vector; inputting the rhythm prompt text into a rhythm encoder for rhythm encoding to obtain a rhythm feature vector; performing feature fusion on the timbre feature vector, the emotion feature vector and the rhythm feature vector to obtain a fused feature vector; and inputting the fusion feature vector into a vocoder for voice conversion to obtain a target synthetic voice. The method can be used for related speech synthesis in business systems such as financial science and technology, medical health, old-age care and the like, and can meet the requirements of users for high-quality speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a speech synthesis method, apparatus, computer device, and storage medium. Background Technology

[0002] As one of the key technologies in the field of human-computer interaction, speech synthesis technology aims to transform text information into natural and fluent speech output. It is widely used in many scenarios such as intelligent voice assistants, audiobooks, voice navigation, virtual anchors, fintech, and medical and elderly care.

[0003] Traditional speech synthesis methods typically only consider the basic semantic information of text, converting text into speech through simple rules or models. The generated speech is relatively simple in terms of timbre, emotion, and rhythm, lacking richness and naturalness, and is difficult to meet users' needs for high-quality speech synthesis.

[0004] This shows that traditional speech synthesis methods suffer from low synthesis quality.

[0005] Application content The purpose of this application is to provide a speech synthesis method, apparatus, computer device, and storage medium to solve the problem of low synthesis quality in traditional speech synthesis methods.

[0006] To address the aforementioned technical problems, this application provides a speech synthesis method, employing the following technical solution: Obtain the voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text, and rhythm prompt text; The timbre prompt text is input into the timbre encoder for timbre encoding to obtain the timbre feature vector; The emotional prompt text is input into the emotional encoder for emotional encoding to obtain an emotional feature vector; The prosodic prompt text is input into the prosodic encoder for prosodic encoding to obtain the prosodic feature vector; The timbre feature vector, the emotion feature vector, and the prosody feature vector are fused to obtain a fused feature vector. The fused feature vector is input into a vocoder for speech conversion to obtain the target synthesized speech.

[0007] To address the aforementioned technical problems, this application also provides a speech synthesis device, which employs the following technical solution: The prompt text acquisition module is used to acquire voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text, and rhythm prompt text; The timbre encoding module is used to input timbre prompt text into the timbre encoder for timbre encoding, thereby obtaining a timbre feature vector; The sentiment encoding module is used to input sentiment prompt text into the sentiment encoder for sentiment encoding, and obtain sentiment feature vectors. The prosody coding module is used to input prosody cue text into the prosody encoder for prosody coding, and obtain prosody feature vectors. The feature fusion module is used to fuse the timbre feature vector, the emotion feature vector, and the prosody feature vector to obtain a fused feature vector. The speech conversion module is used to input the fused feature vector into the vocoder for speech conversion to obtain the target synthesized speech.

[0008] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech synthesis method as described above.

[0009] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech synthesis method described above.

[0010] This application provides a speech synthesis method, comprising: acquiring speech prompt text, wherein the speech prompt text includes timbre prompt text, emotion prompt text, and prosodic prompt text; inputting the timbre prompt text into a timbre encoder for timbre encoding to obtain a timbre feature vector; inputting the emotion prompt text into an emotion encoder for emotion encoding to obtain an emotion feature vector; inputting the prosodic prompt text into a prosodic encoder for prosodic encoding to obtain a prosodic feature vector; fusing the timbre feature vector, the emotion feature vector, and the prosodic feature vector to obtain a fused feature vector; and inputting the fused feature vector into a vocoder for speech conversion to obtain a target synthesized speech. Compared with the prior art, this application can extract features from the speech prompt text from three dimensions: timbre, emotion, and prosodic, and effectively fuse the extracted features to finally synthesize a target speech with a specific timbre, rich emotion, and natural prosodic, thereby meeting users' needs for high-quality speech synthesis. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart illustrating the implementation of the speech synthesis method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the speech synthesis device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the speech synthesis method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the speech synthesis device is generally set in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Continue to refer to Figure 2 The diagram shows a flowchart of an embodiment of the speech synthesis method according to this application. The speech synthesis method described above includes steps S201, S202, S203, and S204.

[0023] In step S201, the voice prompt text is obtained, wherein the voice prompt text includes timbre prompt text, emotion prompt text and rhythm prompt text.

[0024] In the embodiments of this application, users can input voice prompt text through their terminal devices, which is received by the system. The terminal device can be a mobile terminal such as a mobile phone, smartphone, laptop, digital broadcast receiver, PDA (personal digital assistant), PAD (tablet computer), PMP (portable multimedia player), navigation device, etc., or a fixed terminal such as a digital TV, desktop computer, etc. It should be understood that the examples of terminal devices given herein are for convenience of understanding only and are not intended to limit this application.

[0025] In the embodiments of this application, the aforementioned voice prompt text is the specific content that the user wants the system to perform voice synthesis. Specifically, the voice prompt text may be related to financial institutions (such as banks), or it may be related to medical scenarios. It should be understood that the examples of voice prompt texts here are only for ease of understanding and are not intended to limit this application.

[0026] In this embodiment of the application, the voice prompt text includes timbre prompt text, emotion prompt text, and prosodic prompt text. The timbre prompt text is mainly used to specify the timbre features required for synthesized speech, such as descriptive texts like "male voice, deep" or "female voice, sweet." The emotion prompt text is mainly used to express the emotions that the synthesized speech should contain, such as "happy," "sad," or "angry." The prosodic prompt text is mainly used to specify the prosodic characteristics of the synthesized speech, including descriptions of speech rate, intonation, and pauses, such as "moderate speech rate, large intonation fluctuations, and pauses at specific locations."

[0027] In step S202, the timbre prompt text is input into the timbre encoder for timbre encoding to obtain the timbre feature vector.

[0028] In step S203, the emotional prompt text is input into the emotional encoder for emotional encoding to obtain the emotional feature vector.

[0029] In step S204, the prosodic prompt text is input into the prosodic encoder for prosodic encoding to obtain the prosodic feature vector.

[0030] In this embodiment, each encoder is specifically designed for a particular type of prompt text, capable of deeply mining the feature information contained in the text and converting it into a numerical feature vector. The encoder includes a timbre encoder, an emotion encoder, and a prosodic encoder, specifically: (1) The timbre encoder processes the timbre prompt text to obtain a timbre feature vector, which can accurately represent the timbre features of the target speech. As an example, the timbre encoder can be a text vectorization model. (2) The emotion encoder encodes the emotion prompt text and generates an emotion feature vector to reflect the emotional state that the speech should express. As an example, the emotion encoder can be the BERT-FinEmotion fine-tuning model. (3) The prosodic encoder converts the prosodic cue text into a prosodic feature vector to describe the prosodic pattern of the speech. For example, the prosodic encoder can be the RhythmNet model.

[0031] By using three independent encoders, it is ensured that the features of each dimension can be fully and accurately extracted. It should be understood that the examples of timbre encoder, emotion encoder and prosody encoder above are for convenience of understanding only and are not intended to limit this application.

[0032] In step S205, the timbre feature vector, emotion feature vector, and rhythm feature vector are fused to obtain a fused feature vector.

[0033] In this application's embodiments, the purpose of feature fusion is to organically combine features from different dimensions, enabling them to complement and synergize with each other, thereby more comprehensively describing the characteristics of the target speech. Specifically, this application can fuse three feature vectors using a specific fusion algorithm. Examples include weighted fusion, concatenation fusion, or attention-based fusion methods. It should be understood that these examples of fusion algorithms are for illustrative purposes only and are not intended to limit this application. In practical applications, taking weighted fusion as an example, each feature vector is assigned a corresponding weight based on its importance to speech synthesis. The weighted feature vectors are then summed to obtain the fused feature vector. Through reasonable feature fusion, this application can achieve better coordination and unity in timbre, emotion, and prosody in synthesized speech.

[0034] In this embodiment of the application, it is assumed that the timbre feature vector, emotion feature vector, and prosodic feature vector obtained by the three independent encoders are as follows:

[0035]

[0036]

[0037] in, Represents the timbre feature vector. Represents the sentiment feature vector. Represents the prosodic feature vector; Therefore, the process of feature fusion can be represented as follows:

[0038] in, Represents the fused feature vector. This indicates the fusion strategy.

[0039] In the embodiments of this application, the above-mentioned fusion strategy It can be a weighted fusion strategy, a conditional injection strategy, or a cross-attention fusion strategy, specifically: (1) Weighted fusion strategy:

[0040] in, , , These represent the fusion weights; (2) Conditional injection strategy: Inject feature vectors as conditions into different intermediate layers of the TTS model; (3) Cross-attention fusion: By using an attention mechanism to weightedly integrate the three features, the naturalness and consistency of the generated data are improved.

[0041] In step S206, the fused feature vector is input into the vocoder for speech conversion to obtain the target synthesized speech.

[0042] In the embodiments of this application, a vocoder is a model or device that can convert feature vectors into speech signals. It can accurately interpret the information in the fused feature vectors and convert them into natural and fluent target speech with specific timbre, emotion and rhythm features.

[0043] In practical applications, taking a financial business scenario as an example, suppose the voice prompt text obtained by the system is: • Voice prompt text: "Male voice, 35-45 years old, professional financial advisor voice, mid-to-low range, clear enunciation" (matching the trust needs of high-end clients); • Emotional prompt text: "Enthusiastic and positive (profit promotion section), rigorous and objective (risk disclosure section)" (emotional transition in compliance with regulatory requirements); • Pronunciation prompts: "Profit data section: 180 words / minute, tone rising 30%; Risk clause section: 120 words / minute, 0.3-second pause at the end of each sentence" (to enhance the delivery of key information); This application inputs the timbre prompt text into a timbre encoder for timbre encoding, obtaining a 128-dimensional timbre feature vector (including acoustic fingerprints such as age perception, professionalism, and credibility); inputs the emotion prompt text into an emotion encoder for emotion encoding, obtaining a 64-dimensional emotion feature vector (enthusiasm 0.8 / rigor 0.7 / transition smoothness 0.9); inputs the prosody prompt text into a prosody encoder for prosody encoding, obtaining a 96-dimensional prosody feature vector (speech rate mapping value 0.75 / pause pattern encoding 0.62). Then, this application dynamically weights and fuses the 128-dimensional timbre feature vector, the 64-dimensional emotion feature vector, and the 96-dimensional prosody feature vector to obtain a 256-dimensional fused feature vector; finally, the 256-dimensional fused feature vector is input into a vocoder to obtain the target synthesized speech (WAV format) output by the vocoder, thus obtaining synthesized speech that meets the above requirements for the speech prompt text.

[0044] This application provides a speech synthesis method, comprising: acquiring speech prompt text, wherein the speech prompt text includes timbre prompt text, emotion prompt text, and prosodic prompt text; inputting the timbre prompt text into a timbre encoder for timbre encoding to obtain a timbre feature vector; inputting the emotion prompt text into an emotion encoder for emotion encoding to obtain an emotion feature vector; inputting the prosodic prompt text into a prosodic encoder for prosodic encoding to obtain a prosodic feature vector; fusing the timbre feature vector, emotion feature vector, and prosodic feature vector to obtain a fused feature vector; and inputting the fused feature vector into a vocoder for speech conversion to obtain the target synthesized speech. Compared with the prior art, this application can extract features from the speech prompt text from three dimensions: timbre, emotion, and prosodic, and effectively fuse the extracted features to finally synthesize a target speech with a specific timbre, rich emotion, and natural prosodic, thereby meeting users' needs for high-quality speech synthesis.

[0045] In some optional implementations of the embodiments of this application, the step of inputting the timbre prompt text into the timbre encoder for timbre encoding to obtain the timbre feature vector specifically includes the following steps: The timbre prompt text is semantically segmented based on a pre-trained language model to extract key timbre feature labels. Read the acoustic knowledge base and obtain the acoustic feature template corresponding to the timbre key feature label from the acoustic knowledge base; Based on the acoustic feature template, the timbre prompt text is preprocessed for timbre synthesis to obtain a reference timbre segment; The MFCC feature coefficients of the reference timbre segment are extracted using the Mel frequency cepstral coefficient algorithm; The fundamental frequency trajectory parameters of the reference timbre segment are extracted based on the acoustic model; The MFCC feature coefficients and the fundamental frequency trajectory parameters are input into the input text vectorization model to perform timbre nonlinear mapping, thereby obtaining the timbre feature vector.

[0046] In this embodiment, semantic segmentation refers to word segmentation and semantic encoding using a pre-trained language model, outputting a sequence of word vectors, which are then mapped to a label space through a fully connected layer to extract timbre key feature labels. These timbre key feature labels are mainly used to describe the timbre of the synthesized speech. For example, if the timbre prompt text is "male voice, 40 years old, fund manager voice, mid-low range, moderate speaking speed", then the timbre key feature label would be "[gender: male, age: 40, field: fund, range: mid-low, speaking speed: moderate]". It should be understood that the example of semantic segmentation here is only for ease of understanding and is not intended to limit this application.

[0047] In this embodiment of the application, an acoustic knowledge base structure is pre-constructed, which can be shown in the following table:

[0048] Then, this application can query the knowledge base based on the combination of tags to obtain the most matching template.

[0049] In this embodiment, the timbre synthesis preprocessing refers to initializing the speech synthesis engine based on template parameters, then inputting tagged text to generate reference timbre segments, and finally performing prosody correction through a GRU network to ensure that the intonation is ≥90% natural.

[0050] In this embodiment, the reference timbre segment is framed (25ms frame length, 10ms frame shift); then, 13-dimensional MFCC coefficients (including energy and first-order difference) are calculated; finally, a feature matrix (13×T, where T is the number of frames) is output, thereby extracting the MFCC feature coefficients.

[0051] In this embodiment, the fundamental frequency of each frame of the reference timbre segment is calculated using an acoustic model; then, the trajectory is smoothed using the Viterbi algorithm to filter out outliers (such as setting unvoiced segments to zero); finally, the fundamental frequency sequence (1×T, in Hz) is output to extract the fundamental frequency trajectory parameters.

[0052] In this embodiment of the application, the text vectorization model includes an input layer, an encoding layer, a fusion layer, and an output layer, wherein: (1) Input layer: It is mainly used to splice MFCC matrices and fundamental frequency sequences (14×T dimensions). (2) Coding layer: ① Temporal features are extracted using a 2-layer BiLSTM; ②Focus on keyframes (such as accented segments) through self-attention mechanisms; (3) Fusion layer: The weights of MFCC and baseband are mainly adjusted dynamically through the gating unit; (4) Output layer: The main process involves generating a 128-dimensional timbre feature vector (including metadata such as acoustic age and professional level) through a fully connected layer.

[0053] Compared with existing technologies, this application introduces an end-to-end generation method for timbre-based prompt text to high-precision feature vectors through semantic segmentation, acoustic knowledge base matching, and nonlinear mapping.

[0054] In some optional implementations of the embodiments of this application, the step of inputting the emotional prompt text into the emotional encoder for emotional encoding to obtain the emotional feature vector specifically includes the following steps: The sentiment prompt text is segmented and part-of-speech tagged to obtain sentiment prompt word segments; The emotional prompt context is obtained by extracting the context of the emotional prompt words based on the sliding window; The emotional features are extracted from the emotional cue context to obtain the emotional feature vector.

[0055] In this embodiment, the input sentiment prompt text can be segmented using a word segmentation tool (such as Jieba or Stanford). The word segmentation tool divides the continuous text string into individual words based on a preset dictionary and segmentation algorithm. For example, for the sentiment prompt text "This movie is very exciting and impressive," after word segmentation, the result is "This movie is very exciting and impressive." Based on the segmentation, a part-of-speech tagging tool (such as Stanford or Harbin Institute of Technology) is used to tag each word. The purpose of part-of-speech tagging is to determine the grammatical role of each word in the sentence, such as noun, verb, adjective, or adverb. For example, the part-of-speech tagging of the above segmentation result is "This / r / q movie / n is / d / exciting / a, / w makes / v people / n impressed / n deeply / a." Through word segmentation and part-of-speech tagging, the sentiment prompt text can be transformed into a structured data form, providing a foundation for subsequent context extraction and sentiment feature extraction.

[0056] In the embodiments of the present application, a sliding window is a fixed-size window that moves over sequence data and is used to extract local context information. In the present application, a fixed-size sliding window is set, and the window size can be adjusted according to actual needs, such as being set to 3, 5, or 7, etc.; in the present application, with each sentiment prompt tokenization as the center, a sliding window is used to extract context tokenizations within a certain range before and after it. For example, assuming the sliding window size is 5, for the sentiment prompt tokenization "wonderful", its first two tokenizations are "movie very", and the last two tokenizations are ", makes people", then "movie very wonderful, makes people" is the sentiment prompt context extracted with "wonderful" as the center. By means of the sliding window, the local context relationship of the sentiment prompt tokenization in the text can be fully considered, and important information related to sentiment expression can be captured.

[0057] In the embodiments of the present application, the present application uses deep learning models (such as convolutional neural network CNN, recurrent neural network RNN and its variants long short-term memory network LSTM, gated recurrent unit GRU, etc.) to extract sentiment features from the sentiment prompt context. These deep learning models have powerful feature learning capabilities and can automatically learn high-level feature representations from text data. Specifically: (1) Taking CNN as an example: Map each tokenization in the sentiment prompt context to a word vector of a fixed dimension (a pre-trained word vector model can be used, such as Word2Vec, GloVe, etc.), and then combine these word vectors into a matrix as the input of CNN. CNN performs sentiment feature extraction on the input matrix through operations such as convolutional layers and pooling layers, and finally obtains a sentiment feature vector of a fixed dimension. The convolutional kernels in the convolutional layer can capture the feature patterns in the local context, and the pooling layer can reduce the dimension and abstract the features to extract more representative features; (2) Taking LSTM as an example: Similarly, map the tokenizations in the sentiment prompt context to word vectors, and then sequentially input the word vector sequence into the LSTM model. The LSTM model can effectively process the long-term dependencies in the sequence data through its unique gating mechanism (input gate, forget gate, and output gate), and capture the sentiment information in the context. At the last time step of the LSTM model, a hidden state vector containing the entire sentiment prompt context information can be obtained, and this vector is the extracted sentiment feature vector.

[0058] Compared with the prior art, the present application can accurately perform tokenization and part-of-speech tagging on sentiment prompt texts, fully consider the context information of sentiment prompt tokenizations, and extract comprehensive and accurate sentiment feature vectors, thereby improving the performance of sentiment analysis.

[0059] In some optional implementations of the embodiments of this application, the step of extracting emotional features from the emotional cue context to obtain the emotional feature vector specifically includes the following steps: The emotional cue context is input into a bidirectional encoding classifier to obtain the emotional category probability of each emotional category, and the emotional category probability is converted into one-hot encoding; The emotional intensity value of each emotional category is calculated according to the regression model, and the emotional intensity value is converted into an emotional intensity sequence. Contextual features of the emotional cues are obtained based on an attention mechanism; The one-hot encoding, the sentiment intensity sequence, and the contextual features are input into a nonlinear mapping model to fuse sentiment features, thereby obtaining the sentiment feature vector.

[0060] In this embodiment, the preprocessed sentiment cue context is input into a pre-trained BERT classifier. The BERT classifier outputs the sentiment category probability for each sentiment category based on the input sentiment cue context. For example, in a binary sentiment analysis task, it may output the probability of positive and negative sentiment categories; in a multi-class task, it will output multiple probability values ​​for different sentiment categories. Then, the obtained sentiment category probabilities are converted into one-hot encoding. One-hot encoding is a method of representing categorical variables as binary vectors, where each category corresponds to a dimension, where 1 represents a value and 0 represents other dimensions. For example, if there are three sentiment categories A, B, and C, and a sample belongs to category B, its one-hot encoding is [0, 1, 0]. Through one-hot encoding, discrete sentiment category information can be converted into a numerical vector representation, facilitating subsequent feature fusion.

[0061] In this embodiment, a pre-trained regression model is used to calculate the emotional intensity value of each emotional category. The regression model can be a deep learning-based model, such as a Multilayer Perceptron (MLP) or Long Short-Term Memory (LSTM) network, or a traditional machine learning model, such as Support Vector Regression (SVR). These models, by learning from a large amount of labeled emotional data, can accurately predict the intensity value of each emotional category based on the emotional cues and context. For example, for a positive emotional category, a value between 0 and 1 might be output, with a larger value indicating a higher positive emotional intensity. Then, the calculated emotional intensity values ​​of each emotional category are arranged in a certain order to form an emotional intensity sequence. For example, if there are three emotional categories: positive, negative, and neutral, the corresponding emotional intensity values ​​are arranged in the order of positive, negative, and neutral to form a sequence, such as [0.8, 0.2, 0.1]. The emotional intensity sequence can intuitively represent the intensity distribution of different emotional categories.

[0062] In this embodiment, the Transformer attention mechanism is used to process the emotional cues context and obtain contextual features. The Transformer attention mechanism can automatically capture the dependencies and semantic relationships between different words in a text, highlighting important information and suppressing irrelevant information by calculating attention weights. For example, when processing a sentence, the attention mechanism can focus on other words related to the current word, thus better understanding the sentence's semantics. Then, after processing by the Transformer attention mechanism, a feature vector reflecting the internal relationships within the emotional cues context is obtained. This feature vector contains semantic interaction information between words in the context, helping to more accurately understand the emotional expression of the text.

[0063] In this embodiment, the one-hot encoding, sentiment intensity sequence, and context-related feature vector obtained above are input into a nonlinear mapping model for feature fusion. The nonlinear mapping model can be a combination of deep neural networks (DNN), convolutional neural networks (CNN), and recurrent neural networks (RNN), or other models with powerful nonlinear mapping capabilities. The nonlinear mapping model performs complex nonlinear transformations and combinations on the input multi-dimensional features, fusing them into a comprehensive sentiment feature vector. This sentiment feature vector contains information on sentiment category, sentiment intensity, and contextual relevance, and can more comprehensively and accurately reflect the sentiment characteristics of the sentiment cues within the context.

[0064] Compared with existing technologies, this application can extract multi-dimensional features such as sentiment category, sentiment intensity and contextual association from the sentiment cue context, and effectively fuse these features through a nonlinear mapping model to obtain a comprehensive and accurate sentiment feature vector, thereby improving the performance of sentiment analysis.

[0065] In some optional implementations of the embodiments of this application, the step of inputting the prosodic cue text into a prosodic encoder for prosodic encoding to obtain a prosodic feature vector specifically includes the following steps: The prosodic prompt text is segmented into prosodic semantics based on a pre-trained language model, and key prosodic feature labels are extracted. These key prosodic feature labels include speech rate labels, intonation pattern labels, and key segment markers. Call the speech rate template library, and quantify the speech rate of the speech rate tag according to the speech rate template library to obtain the prosodic temporal speech rate curve; Based on the fundamental frequency trajectory generator, the intonation pattern label is modeled to create an intonation curve; The key segment markers are used to mark the pause positions, and the pause duration is obtained; Read the prosodic template library and obtain the acoustic parameter combination corresponding to the prosodic temporal speech rate curve, the intonation curve and the pause duration from the prosodic template library; The prosodic feature vector is obtained by extracting prosodic features from the acoustic parameter combination.

[0066] In this embodiment, the prosodic cue text is input into a pre-trained BERT model, and key prosodic feature labels are extracted from the output of the BERT model. These labels include speech rate labels, intonation pattern labels, and key paragraph markers. Speech rate labels indicate the expected speed of the text, such as fast, medium, or slow; intonation pattern labels describe the intonation variation patterns of the text, such as rising, falling, or level tones; and key paragraph markers identify paragraphs in the text that require special emphasis or have special prosodic features.

[0067] In this embodiment, a speech rate template library is pre-built, storing speech rate quantization standards and temporal variation patterns corresponding to different speech rate tags. For example, a fast speech rate tag may correspond to a higher speech rate value and a specific speech rate variation curve. Then, based on the speech rate template library, this application performs speech rate quantization processing on the extracted speech rate tags. The speech rate tags are converted into specific speech rate values, and a prosodic temporal speech rate curve is generated according to the temporal relationship of the text. This curve can intuitively display the speech rate changes of the text at different points in time.

[0068] In this embodiment, a fundamental frequency trajectory generator is used to model the intonation trajectory of intonation pattern labels. Fundamental frequency is an important physical parameter of intonation; by analyzing intonation pattern labels, the fundamental frequency trajectory generator can simulate the intonation change trajectory of the text. Then, based on the results of the intonation trajectory modeling, a corresponding intonation curve is created. The intonation curve can clearly present the rising, falling, or smooth changes in the text's intonation, providing important intonation information for subsequent prosodic feature fusion.

[0069] In this embodiment, the extracted key segment markers are annotated with pause positions. Based on the semantic and grammatical structure of the text, the positions where pauses are required between and within key segments are determined. Then, according to preset pause rules or through statistical analysis of historical data, an appropriate pause duration is determined for each pause position. Pause duration affects the rhythm and fluency of speech and is an important component of prosodic features.

[0070] In this embodiment, a prosodic template library is pre-constructed, storing acoustic parameter combinations corresponding to different prosodic temporal speech rate curves, intonation curves, and pause durations. Acoustic parameters include pitch, duration, and intensity, which directly affect the timbre and prosodic expression of speech. Then, the generated prosodic temporal speech rate curve, intonation curve, and determined pause duration are used as query conditions to search the prosodic template library and retrieve the corresponding acoustic parameter combinations. In this way, multi-dimensional prosodic features can be associated with specific acoustic parameters.

[0071] In this embodiment, the obtained acoustic parameter combination is further subjected to prosodic feature extraction. Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), or other dimensionality reduction methods can be used, or a deep learning model can be used to learn and extract features from the acoustic parameters. After feature extraction, a prosodic feature vector is obtained that comprehensively reflects multi-dimensional prosodic features such as speech rate, intonation, and pauses. This vector has high dimensionality and rich information, providing a comprehensive and accurate prosodic feature representation for speech processing applications.

[0072] Compared with existing technologies, this application can extract multi-dimensional prosodic features such as speech rate, intonation, and pauses from prosodic prompt text, and obtain corresponding acoustic parameter combinations by matching with a prosodic template library, thereby extracting a comprehensive and accurate prosodic feature vector, providing high-quality prosodic feature representation for speech processing applications.

[0073] In some optional implementations of the embodiments of this application, the step of extracting prosodic features from the acoustic parameter combination to obtain the prosodic feature vector specifically includes the following steps: The fundamental frequency of each frame of the acoustic parameter combination is calculated based on the acoustic model to obtain the fundamental frequency sequence; Calculate the frame-by-frame speech rate of the acoustic parameter combination to obtain the speech rate time-series curve; Identify the pause positions of the acoustic parameter combinations and quantify their durations to obtain a pause position matrix; The fundamental frequency sequence, the speech rate timing curve, and the pause position matrix are concatenated to obtain a concatenated sequence. The concatenated sequence is input into a text vectorization model for nonlinear mapping and vectorization to obtain the prosodic feature vector.

[0074] In this embodiment, the fundamental frequency of each frame of the acoustic parameter combination is calculated using Kaldi's compute-kaldi-pitch-feats tool to obtain a fundamental frequency sequence. Specifically, firstly, an acoustic parameter combination containing the speech signal is obtained. This acoustic parameter combination can be a feature representation of the preprocessed speech signal, such as a signal processed by framing and windowing. Then, the compute-kaldi-pitch-feats tool from the Kaldi open-source speech recognition toolkit is used. This tool employs a specific algorithm to analyze each frame of the speech signal and calculate its fundamental frequency value. The fundamental frequency is an important prosodic feature in a speech signal, reflecting the frequency of vocal cord vibration and closely related to the pitch of the speech. By calculating the fundamental frequency frame by frame, a fundamental frequency sequence composed of the fundamental frequency values ​​of each frame is obtained. This sequence can describe the pitch variation of the speech over the entire time period.

[0075] In this embodiment, the speech rate is calculated for each frame of speech based on framing information derived from acoustic parameter combinations. Speech rate can be calculated by counting the number of syllables or words per unit time. For example, a time window can be pre-defined, the number of syllables within that window can be counted, and then the speech rate can be calculated based on the length of the time window. By calculating the speech rate for each frame of speech, a speech rate sequence that varies over time is obtained, which is then plotted as a time-series curve, i.e., a speech rate time-series curve. This curve visually displays the changes in speech rate at different times, reflecting the speaker's rhythmic speech characteristics.

[0076] In this embodiment, a pause detection algorithm from speech signal processing is used to analyze the combination of acoustic parameters and identify pause locations in speech. The pause detection algorithm can judge based on features such as speech energy and zero-crossing rate. When the speech energy is below a certain threshold and the zero-crossing rate is also at a low level, it can be determined as a pause. After identifying the pause location, the duration of each pause is further quantified, that is, the duration of each pause from start to finish is calculated. The pause location and corresponding duration information are represented in matrix form to obtain the pause location matrix. This matrix can clearly record the distribution and duration information of pauses in speech, which is of great significance for understanding the rhythm and semantic pauses of speech.

[0077] In this embodiment, the obtained fundamental frequency sequence, speech rate temporal curve, and pause position matrix are concatenated. The concatenation method can be selected according to actual needs. For example, the three sequences can be connected end-to-end in chronological order, or other reasonable concatenation methods can be used so that the concatenated sequence can comprehensively contain multiple prosodic feature information such as fundamental frequency, speech rate, and pauses. The formation of the concatenated sequence realizes the effective fusion of different prosodic features, providing a more comprehensive feature representation for subsequent vectorization processing.

[0078] In this embodiment, the concatenated sequence is input into a text vectorization model. The model, through its internal neural network structure, performs nonlinear transformations and abstract extraction of feature information from the concatenated sequence, mapping the high-dimensional concatenated sequence to a low-dimensional vector space, generating a fixed-dimensional prosodic feature vector. This prosodic feature vector effectively represents various prosodic features of speech. Furthermore, through model training and optimization, the prosodic feature vectors of different speech samples exhibit good discriminativeness and representativeness in the vector space, providing strong feature support for subsequent speech processing tasks.

[0079] Compared with existing technologies, this application can comprehensively extract multiple prosodic features of speech and generate vectors that can accurately represent the prosodic features of speech through effective fusion and vectorization processing, providing high-quality feature input for subsequent speech processing tasks.

[0080] In some optional implementations of the embodiments of this application, the step of inputting the fused feature vector into a vocoder for speech conversion to obtain the target synthesized speech specifically includes the following steps: The fused feature vector is converted into a fused feature Mel spectrum based on a nonlinear mapping network; The fused feature Mel spectrum is smoothed using the Viterbi algorithm to obtain a smoothed Mel spectrum; The smoothed Mel spectrum is input into a streaming decoder for waveform conversion to obtain the target synthesized speech.

[0081] In this embodiment, the application first obtains a fused feature vector containing multiple speech feature information, which may originate from multiple dimensions such as text features and prosodic features. Then, the fused feature vector is input into a pre-trained nonlinear mapping network. This nonlinear mapping network employs a multi-layer neural network structure, such as containing multiple fully connected layers and nonlinear activation function layers, and is capable of automatically learning the complex nonlinear relationship between the fused feature vector and the Mel spectrum. Through the forward propagation process of the nonlinear mapping network, the fused feature vector is mapped to the Mel spectrum space to obtain the fused feature Mel spectrum.

[0082] In this embodiment, the Viterbi algorithm is a dynamic programming algorithm commonly used to solve the optimal path problem in Hidden Markov Models. This application treats the fused feature Mel spectrum as a sequence with state transitions, where each Mel spectrum value at each time step corresponds to a state. By defining state transition probabilities and observation probabilities, the Viterbi algorithm is used to find the state sequence among all possible state sequences that maximizes the probability of the observed fused feature Mel spectrum sequence—the optimal path. The fused feature Mel spectrum is then adjusted based on the optimal path to remove noise and outliers, resulting in a smoothed Mel spectrum.

[0083] In this embodiment, the streaming WaveGlow decoder is a speech waveform synthesis model based on streaming generation. It can generate speech waveforms step by step in a streaming manner, improving the efficiency and real-time performance of waveform generation. This application uses smoothed Mel-spectrum data as input. The streaming WaveGlow decoder, through its internal convolutional neural network structure and reversible transform operations, converts the Mel-spectrum information into a continuous speech waveform signal. After processing by the decoder, high-quality, fluent target synthesized speech is finally obtained.

[0084] Compared with existing technologies, this application achieves efficient transformation of fused feature vectors through a nonlinear mapping network, uses the Viterbi algorithm for spectral smoothing, and combines a streaming WaveGlow decoder for efficient waveform transformation, thereby improving the quality and fluency of synthesized speech.

[0085] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0086] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0088] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0089] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech synthesis device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0090] like Figure 3 As shown, the speech synthesis device 200 of this application embodiment includes: The prompt text acquisition module 210 is used to acquire voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text and rhythm prompt text; The timbre encoding module 220 is used to input the timbre prompt text into the timbre encoder for timbre encoding to obtain the timbre feature vector; The emotion encoding module 230 is used to input the emotion prompt text into the emotion encoder for emotion encoding to obtain the emotion feature vector; The prosody coding module 240 is used to input the prosody cue text into the prosody encoder for prosody coding to obtain the prosody feature vector; The feature fusion module 250 is used to fuse the timbre feature vector, the emotion feature vector, and the prosody feature vector to obtain a fused feature vector. The speech conversion module 260 is used to input the fused feature vector into the vocoder for speech conversion to obtain the target synthesized speech.

[0091] In this embodiment, a speech synthesis device 200 is provided, comprising: a prompt text acquisition module 210 for acquiring voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text, and prosody prompt text; a timbre encoding module 220 for inputting the timbre prompt text into a timbre encoder for timbre encoding to obtain a timbre feature vector; an emotion encoding module 230 for inputting the emotion prompt text into an emotion encoder for emotion encoding to obtain an emotion feature vector; a prosody encoding module 240 for inputting the prosody prompt text into a prosody encoder for prosody encoding to obtain a prosody feature vector; a feature fusion module 250 for fusing the timbre feature vector, emotion feature vector, and prosody feature vector to obtain a fused feature vector; and a speech conversion module 260 for inputting the fused feature vector into a vocoder for speech conversion to obtain the target synthesized speech. Compared with existing technologies, this application can extract features from voice prompt text from three dimensions: timbre, emotion, and rhythm, and effectively fuse the extracted features to synthesize target speech with specific timbre, rich emotion, and natural rhythm, so as to meet users' needs for high-quality speech synthesis.

[0092] In some optional implementations of the embodiments of this application, the above-mentioned timbre encoding module includes: The timbre semantic segmentation submodule is used to perform semantic segmentation on the timbre prompt text based on a pre-trained language model and extract key timbre feature labels. The acoustic feature template acquisition submodule is used to read the acoustic knowledge base and acquire the acoustic feature template corresponding to the timbre key feature label from the acoustic knowledge base; The timbre synthesis preprocessing submodule is used to perform timbre synthesis preprocessing on the timbre prompt text according to the acoustic feature template to obtain a reference timbre segment; The MFCC feature coefficient acquisition submodule is used to extract the MFCC feature coefficients of the reference timbre segment according to the Mel frequency cepstral coefficient algorithm; The fundamental frequency trajectory parameter acquisition submodule is used to extract the fundamental frequency trajectory parameters of the reference timbre segment based on the acoustic model; The timbre nonlinear mapping submodule is used to input the MFCC feature coefficients and the fundamental frequency trajectory parameters into the input text vectorization model to perform timbre nonlinear mapping and obtain the timbre feature vector.

[0093] In some optional implementations of the embodiments of this application, the above-mentioned emotion encoding module includes: The sentiment segmentation submodule is used to segment and tag the sentiment prompt text to obtain sentiment prompt word segments; The context extraction submodule is used to extract the context of the sentiment prompt word segmentation based on the sliding window to obtain the sentiment prompt context. The emotion feature extraction submodule is used to extract emotion features from the emotion prompt context to obtain the emotion feature vector.

[0094] In some optional implementations of the embodiments of this application, the above-mentioned emotion feature extraction submodule includes: The emotion classification unit is used to input the emotion cue context into a bidirectional coding classifier to obtain the emotion category probability of each emotion category, and convert the emotion category probability into one-hot encoding; The emotion intensity unit is used to calculate the emotion intensity value of each emotion category according to the regression model, and to convert the emotion intensity value into an emotion intensity sequence. The emotion intensity unit is used to obtain the contextual features of the emotion cue context based on the attention mechanism; The emotion feature fusion unit is used to input the one-hot encoding, the emotion intensity sequence, and the context-related features into a nonlinear mapping model to perform emotion feature fusion and obtain the emotion feature vector.

[0095] In some optional implementations of the embodiments of this application, the prosody coding module includes: The prosodic semantic segmentation submodule is used to perform prosodic semantic segmentation on the prosodic prompt text based on a pre-trained language model and extract key prosodic feature labels, wherein the key prosodic feature labels include speech rate labels, intonation pattern labels and key segment markers. The speech rate quantization submodule is used to call the speech rate template library and perform speech rate quantization on the speech rate tags according to the speech rate template library to obtain the prosodic temporal speech rate curve. The intonation trajectory modeling submodule is used to model the intonation trajectory of the intonation pattern label based on the fundamental frequency trajectory generator and create an intonation curve. The pause position marking submodule is used to mark the pause positions of the key segment markers and obtain the pause duration. The acoustic parameter combination acquisition submodule is used to read the prosodic template library and acquire the acoustic parameter combination corresponding to the prosodic temporal speech rate curve, the intonation curve and the pause duration from the prosodic template library. The prosodic feature extraction submodule is used to extract prosodic features from the acoustic parameter combination to obtain the prosodic feature vector.

[0096] In some optional implementations of the embodiments of this application, the above-mentioned prosodic feature extraction submodule includes: The fundamental frequency sequence calculation unit is used to calculate the fundamental frequency of each frame of the acoustic parameter combination according to the acoustic model, so as to obtain the fundamental frequency sequence; The frame-by-frame speech rate calculation unit is used to calculate the frame-by-frame speech rate of the acoustic parameter combination to obtain the speech rate timing curve. The pause position matrix acquisition unit is used to identify the pause positions of the acoustic parameter combination and quantify the duration to obtain the pause position matrix; A sequence splicing unit is used to splice the fundamental frequency sequence, the speech rate timing curve, and the pause position matrix to obtain a spliced ​​sequence; The prosodic feature vectorization unit is used to input the concatenated sequence into the text vectorization model for nonlinear mapping and vectorization to obtain the prosodic feature vector.

[0097] In some optional implementations of the embodiments of this application, the above-mentioned speech conversion module includes: The feature vector transformation submodule is used to convert the fused feature vector into a fused feature Mel spectrum according to the nonlinear mapping network; The spectrum smoothing submodule is used to perform spectrum smoothing on the fused feature Mel spectrum according to the Viterbi algorithm to obtain a smoothed Mel spectrum. The waveform conversion submodule is used to input the smoothed Mel spectrum into the streaming decoder for waveform conversion to obtain the target synthesized speech.

[0098] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of a computer device according to an embodiment of this application.

[0099] Computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected via a system bus. It should be noted that only computer device 300 with components 310-330 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0100] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0101] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as the hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device 300. In the embodiments of this application, the memory 310 is typically used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for speech synthesis methods. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.

[0102] In some embodiments, processor 320 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 320 is typically used to control the overall operation of computer device 300. In embodiments of this application, processor 320 is used to execute computer-readable instructions stored in memory 310 or process data, such as executing computer-readable instructions for a speech synthesis method.

[0103] The network interface 330 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 300 and other electronic devices.

[0104] The computer device provided in this application can extract features from voice prompt text from three dimensions: timbre, emotion, and rhythm, and effectively fuse the extracted features to synthesize target speech with specific timbre, rich emotion, and natural rhythm, so as to meet users' needs for high-quality speech synthesis.

[0105] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech synthesis method described above.

[0106] The computer-readable storage medium provided in this application can extract features from voice prompt text from three dimensions: timbre, emotion, and rhythm, and effectively fuse the extracted features to synthesize target speech with specific timbre, rich emotion, and natural rhythm, so as to meet users' needs for high-quality speech synthesis.

[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0108] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A speech synthesis method, characterized in that, Includes the following steps: Obtain the voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text, and rhythm prompt text; The timbre prompt text is input into the timbre encoder for timbre encoding to obtain the timbre feature vector; The emotional prompt text is input into the emotional encoder for emotional encoding to obtain an emotional feature vector; The prosodic prompt text is input into the prosodic encoder for prosodic encoding to obtain the prosodic feature vector; The timbre feature vector, the emotion feature vector, and the prosody feature vector are fused to obtain a fused feature vector. The fused feature vector is input into a vocoder for speech conversion to obtain the target synthesized speech.

2. The speech synthesis method according to claim 1, characterized in that, The step of inputting the timbre prompt text into the timbre encoder for timbre encoding to obtain the timbre feature vector specifically includes the following steps: The timbre prompt text is semantically segmented based on a pre-trained language model to extract key timbre feature labels. Read the acoustic knowledge base and obtain the acoustic feature template corresponding to the timbre key feature label from the acoustic knowledge base; Based on the acoustic feature template, the timbre prompt text is preprocessed for timbre synthesis to obtain a reference timbre segment; The MFCC feature coefficients of the reference timbre segment are extracted using the Mel frequency cepstral coefficient algorithm; The fundamental frequency trajectory parameters of the reference timbre segment are extracted based on the acoustic model; The MFCC feature coefficients and the fundamental frequency trajectory parameters are input into the input text vectorization model to perform timbre nonlinear mapping, thereby obtaining the timbre feature vector.

3. The speech synthesis method according to claim 1, characterized in that, The step of inputting the emotional prompt text into the emotional encoder for emotional encoding to obtain the emotional feature vector specifically includes the following steps: The sentiment prompt text is segmented and part-of-speech tagged to obtain sentiment prompt word segments; The emotional prompt context is obtained by extracting the context of the emotional prompt words based on the sliding window; The emotional features are extracted from the emotional cue context to obtain the emotional feature vector.

4. The speech synthesis method according to claim 3, characterized in that, The step of extracting emotional features from the emotional cue context to obtain the emotional feature vector specifically includes the following steps: The emotional cue context is input into a bidirectional encoding classifier to obtain the emotional category probability of each emotional category, and the emotional category probability is converted into one-hot encoding; The emotional intensity value of each emotional category is calculated according to the regression model, and the emotional intensity value is converted into an emotional intensity sequence. Contextual features of the emotional cues are obtained based on an attention mechanism; The one-hot encoding, the sentiment intensity sequence, and the contextual features are input into a nonlinear mapping model to fuse sentiment features, thereby obtaining the sentiment feature vector.

5. The speech synthesis method according to claim 1, characterized in that, The step of inputting the prosodic prompt text into the prosodic encoder for prosodic encoding to obtain the prosodic feature vector specifically includes the following steps: The prosodic prompt text is segmented into prosodic semantics based on a pre-trained language model, and key prosodic feature labels are extracted. These key prosodic feature labels include speech rate labels, intonation pattern labels, and key segment markers. Call the speech rate template library, and quantify the speech rate of the speech rate tag according to the speech rate template library to obtain the prosodic temporal speech rate curve; Based on the fundamental frequency trajectory generator, the intonation pattern label is modeled to create an intonation curve; The key segment markers are used to mark the pause positions, and the pause duration is obtained; Read the prosodic template library and obtain the acoustic parameter combination corresponding to the prosodic temporal speech rate curve, the intonation curve and the pause duration from the prosodic template library; The prosodic feature vector is obtained by extracting prosodic features from the acoustic parameter combination.

6. The speech synthesis method according to claim 5, characterized in that, The step of extracting prosodic features from the acoustic parameter combination to obtain the prosodic feature vector specifically includes the following steps: The fundamental frequency of each frame of the acoustic parameter combination is calculated based on the acoustic model to obtain the fundamental frequency sequence; Calculate the frame-by-frame speech rate of the acoustic parameter combination to obtain the speech rate time-series curve; Identify the pause positions of the acoustic parameter combinations and quantify their durations to obtain a pause position matrix; The fundamental frequency sequence, the speech rate timing curve, and the pause position matrix are concatenated to obtain a concatenated sequence. The concatenated sequence is input into a text vectorization model for nonlinear mapping and vectorization to obtain the prosodic feature vector.

7. The speech synthesis method according to claim 1, characterized in that, The step of inputting the fused feature vector into a vocoder for speech conversion to obtain the target synthesized speech specifically includes the following steps: The fused feature vector is converted into a fused feature Mel spectrum based on a nonlinear mapping network; The fused feature Mel spectrum is smoothed using the Viterbi algorithm to obtain a smoothed Mel spectrum; The smoothed Mel spectrum is input into a streaming decoder for waveform conversion to obtain the target synthesized speech.

8. A speech synthesis device, characterized in that, include: The prompt text acquisition module is used to acquire voice prompt text, wherein the voice prompt text includes timbre prompt text, emotion prompt text, and rhythm prompt text; The timbre encoding module is used to input timbre prompt text into the timbre encoder for timbre encoding, thereby obtaining a timbre feature vector; The sentiment encoding module is used to input sentiment prompt text into the sentiment encoder for sentiment encoding, and obtain sentiment feature vectors. The prosody coding module is used to input prosody cue text into the prosody encoder for prosody coding, and obtain prosody feature vectors. The feature fusion module is used to fuse the timbre feature vector, the emotion feature vector, and the prosody feature vector to obtain a fused feature vector. The speech conversion module is used to input the fused feature vector into the vocoder for speech conversion to obtain the target synthesized speech.

9. A computer device, comprising a memory and a processor, characterized in that, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of the speech synthesis method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech synthesis method as described in any one of claims 1 to 7.