Voice synthesis method, device and equipment based on potential rhythm of diffusion and medium

By adopting a potential rhythmic method based on diffusion in the speech synthesis technology, the problems of phonetic mechanics and lack of emotion in the prior art are solved, and more natural and emotionally rich speech synthesis is achieved, improving the user experience.

CN120220642APending Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510441823.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing speech synthesis technologies are difficult to effectively capture and express the ups and downs of speech, resulting in a mechanical feeling of synthetic speech, lack of emotion, and poor user experience.

Method used

A speech synthesis method based on diffusion-based latent prologue is used to generate more natural and emotionally rich speech through technologies such as hidden coding processing, Mel spectrogram extraction, reference and analytical pronunciation vector construction, vector quantization, etc.

Benefits of technology

It significantly improves the generation speed and smoothness of speech synthesis, enhances the naturalness and emotional expression of speech, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220642A_ABST
    Figure CN120220642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech synthesis, financial science and technology and medical health, and discloses a speech synthesis method, device and equipment of potential rhythm based on diffusion and a medium, and the method comprises the steps: generating a reference rhythm vector corresponding to each frequency band of an initial audio according to a Mel spectrogram and a preset real phoneme duration; according to the text hidden representation and the speaker hidden representation, constructing an analysis phoneme duration corresponding to the reference rhythm vector; constructing an analysis rhythm vector by using the text hidden representation, the speaker hidden representation and the analysis phoneme duration; generating a potential rhythm vector according to the reference rhythm vector and the analysis rhythm vector; and performing speech synthesis by using the text hidden representation, the speaker hidden representation and the potential rhythm vector to obtain a synthesized speech. The time step number is remarkably reduced through model analysis, the generation speed is increased, and meanwhile, the smoothness of the generated voice is remarkably improved through rhythm vector quantization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of speech synthesis, fintech, and healthcare, and particularly to a speech synthesis method, apparatus, device, and medium based on diffusion-based latent prosody. Background Art

[0002] Currently, many financial platforms and healthcare platforms on the market support speech broadcasts of specific content and accurate speech consultation responses based on robot interaction technology. For example, when a user initiates a consultation in a healthcare APP, the consultation question can be autonomously replied to by reading aloud; similarly, a financial APP belonging to a financial platform can also reply to consultation questions by autonomously reading aloud and also has a specific content speech broadcast function. Such autonomous speech interaction methods usually need to obtain synthesized speech according to text and prosody. To achieve this goal, some existing solutions research directions are to design prosody features at coarse-grained and fine-grained levels and construct a hierarchical model.

[0003] However, due to the inherent correlation between prosody features, separately modeling these features may result in synthesized speech being very mechanical and lacking emotion, and it is difficult to achieve the effect that should originally have cadence in human voices, with a mediocre user experience. This undoubtedly greatly reduces the final effect of audiobooks, making many people unacceptable and losing some users. Summary of the Invention

[0004] The present invention provides a speech synthesis method, apparatus, device, and medium based on diffusion-based latent prosody, which significantly reduces the number of time steps through model analysis, improves the generation speed, and at the same time, through prosody vector quantization, the synthesized speech has a significant improvement in smoothness.

[0005] In a first aspect, a speech synthesis method based on diffusion-based latent prosody is provided, including:

[0006] Performing hidden encoding processing on the text to be synthesized to obtain a text hidden representation;

[0007] Generating an initial audio according to the text to be synthesized, and extracting a Mel spectrogram of the initial audio;

[0008] Generating a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration;

[0009] Obtaining a speaker hidden representation of a preset speech segment, and constructing an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation;

[0010] Constructing an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration;

[0011] Generate a potential prosody vector based on the reference prosody vector and the analyzed prosody vector;

[0012] Perform voice synthesis using the text hidden representation, the speaker hidden representation, and the potential prosody vector to obtain synthesized speech.

[0013] In a second aspect, a voice synthesis device based on diffusion-based potential prosody is provided, including:

[0014] An encoding processing module for performing hidden encoding processing on the text to be synthesized to obtain a text hidden representation;

[0015] An extraction module for extracting the Mel spectrogram of the initial audio;

[0016] A vector generation module for generating a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration;

[0017] A construction module for constructing an analyzed phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation;

[0018] A vector construction module for constructing an analyzed prosody vector by using the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration.

[0019] A generation module for generating a potential prosody vector based on the reference prosody vector and the analyzed prosody vector.

[0020] A voice synthesis module for performing voice synthesis using the text hidden representation, the speaker hidden representation, and the potential prosody vector to obtain synthesized speech.

[0021] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned voice synthesis method based on diffusion-based potential prosody are implemented.

[0022] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned voice synthesis method based on diffusion-based potential prosody are implemented.

[0023] In the solution implemented by the above-mentioned speech synthesis method, device, computer device and storage medium based on diffusion-based latent prosody, the text hidden representation helps to enhance the generalization ability of the speech synthesis model. When dealing with unseen text data, the model can quickly learn and infer its implicit features through the text hidden representation, thereby improving the accuracy of prediction and synthesis. By constructing a reference prosody vector, we can fuse speech samples from different sources to improve the consistency and accuracy of their prosody and acoustic features. This can further improve the performance and effect of the prosody analysis and speech synthesis model, enabling it to better meet the needs of users. Through model analysis, the number of time steps is significantly reduced, improving the generation speed. Vector quantization of the initial prosody vector can map high-dimensional vectors to lower-dimensional codebook vectors, thereby improving storage and processing efficiency. Through vector quantization, less important vector components can be filtered out, a small amount of important information can be extracted, and this information can be compressed into a small codebook, reducing information redundancy. Compressing the original high-dimensional initial prosody vector into a codebook vector through vector quantization makes it easier for people to understand and analyze the key features mined in the prosody vector. At the same time, through vector quantization, the smoothness of the generated speech has been significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 FIG. is a schematic diagram of an application environment of a speech synthesis method based on diffusion-based latent prosody in an embodiment of the present invention;

[0026] Figure 2 FIG. is a schematic flowchart of a speech synthesis method based on diffusion-based latent prosody in an embodiment of the present invention;

[0027] Figure 3 FIG. is a schematic structural diagram of a speech synthesis device based on diffusion-based latent prosody in an embodiment of the present invention;

[0028] Figure 4 FIG. is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0029] Figure 5 FIG. is another schematic structural diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] A voice synthesis method based on diffusion-based latent prosody provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server through a network. The server can perform hidden encoding processing on the text to be synthesized to obtain a text hidden representation; generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio; generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration; obtain a speaker hidden representation of a preset speech segment, and construct an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation; construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration; generate a latent prosody vector according to the reference prosody vector and the analysis prosody vector; perform voice synthesis by using the text hidden representation, the speaker hidden representation, and the latent prosody vector to obtain a synthesized voice, and feedback the synthesized voice to the client. The present invention provides a voice synthesis device based on diffusion-based latent prosody. For the synthesized voice service, perform hidden encoding processing on the text to be synthesized to obtain a text hidden representation; generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio; generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration; obtain a speaker hidden representation of a preset speech segment, and construct an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation; construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration; generate a latent prosody vector according to the reference prosody vector and the analysis prosody vector; perform voice synthesis by using the text hidden representation, the speaker hidden representation, and the latent prosody vector to obtain a synthesized voice. Through model analysis, the number of time steps is significantly reduced, and the generation speed is improved. At the same time, through prosody vector quantization, the smoothness of the generated voice is significantly improved. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0032] Please refer to Figure 2 as shown inFigure 2 FIG. 1 is a schematic flow chart of a voice synthesis method based on diffusion-based latent prosody provided by an embodiment of the present invention, including the following steps:

[0033] S1. Perform hidden coding processing on the text to be synthesized to obtain a text hidden representation.

[0034] In the embodiment of the present invention, the hidden coding processing refers to converting the text to be synthesized into a mathematically vector representation that has been encoded and compressed, which is the text hidden representation.

[0035] Specifically, the hidden coding processing usually involves deep learning technologies, such as models like Recurrent Neural Network. First, the text to be synthesized needs to be converted into a sequence of words, characters, or phonemes. Then, the corresponding deep learning model is used to encode and compress these sequences to obtain the text hidden representation.

[0036] In the embodiment of the present invention, performing hidden coding processing on the text to be synthesized to obtain a text hidden representation includes:

[0037] Dividing the text to be synthesized into multiple segments of words;

[0038] Converting the text to be synthesized into a phoneme sequence and a word segmentation sequence according to the word segmentation;

[0039] Concatenating the phoneme sequence and the word segmentation sequence to obtain a combined sequence corresponding to the word segmentation;

[0040] Encoding each of the combined sequences one by one to obtain a text hidden representation.

[0041] In the embodiment of the present invention, the division refers to splitting a text into multiple segments of words, the conversion refers to generating a sequence of corresponding words and vocabulary for the text to be synthesized through word segmentation processing, and converting each word or vocabulary into a corresponding phoneme sequence. The concatenation refers to, after performing word segmentation and phoneme conversion on the input text, concatenating the phoneme sequences corresponding to each word segmentation into a whole sequence in a certain order. The encoding refers to converting the combined sequence into a digital form that can be processed by a computer.

[0042] Specifically, the text is segmented according to pre-set word segmentation rules, such as segmentation based on spaces or punctuation marks; the statistical-based word segmentation method learns from a large corpus, calculates the probability and position information of word occurrences for word segmentation; the deep learning-based word segmentation method uses models such as neural networks for recognition and reasoning to achieve more accurate word segmentation; each word or vocabulary is converted into a corresponding phoneme sequence. By extracting the phoneme sequences corresponding to each segmented word one by one and then using a deep learning model for encoding. Common encoding models include Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM), etc. These models can learn the features and relationships in the phoneme sequences corresponding to each segmented word, and then obtain the corresponding text hidden representation.

[0043] In the speech synthesis scenario in the medical and health field, speech synthesis needs to express highly understandable medical information. Using the text hidden representation can accurately capture key information, and applying the text hidden representation during speech synthesis makes the speech more naturally express medical information and improves the comprehensibility. At the same time, during the patient's inquiry process, the corresponding diagnostic advice is given as a reply to the patient's inquiry by means of self-reading. When dealing with the speech prosody of the diagnostic advice that the patient needs to focus on, the "stress" effect can be increased to attract the patient's attention.

[0044] Similarly, in the financial industry, when processing various customer information and services through speech synthesis technology, since different speech synthesis technologies can lead to significant differences in the speech broadcast effect, it may cause different decisions to be made by transaction managers when processing transactions. If the broadcast of the whole article is full of mechanical feeling and lacks highlighting of key points, it may lead to transaction managers missing important details and making work mistakes. However, using the text hidden representation can accurately capture key information and apply it in speech synthesis, making the speech more naturally express financial information, thus greatly reducing the work mistakes of transaction managers.

[0045] In the embodiments of the present invention, by dividing the text into multiple groups of segmented words and splicing the phoneme sequences corresponding to the segmented words and the segmented word sequences, the combined sequence corresponding to the segmented words contains more detailed and rich text information, which helps to reduce errors caused by information loss during the speech synthesis process and improve the accuracy and naturalness of speech synthesis; encoding the combined sequence corresponding to each group of segmented words to obtain the corresponding text hidden representation, which can divide the originally large-scale text information into multiple small parts, thereby reducing the model calculation amount and model complexity.

[0046] In the embodiments of the present invention, the text hidden representation can capture the abstract features of the text and then provide them to the acoustic model for processing, which helps to improve the processing ability of the acoustic model for different types of input data, thereby improving the effect of speech synthesis. The text hidden representation can capture the abstract features of the text and then provide them to the acoustic model for processing; at the same time, it helps to enhance the generalization ability of the speech synthesis model. When processing unseen text data, the model can quickly learn and infer its implicit features through the text hidden representation, thereby improving the accuracy of prediction and synthesis.

[0047] S2. Generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio.

[0048] In the embodiments of the present invention, the generation refers to converting the text into a speech signal, and the extraction refers to processing the initial audio into the form of Mel Frequency Cepstral Coefficients (MFCC).

[0049] Specifically, according to the preprocessed text information, speech synthesis technology is used to generate the initial audio. The initial audio can be the original speech waveform or the speech waveform synthesized by a speech synthesis model; the initial audio is segmented into a series of shorter audio segments, and each audio segment is windowed with a Hamming window to reduce the edge effect. Each audio segment is Fourier-transformed in the time domain to convert it into an energy distribution diagram in the frequency domain, and the obtained spectral data is mapped to the Mel frequency scale to obtain the corresponding Mel spectrogram.

[0050] In the embodiments of the present invention, the extraction of the Mel spectrogram of the initial audio includes:

[0051] Extract the speech features in the initial audio;

[0052] Calculate the Mel spectral coefficients for each of the speech features to obtain the Mel spectral coefficients corresponding to the speech features;

[0053] Stack the Mel spectral coefficients to obtain a Mel spectrogram.

[0054] In the embodiments of the present invention, the extraction refers to extracting relevant features from the speech signal for tasks such as speech signal recognition, synthesis, and conversion. The Mel spectral coefficient calculation refers to converting the speech signal to the Mel frequency scale and calculating the corresponding Mel spectral coefficients. The matrix stacking is to stack the Mel spectral coefficients of multiple audio frames in a certain order to obtain a two-dimensional matrix.

[0055] Specifically, speech features refer to parameters or algorithms that can reflect the characteristics and structures of speech signals in different aspects. Common speech features include duration, fundamental frequency, formants, Mel spectrogram, and so on. These features that can characterize speech information can be used in fields such as speech recognition, speaker recognition, emotion recognition, speech synthesis, etc. Extracting the speech features in the initial audio means extracting useful features from the original audio signal that can characterize the speech information of the audio. In speech signal preprocessing and feature extraction, the spectral features in speech features are represented using the Mel frequency scale, and then the discrete cosine transform is performed on the signal under the Mel frequency scale to obtain the Mel spectrogram coefficients. After arranging the Mel spectrogram coefficient matrix in a certain regular order, a two-dimensional image representing the spectral features of the speech signal is obtained; usually, the rule and order of matrix stacking refer to stacking the Mel spectrogram coefficients of multiple frames, where each row represents a component of the Mel spectrogram coefficient and each column represents a frame.

[0056] Specifically, the Mel spectrogram is a visualization tool for representing the spectral features of speech signals. It can help researchers more comprehensively understand the characteristics of speech signals in the frequency domain and can also be used in tasks such as classification, recognition, and analysis. Compared with the Mel spectrogram coefficient matrix, the Mel spectrogram is more intuitive and easier to understand, and can intuitively display the spectral differences between different speech signals.

[0057] In the embodiments of the present invention, by calculating the Mel spectrogram coefficients corresponding to the speech features one by one, the frequency domain features of the speech signal can be extracted for subsequent processing and analysis; at the same time, the high-dimensional speech signal is converted into a low-dimensional vector form, reducing the processing difficulty. By stacking these Mel spectrogram coefficients to form a Mel spectrogram, some important features can be more easily extracted, which also helps to enhance the pattern recognition ability. At the same time, the original sound signal is converted into a feature matrix for machine learning tasks, and the spectral features of the sound signal can be intuitively represented, facilitating machine learning tasks and pattern recognition and improving data processing efficiency.

[0058] In the embodiments of the present invention, extracting the Mel spectrogram of the initial audio can convert the speech signal into mathematical features, thereby helping the model obtain more useful information in the speech; because the Mel spectrogram can accurately reflect the intensity distribution of speech at different frequencies, it is a commonly used speech feature extraction method that helps to improve the quality of synthesized speech. The Mel spectrogram is a mathematical feature in matrix form, with strong visualization and scalability, and can be well applied to various models such as neural networks; by extracting the Mel spectrogram as the input of the model, model training and optimization can be conveniently carried out.

[0059] S3. Generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration.

[0060] In the embodiment of the present invention, the generation refers to segmenting the initial audio sampling according to the Mel spectrogram and the true phoneme duration, and then calculating the prosody vector corresponding to each frequency band.

[0061] Specifically, during the process of synthesizing speech, the preset true phoneme duration can be used to divide the speech signal into several relatively uniform time periods (such as a single phoneme); then, by statistically analyzing the Mel spectrogram within each time period, the average value and variance of the Mel spectrogram coefficients within this time period can be obtained, thereby forming a vector describing the prosody characteristics of this frequency band, which is the reference prosody vector corresponding to this frequency band.

[0062] In detail, first, the speech data needs to be preprocessed, including: phoneme alignment (realizing the alignment of phoneme text and speech using natural language processing and automatic alignment tools), and dividing the Mel spectrogram into multiple small segments according to the preset true phoneme duration. The prosody encoder is used to extract the reference prosody vector of each data segment. Therefore, a deep learning model needs to be designed, taking the phoneme duration and the Mel spectrogram as inputs. The deep learning model can adopt a convolutional neural network. Input the preset true phoneme duration and the Mel spectrogram of each data segment into the convolutional neural network, and output the reference prosody vector of each data segment. Through the trained prosody encoder, predict the prosody vector of the frequency band in the initial audio.

[0063] In the analysis of synthetic speech in the field of medical health, speech synthesis technology can be used to reply to patients' consultation questions. This type of autonomous speech interaction method obtains synthetic speech according to text and prosody. However, currently, all the content replied by this speech interaction method when answering patients' questions is full of mechanical feeling, resulting in a poor interaction experience and low user experience and satisfaction with the APP. By adjusting the prosody of the speech through the reference prosody vector, the user experience is made good, and the comfort of the interaction process is increased.

[0064] Similarly, when applying speech synthesis technology to financial field services, when using speech technology to reply to customers' consultation questions, the speech recognition system sometimes has difficulty correctly identifying the key information in the speech and giving corresponding answers, which may lead to customers' dissatisfaction with the question answers and may also guide customers to make wrong judgments, causing economic losses to customers; the reference prosody vector corresponding to each frequency band of the initial audio can improve the accuracy and reliability of speech recognition, and at the same time, the accuracy of speech answers can also be improved through noise reduction, volume adjustment, sentence segmentation, etc.

[0065] S4. Obtain the speaker hidden representation of the preset speech segment, and construct the analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation.

[0066] In the embodiments of the present invention, the obtaining refers to extracting the specific speaker-related characterization hidden therein from the given speech data, and the constructing refers to using the text hidden representation and the speaker hidden representation to construct the length occupied by each phoneme in time.

[0067] Specifically, use the preset speech segment to obtain the characteristic information of the speaker in this speech segment, such as timbre, pronunciation style, etc., and obtain the hidden representation of this speaker from it, and then combine the obtained speaker hidden representation with the text hidden representation to construct a reference prosody vector.

[0068] In the embodiments of the present invention, the constructing the analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation includes:

[0069] Perform phoneme-level expansion on the text hidden representation and the speaker hidden representation respectively to obtain the text phoneme level corresponding to the text hidden representation and the speaker phoneme level corresponding to the speaker hidden representation;

[0070] Integrate the text phoneme level corresponding to the text hidden representation and the speaker phoneme level corresponding to the speaker hidden representation into a comprehensive input variable;

[0071] Perform phoneme duration analysis on the comprehensive input variable to obtain the analysis phoneme duration corresponding to the reference prosody vector.

[0072] In the embodiments of the present invention, the phoneme-level expansion refers to expanding the text hidden representation and the speaker hidden representation to the phoneme level on the basis of them, that is, generating a corresponding hidden representation for each phoneme, and the phoneme duration analysis refers to using the information in the comprehensive input variable to analyze and measure the length occupied by each phoneme in time.

[0073] Specifically, the text phoneme level expansion corresponding to the text hidden representation is to separate the hidden representation of the given text into multiple phoneme-level representations, and each phoneme-level representation corresponds to a text phoneme. In this way, a set of text phoneme level representation vectors containing all text phonemes is obtained, and this set can be used in tasks such as prosody analysis and phoneme recognition; the speaker phoneme level expansion corresponding to the speaker hidden representation is to expand the speaker hidden representation to each phoneme-level representation, that is, generate a corresponding speaker hidden representation for each phoneme. In this way, the speech of the speaker can be more finely modeled according to different phoneme characteristics to improve the performance of tasks such as speaker recognition.

[0074] Specifically, obtaining the phoneme duration analysis in the analysis phoneme duration corresponding to the reference prosody vector means that when generating the reference prosody vector, the duration of each phoneme has been predicted. These duration information will be used as part of the reference prosody vector and applied to tasks such as prosody analysis and speech synthesis. By combining the phoneme duration prediction model with complex models such as the prosody encoder and inputting the comprehensive input variables into the phoneme duration prediction model, a set of rich phoneme duration features can be obtained. These features can be used to generate the reference prosody vector, perform more refined modeling and processing on the speech signal, and thus improve the performance of tasks such as speech generation and natural language processing.

[0075] In addition, there is also a length regulator LR, which uses the phoneme duration to expand the comprehensive input variable h total to the frame level. During the training process, the real duration dur is used, while during the inference process, the predicted duration dur’ is used.

[0076] Length regulator LR:

[0077] - During the training stage, the real phoneme duration dur is used to expand h total :

[0078]

[0079] - During the inference stage, the analyzed phoneme duration dur' is used to expand h total :

[0080]

[0081] - Among them, dur' is the phoneme duration analyzed by the phoneme duration prediction model.

[0082] In the analysis of synthetic speech in the field of medical and health, doctors can use speech synthesis technology to record the condition briefings and medical guidance, or use speech synthesis technology to reply to patients' consultation questions. In speech synthesis technology, by reasonably setting the speech phoneme duration, the naturalness of automatic speech broadcasting can be improved. At the same time, freely adjusting the phoneme duration can avoid the phoneme duration being too short or too long, and improve the user's willingness to use and satisfaction with autonomous speech interaction.

[0083] Similarly, in the field of fintech, financial institutions can provide users with personalized financial consultation services and professional investment advice through autonomous speech interaction technology and autonomous speech systems, and timely provide recommendations for financial products. The analyzed phoneme duration can help the model process the intonation and tone information of different syllables more accurately and flexibly when dealing with autonomous speech interaction, and be closer to artificial voice communication when providing financial consultation services and professional investment advice to customers, bringing a good experience to customers.

[0084] In the embodiments of the present invention, after extending the hidden representation to the phoneme level, each phoneme has a corresponding hidden representation, which can more precisely express the information contained in each phoneme. In this way, the model can more accurately capture the characteristic differences between different phonemes and the individual differences between the voices of different speakers, thereby improving the ability to model speech signals. At the same time, it can help the model process information such as intonation and tone of different syllables more accurately and flexibly during prosody analysis. This is particularly important for improving the quality and fluency of tasks such as natural language generation and speech conversion.

[0085] In the embodiments of the present invention, by constructing a reference prosody vector, we can fuse speech samples from different sources, improving the consistency and accuracy of their prosody and acoustic features. This can further improve the performance and effect of prosody analysis and speech synthesis models, enabling them to better meet the needs of users. Since the reference prosody vector is composed of multiple preset speech samples, it not only contains the speech information in the training set but can also well adapt to new test data, which can improve the generalization ability and applicability of the model, making its application range more extensive.

[0086] S5. Construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration.

[0087] In the embodiments of the present invention, the construction refers to combining the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration in a certain way to construct a vector representation that can express speech prosody information.

[0088] Specifically, the text hidden representation can be used to extract the semantic information of speech and express features such as the words and sentence structure of speech; the speaker hidden representation can be used to extract the speaker personality information of speech and express features such as the voice characteristics and styles of speech; and the analyzed phoneme duration can be used to extract the prosody and acoustic features of speech and express features such as the rhythm, pitch, and intensity of speech. By combining the above three hidden representations, a more comprehensive and accurate speech prosody vector can be obtained.

[0089] In the embodiments of the present invention, the constructing the analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration includes:

[0090] Generate an initial prosody vector according to the combination of the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration;

[0091] Perform vector quantization on the initial prosody vector to obtain the analysis prosody vector.

[0092] In the embodiments of the present invention, the combination generation refers to combining the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration in a certain manner to construct an initial prosody vector that can express the speech prosody information. The vector quantization refers to a method of representing a high-dimensional initial prosody vector with a lower-dimensional codeword.

[0093] Specifically, when processing speech signals, it is often necessary to extract some key feature vectors for subsequent processing. These feature vectors are composed of multiple components. If the original vector form is adopted, the processing machine or algorithm will be greatly restricted. The method of vector quantization is precisely to optimize the performance of feature vectors in storage and processing. Its basic idea is to input all the feature vectors into a specific clustering algorithm, cluster similar vectors into one category, and the central vector of each category is called the codeword in the codebook.

[0094] Specifically, in the construction of the analyzed prosody vector, the method of vector quantization for the initial prosody vector can make the high-dimensional vector be represented as a low-dimensional codeword in a codebook, thereby reducing the complexity and storage overhead, that is, mapping each sub-vector to a vector representation in the codebook.

[0095] In the embodiments of the present invention, the vector quantization of the initial prosody vector to obtain the analyzed prosody vector includes:

[0096] Obtain a codebook, where the codebook includes multiple groups of elements;

[0097] Calculate the Euclidean distance between each initial prosody vector and each element in the codebook to obtain the calculation results of each initial prosody vector and each element in the codebook;

[0098] Select the codebook element with the smallest Euclidean distance from the calculation results, and collect the codebook elements to generate a codebook element set;

[0099] Generate the analyzed prosody vector according to the codebook element set.

[0100] In the embodiments of the present invention, the obtaining refers to the process required to extract specific low-dimensional vectors from the codebook, and the Euclidean distance calculation refers to the process of calculating the Euclidean distance between each initial prosody vector and all vectors in the codebook.

[0101] Specifically, the goal of the vector quantization layer Z is to find the codebook element z i,pros closest to each prosody vector h k . This process can be achieved by minimizing the Euclidean distance:

[0102]

[0103] where zi,pros is the quantized prosody vector (analyzing prosody vector), and ||·||2 represents the Euclidean distance.

[0104] Codebook Z:

[0105] - The codebook Z is a set containing K codeword elements, and each codeword element z k has a dimension of d z :

[0106] Z = {z1, z2, …, z K},

[0107] By performing vector quantization on all initial prosody vectors, an analyzed prosody vector sequence is obtained:

[0108] z pors = {z 1,pros , z 2,pros , …, z i,pros}

[0109] In the embodiments of the present invention, performing vector quantization on the initial prosody vectors can map high-dimensional vectors into lower-dimensional codebook vectors, thereby improving the storage and processing efficiency. Through vector quantization, less important vector components can be filtered out, a small amount of important information can be extracted, and this information can be compressed into a small codebook, reducing information redundancy. By compressing the original high-dimensional initial prosody vectors into codebook vectors through vector quantization, it is easier to help people understand the key features mined in the analyzed prosody vectors.

[0110] In the embodiments of the present invention, by combining the semantic information of text hidden representation, the prosody features of speaker hidden representation, and the syllable-level timing information of analyzing phoneme durations, the prosody features of speech signals can be more accurately described, and the accuracy of prosody performance can be improved; at the same time, speech signals can be more comprehensively analyzed, and some defects in traditional analysis methods can be solved; using the analyzed prosody vector as the feature representation of speech signals can improve the performance and accuracy of speech recognition, thereby improving the practicality and effect of speech recognition technology in the application field.

[0111] S6. Generate a potential prosody vector according to the reference prosody vector and the analyzed prosody vector.

[0112] In the embodiments of the present invention, the generation refers to the process of mapping a high-dimensional prosody vector into a low-dimensional vector space and retaining important prosody features and information as much as possible.

[0113] Specifically, the reference prosody vector and the analyzed prosody vector will be fed into a network of a generative model, such as a variational autoencoder or a generative adversarial network, to generate corresponding latent prosody vectors. In this process, each prosody vector in the feature space can be represented as a latent vector, and the latent vector can be used for tasks such as measuring the similarity between different representations, performing data normalization, and calculating the distance between vectors.

[0114] In an embodiment of the present invention, generating the latent prosody vector according to the reference prosody vector and the analyzed prosody vector includes:

[0115] Performing normalization processing on the reference prosody vector and the analyzed prosody vector respectively to obtain a processing result corresponding to the reference prosody vector and a processing result corresponding to the analyzed prosody vector;

[0116] Calculating the error between the processing result corresponding to the reference prosody vector and the processing result corresponding to the analyzed prosody vector to obtain an error result;

[0117] Combining the error result with the analyzed prosody vector to generate a latent prosody vector.

[0118] In an embodiment of the present invention, the normalization processing refers to the process of scaling the values in these vectors to a specific range according to a ratio, the calculation refers to the process of comparing the differences between the values of these two vectors after normalization processing, and further evaluating the similarity between them, and the combined generation refers to the process of taking the error result as an additional input feature, together with the analyzed prosody vector, and inputting them into the generative model to generate a new latent prosody vector.

[0119] Specifically, find the minimum value and the maximum value in the reference prosody vector and the analyzed prosody vector. For the reference prosody vector and the analyzed prosody vector, the values in these vectors are usually speech-related features such as pseudo power spectrum, fundamental frequency, and energy. Use the minimum value and the maximum value to perform a linear mapping on each value in the vector to limit their values to a specific range. This range is usually [0, 1] or [-1, 1], which is determined according to actual needs. Use the normalized vector as the processing result for subsequent prosody feature modeling and analysis.

[0120] Specifically, obtain the processing result corresponding to the reference prosody vector and the processing result corresponding to the analyzed prosody vector. These results are obtained from the previous processing steps and are the values of two normalized vectors. Define a distance metric, such as Euclidean distance or Manhattan distance. These distance metrics can be used to measure the similarity between two vectors.

[0121] In the synthetic speech analysis in the field of medical and health, some content on the medical platform for treating insomnia and soothing the mood adopts the voice broadcast method. By applying the potential prosody vector technology to the above voice broadcast, it can better simulate the prosody characteristics generated by human speech, thereby maximizing the naturalness and fluency of speech synthesis. At the same time, it can make the speech more expressive and emotionally appealing, better meeting the humanized needs in the field of medical and health.

[0122] Similarly, in the synthetic speech analysis in the field of fintech, speech synthesis is often used in the recommendation of financial products. Using the potential prosody vector technology can improve the naturalness and readability of speech recommendations, thereby better attracting the attention of users and improving the effect of product recommendations; at the same time, using the potential prosody vector technology can make the speech more vivid and natural, enabling users to better understand and accept the services provided.

[0123] In the embodiments of the present invention, by normalizing the vectors, the dimensionality differences are eliminated, and the similarities or differences between them can be compared more objectively; the normalization process can help improve the robustness of the input vectors, thereby better coping with different input conditions; at the same time, the values of the vectors can be mapped to a specific range, which can better extract features.

[0124] In the embodiments of the present invention, generating the potential prosody vector can combine and optimize the features of the reference prosody vector and the analyzed prosody vector, thereby improving the feature extraction ability. The generation process of the potential prosody vector can be optimized according to the task requirements. For example, prior knowledge, multiple modality information, etc. can be added to improve the performance of the model; generating the potential prosody vector can minimize the influence of factors such as different languages, scenarios, speakers, etc., thereby improving the generalization ability of the model for different datasets and tasks.

[0125] S7. Use the text hidden representation, the speaker hidden representation, and the potential prosody vector for speech synthesis to obtain the synthetic speech.

[0126] In the embodiments of the present invention, the speech synthesis refers to the process of generating artificial speech by using computer programs or algorithms for the text hidden representation, the speaker hidden representation, and the potential prosody vector.

[0127] Specifically, using the text hidden representation and the speaker hidden representation, combined with the potential prosody vector, which can specify the prosody characteristics of the output speech, such as intonation, stress, pause, etc., input the text hidden representation, the speaker hidden representation, and the potential prosody vector into the speech synthesis engine to generate a series of speech signals, and then generate a smooth and natural speech according to these signals. Post-process the synthetic speech, such as denoising, volume adjustment, sentence segmentation, etc., to make it more in line with the natural language expression.

[0128] In the synthetic speech analysis in the field of medical and health, by using the text hidden representation and the speaker hidden representation, and combining with the latent prosody vector, the features of text, prosody and speaker can be more accurately extracted during the speech synthesis process, so as to generate a more natural and fluent synthetic speech, enabling the speech information to be more accurately transmitted to users. And in the field of medical and health, it is often necessary to convey emotions and a sense of trust through speech. Reasonably using the text hidden representation, the speaker hidden representation and the latent prosody vector technology can make the speech more expressive and emotionally contagious, so as to better transmit medical and health information and soothe the patient's mood.

[0129] Similarly, in the synthetic speech analysis in the field of fintech, speech synthesis is often used for automated voice customer service and voice interaction. Using the text hidden representation, the speaker hidden representation and the latent prosody vector technology can generate a more natural and fluent synthetic speech, improving the speech experience. Due to the large volume of business and a large number of users in financial institutions, the speech synthesis processed by the model can not only improve the service efficiency, and the user-friendly voice interaction increases the user experience. At the same time, post-processing the synthetic speech, such as denoising, volume adjustment, sentence segmentation, etc., to make it more in line with the natural language expression way, can make the speech clearer and easier to understand, so that users can more easily accept financial information.

[0130] It can be seen that in the above solution, for the synthetic speech service, the text to be synthesized is subjected to hidden coding processing to obtain the text hidden representation; an initial audio is generated according to the text to be synthesized, and the Mel spectrogram of the initial audio is extracted; a reference prosody vector corresponding to each frequency band of the initial audio is generated according to the Mel spectrogram and the preset true phoneme duration; the speaker hidden representation of a preset speech segment is obtained, and the analysis phoneme duration corresponding to the reference prosody vector is constructed according to the text hidden representation and the speaker hidden representation; an analysis prosody vector is constructed by using the text hidden representation, the speaker hidden representation and the analysis phoneme duration; a latent prosody vector is generated according to the reference prosody vector and the analysis prosody vector; speech synthesis is performed by using the text hidden representation, the speaker hidden representation and the latent prosody vector to obtain a synthetic speech. Through model analysis, the number of time steps is significantly reduced, the generation speed is improved, and at the same time, through prosody vector quantization, the smoothness of the generated speech is significantly improved.

[0131] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0132] In one embodiment, a voice synthesis device based on diffusion-based latent prosody is provided, and the voice synthesis device based on diffusion-based latent prosody corresponds one-to-one with the voice synthesis method based on diffusion-based latent prosody in the above embodiment. As Figure 3 shown, the voice synthesis device based on diffusion-based latent prosody includes an encoding processing module 101, an extraction module 102, a vector generation module 103, a construction module 104, a vector construction module 105, a generation module 106, and a voice synthesis module 107. The detailed description of each functional module is as follows:

[0133] The encoding processing module 101 is configured to perform hidden encoding processing on the text to be synthesized to obtain a text hidden representation;

[0134] The extraction module 102 is configured to extract the Mel spectrogram of the initial audio;

[0135] The vector generation module 103 is configured to generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration;

[0136] The construction module 104 is configured to construct an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation;

[0137] The vector construction module 105 is configured to construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration.

[0138] The generation module 106 is configured to generate a latent prosody vector according to the reference prosody vector and the analysis prosody vector.

[0139] The voice synthesis module 107 is configured to perform voice synthesis by using the text hidden representation, the speaker hidden representation, and the latent prosody vector to obtain a synthesized voice.

[0140] In one embodiment, the encoding processing module 101, when performing hidden encoding processing on the text to be synthesized to obtain a text hidden representation, is used for:

[0141] Dividing the text to be synthesized into multiple groups of words;

[0142] Converting the text to be synthesized into a phoneme sequence and a word segmentation sequence according to the word segmentation;

[0143] Concatenating the phoneme sequence and the word segmentation sequence to obtain a combined sequence corresponding to the word segmentation;

[0144] Encoding each combined sequence one by one to obtain a text hidden representation.

[0145] In one embodiment, the extraction module 102 extracts the Mel spectrogram of the initial audio for:

[0146] extracting the speech features in the initial audio;

[0147] calculating the Mel spectrogram coefficients for each of the speech features to obtain the Mel spectrogram coefficients corresponding to the speech features;

[0148] stacking the Mel spectrogram coefficients into a matrix to obtain a Mel spectrogram.

[0149] In one embodiment, the construction module 104 constructs the analysis phoneme duration corresponding to the reference prosody vector based on the text hidden representation and the speaker hidden representation for:

[0150] performing phoneme-level expansion on the text hidden representation and the speaker hidden representation respectively to obtain the text phoneme level corresponding to the text hidden representation and the speaker phoneme level corresponding to the speaker hidden representation;

[0151] aggregating the text phoneme level corresponding to the text hidden representation and the speaker phoneme level corresponding to the speaker hidden representation into a comprehensive input variable;

[0152] performing phoneme duration analysis on the comprehensive input variable to obtain the analysis phoneme duration corresponding to the reference prosody vector.

[0153] In one embodiment, when the vector construction module 105 constructs an analysis prosody vector using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration, it is used for:

[0154] combining and generating an initial prosody vector according to the text hidden representation, the speaker hidden representation, and the analysis phoneme duration;

[0155] performing vector quantization on the initial prosody vector to obtain an analysis prosody vector.

[0156] When performing vector quantization on the initial prosody vector to obtain an analysis prosody vector, it is used for:

[0157] obtaining a codebook, where the codebook includes multiple groups of elements;

[0158] calculating the Euclidean distance between each of the initial prosody vectors and each of the elements in the codebook to obtain the calculation results of each of the initial prosody vectors and each of the elements in the codebook;

[0159] selecting the codebook elements in the results with the smallest Euclidean distance according to the calculation results, and aggregating the codebook elements to generate a codebook element set;

[0160] Generate an analysis prosody vector according to the set of codebook elements.

[0161] In one embodiment, when generating a potential prosody vector according to the reference prosody vector and the analysis prosody vector, the generation module 106 is configured to:

[0162] Normalize the reference prosody vector and the analysis prosody vector respectively to obtain the processing result corresponding to the reference prosody vector and the processing result corresponding to the analysis prosody vector;

[0163] Calculate the error between the processing result corresponding to the reference prosody vector and the processing result corresponding to the analysis prosody vector to obtain an error result;

[0164] Combine the error result with the analysis prosody vector to generate a potential prosody vector.

[0165] The present invention provides a speech synthesis device based on diffusion-based latent prosody. For the synthetic speech service, perform hidden coding processing on the text to be synthesized to obtain a text hidden representation; generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio; generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration; obtain the speaker hidden representation of a preset speech segment, and construct the analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation; construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration; generate a potential prosody vector according to the reference prosody vector and the analysis prosody vector; perform speech synthesis by using the text hidden representation, the speaker hidden representation, and the potential prosody vector to obtain a synthetic speech. Through model analysis, the number of time steps is significantly reduced, the generation speed is improved, and at the same time, through prosody vector quantization, the smoothness of the generated speech is significantly improved.

[0166] For the specific limitations of a speech synthesis device based on diffusion-based latent prosody, reference can be made to the limitations of a speech synthesis method based on diffusion-based latent prosody in the above text, which will not be elaborated here. Each module in the above speech synthesis device based on diffusion-based latent prosody can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0167] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a voice synthesis method based on diffusion-based latent prosody.

[0168] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a voice synthesis method based on diffusion-based latent prosody.

[0169] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0170] Perform hidden encoding processing on the text to be synthesized to obtain a text hidden representation;

[0171] Generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio;

[0172] Generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration;

[0173] Obtain the speaker hidden representation of a preset speech segment, and construct an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation;

[0174] Construct an analysis prosody vector by using the text hidden representation, the speaker hidden representation, and the analysis phoneme duration;

[0175] Generate a potential prosody vector based on the reference prosody vector and the analyzed prosody vector;

[0176] Perform speech synthesis using the text hidden representation, the speaker hidden representation, and the potential prosody vector to obtain synthesized speech.

[0177] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0178] Perform hidden encoding processing on the text to be synthesized to obtain a text hidden representation;

[0179] Generate an initial audio according to the text to be synthesized, and extract the Mel spectrogram of the initial audio;

[0180] Generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel spectrogram and a preset true phoneme duration;

[0181] Obtain the speaker hidden representation of a preset speech segment, and construct an analyzed phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation;

[0182] Construct an analyzed prosody vector using the text hidden representation, the speaker hidden representation, and the analyzed phoneme duration;

[0183] Generate a potential prosody vector based on the reference prosody vector and the analyzed prosody vector;

[0184] Perform speech synthesis using the text hidden representation, the speaker hidden representation, and the potential prosody vector to obtain synthesized speech.

[0185] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0186] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0188] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. If non-company software tools or components appear in the application embodiments, they are only used for illustrative introduction and do not represent actual use; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A method for speech synthesis based on diffusion potential prosody, characterized in that: include: Perform hidden encoding processing on the synthesized text to obtain a hidden representation of the text; Generate an initial audio according to the text to be synthesized, and extract a Mel-spectrogram of the initial audio; Generate a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel-spectrogram and a preset real phoneme duration; Acquire a speaker hidden representation of a preset speech segment, and construct an analysis phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation; constructing an analysis prosody vector using the text hidden representation, the speaker hidden representation and the analysis phoneme duration; generating a potential prosody vector according to the reference prosody vector and the analyzed prosody vector; Speech synthesis is performed using the text hidden representation, the speaker hidden representation and the latent prosody vector to obtain synthesized speech.

2. The method for speech synthesis based on diffusion potential prosody as claimed in claim 1, characterized in that: The step of performing hidden encoding processing on the synthesized text to obtain a hidden representation of the text includes: Dividing the text to be synthesized into multiple groups of words; Converting the to-be-synthesized text into a phoneme sequence and a word segmentation sequence according to the word segmentation; Concatenate the phoneme sequence and the word segmentation sequence to obtain a joint sequence corresponding to the word segmentation; The joint sequences are encoded one by one to obtain a text hidden representation.

3. The method for speech synthesis based on diffusion potential prosody as claimed in claim 1, characterized in that: The step of extracting the Mel-spectrogram of the initial audio comprises: Extracting speech features from the initial audio; Calculating the Mel-spectrogram coefficients of the speech features one by one to obtain the Mel-spectrogram coefficients corresponding to the speech features; The Mel-spectrogram coefficients are stacked in a matrix to obtain a Mel-spectrogram.

4. The method for speech synthesis based on diffusion potential prosody as claimed in claim 1, characterized in that: The constructing the analyzed phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation includes: Performing phoneme-level expansion on the text hidden representation and the speaker hidden representation respectively to obtain a text phoneme level corresponding to the text hidden representation and a speaker phoneme level corresponding to the speaker hidden representation; Aggregating the text phoneme level and the speaker phoneme level into a composite input variable; Performing phoneme duration analysis on the comprehensive input variable to obtain the analyzed phoneme duration corresponding to the reference prosody vector.

5. The method for speech synthesis based on diffusion potential prosody as claimed in claim 1, characterized in that: The step of constructing an analysis prosody vector by using the text hidden representation, the speaker hidden representation and the analysis phoneme duration includes: Generate an initial prosody vector based on the text hidden representation, the speaker hidden representation and the analyzed phoneme duration combination; The initial prosody vector is vector quantized to obtain an analysis prosody vector.

6. The method for speech synthesis based on diffusion potential prosody as claimed in claim 5, characterized in that: The step of performing vector quantization on the initial rhythm vector to obtain an analysis rhythm vector includes: Obtain a codebook containing multiple groups of elements; Performing Euclidean distance calculation on each of the initial prosody vectors and each element in the codebook to obtain calculation results of each of the initial prosody vectors and each element in the codebook; Selecting a codebook element with the smallest Euclidean distance according to the calculation result, and aggregating the codebook elements to generate a codebook element set; An analysis prosody vector is generated according to the set of codebook elements.

7. The method for speech synthesis based on diffusion potential prosody as claimed in claim 1, characterized in that: The step of generating a potential prosody vector according to the reference prosody vector and the analyzed prosody vector comprises: Normalizing the reference prosody vector and the analysis prosody vector respectively to obtain a processing result corresponding to the reference prosody vector and a processing result corresponding to the analysis prosody vector; Calculating an error between a processing result corresponding to the reference prosody vector and a processing result corresponding to the analysis prosody vector to obtain an error result; The error result is combined with the analyzed prosody vector to generate a latent prosody vector.

8. A speech synthesis device based on diffusion potential prosody, characterized in that: include: An encoding processing module is used to perform hidden encoding processing on the synthesized text to obtain a hidden representation of the text; An extraction module, used for extracting a Mel-spectrogram of the initial audio; A vector generation module, used for generating a reference prosody vector corresponding to each frequency band of the initial audio according to the Mel-spectrogram and a preset real phoneme duration; A construction module, configured to construct the analyzed phoneme duration corresponding to the reference prosody vector according to the text hidden representation and the speaker hidden representation; A vector construction module, used for constructing an analysis prosody vector using the text hidden representation, the speaker hidden representation and the analysis phoneme duration; A generating module, configured to generate a potential prosody vector according to the reference prosody vector and the analyzed prosody vector; The speech synthesis module is used to perform speech synthesis using the text hidden representation, the speaker hidden representation and the potential prosody vector to obtain synthesized speech.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the speech synthesis method based on diffuse latent prosody are implemented as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the speech synthesis method based on diffuse latent prosody are implemented as claimed in any one of claims 1 to 7.

Citation Information

Cited By

  • Discrete audio feature generation method and device and audio data word segmentation device training method and device

    CN120748434A