Speech synthesis method and device, electronic equipment and storage medium

By embedding processing and acoustic feature conversion of text data, combined with spectrum conversion and nesting merging technology, the problem of speech quality degradation in speech synthesis is solved, and higher quality speech synthesis is achieved.

CN120356455APending Publication Date: 2025-07-22PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510631827.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing speech synthesis method, the speech feature generation process is independent of the text parsing process, resulting in the inability to perceive the context information of the input text, resulting in a decline in speech quality.

Method used

By embedding the target text data, the text embedding vector sequence is generated, and the target speech synthesis model is used for acoustic feature conversion and spectrum conversion. Combined with the acoustic feature alternating update and spectrum nesting merging technology, the Mel spectrum sequence is generated, and finally speech synthesis is performed.

Benefits of technology

Improve the coherence and nature of speech synthesis and improve the quality of speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356455A_ABST
    Figure CN120356455A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech synthesis method and device, electronic equipment and a storage medium, belongs to the technical field of speech synthesis, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: performing embedding processing on target text data to obtain a starting text embedding vector and a subsequent text embedding vector; based on a preset target speech synthesis model, performing acoustic feature conversion on the initial text embedding vector to obtain an initial acoustic feature vector; on the basis of the initial acoustic feature vector and the subsequent text embedding vector, feature alternate updating is carried out, and a subsequent acoustic feature vector is obtained; performing spectrum conversion on the initial acoustic feature vector to obtain an initial Mel spectrum; based on the target speech synthesis model and the initial Mel spectrum, performing spectrum alternating conversion on the subsequent acoustic feature vector to obtain a subsequent Mel spectrum; and performing speech synthesis on the initial Mel spectrum and the subsequent Mel spectrum based on the target speech synthesis model. According to the embodiment of the invention, the speech quality of speech synthesis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and is applicable to the fields of fintech and healthcare, and particularly relates to a speech synthesis method and apparatus, an electronic device, and a storage medium. Background Art

[0002] Speech synthesis is used to convert text information into speech information. For example, in the field of fintech, a financial intelligent customer service can perform speech synthesis on a reply text to complete the reply to a customer's question, which can help the customer quickly understand financial-related information. For example, in the field of healthcare, a doctor can use a medical speech assistant to perform speech synthesis on a patient's examination report, so that the patient's examination report is output through speech, thereby helping the doctor quickly understand the patient's information.

[0003] Currently, the method of speech synthesis is usually to perform text parsing on the input text, then generate speech features, and finally perform speech synthesis based on the speech features. However, in this speech synthesis method, both the speech feature generation process and the text parsing process are independent, resulting in the inability to perceive the context information of the input text during the speech synthesis process, causing the speech quality of the speech synthesis to decline. Therefore, how to improve the speech quality of speech synthesis has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a speech synthesis method and apparatus, an electronic device, and a storage medium, aiming to improve the speech quality of speech synthesis.

[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes a speech synthesis method, and the method includes:

[0006] Obtain target text data;

[0007] Perform embedding processing on the target text data to obtain a text embedding vector sequence, where the text embedding vector sequence includes a starting text embedding vector and subsequent text embedding vectors;

[0008] Based on a preset target speech synthesis model, perform acoustic feature conversion on the starting text embedding vector to obtain a starting acoustic feature vector;

[0009] Based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vectors, perform alternating update of acoustic features to obtain subsequent acoustic feature vectors;

[0010] Based on the target speech synthesis model, perform spectral conversion on the starting acoustic feature vector to obtain a starting Mel spectrogram;

[0011] Based on the target voice synthesis model and the initial Mel spectrogram, perform spectral alternating conversion on the subsequent acoustic feature vectors to obtain subsequent Mel spectrograms;

[0012] Perform spectral nested merging on the initial Mel spectrogram and the subsequent Mel spectrograms to obtain a Mel spectrogram sequence;

[0013] Based on the target voice synthesis model, perform voice synthesis on the Mel spectrogram sequence.

[0014] In some embodiments, the method further includes training the target voice synthesis model, including:

[0015] Obtain voice sample data and text sample data;

[0016] Extract features from the voice sample data to obtain voice feature vectors;

[0017] Perform random masking on the text sample data to obtain masked regions and unmasked regions;

[0018] Perform embedding processing on the masked regions to obtain masked phrase embedding vectors;

[0019] Perform embedding processing on the unmasked regions to obtain masked text embedding vectors;

[0020] Based on the masked text embedding vectors, the masked phrase embedding vectors, and the voice feature vectors, perform model training on a preset initial voice synthesis model to obtain the target voice synthesis model.

[0021] In some embodiments, the performing model training on a preset initial voice synthesis model based on the masked text embedding vectors, the masked phrase embedding vectors, and the voice feature vectors to obtain the target voice synthesis model includes:

[0022] Based on the initial voice synthesis model, perform vector conversion on the masked text embedding vectors to obtain predicted voice feature vectors;

[0023] Based on the initial voice synthesis model, perform masked phrase prediction on the masked text embedding vectors to obtain predicted phrase embedding vectors;

[0024] Based on the initial voice synthesis model and the masked text embedding vectors, perform sequence prediction to obtain predicted text embedding vectors;

[0025] Based on the initial voice synthesis model, perform loss calculation on the voice feature vectors, the masked phrase embedding vectors, the masked text embedding vectors, the predicted voice feature vectors, the predicted phrase embedding vectors, and the predicted text embedding vectors to obtain total model loss data;

[0026] Based on the total model loss data, adjust the parameters of the initial speech synthesis model to obtain the target speech synthesis model.

[0027] In some embodiments, calculating losses for the speech feature vector, the masked phrase embedding vector, the masked text embedding vector, the predicted speech feature vector, the predicted phrase embedding vector, and the predicted text embedding vector based on the initial speech synthesis model to obtain total model loss data, including:

[0028] Calculating a speech feature loss data based on the initial speech synthesis model for the predicted speech feature vector and the speech feature vector;

[0029] Calculating a masked phrase loss data based on the initial speech synthesis model for the predicted phrase embedding vector and the masked phrase embedding vector;

[0030] Calculating a text prediction loss data based on the initial speech synthesis model for the predicted text embedding vector and the masked text embedding vector;

[0031] Merging the speech feature loss data, the masked phrase loss data, and the text prediction loss data to obtain the total model loss data.

[0032] In some embodiments, performing acoustic feature alternating updates based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vector to obtain a subsequent acoustic feature vector, including:

[0033] Concatenating the starting acoustic feature vector and the subsequent text embedding vector based on the target speech synthesis model to obtain a concatenated feature vector;

[0034] Performing an acoustic feature transformation on the concatenated feature vector based on the target speech synthesis model to obtain the subsequent acoustic feature vector.

[0035] In some embodiments, performing an embedding process on the target text data to obtain a sequence of text embedding vectors, including:

[0036] Performing a word segmentation process on the target text data to obtain segmented text blocks;

[0037] Performing a vector mapping on the segmented text blocks to obtain the sequence of text embedding vectors.

[0038] In some embodiments, performing speech synthesis on the Mel spectrogram sequence based on the target speech synthesis model, including:

[0039] Based on the target speech synthesis model, perform waveform conversion on the Mel spectrum sequence to obtain a speech time-domain waveform;

[0040] Based on the target speech synthesis model, perform signal conversion on the speech time-domain waveform to obtain a speech signal;

[0041] Based on the target speech synthesis model, perform audio synthesis on the speech signal to obtain a target synthesized speech.

[0042] To achieve the above object, a second aspect of the embodiments of the present application proposes a speech synthesis device, the device includes:

[0043] A data acquisition module, configured to acquire target text data;

[0044] A text embedding module, configured to perform embedding processing on the target text data to obtain a text embedding vector sequence, where the text embedding vector sequence includes a starting text embedding vector and subsequent text embedding vectors;

[0045] A feature conversion module, configured to perform acoustic feature conversion on the starting text embedding vector based on a preset target speech synthesis model to obtain a starting acoustic feature vector;

[0046] An interleaved encoding module, configured to perform alternating update of acoustic features based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vectors to obtain subsequent acoustic feature vectors;

[0047] A spectrum conversion module, configured to perform spectrum conversion on the starting acoustic feature vector based on the target speech synthesis model to obtain a starting Mel spectrum;

[0048] A spectrum alternating conversion module, configured to perform spectrum alternating conversion on the subsequent acoustic feature vectors based on the target speech synthesis model and the starting Mel spectrum to obtain subsequent Mel spectra;

[0049] A spectrum nested merging module, configured to perform spectrum nested merging on the starting Mel spectrum and the subsequent Mel spectra to obtain a Mel spectrum sequence;

[0050] A speech synthesis module, configured to perform speech synthesis on the Mel spectrum sequence based on the target speech synthesis model.

[0051] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0052] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described in the first aspect above.

[0053] The voice synthesis method, device, electronic device and storage medium provided by the present application perform embedding processing on the obtained target text data to generate a text embedding vector sequence including a starting text embedding vector and subsequent text embedding vectors, which can provide structured vector input. Further, the starting text embedding vector is subjected to acoustic feature conversion by using a target voice synthesis model to generate a starting acoustic feature vector, and through an acoustic feature alternating update mechanism, combined with subsequent text embedding vectors, subsequent acoustic feature vectors are obtained, improving the dynamics and coherence in the voice synthesis process. Secondly, based on the target voice synthesis model, the starting acoustic feature vector is subjected to spectrum conversion to obtain a starting Mel spectrum, and then combined with the starting Mel spectrum, the subsequent acoustic feature vectors are subjected to spectrum alternating conversion to obtain subsequent Mel spectra, and through a spectrum nested merging technique, a complete Mel spectrum sequence is generated, ensuring the coherence and naturalness of the synthesized voice. Finally, based on the target voice synthesis model, voice synthesis is performed on the Mel spectrum sequence, ensuring the coherence and naturalness of the synthesized voice, thereby improving the quality of the synthesized voice. Description of the Drawings

[0054] Figure 1 is a flowchart of the voice synthesis method provided by the embodiments of the present application;

[0055] Figure 2 is Figure 1 a flowchart of step S102 in

[0056] Figure 3 is a flowchart of the voice synthesis method provided by another embodiment of the present application;

[0057] Figure 4 is Figure 3 a flowchart of step S306 in

[0058] Figure 5 is Figure 4 a flowchart of step S404 in

[0059] Figure 6 is Figure 1 a flowchart of step S104 in

[0060] Figure 7 is Figure 6 a flowchart of step S108 in

[0061] Figure 8It is a schematic structural diagram of the speech synthesis device provided by an embodiment of the present application;

[0062] Figure 9 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] It should be noted that although functional module division is performed in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different sequence in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0066] First, several nouns involved in the present application are analyzed:

[0067] Speech synthesis system: A speech synthesis system is a technical system that converts text information into speech information. It realizes this process through the following steps: First, preprocess the input text, including word segmentation, grammar analysis, etc., and then convert the processed text into a text embedding vector; then use the speech synthesis model to convert the text embedding vector into an acoustic feature vector, such as Mel spectrogram; finally, convert the acoustic feature vector into an audible speech signal through a vocoder. This system is widely used in fields such as voice assistants, navigation software, and audiobooks, and can generate natural and fluent speech output in real time, improving the user's interaction experience.

[0068] Bag-of-words model: The bag-of-words model is a simple text representation method. It regards text as a set of words, ignores the order and grammatical structure of words, and only focuses on the frequency of word occurrence. By counting the occurrence times of each word in the text, the text can be converted into a vector form, which is suitable for natural language processing tasks such as text classification and clustering.

[0069] Word Embedding Model: A word embedding model is a technique that maps words in text to a low-dimensional vector space. It learns the vector representations of each word through training, and these vectors can capture the semantic information and context relationships of words. Word embedding models are widely used in natural language processing tasks such as text classification, sentiment analysis, machine translation, and speech recognition.

[0070] Speech synthesis is used to convert text information into speech information. For example, in the fintech field, financial intelligent customer service can perform speech synthesis on the reply text to complete the response to customers' questions, which can help customers quickly understand financial-related information. For example, in the medical field, doctors can use medical speech assistants to perform speech synthesis on patients' examination reports, so that the patients' examination reports are output as speech, thereby helping doctors quickly understand patients' information.

[0071] Currently, the method of speech synthesis usually parses the input text, then generates speech features, and finally performs speech synthesis based on the speech features. However, in this speech synthesis method, both the speech feature generation process and the text parsing process are independent, resulting in the inability to perceive the context information of the input text during the speech synthesis process, causing the speech quality of the speech synthesis to decline. Therefore, how to improve the speech quality of speech synthesis has become a technical problem to be solved urgently.

[0072] Based on this, the embodiments of the present application provide a speech synthesis method, apparatus, electronic device, and storage medium, aiming to improve the speech quality of speech synthesis.

[0073] The speech synthesis method, apparatus, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the speech synthesis method in the embodiments of the present application is described.

[0074] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application system.

[0075] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] The speech synthesis method provided by the embodiments of this application relates to the field of speech synthesis technology and is applicable to the fields of fintech and healthcare. The speech synthesis method provided by the embodiments of this application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the speech synthesis method, etc., but is not limited to the above forms.

[0077] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0078] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained by means of a pop-up window or jumping to a confirmation page, etc. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0079] Figure 1 is an optional flowchart of the speech synthesis method provided by the embodiments of this application. This method can be used in a speech synthesis system. Figure 1The method in [it] may include but is not limited to steps S101 to S108.

[0080] Step S101, obtain target text data;

[0081] Step S102, perform embedding processing on the target text data to obtain a text embedding vector sequence, where the text embedding vector sequence includes a starting text embedding vector and subsequent text embedding vectors;

[0082] Step S103, based on a preset target speech synthesis model, perform acoustic feature conversion on the starting text embedding vector to obtain a starting acoustic feature vector;

[0083] Step S104, based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vectors, perform alternating update of acoustic features to obtain subsequent acoustic feature vectors;

[0084] Step S105, based on the target speech synthesis model, perform spectral conversion on the starting acoustic feature vector to obtain a starting Mel spectrogram;

[0085] Step S106, based on the target speech synthesis model and the starting Mel spectrogram, perform alternating spectral conversion on the subsequent acoustic feature vectors to obtain subsequent Mel spectrograms;

[0086] Step S107, perform spectral nested merging on the starting Mel spectrogram and the subsequent Mel spectrograms to obtain a Mel spectrogram sequence;

[0087] Step S108, based on the target speech synthesis model, perform speech synthesis on the Mel spectrogram sequence.

[0088] Steps S101 to S108 shown in the embodiments of the present application, by performing embedding processing on the obtained target text data to generate a text embedding vector sequence including a starting text embedding vector and subsequent text embedding vectors, can provide structured vector input. Further, using the target speech synthesis model to perform acoustic feature conversion on the starting text embedding vector to generate a starting acoustic feature vector, and through the acoustic feature alternating update mechanism, combined with the subsequent text embedding vectors, to obtain subsequent acoustic feature vectors, which improves the dynamics and coherence in the speech synthesis process. Secondly, based on the target speech synthesis model, perform spectral conversion on the starting acoustic feature vector to obtain a starting Mel spectrogram, and then combined with the starting Mel spectrogram, perform alternating spectral conversion on the subsequent acoustic feature vectors to obtain subsequent Mel spectrograms, and through the spectral nested merging technology, generate a complete Mel spectrogram sequence, ensuring the coherence and naturalness of the synthesized speech. Finally, based on the target speech synthesis model, perform speech synthesis on the Mel spectrogram sequence, ensuring the coherence and naturalness of the synthesized speech, thereby improving the quality of the synthesized speech.

[0089] In step S101 of some embodiments, the target text data refers to the text content that needs to be synthesized into speech. For example, in the stock market analysis scenario, the target text data can be a text of a stock market analysis report. In the medical examination scenario, the target text data can be a text of a medical examination report.

[0090] In the embodiments of the present application, the target text data is the text data input into the speech synthesis system. The speech synthesis system can obtain the target text data by reading the text content that needs to be synthesized into speech input by the user. For example, in the stock market analysis scenario, the speech synthesis system can obtain the text of the stock market analysis report that needs to be synthesized into speech by reading the analysis report content input by the stock market analyst. In the medical examination scenario, the speech synthesis system can obtain the text of the medical examination report that needs to be synthesized into speech by reading the examination report content input by the doctor.

[0091] In step S102 of some embodiments, the text embedding vector sequence is a vector representation form of the target text data. The starting text embedding vector is the first vector in the text embedding vector sequence, representing the beginning part of the target text data. For example, in the stock market analysis scenario, when the text embedding vector sequence is "Today, stock market, overall, shows, an, upward, trend, Shanghai Composite Index, increase, is, 1.2%, Shenzhen Component Index, increase, is, 2.3%", the starting text embedding vector can be "Today". In the medical examination scenario, when the text embedding vector sequence is "Patient, Li Mou, 45 years old, male, electrocardiogram, examination, shows, heart rate, normal, no, abnormal, waveform", the starting text embedding vector can be "Patient". The subsequent text embedding vectors refer to the other vectors in the text embedding vector sequence except the starting text embedding vector. For example, in the stock market analysis scenario, when the text embedding vector sequence is "Today, stock market, overall, shows, an, upward, trend, Shanghai Composite Index, increase, is, 1.2%, Shenzhen Component Index, increase, is, 2.3%", the subsequent text embedding vectors are "stock market, overall, shows, an, upward, trend, Shanghai Composite Index, increase, is, 1.2%, Shenzhen Component Index, increase, is, 2.3%". In the medical examination scenario, when the subsequent text embedding vectors are "Patient, Li Mou, 45 years old, male, electrocardiogram, examination, shows, heart rate, normal, no, abnormal, waveform", the subsequent text embedding vectors are "Li Mou, 45 years old, male, electrocardiogram, examination, shows, heart rate, normal, no, abnormal, waveform".

[0092] In the embodiments of the present application, the target text data can be segmented to obtain multiple phrases, and then all the phrases are mapped to the vector space according to the arrangement order of each phrase in the target text data, and the text embedding vector sequence can be obtained.

[0093] For details, please refer toFigure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S202:

[0094] Step S201, perform word segmentation on the target text data to obtain segmented text blocks;

[0095] Step S202, perform vector mapping on the segmented text blocks to obtain a sequence of text embedding vectors.

[0096] In step S201 of some embodiments, the segmented text blocks refer to the words or phrases in the target text data. For example, in the stock market analysis scenario, when the target text data is "The overall stock market showed an upward trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%", the segmented text blocks can be "today, stock market, overall, showed, upward, trend, Shanghai Composite Index, rise, by, 1.2%, Shenzhen Component Index, rise, by, 2.3%". In the medical examination scenario, when the target text data is "Patient Li, a 45-year-old male, the electrocardiogram examination shows normal heart rate and no abnormal waveforms", the segmented text blocks can be "patient, Li, 45 years old, male, electrocardiogram, examination, shows, heart rate, normal, no, abnormal, waveforms".

[0097] In step S202 of some embodiments, each segmented text block can be mapped to a pre-constructed vector space using a traditional machine learning model or a neural network model to obtain a sequence of text embedding vectors. Among them, the traditional machine learning model can be a bag-of-words model, etc., and the neural network model can be a word embedding model, a bidirectional encoder representation model, etc.

[0098] In steps S201 and S202 illustrated in the embodiments of the present application, by performing word segmentation on the target text data, the target text data is decomposed into smaller semantic units, that is, segmented text blocks, which improves the fineness of text processing and lays a foundation for vector mapping. Further, by performing vector mapping on the segmented text blocks to generate a sequence of text embedding vectors, the target text data can be effectively represented and calculated in the vector space, enhancing the processability of the target text data in the model and improving the efficiency of text analysis.

[0099] In step S103 of some embodiments, the target speech synthesis model refers to a deep learning model used to convert text embedding vectors into speech signals. The starting acoustic feature vector is an acoustic feature vector generated by the target speech synthesis model according to the starting text embedding vector.

[0100] In the implementation of the present application, by inputting the above-mentioned starting text embedding vector into the target speech synthesis model, the target speech synthesis model can convert the starting text embedding vector into a starting acoustic feature vector according to the conversion method of text vector to acoustic feature learned during training.

[0101] It should be noted that to ensure the accuracy of the acoustic feature conversion of the target speech synthesis model for the starting text embedding vector, the target speech synthesis model must be a trained speech synthesis model.

[0102] Specifically, please refer to Figure 3 , in some embodiments, the training method of the target speech synthesis model may include, but is not limited to, steps S301 to S306:

[0103] Step S301, obtaining speech sample data and text sample data;

[0104] Step S302, extracting features from the speech sample data to obtain speech feature vectors;

[0105] Step S303, randomly masking the text sample data to obtain a masked area and an unmasked area;

[0106] Step S304, performing embedding processing on the masked area to obtain masked phrase embedding vectors;

[0107] Step S305, performing embedding processing on the unmasked area to obtain masked text embedding vectors;

[0108] Step S306, based on the masked text embedding vectors, masked phrase embedding vectors, and speech feature vectors, training a preset initial speech synthesis model to obtain the target speech synthesis model.

[0109] In step S301 of some embodiments, the text sample data refers to the text instances used to train the speech synthesis model. The speech sample data refers to the speech instances corresponding to the text sample data.

[0110] In the embodiments of the present application, the text sample data and its corresponding speech sample data can be obtained by web crawling, that is, the speech sample data, or can be obtained by manual writing and reading to obtain the text sample data and its corresponding speech sample data.

[0111] In step S302 of some embodiments, the speech feature vector refers to the vector representation of the speech features, where the semantic features are various parameters characterizing the properties of the speech signal, such as acoustic features, spectral features, speech quality features, etc.

[0112] In the embodiments of the present application, by extracting the acoustic features, spectral features, speech quality features, etc. in the speech sample data, the speech features of the speech sample data can be obtained, and then by performing vector mapping on the extracted speech features, speech feature vectors can be obtained.

[0113] In steps S303 to S305 of some embodiments, the masked region refers to the masked text part in the text sample data, and the non-masked region refers to the unmasked text part in the text sample data. For example, in the stock market analysis scenario, when the text sample data is "The overall stock market showed an upward trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%", and the randomly masked phrase is "upward", the masked region is "upward", and the non-masked region is "The overall stock market showed an XX trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%", where XX represents the masked region "upward". In the medical examination scenario, when the text sample data is "Patient Li, a 45-year-old male, the electrocardiogram examination shows that the heart rate is normal and there are no abnormal waveforms", and the randomly masked phrase is "normal", the masked region is "normal", and the non-masked region is "Patient Li, a 45-year-old male, the electrocardiogram examination shows that the heart rate is XX and there are no abnormal waveforms", where XX represents the masked region "normal". The masked phrase embedding vector refers to the vector representation form of the text in the masked region. The masked text embedding vector refers to the vector representation form of the text in the non-masked region.

[0114] In the embodiments of the present application, by performing word segmentation on the text sample data, multiple words or phrases can be obtained. Then, by randomly masking the multiple words or phrases, random masking of the text sample data can be achieved, thereby obtaining the masked region and the non-masked region. Further, models such as the bag-of-words model, the word embedding model, or the bidirectional encoder representation model can be used to perform vector mapping on the text content in the masked region to obtain the masked phrase embedding vector. Similarly, models such as the bag-of-words model, the word embedding model, or the bidirectional encoder representation model can be used to perform vector mapping on the text content in the non-masked region to obtain the masked text embedding vector.

[0115] In step S306 of some embodiments, the initial speech synthesis model refers to a speech synthesis model that has not been trained with specific data. It should be noted that the speech synthesis accuracy of the initial speech synthesis model is relatively low.

[0116] After obtaining the masked text embedding vector, the masked phrase embedding vector, and the speech feature vector, the masked text embedding vector, the masked phrase embedding vector, and the speech feature vector are input into a pre-constructed initial speech synthesis model. The initial speech synthesis model can perform forward propagation and backward propagation on the masked text embedding vector, the masked phrase embedding vector, and the speech feature vector through the method of supervised learning, and a target speech synthesis model can be obtained.

[0117] Specifically, please refer to Figure 4 , in some embodiments, step S306 may include but is not limited to steps S401 to S405:

[0118] Step S401: Based on the initial speech synthesis model, perform vector conversion on the masked text embedding vector to obtain a predicted speech feature vector;

[0119] Step S402: Based on the initial speech synthesis model, perform masked phrase prediction on the masked text embedding vector to obtain a predicted phrase embedding vector;

[0120] Step S403: Based on the initial speech synthesis model and the masked text embedding vector, perform sequence prediction to obtain a predicted text embedding vector;

[0121] Step S404: Based on the initial speech synthesis model, calculate the loss for the speech feature vector, masked phrase embedding vector, masked text embedding vector, predicted speech feature vector, predicted phrase embedding vector, and predicted text embedding vector to obtain the total model loss data;

[0122] Step S405: Based on the total model loss data, adjust the parameters of the initial speech synthesis model to obtain the target speech synthesis model.

[0123] In steps S401 to S403 of some embodiments, the predicted speech feature vector refers to the form of the speech feature vector of the masked text embedding vector calculated by the initial speech synthesis model based on its initial parameters and initial functions. The predicted phrase embedding vector refers to the missing phrase vector of the masked text embedding vector calculated by the initial speech synthesis model based on the initial parameters and initial functions. The predicted text embedding vector refers to the semantic information of the masked text embedding vector calculated by the initial speech synthesis model based on the initial parameters and initial functions.

[0124] It should be noted that the initial parameters can be the initial phrase prediction parameters for the missing part in the predicted text vector and the initial semantic information understanding parameters for understanding the semantic information in the text vector. The initial functions can be the initial text-to-speech conversion functions that convert the text vector into a speech feature vector.

[0125] In the implementation of this application, the initial speech synthesis model can use the initial conversion function to map the masked text embedding vector to the speech feature space to obtain the predicted speech feature vector. Similarly, the initial speech synthesis model can perform semantic understanding on the masked text embedding vector according to the initial semantic information understanding parameters to obtain the predicted text embedding vector. Further, the initial speech synthesis model predicts the missing phrase vector in the masked text embedding vector according to the semantic information of the masked text embedding vector it understands and the initial phrase prediction parameters, so as to obtain the predicted phrase embedding vector.

[0126] In step S404 of some embodiments, the total model loss data refers to the total error of the initial speech synthesis model during training.

[0127] In the embodiments of the present application, the predicted speech feature vector, the predicted phrase embedding vector, and the predicted text embedding vector can be compared one by one with the input speech feature vector, the masked phrase embedding vector, and the masked text embedding vector to obtain the model training error for each dimension. Then, by combining the model training errors for each dimension, the total model loss data can be obtained.

[0128] Specifically, please refer to Figure 5 , in some embodiments, step S404 may include but is not limited to steps S501 to S504:

[0129] Step S501, based on the initial speech synthesis model, calculate the loss between the predicted speech feature vector and the speech feature vector to obtain the speech feature loss data;

[0130] Step S502, based on the initial speech synthesis model, calculate the loss between the predicted phrase embedding vector and the masked phrase embedding vector to obtain the masked phrase loss data;

[0131] Step S503, based on the initial speech synthesis model, calculate the loss between the predicted text embedding vector and the masked text embedding vector to obtain the text prediction loss data;

[0132] Step S504, merge the speech feature loss data, the masked phrase loss data, and the text prediction loss data to obtain the total model loss data.

[0133] In steps S501 to S503 of some embodiments, the speech feature loss data represents the vector gap between the predicted speech feature vector and the speech feature vector. The masked phrase loss data represents the vector gap between the predicted phrase embedding vector and the masked phrase embedding vector. The text prediction loss data represents the vector gap between the predicted text embedding vector and the masked text embedding vector.

[0134] In the embodiments of the present application, the initial speech synthesis model also includes a loss function for calculating the gap between the model prediction data and the actual input data. Among them, the model prediction data can be the above-mentioned predicted speech feature vector, predicted phrase embedding vector, and predicted text embedding vector, and the actual input data can be the above-mentioned speech feature vector, masked phrase embedding vector, and masked text embedding vector. Therefore, the initial speech synthesis model can use this loss function to calculate the vector gap between the predicted speech feature vector and the speech feature vector to obtain the speech feature loss data, use this loss function to calculate the masked phrase loss data representing the vector gap between the predicted phrase embedding vector and the masked phrase embedding vector to obtain the masked phrase loss data, and use this loss function to calculate the text prediction loss data representing the vector gap between the predicted text embedding vector and the masked text embedding vector to obtain the text prediction loss data.

[0135] In step S504 of some embodiments, by performing a summation operation on the above-mentioned speech feature loss data, masked phrase loss data, and text prediction loss data, the total loss of the initial speech synthesis model, that is, the model total loss data, can be obtained.

[0136] In steps S501 to S504 illustrated in the embodiments of the present application, by comparing the predicted speech feature vector and the speech feature vector, the speech feature loss data is calculated, which can accurately quantify the error of the initial speech synthesis model in restoring speech features, thereby helping the initial speech synthesis model to optimize the naturalness and sound quality of speech synthesis in a targeted manner. By comparing the predicted phrase embedding vector and the masked phrase embedding vector, the masked phrase loss data is obtained, which can improve the prediction ability of the initial speech synthesis model for text missing areas. Secondly, by comparing the predicted text embedding vector and the masked text embedding vector, the text prediction loss data is obtained, which can enhance the initial speech synthesis model's understanding of the overall semantics of the text and the ability to generate coherence. Finally, by merging the speech feature loss data, masked phrase loss data, and text prediction loss data to form the model total loss data, a scientific basis can be provided for adjusting the model parameters of the initial speech synthesis model, thereby achieving an overall improvement in the speech quality, prediction of text missing areas, and semantic coherence of the initial speech synthesis model.

[0137] In step S405 of some embodiments, according to the model total loss data, the predicted speech feature vector, predicted phrase embedding vector, predicted text embedding vector, and model total loss data can be backpropagated within the initial speech synthesis model, and the above-mentioned initial phrase prediction parameters, initial semantic information understanding parameters, initial text-to-speech conversion function, etc. can be modified, so as to realize the parameter adjustment of the initial speech synthesis model, and thus obtain the trained target speech synthesis model.

[0138] In steps S401 to S405 illustrated in the embodiments of the present application, the initial speech synthesis model generates a predicted speech feature vector by performing vector conversion on the masked text embedding vector, obtains a predicted phrase embedding vector by performing masked phrase prediction on the masked text embedding vector, and obtains a predicted text embedding vector by performing sequence prediction on the masked text embedding vector, which can reflect the speech synthesis ability of the initial speech synthesis model before training. Further, according to the initial speech synthesis model, loss calculations are performed on the speech feature vector, masked phrase embedding vector, masked text embedding vector, predicted speech feature vector, predicted phrase embedding vector, and predicted text embedding vector to obtain the model total loss data, and then according to the model total loss data, the parameters of the initial speech synthesis model are adjusted to obtain the target speech synthesis model, which can improve the naturalness, accuracy, and coherence of the target speech synthesis model in speech synthesis, thereby ensuring that the synthesized speech not only has excellent sound quality but also can accurately and smoothly convey the information in the text.

[0139] In steps S301 to S306 illustrated in the embodiments of the present application, by obtaining voice sample data and text sample data, a data basis is provided for training the initial voice synthesis model. Further, feature extraction is performed on the voice sample data to obtain voice feature vectors, which can accurately capture the key characteristics of the voice sample data. Secondly, the text sample data is randomly masked and embedded respectively to obtain masked phrase embedding vectors and masked text embedding vectors, which can enhance the initial voice synthesis model's grasp of text details and improve the initial voice synthesis model's understanding ability of the overall semantics of the text. Finally, according to the masked text embedding vectors, masked phrase embedding vectors, and voice feature vectors, the initial voice synthesis model is trained to obtain the target voice synthesis model, which can improve the naturalness, accuracy, and coherence of voice synthesis of the target voice synthesis model.

[0140] In step S104 of some embodiments, the subsequent acoustic feature vector refers to the acoustic feature vector generated by the target voice synthesis model according to the starting acoustic feature vector and the subsequent text embedding vector. For example, in the stock market analysis scenario, when the text sample data is "The overall stock market showed an upward trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%", the subsequent acoustic feature vector can be the acoustic feature vector generated according to the vectors corresponding to "today" and "stock market". In the medical examination scenario, when the text sample data is "Patient Li, a 45-year-old male, the electrocardiogram examination shows normal heart rate and no abnormal waveforms", the subsequent acoustic feature vector can be the acoustic feature vector generated according to the vectors corresponding to "patient" and "Li".

[0141] In the embodiments of the present application, by combining the starting acoustic feature vector and the subsequent text embedding vector, a concatenated feature vector can be obtained, and then using the target voice synthesis model, acoustic feature conversion is performed on the concatenated feature vector to obtain the subsequent acoustic feature vector.

[0142] Specifically, please refer to Figure 6 , in some embodiments, step S104 may include but is not limited to steps S601 to S602:

[0143] Step S601, based on the target voice synthesis model, vector concatenation is performed on the starting acoustic feature vector and the subsequent text embedding vector to obtain a concatenated feature vector;

[0144] Step S602, based on the target voice synthesis model, acoustic feature conversion is performed on the concatenated feature vector to obtain the subsequent acoustic feature vector.

[0145] In step S601 of some embodiments, the concatenated feature vector refers to the vector concatenation representation of the starting acoustic feature vector and the subsequent text embedding vector.

[0146] In the embodiments of the present application, after obtaining the starting acoustic feature vector, the target speech synthesis model will select the first vector according to the vector sorting in the subsequent text embedding vectors and splice it with the starting acoustic feature vector to obtain a spliced feature vector. For example, in the stock market analysis scenario, when the starting acoustic feature vector is "Today", and the subsequent text embedding vectors are "stock market, overall, shows, an, upward, trend, the Shanghai Composite Index, has, a, gain, of, 1.2%, the Shenzhen Component Index, has, a, gain, of, 2.3%", splicing "Today" with "stock market" can obtain a spliced feature vector. In the medical examination scenario, when the starting acoustic feature vector is "patient", and the subsequent text embedding vectors are "Li, 45 years old, male, electrocardiogram, examination, shows, heart rate, normal, no, abnormal, waveform", splicing "patient" with "Li" can obtain a spliced feature vector.

[0147] In step S602 of some embodiments, after obtaining the spliced feature vector, the target speech synthesis model can obtain the subsequent acoustic feature vector by converting the spliced feature vector into an acoustic feature.

[0148] It should be noted that after processing the first text embedding vector in the subsequent text embedding vectors, it is necessary to splice the acoustic feature vector converted from this text embedding vector with the second text embedding vector in the subsequent text embedding vectors, and then perform acoustic feature conversion on the spliced vector, and loop in turn until all the text embedding vectors in the subsequent text embedding vectors are converted into acoustic feature vectors. Nesting and combining the acoustic feature vectors converted from the subsequent text embedding vectors can obtain the subsequent acoustic feature vector.

[0149] In steps S601 to S602 illustrated in the embodiments of the present application, by splicing the starting acoustic feature vector with the subsequent text embedding vectors to form a spliced feature vector, it is possible to fuse the existing acoustic information with the new text content, providing a comprehensive feature basis for subsequent speech generation. Further, the target speech synthesis model performs acoustic feature conversion on the spliced feature vector to obtain the subsequent acoustic feature vector, so that the synthesized speech output by the target speech synthesis model can not only accurately reflect the text content, but also achieve smooth voice transition and natural voice expression, greatly improving the user experience.

[0150] In step S105 of some embodiments, the starting Mel spectrogram is a Mel spectrogram generated by the target speech synthesis model according to the starting acoustic feature vector, and the starting Mel spectrogram can represent the spectral information of the initial part of the speech signal.

[0151] In the embodiments of the present application, after receiving the starting acoustic feature vector, the target speech synthesis model can map the starting acoustic feature vector in the Mel spectrum dimension through an activation function or the like, so that the starting acoustic feature vector is converted into a starting Mel spectrum.

[0152] In step S106 of some embodiments, the subsequent Mel spectrum is the Mel spectrum generated by the target speech synthesis model according to the subsequent acoustic feature vectors, and the subsequent Mel spectrum can represent the spectral information of the subsequent part of the speech signal.

[0153] In the embodiments of the present application, the starting Mel spectrum can be concatenated with the first acoustic feature vector in the subsequent acoustic feature vectors to obtain a concatenated input vector, and then the target speech synthesis model is used to perform Mel spectrum conversion on the concatenated input vector to obtain the Mel spectrum representation of the first acoustic feature vector in the subsequent acoustic feature vectors. Further, the Mel spectrum representation is concatenated with the second acoustic feature vector in the subsequent acoustic feature vectors and Mel spectrum conversion is performed to obtain the Mel spectrum representation of the second acoustic feature vector in the subsequent acoustic feature vectors. This cycle continues until the last acoustic feature vector in the subsequent acoustic feature vectors is converted into a Mel spectrum, and the subsequent Mel spectrum can be obtained.

[0154] In step S107 of some embodiments, the Mel spectrum sequence is a sequence composed of the starting Mel spectrum and the subsequent Mel spectra arranged in chronological order, and the Mel spectrum sequence can represent the spectral information of the entire speech signal.

[0155] In the embodiments of the present application, by sorting and merging the Mel spectrograms in the above-mentioned starting Mel spectrogram and subsequent Mel spectrograms according to the generation time of the Mel spectrograms, a Mel spectrogram sequence can be obtained. For example, in the stock market analysis scenario, when the starting Mel spectrogram is the spectrogram representation of "today", and the subsequent text embedding vector is the spectrogram representation of "stock market, overall, showing, an, upward, trend, Shanghai Composite Index, increase, of, 1.2%, Shenzhen Component Index, increase, of, 2.3%", sorting and merging the spectrogram representation of "today" and the spectrogram representation of "stock market, overall, showing, an, upward, trend, Shanghai Composite Index, increase, of, 1.2%, Shenzhen Component Index, increase, of, 2.3%" according to the generation time of the Mel spectrograms can obtain the Mel spectrogram sequence "Today, the overall stock market shows an upward trend. The increase of the Shanghai Composite Index is 1.2%, and the increase of the Shenzhen Component Index is 2.3%". In the medical examination scenario, when the starting Mel spectrogram is the spectrogram representation of "patient", and the subsequent Mel spectrogram is the spectrogram representation of "Li Mou, 45 years old, male, electrocardiogram, examination, shows, normal, heart, rate, no, abnormal, waveform", sorting and merging the spectrogram representation of "patient" and the spectrogram representation of "Li Mou, 45 years old, male, electrocardiogram, examination, shows, normal, heart, rate, no, abnormal, waveform" according to the generation time of the Mel spectrograms can obtain the Mel spectrogram sequence "Patient Li Mou, 45-year-old male, electrocardiogram examination shows normal heart rate and no abnormal waveform".

[0156] In step S108 of some embodiments, after obtaining the Mel spectrogram sequence, the target voice synthesis model can convert the Mel spectrogram sequence into speech through a vocoder.

[0157] Specifically, please refer to Figure 7 , in some embodiments, step S108 may include but is not limited to steps S701 to S703:

[0158] Step S701, based on the target voice synthesis model, perform waveform conversion on the Mel spectrogram sequence to obtain a speech time-domain waveform;

[0159] Step S702, based on the target voice synthesis model, perform signal conversion on the speech time-domain waveform to obtain a speech signal;

[0160] Step S703, based on the target voice synthesis model, perform audio synthesis on the speech signal to obtain the target synthesized speech.

[0161] In step S701 of some embodiments, the speech time-domain waveform represents the waveform of the speech signal in the time domain.

[0162] In the embodiments of the present application, the vocoder built into the target voice synthesis model can be used to convert the Mel spectrogram into the voice time-domain waveform. It should be noted that the vocoder built into the target voice synthesis model can be HiFi-GAN (High-Fidelity Generative Adversarial Networks) or WaveGlow (WaveGlow: Probabilistic Wavelet Flow for Audio Synthesis).

[0163] In step S702 of some embodiments, the voice signal refers to the analog signal representation of the voice.

[0164] In the embodiments of the present application, the target voice synthesis model can convert the digital-form voice time-domain waveform into a continuous analog signal, that is, a voice signal, by performing digital-to-analog conversion on the voice time-domain waveform.

[0165] In step S703 of some embodiments, the target synthesized voice refers to the voice manifestation form of the target text data. For example, in the stock market analysis scenario, when the target text data is "The overall stock market showed an upward trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%", the target synthesized voice can be the voice form of "The overall stock market showed an upward trend today, with the Shanghai Composite Index rising by 1.2% and the Shenzhen Component Index rising by 2.3%". In the medical examination scenario, when the target text data is "Patient Li, a 45-year-old male, the electrocardiogram examination shows normal heart rate and no abnormal waveform", the target synthesized voice can be the voice form of "Patient Li, a 45-year-old male, the electrocardiogram examination shows normal heart rate and no abnormal waveform".

[0166] In the embodiments of the present application, the target synthesized voice with better voice quality can be generated by processing the voice signal such as filtering, duration adjustment, and volume normalization.

[0167] In steps S701 to S703 shown in the embodiments of the present application, the Mel spectrogram sequence is converted into a waveform using the target voice synthesis model to obtain the characteristics of the voice corresponding to the Mel spectrogram sequence in the time domain dimension. Further, the voice time-domain waveform is subjected to signal conversion to obtain a voice signal, ensuring the availability of the signal. Finally, through the audio synthesis step, the voice signal is further optimized and processed to generate the target synthesized voice, improving the naturalness and fluency of the voice synthesis and also enhancing the voice quality of the voice synthesis.

[0168] This application generates a text embedding vector sequence containing a starting text embedding vector and subsequent text embedding vectors through embedding processing on the obtained target text data, providing a structured vector input. Further, using the target speech synthesis model, the starting text embedding vector is acoustically feature-transformed to generate a starting acoustic feature vector, and through an acoustic feature alternating update mechanism, combined with the subsequent text embedding vectors, subsequent acoustic feature vectors are obtained, improving the dynamics and coherence in the speech synthesis process. Secondly, based on the target speech synthesis model, the starting acoustic feature vector is spectrogram-transformed to obtain a starting Mel spectrogram, and then combined with the starting Mel spectrogram, the subsequent acoustic feature vectors are spectrogram-alternatively transformed to obtain subsequent Mel spectrograms, and through a spectrogram nested merging technique, a complete Mel spectrogram sequence is generated, ensuring the coherence and naturalness of the synthesized speech. Finally, based on the target speech synthesis model, speech synthesis is performed on the Mel spectrogram sequence, ensuring the coherence and naturalness of the synthesized speech, thereby improving the quality of the synthesized speech.

[0169] Please refer to Figure 8 , this embodiment of the application also provides a speech synthesis device that can implement the above speech synthesis method. The device includes:

[0170] A data acquisition module 801 for acquiring target text data;

[0171] A text embedding module 802 for performing embedding processing on the target text data to obtain a text embedding vector sequence, where the text embedding vector sequence includes a starting text embedding vector and subsequent text embedding vectors;

[0172] A feature transformation module 803 for acoustically feature-transforming the starting text embedding vector based on a preset target speech synthesis model to obtain a starting acoustic feature vector;

[0173] An interleaved coding module 804 for alternately updating acoustic features based on the target speech synthesis model, the starting acoustic feature vector, and subsequent text embedding vectors to obtain subsequent acoustic feature vectors;

[0174] A spectrogram transformation module 805 for spectrogram-transforming the starting acoustic feature vector based on the target speech synthesis model to obtain a starting Mel spectrogram;

[0175] A spectrogram alternating transformation module 806 for spectrogram-alternatively transforming the subsequent acoustic feature vectors based on the target speech synthesis model and the starting Mel spectrogram to obtain subsequent Mel spectrograms;

[0176] A spectrogram nested merging module 807 for spectrogram-nested merging the starting Mel spectrogram and subsequent Mel spectrograms to obtain a Mel spectrogram sequence;

[0177] A voice synthesis module 808, configured to perform voice synthesis on a Mel spectrogram sequence based on a target voice synthesis model.

[0178] The specific implementation manner of this voice synthesis device is basically the same as the specific embodiments of the above voice synthesis method, and will not be elaborated here.

[0179] An embodiment of this application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above voice synthesis method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0180] Please refer to Figure 9 , Figure 9 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0181] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;

[0182] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and are called by the processor 901 to execute the voice synthesis method of the embodiments of this application;

[0183] An input / output interface 903, configured to implement information input and output;

[0184] A communication interface 904, configured to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0185] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0186] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0187] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above voice synthesis method is implemented.

[0188] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0189] The voice synthesis method, voice synthesis device, electronic device, and storage medium provided by the embodiment of the present application obtain target text data, perform embedding processing on the target text data to obtain a text embedding vector sequence, where the text embedding vector sequence includes a starting text embedding vector and subsequent text embedding vectors. Further, based on a preset target voice synthesis model, an acoustic feature conversion is performed on the starting text embedding vector to obtain a starting acoustic feature vector, and then based on the target voice synthesis model, the starting acoustic feature vector, and the subsequent text embedding vectors, an acoustic feature alternating update is performed to obtain subsequent acoustic feature vectors. Secondly, based on the target voice synthesis model, a spectrum conversion is performed on the starting acoustic feature vector to obtain a starting Mel spectrum, and then based on the target voice synthesis model and the starting Mel spectrum, a spectrum alternating conversion is performed on the subsequent acoustic feature vectors to obtain subsequent Mel spectra, and the starting Mel spectrum and the subsequent Mel spectra are spectrally nested and merged to obtain a Mel spectrum sequence. Finally, based on the target voice synthesis model, voice synthesis is performed on the Mel spectrum sequence.

[0190] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0191] Those skilled in the art can understand that the technical solutions shown in the figures do not limit the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0193] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0194] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0195] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0196] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0197] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0198] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0199] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0200] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A speech synthesis method, characterized in that, The method includes: Obtain target text data; Perform embedding processing on the target text data to obtain a sequence of text embedding vectors, where the sequence of text embedding vectors includes a starting text embedding vector and subsequent text embedding vectors; Based on a preset target speech synthesis model, perform acoustic feature conversion on the starting text embedding vector to obtain a starting acoustic feature vector; Based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vectors, perform alternating updates of acoustic features to obtain subsequent acoustic feature vectors; Based on the target speech synthesis model, perform spectral conversion on the starting acoustic feature vector to obtain a starting Mel spectrogram; Based on the target speech synthesis model and the starting Mel spectrogram, perform alternating spectral conversion on the subsequent acoustic feature vectors to obtain subsequent Mel spectrograms; Perform spectral nested merging on the starting Mel spectrogram and the subsequent Mel spectrograms to obtain a sequence of Mel spectrograms; Based on the target speech synthesis model, perform speech synthesis on the sequence of Mel spectrograms.

2. The method according to claim 1, characterized in that, The method further includes training the target speech synthesis model, including: Obtain speech sample data and text sample data; Extract features from the speech sample data to obtain speech feature vectors; Perform random masking on the text sample data to obtain a masked region and an unmasked region; Perform embedding processing on the masked region to obtain masked phrase embedding vectors; Perform embedding processing on the unmasked region to obtain masked text embedding vectors; Based on the masked text embedding vectors, the masked phrase embedding vectors, and the speech feature vectors, perform model training on a preset initial speech synthesis model to obtain the target speech synthesis model.

3. The method according to claim 2, wherein The performing model training on a preset initial speech synthesis model based on the masked text embedding vectors, the masked phrase embedding vectors, and the speech feature vectors to obtain the target speech synthesis model includes: Based on the initial speech synthesis model, perform vector conversion on the masked text embedding vectors to obtain predicted speech feature vectors; Based on the initial speech synthesis model, perform masked phrase prediction on the masked text embedding vectors to obtain predicted phrase embedding vectors; Based on the initial speech synthesis model and the masked text embedding vectors, perform sequence prediction to obtain predicted text embedding vectors; Based on the initial speech synthesis model, perform loss calculation on the speech feature vectors, the masked phrase embedding vectors, the masked text embedding vectors, the predicted speech feature vectors, the predicted phrase embedding vectors, and the predicted text embedding vectors to obtain total model loss data; Based on the total model loss data, perform parameter adjustment on the initial speech synthesis model to obtain the target speech synthesis model.

4. The method according to claim 3, characterized in that The performing loss calculation on the speech feature vectors, the masked phrase embedding vectors, the masked text embedding vectors, the predicted speech feature vectors, the predicted phrase embedding vectors, and the predicted text embedding vectors based on the initial speech synthesis model to obtain total model loss data includes: Based on the initial speech synthesis model, calculate the loss between the predicted speech feature vector and the speech feature vector to obtain speech feature loss data; Based on the initial speech synthesis model, calculate the loss between the predicted phrase embedding vector and the masked phrase embedding vector to obtain masked phrase loss data; Based on the initial speech synthesis model, calculate the loss between the predicted text embedding vector and the masked text embedding vector to obtain text prediction loss data; Merge the speech feature loss data, the masked phrase loss data, and the text prediction loss data to obtain the total model loss data.

5. The method according to claim 1, characterized in that The acoustic feature alternating update based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vector to obtain the subsequent acoustic feature vector includes: Based on the target speech synthesis model, concatenate the starting acoustic feature vector and the subsequent text embedding vector to obtain a concatenated feature vector; Based on the target speech synthesis model, perform acoustic feature conversion on the concatenated feature vector to obtain the subsequent acoustic feature vector.

6. The method according to any one of claims 1-5, characterized in that, The embedding process for the target text data to obtain a sequence of text embedding vectors includes: Perform word segmentation on the target text data to obtain segmented text blocks; Perform vector mapping on the segmented text blocks to obtain the sequence of text embedding vectors.

7. The method according to any one of claims 1-5, characterized in that, The speech synthesis based on the target speech synthesis model for the Mel spectrogram sequence includes: Based on the target speech synthesis model, perform waveform conversion on the Mel spectrogram sequence to obtain a speech time-domain waveform; Based on the target speech synthesis model, perform signal conversion on the speech time-domain waveform to obtain a speech signal; Based on the target speech synthesis model, perform audio synthesis on the speech signal to obtain the target synthesized speech.

8. A voice synthesis device, characterized in that, The device includes: A data acquisition module for acquiring target text data; A text embedding module for performing an embedding process on the target text data to obtain a sequence of text embedding vectors, where the sequence of text embedding vectors includes a starting text embedding vector and a subsequent text embedding vector; A feature conversion module for performing acoustic feature conversion on the starting text embedding vector based on a preset target speech synthesis model to obtain a starting acoustic feature vector; An interleaved encoding module for performing acoustic feature alternating update based on the target speech synthesis model, the starting acoustic feature vector, and the subsequent text embedding vector to obtain a subsequent acoustic feature vector; A spectrogram conversion module for performing spectrogram conversion on the starting acoustic feature vector based on the target speech synthesis model to obtain a starting Mel spectrogram; A spectrogram alternating conversion module for performing spectrogram alternating conversion on the subsequent acoustic feature vector based on the target speech synthesis model and the starting Mel spectrogram to obtain a subsequent Mel spectrogram; A spectrogram nested merge module for performing spectrogram nested merge on the starting Mel spectrogram and the subsequent Mel spectrogram to obtain a Mel spectrogram sequence; A speech synthesis module for performing speech synthesis on the Mel spectrogram sequence based on the target speech synthesis model.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the voice synthesis method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the voice synthesis method according to any one of claims 1 to 7 is implemented.