Speech synthesis method and device, electronic equipment and storage medium

By employing a two-stage approach involving phoneme sequence duration prediction and feature generation, the problems of time consumption and low naturalness in zero-sample speech synthesis are solved, thus achieving high-quality speech synthesis.

CN120877708APending Publication Date: 2025-10-31BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510926649.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing zero-shot speech synthesis methods are time-consuming and produce synthesized speech with low naturalness and high word error rate, resulting in poor speech quality.

Method used

By acquiring the text to be synthesized and the prompt speech, the phoneme sequence duration is predicted, the phoneme sequence length is adjusted, and the target semantic and acoustic features are generated by combining feature extraction and prior distribution. Finally, the target synthesized speech is generated, and a two-stage method of text-to-semantics and semantics-to-acoustics is adopted without phoneme-level duration prediction.

Benefits of technology

It improves speech synthesis speed, reduces word error rate, and enhances the naturalness and quality of synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877708A_ABST
    Figure CN120877708A_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-synthesized text and prompt speech; performing duration prediction based on a phoneme sequence of the to-be-synthesized text and the prompt voice to obtain playing duration information of a target synthesized voice; adjusting the sequence length of the phoneme sequence based on the playing time length information to obtain a target phoneme sequence; the sequence length of the target phoneme sequence is matched with the playing duration information; performing feature extraction on the target phoneme sequence, and generating target semantic features based on a feature extraction result and prior distribution; target acoustic features are generated based on the target semantic features and the prior distribution, and the target synthetic speech is generated based on the target acoustic features. The speech synthesis speed is improved, the naturalness of the synthetic speech is high, the word error rate is low, and the quality of the synthetic speech is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a speech synthesis method, apparatus, electronic device, and storage medium. Background Technology

[0002] Text-to-Speech (TTS) is a technology that converts text information into natural speech. By simulating the characteristics of human speech, such as pitch, speed, tone, and intonation, it enables computers to express text information like humans, achieving a more humanized and intelligent interactive experience.

[0003] In recent years, with the development of Large Language Model (LLM) technology, zero-shot speech synthesis (ZTS) has been widely studied. Leveraging the in-context learning capabilities of LLM, it can adapt to new timbres without additional training. However, ZTS methods in related technologies are not only time-consuming in the speech synthesis process, but also suffer from low naturalness and high word error rates, resulting in poor quality synthesized speech. Summary of the Invention

[0004] This disclosure provides a speech synthesis method, apparatus, electronic device, and storage medium to at least solve the problems in related technologies where the speech synthesis process is not only time-consuming, but also results in low naturalness and a high word error rate in the synthesized speech, leading to poor quality of the synthesized speech. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a speech synthesis method is provided, comprising: Obtain the text to be synthesized and the prompt speech; Based on the phoneme sequence of the text to be synthesized and the prompt speech, the playback duration information of the target synthesized speech is obtained by predicting the duration. The sequence length of the phoneme sequence is adjusted based on the playback duration information to obtain the target phoneme sequence; the sequence length of the target phoneme sequence is matched with the playback duration information. Feature extraction is performed on the target phoneme sequence, and target semantic features are generated based on the feature extraction results and prior distribution; Target acoustic features are generated based on the target semantic features and the prior distribution, and target synthesized speech is generated based on the target acoustic features.

[0005] In one exemplary embodiment, the step of predicting the playback duration of the target synthesized speech based on the phoneme sequence of the text to be synthesized and the prompt speech includes: Text encoding features are obtained by performing text encoding processing based on the phoneme sequence of the text to be synthesized; Based on the prompting voice, the acoustic features of the target object to which the prompting voice belongs are extracted, and the acoustic features are style-coded to obtain the style-coded features of the target object; The style encoding features of the target object are fused with the text encoding features to obtain fused encoding features; Based on the fused coding features and the prior distribution, the playback duration information of the target synthesized speech is generated.

[0006] In one exemplary embodiment, fusing the style encoding features of the target object with the text encoding features to obtain fused encoding features includes: Self-attention processing is performed on the text encoding features to obtain key text features; The style encoding features of the target object are subjected to affine coupling processing to obtain scaling and translation coefficients; The key text features are subjected to instance normalization to obtain instance normalized text features; The normalized text features of the instance are linearly adjusted based on the scaling and translation coefficients to obtain the fused encoding features.

[0007] In one exemplary implementation, generating the playback duration information of the target synthesized speech based on the fused coding features and the prior distribution includes: Determine the transcript sequence of the text corresponding to the prompt speech; Determine the ratio between the sequence length of the transcribed audio sequence and the duration of the prompt speech; The fused encoded features are concatenated with the ratio; Based on the splicing result and the prior distribution, the playback duration information of the target synthesized speech is generated.

[0008] In one exemplary implementation, generating target semantic features based on the feature extraction results and prior distribution includes: Based on the result of the feature extraction, the sampling result sampled from the prior distribution is mapped using the first distribution mapping relationship to obtain the target semantic features; the first distribution mapping relationship indicates the mapping relationship between the prior distribution and the speech feature distribution.

[0009] In an exemplary implementation, generating target acoustic features based on the target semantic features and the prior distribution includes: Based on the prompting voice, extract the voiceprint features of the target object to which the prompting voice belongs; Based on the target semantic features and the voiceprint features, the sampling results sampled from the prior distribution are mapped using a second distribution mapping relationship to obtain the target acoustic features; the second distribution mapping relationship indicates the mapping relationship between the prior distribution and the acoustic feature distribution.

[0010] In one exemplary embodiment, the speech synthesis method is implemented based on a speech synthesis model, the speech synthesis model including a duration prediction sub-model, and the method further includes training the duration prediction sub-model, wherein training the duration prediction sub-model includes: Obtain the first sample speech and the first sample transcribed text corresponding to the first sample speech; Determine the first sample phoneme sequence of the transcribed text of the first sample, and the first sample acoustic features of the speech of the first sample; The first sample phoneme sequence is input into the text encoding network of the duration prediction sub-model to obtain the sample text encoding features; the first sample acoustic features are input into the style encoding network of the duration prediction sub-model to obtain the sample style encoding features of the sample object to which the first sample speech belongs. The sample text encoding features and the sample style encoding features are input into a fusion network for fusion to obtain sample fusion features; The sample fusion features are used as conditional inputs to the first generative model based on stream matching for speech duration generation processing. The duration prediction sub-model is then trained based on the results of the speech duration generation processing to obtain the trained duration prediction sub-model.

[0011] In one exemplary embodiment, the speech synthesis model further includes a text-to-semantic sub-model and a semantic-to-acoustic feature sub-model, and the method further includes: Obtain the second sample speech and the corresponding second sample transcribed text; Determine the sample speech features corresponding to the second sample speech and the second sample phoneme sequence corresponding to the second sample text; perform partial masking on the sample speech features to obtain sample masked speech features; The second sample speech and the second sample phoneme sequence are input into the trained duration prediction sub-model to predict the speech duration and obtain the predicted sample playback duration; the sequence length of the second sample phoneme sequence is adjusted based on the sample playback duration to obtain the target sample phoneme sequence; The target sample phoneme sequence is input into the feature extraction network of the text-to-semantic sub-model to extract sample features and obtain sample phoneme features; the sample phoneme features and the sample mask speech features are used as conditions to input into the second generative model of the text-to-semantic sub-model based on stream matching to generate semantic features, and the text-to-semantic sub-model is trained based on the result of the semantic feature generation process to obtain the trained text-to-semantic sub-model; Acquire the acoustic features of the second sample speech, partially mask the acoustic features to obtain the masked acoustic features; input the sample speech features and the masked acoustic features as conditions into a third generative model based on stream matching for acoustic feature generation processing, and train the third generative model based on the results of the acoustic feature generation processing to obtain a trained semantic-to-acoustic feature sub-model.

[0012] According to a second aspect of the present disclosure, a speech synthesis apparatus is provided, comprising: The acquisition unit is configured to acquire the text to be synthesized and the prompt speech; The duration prediction unit is configured to perform duration prediction based on the phoneme sequence of the text to be synthesized and the prompt speech to obtain the playback duration information of the target synthesized speech; A phoneme sequence adjustment unit is configured to adjust the sequence length of the phoneme sequence based on the playback duration information to obtain a target phoneme sequence; the sequence length of the target phoneme sequence matches the playback duration information. The text-to-semantic unit is configured to perform feature extraction on the target phoneme sequence and generate target semantic features based on the result of the feature extraction and the prior distribution. The synthesized speech generation unit is configured to generate target acoustic features based on the target semantic features and the prior distribution, and generate the target synthesized speech based on the target acoustic features.

[0013] In one exemplary embodiment, the duration prediction unit includes: The text encoding unit is configured to perform text encoding processing based on the phoneme sequence of the text to be synthesized, to obtain text encoding features; The style coding unit is configured to perform the following operations: extract the acoustic features of the target object to which the prompting voice belongs based on the prompting voice, perform style coding processing on the acoustic features, and obtain the style coding features of the target object; The fusion unit is configured to perform the fusion of the style coding features of the target object with the text coding features to obtain fused coding features; The duration information generation unit is configured to generate playback duration information of the target synthesized speech based on the fused coding features and the prior distribution.

[0014] In one exemplary embodiment, the fusion unit includes: An attention processing unit is configured to perform self-attention processing on the text encoded features to obtain key text features; An affine coupling unit is configured to perform affine coupling processing on the style-encoded features of the target object to obtain scaling and translation coefficients. The instance normalization unit is configured to perform instance normalization processing on the key text features to obtain instance normalized text features. The linear adjustment unit is configured to perform linear adjustment on the instance normalized text features based on the scaling and translation coefficients to obtain the fused encoded features.

[0015] In one exemplary embodiment, the duration information generation unit includes: The first determining unit is configured to determine the transcript sequence of the transcribed text corresponding to the prompt speech. The second determining unit is configured to determine the ratio between the sequence length of the transcribing element sequence and the speech duration of the prompt speech; The splicing unit is configured to perform the splicing of the fused encoded features with the ratio; The duration information generation subunit is configured to generate playback duration information of the target synthesized speech based on the splicing result and the prior distribution.

[0016] In an exemplary embodiment, the text-to-semantic unit, when generating target semantic features based on the feature extraction results and prior distribution, is specifically configured to perform: based on the feature extraction results, using a first distribution mapping relationship to map the sampling results sampled from the prior distribution to obtain the target semantic features; the first distribution mapping relationship indicates the mapping relationship between the prior distribution and the speech feature distribution.

[0017] In an exemplary embodiment, the speech synthesis unit, when generating target acoustic features based on the target semantic features and the prior distribution, is specifically configured to perform the following: extracting the voiceprint features of the target object to which the prompting voice belongs based on the prompting voice; and mapping the sampling results sampled from the prior distribution using a second distribution mapping relationship based on the target semantic features and the voiceprint features to obtain the target acoustic features; wherein the second distribution mapping relationship indicates the mapping relationship between the prior distribution and the acoustic feature distribution.

[0018] In an exemplary embodiment, the apparatus further includes a first training unit configured to perform: acquiring a first sample speech and a first sample transcribed text corresponding to the first sample speech; determining a first sample phoneme sequence of the first sample transcribed text and a first sample acoustic feature of the first sample speech; inputting the first sample phoneme sequence into the text encoding network of the duration prediction sub-model to obtain sample text encoding features; inputting the first sample acoustic features into the style encoding network of the duration prediction sub-model to obtain sample style encoding features of the sample object to which the first sample speech belongs; inputting the sample text encoding features and the sample style encoding features into a fusion network for fusion to obtain sample fusion features; inputting the sample fusion features as conditions into a first generative model based on stream matching for speech duration generation processing, and training the duration prediction sub-model based on the result of the speech duration generation processing to obtain a trained duration prediction sub-model.

[0019] In one exemplary embodiment, the apparatus further includes a second training unit configured to perform: acquiring a second sample speech and a second sample transcribed text corresponding to the second sample speech; determining sample speech features corresponding to the second sample speech and a second sample phoneme sequence corresponding to the second sample text; partially masking the sample speech features to obtain sample masked speech features; inputting the second sample speech and the second sample phoneme sequence into the trained duration prediction submodel to predict speech duration, obtaining a predicted sample playback duration; adjusting the sequence length of the second sample phoneme sequence based on the sample playback duration to obtain a target sample phoneme sequence; and inputting the target sample phoneme sequence into the feature extraction network of the text-to-semantic submodel for further processing. Sample features are extracted to obtain sample phoneme features; the sample phoneme features and the sample mask speech features are used as conditions to input into the second generative model based on flow matching of the text-to-semantic sub-model for semantic feature generation processing, and the text-to-semantic sub-model is trained based on the result of the semantic feature generation processing to obtain a trained text-to-semantic sub-model; sample acoustic features of the second sample speech are obtained, and the sample acoustic features are partially masked to obtain sample mask acoustic features; the sample speech features and the sample mask acoustic features are used as conditions to input into the third generative model based on flow matching for acoustic feature generation processing, and the third generative model is trained based on the result of the acoustic feature generation processing to obtain a trained semantic-to-acoustic feature sub-model.

[0020] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the speech synthesis method in the first aspect described above.

[0021] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the speech synthesis method of the first aspect described above.

[0022] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the speech synthesis method of the first aspect described above.

[0023] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: By acquiring the text to be synthesized and the prompt speech, the phoneme sequence corresponding to the text to be synthesized is determined. Then, based on the phoneme sequence and the prompt speech, the speech duration is predicted to obtain the playback duration information of the target synthesized speech. Based on the playback duration information, the phoneme sequence is adjusted to obtain a target phoneme sequence whose sequence length matches the playback duration information. Feature extraction is performed on the target phoneme sequence. Based on the feature extraction results and prior distribution, target semantic features are generated. Based on the target semantic features and prior distribution, target acoustic features are generated. Based on the target acoustic features, target synthesized speech is generated. Thus, by combining the phoneme sequence of the text to be synthesized and the prompt speech to predict the playback duration information at the sentence level, a speech synthesis method that does not require phoneme-level duration prediction in two stages (text to semantics and semantics to acoustics) is adopted. This avoids the problems of long time consumption and limited naturalness of synthesized speech caused by phoneme-level duration prediction, as well as the problem of high word error rate caused by predicting playback duration only using text. Thus, while improving the speech synthesis speed, the quality of synthesized speech is greatly improved due to the high naturalness of synthesized speech and the low word error rate.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0026] Figure 1 This is a schematic diagram illustrating the implementation environment of a speech synthesis method according to an exemplary embodiment; Figure 2This is a flowchart illustrating a speech synthesis method according to an exemplary embodiment; Figure 3 This is a flowchart illustrating another speech synthesis method according to an exemplary embodiment; Figure 4 This is a schematic diagram of the structure of a speech synthesis model according to an exemplary embodiment; Figure 5 This is a block diagram illustrating a speech synthesis apparatus according to an exemplary embodiment; Figure 6 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0028] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0029] Please see Figure 1 The diagram illustrates an implementation environment for a speech synthesis method according to an exemplary embodiment. The implementation environment may include a terminal 110 and a server 120, which can communicate with each other via a wired network or a wireless network.

[0030] Terminal 110 can be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. Terminal 110 may have client software, such as an application (App), installed that provides speech synthesis functionality. This application can be a dedicated speech synthesis application or other applications with speech synthesis capabilities, such as live streaming applications with speech synthesis functionality. Users of Terminal 110 can log in to the application using pre-registered user information, which may include an account and password.

[0031] Server 120 can be a server that provides speech synthesis services for applications in terminal 110. It should be noted that the server involved in this embodiment can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0032] Specifically, server 120 can store a speech synthesis model that can achieve zero-shot speech synthesis. In this embodiment, the speech synthesis model is a generative model based on stream matching. Server 120 can train and update the speech synthesis model according to a predetermined period. When terminal 110 needs to perform speech synthesis, it can send the corresponding text to be synthesized and the prompt voice to server 120. Server 120 calls the speech synthesis model to perform speech synthesis based on the text to be synthesized and the prompt voice, and returns the synthesized speech to terminal 110. The content of the synthesized speech comes from the text to be synthesized, and the timbre, emotion, rhythm, and other sound information in the synthesized speech comes from the prompt voice.

[0033] Understandably, terminal 110 can also download the speech synthesis model from server 120 and store it locally. When speech synthesis is required, terminal 110 can directly call the locally stored speech synthesis model to perform speech synthesis.

[0034] Figure 2 This is a flowchart illustrating a speech synthesis method according to an exemplary embodiment, such as... Figure 2 As shown, the method may include the following steps S201 to S209: In step 201, the text to be synthesized and the prompt voice are obtained.

[0035] The text to be synthesized is the text corresponding to the speech content to be synthesized. The text to be synthesized can be Chinese text, English text, mixed Chinese and English text, or text in other languages.

[0036] The prompt voice is the audio corresponding to the timbre, rhythm, emotion, and other sound information of the synthesized speech. The playback duration of this audio can be 5 seconds, 10 seconds, or 20 seconds, etc. The "speaker" corresponding to this prompt voice is the target object to which the prompt voice belongs. It can be understood that the target object can be a person or animal with biological characteristics, or any object that can emit audio with speech content, such as a virtual character.

[0037] In practice, the terminal can display a speech synthesis interface to the user. This interface includes a text input area and a prompt voice input area. The text input area is used to input the text to be synthesized, and the prompt voice input area is used to input the prompt voice. The prompt voice can be input in real-time recording or from an uploaded existing audio text. The terminal can respond to a speech synthesis command triggered by this interface by sending a speech synthesis request to the server based on the input information from the text input area and the prompt voice input area. This request carries the text to be synthesized from the text input area and the prompt voice from the prompt voice input area. The server then receives the request, parses it, and obtains the text to be synthesized and the prompt voice.

[0038] In step S203, the playback duration information of the target synthesized speech is obtained by predicting the duration based on the phoneme sequence of the text to be synthesized and the prompt speech.

[0039] A phoneme is the smallest unit of speech determined by the natural properties of speech. For example, in Chinese pronunciation, an initial consonant or a final vowel can be considered a phoneme. In other languages, each pronunciation is also equivalent to a phoneme. A phoneme sequence is a sequence of phonemes in the text to be synthesized, arranged in the order they appear in the text. The phoneme sequence of the text to be synthesized can be obtained by performing appropriate transformations on the text.

[0040] The target synthesized speech is the final synthesized speech. Its content originates from the text to be synthesized, while its timbre, emotion, rhythm, and other auditory information come from the prompting speech. The playback duration information indicates the number of frames in the target synthesized speech, representing the length of the entire sentence.

[0041] In some exemplary implementations, such as Figure 3 As shown, step S203 above, when performing duration prediction based on the phoneme sequence of the text to be synthesized and the prompt speech to obtain the playback duration information of the target synthesized speech, may include: In step S301, text encoding processing is performed based on the phoneme sequence of the text to be synthesized to obtain text encoding features.

[0042] Specifically, the phoneme sequence of the text to be synthesized can be embedded first to obtain a phoneme embedding vector sequence, which includes the embedding vector of each phoneme. Then, the phoneme embedding vector sequence can be processed by text encoding to obtain text encoding features.

[0043] Among them, text encoding processing can be based on self-attention mechanism. Specifically, a multi-layer self-attention network can be used to perform text encoding based on self-attention mechanism in sequence, and the text encoding result output by the last layer of self-attention network can be obtained.

[0044] To more accurately represent the semantic information corresponding to the phoneme pronunciation space, text encoding processing can use a multi-layer text encoding network based on an attention mechanism and a pre-trained BERT model to encode the phoneme sequence, thereby obtaining the text encoding result output by the multi-layer text encoding network based on the attention mechanism and the text encoding result output by the pre-trained BERT model. Then, the two text encoding results are added together, and the result of the addition is used as the text encoding processing to obtain the text encoding feature.

[0045] In step S303, based on the prompting voice, the acoustic features of the target object to which the prompting voice belongs are extracted, and the acoustic features are style-coded to obtain the style-coded features of the target object.

[0046] Specifically, the acoustic features of the target object to which the prompting speech belongs can be a Mel spectrum.

[0047] Among them, style coding processing of the acoustic features of the target object to which the prompting voice belongs can extract the style information of the target object, including prosody, rhythm, speech rate and timbre, that is, the style coding features of the target object can represent the style information of the target object.

[0048] In practice, style coding can be performed using a style coding network. This involves inputting the acoustic features of the target object to which the prompt speech belongs into the style coding network for style coding to obtain the output style-coded features. The style coding network can consist of convolutional layers, attention layers, and average pooling layers.

[0049] In step S305, the style encoding features of the target object are fused with the text encoding features to obtain fused encoding features.

[0050] Specifically, a cross-attention mechanism can be used to fuse the style encoding features and text encoding features of the target object to obtain fused encoding features. In the cross-attention mechanism, the query Q is the style encoding feature, and the key K and value V are the text encoding features.

[0051] In some exemplary embodiments, in order to improve the quality of synthesized speech, step S305 above may include, when performing the fusion of the style coding features of the target object with the text coding features to obtain fused coding features: Self-attention processing is performed on the text encoding features to obtain key text features; The style encoding features of the target object are subjected to affine coupling processing to obtain scaling and translation coefficients; The key text features are subjected to instance normalization to obtain instance normalized text features; The normalized text features of the instance are linearly adjusted based on the scaling and translation coefficients to obtain the fused encoding features.

[0052] Specifically, self-attention processing can be achieved using the following formula (1): (1) in, Represents the query matrix. Represents the key matrix, Represents a value matrix, The softmax value represents the number of channels in the feature sequence within the self-attention mechanism. ) represents normalization, and T represents transpose. In this embodiment of the disclosure, when performing self-attention processing on text encoding features, the query matrix... Key matrix Sum matrix All of these originate from text encoding features, and are thus derived from the formulas described above. This is the key text feature h.

[0053] Instance normalization of key text features ensures that these features follow a standard normal distribution. Specifically, instance normalization of key text feature h can be performed as follows: ,in, The mean of the key text feature h is represented. The standard deviation of the key text feature h is represented.

[0054] In this process, affine coupling processing of the style encoding features of the target object can input these style encoding features s into the affine coupling transformation layer to obtain the gain. and bias The gain That is, the bias is used as a scaling factor. That is, it serves as the translation coefficient.

[0055] Specifically, the normalized text features of the instance are linearly adjusted based on the scaling and translation coefficients to obtain the fused encoded features. This can be expressed as formula (2): (2).

[0056] In the above implementation, key text features are obtained by performing self-attention processing on the text encoding features, and then instance normalization is performed on the key text features to obtain instance normalized text features. Affine coupling processing is performed on the style encoding features to obtain scaling and translation coefficients. The instance normalized text features are then linearly adjusted using the scaling and translation coefficients to obtain fused encoding features. This results in fused encoding features that contain style information of the target object while having richer text content, which is beneficial to improving the quality of synthesized speech.

[0057] In step S307, based on the fusion coding features and prior distribution, the playback duration information of the target synthesized speech is generated.

[0058] Specifically, the fused encoded features can be used as conditional inputs to a trained first generative model based on stream matching to generate playback time information for the target synthesized speech. This trained first generative model based on stream matching generates a distribution of sentence duration based on the sampling results of the prior distribution and the fused encoded features, thus obtaining the playback duration information. The prior distribution can be a Gaussian distribution. The trained first generative model based on stream matching can be a multi-layer stacked structure, with each layer including: a normalization layer (actnorm), an invertible 1x1 convolution, and a coupling layer. This trained first generative model based on stream matching is used to transform the Gaussian distribution into the target distribution, which is the distribution of sentence duration.

[0059] In some exemplary embodiments, step S307 above, when generating playback duration information of the target synthesized speech based on the fusion coding features and prior distribution, may include: determining the transcribed text sequence corresponding to the prompt speech; determining the ratio between the sequence length of the transcribed text sequence and the speech duration of the prompt speech; concatenating the fusion coding features with the ratio; and generating playback duration information of the target synthesized speech based on the concatenation result and the prior distribution.

[0060] Specifically, the text corresponding to the prompt speech is first determined to obtain the transcribed text. Then, the phoneme sequence corresponding to the transcribed text is determined to obtain the transcribing phoneme sequence. The ratio between the sequence length of the transcribing phoneme sequence and the speech duration of the prompt speech is calculated. The fused coding features are then concatenated with this ratio. The concatenation result is used as a conditional input to the trained first generative model based on stream matching to generate the playback duration information of the target synthesized speech.

[0061] The above implementation calculates the ratio between the sequence length of the transcribed phoneme sequence and the speech duration of the prompt speech, and concatenates the fused coding features with this ratio. Based on the concatenation result and prior distribution, it generates playback duration information of the target synthesized speech, thereby making full use of the clearer relationship between phonemes and pronunciation length, improving the prediction accuracy of playback duration information, and helping to improve the quality of synthesized speech.

[0062] The above implementation method performs text encoding processing on the phoneme sequence of the text to be synthesized to obtain text encoding features, and performs style encoding processing on the acoustic features of the prompt speech to obtain style encoding features. Then, the text encoding features and style encoding features are fused to obtain fused encoding features. Based on the fused encoding features and prior distribution, playback duration information is generated. This method considers both the phoneme sequence and the prompt speech, and makes full use of the prompt speech to generate sentence-level playback duration information, ensuring the accuracy of the playback duration information, avoiding phenomena such as missing words or extra words, and improving the quality of synthesized speech.

[0063] In step S205, the sequence length of the phoneme sequence is adjusted based on the playback duration information to obtain a target phoneme sequence, wherein the sequence length of the target phoneme sequence matches the playback duration information.

[0064] Specifically, the length of the phoneme sequence can be adjusted to match the number of frames corresponding to the playback duration information. In practice, preset symbols are used to supplement the phoneme sequence to the estimated number of frames corresponding to the entire sentence playback duration information. For example, the preset symbols could be... <pad>symbol.

[0065] In step S207, feature extraction is performed on the target phoneme sequence, and target semantic features are generated based on the result of the feature extraction and the prior distribution.

[0066] Specifically, the target phoneme sequence can be convolved to extract features and obtain the feature extraction results.

[0067] The prior distribution can be a Gaussian distribution.

[0068] When generating target semantic features, the results of the above feature extraction can be used as conditions to input into a trained second generative model based on flow matching. The trained second generative model based on flow matching generates target semantic features based on the sampling results from the prior distribution and the results of the above feature extraction, and the target semantic features are continuous representations.

[0069] The trained second generative model based on flow matching can be a multi-layer stacked structure, each layer including: a normalization layer (actnorm), an invertible 1x1 convolution, and a coupling layer. The trained second generative model based on flow matching is used to transform the Gaussian distribution into a continuous semantic representation space under the guidance of the feature extraction results of the target phoneme sequence. The target semantic features are features in the continuous semantic representation space.

[0070] In an exemplary embodiment, in order to improve the accuracy of the target semantic features, step S207 above may include generating the target semantic features based on the result of feature extraction and the prior distribution: based on the result of feature extraction, using a first distribution mapping relationship to map the sampling result sampled from the prior distribution to obtain the target semantic features, wherein the first distribution mapping relationship indicates the mapping relationship between the prior distribution and the speech feature distribution.

[0071] Specifically, the first distribution mapping relationship can be the transformation relationship corresponding to the trained second generative model based on flow matching. The speech feature distribution can be a continuous Hubert semantic representation space obtained based on self-supervised learning. Then, guided by the feature extraction results, the target semantic features obtained by mapping the sampling results sampled from the prior distribution using the first distribution mapping relationship are the target Hubert features. This avoids information loss during the transformation process, ensures the accuracy of the target semantic features, and helps improve the quality of synthesized speech.

[0072] In step S209, target acoustic features are generated based on the target semantic features and the prior distribution, and target synthesized speech is generated based on the target acoustic features.

[0073] Specifically, the target semantic features can be used as conditional inputs to a trained flow-matching-based third generative model, which generates target acoustic features based on sampling results from a prior distribution and the target semantic features. For example, the prior distribution could be a Gaussian distribution, and the target acoustic features could be a Mel spectrum.

[0074] The trained third generative model based on flow matching can be a multi-layer stacked structure, each layer including: a normalization layer (actnorm), an invertible 1x1 convolution, and a coupling layer. The trained third generative model based on flow matching is used to transform a Gaussian distribution into an acoustic feature space under the guidance of target semantic features, where the target acoustic features are features in the acoustic feature space.

[0075] In some exemplary embodiments, step S209 above, when generating target acoustic features based on the target semantic features and the prior distribution, may include: Based on the prompting voice, extract the voiceprint features of the target object to which the prompting voice belongs; Based on the target semantic features and the voiceprint features, the sampling results sampled from the prior distribution are mapped using a second distribution mapping relationship to obtain the target acoustic features; the second distribution mapping relationship indicates the mapping relationship between the prior distribution and the acoustic feature distribution.

[0076] Specifically, a pre-trained voiceprint recognition model can be used to perform voiceprint recognition on the prompt speech to obtain the voiceprint features of the target pair to which the prompt speech belongs. The pre-trained voiceprint recognition model can be the ECAPA-TDNN model.

[0077] The second distribution mapping relationship can be the transformation relationship corresponding to the trained third generative model based on flow matching. It can concatenate the target semantic features and the voiceprint features of the target object, and use the concatenated features as a condition input to the trained third generative model based on flow matching. Then, under the guidance of the target semantic features and the voiceprint features of the target object, the sampling results sampled from the Gaussian distribution are mapped using the second distribution mapping relationship to obtain the target acoustic features, thereby improving the accuracy of the target acoustic features and improving the quality of synthesized speech.

[0078] Specifically, when generating target synthesized speech based on target acoustic features, these features can be input into a vocoder. The vocoder then reconstructs the speech waveform from these features, thus obtaining the target synthesized speech. The vocoder can be WaveNet, Griffin-Lim, a single-layer recurrent neural network model like WaveRNN, or a Parallel WaveGan based on a non-autoregressive network, etc., to achieve better sound quality and a sound quality effect close to that of a real person speaking.

[0079] The speech synthesis method of this disclosure can be implemented based on a speech synthesis model, which is a generative model based on stream matching. Specifically, as shown... Figure 4 The diagram shows the structure of a speech synthesis model, which includes a text-to-semantic sub-model, a semantic-to-acoustic feature sub-model, a duration prediction sub-model, and a vocoder.

[0080] The duration prediction sub-model can include a text encoding network, a style encoding network, a fusion network, and a first generative model based on stream matching. The text encoding network can include a multi-layer (e.g., 6-layer) text encoding network based on an attention mechanism and a pre-trained BERT model, used to separately encode the phoneme sequences of the text to be synthesized to obtain their respective output text encoding results. These two text encoding results are then added together to obtain the text encoding features obtained by the text encoding network. The style encoding network takes the acoustic features (e.g., Mel spectrum) of the target object to which the cues belong as input for style encoding to obtain the style encoding features of that target object. The fusion network is used to fuse text encoding features and style encoding features to obtain fused encoding features. These fused encoding features will be used as conditional inputs to the first generative model based on flow matching. Therefore, the fusion network can also be called a conditional encoding network. Specifically, the conditional encoding network can be composed of multiple attention mechanism modules with adaptive layer normalization (ALN). Each attention mechanism module with ALN includes an attention layer, an affine coupling layer, and an ALN ​​layer. The text encoding features are input to the attention layer for self-attention processing to obtain the key text features of the output. The style encoding features are input to the affine coupling layer for affine coupling processing to obtain the gain as a scaling factor and the bias as a translation factor of the output. The outputs of the attention layer and the affine coupling layer are fed into the ALN layer for adaptive normalization processing. The specific adaptive normalization processing can be found in the aforementioned formula (2), thereby obtaining the fused encoding features output by the ALN layer. The first generative model based on stream matching is used to generate the distribution of sentence duration based on the sampling results of the prior distribution and the fused coding features, thereby obtaining the whole sentence playback duration information of the synthesized speech. For details, please refer to the relevant description of step S307 above, which will not be repeated here.

[0081] The text-to-semantic sub-model may include a feature extraction network and a flow-matching-based second generative model. The feature extraction network extracts features from the target phoneme sequence, and the result of this feature extraction is used as conditional input to the flow-matching-based second generative model. For example, the feature extraction network may have an improved neural network structure capable of capturing multi-level text features, ensuring accurate alignment between text and speech. The flow-matching-based second generative model generates target semantic features in a continuous Hubert semantic representation space based on sampling results from a prior distribution and the feature extraction results. For details, please refer to the relevant description in step S207 above, which will not be repeated here.

[0082] The semantic-to-acoustic feature sub-model can be a third generative model based on stream matching, used to generate target acoustic features (such as the target Mel spectrum) from sampling results from a prior distribution and target semantic features. A vocoder is used to restore the target acoustic features (such as the target Mel spectrum) to a playable speech waveform, thus obtaining the target synthesized speech.

[0083] This embodiment combines the phoneme sequence of the text to be synthesized with the playback duration information at the sentence level of the prompt speech prediction, and then adopts a stream matching-based speech synthesis method that does not require phoneme-level duration prediction in two stages: text to semantics and semantics to acoustics. This not only avoids the problem of poor stability caused by the difficulty of directly converting the text mode to the acoustic mode, but also improves the speech synthesis speed, and the synthesized speech has high naturalness and low word error rate, which greatly improves the quality of the synthesized speech.

[0084] The training process of the above speech synthesis feature model is briefly described below. This training process may include training the duration prediction sub-model, as well as training the text-to-semantic sub-model and the semantic-to-acoustic feature sub-model.

[0085] The training process for the duration prediction sub-model may include the following steps: Obtain the first sample speech and the first sample transcribed text corresponding to the first sample speech; Determine the first sample phoneme sequence of the transcribed text of the first sample, and the first sample acoustic features of the speech of the first sample; The first sample phoneme sequence is input into the text encoding network of the duration prediction sub-model to obtain the sample text encoding features; the first sample acoustic features are input into the style encoding network of the duration prediction sub-model to obtain the sample style encoding features of the sample object to which the first sample speech belongs. The sample text encoding features and the sample style encoding features are input into a fusion network for fusion to obtain sample fusion features; The sample fusion features are used as conditional inputs to the first generative model based on stream matching for speech duration generation processing. The duration prediction sub-model is then trained based on the results of the speech duration generation processing to obtain the trained duration prediction sub-model.

[0086] Specifically, the loss function for training the duration prediction sub-model can be shown in the following formula (3): (3) Where E represents the mathematical expectation over the model prediction error on the full probability distribution; The first generative model based on flow matching represents the time step The output represents the flow rate. and These represent the true duration and Gaussian noise, respectively. Represents the duration of the noisy signal, i.e. .

[0087] The above implementation method trains a duration prediction sub-model, enabling the trained duration prediction sub-model to combine the prompt speech and the input phoneme sequence to achieve accurate sentence-level duration prediction. This improves the accuracy of the whole sentence playback duration information of the synthesized speech and the efficiency of speech synthesis, thereby helping to improve the quality of synthesized speech.

[0088] Training the text-to-semantic sub-model and the semantic-to-acoustic feature sub-model may include the following steps: Obtain the second sample speech and the corresponding second sample transcribed text; Determine the sample speech features corresponding to the second sample speech and the second sample phoneme sequence corresponding to the second sample text; perform partial masking on the sample speech features to obtain sample masked speech features; The second sample speech and the second sample phoneme sequence are input into the trained duration prediction sub-model to predict the speech duration and obtain the predicted sample playback duration; the sequence length of the second sample phoneme sequence is adjusted based on the sample playback duration to obtain the target sample phoneme sequence; The target sample phoneme sequence is input into the feature extraction network of the text-to-semantic sub-model to extract sample features and obtain sample phoneme features; the sample phoneme features and the sample mask speech features are used as conditions to input into the second generative model of the text-to-semantic sub-model based on stream matching to generate semantic features, and the text-to-semantic sub-model is trained based on the result of the semantic feature generation process to obtain the trained text-to-semantic sub-model; Acquire the acoustic features of the second sample speech, partially mask the acoustic features to obtain the masked acoustic features; input the sample speech features and the masked acoustic features as conditions into a third generative model based on stream matching for acoustic feature generation processing, and train the third generative model based on the results of the acoustic feature generation processing to obtain a trained semantic-to-acoustic feature sub-model.

[0089] Specifically, the loss function for training the text-to-semantic sub-model is shown in the following formula (4): (4) Where E represents the mathematical expectation over the model prediction error over the full probability distribution; The second generative model based on flow matching represents the time step The output represents the flow rate; and These represent the sample speech features (such as the true Hubert representation) and the prior distribution (such as Gaussian noise), respectively. Represents noisy speech features (such as noisy Hubert representation), i.e. .

[0090] Specifically, the loss function for training the semantic-to-acoustic feature sub-model is shown in the following formula (5): (5) Where E represents the mathematical expectation over the model prediction error over the full probability distribution; This represents a third generative model based on flow matching at time step The output represents the flow rate; and These represent the acoustic features of the sample (such as the true Mel representation) and the prior distribution (such as Gaussian noise), respectively. Represents noisy characteristics (such as noisy Mel characterization), i.e. .

[0091] The above implementation method achieves a two-stage streaming matching speech synthesis method that does not require phoneme-level duration prediction by training a text-to-semantic sub-model and a semantic-to-acoustic feature sub-model. This not only avoids the problem of poor stability caused by the difficulty of directly converting the text modality to the acoustic modality, but also improves the speech synthesis speed. Furthermore, the synthesized speech has high naturalness and low word error rate, which greatly improves the quality of the synthesized speech.

[0092] Figure 5 This is a block diagram illustrating a speech synthesis apparatus according to an exemplary embodiment. (Refer to...) Figure 5 The speech synthesis device 500 includes: The acquisition unit 510 is configured to acquire the text to be synthesized and the prompt speech; The duration prediction unit 520 is configured to perform duration prediction based on the phoneme sequence of the text to be synthesized and the prompt speech to obtain the playback duration information of the target synthesized speech; The phoneme sequence adjustment unit 530 is configured to adjust the sequence length of the phoneme sequence based on the playback duration information to obtain a target phoneme sequence; the sequence length of the target phoneme sequence matches the playback duration information. The text-to-semantic unit 540 is configured to perform feature extraction on the target phoneme sequence and generate target semantic features based on the result of the feature extraction and the prior distribution. The synthesized speech generation unit 550 is configured to generate target acoustic features based on the target semantic features and the prior distribution, and generate the target synthesized speech based on the target acoustic features.

[0093] In one exemplary embodiment, the duration prediction unit 520 includes: The text encoding unit is configured to perform text encoding processing based on the phoneme sequence of the text to be synthesized, to obtain text encoding features; The style coding unit is configured to perform the following operations: extract the acoustic features of the target object to which the prompting voice belongs based on the prompting voice, perform style coding processing on the acoustic features, and obtain the style coding features of the target object; The fusion unit is configured to perform the fusion of the style coding features of the target object with the text coding features to obtain fused coding features; The duration information generation unit is configured to generate playback duration information of the target synthesized speech based on the fused coding features and the prior distribution.

[0094] In one exemplary embodiment, the fusion unit includes: An attention processing unit is configured to perform self-attention processing on the text encoded features to obtain key text features; An affine coupling unit is configured to perform affine coupling processing on the style-encoded features of the target object to obtain scaling and translation coefficients. The instance normalization unit is configured to perform instance normalization processing on the key text features to obtain instance normalized text features. The linear adjustment unit is configured to perform linear adjustment on the instance normalized text features based on the scaling and translation coefficients to obtain the fused encoded features.

[0095] In one exemplary embodiment, the duration information generation unit includes: The first determining unit is configured to determine the transcript sequence of the transcribed text corresponding to the prompt speech. The second determining unit is configured to determine the ratio between the sequence length of the transcribing element sequence and the speech duration of the prompt speech; The splicing unit is configured to perform the splicing of the fused encoded features with the ratio; The duration information generation subunit is configured to generate playback duration information of the target synthesized speech based on the splicing result and the prior distribution.

[0096] In an exemplary embodiment, the text-to-semantic unit 540, when generating target semantic features based on the feature extraction results and prior distribution, is specifically configured to perform: based on the feature extraction results, using a first distribution mapping relationship to map the sampling results sampled from the prior distribution to obtain the target semantic features; the first distribution mapping relationship indicates the mapping relationship between the prior distribution and the speech feature distribution.

[0097] In an exemplary embodiment, the synthesized speech generation unit 550, when generating target acoustic features based on the target semantic features and the prior distribution, is specifically configured to perform the following: extracting the voiceprint features of the target object to which the prompting speech belongs based on the prompting speech; and mapping the sampling results sampled from the prior distribution using a second distribution mapping relationship based on the target semantic features and the voiceprint features to obtain the target acoustic features; the second distribution mapping relationship indicates the mapping relationship between the prior distribution and the acoustic feature distribution.

[0098] In an exemplary embodiment, the apparatus further includes a first training unit configured to perform: acquiring a first sample speech and a first sample transcribed text corresponding to the first sample speech; determining a first sample phoneme sequence of the first sample transcribed text and a first sample acoustic feature of the first sample speech; inputting the first sample phoneme sequence into the text encoding network of the duration prediction sub-model to obtain sample text encoding features; inputting the first sample acoustic features into the style encoding network of the duration prediction sub-model to obtain sample style encoding features of the sample object to which the first sample speech belongs; inputting the sample text encoding features and the sample style encoding features into a fusion network for fusion to obtain sample fusion features; inputting the sample fusion features as conditions into a first generative model based on stream matching for speech duration generation processing, and training the duration prediction sub-model based on the result of the speech duration generation processing to obtain a trained duration prediction sub-model.

[0099] In one exemplary embodiment, the apparatus further includes a second training unit configured to perform: acquiring a second sample speech and a second sample transcribed text corresponding to the second sample speech; determining sample speech features corresponding to the second sample speech and a second sample phoneme sequence corresponding to the second sample text; partially masking the sample speech features to obtain sample masked speech features; inputting the second sample speech and the second sample phoneme sequence into the trained duration prediction submodel to predict speech duration, obtaining a predicted sample playback duration; adjusting the sequence length of the second sample phoneme sequence based on the sample playback duration to obtain a target sample phoneme sequence; and inputting the target sample phoneme sequence into the feature extraction network of the text-to-semantic submodel for further processing. Sample features are extracted to obtain sample phoneme features; the sample phoneme features and the sample mask speech features are used as conditions to input into the second generative model based on flow matching of the text-to-semantic sub-model for semantic feature generation processing, and the text-to-semantic sub-model is trained based on the result of the semantic feature generation processing to obtain a trained text-to-semantic sub-model; sample acoustic features of the second sample speech are obtained, and the sample acoustic features are partially masked to obtain sample mask acoustic features; the sample speech features and the sample mask acoustic features are used as conditions to input into the third generative model based on flow matching for acoustic feature generation processing, and the third generative model is trained based on the result of the acoustic feature generation processing to obtain a trained semantic-to-acoustic feature sub-model.

[0100] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0101] In one exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the steps of any of the speech synthesis methods provided in the above embodiments.

[0102] The electronic device can be a terminal, a server, or a similar computing device; for example, it can run on a server. Figure 6 This is a hardware structure block diagram of a server running a speech synthesis method provided in an embodiment of the present invention, such as... Figure 6 As shown, the server 600 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 610 (CPUs 610 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 630 for storing data, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 623 or data 622. The memory 630 and storage media 620 may be temporary or persistent storage. The program stored in the storage media 620 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 610 may be configured to communicate with the storage media 620 and execute the series of instruction operations stored in the storage media 620 on the server 600. Server 600 may also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input / output interfaces 640, and / or one or more operating systems 621, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0103] The input / output interface 640 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 600. In one example, input / output interface 640 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, input / output interface 640 may be a radio frequency (RF) module used for wireless communication with the Internet.

[0104] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 600 may also include... Figure 6 The more or fewer components shown, or having the same Figure 6 The different configurations shown.

[0105] In one exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 630 including instructions, which can be executed by a processor 610 of the device 600 to perform the speech synthesis method of the present disclosure. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0106] In one exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the speech synthesis method provided in the embodiments of this disclosure.

[0107] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0108] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.< / pad>

Claims

1. A speech synthesis method, characterized in that, include: Obtain the text to be synthesized and the prompt speech; Based on the phoneme sequence of the text to be synthesized and the prompt speech, the playback duration information of the target synthesized speech is obtained by predicting the duration. The sequence length of the phoneme sequence is adjusted based on the playback duration information to obtain the target phoneme sequence; the sequence length of the target phoneme sequence is matched with the playback duration information. Feature extraction is performed on the target phoneme sequence, and target semantic features are generated based on the feature extraction results and prior distribution; Target acoustic features are generated based on the target semantic features and the prior distribution, and target synthesized speech is generated based on the target acoustic features.

2. The speech synthesis method according to claim 1, characterized in that, The step of predicting the playback duration of the target synthesized speech based on the phoneme sequence of the text to be synthesized and the prompt speech includes: Text encoding features are obtained by performing text encoding processing based on the phoneme sequence of the text to be synthesized; Based on the prompting voice, the acoustic features of the target object to which the prompting voice belongs are extracted, and the acoustic features are style-coded to obtain the style-coded features of the target object; The style encoding features of the target object are fused with the text encoding features to obtain fused encoding features; Based on the fused coding features and the prior distribution, the playback duration information of the target synthesized speech is generated.

3. The speech synthesis method according to claim 2, characterized in that, The step of fusing the style encoding features of the target object with the text encoding features to obtain the fused encoding features includes: Self-attention processing is performed on the text encoding features to obtain key text features; The style encoding features of the target object are subjected to affine coupling processing to obtain scaling and translation coefficients; The key text features are subjected to instance normalization to obtain instance normalized text features; The normalized text features of the instance are linearly adjusted based on the scaling and translation coefficients to obtain the fused encoding features.

4. The speech synthesis method according to claim 2, characterized in that, The step of generating playback duration information for the target synthesized speech based on the fused coding features and the prior distribution includes: Determine the transcript sequence of the text corresponding to the prompt speech; Determine the ratio between the sequence length of the transcribed audio sequence and the duration of the prompt speech; The fused encoded features are concatenated with the ratio; Based on the splicing result and the prior distribution, the playback duration information of the target synthesized speech is generated.

5. The speech synthesis method according to any one of claims 1 to 4, characterized in that, The generation of target semantic features based on the feature extraction results and prior distribution includes: Based on the result of the feature extraction, the sampling result sampled from the prior distribution is mapped using the first distribution mapping relationship to obtain the target semantic features; the first distribution mapping relationship indicates the mapping relationship between the prior distribution and the speech feature distribution.

6. The speech synthesis method according to claim 5, characterized in that, The generation of target acoustic features based on the target semantic features and the prior distribution includes: Based on the prompting voice, extract the voiceprint features of the target object to which the prompting voice belongs; Based on the target semantic features and the voiceprint features, the sampling results sampled from the prior distribution are mapped using a second distribution mapping relationship to obtain the target acoustic features; the second distribution mapping relationship indicates the mapping relationship between the prior distribution and the acoustic feature distribution.

7. The speech synthesis method according to any one of claims 1 to 6, characterized in that, The speech synthesis method is based on a speech synthesis model, which includes a duration prediction sub-model. The method further includes training the duration prediction sub-model, wherein training the duration prediction sub-model includes: Obtain the first sample speech and the first sample transcribed text corresponding to the first sample speech; Determine the first sample phoneme sequence of the transcribed text of the first sample, and the first sample acoustic features of the speech of the first sample; The first sample phoneme sequence is input into the text encoding network of the duration prediction sub-model to obtain the sample text encoding features; the first sample acoustic features are input into the style encoding network of the duration prediction sub-model to obtain the sample style encoding features of the sample object to which the first sample speech belongs. The sample text encoding features and the sample style encoding features are input into a fusion network for fusion to obtain sample fusion features; The sample fusion features are used as conditional inputs to the first generative model based on stream matching for speech duration generation processing. The duration prediction sub-model is then trained based on the results of the speech duration generation processing to obtain the trained duration prediction sub-model.

8. The speech synthesis method according to claim 7, characterized in that, The speech synthesis model further includes a text-to-semantic sub-model and a semantic-to-acoustic feature sub-model, and the method further includes: Obtain the second sample speech and the corresponding second sample transcribed text; Determine the sample speech features corresponding to the second sample speech and the second sample phoneme sequence corresponding to the second sample text; perform partial masking on the sample speech features to obtain sample masked speech features; The second sample speech and the second sample phoneme sequence are input into the trained duration prediction sub-model to predict the speech duration and obtain the predicted sample playback duration; the sequence length of the second sample phoneme sequence is adjusted based on the sample playback duration to obtain the target sample phoneme sequence; The target sample phoneme sequence is input into the feature extraction network of the text-to-semantic sub-model to extract sample features and obtain sample phoneme features; the sample phoneme features and the sample mask speech features are used as conditions to input into the second generative model of the text-to-semantic sub-model based on stream matching to generate semantic features, and the text-to-semantic sub-model is trained based on the result of the semantic feature generation process to obtain the trained text-to-semantic sub-model; Acquire the acoustic features of the second sample speech, partially mask the acoustic features to obtain the masked acoustic features; input the sample speech features and the masked acoustic features as conditions into a third generative model based on stream matching for acoustic feature generation processing, and train the third generative model based on the results of the acoustic feature generation processing to obtain a trained semantic-to-acoustic feature sub-model.

9. A speech synthesis device, characterized in that, include: The acquisition unit is configured to acquire the text to be synthesized and the prompt speech; The duration prediction unit is configured to perform duration prediction based on the phoneme sequence of the text to be synthesized and the prompt speech to obtain the playback duration information of the target synthesized speech; A phoneme sequence adjustment unit is configured to adjust the sequence length of the phoneme sequence based on the playback duration information to obtain a target phoneme sequence; the sequence length of the target phoneme sequence matches the playback duration information. The text-to-semantic unit is configured to perform feature extraction on the target phoneme sequence and generate target semantic features based on the result of the feature extraction and the prior distribution. The synthesized speech generation unit is configured to generate target acoustic features based on the target semantic features and the prior distribution, and generate the target synthesized speech based on the target acoustic features.

10. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the speech synthesis method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the speech synthesis method as described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Parallelized Tacotron: non-autoregressive and controllable TTS

    CN116457870A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN117711371A

  • Speech synthesis method and device, equipment and storage medium

    CN119811363A

  • Voice generation model construction method and device, electronic equipment and readable medium

    CN119920230A