Speech synthesis generation method, electronic device, and storage medium
By adding and removing noise from the initial speech data, a combination of target speech features and prosodic features is generated, which solves the problem of insufficient correlation between prosodic features and speech representation in the existing technology and improves the naturalness and expressiveness of speech synthesis.
Patent Information
- Application Number
- CN202411904686.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-12-23
AI Technical Summary
Existing speech synthesis technologies fail to fully learn the interrelationship between prosodic features and speech representation when predicting prosodic features, resulting in limitations in the naturalness and expressiveness of speech synthesis results.
By acquiring the speech features and prosodic features of the initial speech data, splicing them together and adding noise, and then using a diffusion model for denoising, a combination of the target speech features and prosodic features is generated, thereby improving the naturalness and expressiveness of the generated speech.
By collaboratively constructing linguistic and prosodic features, the naturalness and expressiveness of speech generation are significantly improved.
Smart Images

Figure CN119864006B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments disclosed in the present application relate to the technical field of artificial intelligence, and more particularly to a speech synthesis generation method, an electronic device and a storage medium. BACKGROUND
[0002] In recent years, speech synthesis technology has developed significantly, and high expressiveness and high naturalness have increasingly become the focus of speech synthesis research. In speech synthesis, prosodic features are the key to the expressiveness and naturalness of speech. At present, various prosodic features are usually predicted based on text information, and then acoustic features are reconstructed using prosodic features and text information. However, there is still a certain difference between the distribution of the predicted prosodic features and the real prosodic features, which affects the effect of speech synthesis. SUMMARY
[0003] According to the embodiments of the present application, a speech synthesis generation method, an electronic device and a storage medium are proposed to solve the above problems.
[0004] A first aspect of the present application discloses a speech synthesis generation method, comprising: obtaining initial speech features and initial prosodic features corresponding to initial speech data; concatenating the initial speech features and the initial prosodic features to obtain an initial noise-added object; adding noise to the initial noise-added object to obtain a noise-added object; inputting the noise-added object and a phoneme sequence corresponding to the initial speech data into a diffusion model to denoise the noise-added object and obtain a target object, wherein the target object comprises a combination of target speech features and target prosodic features; and obtaining target speech data corresponding to the target object.
[0005] In some embodiments, the adding noise to the initial noise-added object to obtain a noise-added object comprises: inputting the initial noise-added object into a noise scheduler to obtain the noise-added object, wherein the noise intensity of the initial noise-added object at each noise-added time step is different and gradually increases to 1, so that the noise-added object is added with a specific proportion of noise.
[0006] In some embodiments, the diffusion model comprises a preprocessing network and a denoising network; and inputting the noise-added object and the phoneme sequence corresponding to the initial speech data into the diffusion model to denoise the noise-added object and obtain a target object comprises: inputting the noise-added object and the phoneme sequence corresponding to the initial speech data into the preprocessing network to obtain a feature time sequence, wherein the feature time sequence comprises a combination of text features corresponding to the phoneme sequence and noise-added features corresponding to the noise-added object; and inputting the feature time sequence into the denoising network for denoising to obtain the target object.
[0007] In some embodiments, the pre-processing network comprises a phoneme encoder and a pre-processing sub-network; inputting the phoneme sequence corresponding to the noisy object and the initial speech data into the pre-processing network to obtain a feature time sequence, comprises: inputting the phoneme sequence into the phoneme encoder to obtain the text feature; inputting the noisy object into the pre-processing sub-network to obtain the noisy feature; and splicing the text feature and the noisy feature to obtain the feature time sequence.
[0008] In some embodiments, the phoneme encoder comprises an embedding layer and a Transformer layer; inputting the phoneme sequence into the phoneme encoder to obtain the text feature, comprises: encoding the identity of the phoneme sequence by using the embedding layer to obtain a phoneme embedding feature; and processing the phoneme embedding feature by using the Transformer layer to obtain the text feature.
[0009] In some embodiments, the pre-processing sub-network comprises a convolutional layer and a normalization layer; inputting the noisy object into the pre-processing sub-network to obtain the noisy feature, comprises: processing the noisy object by using the convolutional layer and the normalization layer to obtain the noisy feature, wherein the dimensions of the noisy feature and the text feature are consistent.
[0010] In some embodiments, a feature length of the initial speech feature is obtained; wherein the time length of the initial noisy object is the sum of the feature length of the initial speech feature and the feature length of the initial prosody feature.
[0011] In some embodiments, the feature length of the initial speech feature is obtained by inputting a phoneme sequence corresponding to the initial speech data into a mel-spectrum prediction model to obtain a mel-spectrum feature corresponding to the phoneme sequence, so as to determine the feature length of the initial speech feature.
[0012] In some embodiments, the prosody feature corresponding to the initial speech data comprises at least one of an energy feature, a duration feature and a pitch feature, and the prosody feature is extracted based on a phoneme sequence corresponding to the initial speech data.
[0013] The second aspect of the present application discloses an electronic device comprising a memory and a processor coupled to each other, wherein the processor is configured to execute program instructions stored in the memory to implement the speech synthesis generation method of the first aspect.
[0014] The third aspect of the present application discloses a non-volatile computer readable storage medium having program instructions stored thereon, wherein the program instructions are executed by a processor to implement the speech synthesis generation method of the first aspect.
[0015] The beneficial effects of the present application are: obtaining initial speech features and initial prosody features corresponding to the initial speech data, splicing the initial speech features and the initial prosody features to obtain an initial to-be-noised object, obtaining a noised object by adding noise to the initial to-be-noised object, inputting the noised object and a phoneme sequence corresponding to the initial speech data into a diffusion model to denoise the noised object to obtain a target object, wherein the target object includes a combination of target speech features and target prosody features, and further obtaining target speech data corresponding to the target object, thereby improving the naturalness and expressiveness of speech generation in the process of cooperatively constructing language features and prosody features. BRIEF DESCRIPTION OF DRAWINGS
[0016] The present application will be further described below in conjunction with the accompanying drawings and embodiments. In the drawings:
[0017] Figure 1 is a flowchart of a speech synthesis generation method according to an embodiment of the present application;
[0018] Figure 2 is a flowchart of a speech synthesis generation method according to an embodiment of the present application;
[0019] Figure 3 is a structural schematic diagram of an electronic device according to an embodiment of the present application;
[0020] Figure 4 is a structural schematic diagram of a non-volatile computer readable storage medium according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In the present application, the phrase "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various places in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment that is not mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0022] The term "and / or" in the present application is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents an "or" relationship between the associated objects. In addition, "multiple" in this paper means two or more than two. In addition, the term "at least one" in this paper means any one of multiple or any combination of at least two of multiple, for example, including at least one of A, B and C, which means including any one or more elements selected from the set consisting of A, B and C. In addition, the terms "first", "second", "third" in the present application are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.
[0023] At present, many speech synthesis methods model the prosody of speech, and usually rely on text input when predicting prosodic features of speech. The predicted prosodic features are often strongly correlated with text features, and the mutual relationship between various prosodic features and prosodic features and speech representation is not fully learned, thereby limiting the naturalness and expressiveness of speech synthesis effect.
[0024] Therefore, the present application provides a speech synthesis generation method, an electronic device and a storage medium to improve the naturalness and expressiveness of speech synthesis effect.
[0025] In order to enable those skilled in the art to better understand the technical solutions of the present application, the technical solutions of the present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0026] Please refer to Figure 1 , Figure 1 is a flowchart of the speech synthesis generation method of the embodiment of the present application. The execution subject of the method can be an electronic device with computing function, for example, microcomputer, server, and mobile device such as notebook computer and tablet computer, etc.
[0027] It should be noted that the method of the present application is not limited to the order of the flowchart shown in Figure 1 .
[0028] In some possible implementation manners, the method can be realized by the processor calling the computer readable instructions stored in the memory, as shown in Figure 1 , the method can include the following steps:
[0029] S11: obtaining initial speech features and initial prosodic features corresponding to initial speech data.
[0030] The initial speech data is processed to obtain initial speech features and initial prosody features corresponding to the initial speech data. The initial speech data can be text, and the initial speech data is processed, for example, the text is converted into a phoneme sequence y, and then the corresponding initial speech features V0 and initial prosody features R0 are obtained based on the phoneme sequence y. The phoneme refers to the basic sound unit in pronunciation, and the initial speech features V0 can be used to represent the audio signal corresponding to the phoneme sequence, and the initial prosody features R0 can be used to represent the duration, pitch and energy of each phoneme in the phoneme sequence y.
[0031] S12: The initial speech features and the initial prosody features are spliced to obtain an initial to-be-noised object.
[0032] The initial speech features and the initial prosody features are spliced, for example, the initial speech features V0 and the initial prosody features R0 are spliced in time sequence to obtain an initial to-be-noised object X0, and the initial to-be-noised object X0 includes {V0, R0}.
[0033] S13: The initial to-be-noised object is noised to obtain a noised object.
[0034] The initial to-be-noised object is noised, for example, at each time step, noise is added to the initial to-be-noised object X0 to make the initial to-be-noised object tend to be completely randomized to obtain a noised object X t , wherein the noised object X t may include {V t , R t}, V t represents the initial speech features after noising, and R t represents the initial prosody features after noising.
[0035] S14: The noised object and the phoneme sequence corresponding to the initial speech data are input into a diffusion model to denoise the noised object to obtain a target object, wherein the target object includes a combination of target speech features and target prosody features.
[0036] The noised object and the phoneme sequence corresponding to the initial speech data are input into a diffusion model, that is, the phoneme sequence corresponding to the initial speech data is used as a control condition, and the diffusion model is used to denoise the noised object X t to obtain a target object X0'. In some examples, the target object X0' includes a combination of target speech features and target prosody features, for example, based on the noised object X tThe phoneme sequence y is used to predict the target speech feature V' and the target prosody feature R' corresponding to the initial speech data by using the diffusion model. The target speech feature V' and the target prosody feature R' can be obtained by predicting the joint distribution of the speech feature and the prosody feature corresponding to the reference initial speech data. It can be understood that the obtained target object X0' includes the joint distribution information of the speech feature and the prosody feature.
[0037] S15: Obtain target speech data corresponding to the target object.
[0038] Based on the target object, target speech data containing a combination of target speech features and target prosody features is obtained, for example, speech conversion is performed on the target object x'0 to generate target audio corresponding to the obtained text. The obtained target speech data contains a combination of target speech features V' and target prosody features R', that is, the target audio is converted from the speech feature associated with the prosody feature.
[0039] In this embodiment, the initial speech feature and the initial prosody feature corresponding to the initial speech data are obtained, which are spliced to obtain an initial to-be-noised object. The to-be-noised object is denoised to obtain a target object by inputting the to-be-noised object and the phoneme sequence corresponding to the initial speech data to the diffusion model. The target object includes a combination of target speech features and target prosody features. Further, based on the target object, target speech data containing a combination of target speech features and target prosody features is obtained. In the process of cooperatively constructing language features and prosody features, the naturalness and expressiveness of speech generation are improved.
[0040] In some embodiments, the prosody feature corresponding to the initial speech data includes at least one of an energy feature, a duration feature, and a pitch feature, and the prosody feature is extracted based on the phoneme sequence corresponding to the initial speech data.
[0041] The initial prosody feature corresponding to the initial speech data is obtained, for example, the text is converted into a phoneme sequence y by using a text front-end tool, and then the corresponding initial prosody feature R0 is obtained based on the phoneme sequence y. The initial prosody feature R0 can include an energy feature e0, a duration feature d0, and a pitch feature p0.
[0042] The initial prosody feature R0 is extracted based on the initial speech data corresponding to the phoneme sequence y, for example, the open source MFA (Montreal Forced Aligner) alignment tool is used to obtain the time length corresponding to the phoneme, that is, the time length feature d0; the open source RMVPE (Robust Model for Vocal Pitch Estimation) human voice pitch model is used to extract the average pitch corresponding to each phoneme, that is, the pitch feature p0; the librosa tool is used to extract the average energy corresponding to each phoneme, that is, the energy feature e0.
[0043] In some examples, the initial speech feature corresponding to the initial speech data is obtained, for example, the phoneme sequence y is predicted by using a mel-spectrum length predictor, to obtain the mel-spectrum feature m0 corresponding to the phoneme sequence y as the corresponding initial speech feature, wherein the mel-spectrum feature m0 can include the mel-spectrum length corresponding to the phoneme sequence y.
[0044] In some embodiments, the initial to-be-noised object is noised to obtain a noised object, comprising: inputting the initial to-be-noised object into a noise scheduler to obtain the noised object, wherein the noise intensity of the initial to-be-noised object at each noising time step is different and gradually increases to 1, so that the noised object is noised by a specific proportion of noise.
[0045] Wherein, the noise scheduler (for example, noise scheduler) can be used to control the noising intensity, and define the noise intensity corresponding to each time step t. The initial to-be-noised object X0 is input into the noise scheduler, and then the noised object X t is obtained, for example, at each time step t, a specific proportion of Gaussian noise is added to X0{m0, d0, p0, e0} according to the β t value preset by the noise scheduler, to obtain the noised object X t {m t ,d t ,p t ,e t}, the noising process can be represented as follows:
[0046]
[0047] Wherein, the noise intensity of the initial to-be-noised object X0 at each noising time step t is different and gradually increases to 1, so that the noised object X t gradually tends to complete randomization.
[0048] In some embodiments, the diffusion model includes a preprocessing network and a denoising network, wherein the denoising network can be a DiT (Diffusion Transformer) network.
[0049] The noisy object and the phoneme sequence corresponding to the initial speech data are input into the diffusion model to denoise the noisy object to obtain the target object, including: inputting the noisy object and the phoneme sequence corresponding to the initial speech data into the preprocessing network to obtain a feature time sequence, wherein the feature time sequence includes a combination of a text feature corresponding to the phoneme sequence and a noisy feature corresponding to the noisy object; inputting the feature time sequence into the denoising network to perform denoising to obtain the target object.
[0050] The noisy object and the phoneme sequence corresponding to the initial speech data are input into the preprocessing network, for example, the noisy object X t and the phoneme sequence y corresponding to the initial speech data are spliced on the time sequence through the preprocessing network to obtain a feature time sequence, wherein the feature time sequence includes a combination of a text feature corresponding to the phoneme sequence y and a noisy feature corresponding to the noisy object, for example, {y, m t ,d t ,p t ,e t}. Further, the feature time sequence is input into the denoising network to perform denoising to obtain the target object, for example, the feature time sequence is input into the DiT network together with the time step to perform denoising and predict the target object x′0{m′0,d′0,p′0,e0′0}.
[0051] The training process of the diffusion model can be represented as:
[0052] p(m′0,d′0,p′0,e0′|y,m t ,d t ,p t ,e t ,t)
[0053] y represents the phoneme sequence, m t ,d t ,p t ,e t represent the noisy mel-spectrogram feature, duration feature, pitch feature, and energy feature, respectively, and m′0,d′0,p′0,e0′ represent the denoised mel-spectrogram feature, duration feature, pitch feature, and energy feature, respectively.
[0054] The corresponding loss function can be represented as:
[0055]
[0056] wherein ∈ θ represents the denoising network, and ε represents the noise added to the noisy object X t .
[0057] It can be understood that the noisy object X tRemoving noise, i.e., at each time step t, the denoising network can predict the component of the current noise and subtract it from the noisy feature to recover the clean feature. This process realizes the restoration of noisy speech to clean speech and the reconstruction of prosodic features. That is, at each time step t, not only the clean mel-spectrogram can be generated, but also the corresponding duration feature, pitch feature and energy feature are jointly generated, which improves the consistency between speech and prosodic features.
[0058] In the present embodiment, by learning the correlation between speech features and various prosodic features, i.e., the same degree of noise is added to speech and prosodic features during training (both are time step t), the clean speech and prosodic features are predicted by the denoising network, so that the joint distribution of speech and prosodic features can be learned, which improves the naturalness and expressiveness of speech generation.
[0059] In some embodiments, the preprocessing network includes a phoneme encoder and a preprocessing subnetwork. The phoneme sequence corresponding to the noisy object and the initial speech data is input into the preprocessing network to obtain a feature time sequence, including: inputting the phoneme sequence into the phoneme encoder to obtain a text feature; inputting the noisy object into the preprocessing subnetwork to obtain a noisy feature; and splicing the text feature and the noisy feature to obtain the feature time sequence.
[0060] The phoneme sequence is input into the phoneme encoder to obtain a text feature, for example, the phoneme sequence y is encoded by the phoneme encoder to obtain a corresponding text feature. The noisy object is input into the preprocessing subnetwork to obtain a noisy feature, for example, the noisy object X t is mapped to generate a noisy feature of a target dimension. Further, the text feature and the noisy feature are spliced, for example, the text feature and the noisy feature are spliced on the realization sequence to obtain a feature time sequence, which is input into the denoising network together with the time step.
[0061] In some embodiments, the phoneme encoder includes an embedding layer and a Transformer layer, for example, the phoneme encoder includes an embedding layer and a 6-layer Transformer.
[0062] The phoneme sequence is input into the phoneme encoder to obtain a text feature, including: encoding the identity of the phoneme sequence by the embedding layer to obtain a phoneme embedding feature; and processing the phoneme embedding feature by the Transformer layer to obtain the text feature.
[0063] The phoneme sequence identifier is encoded by using an embedding layer to obtain phoneme embedding features, for example, the phoneme sequence identifier (phoneme id) is encoded by using an embedding layer to obtain 512-dimensional phoneme embedding features (phoneme embedding). Further, the phoneme embedding features are processed by using a Transformer layer to obtain text features, for example, the phoneme embedding features are further processed by using a 6-layer Transformer to obtain corresponding text features.
[0064] In some embodiments, the preprocessing subnetwork includes a convolutional layer and a normalization layer, for example, the preprocessing subnetwork includes 3 one-dimensional convolutional layers and normalization layers.
[0065] The noisy object is input into the preprocessing subnetwork to obtain noisy features, including: processing the noisy object by using a convolutional layer and a normalization layer to obtain noisy features, wherein the dimensions of the noisy features and the text features are consistent.
[0066] The noisy object is input into the preprocessing subnetwork, i.e., the noisy object is processed by using a convolutional layer and a normalization layer to obtain noisy features, for example, the noisy object X t The 80-dimensional projection is to 512-dimensional, and then the 512-dimensional noisy features are obtained.
[0067] In some embodiments, the feature length of the initial speech feature is obtained; wherein the time length of the initial noisy object is the sum of the feature length of the initial speech feature and the feature length of the initial prosody feature.
[0068] The feature length of the initial speech feature is obtained, for example, the feature length of the initial speech feature is determined to be T1. Further, the time length of the initial noisy object can be determined according to the feature length of the initial speech feature, and the time length of the initial noisy object X0 is the sum of the feature length of the initial speech feature and the feature length of the initial prosody feature. The feature length of the initial prosody feature is determined based on the length T2 of the phoneme sequence y, for example, the initial prosody feature includes energy features, duration features and pitch features, and since the lengths of the duration, pitch and energy are the same as the length of the phoneme sequence, accordingly, the feature length of the initial prosody feature is 3 times the length of the phoneme sequence, i.e., 3*T2. At this time, the time sequence length of the initial noisy object X0 is T1+3*T2.
[0069] In some embodiments, the feature length of the initial speech feature is obtained, including: inputting the phoneme sequence corresponding to the initial speech data into a mel-spectrum prediction model to obtain the mel-spectrum feature corresponding to the phoneme sequence, to determine the feature length of the speech feature.
[0070] The Mel-spectrum prediction model is an Encoder-Decoder Transformer structure. The phoneme sequence corresponding to the initial speech data is input into the Mel-spectrum prediction model to obtain the Mel-spectrum features corresponding to the phoneme sequence. For example, the text encoder in the Mel-spectrum prediction model encodes the phoneme sequence y to obtain the phoneme embedding. The phoneme embedding is then decoded to obtain the Mel-spectrum length. The obtained Mel-spectrum length is used as the Mel-spectrum feature corresponding to the phoneme sequence y, thereby determining the feature length of the initial speech features. The training of the Mel-spectrum prediction model can be optimized using a loss function.
[0071] In some examples, initial speech features and initial prosodic features corresponding to the initial speech data are obtained. For example, based on the phoneme sequence y, the corresponding Mel spectrum feature m0, energy feature e0, duration feature d0, and pitch feature p0 are obtained. For instance, the Mel spectrum feature m0 and energy feature e0 have an 80-dimensional dimension, while the duration feature d0 and pitch feature p0 have a 1-dimensional dimension. The duration feature d0 and pitch feature p0 need to be copied 80 times to obtain 80-dimensional duration feature d0 and pitch feature p0. Further, by unifying the dimensions of the Mel spectrum feature m0 to 80 dimensions, and concatenating the energy feature e0, duration feature d0, and pitch feature p0 in time sequence, the initial object to be noised, X0, can be obtained. At this point, the time series length of the initial object to be noised, X0, is three times the length of the phoneme sequence plus the length of the Mel spectrum.
[0072] Understandably, in some examples, the length of the Mel spectrum is determined using a Mel spectrum length predictor via a phoneme sequence. Since the lengths of duration, pitch, and energy are the same as the length of the phoneme sequence, the object to be added with noise, X, can then be determined. t The total length. Furthermore, the noisy object X with completely Gaussian noise. t By iterating N steps in reverse diffusion (e.g., set to 1000), speech duration, pitch, and energy features can be jointly generated.
[0073] To facilitate understanding, the speech synthesis generation method of this application embodiment is illustrated with examples, such as... Figure 2 As shown, Figure 2 This is a schematic flowchart of a speech synthesis generation method according to an embodiment of this application.
[0074] For example, initial speech features and initial prosodic features corresponding to the text are obtained. The initial speech features include Mel spectrum features, and the initial prosodic features include duration features, pitch features, and energy features. The initial speech features and initial prosodic features are concatenated to obtain the initial object to be denoised X0{m0,d0,p0,e0}. The initial object to be denoised X0 is then input into a noise scheduler to obtain the object to be denoised X. t {m td t ,p t ,e t}. The phoneme sequence y is encoded by using the phoneme encoder to obtain a 512-dimensional text feature, the noisy object X t is projected by using the preprocessing subnetwork to obtain a 512-dimensional noisy feature, the text feature and the noisy feature are spliced to obtain a feature time sequence, and the feature time sequence is input into the denoising network to obtain the target object x′0{m′0,d′0,p′0,e0′}. Then, the target object x′0 can be used to obtain the target speech data corresponding to the text.
[0075] Those skilled in the art can understand that, in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.
[0076] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of an electronic device of an embodiment of the present application. The electronic device 30 includes a memory 31 and a processor 32 coupled with each other. The processor 32 is configured to execute program instructions stored in the memory 31 to implement the steps of the speech synthesis generation method embodiment described above. In a specific implementation scenario, the electronic device 30 can include but is not limited to a microcomputer, a server, without limitation.
[0077] Specifically, the processor 32 is configured to control itself and the memory 31 to implement the steps of the speech synthesis generation method embodiment described above. The processor 32 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 32 can be an integrated circuit chip with a signal processing capability. The processor 32 can also be a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), an FPGA (Field-Programmable Gate Array, field programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 32 can be implemented by an integrated circuit chip together.
[0078] Please refer to Figure 4 , Figure 4Fig. 1 is a structural schematic diagram of a non-volatile computer readable storage medium according to an embodiment of the present application. The non-volatile computer readable storage medium 40 is used to store a computer program 401, which, when executed by a processor, for example, the processor 32 described above Figure 3 The processor 32 in the embodiment executes the steps of the method for speech synthesis generation.
[0079] The above description of the various embodiments tends to emphasize the differences between the various embodiments, and the same or similar parts can be mutually referred to, and for brevity, will not be repeated here.
[0080] In several embodiments provided in the present application, it should be understood that the disclosed method and related device can be implemented in other ways. For example, the above-described device implementation is only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication between the shown or discussed mutual units can be indirect coupling or communication through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0081] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0082] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that makes a contribution to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0083] Those skilled in the art will readily recognize a variety of modifications and changes that can be made to the embodiments without departing from the true scope of the application, which is indicated by the appended claims.
Claims
1. A speech synthesis generation method characterized by, The method comprises: obtaining initial speech features and initial prosody features corresponding to initial speech data; concatenating the initial speech features and the initial prosody features to obtain an initial object to be added with noise; adding noise to the initial object to be added with noise to obtain a noise-added object; inputting the noise-added object and a phoneme sequence corresponding to the initial speech data into a diffusion model to denoise the noise-added object to obtain a target object, wherein the diffusion model comprises a preprocessing network and a denoising network, and the target object comprises a combination of target speech features and target prosody features; obtaining target speech data corresponding to the target object; the method further comprises: inputting the noise-added object and the phoneme sequence corresponding to the initial speech data into the preprocessing network to obtain a feature time sequence, wherein the feature time sequence comprises a combination of text features corresponding to the phoneme sequence and noise-added features corresponding to the noise-added object; inputting the feature time sequence into the denoising network to perform denoising to obtain the target object.
2. The method of claim 1, wherein: the method further comprises: inputting the initial object to be added with noise into a noise scheduler to obtain the noise-added object, wherein the initial object to be added with noise has different noise intensities at each noise-added time step and gradually increases to 1, so that the noise-added object is added with noise of a specific proportion.
3. The method of claim 1, wherein, the preprocessing network comprises a phoneme encoder and a preprocessing subnetwork; inputting the noise-added object and the phoneme sequence corresponding to the initial speech data into the preprocessing network to obtain a feature time sequence comprises: inputting the phoneme sequence into the phoneme encoder to obtain the text features; inputting the noise-added object into the preprocessing subnetwork to obtain the noise-added features; concatenating the text features and the noise-added features to obtain the feature time sequence.
4. The method of claim 3, wherein, the phoneme encoder comprises an embedding layer and a Transformer layer; inputting the phoneme sequence into the phoneme encoder to obtain the text features comprises: encoding the identity of the phoneme sequence using the embedding layer to obtain phoneme embedding features; processing the phoneme embedding features using the Transformer layer to obtain the text features.
5. The method of claim 3, wherein, the preprocessing subnetwork comprises a convolutional layer and a normalization layer; inputting the noise-added object into the preprocessing subnetwork to obtain the noise-added features comprises: processing the noise-added object using the convolutional layer and the normalization layer to obtain the noise-added features, wherein the noise-added features have the same dimension as the text features.
6. The method of claim 1, wherein, The method further comprises: obtaining a feature length of the initial speech features; wherein the time length of the initial object to be added with noise is the sum of the feature length of the initial speech features and the feature length of the initial prosody features.
7. The method of claim 6, wherein, the method further comprises: inputting the phoneme sequence corresponding to the initial speech data into a mel-spectrum prediction model to obtain mel-spectrum features corresponding to the phoneme sequence, so as to determine a feature length of the initial speech features.
8. The method of any one of claims 1-7, characterized in that, The prosody features corresponding to the initial speech data include at least one of energy features, duration features, and pitch features, and the prosody features are extracted based on the phoneme sequence corresponding to the initial speech data.
9. An electronic device, comprising: The device comprises a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the speech synthesis generation method of any one of claims 1-8.
10. A non-transitory computer readable storage medium having stored thereon program instructions, the program instructions comprising instructions for causing a processor to: 5 perform the method of any one of claims 1-9. The program instructions, when executed by the processor, implement the speech synthesis generation method of any one of claims 1-8.
Citation Information
Patent Citations
Speech synthesis model training method, speech synthesis method and task platform
CN118898986A