Speech synthesis methods, devices, electronic equipment and storage media
By using a normalized exponential function and a length prediction model in the speech synthesis process, combined with an upsampling learning model, the complexity and error problems caused by dependence on external signals in existing technologies are solved, and efficient unsupervised speech synthesis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-10
- Publication Date
- 2026-03-06
AI Technical Summary
Existing speech synthesis techniques rely on externally provided supervised duration signals, which increases the complexity of the training process and reduces the synthesis effect, and the errors cannot be eliminated.
By acquiring the data to be processed, the duration probability is calculated using a normalized exponential function. Then, by combining the length prediction model and the upsampling learning model, the duration of each phoneme is determined, thus achieving unsupervised speech synthesis.
This reduces training complexity, decreases reliance on external models, and ensures the stability and accuracy of the synthesized results.
Smart Images

Figure CN115547294B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a speech synthesis method, apparatus, electronic device and storage medium. Background Technology
[0002] With the development of multimedia communication technology and artificial intelligence, speech synthesis and speech recognition technologies have become key technologies for human-computer voice communication. In some application scenarios, due to specific application requirements such as confidentiality and personalization, it is necessary to use audio processing technology to synthesize target speech from user-input text or speech.
[0003] In related technologies, the aforementioned speech synthesis technology can be implemented through machine learning models. However, the models that implement speech synthesis often rely on externally provided supervised duration signals. This requirement increases the complexity of the training process and makes the model sensitive to the performance of the model that provides supervised duration signals. It also introduces more errors and reduces the synthesis effect. Summary of the Invention
[0004] In view of this, this application proposes a speech synthesis method, apparatus, electronic device and storage medium, thereby providing an unsupervised speech synthesis model with performance no less than that of current supervised speech synthesis models, thereby reducing training complexity and dependence on external models, and ultimately ensuring synthesis effect.
[0005] To achieve the above objectives, this application provides a speech synthesis method, comprising:
[0006] Acquire data to be processed, wherein the data to be processed contains at least one phoneme information of the target text and the target timbre;
[0007] The normalized exponential function is calculated on the data to be processed to obtain the duration probability corresponding to each phoneme information. The duration probability is multiplied by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme information. The synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model.
[0008] The duration and the data to be processed are input into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0009] Based on the same technical concept, this application also provides a speech synthesis device, including:
[0010] The acquisition module is used to acquire data to be processed, wherein the data to be processed includes at least one phoneme information of the target text and the target timbre;
[0011] The calculation module is used to perform a normalized exponential function calculation on the data to be processed to obtain the duration probability corresponding to each phoneme information, and multiply the duration probability by the synthesis length obtained through the data to be processed to obtain the duration corresponding to each phoneme information; the synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model;
[0012] The processing module is used to input the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0013] Based on the same concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the above.
[0014] Based on the same concept, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the method described in any of the above-mentioned methods.
[0015] As can be seen from the above, this application provides a speech synthesis method, apparatus, electronic device, and storage medium, comprising: acquiring data to be processed; calculating the duration probability corresponding to each phoneme by performing a normalized exponential function on the data to be processed; multiplying the duration probability by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme; and inputting the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre. This application, by determining the duration probability through a normalized exponential function when using a duration model for duration prediction, and combining it with the synthesis length obtained from a length prediction model, can determine the duration of each corresponding phoneme, thereby synthesizing the target audio. This unsupervised approach reduces training complexity and dependence on external models, and ultimately ensures the synthesis effect. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A schematic flowchart of a speech synthesis method provided in an embodiment of this application;
[0018] Figure 2A flowchart illustrating the actual model application of a speech synthesis method provided in this application embodiment;
[0019] Figure 3 A flowchart illustrating the training method for the length prediction model provided in this application embodiment;
[0020] Figure 4 This is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of this application;
[0021] Figure 5 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this application's technical solution, based on the prompt message.
[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.
[0026] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0027] To make the objectives, technical solutions, and advantages of this specification clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0028] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element, object, or method step preceding the term covers the element, object, or method step listed after the term and its equivalents, without excluding other elements, objects, or method steps. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0030] The principles and spirit of this application will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement this application, and are not intended to limit the scope of this application in any way. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0031] According to embodiments of this application, a speech synthesis method, apparatus, electronic device, and storage medium are proposed.
[0032] In this article, it is important to understand that any number of elements in the accompanying figures is for illustrative purposes and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0033] The principles and spirit of this application will be explained in detail below with reference to several representative embodiments.
[0034] In related technologies, speech synthesis is generally implemented using TTS (Text-to-Speech Synthesis) models. Current TTS models rely on an external aligner that provides supervised duration signals. This requirement increases the complexity of the TTS model training process and makes the model sensitive to the performance of the aligner. Although these aligners can usually provide reasonable alignment, they are not necessarily the optimal form of decoder because they cannot be jointly trained with the TTS model. In specific applications, such as the Parallel Tacotron model and related TTS models, a length regulator upsamples the encoder output based on the duration corresponding to each factor in the data to be processed, which requires rounding the duration. This rounding introduces two problems. First, it injects a rounding error. Although a simple rounding algorithm can minimize the rounding error in related technologies, this error still exists and needs to be handled by external means such as networks. Second, the rounding operation used by the Parallel Tacotron model is non-differentiable, so the error gradient does not propagate through the operation. In summary, current speech synthesis techniques that rely on externally provided supervised duration signals not only increase the complexity of the overall training process, but also make the model sensitive to the performance of the model that provides the supervised duration signal because the error cannot be eliminated, thus reducing the synthesis effect.
[0035] In light of the above practical considerations, this application proposes a speech synthesis scheme, comprising: acquiring data to be processed; calculating the duration probability corresponding to each phoneme by performing a normalized exponential function on the data to be processed; multiplying the duration probability by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme; and inputting the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre. This application determines the duration probability by using a normalized exponential function when using a duration model for duration prediction, and combines this with the synthesis length obtained from a length prediction model, thereby determining the duration of each corresponding phoneme and synthesizing the target audio. This unsupervised approach reduces training complexity and dependence on external models, ultimately ensuring the synthesis effect.
[0036] First, this application provides a speech synthesis method. This speech synthesis method is applied to a terminal device, which includes, but is not limited to, desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of performing the aforementioned functions. In this application embodiment, the terminal device also deploys a pre-trained length prediction model and a learned upsampling model.
[0037] refer to Figure 1 The speech synthesis method of this embodiment may include the following steps:
[0038] Step 101: Obtain the data to be processed, which includes at least one phoneme information of the target text and the target timbre.
[0039] In this embodiment, the data to be processed is a collection of data generated by combining relevant speech synthesis materials. It generally includes the target text or target speech corresponding to the speech to be generated, such as broadcasting a piece of text (target text) in speech form or broadcasting a piece of speech in another speech form. For these target texts, each phoneme information corresponding to each character can be determined. For example, for an English text, the phonetic symbol corresponding to each word can be used as phoneme information; or for a Chinese text, the pinyin corresponding to each Chinese character can be used as phoneme information. The data to be processed generally also needs to include the target timbre, which is the timbre of the final target speech corresponding to the speech synthesis. It can generally be the voice information input by a specific voice input user or recorded speech material.
[0040] In some embodiments, the two types of data are input into their respective encoders, then combined and subjected to multiple convolutions to obtain the data to be processed. In a specific embodiment, such as... Figure 2 As shown, in the relevant Parallel Tacotron model or Parallel Tacotron 2 model, the target text containing phoneme information is input to the text encoder, and the target timbre is input to the voiceprint encoder. The two are then combined and input into the convolutional block (4LConv Block) of the duration prediction model to obtain the data to be processed, V, where V = {v1, ..., v...}. K}, where K can be understood as the number of phonemes in the dataset to be processed, v k It is an M×1 column vector.
[0041] Step 102: Calculate the normalized exponential function on the data to be processed to obtain the duration probability corresponding to each phoneme information. Multiply the duration probability by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme information. The synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model.
[0042] In this embodiment, the data to be processed is first calculated using a normalized exponential function (Softmax). The result of this calculation represents the duration probability of each phoneme in the data, where the duration probability can be understood as the proportion of phoneme length in the synthesized speech. Simultaneously, the data to be processed is input into a pre-trained length prediction model to predict the total audio length of the synthesized speech, i.e., the synthesis length. Then, multiplying the duration probability by the synthesis length yields the duration corresponding to each phoneme, thereby approximating the correspondence between each phoneme (between sound and text or between sound and the timeline).
[0043] In some embodiments, such as Figure 2 As shown, the data V to be processed is mapped (projected) according to relevant techniques. Then, the duration probability P corresponding to each phoneme (an item in the data V) is calculated using Softmax. Simultaneously, the output length of the data V is predicted based on a length prediction model. Finally, the duration probability P is multiplied by the output length to obtain the duration D.
[0044] Step 103: Input the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0045] In this embodiment, the duration D corresponding to each factor information calculated in the previous step is input into a pre-trained upsampling learning model. This upsampling learning model can be similar to the upsampling learning models of related technologies such as Parallel Tacotron or Parallel Tacotron 2, or it can be appropriately adjusted based on the upsampling learning models of Parallel Tacotron or Parallel Tacotron 2 to improve the alignment effect (between sound and text or between sound and the timeline). After further processing of the data by the upsampling learning model, the final output data, after decoding, yields the speech data of the target text that conforms to the target timbre.
[0046] In addition, the generated voice data can be output. This output can be used to store, display, use, or further process the voice data. The specific output method for the voice data can be flexibly selected according to different application scenarios and implementation needs.
[0047] For example, in the application scenario where the method of this embodiment is executed on a single device, the voice data can be directly output on the display component (monitor, projector, etc.) of the current device in a display manner, so that the operator of the current device can directly see the content of the voice data on the display component.
[0048] For example, in application scenarios where the method of this embodiment is executed on a system composed of multiple devices, voice data can be sent to other preset devices within the system as receivers, i.e., synchronous terminals, via any data communication method (wired connection, NFC, Bluetooth, Wi-Fi, cellular network, etc.), so that the synchronous terminals can perform subsequent processing. Optionally, the synchronous terminal can be a preset server, which is generally located in the cloud and serves as a data processing and storage center, capable of storing and distributing voice data; wherein, the receivers of the distribution are terminal devices, and the owners or operators of these terminal devices can be providers of target timbres, operators of speech synthesis, users of speech synthesis software, etc.
[0049] For example, in the application scenario where the method of this embodiment is executed on a system composed of multiple devices, voice data can be directly sent to a preset terminal device through any data communication method. The terminal device can be one or more of the devices listed in the preceding paragraphs.
[0050] As can be seen from the above, a speech synthesis method according to an embodiment of this application includes: acquiring data to be processed; calculating the duration probability corresponding to each phoneme by performing a normalized exponential function on the data to be processed; multiplying the duration probability by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme; and inputting the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre. This application determines the duration probability by using a normalized exponential function when using a duration model for duration prediction, and combines this with the synthesis length obtained from a length prediction model, thereby determining the duration of each corresponding phoneme and synthesizing the target audio. This unsupervised approach reduces training complexity and dependence on external models, and ultimately ensures the synthesis effect.
[0051] As an optional implementation, embodiments of this application also include a method for training a length prediction model.
[0052] refer to Figure 3The training method for the length prediction model in this embodiment may include the following steps:
[0053] Step 201: Obtain training data, determine the training duration probability and training target length corresponding to the training data, and multiply the training duration probability by the training target length to obtain the training duration.
[0054] In this embodiment, to train the length prediction model, training data must first be acquired. This training data may include the data to be processed for training, the corresponding known duration probabilities (i.e., training duration probabilities), the known length information (i.e., the training target length), and so on. The product of the two is the training duration corresponding to each phoneme information in the data to be processed for training.
[0055] Step 202: Predict the frame length of each phoneme information in the training data to generate the training synthesis length corresponding to each phoneme information, and train the length prediction model by checking whether the result of multiplying the training duration probability by the training synthesis length and the training duration meet a preset condition.
[0056] In this embodiment, after determining the training duration probability and training duration, the corresponding training synthesis length can be generated by predicting the frame length of each phoneme based on a preset model. The frame length of phoneme information can be understood as the duration of different phoneme occurrences. Different phonemes (whether English vowels and consonants or Chinese finals and initials, etc.) have different pronunciation durations. Furthermore, the pronunciation duration may differ in different words. Additionally, the pronunciation habits of the target voice's speaker can also lead to different frame lengths of phoneme information. Therefore, to align phoneme information by considering these factors, it is necessary to know the duration of each phoneme in the target voice. Of course, more detailed division can also be performed based on each word.
[0057] As an optional implementation, a loss function corresponding to the training target length and the training synthesis length can be set during the training of the length prediction model to accurately constrain the training results. That is, during the training of the length prediction model, a first loss function is set to satisfy a preset threshold for the error between the training synthesis length and the training target length, so as to constrain the length prediction model through the first loss function.
[0058] In a specific embodiment, multiplying the training duration probability P by the training target length yields the corresponding training duration D. Since the sum of all training duration probabilities P is 1, it can be deduced that the sum of all training durations D is the training target length. Based on this, a first loss function can be further defined, limiting the difference between the training target length and the training synthesis length to whether it satisfies a preset threshold, i.e., Loss1 = L1(output_length, target_length).
[0059] As an optional implementation, inputting the duration and the data to be processed into a pre-trained upsampling learning model includes: calculating the sampling boundary distance of each phoneme information at a set time based on the duration; performing a one-dimensional convolution on the data to be processed; inputting the sampling boundary distance and the convolved data to be processed into a preset multilayer perceptron learning model to obtain intermediate data and an auxiliary context tensor; calculating an intermediate attention matrix by using a Gaussian noise function and a sigmoid function on the intermediate data; and calculating the upsampling result of the upsampling learning model using the intermediate attention matrix, the auxiliary context tensor, and the data to be processed. This allows for further sampling operations on the duration and the data to be processed using the upsampling learning model.
[0060] In this embodiment, as Figure 2 As shown, in the relevant Parallel Tacotron 2 model, boundary calculations can be performed according to the corresponding boundary network calculation rules. Firstly, the boundary of any element in the data V to be processed can be represented as... Therefore, the duration can be expressed as:
[0061] e k =s k +D k
[0062] Then, the boundary is mapped into two T×K network matrices S and E, thus giving the sampling boundary distance of any element at time t:
[0063] S tk =ts k E tk =e k -t
[0064] Where S tk and E tk These are the (t,k)-th elements of the network matrices S and E, respectively.
[0065] Next, a one-dimensional convolution is performed on the data to be processed to obtain Conv1D(V). The distance between the convolved data and the previously obtained sampling boundary is then input into the multilayer perceptron learning model, which can then calculate an intermediate data MLP(S,E,Conv1D(V)), where MLP() represents the multilayer perceptron learning model. Simultaneously, an AT×K×P auxiliary context tensor C = MLP(S,E,Conv1D(V)) can be obtained using the multilayer perceptron learning model. Then, for the intermediate data MLP(S,E,Conv1D(V)), Gaussian noise and sigmoid are used to replace the softmax function in the relevant Parallel Tacotron 2 model to obtain the intermediate attention matrix W, i.e., W = Gaussian noise(MLP(S,E,Conv1D(V))) + sigmoid(MLP(S,E,Conv1D(V))).
[0066] Finally, following the calculation rules of subsequent steps in the relevant Parallel Tacotron 2 model, the intermediate attention matrix W, auxiliary context tensor C, and the data to be processed V are used to calculate the final upsampling result O. The upsampling result O is then decoded to obtain the final speech data result.
[0067] As an optional implementation, a loss function constraining the intermediate attention matrix can be set during the training of the upsampling learning model to accurately constrain the training results, improve the speech stability of the output, and prevent the output speech from fluctuating in pitch. Specifically, during training, the upsampling learning model sets a second loss function where the sum of the second dimension of the intermediate attention matrix is 1, thereby constraining the upsampling learning model; and during training, the upsampling learning model sets a third loss function where the sum of the first dimension of the intermediate attention matrix is the duration, thereby constraining the upsampling learning model.
[0068] In this embodiment, the second loss function is Loss2 = L1(sum(W,2),1), and the third loss function is Loss3 = L1(sum(W,1),D). The second loss function can bring the values of the intermediate attention matrix W towards 0 or 1, thus approximating a Bernoulli distribution for the intermediate attention matrix W. The third loss function ensures that the intermediate attention matrix W is strictly upsampled according to the duration D, providing approximate monotonicity.
[0069] As an optional implementation, the step of obtaining the data to be processed includes: encoding the phoneme information and the target timbre; and performing multiple convolutions on the encoded phoneme information and the target timbre to obtain the data to be processed.
[0070] It should be noted that the method in this application embodiment can be executed by a single device, such as a computer or server. The method in this application embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this application embodiment, and the multiple devices will interact with each other to complete the method described.
[0071] It should be noted that the above description describes specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0072] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a speech synthesis device.
[0073] refer to Figure 4 The speech synthesis device includes:
[0074] The acquisition module 410 is used to acquire data to be processed, the data to be processed including at least one phoneme information of the target text and the target timbre.
[0075] The calculation module 420 is used to perform a normalized exponential function calculation on the data to be processed to obtain the duration probability corresponding to each phoneme information, and multiply the duration probability by the synthesis length obtained through the data to be processed to obtain the duration corresponding to each phoneme information; the synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model.
[0076] The processing module 430 is used to input the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0077] In some optional embodiments, the computing module 420 is further configured to:
[0078] Acquire training data, determine the training duration probability and training target length corresponding to the training data, and multiply the training duration probability by the training target length to obtain the training duration.
[0079] The training synthesis length corresponding to each phoneme information is generated by predicting the frame length of each phoneme information in the training data, and the length prediction model is trained by checking whether the result of multiplying the training duration probability by the training synthesis length and the training duration meet a preset condition.
[0080] In some optional embodiments, during training, the length prediction model is configured with a first loss function that satisfies a preset threshold error between the training synthesized length and the training target length, so as to constrain the length prediction model through the first loss function.
[0081] In some optional embodiments, the processing module 430 is further configured to:
[0082] Calculate the sampling boundary distance of each phoneme information at a set time based on the duration;
[0083] One-dimensional convolution is performed on the data to be processed, and the sampling boundary distance and the convolved data to be processed are input into a preset multilayer perceptron learning model to obtain intermediate data and auxiliary context tensor.
[0084] The intermediate attention matrix is obtained by calculating the intermediate data using a Gaussian noise function and a sigmoid function.
[0085] The upsampling result of the upsampling learning model is calculated using the intermediate attention matrix, the auxiliary context tensor, and the data to be processed.
[0086] In some optional embodiments, during training, the upsampling learning model is configured with a second loss function whose dimension 2 of the intermediate attention matrix is summed to 1, so as to constrain the upsampling learning model through the second loss function.
[0087] In some optional embodiments, during training, the upsampling learning model sets the 1-dimensional sum of the intermediate attention matrix as a third loss function of the duration to constrain the upsampling learning model through the third loss function.
[0088] In some optional embodiments, the acquisition module 410 is further configured to:
[0089] The phoneme information and the target timbre are encoded;
[0090] The encoded phoneme information and the target timbre are convolved multiple times to obtain the data to be processed.
[0091] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.
[0092] The apparatus described above is used to implement the corresponding speech synthesis methods in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0093] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech synthesis method as described in any of the above embodiments.
[0094] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0095] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0096] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0097] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0098] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0099] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0100] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0101] The electronic devices described above are used to implement the corresponding speech synthesis methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0102] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the speech synthesis method as described in any of the above embodiments.
[0103] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0104] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the speech synthesis method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0105] According to one or more embodiments of this disclosure, Example 1 provides a speech synthesis method, including:
[0106] Acquire data to be processed, wherein the data to be processed contains at least one phoneme information of the target text and the target timbre;
[0107] The normalized exponential function is calculated on the data to be processed to obtain the duration probability corresponding to each phoneme information. The duration probability is multiplied by the synthesis length obtained from the data to be processed to obtain the duration corresponding to each phoneme information. The synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model.
[0108] The duration and the data to be processed are input into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0109] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, the method further comprising training the length prediction model by:
[0110] Acquire training data, determine the training duration probability and training target length corresponding to the training data, and multiply the training duration probability by the training target length to obtain the training duration.
[0111] The training synthesis length corresponding to each phoneme information is generated by predicting the frame length of each phoneme information in the training data, and the length prediction model is trained by checking whether the result of multiplying the training duration probability by the training synthesis length and the training duration meet a preset condition.
[0112] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein during training, the length prediction model sets a first loss function that satisfies a preset threshold error between the training synthesized length and the training target length, so as to constrain the length prediction model through the first loss function.
[0113] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein inputting the duration and the data to be processed into a pre-trained upsampling learning model includes:
[0114] Calculate the sampling boundary distance of each phoneme information at a set time based on the duration;
[0115] One-dimensional convolution is performed on the data to be processed, and the sampling boundary distance and the convolved data to be processed are input into a preset multilayer perceptron learning model to obtain intermediate data and auxiliary context tensor.
[0116] The intermediate attention matrix is obtained by calculating the intermediate data using a Gaussian noise function and a sigmoid function.
[0117] The upsampling result of the upsampling learning model is calculated using the intermediate attention matrix, the auxiliary context tensor, and the data to be processed.
[0118] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein during training, the upsampling learning model sets a second loss function whose dimension 2 of the intermediate attention matrix is summed to 1, so as to constrain the upsampling learning model by means of the second loss function.
[0119] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 4, wherein during training, the upsampling learning model sets the 1-dimensional sum of the intermediate attention matrix to a third loss function of the duration, so as to constrain the upsampling learning model by the third loss function.
[0120] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein obtaining the data to be processed includes:
[0121] The phoneme information and the target timbre are encoded;
[0122] The encoded phoneme information and the target timbre are convolved multiple times to obtain the data to be processed.
[0123] According to one or more embodiments of this disclosure, Example 8 provides a speech synthesis apparatus, including:
[0124] The acquisition module is used to acquire data to be processed, wherein the data to be processed includes at least one phoneme information of the target text and the target timbre;
[0125] The calculation module is used to perform a normalized exponential function calculation on the data to be processed to obtain the duration probability corresponding to each phoneme information, and multiply the duration probability by the synthesis length obtained through the data to be processed to obtain the duration corresponding to each phoneme information; the synthesis length is obtained by inputting the data to be processed into a pre-trained length prediction model;
[0126] The processing module is used to input the duration and the data to be processed into a pre-trained upsampling learning model to obtain speech data of the target text that conforms to the target timbre.
[0127] According to one or more embodiments of this disclosure, Example 9 provides the apparatus of Example 8, the computing module, which is further configured to:
[0128] Acquire training data, determine the training duration probability and training target length corresponding to the training data, and multiply the training duration probability by the training target length to obtain the training duration.
[0129] The training synthesis length corresponding to each phoneme information is generated by predicting the frame length of each phoneme information in the training data, and the length prediction model is trained by checking whether the result of multiplying the training duration probability by the training synthesis length and the training duration meet a preset condition.
[0130] According to one or more embodiments of this disclosure, Example 10 provides the apparatus of Example 9, wherein during training, the length prediction model sets a first loss function to satisfy a preset threshold for the error between the training synthesized length and the training target length, so as to constrain the length prediction model through the first loss function.
[0131] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 8, wherein the processing module is further configured to:
[0132] Calculate the sampling boundary distance of each phoneme information at a set time based on the duration;
[0133] One-dimensional convolution is performed on the data to be processed, and the sampling boundary distance and the convolved data to be processed are input into a preset multilayer perceptron learning model to obtain intermediate data and auxiliary context tensor.
[0134] The intermediate attention matrix is obtained by calculating the intermediate data using a Gaussian noise function and a sigmoid function.
[0135] The upsampling result of the upsampling learning model is calculated using the intermediate attention matrix, the auxiliary context tensor, and the data to be processed.
[0136] According to one or more embodiments of this disclosure, Example 12 provides an apparatus of Example 11, wherein during training, the upsampling learning model sets a second loss function whose dimension 2 of the intermediate attention matrix sums to 1, so as to constrain the upsampling learning model by means of the second loss function.
[0137] According to one or more embodiments of this disclosure, Example 13 provides an apparatus of Example 11, wherein during training, the upsampling learning model sets the dimension 1 summation of the intermediate attention matrix to a third loss function of the duration to constrain the upsampling learning model by the third loss function.
[0138] According to one or more embodiments of this disclosure, Example 14 provides the apparatus of Example 8, wherein the acquisition module is further configured to:
[0139] The phoneme information and the target timbre are encoded;
[0140] The encoded phoneme information and the target timbre are convolved multiple times to obtain the data to be processed.
[0141] According to one or more embodiments of the present disclosure, Example 15 provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of Examples 1 to 7.
[0142] According to one or more embodiments of the present disclosure, Example 16 provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method as described in any one of Examples 1 to 7.
[0143] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0144] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0145] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0146] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A speech synthesis method characterized by, The method comprises the following steps: obtaining to-be-processed data, wherein the to-be-processed data comprises at least one phoneme information of target text and a target voice tone; performing normalized exponential function calculation on the to-be-processed data to obtain a duration probability corresponding to each phoneme information, multiplying the duration probability by a synthesis length obtained through the to-be-processed data to obtain a duration corresponding to each phoneme information; the synthesis length is obtained by inputting the to-be-processed data into a pre-trained length prediction model; inputting the duration and the to-be-processed data into a pre-trained up-sampling learning model to obtain voice data of the target text conforming to the target voice tone; the step of inputting the duration and the to-be-processed data into the pre-trained up-sampling learning model comprises the following steps: calculating a sampling boundary distance of each phoneme information at a set time according to the duration; performing one-dimensional convolution on the to-be-processed data, inputting the sampling boundary distance and the to-be-processed data after convolution into a pre-set multi-layer perception learning model to obtain intermediate data and an auxiliary context tensor; performing calculation on the intermediate data through a Gaussian noise function and an S-shaped function to obtain an intermediate attention matrix; calculating the up-sampling result of the up-sampling learning model through the intermediate attention matrix, the auxiliary context tensor and the to-be-processed data.
2. The method of claim 1, wherein, The method further comprises training the length prediction model through the following method: obtaining training data, determining a training duration probability corresponding to the training data and a training target length, multiplying the training duration probability by the training target length to obtain a training duration; predicting the frame length of each phoneme information of the training data to generate a training synthesis length corresponding to each phoneme information, and training the length prediction model according to whether a preset condition is met between the result of multiplying the training duration probability by the training synthesis length and the training duration.
3. The method of claim 2, wherein, When training the length prediction model, a first loss function is set for an error between the training synthesis length and the training target length to meet a preset threshold, so as to constrain the length prediction model through the first loss function.
4. The method of claim 1, wherein, When training the up-sampling learning model, a second loss function is set for a dimension 2 cumulative sum of the intermediate attention matrix to be 1, so as to constrain the up-sampling learning model through the second loss function.
5. The method of claim 1, wherein, When training the up-sampling learning model, a third loss function is set for a dimension 1 cumulative sum of the intermediate attention matrix to be the duration, so as to constrain the up-sampling learning model through the third loss function.
6. The method of claim 1, wherein, The step of obtaining to-be-processed data comprises the following steps: encoding the phoneme information and the target voice tone; performing multiple convolutions on the encoded phoneme information and the target voice tone to obtain the to-be-processed data.
7. A speech synthesis apparatus characterized by comprising: The method comprises the following steps: an obtaining module is configured to obtain to-be-processed data, wherein the to-be-processed data comprises at least one phoneme information of target text and a target voice tone; The computing module is configured to perform a normalized exponential function calculation on the to-be-processed data to obtain a duration probability corresponding to each phoneme information, multiply the duration probability with a synthesized length obtained by the to-be-processed data, and obtain a duration corresponding to each phoneme information; the synthesized length is obtained by inputting the to-be-processed data into a pre-trained length prediction model; The processing module is configured to input the duration and the to-be-processed data into a pre-trained up-sampling learning model to obtain voice data of the target text in the target voice style; The inputting of the duration and the to-be-processed data into the pre-trained up-sampling learning model comprises: calculating a sampling boundary distance of each phoneme information at a set time according to the duration; performing one-dimensional convolution on the to-be-processed data, inputting the sampling boundary distance and the to-be-processed data after convolution into a pre-set multi-layer perception learning model to obtain intermediate data and an auxiliary context tensor; performing calculation on the intermediate data by using a Gaussian noise function and an S-shaped function to obtain an intermediate attention matrix; calculating the up-sampling result of the up-sampling learning model by using the intermediate attention matrix, the auxiliary context tensor and the to-be-processed data.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method according to any one of claims 1 to 6 when executing the program.
9. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium stores computer instructions for causing the computer to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice synthesis method and device
CN108806665A
Speech synthesis method and device, storage medium and electronic equipment
CN111653266A