Voice Processing Method, Device, Computer-Readable Storage Medium, and Computer Equipment
The adaptive voice synthesis method addresses the inefficiencies of cloud-trained models by pruning and training device-specific models, optimizing voice synthesis for diverse devices and enhancing user experience through improved performance and resource efficiency.
Patent Information
- Application Number
- CN202111620262.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The synthesis capability of existing speech synthesis models on the device side cannot be optimal, because the cloud-trained models cannot adapt to the diversity and computing power differences on the device side, resulting in long wait times and poor synthesis effect.
On the server side, the initial speech synthesis model is cropped and trained based on the terminal's performance data, the target speech synthesis model is generated, and deployed to the device side to ensure that the model matches the performance of the terminal.
It improves the efficiency and effectiveness of voice synthesis, reduces waiting time, improves user experience, and reduces the business pressure of the server.
Smart Images

Figure CN114267322B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a voice processing method, apparatus, computer-readable storage medium, and computer device. Background Art
[0002] In recent years, with the development of artificial intelligence technology, human-computer interaction has become increasingly frequent. The interaction method using voice as the medium has gradually become the mainstream form in the field of human-computer interaction, and the importance of voice synthesis technology has become increasingly prominent. Voice synthesis technology is a signal processing technology that converts text into voice, and it can endow machines with the ability to speak like humans. Voice cloning is an application of voice synthesis technology. Voice cloning can quickly realize user voice customization and has broad application prospects in fields such as voice assistants and navigation. Voice cloning refers to the process in which a synthesis training system uses the voice of a specific user for model training, and then uses the trained voice synthesis model to implement the conversion from text to the voice of this specific user.
[0003] Since the training of voice synthesis models requires a large amount of computing power and a long training time, current voice synthesis models are all trained in the cloud with better hardware conditions. For example, users upload audio to the cloud, and the trained voice synthesis model in the cloud is downloaded to the device side. In this way, the voice synthesis ability is built into the device side, and subsequent synthesis processes are all carried out on the device side.
[0004] With the development of current deep learning technology and hardware, mainstream voice synthesis models basically use neural networks to implement. To achieve the optimal effect, such voice synthesis models usually need to be trained for several hours to dozens of hours, and the waiting time for users is too long. In addition, there are various types of devices on the device side. Therefore, using a set of voice synthesis models trained in the cloud will result in the inability to run the synthesis ability on the device side or the inability to achieve the optimal synthesis effect. Summary of the Invention
[0005] Embodiments of this application provide a voice processing method, apparatus, computer-readable storage medium, and computer device, which can obtain a target voice synthesis model that matches the performance data of the terminal, and thus provide a voice synthesis service that conforms to the performance data of the terminal based on the target voice synthesis model, improving the user experience.
[0006] Embodiments of this application provide a voice processing method, including:
[0007] Obtain a target text and a target speech synthesis model, where the target speech synthesis model is obtained by pruning a target network module in an initial speech synthesis model according to target performance data of a terminal corresponding to a target pronunciation object, and training a to-be-trained speech synthesis model obtained by the pruning process with speech data of the target pronunciation object, and the speech data has a target timbre;
[0008] Use the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre.
[0009] An embodiment of the present application further provides a speech processing method, including:
[0010] According to a speech synthesis service request from a terminal, determine the target performance data of the terminal and the speech data of a target pronunciation object, and the speech data has a target timbre;
[0011] According to the target performance data, prune a target network module in an initial speech synthesis model to obtain a to-be-trained speech synthesis model;
[0012] Based on the speech data, train the to-be-trained speech synthesis model to obtain a target speech synthesis model matched with the terminal, so that the terminal processes a target text based on the target speech synthesis model to obtain synthesized speech data with a target timbre.
[0013] An embodiment of the present application further provides a speech processing device, including:
[0014] An obtaining module, configured to obtain a target text and a target speech synthesis model, where the target speech synthesis model is obtained by pruning a target network module in an initial speech synthesis model according to target performance data of a terminal corresponding to a target pronunciation object, and training a to-be-trained speech synthesis model obtained by the pruning process with speech data of the target pronunciation object, and the speech data has a target timbre;
[0015] A synthesis module, configured to use the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre.
[0016] An embodiment of the present application further provides a speech processing device, including:
[0017] A first receiving module, configured to receive a speech synthesis service request from a terminal;
[0018] A first determining module, configured to determine the target performance data of the terminal and the speech data of a target pronunciation object according to the speech synthesis service request from the terminal, and the speech data has a target timbre;
[0019] A clipping module, configured to clip a target network module in an initial speech synthesis model according to the target performance data to obtain a speech synthesis model to be trained;
[0020] A training module, configured to train the speech synthesis model to be trained based on the speech data to obtain a target speech synthesis model matching the terminal, so that the terminal processes a target text based on the target speech synthesis model to obtain synthesized speech data with a target voice color.
[0021] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which is suitable for being loaded by a processor to execute the steps in the speech processing method described in any of the above embodiments.
[0022] An embodiment of the present application further provides a computer device including a memory and a processor. The memory stores a computer program, and the processor executes the steps in the speech processing method described in any of the above embodiments by calling the computer program stored in the memory.
[0023] The speech processing method, device, computer-readable storage medium, and computer device provided in the embodiments of the present application determine target performance data and speech data of a target pronunciation object according to a speech synthesis service request from a terminal. The speech data has a target voice color. The target network module in the initial speech synthesis model is clipped according to the target performance data to obtain a speech synthesis model to be trained, and the speech synthesis model to be trained is trained using the speech data to obtain a target speech synthesis model, so that the terminal uses the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with a target voice color. The embodiments of the present application can clip the target network module in the initial speech synthesis model according to the target performance data of the terminal to obtain a speech synthesis model to be trained that matches the target performance data of the terminal, and then train the speech synthesis model to be trained using the speech data to obtain a target speech synthesis model that matches the performance data of the terminal, so as to provide a speech synthesis service that meets the target performance data of the terminal based on the target speech synthesis model and improve the user experience. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1a It is a schematic diagram of an application scenario of the voice processing method in the embodiment of the present application.
[0026] Figure 1b It is another schematic diagram of an application scenario of the voice processing method in the embodiment of the present application.
[0027] Figure 2 It is a schematic flowchart of the voice processing method provided by the embodiment of the present application.
[0028] Figure 3 It is a schematic flowchart of the voice processing method provided by the embodiment of the present application.
[0029] Figure 4 It is a schematic diagram of the target voice synthesis model provided by the embodiment of the present application.
[0030] Figure 5 It is a schematic sub - flowchart of the voice processing method provided by the embodiment of the present application.
[0031] Figure 6 It is a schematic diagram of the cropping process provided by the embodiment of the present application.
[0032] Figure 7 It is another schematic flowchart of the voice processing method provided by the embodiment of the present application.
[0033] Figure 8 It is a schematic flowchart of the voice processing method provided by another embodiment of the present application.
[0034] Figure 9 It is a schematic flowchart of the voice processing method provided by another embodiment of the present application.
[0035] Figure 10 It is another schematic flowchart of the voice processing method provided by another embodiment of the present application.
[0036] Figure 11 It is yet another schematic flowchart of the voice processing method provided by another embodiment of the present application.
[0037] Figure 12 It is a schematic diagram of the structure of the voice processing device provided by the embodiment of the present application.
[0038] Figure 13 It is another schematic diagram of the structure of the voice processing device provided by the embodiment of the present application.
[0039] Figure 14 It is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0041] The embodiments of the present application provide a voice processing method, apparatus, computer-readable storage medium, and computer device. Specifically, the voice processing method in the embodiments of the present application can be executed by a computer device. Among them, the computer device can be a terminal or a server, etc., and can also be jointly executed by a terminal and a server. The terminal can be a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer (PC), a smart vehicle terminal, a robot, or a device with voice functions such as a similar robot function. The server can be an independent physical server, a service node in a blockchain system, a server cluster composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0042] Figure 1a and Figure 1b is a schematic diagram of the application scenario of the voice processing method provided by the embodiments of the present application. Among them, communication is carried out between the terminal and the server. The terminal and the server can be connected and communicate through communication connection methods such as Bluetooth, USB (Universal Serial Bus), or network. In the embodiments of the present application, the connection and communication through the network are taken as an example for description.
[0043] Among them, the training of the speech synthesis model requires a large amount of computing power and a long training time. Therefore, the training of the speech synthesis model is mostly completed on the server side. Since the speech synthesis model can be implemented using different neural networks, if the effect of such a speech synthesis model is to be optimized, it usually takes several hours to dozens of hours, and the waiting time is too long. In addition, there are various devices connected to the server, and the computing power supported by each device varies greatly. The size of the speech synthesis model that different devices can support is also different. If the server uses a model structure with a fixed size, the synthesis ability cannot run on devices with low computing power, and high-end computing power devices cannot fully utilize their computing power, resulting in the synthesis effect not being optimal. To solve this problem, several different model structures or different model sizes can be configured on the server side, and then selected according to the computing power of the user device. For example, a matching model is selected from several different model structures or different model sizes, and after selecting the matching model, the speech data of the user device is trained to obtain a target model that matches the computing power of the user device. In this case, multiple model systems need to be designed on the server side.
[0044] This application proposes a new technical solution to achieve the matching of the target speech synthesis model and the target performance data of the terminal.
[0045] As Figure 1a shown, the target pronunciation object corresponding to the terminal sends a speech synthesis service request to the server. The server receives the speech synthesis service request and determines the speech data and the target performance data of the terminal based on the speech synthesis service request. The server performs pruning processing on the target network module in the pre-trained initial speech synthesis model according to the target performance data to obtain a speech synthesis model to be trained, and then trains the speech synthesis model to be trained according to the speech data to obtain a target speech synthesis model, and deploys the target speech synthesis model to the synthesis service / speech synthesis service of the server. Among them, the initial speech synthesis model can also be understood as a full-scale speech synthesis model. When the terminal wants to implement speech synthesis, it sends a target text to the server. The synthesis service of the server selects the corresponding target speech synthesis model to perform speech synthesis processing on the target text, obtains the synthesized speech data of the target timbre corresponding to the target pronunciation object, and returns the synthesized speech data to the terminal device for display and / or playback on the terminal device.
[0046] As Figure 1bAs shown, after the server trains the pre-trained speech synthesis model obtained by cropping and processing the speech data pair, and after obtaining the target speech synthesis model, the server sends the target speech synthesis model to the terminal. In this way, the synthesis ability of the target speech synthesis model is built into the terminal. When the terminal wants to perform speech synthesis, it sends the target text to the built-in target speech synthesis model to call the target speech synthesis model to perform speech synthesis processing on the target text, obtaining synthesized speech data with the target timbre, and displaying the synthesized speech data on the terminal. This method can greatly relieve the service pressure on the server.
[0047] In one embodiment, also refer to the Figure 1b corresponding application scenario. The full amount of parameter data corresponding to the initial speech synthesis model and the full amount of computational graphs corresponding to the initial speech synthesis model will be built into the terminal. When the server trains the speech synthesis model to be trained obtained by cropping and processing the speech data pair, and after obtaining the target speech synthesis model, the server sends the parameter change data in the target speech synthesis model and the target computational graph of the target speech synthesis model to the terminal. The terminal performs corresponding parameter replacement processing on the full amount of parameter data according to the parameter change data, and performs corresponding computational graph replacement processing on the full amount of computational graphs according to the target computational graph, obtaining the target speech synthesis model. In this case, even if the model structure corresponding to the initial speech synthesis model on the server side changes, the terminal does not need to be updated; and since only the parameter change data and the target computational graph are sent to the terminal, instead of sending the entire target speech synthesis model, the network resources and consumption of the server can be further reduced.
[0048] The above describes some application scenarios of the embodiments of the present application. Next, a speech processing method, device, computer-readable storage medium, and computer device provided by the embodiments of the present application will be described in detail. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.
[0049] In one embodiment, as Figure 2 and Figure 3 shown, it is a schematic flowchart of the speech processing method provided by the embodiments of the present application. This speech processing method is applied to the server side. As Figure 2 shown, this speech processing method includes the following steps.
[0050] 101. According to the speech synthesis service request from the terminal, determine the target performance data of the terminal and the speech data of the target pronunciation object, and this speech data has the target timbre.
[0051] Among them, the target pronunciation object can be a real user object, a virtual user object, or a machine pronunciation object, etc. The real user object includes user objects holding terminals, such as users holding smartphones, etc.; the virtual user object can be a user object in virtual games, virtual anchors, virtual teachers, personal voice packs for map navigation, etc. with voice functions; the machine pronunciation object includes pronunciation objects corresponding to robots or machines similar to robots with voice functions, etc. In some embodiments, the machine pronunciation object can also be understood as a kind of virtual user object.
[0052] The voice data of the target pronunciation object refers to the voice data emitted by the target pronunciation object, such as the voice data emitted by a user holding a smartphone, the voice data emitted by a robot, the voice data emitted by a virtual teacher, etc. The voice data has the target timbre corresponding to the target pronunciation object.
[0053] The terminal corresponding to the target pronunciation object includes smartphones, devices running virtual games, devices running application programs such as virtual teachers, robots, or machines similar to robots.
[0054] The target performance data of the terminal includes at least one performance index information, and the performance index information can be the computing power information characterizing the terminal, such as the computing power information supported by the terminal hardware. In one embodiment, the target performance data of the terminal can be represented by how many million integer instructions per second (DMIPS) are executed, where how many million integer instructions per second can also be understood as the computing power information supported by the terminal hardware. It can be understood that the target performance data can also be represented by other performance indexes, which will not be limited here.
[0055] In one embodiment, a corresponding APP or a corresponding small program is built into the terminal, and the interface corresponding to the voice synthesis service request is displayed through the corresponding APP / small program. On this interface, there is a voice control. Triggering this voice control records the target pronunciation object. After the recording is completed, a voice synthesis server request is generated and sent to the server. In this way, the voice data of the target pronunciation object is carried in the voice synthesis service request. In one embodiment, on the interface of the corresponding APP / small program displayed on the terminal, the voice synthesis service control is triggered to generate a voice synthesis service request and send it to the user. After the server receives the voice synthesis service request, it displays an interface for entering voice to the user for the user to input voice data and send the voice data to the server. In other embodiments, there can also be other application scenarios, which will not be elaborated here.
[0056] In one embodiment, the voice data of the target pronunciation object and the target performance data of the corresponding terminal are pre-stored in the server. Thus, when a voice synthesis service request is received, the voice data of the target pronunciation object and the target performance data can be directly obtained from the server.
[0057] In one embodiment, the performance data of various different terminal models are pre-stored in the server. When a voice synthesis service request is received, the voice data in the voice synthesis service request is obtained, and the target performance data corresponding to the terminal model of the target pronunciation object is determined from the pre-stored performance data.
[0058] In some other embodiments, the voice data of the target pronunciation object and the target performance data of the corresponding terminal can be obtained from the terminal. For example, the terminal sends a voice synthesis service request to the server, and the voice data of the target pronunciation object is carried in the voice synthesis service request; after sending the voice synthesis service request, the terminal obtains the target performance data corresponding to the terminal to send the target performance data to the server, or the terminal obtains the corresponding target performance data when the recording function is turned on to send the target performance data to the server, etc.
[0059] In one embodiment, the target performance data can also be carried in the voice synthesis service request and sent to the server together with the voice data.
[0060] 102. According to the target performance data, perform pruning processing on the target network module in the initial voice synthesis model to obtain a voice synthesis model to be trained.
[0061] The initial voice synthesis model can be obtained through pre-training and stored in the server, so that the initial voice synthesis model can be directly obtained.
[0062] Please refer to Figure 4 , Figure 4 Although it is a schematic structural diagram of the target voice synthesis model, the structure of the initial voice synthesis model after pre-training is the same as that of the target voice synthesis model, and also includes an acoustic model and a vocoder. Among them, the acoustic model includes an encoder, an attention network, and a decoder. Among them, the specific network structures of the encoder and the decoder can be network structures such as Long Short-Term Memory (LSTM) neural network and Convolutional Neural Networks (CNN). In the embodiments of the present application, the network structure corresponding to the encoder is called an encoding neural network, and the network structure corresponding to the decoder is called a decoding neural network.
[0063] The target network module in the initial speech synthesis model includes one or more of an encoding neural network, an attention network, a decoding neural network, and a vocoder-corresponding vocoder neural network (if the vocoder uses a neural network to complete the response function). It can be understood that one or more corresponding network structures in the encoding neural network, attention network, decoding neural network, and vocoder-corresponding vocoder neural network in the initial speech synthesis model are subject to pruning processing.
[0064] Among them, the pruning processing can be understood as pruning / discarding / skipping certain network layers in the target network module. For example, the number of network layers in a certain network module is 20 layers. During training, the second layer is discarded for the first time, and then the target network module with the second layer discarded is used for training. The seventh to ninth layers are discarded for the second time (the second layer is still there at the second time), and then the target network module with the seventh to ninth layers discarded is used for training, etc. Among them, it can be random pruning processing or pruning processing according to some pruning rules / laws. For example, pruning is performed at intervals of a preset number of network layers and then pruning is performed again at intervals of a preset number of network layers, etc.
[0065] The initial speech synthesis model is a model pre-trained based on the speech samples of multiple pronunciation objects. During the pre-training process of the initial speech synthesis model (such as during each round of training), the network layers of the network modules therein are also subject to pruning processing.
[0066] In one embodiment, as Figure 5 shown, the steps of pre-training to obtain the initial speech synthesis model may include the following steps 1021 to 1024.
[0067] 1021. Obtain the speech samples of multiple pronunciation objects and the original speech synthesis model, where the multiple pronunciation objects have different timbres.
[0068] The multiple pronunciation objects can be multiple pronunciation objects of the same type, such as all real user objects, or multiple pronunciation objects of different types, such as including both real user objects and virtual user objects, etc. Each pronunciation object has a different timbre. The original speech synthesis model in this step is a speech synthesis model that has not been pre-trained.
[0069] 1022. Input the speech samples of the multiple pronunciation objects into the original speech synthesis model, and perform pruning processing on the network modules in the original speech synthesis model, so as to use the original speech synthesis model after pruning processing to perform acoustic processing on the speech samples to obtain training acoustic features.
[0070] The network modules for pruning in the original speech synthesis model include one or more of an encoding neural network, an attention network, a decoding neural network, and a vocoder corresponding vocoder neural network (if the vocoder uses a neural network to complete the response function). Understandably, pruning processing is performed on one or more corresponding network structures of the encoding neural network, the attention network, the decoding neural network, and the vocoder corresponding vocoder neural network.
[0071] As Figure 6 described, it is a schematic diagram of the pruning process provided by the embodiments of the present application. Taking the network module as the decoding neural network as an example, assuming that the decoding neural network is a 9-layer network structure, after pruning the decoding neural network, the decoding neural network can be pruned to 1 layer, 3 layers, 6 layers, etc.
[0072] Next, the original speech synthesis model after pruning processing is used to perform acoustic feature processing on the speech sample to obtain training acoustic features. Specifically, it includes: using the original speech synthesis model (the encoder in it) to extract semantic features from the speech sample to obtain training semantic features, where the training semantic features include context information, and using the original speech synthesis model (the attention network and the decoder in it) to perform timbre processing on the training semantic features to obtain training acoustic features. Among them, the training acoustic features can be Mel spectrum features. For example, MFCC (Mel Frequency Cepstrum Coefficient) is extracted from the speech sample, and then the training semantic features of the speech sample are determined according to the Mel cepstrum coefficients; timbre processing is performed on the training semantic features to obtain Mel spectrum features. The training acoustic features can also be other acoustic features.
[0073] Furthermore, obtain the phoneme sequence corresponding to the text in the speech sample; use the original speech synthesis model (the encoder in it) to extract semantic features from the phoneme sequence to obtain training semantic features.
[0074] Among them, a phoneme refers to the pronunciation sound unit of a character or word in the text, such as the initial consonant and the final consonant in Chinese characters, etc. Correspondingly, a phoneme sequence can refer to a sequence composed of multiple phonemes. In one embodiment, the text in the speech data of the target pronunciation object is extracted, and then the text is parsed, such as performing word segmentation processing on the text, and then phonetic annotation is performed on the multiple word segments and / or characters obtained by the word segmentation processing to obtain the phonemes of each word segment and / or character, and the obtained phonemes are combined to obtain a phoneme sequence.
[0075] For example, a text in the speech data is: I love my mom. Correspondingly, the word segmentation is "I", "love", "mom". When obtaining the phonemes corresponding to these word segments, combine the phonemes corresponding to these word segments to obtain the phoneme sequence w o ai m a m a of this text.
[0076] 1023. Use the original speech synthesis model to perform speech synthesis processing on the training acoustic features to obtain the training predicted speech.
[0077] The original speech synthesis model refers to the speech synthesis model after cropping processing. Use the (vocoder in) the original speech synthesis model after cropping processing to perform speech synthesis processing on the training acoustic features to obtain the training predicted speech.
[0078] 1024. Based on the loss value between the training predicted speech and the corresponding speech sample, adjust the network parameters in the original speech synthesis model to obtain the initial speech synthesis model after pre-training.
[0079] In this way, the training is completed to obtain the initial speech synthesis model.
[0080] For example, assume there are 1000 speech samples. Then, select 10 speech samples from the 1000 speech samples in one training and input the 10 speech samples into the untrained original speech synthesis model. At this time, perform cropping processing on the network module in the original speech synthesis model. For example, discard the third layer of the decoding neural network. In this way, after the data enters the decoding neural network and is processed by the first network layer and the second network layer, it then enters the fourth network layer for processing until the decoding neural network finishes processing to obtain the training acoustic features corresponding to the 10 speech samples. Perform speech synthesis processing on the training acoustic features corresponding to the 10 speech samples to obtain 10 different training predicted speeches. Compare the 10 training predicted speeches with the corresponding 10 speech samples respectively, and based on the loss value between the corresponding training predicted speech and the speech sample, adjust the network parameters in the original speech synthesis model.
[0081] In the subsequent training, 10 speech samples are taken out from the 1000 samples (these 10 speech samples are different from those in the last training), and the 10 speech samples are input into the original speech synthesis model. At this time, the network modules in the original speech synthesis model are trimmed, for example, the sixth to ninth layers of the decoding neural network are discarded. In this way, after the data enters the decoding neural network, it is processed by the fifth network layer, and then enters the tenth network layer for processing until the decoding neural network is processed, so as to obtain the training acoustic features corresponding to the 10 speech samples. The training acoustic features corresponding to the 10 speech samples are processed for speech synthesis to obtain 10 different training predicted speech. Based on the loss value between the corresponding training predicted speech and the speech sample, the network parameters in the original speech synthesis model are adjusted.
[0082] In this way, all 1000 speech samples are trained once to complete one round of training. The 1000 speech samples are trained multiple times in the same way to obtain an initial speech synthesis model.
[0083] According to the above steps, the untrained original speech synthesis model is pre-trained to obtain the pre-trained initial speech synthesis model of the present application. It should be noted that only one model (an original speech synthesis model) is trained in the present application, and the training of the model can be completed at one time. In addition, during the training process, the network modules in the original speech synthesis model are randomly pruned. In this way, the initial speech synthesis model obtained by pre-training can also have the ability of a complete model after pruned. This is the reason why the target network module in the pre-trained initial speech synthesis model can be pruned.
[0084] It should be noted that since the acoustic model in the original speech synthesis model is relatively important, in some embodiments, the acoustic model in the original speech synthesis model can be trained separately, and then the vocoder can be trained separately, and then the trained acoustic model and vocoder can be merged to obtain the pre-trained initial speech synthesis model.
[0085] For example, when training the acoustic model in the original speech synthesis model alone, it may include: obtaining speech samples and an acoustic model of multiple pronunciation objects, where the multiple pronunciation objects have different timbres; inputting the speech samples into the acoustic model, performing pruning processing on the network module in the acoustic model, using the pruned acoustic model to perform acoustic feature processing on the speech samples to obtain first training acoustic features; performing audio signal processing on the speech samples to obtain second training acoustic features, and updating the parameter data of the neural network in the acoustic model based on the loss value between the first training acoustic features and the second training acoustic features to obtain a pre-trained acoustic model. Among them, the step of using the pruned acoustic model to perform acoustic feature processing on the speech samples to obtain first training acoustic features includes: using the pruned acoustic model to extract semantic features from the speech samples to obtain training semantic features, and performing timbre processing on the training semantic features to obtain first training acoustic features.
[0086] When training the vocoder, the input is the training acoustic features and the output is the synthesized speech data. In one embodiment, when training the vocoder, the network layers in the vocoder can also be randomly pruned so that the trained vocoder also has the capabilities of the complete model. Thus, in the step of pruning the target network module in the initial speech synthesis model, the target network module also includes the vocoder.
[0087] It should be noted that the initial speech synthesis model in the embodiments of the present application can adopt a complete training method, such as inputting a phoneme sequence and outputting training predicted speech; or it can be trained in a two-stage manner, such as separately training the acoustic model and the vocoder, and then combining the acoustic model and the vocoder to obtain the initial speech synthesis model.
[0088] The above embodiments describe the process of pre-training to obtain the initial speech synthesis model. Next, how to use the pre-trained initial speech synthesis model will be described.
[0089] In one embodiment, the step of 102 above includes: determining the target number of layers of the target network module in the initial speech synthesis model according to the target performance data; pruning the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain a speech synthesis model to be trained.
[0090] Determining the target number of layers of the target network module according to the target performance data and pruning the number of layers of the target network module in the initial speech synthesis model to the target number of layers. Thus, the obtained speech synthesis model to be trained matches the target performance data of the terminal, enabling the speech synthesis model to be trained to achieve the best speech synthesis effect on the terminal and improving the user experience.
[0091] In one embodiment, the step of determining the target number of layers of the target network module in the initial speech synthesis model according to the target performance data includes: obtaining the total complexity information corresponding to the initial speech synthesis model, the module complexity information corresponding to the target network module, the number of layers corresponding to the target network module, and the performance data of the terminal matched with the initial speech synthesis model; determining the target number of layers of the target network module in the initial speech synthesis model according to the total complexity information, the module complexity information, the number of layers, the performance data, and the target performance data. Among them, the total complexity information corresponding to the initial speech synthesis model and the module complexity information corresponding to the target network module can be quantified using time complexity. For example, the amount of computation corresponding to the initial speech synthesis model is used as the total complexity information corresponding to the initial speech synthesis model, and the amount of computation corresponding to the target network module is used as the module complexity information corresponding to the target network module. Among them, the amount of computation of the model or the amount of computation of the network module can be determined in the existing manner.
[0092] Among them, if the target network module is an encoding neural network, obtain the module complexity information required by the encoding neural network corresponding to the encoder and the number of layers corresponding to the encoding neural network. If the target network module is a decoding neural network and an encoding neural network, obtain the module complexity information required by the corresponding encoding neural network and decoding neural network (together) and the number of layers corresponding to the encoding neural network and decoding neural network (together), and so on. Here, no further examples are given one by one. Among them, in the acoustic model, generally, the decoding neural network is preferably pruned, rather than the encoding neural network and the attention network. Therefore, here, taking the target network module as the decoding neural network and pruning the decoding neural network as an example to illustrate how to determine the target number of layers of the decoding neural network.
[0093] Assume that the performance data supported by the model has a linear relationship with the model complexity, or in other words, it is proportional. Let the total complexity information of the initial speech synthesis model be A. It is necessary to perform pruning on the decoding neural network corresponding to the decoder. Correspondingly, the module complexity information of the decoding neural network part in the initial speech synthesis model is D, the number of layers of the decoding neural network is m, and the performance data of the terminal corresponding to the initial speech synthesis model is DMF. The target performance data of the terminal corresponding to the target pronunciation object is DMX, and the model complexity supported by DMX is A'. Among them, the performance data of the terminal corresponding to the initial speech synthesis model refers to the performance data corresponding to the terminal with the lowest standard that can run the initial speech synthesis model, which can be determined in the following way: First, determine the terminal with the lowest standard that can run the initial speech synthesis model, and then execute a performance test program on this lowest-standard terminal to determine the performance data of the terminal (such as computing power information). For example, the number of million integer instructions executed per second by this terminal is used as the performance data of the terminal. According to the linear relationship between the performance data supported by the model and the model complexity, the following formula (1) holds.
[0094]
[0095] A' includes two parts, the module complexity information of the decoding neural network after pruning and the complexity information of other parts. The complexity information of other parts is A - D. Assume that the decoding neural network needs to be pruned to n layers, where n is the target number of layers after pruning of the decoding neural network. Then the following formula (2) holds.
[0096]
[0097] According to formula (1) and formula (2), the following formula (3) can be obtained.
[0098]
[0099] According to formula (3), the target number of layers of the target network module can be obtained. It should be noted that other methods can also be used to determine the target number of layers.
[0100] After obtaining the target number of layers, prune the number of layers of the target network module in the initial speech synthesis model to cut the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain the speech synthesis model to be trained. For example, if the target number of layers of the decoding neural network is 10, then cut the number of layers of the decoding neural network to 10 layers. Among them, the pruning can be random pruning or pruning according to a certain rule / law. The obtained speech synthesis model to be trained matches the target performance data of the terminal corresponding to the target pronunciation object, so that the speech synthesis model can achieve the best effect.
[0101] 103. Based on the voice data, train the voice synthesis model to be trained to obtain the target voice synthesis model that matches the terminal. Obtaining the matching target voice synthesis model enables the terminal to process the target text based on the target voice synthesis model to obtain the synthesized voice data with the target voice color.
[0102] The voice synthesis model to be trained obtained after the cropping process cannot be directly used to obtain the synthesized voice data of the target pronunciation object. Therefore, it is necessary to use the voice data to train the voice synthesis model to be trained to update the parameter data in the voice synthesis model to be trained.
[0103] In one embodiment, the step 103 above includes: determining the number of update rounds of the voice synthesis model to be trained according to the target number of layers of the target network module, and training the voice synthesis model to be trained according to the voice data until the number of training rounds reaches the number of update rounds to obtain the target voice synthesis model.
[0104] Among them, assuming that the number of layers of the voice synthesis model is linearly related to the number of update rounds, based on this assumption, the step of determining the number of update rounds of the voice synthesis model to be trained according to the target number of layers of the target network module includes: obtaining the number of training rounds of the target network module and the number of layers corresponding to the target network module when the initial voice synthesis model is pre-trained, and determining the number of update rounds of the target network module according to the number of training rounds and layers of the target network module and the target number of layers, and determining this number of update rounds as the number of update rounds of the voice synthesis model. For example, when the initial voice synthesis model is pre-trained, the number of training rounds of the target network module is 10 rounds, and the corresponding number of layers is 50 layers. When the target number of layers is 40 layers, the corresponding number of update rounds is 8 rounds.
[0105] After determining the number of update rounds, train the voice synthesis model to be trained according to the voice data until the number of training rounds reaches the determined number of update rounds, and stop training to obtain the target voice synthesis model.
[0106] In some embodiments, stopping the training of the voice synthesis model to be trained may not be determined according to the number of update rounds, but according to whether the loss function of the voice synthesis model to be trained converges, or whether the loss value of the loss function of the voice synthesis model to be trained is within the corresponding threshold range. If so, stop training to obtain the target voice synthesis model.
[0107] In one embodiment, the step of training the speech synthesis model to be trained based on speech data includes: obtaining a phoneme sequence corresponding to the text in the speech data; using an encoder in the speech synthesis model to extract semantic features from the phoneme sequence to obtain semantic features, which are represented in the form of vectors and include context information; using an attention network and a decoder in the speech synthesis model to perform timbre processing on the semantic features to obtain acoustic features; using a vocoder in the speech synthesis model to perform speech synthesis processing on the acoustic features to obtain predicted speech; and adjusting network parameters in the speech synthesis model based on the loss value between the predicted speech and the speech data to obtain a target speech synthesis model. The target speech synthesis model is a speech synthesis model that can achieve cloning of the target pronunciation object's speech.
[0108] In one embodiment, the acoustic model in the speech synthesis model can also be trained, and after merging the trained acoustic model and the vocoder, a target speech synthesis model is obtained. Among them, training the acoustic model in the speech synthesis model to obtain a trained acoustic model includes: obtaining a phoneme sequence corresponding to the text in the speech data; encoding and decoding the phoneme sequence through the acoustic model to obtain first acoustic features; performing audio signal processing on the speech data to obtain second acoustic features; comparing the first acoustic features and the second acoustic features, and updating the parameter data of the acoustic model according to the comparison result to obtain a trained acoustic model. Among them, the step of encoding and decoding the phoneme sequence through the acoustic model to obtain first acoustic features includes: using an encoder in the acoustic model to extract semantic features from the phoneme sequence to obtain semantic features, which are represented in the form of vectors and include context information; using an attention network and a decoder in the acoustic model to perform target timbre processing on the semantic features to obtain first acoustic features.
[0109] After training the speech synthesis model to be trained to obtain a target speech synthesis model that matches the terminal, the target speech synthesis model can be used for speech synthesis processing.
[0110] In one embodiment, obtaining a matching target speech synthesis model enables the terminal to process the target text based on the target speech synthesis model to obtain synthesized speech data with the target voice color, which can be applied to the following scenarios: The terminal sends the target text to the server, and the server calls the target speech synthesis model according to the target text to perform speech synthesis processing to obtain synthesized speech data, and sends the synthesized speech data to the terminal to display and / or play the synthesized speech data on the terminal.
[0111] In one embodiment, obtaining a matching target speech synthesis model enables the terminal to process the target text based on the target speech synthesis model to obtain synthesized speech data with a target voice color, which can be applied to the following scenarios: after the server obtains a target speech synthesis model matching the terminal, it sends the target speech synthesis model to the terminal so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with a target voice color.
[0112] Correspondingly, in this embodiment, the speech processing method further includes step 104.
[0113] 104. Send the target speech synthesis model to the terminal so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with a target voice color.
[0114] The target voice color is the voice color of the target pronunciation object. Among them, the target text can be a paragraph, a story, a sentence, one or more lines of text, a phrase or multiple phrases, etc. At the same time, the target text can cover various fields, such as technology, sports, leisure and entertainment, food, and literature, etc.
[0115] In the above embodiment, the target network module in the initial speech synthesis model can be trimmed according to the target performance data of the terminal to obtain a speech synthesis model to be trained that matches the target performance data of the terminal, and then the speech synthesis model to be trained is trained using speech data to obtain a target speech synthesis model that matches the performance data of the terminal, so as to provide a speech synthesis service that meets the performance data of the terminal based on the target speech synthesis model and improve the user experience.
[0116] In one embodiment, as Figure 7 shown, the steps of the above 104 include 1041 to 1042.
[0117] 1041. Extract the parameter change data in the target speech synthesis model.
[0118] Among them, the parameter change data in the target speech synthesis model is based on the initial speech synthesis model. In one embodiment, since the network structure of the target network module is trimmed, therefore, the parameter data corresponding to the target network module and the parameter data of other network module changes caused by the trimming process of the target network module are used as the parameter change data. Preferably, the parameter change data determined in this way is used, and this way can improve the efficiency of the terminal to generate the target speech synthesis model. In another embodiment, the data of the parameter change part in the target network module is used as the parameter change data. For example, if the tenth layer is trimmed in the decoding neural network, then the parameter data of all network layers from the ninth layer to the back of the decoding neural network is used as the parameter change data.
[0119] Among them, the target computation graph corresponding to the target network module can be implemented by calling a preset function interface.
[0120] 1042, send the parameter change data and the target computation graph corresponding to the target network module to the terminal, so that the terminal generates a target voice synthesis model according to the parameter change data and the target computation graph, and performs voice synthesis processing on the target text according to the target voice synthesis model to obtain synthesized voice data with a target timbre.
[0121] After determining the parameter change data and the target computation graph corresponding to the target network module, send the parameter change data and the target computation graph to the terminal, as Figure 3 shown. In one embodiment, before sending the parameter change data and the target computation graph to the terminal, it is necessary to package the parameter change data and the target computation graph. Further, encryption can be performed, and then packaging, and then sending to the terminal. The terminal generates a target voice synthesis model according to the parameter change data and the target computation graph. It should be noted that the target voice synthesis model generated by this terminal is the same as the target voice synthesis model generated on the server.
[0122] After the terminal generates the target voice synthesis model, perform voice synthesis processing on the target text according to the target voice synthesis model to obtain synthesized voice data with a target timbre. Display and / or play the synthesized voice data on the terminal interface.
[0123] In this embodiment, the server does not send the trained target voice synthesis model to the terminal, but sends the parameter change data in the target voice synthesis model and the target computation graph corresponding to the target network module to the terminal, which can reduce the network resources and consumption of the server. At the same time, even if the model structure corresponding to the initial voice synthesis model on the server side changes, the terminal side does not need to be updated.
[0124] The above part describes the voice processing method applied to the server side. Next, the voice processing method applied to the terminal will be described.
[0125] As Figure 8 shown, it is a schematic flowchart of the voice processing method provided by the embodiment of the present application. Among them, this voice processing method is applied to the terminal, and this voice processing method includes the following steps.
[0126] 201, obtain a target text and a target voice synthesis model. The target voice synthesis model is obtained by cropping the target network module in the initial voice synthesis model according to the target performance data of the terminal corresponding to the target pronunciation object, and training the cropped voice synthesis model to be trained with the voice data of the target pronunciation object. The voice data has a target timbre.
[0127] Among them, the target text is determined on the terminal. For example, the target text is input / pasted in the input box of the terminal interface; or a document is loaded on the terminal interface to extract the target text in the document; or a picture is loaded to extract the target text in the picture; or even the target text in a video is extracted, etc. The target text includes a paragraph, a story, a sentence, one or more lines of text, a phrase or multiple phrases, etc., without specific limitation.
[0128] Among them, the target speech synthesis model is sent by the server. The server performs pruning processing on the target network module in the initial speech synthesis model according to the target performance data of the terminal of the target pronunciation object, and uses the speech data of the target pronunciation object to train the to-be-trained speech synthesis model obtained by the pruning processing to obtain the target speech synthesis model. Among them, the speech data of the target pronunciation object has the target timbre. Specifically, reference can be made to the corresponding description on the server side in the above text, which will not be elaborated here.
[0129] In an embodiment, the initial speech synthesis model can be obtained after pre-training. Specifically, for the steps to obtain the pre-trained initial speech synthesis model, reference can be made to the corresponding description in the above text, which will not be elaborated here.
[0130] 202. Use the target speech synthesis model to perform speech synthesis processing on the target text to obtain the synthesized speech data with the target timbre.
[0131] In an embodiment, the steps of 202 include: obtaining the target phoneme sequence corresponding to the target text; performing acoustic processing on the target phoneme sequence through the target speech synthesis model to obtain the target acoustic features including the target timbre; using the target speech synthesis model to perform speech synthesis processing on the target acoustic features to obtain the synthesized speech data with the target timbre. For the implementation of the specific steps, reference can be made to the description of the corresponding steps in the above text and Figure 4 as shown, which will not be elaborated here. Further, the step of performing acoustic processing on the target phoneme sequence through the target speech synthesis model to obtain the target acoustic features including the target timbre includes: performing semantic feature extraction on the target phoneme sequence through the encoder in the target speech synthesis model to obtain the target semantic features, and the target speech features include context information; using the attention network and decoder in the target speech synthesis model to perform target timbre processing on the semantic features to obtain the target acoustic features.
[0132] The target speech synthesis model in this embodiment is obtained by pruning the initial speech synthesis model according to the target performance data of the terminal and training the to-be-trained speech synthesis model after pruning with the speech data of the target pronunciation object. In this way, the parameter data in the to-be-trained speech synthesis model is updated with the speech data of the target pronunciation object, so that the terminal provides a speech synthesis service that meets the target performance data of the terminal based on this target speech synthesis model, improving the user experience.
[0133] In one embodiment, as Figure 9 shown, before obtaining the target text, the speech processing method further includes the following steps 201a to 201c.
[0134] 201a, generate a speech synthesis service request and obtain the target performance data of the terminal.
[0135] Specifically, for the generation of the speech synthesis service request, please refer to the corresponding description in the above text and will not be elaborated here. In addition, in this embodiment, it is clearly defined to obtain the corresponding target performance data on the terminal.
[0136] In one embodiment, a corresponding APP / miniapp is built in the terminal or loaded. When the terminal loads the APP / miniapp for the first time, the performance determination program built in the APP / miniapp is automatically executed to determine the target performance data of the terminal through the performance determination program and save the target performance data. Or the target performance data is re-determined and saved every preset time when the APP / miniapp is loaded. It is also possible that the terminal sends the terminal model information to the server, obtains the target performance data corresponding to the terminal model information from the server, and saves it.
[0137] Correspondingly, the step of obtaining the target performance data of the terminal includes: obtaining the pre-saved target performance data of the terminal; or executing a performance determination program to determine the target performance data of the terminal, that is, by executing a performance determination program on the terminal to obtain the target performance data of the terminal.
[0138] To avoid the influence of other applications currently running on the terminal on the determination of the target performance data when executing the performance determination program, in one embodiment, the above step of executing a performance determination program to determine the target performance data of the terminal includes: obtaining other applications currently running on the terminal, and pausing or closing other applications; executing a performance determination program to determine the target performance data of the terminal; starting other applications. That is, before executing the performance determination program, pause or close other applications currently running on the terminal, and automatically start the paused or closed other applications after executing the performance determination program.
[0139] 201b, send the speech synthesis service request and the target performance data to the server, so that the server performs pruning processing on the target network module in the initial speech synthesis model according to the target performance data, and uses the speech data in the speech synthesis service request to train the pruned speech synthesis model to be trained to obtain the target speech synthesis model, where the speech data has a target timbre.
[0140] Among them, the speech data is carried in the speech synthesis service request. Sending the speech synthesis service request and the target performance data to the server triggers the server to execute the corresponding steps described in the above embodiments to obtain the target speech synthesis model.
[0141] 201c, receive the target speech synthesis model returned by the server.
[0142] 201, obtain the target text and the target speech synthesis model.
[0143] 202, use the target speech synthesis model to perform speech synthesis processing on the target text to obtain the synthesized speech data with the target timbre.
[0144] In this embodiment, it is further limited that the terminal determines the target performance data and uploads the target performance data to the server.
[0145] In one embodiment, as Figure 10 and Figure 11 shown, it is another flowchart of the speech processing method provided by the embodiment of the present application. Please refer to Figure 10 , the speech processing method includes the following steps.
[0146] 301, obtain the parameter change data in the target speech synthesis model and the target computation graph corresponding to the target network module in the target speech synthesis model, where the parameter change data and the computation graph are determined and sent by the server.
[0147] After the server obtains the target speech synthesis model, extract the parameter change data in the target speech synthesis model and the target computation graph corresponding to the target network module in the target speech synthesis model, and send the parameter change data and the computation graph to the terminal, and the terminal receives the parameter change data and the computation graph.
[0148] In one embodiment, if the server sends the parameter change data and the target computation graph to the terminal after packing, then after receiving, the terminal needs to unpack to obtain the parameter change data and the target computation graph.
[0149] In one embodiment, if the server encrypts the parameter change data and the target computation graph and then packs and sends them to the terminal, after receiving them, the terminal also needs to unpack them to obtain the encrypted parameter change data and the target computation graph, and then decrypt them to obtain the parameter change data and the target computation graph.
[0150] 302. Generate a target voice synthesis model that matches the terminal according to the parameter change data and the target computation graph.
[0151] In one embodiment, the step of 302 includes: obtaining the full-scale parameter data and the full-scale computation graph saved on the terminal, where the full-scale parameter data and the full-scale computation graph are the parameter data and the computation graph corresponding to the initial voice synthesis model after pre-training; updating the full-scale parameter data and the full-scale computation graph according to the parameter change data and the target computation graph to generate a target voice synthesis model that matches the terminal.
[0152] In this embodiment, a corresponding APP is built in the terminal device or a small program is loaded. In the APP / small program, the full-scale parameter data corresponding to the pre-trained initial voice synthesis model and the full-scale computation graph corresponding to the initial voice synthesis model are built in correspondingly. Obtain the full-scale parameter data and the full-scale computation graph corresponding to the initial voice synthesis model.
[0153] In one embodiment, the step of updating the full-scale parameter data and the full-scale computation graph according to the parameter change data and the target computation graph to generate a target voice synthesis model that matches the terminal includes: using the parameter change data to replace the corresponding parameter data in the full-scale parameter data, and using the target computation graph to replace the computation graph corresponding to the target network module in the full-scale computation graph to generate a target voice synthesis model that matches the terminal. It should be noted that the target voice synthesis model is the same as the target voice synthesis model trained on the server.
[0154] 303. Obtain the target text. For the specific method of obtaining the target text, please refer to the corresponding description in the above text.
[0155] 304. Use the target voice synthesis model to perform voice synthesis processing on the target text to obtain synthesized voice data with the target timbre.
[0156] In one embodiment, a synthesis engine is built in the APP / small program. The synthesis engine is used to provide a first interface to generate a voice synthesis service request, provide a second interface to input the target text, etc. The synthesis engine realizes functions such as calling the generated target voice synthesis model to perform voice synthesis processing, and calling the full-scale parameter data and the full-scale computation graph to generate the target voice synthesis model, etc.
[0157] Correspondingly, the target text is input into the synthesis engine, and the target speech synthesis model is called through the synthesis engine to utilize the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target tone color, as Figure 11 shown. The synthesis engine is also used to execute performance determination programs and the like.
[0158] In the above embodiments, the appropriate matching target speech synthesis model is automatically determined according to the target performance data of the terminal, and by combining the built-in initial speech synthesis model in the APP / miniprogram and partial model update on the server side, the download of the model update parameters is minimized, the response speed and efficiency are improved, and the user experience is enhanced. The solutions in the above embodiments can be applied to the training of any speech synthesis model based on a deep neural network.
[0159] In the above embodiments, each embodiment has different emphases. For the sake of brevity, for the parts not described in detail in one embodiment, please refer to the corresponding descriptions in other embodiments.
[0160] All the above technical solutions can be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated herein one by one.
[0161] To facilitate better implementation of the speech processing method running on the server side in the embodiments of the present application, the embodiments of the present application also provide a speech processing device. Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of the speech processing device provided in the embodiments of the present application. The speech processing device is integrated in the server. The speech processing device 400 may include a first receiving module 401, a first determining module 402, a cropping module 403, and a training module 404.
[0162] The first receiving module 401 is configured to receive a speech synthesis service request from the terminal.
[0163] The first determining module 402 is configured to determine the target performance data of the terminal and the speech data of the target pronunciation object according to the speech synthesis service request from the terminal, and the speech data has the target tone color.
[0164] The cropping module 403 is configured to crop the target network module in the initial speech synthesis model according to the target performance data to obtain a speech synthesis model to be trained.
[0165] The training module 404 is configured to train the speech synthesis model to be trained based on the speech data to obtain the target speech synthesis model matched with the terminal, so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with the target speech color.
[0166] In one embodiment, the voice processing device further includes a first sending module 405, configured to send the target voice synthesis model to the terminal, so that the terminal processes the target text based on the target voice synthesis model to obtain synthesized voice data with a target voice color.
[0167] In one embodiment, the cropping module 403 is specifically configured to determine the target number of layers of the target network module in the initial voice synthesis model according to the target performance data; crop the number of layers of the target network module in the initial voice synthesis model to the target number of layers to obtain a voice synthesis model to be trained.
[0168] In one embodiment, when the cropping module 403 executes the step of determining the target number of layers of the target network module in the initial voice synthesis model according to the target performance data, it obtains the total complexity information required by the initial voice synthesis model, the module complexity information required by the target network module, the number of layers corresponding to the target network module, and the performance data of the terminal matching the initial voice synthesis model; determines the target number of layers of the target network module in the initial voice synthesis model according to the total complexity information, the module complexity information, the number of layers, the performance data, and the target performance data.
[0169] In one embodiment, the initial voice synthesis model includes an acoustic model, and the acoustic model includes an encoding neural network corresponding to an encoder, an attention network, and a decoding neural network corresponding to a decoder. The target network module in the initial voice synthesis module includes one or more of the encoding neural network, the attention network, and the decoding neural network.
[0170] In one embodiment, when the training module 404 executes the step of training the voice synthesis model to be trained based on the voice data, it specifically executes: determining the number of update rounds of the voice synthesis model to be trained according to the target number of layers; training the voice synthesis model to be trained according to the voice data until the number of training rounds reaches the number of update rounds to obtain a target voice synthesis model.
[0171] In one embodiment, when the training module 404 executes the step of training the voice synthesis model to be trained according to the voice data, it specifically executes: obtaining the phoneme sequence corresponding to the text in the voice data; performing acoustic processing on the phoneme sequence by using the voice synthesis model to be trained to obtain acoustic features; performing voice synthesis processing on the acoustic features by using the voice synthesis model to be trained to obtain predicted voice; adjusting the network parameters in the voice synthesis model to be trained based on the loss value between the predicted voice and the voice data.
[0172] In one embodiment, the first sending module 405 specifically performs: extracting the parameter change data in the target voice synthesis model; sending the parameter change data and the target computation graph corresponding to the target network module to the terminal, so that the terminal generates a target voice synthesis model according to the parameter change data and the target computation graph, and performs voice synthesis processing on the target text according to the target voice synthesis model to obtain synthesized voice data with the target timbre.
[0173] In one embodiment, the voice processing device 400 further includes a pre-training module 406. The pre-training module 406 is used to pre-train the initial voice synthesis model. Specifically, it is used to obtain voice samples of multiple pronunciation objects and the original voice synthesis model, and the multiple pronunciation objects have different timbres; input the voice samples of the multiple pronunciation objects into the original voice synthesis model, and perform pruning processing on the network module in the original voice synthesis model, so as to perform acoustic processing on the voice samples by using the original voice synthesis model after pruning processing to obtain training acoustic features; perform voice synthesis processing on the training acoustic features by using the original voice synthesis model to obtain training predicted voices; adjust the network parameters in the original voice synthesis model based on the loss value between the training predicted voices and the corresponding voice samples to obtain the pre-trained initial voice synthesis model.
[0174] In one embodiment, the pre-training module 406 is further used to pre-train the acoustic model, and merge the pre-trained acoustic model and the vocoder to obtain the initial voice synthesis model.
[0175] Meanwhile, in order to better implement the voice processing method running on the terminal in the embodiments of the present application, the embodiments of the present application also provide a voice processing device. Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of the voice processing device provided in the embodiments of the present application. The voice processing device is integrated in the terminal. The voice processing device 500 may include a first acquisition module 501 and a synthesis module 502.
[0176] The first acquisition module 501 is used to acquire a target text and a target voice synthesis model. The target voice synthesis model is obtained by performing pruning processing on the target network module in the initial voice synthesis model according to the target performance data of the terminal corresponding to the target pronunciation object, and training the obtained voice synthesis model to be trained by using the voice data of the target pronunciation object. The voice data has the target timbre.
[0177] The synthesis module 502 is used to perform voice synthesis processing on the target text by using the target voice synthesis model to obtain synthesized voice data with the target timbre.
[0178] In one embodiment, the synthesis module 502 is specifically configured to obtain a target phoneme sequence corresponding to the target text; perform acoustic processing on the target phoneme sequence through the target speech synthesis model to obtain target acoustic features including the target timbre; and perform speech synthesis processing on the target acoustic features by using the target speech synthesis model to obtain synthesized speech data with the target timbre.
[0179] In one embodiment, as Figure 13 shown, the speech processing device 500 further includes a request generation module 503, a second acquisition module 504, a second sending module 505, and a second receiving module 506. Among them, the generation module 503 is configured to generate a speech synthesis service request. The second acquisition module 504 is configured to acquire target performance data of the terminal. The second sending module 505 is configured to send the speech synthesis service request and the target performance data to the server, so that the server performs pruning processing on a target network module in the initial speech synthesis model according to the target performance data, and trains the to-be-trained speech synthesis model obtained by the pruning processing by using the speech data to obtain a target speech synthesis model. The second receiving module 506 is configured to receive the target speech synthesis model returned by the server.
[0180] In one embodiment, when the second acquisition module 504 executes the step of acquiring the target performance data of the terminal, it specifically executes: acquiring the target performance data of the terminal pre-stored; or executing a performance determination program to determine the target performance data of the terminal.
[0181] In one embodiment, when the second acquisition module 504 executes the step of executing the performance determination program to determine the target performance data of the terminal, it specifically executes: acquiring other application programs currently running on the terminal, and pausing or closing the other application programs; executing the performance determination program to determine the target performance data of the terminal; and starting the other application programs.
[0182] In one embodiment, when the first acquisition module 501 executes the step of acquiring the target speech synthesis model, it specifically executes: acquiring parameter change data in the target speech synthesis model and a target computational graph corresponding to the target network module, where the parameter change data and the target computational graph are determined and sent by the server.
[0183] The second generation module 507 is configured to generate a target speech synthesis model matching the terminal according to the parameter change data and the target computational graph.
[0184] All the above technical solutions can be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated herein one by one.
[0185] Correspondingly, an embodiment of the present application further provides a computer device, which may be a terminal or a server. As Figure 14 shown, Figure 14 is a schematic structural diagram of the computer device provided by the embodiment of the present application. The computer device 600 includes a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, and a computer program stored on the memory 602 and executable on the processor. Among them, the processor 601 is electrically connected to the memory 602. Those skilled in the art can understand that the structural diagram of the computer device shown in the figure does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0186] The processor 601 is the control center of the computer device 600, connecting various parts of the entire computer device 600 through various interfaces and lines. By running or loading software programs (computer programs) and / or modules stored in the memory 602, and calling data stored in the memory 602, it executes various functions of the computer device 600 and processes data, thereby monitoring the entire computer device 600.
[0187] In the embodiment of the present application, the processor 601 in the computer device 600 will load the instructions corresponding to the processes of one or more application programs into the memory 602 according to the following steps, and the processor 601 will run the application programs stored in the memory 602 to implement various functions:
[0188] According to the speech synthesis service request from the terminal, determine the target performance data of the terminal and the speech data of the target pronunciation object, and the speech data has a target timbre; according to the target performance data, perform pruning processing on the target network module in the initial speech synthesis model to obtain a speech synthesis model to be trained; based on the speech data, train the speech synthesis model to be trained to obtain a target speech synthesis model matching the terminal, so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with the target timbre. Or
[0189] Obtain a target text and a target speech synthesis model, where the target speech synthesis model is obtained by performing pruning processing on the target network module in the initial speech synthesis model according to the target performance data of the terminal corresponding to the target pronunciation object, and training the pruned speech synthesis model to be trained with the speech data of the target pronunciation object, and the speech data has a target timbre; use the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre.
[0190] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0191] Optionally, as Figure 14 shown, the computer device 600 further includes: a touch display screen 603, a radio frequency circuit 604, an audio circuit 605, an input unit 606, and a power supply 607. Among them, the processor 601 is electrically connected to the touch display screen 603, the radio frequency circuit 604, the audio circuit 605, the input unit 606, and the power supply 607 respectively. Those skilled in the art can understand that Figure 14 the structure of the computer device shown in
[0192] The touch display screen 603 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The touch display screen 603 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits it to the processor 601 to determine the type of touch event. Subsequently, the processor 601 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 603 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 603 can also be used as a part of the input unit 606 to implement the input function.
[0193] In the embodiments of the present application, the touch display screen 603 is used to present a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface.
[0194] The radio frequency circuit 604 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other computer devices, and to send and receive signals between the network device or other computer devices.
[0195] The audio circuit 605 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 605 can transmit the electrical signal converted from the received audio data to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 605 and then converted into audio data. After the audio data is output to the processor 601 for processing, it is sent through the radio frequency circuit 604 to, for example, another computer device, or the audio data is output to the memory 602 for further processing. The audio circuit 605 may also include an earphone jack to provide communication between a peripheral earphone and the computer device.
[0196] The input unit 606 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0197] The power supply 607 is used to supply power to each component of the computer device 600. Optionally, the power supply 607 can be logically connected to the processor 601 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 607 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0198] Although Figure 14 not shown in the figure, the computer device 600 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0199] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0200] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0201] To this end, an embodiment of the present application provides a computer-readable storage medium storing multiple computer programs that can be loaded by a processor to execute the steps in any voice processing method running on the server side or any voice processing method running on the terminal provided by the embodiments of the present application. For example, the computer program can execute the following steps:
[0202] According to a voice synthesis service request from a terminal, determine the target performance data of the terminal and the voice data of the target pronunciation object, where the voice data has a target timbre; according to the target performance data, perform pruning processing on the target network module in the initial voice synthesis model to obtain a to-be-trained voice synthesis model; based on the voice data, train the to-be-trained voice synthesis model to obtain a target voice synthesis model matching the terminal, so that the terminal processes a target text based on the target voice synthesis model to obtain synthesized voice data with the target timbre. Or
[0203] Obtain a target text and a target voice synthesis model, where the target voice synthesis model is obtained by performing pruning processing on the target network module in the initial voice synthesis model according to the target performance data of the terminal corresponding to the target pronunciation object, and training the to-be-trained voice synthesis model obtained by the pruning processing using the voice data of the target pronunciation object, where the voice data has a target timbre; use the target voice synthesis model to perform voice synthesis processing on the target text to obtain synthesized voice data with the target timbre.
[0204] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0205] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0206] Since the computer program stored in the storage medium can execute the steps in any voice processing method provided by the embodiments of the present application, the beneficial effects achievable by any voice processing method provided by the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments and will not be elaborated here.
[0207] The above has introduced in detail a voice processing method, device, storage medium, and computer device provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A voice processing method, characterized in that, Including: Obtain a target text and a target speech synthesis model, where the target speech synthesis model is obtained by pruning a target network module in an initial speech synthesis model according to target performance data of a terminal corresponding to a target pronunciation object, and training a to-be-trained speech synthesis model obtained by the pruning process with speech data of the target pronunciation object, and the speech data has a target timbre; Use the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre; Among them, the to-be-trained speech synthesis model is obtained by the following method: Determine the target number of layers of the target network module in the initial speech synthesis model according to the target performance data; Prune the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain a to-be-trained speech synthesis model.
2. The voice processing method according to claim 1, wherein Before the step of obtaining the target text, it further includes: Generate a speech synthesis service request and obtain the target performance data of the terminal; Send the speech synthesis service request and the target performance data to a server, so that the server prunes the target network module in the initial speech synthesis model according to the target performance data, trains the to-be-trained speech synthesis model obtained by the pruning process with the speech data to obtain a target speech synthesis model, and receive the target speech synthesis model returned by the server.
3. The voice processing method according to claim 2, wherein The step of obtaining the target performance data of the terminal includes: Obtain the pre-saved target performance data of the terminal; or Execute a performance determination program to determine the target performance data of the terminal.
4. The voice processing method according to claim 3, wherein The step of executing the performance determination program to determine the target performance data of the terminal includes: Obtain other applications currently running on the terminal and pause or close the other applications; Execute a performance determination program to determine the target performance data of the terminal; Start the other applications.
5. The voice processing method according to claim 1, characterized in that The step of obtaining the target speech synthesis model includes: Obtain parameter change data in the target speech synthesis model and a target computational graph corresponding to the target network module, where the parameter change data and the target computational graph are determined and sent by the server; The speech processing method further includes: generating a target speech synthesis model matching the terminal according to the parameter change data and the target computational graph.
6. The voice processing method according to claim 5, wherein The step of generating a target speech synthesis model matching the terminal according to the parameter change data and the target computational graph includes: Obtain full-scale parameter data and a full-scale computational graph saved on the terminal, where the full-scale parameter data and the full-scale computational graph are parameter data and computational graphs corresponding to the initial speech synthesis model; Update the full-scale parameter data and the full-scale computational graph according to the parameter change data and the target computational graph to generate a target speech synthesis model matching the terminal.
7. The voice processing method according to claim 1, wherein The step of using the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre includes: Obtain a target phoneme sequence corresponding to the target text; Acoustically process the target phoneme sequence through the target speech synthesis model to obtain target acoustic features including the target timbre; Perform speech synthesis processing on the target acoustic features using the target speech synthesis model to obtain synthesized speech data with the target timbre.
8. A voice processing method, characterized in that Includes: According to a speech synthesis service request from a terminal, determine the target performance data of the terminal and the speech data of the target pronunciation object, the speech data having a target timbre; According to the target performance data, perform pruning processing on the target network module in the initial speech synthesis model to obtain a speech synthesis model to be trained; Based on the speech data, train the speech synthesis model to be trained to obtain the target speech synthesis model matched by the terminal, so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with the target speech timbre; Among them, the step of performing pruning processing on the target network module in the pre-trained initial speech synthesis model according to the target performance data to obtain a speech synthesis model to be trained includes: Determine the target number of layers of the target network module in the initial speech synthesis model according to the target performance data; Prune the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain a speech synthesis model to be trained.
9. The voice processing method according to claim 8, wherein The step of determining the target number of layers of the target network module in the initial speech synthesis model according to the target performance data includes: Obtain the total complexity information corresponding to the initial speech synthesis model, the module complexity information corresponding to the target network module, the number of layers corresponding to the target network module, and the performance data of the terminal matched with the initial speech synthesis model; According to the total complexity information, the module complexity information, the number of layers, the performance data, and the target performance data, determine the target number of layers of the target network module in the initial speech synthesis model.
10. The voice processing method according to claim 8, wherein The step of training the speech synthesis model to be trained based on the speech data includes: Determine the update rounds of the speech synthesis model to be trained according to the target number of layers; Train the speech synthesis model to be trained according to the speech data until the number of training rounds reaches the update rounds to obtain the target speech synthesis model.
11. The voice processing method according to claim 8, wherein The pre-training of the initial speech synthesis model includes the following steps: Obtain speech samples of multiple pronunciation objects and the original speech synthesis model, and the multiple pronunciation objects have different timbres; Input the speech samples of multiple pronunciation objects into the original speech synthesis model, and perform pruning processing on the network module in the original speech synthesis model, so as to acoustically process the speech samples using the pruned original speech synthesis model to obtain training acoustic features; Perform speech synthesis processing on the training acoustic features using the original speech synthesis model to obtain training predicted speech; Based on the loss value between the training predicted speech and the corresponding speech sample, adjust the network parameters in the original speech synthesis model to obtain the pre-trained initial speech synthesis model.
12. The voice processing method according to claim 8, wherein The initial speech synthesis model includes an acoustic model, and the pre-training of the acoustic model includes the following steps: Obtain speech samples and an acoustic model of multiple pronunciation objects, where the multiple pronunciation objects have different timbres; Input the speech samples into the acoustic model, perform pruning processing on the network modules in the acoustic model, and use the pruned acoustic model to perform acoustic feature processing on the speech samples to obtain first training acoustic features; Perform audio signal processing on the speech samples to obtain second training acoustic features; Based on the loss values of the first training acoustic features and the second training acoustic features, update the parameter data of the neural network in the acoustic model to obtain a pre-trained acoustic model.
13. A voice processing device, characterized in that, Including: An acquisition module, configured to acquire a target text and a target speech synthesis model. The target speech synthesis model is obtained by performing pruning processing on the target network module in the pre-trained initial speech synthesis model according to the target performance data of the terminal corresponding to the target pronunciation object, and training the to-be-trained speech synthesis model obtained by the pruning processing using the speech data of the target pronunciation object, where the speech data has a target timbre; A synthesis module, configured to use the target speech synthesis model to perform speech synthesis processing on the target text to obtain synthesized speech data with the target timbre; Wherein, the to-be-trained speech synthesis model is obtained by the following method: Determine the target number of layers of the target network module in the initial speech synthesis model according to the target performance data; Prune the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain a to-be-trained speech synthesis model.
14. A voice processing device, characterized in that, Including: A first receiving module, configured to receive a speech synthesis service request from a terminal; A first determination module, configured to determine the target performance data of the terminal and the speech data of the target pronunciation object according to the speech synthesis service request from the terminal, where the speech data has a target timbre; A pruning module, configured to perform pruning processing on the target network module in the pre-trained initial speech synthesis model according to the target performance data to obtain a to-be-trained speech synthesis model; A training module, configured to train the to-be-trained speech synthesis model based on the speech data to obtain a target speech synthesis model matched to the terminal, so that the terminal processes the target text based on the target speech synthesis model to obtain synthesized speech data with a target voice color; Wherein, the pruning module is configured to: Determine the target number of layers of the target network module in the initial speech synthesis model according to the target performance data; Prune the number of layers of the target network module in the initial speech synthesis model to the target number of layers to obtain a to-be-trained speech synthesis model.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the speech processing method according to any one of claims 1-7 or execute the steps in the speech processing method according to any one of claims 8-12.
16. A computer device, characterized in that, The computer device includes a memory and a processor. A computer program is stored in the memory. The processor executes the steps in the voice processing method according to any one of claims 1-7 or executes the steps in the voice processing method according to any one of claims 8-12 by calling the computer program stored in the memory.
Citation Information
Patent Citations
End-to-end speech synthesis method based on WaveRNN
CN110473515A
Speech synthesis method, related equipment, device thereof and medium
CN113488020A