Training Method, Device, Equipment, Medium and Program Product of Vocoder Model

By generating a simulated audio training vocoder model, the problem of insufficient audio coverage of real people is solved, and efficient speech synthesis effect is achieved in a wide range of categories.

CN115223538BActive Publication Date: 2025-07-25SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210824486.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-07-25
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

In the prior art, the vocoder model requires a large amount of real human voice audio for training, but real human voice audio is difficult to cover all possible vocal categories, resulting in poor training results and synthetic voice distortion.

Method used

By obtaining at least two sample acoustic features of the sample audio, simulated acoustic features are generated and fusion, the voice coder model is used for speech synthesis processing, the training loss is calculated and the model parameters are updated, and the generated simulated audio is used for training.

Benefits of technology

No need for a large amount of real audio training, the simulated audio categories cover a wide range of and easy to obtain, improving the training effect of the vocoder model and ensuring good speech synthesis quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223538B_ABST
    Figure CN115223538B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment, medium and program product for a vocoder model, belonging to the technical field of machine learning. It includes: obtaining at least two sample acoustic features corresponding to a sample audio; performing simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; fusing the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio; using the vocoder model to perform speech synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio; calculating the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; and updating the model parameters of the vocoder model according to the training loss. Through the above method, it is not necessary to use a large number of real audios for training, and the category coverage range of the simulated audio is wide, so that the vocoder model has a good speech synthesis effect for audios in a very wide audio domain range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of machine learning, and particularly to a method, device, equipment, medium and program product for training a vocoder model. Background Art

[0002] A vocoder model refers to a model that restores the feature parameters of acoustic features transmitted through an information channel into the original speech waveform. The quality of the vocoder model directly determines the quality of the finally synthesized speech signal.

[0003] In the related art, taking human voice audio as an example, by using a large number of real human voice audios, the real human voice audios are disassembled into Mel-spectrum features, and the disassembled Mel-spectrum features are subjected to speech synthesis through a vocoder model to obtain synthesized human voice audios. Based on the real human voice audios and the synthesized human voice audios, the value of the loss function is calculated, and then the model parameters of the vocoder model are updated. After multiple trainings, a trained vocoder model is obtained.

[0004] In the above related art, a large number of real human voice audios are required to train the vocoder model. However, it is difficult for the obtained real human voice audios to cover all possible human voice categories, resulting in poor training effects of the vocoder model and easy distortion of the synthesized human voice audios by the vocoder model. Summary of the Invention

[0005] The present application provides a method, device, equipment, medium and program product for training a vocoder model, which can improve the speech synthesis performance of the vocoder model. The technical solution is as follows:

[0006] According to one aspect of the present application, a method for training a vocoder model is provided. The method includes:

[0007] Obtain at least two sample acoustic features corresponding to the sample audio;

[0008] Perform simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features;

[0009] Fuse the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio;

[0010] Use the vocoder model to perform speech synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio;

[0011] Calculate the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio;

[0012] Update the model parameters of the vocoder model according to the training loss.

[0013] According to one aspect of the present application, a voice synthesis method is provided. The method includes:

[0014] Obtain the Mel spectrum features corresponding to the target audio;

[0015] Input the Mel spectrum features into a vocoder model for voice synthesis processing to obtain a target synthesized audio corresponding to the target audio, where the vocoder model is the vocoder model as described above.

[0016] According to one aspect of the present application, a training device for a vocoder model is provided. The device includes:

[0017] An acquisition module, configured to acquire at least two sample acoustic features corresponding to a sample audio;

[0018] A generation module, configured to perform simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features;

[0019] A fusion module, configured to fuse the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio;

[0020] A voice synthesis module, configured to perform voice synthesis processing on the simulated audio by using the vocoder model to generate a simulated synthesized audio corresponding to the simulated audio;

[0021] A calculation module, configured to calculate a training loss of the vocoder model based on the simulated audio and the simulated synthesized audio;

[0022] An update module, configured to update model parameters of the vocoder model according to the training loss.

[0023] According to one aspect of the present application, a voice synthesis device is provided. The device includes: an acquisition module, configured to acquire Mel spectrum features corresponding to a target audio;

[0024] A voice synthesis module, configured to input the Mel spectrum features into a vocoder model for voice synthesis processing to obtain a target synthesized audio corresponding to the target audio, where the vocoder model is the vocoder model as described above.

[0025] According to another aspect of the present application, a computer device is provided. The computer device includes: a processor and a memory. At least one computer program is stored in the memory, and at least one computer program is loaded and executed by the processor to implement the training method of the vocoder model or the voice synthesis method as described in the above aspects.

[0026] According to another aspect of the present application, there is provided a computer storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the training method of the vocoder model or the voice synthesis method described in the above aspect.

[0027] According to another aspect of the present application, there is provided a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device from the computer-readable storage medium, so that the computer device executes the training method of the vocoder model or the voice synthesis method described in the above aspect.

[0028] The beneficial effects brought by the technical solution provided by the present application at least include:

[0029] By obtaining at least two sample acoustic features corresponding to a sample audio; performing simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; fusing the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio; using a vocoder model to perform voice synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio; calculating the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; and updating the model parameters of the vocoder model according to the training loss. The training method of the vocoder model provided by the present application can directly use the generated simulated audio for training without using a large number of real audios for training, and the category coverage range of the simulated audio is wide, the quantity is sufficient, and it is easy to obtain. Based on this, the training effect of the vocoder model can be improved, so that the vocoder model has a good voice synthesis effect for audios in a very wide audio domain range. Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0031] Figure 1 It is a schematic diagram of a method for training a vocoder model provided by an exemplary embodiment of the present application;

[0032] Figure 2 It is a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;

[0033] Figure 3It is a flowchart of a method for training a vocoder model provided by an exemplary embodiment of the present application;

[0034] Figure 4 It is a flowchart of a method for training a vocoder model provided by an exemplary embodiment of the present application;

[0035] Figure 5 It is a schematic diagram of an application scenario of a vocoder model provided by an exemplary embodiment of the present application;

[0036] Figure 6 It is a schematic diagram of a sample acoustic feature provided by an exemplary embodiment of the present application;

[0037] Figure 7 It is a schematic diagram of a method for training a vocoder model provided by an exemplary embodiment of the present application;

[0038] Figure 8 It is a framework diagram of the generation of a training system for a vocoder model and the training of the vocoder model provided by an exemplary embodiment of the present application;

[0039] Figure 9 It is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application;

[0040] Figure 10 It is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application;

[0041] Figure 11 It is a block diagram of a training device for a vocoder model provided by an exemplary embodiment of the present application;

[0042] Figure 12 It is a block diagram of a speech synthesis device provided by an exemplary embodiment of the present application;

[0043] Figure 13 It is a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0044] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0045] The embodiments of the present application provide a method for training a vocoder model. For ease of understanding, the following explains several terms related to the present application.

[0046] 1) Artificial Intelligence (AI)

[0047] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0048] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The display device with an image acquisition component shown in this application is mainly related to computer vision technology, machine learning / deep learning, autonomous driving, intelligent transportation, and other directions.

[0049] 2) Machine Learning (ML)

[0050] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0051] 3) Computer Vision Technology (CV)

[0052] Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace the human eye for target recognition and measurement, and further performs graphic processing to make the computer process into images that are more suitable for human eye observation or transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as the training of vocoder models, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and intelligent transportation.

[0053] The embodiments of the present application provide a technical solution for a method of training a vocoder model, as Figure 1 shown. This method can be executed by a computer device, and the computer device can be a terminal or a server.

[0054] Exemplarily, the computer device obtains a sample audio 101 and decomposes the sample audio 101 into at least two sample acoustic features through an encoding network 102.

[0055] Optionally, the sample acoustic features include at least two of a harmonic distribution feature 103, a harmonic amplitude feature 104, a fundamental frequency feature 105, and a time-varying filtered noise feature 106, but are not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0056] Exemplarily, the computer device performs simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; the computer device fuses the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio 115; the computer device uses a vocoder model 116 to perform speech synthesis processing on the simulated audio 115 to generate a simulated synthesized audio 117 corresponding to the simulated audio 115; the computer device calculates the training loss of the vocoder model 116 based on the simulated audio 115 and the simulated synthesized audio 117; the computer device updates the model parameters of the vocoder model 116 according to the training loss.

[0057] The simulated acoustic feature refers to an acoustic feature that has the same or similar change pattern as the sample acoustic feature.

[0058] In a possible implementation, after obtaining at least two sample acoustic features corresponding to the sample audio, the computer device determines the change pattern corresponding to the sample acoustic features through the feature extraction network 107, and constructs an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network 108, so as to obtain the simulated acoustic feature corresponding to the sample acoustic features.

[0059] The change pattern is used to represent the change trend of the sample acoustic features. For example, spikes periodically appear in the sample acoustic features, and the amplitude values of the spikes show periodic changes.

[0060] As Figure 1 shown, the computer device decomposes the sample audio 101 into a harmonic distribution feature 103, a harmonic amplitude feature 104, a fundamental frequency feature 105, and a time-varying filtered noise feature 106 through the encoding network 102; the computer device determines the change trend of the corresponding sample acoustic features through the feature extraction network 107, and constructs an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network 108, that is, obtains a simulated harmonic distribution feature 109, a simulated harmonic amplitude feature 110, a simulated fundamental frequency feature 111, and a simulated time-varying filtered noise feature 112.

[0061] In a possible implementation, the computer device performs a fusion process on the simulated acoustic features corresponding to different sample acoustic features to obtain a fused simulated audio; the computer device performs a decoding process on the fused simulated audio through the decoding network 114 to generate a simulated audio 115.

[0062] Optionally, at least one simulated acoustic feature corresponds to one sample acoustic feature. The computer device freely combines the simulated acoustic features corresponding to different sample acoustic features and performs a fusion process to obtain a fused simulated audio; the computer device performs a decoding process on the fused simulated audio through the decoding network 114 to generate a simulated audio 115.

[0063] In a possible implementation, the computer device decomposes the simulated audio 115 into Mel spectrum features; the computer device uses a vocoder model 116 to perform speech synthesis processing on the Mel spectrum features to generate a simulated synthesized audio 117 corresponding to the simulated audio 115.

[0064] Exemplarily, the computer device calculates the training loss of the vocoder model 116 based on the simulated audio 115 and the simulated synthesized audio 117; and updates the model parameters of the vocoder model 116 according to the training loss.

[0065] In summary, the method provided in this embodiment obtains at least two sample acoustic features corresponding to a sample audio; performs simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; fuses the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio; uses a vocoder model to perform speech synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio; calculates the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; and updates the model parameters of the vocoder model according to the training loss. The training method for the vocoder model provided in this application can directly use the generated simulated audio for training without using a large number of real audios for training, and the simulated audio has a wide coverage range of categories, sufficient quantity, and is easy to obtain. Based on this, the training effect of the vocoder model can be improved, so that the vocoder model has a good speech synthesis effect for audios in a very wide audio domain range.

[0066] Figure 2 FIG. shows a schematic architecture diagram of a computer system provided in an embodiment of this application. The computer system may include: a terminal 100 and a server 200.

[0067] The terminal 100 may be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a personal computer (PC), a smart voice interaction device, a smart home appliance, an in-vehicle terminal, an aircraft, an unmanned vending terminal, etc. A client of a target application program may be installed and run in the terminal 100. The target application program may be an application program for training a reference vocoder model or other application programs providing a training function for the vocoder model. This application does not make any limitations in this regard. In addition, this application does not make any limitations on the form of the target application program, including but not limited to an application program (App), a small program, etc. installed in the terminal 100, and may also be in the form of a web page.

[0068] The server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a content delivery network (CDN), and a cloud server of basic cloud computing services such as a big data and artificial intelligence platform. The server 200 may be the background server of the above target application program for providing background services for the client of the target application program.

[0069] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backing, which can only be achieved through cloud computing.

[0070] In some embodiments, the above server can also be implemented as a node in a blockchain system. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0071] Communication can be carried out between the terminal 100 and the server 200 through a network, such as a wired or wireless network.

[0072] For the training method or voice synthesis method of the vocoder model provided by the embodiments of the present application, the execution subject of each step can be a computer device, and a computer device refers to an electronic device with data computing, processing, and storage capabilities. Taking Figure 2 the implementation environment of the shown solution as an example, the training method or voice synthesis method of the vocoder model can be executed by the terminal 100 (such as the client of the target application installed and running in the terminal 100 executing the training method or voice synthesis method of the vocoder model), or can be executed by the server 200, or can be executed by the interaction and cooperation between the terminal 100 and the server 200. The present application does not make any limitations in this regard.

[0073] Figure 3 is a flowchart of the training method of the vocoder model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be Figure 2 the terminal 100 or the server 200 in

[0074] Step 302: Obtain at least two types of sample acoustic features corresponding to the sample audio.

[0075] Exemplarily, an acoustic feature refers to a physical quantity representing the acoustic characteristics of an audio, and is also a general term for the acoustic manifestations of various elements of the audio.

[0076] Optionally, the sample acoustic features include at least two of the harmonic distribution feature, the harmonic amplitude feature, the fundamental frequency feature, and the time-varying filtered noise feature, but are not limited thereto. The embodiments of the present application do not make specific limitations in this regard. In the present application, the above four acoustic features are taken as examples for description.

[0077] Among them, the method for obtaining the sample audio includes at least one of the following situations:

[0078] 1. The computer device receives the sample audio. For example, the terminal is the terminal that initiates audio recording. The audio is recorded through the terminal, and after the recording is completed, the audio is used as the sample audio.

[0079] 2. The computer device obtains the sample audio from a pre-stored database.

[0080] It should be noted that the above methods for obtaining the sample audio are only illustrative examples, and the embodiments of the present application do not limit this.

[0081] Step 304: Perform simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features.

[0082] The simulated acoustic feature refers to an acoustic feature simulated by a computer device.

[0083] The acoustic feature refers to a feature obtained by decomposing an audio.

[0084] Optionally, the simulated acoustic features include at least two of the simulated harmonic distribution feature, the simulated harmonic amplitude feature, the simulated fundamental frequency feature, and the simulated time-varying filtered noise feature, but are not limited thereto. The embodiments of the present application do not make specific limitations in this regard.

[0085] Exemplarily, the computer device performs simulated acoustic feature generation processing based on the obtained sample acoustic features to obtain simulated acoustic features corresponding to the respective sample acoustic features.

[0086] For example, the computer device performs simulated acoustic feature generation processing based on the harmonic distribution feature to obtain the simulated harmonic distribution feature corresponding to the harmonic distribution feature.

[0087] The computer device performs simulated acoustic feature generation processing based on the harmonic amplitude feature to obtain the simulated harmonic amplitude feature corresponding to the harmonic amplitude feature.

[0088] The computer device performs analog acoustic feature generation processing based on the fundamental frequency feature to obtain the analog fundamental frequency feature corresponding to the fundamental frequency feature.

[0089] The computer device performs analog acoustic feature generation processing based on the time-varying filtered noise feature to obtain the analog time-varying filtered noise feature corresponding to the time-varying filtered noise feature.

[0090] Step 306: Fuse the analog acoustic features corresponding to different sample acoustic features to obtain an analog audio.

[0091] Exemplarily, in the case of obtaining the analog acoustic features corresponding to different sample acoustic features, the computer device fuses the analog acoustic features corresponding to different sample acoustic features to obtain the corresponding analog audio.

[0092] For example, the computer device obtains the analog harmonic distribution feature corresponding to the harmonic distribution feature, the analog harmonic amplitude feature corresponding to the harmonic amplitude feature, the analog fundamental frequency feature corresponding to the fundamental frequency feature, and the analog time-varying filtered noise feature corresponding to the time-varying filtered noise feature. The computer device fuses the analog harmonic distribution feature, the analog harmonic amplitude feature, the analog fundamental frequency feature, and the analog time-varying filtered noise feature to obtain the corresponding analog audio; or, the computer device fuses at least two of the analog harmonic distribution feature, the analog harmonic amplitude feature, the analog fundamental frequency feature, and the analog time-varying filtered noise feature to obtain the corresponding analog audio.

[0093] Step 308: Use a vocoder model to perform speech synthesis processing on the analog audio to generate an analog synthesized audio corresponding to the analog audio.

[0094] A vocoder model refers to a model that restores the feature parameters of the acoustic features transmitted through the information channel into the original speech waveform.

[0095] An analog synthesized audio refers to an audio obtained through speech synthesis processing using a vocoder model.

[0096] Exemplarily, the computer device uses a vocoder model to perform speech synthesis processing on the analog audio to generate an analog synthesized audio corresponding to the analog audio.

[0097] Step 310: Calculate the training loss of the vocoder model based on the analog audio and the analog synthesized audio.

[0098] Exemplarily, the computer device calculates the training loss of the vocoder model based on the analog audio and the analog synthesized audio.

[0099] The training loss refers to the difference value between the input and output of the vocoder model, and the performance of the vocoder model is measured by the training loss.

[0100] Step 312: Update the model parameters of the vocoder model according to the training loss.

[0101] Exemplarily, the computer device updates the model parameters of the vocoder model according to the training loss.

[0102] Updating the model parameters refers to updating the network parameters in the vocoder model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but is not limited thereto, and the embodiments of the present application do not make limitations in this regard.

[0103] It can be understood that taking human voice audio as an example, the sample audio uses real human voice audio. By obtaining the acoustic features of the real human voice audio and performing simulated acoustic feature generation processing on the acoustic features, the simulated acoustic features are obtained. The computer device fuses the simulated acoustic features corresponding to different sample acoustic features to finally obtain the simulated audio. By extracting the acoustic features of the real human voice audio and imitating the acoustic features, the obtained simulated audio is closer to the real human voice audio, thus solving the problem that it is difficult for real human voice audio to cover all possible human voice categories caused by training with a large number of real audio, and the category coverage of the simulated audio is wide, easy to obtain, and sufficient in quantity, greatly reducing the consumption of manpower, material resources, and financial resources brought about by collecting real human voice audio.

[0104] It can be understood that the training method of the vocoder model proposed in the embodiments of the present application can be applied to any learning-based vocoder.

[0105] In summary, the method provided in this embodiment includes: obtaining at least two sample acoustic features corresponding to the sample audio; performing simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; fusing the simulated acoustic features corresponding to different sample acoustic features to obtain the simulated audio; using the vocoder model to perform speech synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio; calculating the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; and updating the model parameters of the vocoder model according to the training loss. The training method of the vocoder model provided in the present application can directly use the generated simulated audio for training without using a large number of real audio for training, and the category coverage of the simulated audio is wide, the quantity is sufficient, and it is easy to obtain. Based on this, the training effect of the vocoder model can be improved, so that the vocoder model has a good speech synthesis effect for audio with a very wide audio domain range.

[0106] An embodiment of the present application provides a training system for a vocoder model. The training system for the vocoder model runs a vocoder model and a simulated audio generation model; the simulated audio generation model includes an encoding network, a feature extraction network, a simulated acoustic feature generation network, and a decoding network.

[0107] The computer device obtains a sample audio and decomposes the sample audio into at least two sample acoustic features through the encoding network.

[0108] The computer device performs simulated acoustic feature generation processing on the sample acoustic features through the simulated acoustic feature generation network to obtain simulated acoustic features corresponding to the sample acoustic features; the computer device fuses the simulated acoustic features corresponding to different sample acoustic features through the decoding network to obtain a simulated audio; the computer device performs speech synthesis processing on the simulated audio using the vocoder model to generate a simulated synthesized audio corresponding to the simulated audio; the computer device calculates the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; the computer device updates the model parameters of the vocoder model according to the training loss.

[0109] Based on the training system for the vocoder model, the following training method for the vocoder model is provided.

[0110] Figure 4 It is a flowchart of the training method for the vocoder model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be Figure 2 the terminal 100 or the server 200 in

[0111] Step 402: Obtain at least two sample acoustic features corresponding to the sample audio.

[0112] Exemplarily, the training system for the vocoder model runs a vocoder model and a simulated audio generation model.

[0113] The simulated audio generation model refers to a model for generating simulated audio based on sample audio.

[0114] Optionally, the simulated audio generation model includes an encoding network, a feature extraction network, a simulated acoustic feature generation network, and a decoding network.

[0115] Exemplarily, the computer device obtains a sample audio and decomposes the sample audio into at least two sample acoustic features through the encoding network.

[0116] The formula for the encoding network to decompose the sample audio into at least two sample acoustic features can be expressed as:

[0117] Enc(x)=H amp ,H dist ,H f ,N

[0118] In the formula, Enc() represents the encoding network, x is the sample audio, and H dist is the harmonic distribution feature, H amp is the harmonic amplitude feature, H f is the fundamental frequency feature, and N is the time-varying filtered noise feature.

[0119] The acoustic feature refers to the physical quantity representing the acoustic characteristics of the audio, and is also a general term for the acoustic manifestations of various elements of the audio.

[0120] Optionally, the sample acoustic features include at least two of the harmonic distribution feature, the harmonic amplitude feature, the fundamental frequency feature, and the time-varying filtered noise feature, but are not limited thereto. The embodiments of the present application do not make specific limitations in this regard. In the present application, the above four acoustic features are taken as examples for description.

[0121] Exemplarily, the vocoder model is applicable to at least one of different speakers, genders, accents, singing scenarios, or instrument scenarios, but is not limited thereto. The embodiments of the present application do not make specific limitations on the application scope of the vocoder.

[0122] As Figure 5 shown in the schematic diagram of the application scenario of the vocoder model, the vocoder model 501 can be applied to scenarios such as text-to-speech conversion, singing synthesis, voice conversion, cross-lingual speech translation, etc. The trained vocoder model 501 converts the acoustic features or spectral features into the final synthesized audio 502. In addition, due to the generality of the trained vocoder model 501, it can be used under diverse conditions, including different speakers, genders, accents, singing, instruments, etc.

[0123] Step 404: Determine the change pattern corresponding to the sample acoustic features through the feature extraction network, and construct acoustic features with the same change pattern as the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features.

[0124] The simulated acoustic feature generation network refers to a network that constructs acoustic features with the same change pattern as the corresponding sample acoustic features based on the sample acoustic features.

[0125] The change pattern refers to the change trend of the sample acoustic features.

[0126] For example, Figure 6A schematic diagram showing sample acoustic features is presented. Taking the harmonic amplitude feature as an example, the abscissa in the figure represents the number of points, and the ordinate represents the amplitude value. In a computer device, audio exists as discrete points, and the number of points represents the number of points that appear within one second. The computer device determines the change pattern corresponding to the harmonic amplitude feature. For example, the amplitude value is relatively stable and the amplitude value is 0 in the interval from 20 points to 50 points, and the amplitude value corresponding to 60 points is 0.46, etc. In the embodiments of the present application, the change pattern corresponding to the harmonic amplitude feature is determined by the amplitude value and the flatness, but is not limited thereto. The embodiments of the present application do not make specific limitations on the determination method of the change pattern.

[0127] Exemplarily, the computer device determines the change pattern corresponding to the sample acoustic feature through a feature extraction network, constructs an acoustic feature identical to the change pattern corresponding to the sample acoustic feature through a simulated acoustic feature generation network, and obtains the simulated acoustic feature corresponding to the sample acoustic feature.

[0128] Optionally, the computer device constructs an acoustic feature identical to the first change pattern corresponding to the harmonic distribution feature through a simulated acoustic feature generation network, and obtains the simulated harmonic distribution feature corresponding to the harmonic distribution feature.

[0129] The first change pattern is used to represent the change trend corresponding to the harmonic distribution feature. The computer device obtains the first change pattern corresponding to the harmonic distribution feature by analyzing the change trend corresponding to the harmonic distribution feature.

[0130] Optionally, the computer device constructs an acoustic feature identical to the second change pattern corresponding to the harmonic amplitude feature through a simulated acoustic feature generation network, and obtains the simulated harmonic amplitude feature corresponding to the harmonic amplitude feature.

[0131] The second change pattern is used to represent the change trend corresponding to the harmonic amplitude feature. The computer device determines the change pattern corresponding to the harmonic amplitude feature by analyzing the change trend corresponding to the harmonic amplitude feature, using the amplitude value and the flatness, so as to obtain the second change pattern corresponding to the harmonic amplitude feature.

[0132] Optionally, the computer device constructs an acoustic feature identical to the third change pattern corresponding to the fundamental frequency feature through a simulated acoustic feature generation network, and obtains the simulated fundamental frequency feature corresponding to the fundamental frequency feature.

[0133] The third change pattern is used to represent the change trend corresponding to the fundamental frequency feature. The computer device obtains the third change pattern corresponding to the fundamental frequency feature by analyzing the change trend corresponding to the fundamental frequency feature.

[0134] Optionally, the computer device constructs an acoustic feature identical to the fourth change pattern corresponding to the time-varying filtered noise feature through a simulated acoustic feature generation network, and obtains the simulated time-varying filtered noise feature corresponding to the time-varying filtered noise feature.

[0135] The fourth variation pattern is used to represent the variation trend corresponding to the time-varying filtered noise feature. The computer device analyzes the variation trend corresponding to the time-varying filtered noise feature to obtain the fourth variation pattern corresponding to the time-varying filtered noise feature.

[0136] It can be understood that the above harmonic distribution feature, harmonic amplitude feature, fundamental frequency feature, and time-varying filtered noise feature are only four acoustic features exemplified in the embodiments of the present application, but are not limited thereto. More acoustic features can be obtained by decomposing the sample audio, and the embodiments of the present application do not make specific limitations thereto.

[0137] It can be understood that the determination of the variation pattern corresponding to the above acoustic features can be achieved from multiple perspectives. Taking the harmonic amplitude feature as an example, the computer device determines the variation pattern corresponding to the harmonic amplitude feature with the amplitude value and flatness as consideration parameters, but is not limited thereto. The embodiments of the present application do not make specific limitations on the determination method of the variation pattern.

[0138] Step 406: Fuse the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio.

[0139] Exemplarily, in the case of obtaining the simulated acoustic features corresponding to different sample acoustic features, the computer device fuses the simulated acoustic features corresponding to different sample acoustic features to obtain the corresponding simulated audio.

[0140] In a possible implementation manner, in the case of obtaining the simulated acoustic features corresponding to different sample acoustic features, the computer device performs a fusion process on the simulated acoustic features corresponding to different sample acoustic features to obtain a fused simulated audio; the computer device performs a decoding process on the fused simulated audio through a decoding network to generate a simulated audio.

[0141] The formula for the decoding network to perform a fusion process on the simulated acoustic features to generate a simulated audio can be expressed as:

[0142] X’ = Dec(H amp ’, H dist ’, H f ’, N’)

[0143] In the formula, Dec() represents the decoding network, x’ is the simulated audio, H dist ’ is the simulated harmonic distribution feature, H amp ’ is the simulated harmonic amplitude feature, H f ’ is the simulated fundamental frequency feature, and N’ is the simulated time-varying filtered noise feature.

[0144] For example, a computer device obtains an analog harmonic distribution feature corresponding to a harmonic distribution feature, an analog harmonic amplitude feature corresponding to a harmonic amplitude feature, an analog fundamental frequency feature corresponding to a fundamental frequency feature, and an analog time-varying filtered noise feature corresponding to a time-varying filtered noise feature. The computer device fuses the analog harmonic distribution feature, the analog harmonic amplitude feature, the analog fundamental frequency feature, and the analog time-varying filtered noise feature to obtain a corresponding analog audio; alternatively, the computer device fuses at least two of the analog harmonic distribution feature, the analog harmonic amplitude feature, the analog fundamental frequency feature, and the analog time-varying filtered noise feature to obtain a corresponding analog audio.

[0145] Step 408: Use a vocoder model to perform speech synthesis processing on the analog audio to generate an analog synthesized audio corresponding to the analog audio.

[0146] A vocoder model refers to a model that restores the characteristic parameters of acoustic features transmitted through an information channel into an original speech waveform.

[0147] An analog synthesized audio refers to an audio obtained by performing speech synthesis processing through a vocoder model.

[0148] Exemplarily, a computer device uses a vocoder model to perform speech synthesis processing on the analog audio to generate an analog synthesized audio corresponding to the analog audio.

[0149] In a possible implementation, the computer device decomposes the analog audio into Mel-spectrum features; the computer device uses a vocoder model to perform speech synthesis processing on the Mel-spectrum features to generate an analog synthesized audio corresponding to the analog audio.

[0150] Step 410: Calculate the training loss of the vocoder model based on the analog audio and the analog synthesized audio.

[0151] Exemplarily, a computer device calculates the training loss of the vocoder model based on the analog audio and the analog synthesized audio.

[0152] The training loss refers to the difference value between the input and output of the vocoder model, and the performance of the vocoder model can be measured by the training loss.

[0153] Step 412: Update the model parameters of the vocoder model according to the training loss.

[0154] Exemplarily, a computer device updates the model parameters of the vocoder model according to the training loss.

[0155] Model parameter update refers to updating the network parameters in the vocoder model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but not limited to this, and the embodiments of the present application do not make any limitations in this regard.

[0156] Based on the loss function value, the loss function value is used as a training metric to update the model parameters in the vocoder model until the loss function value converges, thereby obtaining a trained vocoder model.

[0157] The convergence of the loss function value means that the loss function value no longer changes, or the error difference between two adjacent iterations during the training of the vocoder model is less than a preset value, or the number of training times of the vocoder model reaches at least one of the preset numbers, but is not limited thereto. The embodiments of the present application do not limit this.

[0158] Optionally, the target condition satisfied by the training can be that the number of training iterations of the initial model reaches the target number, and those skilled in the art can preset the number of training iterations. Or, the target condition satisfied by the training can be that the loss value satisfies the target threshold condition, such as the loss value being less than 0.00001, but is not limited thereto. The embodiments of the present application do not limit this.

[0159] In summary, the method provided in this embodiment obtains at least two sample acoustic features corresponding to the sample audio, and the sample acoustic features include at least two of the harmonic distribution feature, the harmonic amplitude feature, the fundamental frequency feature, and the time-varying filtered noise feature; performs simulated acoustic feature generation processing on the sample acoustic features based on the simulated acoustic feature generation network in the simulated audio generation model to obtain simulated acoustic features corresponding to the sample acoustic features; fuses the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio; uses the vocoder model to perform speech synthesis processing on the simulated audio to generate a simulated synthesized audio corresponding to the simulated audio; calculates the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; and updates the model parameters of the vocoder model according to the training loss. The training method of the vocoder model provided by the present application can directly use the generated simulated audio for training without using a large number of real audios for training. This training method can be applied to any vocoder model, and the simulated audio has a wide coverage range of categories, a sufficient number, and is easy to obtain. Based on this, the training effect of the vocoder model can be improved, so that the vocoder model has a good speech synthesis effect for audios in a very wide audio domain range.

[0160] Figure 7 It is a schematic diagram of a training method of a vocoder model provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be Figure 2 the terminal 100 or the server 200 in

[0161] In the communication field, the quality of the vocoder model directly determines the quality of the finally synthesized speech signal. Training the vocoder model can obtain a higher-quality speech signal. For example Figure 7As shown, the training system of the vocoder model runs the vocoder model 704 and the simulated audio generation model 702.

[0162] Optionally, the simulated audio generation model 702 includes an encoding network, a feature extraction network, a simulated acoustic feature generation network, and a decoding network.

[0163] The computer device obtains the sample audio 701 and inputs the sample audio 701 into the simulated audio generation model 702.

[0164] The computer device decomposes the sample audio 701 into at least two sample acoustic features through the encoding network in the simulated audio generation model 702.

[0165] The computer device determines the change pattern corresponding to the sample acoustic features through the feature extraction network, and constructs acoustic features identical to the change pattern corresponding to the sample acoustic features through the simulated acoustic feature generation network in the simulated audio generation model 702, obtaining the simulated acoustic features corresponding to the sample acoustic features.

[0166] The computer device fuses the simulated acoustic features corresponding to different sample acoustic features through the decoding network in the simulated audio generation model 702 to obtain at least one simulated audio 703.

[0167] The computer device inputs the obtained simulated audio 703 into the vocoder model 704 for speech synthesis processing to generate a simulated synthesized audio 705 corresponding to the simulated audio 703.

[0168] The computer device calculates the training loss of the vocoder model 704 based on the simulated audio 703 and the simulated synthesized audio 705; and updates the model parameters of the vocoder model 704 according to the training loss.

[0169] In summary, the training method of the vocoder model provided in this application can directly use the generated simulated audio for training, without the need to use a large number of real audios for training, and the category coverage of the simulated audio is wide, the quantity is sufficient, and it is easy to obtain. Based on this, the training effect of the vocoder model can be improved, so that the vocoder model has a good speech synthesis effect for audios in a very wide audio domain range.

[0170] The training method of the vocoder model involved in this application can be implemented based on the training system of the vocoder model. This solution includes the generation stage of the training system of the vocoder model and the training stage of the vocoder model. Figure 8 It is a framework diagram showing the generation of a training system for a vocoder model and the training of a vocoder model according to an exemplary embodiment of this application, as Figure 8As shown, in the generation stage of the training system of the vocoder model, after the training system generation device 810 of the vocoder model obtains the training system of the vocoder model through a pre-set training sample data set, the training result of the vocoder model is generated based on the training system of the vocoder model. In the training stage of the vocoder model, the training device 820 of the vocoder model processes the input sample audio based on the training system of the vocoder model to obtain the training result of the vocoder model.

[0171] Among them, the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can be computer devices. For example, the computer device can be a fixed computer device such as a personal computer or a server, or the computer device can also be a mobile computer device such as a tablet computer or an e-book reader.

[0172] Optionally, the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can be the same device, or the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can also be different devices. And when the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model are different devices, the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can be devices of the same type. For example, the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can both be servers; or the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model can also be devices of different types. For example, the training device 820 of the vocoder model can be a personal computer or a terminal, and the training system generation device 810 of the vocoder model can be a server, etc. The specific types of the training system generation device 810 of the vocoder model and the training device 820 of the vocoder model are not limited in the embodiments of the present application.

[0173] The above embodiments have described the training method of the vocoder model. Next, the speech synthesis method based on the vocoder model will be described. The speech synthesis method provided by the embodiments of the present application can also be applied to scenarios with a small number of samples (i.e., small samples).

[0174] Figure 9 is a flowchart of the speech synthesis method provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be Figure 2 the terminal 100 or the server 200 in

[0175] Step 902: Obtain the Mel spectrum features corresponding to the target audio.

[0176] Exemplarily, the Mel-spectrum feature is a feature for describing audio that is widely used in automatic speech and speaker recognition.

[0177] Among them, the ways to obtain the target audio include at least one of the following situations:

[0178] 1. The computer device receives the target audio. For example, the terminal is the terminal that initiates audio recording. The audio is recorded through the terminal, and after the recording ends, the audio is used as the target audio.

[0179] 2. The computer device obtains the target audio from the stored database.

[0180] It should be noted that the above ways to obtain the target audio are only illustrative examples, and the embodiments of the present application do not limit this.

[0181] Step 904: Input the Mel-spectrum feature into the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio.

[0182] The vocoder model refers to a model that restores the feature parameters of the acoustic features transmitted through the information channel into the original speech waveform.

[0183] The target synthesized audio refers to the audio obtained through speech synthesis processing by the vocoder model.

[0184] Exemplarily, the computer device performs speech synthesis processing on the Mel-spectrum feature to obtain the target synthesized audio corresponding to the target audio.

[0185] In summary, the method provided in this embodiment, by obtaining the Mel-spectrum feature corresponding to the target audio and inputting the Mel-spectrum feature into the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio, can obtain a target synthesized audio with higher quality based on the trained vocoder model.

[0186] Figure 10 is a flowchart of a speech synthesis method provided by an exemplary embodiment of the present application. This method can be executed by a computer device, and the computer device can be Figure 2 the terminal 100 or the server 200 in. This method includes:

[0187] Step 1002: Obtain the target audio and decompose the target audio into Mel-spectrum features through the encoding network.

[0188] Exemplarily, the vocoder model includes an encoding network and a speech synthesis network. The computer device obtains the target audio; the computer device decomposes the target audio into Mel-spectrum features through the encoding network.

[0189] Exemplarily, the vocoder model is applicable to at least one of different speakers, genders, accents, singing scenarios, or instrument scenarios, but is not limited thereto. The embodiments of the present application do not specifically limit the application scope of the vocoder.

[0190] Step 1004: Input the Mel spectrum features into the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio.

[0191] Exemplarily, the computer device inputs the Mel spectrum features into the speech synthesis network in the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio.

[0192] To verify the effectiveness of the training method of the vocoder model proposed in the embodiments of the present application, a control vocoder model and an improved vocoder model (i.e., the vocoder model in this solution) are selected for a comparative experiment in scenarios including different speakers, singing scenarios, and instrument scenarios, and a subjective index is used for comparison. The specific situation is as follows in Table 1:

[0193] Table 1 Comparison Table of Subjective Test Results of Vocoder Models

[0194]

[0195] In the table, (Mean Opinion Score, MOS) is the mean opinion score, and the larger the value of MOS, the better the speech synthesis effect of the vocoder model. (Confidence Interval, CI) is the confidence interval.

[0196] It can be seen from Table 1 that the vocoder model provided in the embodiments of the present application has a greater effect improvement compared to the control vocoder model, and the speech quality synthesized by the vocoder model is improved.

[0197] To verify the effectiveness of the training method of the vocoder model proposed in the embodiments of the present application, a control vocoder model and an improved vocoder model (i.e., the vocoder model in this solution) are selected for a comparative experiment in scenarios including different speakers and singing scenarios, and an objective index is used for comparison. The specific situation is as follows in Table 2:

[0198] Table 2 Comparison Table of Objective Test Results of Vocoder Models

[0199]

[0200] In the table, the larger the Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI), the better the speech synthesis effect of the vocoder model.

[0201] As can be seen from Table 2, the vocoder model provided by the embodiments of the present application has a significant improvement in effect compared to the control vocoder model, improving the speech quality synthesized by the vocoder model.

[0202] In summary, the method provided in this embodiment decomposes the target audio into Mel-spectrum features through an encoding network, and inputs the Mel-spectrum features into the speech synthesis network in the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio. Based on the trained vocoder model, a target synthesized audio with higher quality can be obtained.

[0203] Figure 11 The structural schematic diagram of the training device of the vocoder model provided by an exemplary embodiment of the present application is shown. This device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:

[0204] An acquisition module 1101, configured to acquire at least two sample acoustic features corresponding to the sample audio;

[0205] A generation module 1102, configured to perform simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features;

[0206] A fusion module 1103, configured to fuse the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio;

[0207] A speech synthesis module 1104, configured to perform speech synthesis processing on the simulated audio by using the vocoder model to generate a simulated synthesized audio corresponding to the simulated audio;

[0208] A calculation module 1105, configured to calculate the training loss of the vocoder model based on the simulated audio and the simulated synthesized audio;

[0209] An update module 1106, configured to update the model parameters of the vocoder model according to the training loss.

[0210] In a possible implementation manner, the training system of the vocoder model runs the vocoder model and a simulated audio generation model; wherein, the simulated audio generation model includes a simulated acoustic feature generation network and a feature extraction network; the generation module 1102 is configured to determine the change pattern corresponding to the sample acoustic features through the feature extraction network, and the change pattern is used to represent the change trend of the sample acoustic features; and construct acoustic features with the same change pattern as the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features.

[0211] In a possible implementation manner, the generating module 1102 is configured to construct acoustic features identical to a first variation pattern corresponding to the harmonic distribution feature through the simulated acoustic feature generation network, so as to obtain simulated harmonic distribution features corresponding to the harmonic distribution feature.

[0212] In a possible implementation manner, the generating module 1102 is configured to construct acoustic features identical to a second variation pattern corresponding to the harmonic amplitude feature through the simulated acoustic feature generation network, so as to obtain simulated harmonic amplitude features corresponding to the harmonic amplitude feature.

[0213] In a possible implementation manner, the generating module 1102 is configured to construct acoustic features identical to a third variation pattern corresponding to the fundamental frequency feature through the simulated acoustic feature generation network, so as to obtain simulated fundamental frequency features corresponding to the fundamental frequency feature.

[0214] In a possible implementation manner, the generating module 1102 is configured to construct acoustic features identical to a fourth variation pattern corresponding to the time-varying filtered noise feature through the simulated acoustic feature generation network, so as to obtain simulated time-varying filtered noise features corresponding to the time-varying filtered noise feature.

[0215] In a possible implementation manner, the fusing module 1103 is configured to perform a fusing process on the simulated acoustic features corresponding to different sample acoustic features to obtain a fused simulated audio; and perform a decoding process on the fused simulated audio through the decoding network to generate the simulated audio.

[0216] The simulated acoustic features include at least two of simulated harmonic distribution features, simulated harmonic amplitude features, simulated fundamental frequency features, and simulated time-varying filtered noise features.

[0217] In a possible implementation manner, the obtaining module 1101 is configured to obtain the sample audio and decompose the sample audio into at least two sample acoustic features through the encoding network.

[0218] The sample acoustic features include at least two of harmonic distribution features, harmonic amplitude features, fundamental frequency features, and time-varying filtered noise features.

[0219] In a possible implementation manner, the speech synthesis module 1104 is configured to decompose the simulated audio into Mel spectrum features; and perform speech synthesis processing on the Mel spectrum features by using the vocoder model to generate the simulated synthesized audio corresponding to the simulated audio.

[0220] Figure 12The structural schematic diagram of a voice synthesis device provided by an exemplary embodiment of the present application is shown. This device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:

[0221] An acquisition module 1201, configured to acquire Mel spectrum features corresponding to a target audio;

[0222] A voice synthesis module 1202, configured to input the Mel spectrum features into a vocoder model for voice synthesis processing to obtain a target synthesized audio corresponding to the target audio, where the vocoder model is the vocoder model described in the above embodiment.

[0223] In a possible implementation manner, the vocoder model includes an encoding network. The acquisition module 1201 is configured to acquire the target audio; and decompose the target audio into Mel spectrum features through the encoding network.

[0224] In a possible implementation manner, the vocoder model includes a voice synthesis network; the voice synthesis module 1202 is configured to input the Mel spectrum features into the voice synthesis network for voice synthesis processing to obtain a target synthesized audio corresponding to the target audio.

[0225] Figure 13 The structural block diagram of a computer device 1300 shown by an exemplary embodiment of the present application is shown. This computer device can be implemented as the server in the above solution of the present application. The image computer device 1300 includes a central processing unit (CPU) 1301, a system memory 1304 including a random access memory (RAM) 1302 and a read-only memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the central processing unit 1301. The image computer device 1300 further includes a mass storage device 1306 for storing an operating system 1309, application programs 1310, and other program modules 1311.

[0226] The mass storage device 1306 is connected to the central processing unit 1301 through a mass storage controller (not shown) connected to the system bus 1305. The mass storage device 1306 and its associated computer-readable medium provide non-volatile storage for the image computer device 1300. That is to say, the mass storage device 1306 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0227] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only registers (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above system memory 1304 and mass storage device 1306 can be collectively referred to as memory.

[0228] According to various embodiments of the present disclosure, the image computer device 1300 may also be operated by connecting to a remote computer on a network such as the Internet. That is, the image computer device 1300 can be connected to the network 1308 through the network interface unit 1307 connected to the system bus 1305, or in other words, the network interface unit 1307 can also be used to connect to other types of networks or remote computer systems (not shown).

[0229] The memory further includes at least one segment of computer program, and the at least one segment of computer program is stored in the memory. The central processing unit 1301 implements all or part of the steps in the training method or speech synthesis method of the vocoder model shown in the above various embodiments by executing the at least one segment of program.

[0230] The embodiments of the present application further provide a computer device, which includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the training method or speech synthesis method of the vocoder model provided in the above method embodiments.

[0231] The embodiments of the present application further provide a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by the processor to implement the training method or speech synthesis method of the vocoder model provided in the above method embodiments.

[0232] The embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device, so that the computer device executes to implement the training method or voice synthesis method of the vocoder model provided in the above method embodiments.

[0233] It can be understood that in the specific implementation of the present application, for the data, historical data, portraits and other user data processing related to the user identity or characteristics, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0234] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0235] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.

[0236] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a vocoder model, characterized in that The method includes: Obtaining at least two sample acoustic features corresponding to a sample audio, where the sample acoustic features include at least two of a harmonic distribution feature, a harmonic amplitude feature, a fundamental frequency feature, and a time-varying filtered noise feature; Performing simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features; Fusing the simulated acoustic features corresponding to different sample acoustic features, and decoding the fused simulated audio obtained by the fusion processing to obtain a simulated audio; Decomposing the simulated audio into Mel spectrum features; Using the vocoder model to perform speech synthesis processing on the Mel spectrum features to generate a simulated synthesized audio corresponding to the simulated audio; Calculating a training loss of the vocoder model based on the simulated audio and the simulated synthesized audio; Updating model parameters of the vocoder model according to the training loss.

2. The method according to claim 1, characterized in that, The training system of the vocoder model runs the vocoder model and a simulated audio generation model; wherein, the simulated audio generation model includes a simulated acoustic feature generation network and a feature extraction network; The performing simulated acoustic feature generation processing based on the sample acoustic features to obtain simulated acoustic features corresponding to the sample acoustic features includes: Determining a change pattern corresponding to the sample acoustic features through the feature extraction network, where the change pattern is used to represent a change trend of the sample acoustic features; Constructing an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features.

3. The method according to claim 2, wherein The sample acoustic features include a harmonic distribution feature; The constructing an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features includes: Constructing an acoustic feature with the same first change pattern as that corresponding to the harmonic distribution feature through the simulated acoustic feature generation network to obtain a simulated harmonic distribution feature corresponding to the harmonic distribution feature.

4. The method according to claim 2, wherein The sample acoustic features include a harmonic amplitude feature; The constructing an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features includes: Constructing an acoustic feature with the same second change pattern as that corresponding to the harmonic amplitude feature through the simulated acoustic feature generation network to obtain a simulated harmonic amplitude feature corresponding to the harmonic amplitude feature.

5. The method according to claim 2, wherein The sample acoustic features include a fundamental frequency feature; The constructing an acoustic feature with the same change pattern as that corresponding to the sample acoustic features through the simulated acoustic feature generation network to obtain the simulated acoustic features corresponding to the sample acoustic features includes: Constructing an acoustic feature with the same third change pattern as that corresponding to the fundamental frequency feature through the simulated acoustic feature generation network to obtain a simulated fundamental frequency feature corresponding to the fundamental frequency feature.

6. The method according to claim 2, wherein The sample acoustic features include a time-varying filtered noise feature; Constructing an acoustic feature with the same change pattern as that corresponding to the sample acoustic feature through the simulated acoustic feature generation network to obtain the simulated acoustic feature corresponding to the sample acoustic feature, includes: Constructing an acoustic feature with the same fourth change pattern as that corresponding to the time-varying filtered noise feature through the simulated acoustic feature generation network to obtain the simulated time-varying filtered noise feature corresponding to the time-varying filtered noise feature.

7. The method according to claim 2, wherein The simulated audio generation model further includes a decoding network; Fusing the simulated acoustic features corresponding to different sample acoustic features to obtain a simulated audio, includes: Performing a fusion process on the simulated acoustic features corresponding to different sample acoustic features to obtain a fused simulated audio; Performing a decoding process on the fused simulated audio through the decoding network to generate the simulated audio.

8. The method according to claim 7, wherein The simulated acoustic feature includes at least two of a simulated harmonic distribution feature, a simulated harmonic amplitude feature, a simulated fundamental frequency feature, and a simulated time-varying filtered noise feature.

9. The method according to claim 2, wherein The simulated audio generation model further includes an encoding network and a feature extraction network; Obtaining at least two sample acoustic features corresponding to a sample audio, includes: Obtaining the sample audio and decomposing the sample audio into at least two sample acoustic features through the encoding network.

10. A voice synthesis method, characterized in that, The method includes: Obtaining the Mel spectrum feature corresponding to a target audio; Inputting the Mel spectrum feature into a vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio, wherein the vocoder model is the vocoder model according to any one of claims 1 to 9.

11. The method according to claim 10, wherein The vocoder model includes an encoding network; Obtaining the Mel spectrum feature corresponding to a target audio, includes: Obtaining the target audio; Decomposing the target audio into the Mel spectrum feature through the encoding network.

12. The method according to claim 11, wherein The vocoder model includes a speech synthesis network; Inputting the Mel spectrum feature into the vocoder model for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio, includes: Inputting the Mel spectrum feature into the speech synthesis network for speech synthesis processing to obtain the target synthesized audio corresponding to the target audio.

13. A training device for a vocoder model, characterized in that The apparatus includes: An obtaining module, configured to obtain at least two sample acoustic features corresponding to a sample audio, where the sample acoustic feature includes at least two of a harmonic distribution feature, a harmonic amplitude feature, a fundamental frequency feature, and a time-varying filtered noise feature; A generating module, configured to perform simulated acoustic feature generation processing based on the sample acoustic features to obtain the simulated acoustic features corresponding to the sample acoustic features; A fusing module, configured to perform a fusion process on the simulated acoustic features corresponding to different sample acoustic features, and perform a decoding process on the fused simulated audio obtained by the fusion process to obtain a simulated audio; A speech synthesis module, configured to decompose the simulated audio into Mel spectrum features; Using the vocoder model to perform speech synthesis processing on the Mel spectrum features to generate the simulated synthesized audio corresponding to the simulated audio; A calculation module, configured to calculate a training loss of the vocoder model based on the analog audio and the analog synthesized audio; An update module, configured to update model parameters of the vocoder model according to the training loss.

14. A voice synthesis device, characterized in that, The apparatus includes: An acquisition module, configured to acquire Mel-spectrum features corresponding to target audio; A speech synthesis module, configured to input the Mel-spectrum features into a vocoder model for speech synthesis processing to obtain target synthesized audio corresponding to the target audio, where the vocoder model is the vocoder model according to any one of claims 1 to 11.

15. A computer device, characterized in that, The computer device includes: a processor and a memory. At least one computer program is stored in the memory, and at least one of the computer programs is loaded and executed by the processor to implement the training method of the vocoder model according to any one of claims 1 to 9, or the speech synthesis method according to any one of claims 10 to 12.

16. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and at least one computer program is loaded and executed by a processor to implement the training method of the vocoder model according to any one of claims 1 to 9, or the speech synthesis method according to any one of claims 10 to 12.

17. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device, so that the computer device executes the training method of the vocoder model according to any one of claims 1 to 9, or the speech synthesis method according to any one of claims 10 to 12.