Voice Processing Method, Apparatus, Electronic Device, and Storage Medium

By using a small amount of noise-containing real speech and pre-trained acoustic models in the speech cloning technology, the high requirements for corpus quality and quantity in the prior art are solved, and high-quality speech cloning in low resources is achieved.

CN115497451BActive Publication Date: 2025-06-24WWZN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211124413.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-06-24
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

Existing speech cloning technologies require a large number of high-quality corpus and it is difficult to obtain high-similar and natural speech in noisy environments.

Method used

By obtaining a small amount of real speech containing noise, extracting noise features to generate mask information, using a pre-trained acoustic model to generate acoustic features based on text and mask information, and updating model parameters to achieve speech cloning.

Benefits of technology

It realizes high-quality voice cloning under low resources, simplifies the voice cloning process, and reduces the requirements for corpus quality and quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497451B_ABST
    Figure CN115497451B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech processing method, apparatus, electronic device, and storage medium. The speech processing method according to an embodiment of the present disclosure includes: obtaining a first text and a first true speech of a first speaker, where the content of the first true speech is the same as the content of the first text; obtaining first mask information indicating noise characteristics in the first true speech; generating first acoustic features corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model; extracting second acoustic features of the first speaker from the first true speech; and updating parameters of the acoustic model according to the first acoustic features and the second acoustic features. The present disclosure can achieve high-quality speech cloning under low-resource conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a voice processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Voice cloning refers to a technology in which a computer can automatically generate any voice of a target speaker based on the existing voice data of the target speaker. Currently, existing voice cloning technologies require a large amount of corpus of the target speaker and require that the corpus does not contain any noise. However, in practical applications, the corpus of the target speaker often not only contains various noises, but also the number of corpus is limited. Therefore, it is difficult for current voice cloning technologies to obtain voices with high similarity, high naturalness, and low noise. Summary of the Invention

[0003] To solve at least one of the above technical problems, the present disclosure provides a voice processing method, apparatus, electronic device, and storage medium.

[0004] The first aspect of the present disclosure provides a voice processing method, including:

[0005] Obtaining a first text and first real voice of a first speaker, where the content of the first real voice is the same as the content of the first text;

[0006] Obtaining first mask information indicating noise characteristics in the first real voice;

[0007] Generating first acoustic features corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model;

[0008] Extracting second acoustic features of the first speaker from the first real voice;

[0009] Updating parameters of the acoustic model according to the first acoustic features and the second acoustic features.

[0010] In some embodiments of the present disclosure, the voice processing method further includes:

[0011] Obtaining a second text and pre-configured second mask information, where the second mask information is a clean mask;

[0012] Generating second acoustic features corresponding to the second text based on the second mask information and the second text by using the acoustic model with updated parameters;

[0013] Synthesizing the second acoustic features into a first voice, where the content of the first voice is the same as the second text and the second voice has the timbre characteristics of the first speaker.

[0014] In some embodiments of the present disclosure, synthesizing the second acoustic feature into the first speech includes: synthesizing the second acoustic feature into the first speech by using a pre-trained vocoder, where the vocoder is trained according to the audio data of the first speaker.

[0015] In some embodiments of the present disclosure, generating a spectral frame corresponding to the first text and generating a first acoustic feature by using a pre-trained acoustic model based on the first text and the first mask information includes:

[0016] Obtaining a first text feature vector corresponding to the first text by using an encoder in the acoustic model;

[0017] Generating the first acoustic feature by using a decoder in the acoustic model according to the first text feature vector and the first mask information.

[0018] In some embodiments of the present disclosure, generating the first acoustic feature by using a decoder in the acoustic model according to the first text feature vector and the first mask information includes:

[0019] Performing processing of an attention network in the decoder by using the first text feature vector to obtain a first attention vector corresponding to the first text;

[0020] Based on the first attention vector and the previous spectral frame, sequentially performing processing of an LSTM and a linear projection layer in the decoder to obtain the current spectral frame;

[0021] Performing processing of a post-processing network in the decoder based on the current spectral frame and the first mask information to optimize the current spectral frame;

[0022] After obtaining all spectral frames corresponding to the first text, splicing all spectral frames to obtain the first acoustic feature.

[0023] In some embodiments of the present disclosure, the acoustic model is trained according to the corpora of multiple second speakers, and the corpora of the multiple second speakers include: clean real speech and real speech with noise.

[0024] A second aspect of the present disclosure provides a speech processing apparatus, including:

[0025] An acquisition unit, configured to acquire a first text and a first real speech of a first speaker, where the content of the first real speech is the same as the content of the first text, and the first real speech contains noise;

[0026] A noise processing unit, configured to acquire first mask information indicating a noise feature in the first real speech;

[0027] A first acoustic feature unit, configured to generate a first acoustic feature corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model;

[0028] A second acoustic feature unit, configured to extract a second acoustic feature of the first speaker from the first real speech;

[0029] A parameter update unit, configured to update parameters of the acoustic model according to the first acoustic feature and the second acoustic feature, so that the acoustic model can be used to clone the speech of the first speaker.

[0030] In some embodiments of the present disclosure, the acquisition unit is further configured to acquire a second text and a pre-configured second mask information, where the second mask information is a clean mask; the first acoustic feature unit is further configured to generate a second acoustic feature corresponding to the second text based on the second mask information and the second text by using the acoustic model with updated parameters; the speech processing device further includes: a speech generation unit, configured to synthesize the second acoustic feature into a first speech, where the content of the first speech is the same as that of the second text and the second speech has the timbre feature of the first speaker.

[0031] A third aspect of the present disclosure provides an electronic device, including:

[0032] A memory, where the memory stores execution instructions; and

[0033] A processor, where the processor executes the execution instructions stored in the memory, so that the processor executes the above-mentioned speech processing method.

[0034] A fourth aspect of the present disclosure provides a readable storage medium, where execution instructions are stored in the readable storage medium, and when the execution instructions are executed by a processor, they are used to implement the above-mentioned speech processing method.

[0035] Embodiments of the present disclosure only need a small amount of corpus with low quality requirements to implement voice cloning of a speaker, and can be directly implemented by low-resource hardware, which is not only simple and fast to implement. Moreover, for various users who need to clone speech, they only need to record a small amount of audio with low quality to automatically implement their voice cloning. Description of the Drawings

[0036] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure, and the drawings are included in this specification and form a part of this specification.

[0037] Figure 1It is a schematic flowchart of a voice processing method according to some embodiments of the present disclosure.

[0038] Figure 2 It is a schematic block diagram of a voice processing device implemented in hardware using a processing system according to an embodiment of the present disclosure.

[0039] Explanation of Reference Numerals

[0040] 200 Voice processing model

[0041] 300 Bus

[0042] 400 Processor

[0043] 500 Memory

[0044] 600 Various other circuits Specific embodiments

[0045] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of description, only the parts related to the present disclosure are shown in the drawings.

[0046] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.

[0047] Unless otherwise specified, the exemplary embodiments / Examples shown are understood to provide exemplary features of various details of some ways that can implement the technical concept of the present disclosure in practice. Therefore, unless otherwise specified, without departing from the technical concept of the present disclosure, the features of various embodiments / Examples can be combined, separated, interchanged, and / or rearranged additionally.

[0048] In the drawings, cross-hatching and / or shading are generally used to make the boundaries between adjacent components clear. Thus, unless stated, the presence or absence of cross-hatching or shading does not convey or imply any preference or requirement for the specific material, material properties, dimensions, proportions, commonality between the components shown, and / or any other characteristics, attributes, properties, etc. of the components. Additionally, in the drawings, for the purpose of clarity and / or description, the dimensions and relative dimensions of the components may be exaggerated. When the exemplary embodiments can be implemented differently, the specific process sequences can be executed in a different order than described. For example, two consecutively described processes can be executed substantially simultaneously or in an order opposite to the described order. Moreover, the same reference numerals denote the same components.

[0049] When a component is referred to as being "on" or "above" another component, "connected to" or "coupled to" another component, the component can be directly on the other component, directly connected to or directly coupled to the other component, or there can be an intermediate component. However, when a component is referred to as being "directly on" another component, "directly connected to" or "directly coupled to" another component, there is no intermediate component. For this reason, the term "connected" can refer to a physical connection, an electrical connection, etc., and can have or not have an intermediate component.

[0050] The terms used herein are for the purpose of describing particular embodiments and are not intended to be limiting. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are also intended to include the plural forms. In addition, when the terms "comprising" and / or "including" and their variants are used in this specification, it is stated that there are the stated features, integers, steps, operations, components, assemblies and / or groups thereof, but does not preclude the presence or addition of one or more other features, integers, steps, operations, components, assemblies and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about" and other similar terms are used as approximate terms and not as terms of degree, so they are used to explain the inherent deviations of measured, calculated and / or provided values that would be recognized by a person of ordinary skill in the art.

[0051] Term Explanation in this article:

[0052] Neural network vocoder (LPCNet): A neural network-based vocoder capable of synthesizing acoustic features such as Mel spectrograms into audio.

[0053] Text encoder: Comprising a character embedding convolutional neural network and a bi-directional long short-term memory (bi-directional LSTM) network connected in sequence, the convolutional neural network can include 3 convolutional layers connected in sequence. This encoder can be used to encode text into a vector.

[0054] Decoder: It includes a preprocessing network (Pre-Net), an attention network (Stepwise Attention), a two-layer LSTM (2LSTM Layers) (i.e., two stacked LSTMs), a linear projection layer (Linear Projection), a linear projection layer with activation function (Linear Projection+Sigmoid), a postprocessing network (Post-Net), and a decoding output network. The output data of the preprocessing network and the output data of the attention network are input into the two-layer LSTM. The output data of the two-layer LSTM are input into the linear projection layer and the linear projection layer with activation function. The output data of the linear projection layer are input into the postprocessing network. The output data of the postprocessing network and the output data of the linear projection layer are simultaneously input into the decoding output network. The output data of the decoding output network are the output data of the decoder, that is, acoustic features such as, for example, a mel spectrogram. For example, the postprocessing network (Post-Net) can be composed of 5 convolutional layers, and the postprocessing network can correct the spectral frames output by the linear projection layer. The decoder is usually implemented as an autoregressive recurrent neural network.

[0055] Currently, the related technologies of voice cloning mainly include the following:

[0056] 1) Speaker-adaptive voice cloning: An adaptive training is performed on a multi-speaker text-to-speech (TTS) model to train a model that can synthesize the voice of the target speaker. The voice of the target speaker is cloned through this model. Although this method can clone a voice that is relatively close to the real voice of the target speaker and has a high naturalness, this method requires a large amount of corpus of the target speaker and has high requirements for the quality of the corpus of the target speaker. For example, the corpus should not contain too much noise.

[0057] 2) Speaker-encoded voice cloning: Based on a pre-trained independent model, the timbre embedding information of the target speaker is inferred using a large amount of corpus of the target speaker, and the timbre embedding information of the target speaker is provided to a multi-speaker text-to-speech system, and the voice of the target speaker is synthesized by this text-to-speech system. The voice obtained by this method is poor in both naturalness and similarity, also requires a large amount of corpus of the target speaker, and since an independent model needs to be trained and two models are allowed to run simultaneously, the resource consumption is higher.

[0058] In summary, the current voice cloning technologies have problems such as high requirements for corpus quality, large demand for corpus quantity, low voice quality, and high resource consumption. In view of this, the embodiments of the present disclosure provide the following voice processing methods, devices, electronic devices, and storage media, which train an existing acoustic model based on a small amount of corpus of the speaker to make it have the ability of noise suppression and voice cloning at the same time, so as to achieve high-quality voice cloning in low-resource situations.

[0059] Figure 1 The flowchart shows the voice processing method according to some embodiments of the present disclosure.

[0060] As Figure 1 shown, the voice processing method according to the embodiments of the present disclosure may include:

[0061] Step S12, obtaining a first text and a first true voice of a first speaker, where the content of the first true voice is the same as the content of the first text;

[0062] In a specific application, the first speaker may input the first text to the electronic device through a human-computer interaction interface or the like, and at the same time use the microphone of the electronic device to record the first true voice with the same content as the first text. The first true voice may include various noises such as environmental noise and the voices of others, and the amount and level of the noise are not limited.

[0063] Here, the electronic device may be a mobile terminal such as a mobile phone, a tablet computer or the like, or a fixed terminal such as a computer. The electronic device only needs to have the ability to run the following acoustic model and allow the user to input text and record audio.

[0064] In the embodiments of the present disclosure, the data volume of the first text and the first true voice may be small and can be flexibly controlled by the user (for example, the first speaker). For example, the first text may include about 20 sentences of similar size, and the first true voice may be, for example, about 2 minutes of audio data.

[0065] Step S14, obtaining first mask information indicating the noise characteristics in the first true voice;

[0066] In the embodiments of the present disclosure, various applicable noise extraction methods may be used to extract noise information from the first true voice. The noise information here includes information on various noises such as environmental noise, the voices of others, and background sounds.

[0067] In some embodiments, a pre-trained or existing noise extraction model (for example, RNNOISE, etc.) may be used to extract mask information indicating noise characteristics from the first true voice. For example, the noise extraction model may be a deep learning model such as a neural network, which may be trained with data including noisy audio and noise-free audio. Specifically, the voice feature vector of the first true voice of the first speaker may be extracted through a convolutional neural network or the like, and then the deep learning model may be used to process the voice feature vector of the first true voice to extract the noise information of the first true voice.

[0068] In some embodiments, the value of the mask information can be between [0, 1]. When the value of the mask information is 0, it indicates that the noise is the largest. When the value of the mask information is 1, the noise is the smallest. The closer the value of the mask information is to 0, the greater the noise. The closer the value of the mask information is to 1, the smaller the noise.

[0069] Step S16: Use the pre-trained acoustic model to generate the first acoustic features corresponding to the first text based on the first text and the first mask information;

[0070] In some embodiments, the acoustic model can adopt the acoustic model structure in speech synthesis models such as Tacotron, Tacotron2, etc. That is, the acoustic model can include an encoder and a decoder. The structure of the encoder can adopt the structure of the text encoder described above, and the structure of the decoder can adopt the structure of the decoder described above. The present disclosure embodiments do not limit the network architecture of the acoustic model.

[0071] Taking the encoder + decoder architecture as an example, step S16 can include:

[0072] Step a1: Use the encoder in the acoustic model to obtain the first text feature vector corresponding to the first text;

[0073] Step a2: Use the decoder in the acoustic model to generate the first acoustic features according to the first text feature vector and the first mask information.

[0074] Still taking the decoder architecture described above as an example, step a2 can include:

[0075] Step a21: Use the first text feature vector to perform the processing of the attention network in the decoder to obtain the first attention vector corresponding to the first text;

[0076] Step a22: Based on the first attention vector and the previous spectral frame, sequentially perform the processing of the LSTM and the linear projection layer in the decoder to obtain the current spectral frame;

[0077] Step a23: Based on the current spectral frame and the first mask information, perform the processing of the post-processing network in the decoder to optimize the current spectral frame;

[0078] Step a24: After obtaining all the spectral frames corresponding to the first text, splice all the spectral frames to obtain the first acoustic features.

[0079] In this way, while the post-processing network performs fine-tuning on the spectral frames, the mask information and the spectral frames can be fused to suppress the noise features in the spectral frames, thereby obtaining spectral frames with better quality, and further improving the quality of the final synthesized speech.

[0080] In some embodiments, the acoustic model may further include: a duration prediction model, which can be used to estimate the phoneme duration of text. The duration prediction model can be a deep learning model such as a neural network, and is obtained through pre-training.

[0081] In step S16, before step a2, it may further include: obtaining the phoneme duration of the first text; expanding the first text feature vector according to the phoneme duration of the first text. In this way, the decoder does not need to use an autoregressive recurrent neural network, and can terminate the generation of spectral frames in real time through the expanded first text feature vector.

[0082] In some embodiments, the acoustic model can be trained according to the corpora of multiple second speakers, and the corpora of the multiple second speakers may include: clean real speech and real speech with noise.

[0083] In some embodiments, the acoustic model can be pre-trained through the following steps:

[0084] Step b1, construct an acoustic model;

[0085] Step b2, obtain sample data;

[0086] The sample data includes text samples and the corpora of multiple second speakers corresponding to the text samples. Each corpus of a second speaker contains real speech samples corresponding to one or more text samples and labeled with speaker information. The real speech samples can be clean speech or can contain noise.

[0087] Preferably, the noise contents of different real speech samples in the corpora of multiple second speakers can be different. For example, in the corpus, a certain proportion of real speech samples can be clean speech, another proportion of real speech samples can contain low-level noise, and still a certain proportion of real speech samples can contain more noise, etc. In this way, it is beneficial for the acoustic model to better learn the characteristics of different noises and generate clean acoustic features under various noisy corpora.

[0088] Step b3, use the sample data to train the above acoustic model so that the acoustic model can learn the acoustic features of other speakers using corpora with uneven quality and different noise contents.

[0089] The training process of step b3 is basically the same as the process of steps S12 to S110 described above, and will not be elaborated here.

[0090] Step S18, extract the second acoustic features of the first speaker from the first real speech;

[0091] For example, the second acoustic feature can be extracted from the first real speech through, for example, a pre-trained speech feature recognition model or an existing recognition tool, etc.

[0092] Step S110: Update the parameters of the acoustic model according to the first acoustic feature and the second acoustic feature.

[0093] For example, the update of the acoustic model parameters can be achieved through algorithms such as gradient descent.

[0094] For example, loss functions such as mean square error (MSE) can be adopted when updating the acoustic model parameters.

[0095] Through steps S12 to S110, it is not necessary to perform noise reduction processing on the audio of the first speaker, and the acoustic model can be made to have the function of generating the acoustic features of the first speaker by using a small amount of corpus of the first speaker. In this way, by combining the acoustic model with the vocoder described below, the speech synthesis of the first speaker can be realized, that is, the voice cloning of the first speaker is realized.

[0096] In some embodiments, after step S110, it may further include:

[0097] Step S112: Obtain the second text and the pre-configured second mask information, where the second mask information is a clean mask;

[0098] The second text refers to the text of the speech to be cloned or synthesized. The second text can be input by the user, actively obtained by the electronic device from an external device, or actively provided by the external device to the electronic device.

[0099] The second mask information can be pre-stored in a specified storage space, and the second mask information can be read from the specified storage space during the voice cloning process.

[0100] For example, the values of all elements in the second mask information can be 1 to obtain a clean synthesized speech through the second mask information.

[0101] Step S114: Use the updated acoustic model to generate the second acoustic feature corresponding to the second text based on the second mask information and the second text;

[0102] Step S116: Synthesize the second acoustic feature into the first speech, where the content of the first speech is the same as the second text and the second speech has the timbre characteristics of the first speaker.

[0103] For example, a pre-trained vocoder can be used to synthesize the second acoustic feature into the first speech, and the vocoder can be trained based on the audio data of the first speaker. The vocoder can be, but is not limited to, the LPCNet described above. In a specific application, the vocoder can be trained independently and installed in hardware such as an electronic device for use after training.

[0104] In steps S112 to S114, by providing clean mask information, high-quality voice cloning of any content for the target speaker can be achieved on a user side such as an electronic device.

[0105] Since only a small amount of corpus with low quality requirements is needed in the embodiments of the present disclosure to achieve voice cloning of a speaker, it can be directly implemented through the above-mentioned electronic device or other similar low-resource hardware. It is not only simple and fast to implement, but also can be automatically implemented. For various users who need to clone voice, only a small amount of audio with low quality needs to be recorded to automatically achieve their voice cloning.

[0106] Figure 2 FIG. 200 is a schematic structural diagram of a voice processing device implemented in hardware using a processing system according to an embodiment of the present disclosure.

[0107] The device may include corresponding modules for performing each or several steps in the above flowchart. Therefore, each step or several steps in the above flowchart can be performed by the corresponding modules, and the device may include one or more of these modules. The modules can be one or more hardware modules specifically configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer-readable medium for implementation by the processor, or implemented through a certain combination.

[0108] This hardware structure can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 300 connects various circuits including one or more processors 400, a memory 500, and / or hardware modules together. The bus 300 can also connect various other circuits 600 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0109] The bus 300 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one connection line is used in this figure, but it does not mean that there is only one bus or one type of bus.

[0110] Any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present disclosure belong. The processor executes the various methods and processes described above. For example, the method embodiments in the present disclosure can be implemented as software programs, which are tangibly contained in a machine-readable medium, such as a memory. In some embodiments, part or all of the software program can be loaded and / or installed via the memory and / or a communication interface. When the software program is loaded into the memory and executed by the processor, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the processor can be configured to execute one of the above methods by any other suitable means (e.g., by means of firmware).

[0111] The logic and / or steps represented in the flowchart or described in other ways herein can be specifically implemented in any readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices.

[0112] For the purposes of this specification, a "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the readable storage medium include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM). Additionally, the readable storage medium can even be paper or other suitable medium on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing it as appropriate, and then storing it in a memory.

[0113] It should be understood that various parts of the present disclosure can be implemented in hardware, software, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0114] Those of ordinary skill in the art of the present technology can understand that all or part of the steps for implementing the above embodiments of the method can be completed by a program instructing relevant hardware, and the program can be stored in a readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0115] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0116] Figure 2 It is a schematic structural diagram of a voice processing device 200 according to an embodiment of the present disclosure.

[0117] As Figure 2 shown, the voice processing device 200 can include:

[0118] An acquisition unit 202, configured to acquire a first text and a first true speech of a first speaker, where the content of the first true speech is the same as that of the first text, and the first true speech contains noise;

[0119] A noise processing unit 204, configured to acquire first mask information indicating noise characteristics in the first true speech;

[0120] A first acoustic feature unit 206, configured to generate first acoustic features corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model;

[0121] A second acoustic feature unit 208, configured to extract second acoustic features of the first speaker from the first true speech;

[0122] A parameter update unit 210, configured to update parameters of the acoustic model according to the first acoustic features and the second acoustic features, so that the acoustic model can be used to clone the speech of the first speaker.

[0123] In some embodiments, the acquisition unit 202 may further be configured to acquire a second text and pre-configured second mask information, and the second mask information is a clean mask;

[0124] The first acoustic feature unit 206 may further be configured to generate second acoustic features corresponding to the second text based on the second mask information and the second text by using the acoustic model with updated parameters;

[0125] The voice processing device 200 may further include: a voice generation unit 212, configured to synthesize the second acoustic features into a first voice, where the content of the first voice is the same as that of the second text and the second voice has the timbre characteristics of the first speaker.

[0126] The present disclosure further provides an electronic device, including: a memory storing execution instructions; and a processor or other hardware module, where the processor or other hardware module executes the execution instructions stored in the memory, so that the processor or other hardware module executes the above voice processing method.

[0127] The present disclosure further provides a readable storage medium storing execution instructions, where the execution instructions, when executed by a processor, are used to implement the above voice processing method.

[0128] In the description of this specification, the description referring to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0129] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0130] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. A voice processing method, characterized in that, Including: Obtain a first text and first true speech of a first speaker, where the content of the first true speech is the same as the content of the first text; Obtain first mask information indicating noise characteristics in the first true speech; Generate first acoustic features corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model, including: obtaining a first text feature vector corresponding to the first text by using an encoder in the acoustic model; generating first acoustic features by using a decoder in the acoustic model according to the first text feature vector and the first mask information, including: performing processing of an attention network in the decoder by using the first text feature vector to obtain a first attention vector corresponding to the first text; sequentially performing processing of an LSTM and a linear projection layer in the decoder based on the first attention vector and a previous spectral frame to obtain a current spectral frame; performing processing of a post-processing network in the decoder based on the current spectral frame and the first mask information to optimize the current spectral frame; after obtaining all spectral frames corresponding to the first text, splicing all spectral frames to obtain first acoustic features; Extract second acoustic features of the first speaker from the first true speech; Update parameters of the acoustic model according to the first acoustic features and the second acoustic features; The acoustic model is trained according to corpora of multiple second speakers, The corpora of the multiple second speakers include: clean true speech and true speech containing noise.

2. The voice processing method according to claim 1, wherein Also including: Obtain a second text and pre-configured second mask information, where the second mask information is a clean mask; Generate second acoustic features corresponding to the second text based on the second mask information and the second text by using the acoustic model with updated parameters; Synthesize the second acoustic features into first speech, where the content of the first speech is the same as the second text and the first speech has the timbre characteristics of the first speaker.

3. The voice processing method according to claim 2, characterized in that, The synthesizing the second acoustic features into first speech includes: synthesizing the second acoustic features into first speech by using a pre-trained vocoder, and the vocoder is trained according to audio data of the first speaker.

4. A voice processing device, characterized in that, Including: An obtaining unit, configured to obtain a first text and first true speech of a first speaker, where the content of the first true speech is the same as the content of the first text, and the first true speech contains noise; A noise processing unit, configured to obtain first mask information indicating noise characteristics in the first true speech; The first acoustic feature unit is configured to generate a first acoustic feature corresponding to the first text based on the first text and the first mask information by using a pre-trained acoustic model, including: obtaining a first text feature vector corresponding to the first text by using an encoder in the acoustic model; generating the first acoustic feature by using a decoder in the acoustic model according to the first text feature vector and the first mask information, including: performing processing on an attention network in the decoder by using the first text feature vector to obtain a first attention vector corresponding to the first text; sequentially performing processing on an LSTM and a linear projection layer in the decoder based on the first attention vector and a previous spectral frame to obtain a current spectral frame; performing processing on a post-processing network in the decoder based on the current spectral frame and the first mask information to optimize the current spectral frame; after obtaining all spectral frames corresponding to the first text, splicing all the spectral frames to obtain the first acoustic feature; The second acoustic feature unit is configured to extract a second acoustic feature of the first speaker from the first real speech; The parameter update unit is configured to update parameters of the acoustic model according to the first acoustic feature and the second acoustic feature, so that the acoustic model can be used to clone the voice of the first speaker; The acoustic model is trained according to corpora of multiple second speakers, and the corpora of the multiple second speakers include: clean real speech and real speech with noise.

5. The voice processing device according to claim 4, wherein, The obtaining unit is further configured to obtain a second text and pre-configured second mask information, and the second mask information is a clean mask; The first acoustic feature unit is further configured to generate a second acoustic feature corresponding to the second text based on the second mask information and the second text by using the acoustic model with updated parameters; The voice processing device further includes: a voice generation unit configured to synthesize the second acoustic feature into a first voice, where the content of the first voice is the same as that of the second text and the first voice has the timbre feature of the first speaker.

6. An electronic device, characterized in that, including: a memory that stores execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the voice processing method according to any one of claims 1 to 4.

7. A readable storage medium, characterized in that, Execution instructions are stored in the readable storage medium, and when the execution instructions are executed by the processor, they are used to implement the voice processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speech synthesis model training method and speech synthesis method

    CN112634856A

  • Speech synthesis system training method and device, computer equipment and storage medium

    CN113299269A