Speech synthesis method, speech synthesis device and electronic device
By processing pitch features step by step to generate high-quality voice signals, the problem of speech synthesis speed and quality in lightweight deployment structures is solved, and efficient speech synthesis effect is achieved.
Patent Information
- Application Number
- CN202210518612.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-05-13
AI Technical Summary
It is difficult for existing speech synthesis technology to achieve high-quality speech synthesis in lightweight deployment structural scenarios, and the existing neural networks are slow to decode or have large parameters, which cannot meet the needs of synthesis speed and effect.
Pitch is used as an acoustic feature, and high-quality voice signals are generated through step-by-step processing by prevocoder and corrector, avoiding one-time conversion of complex neural networks and using simple speech synthesizers and deep neural networks for correction.
Generate high-quality voice signals in a lightweight deployment structure, improve decoding speed and sound quality, reduce network parameters, and is suitable for voice interactions of smart devices.
Smart Images

Figure CN114974207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice synthesis method, a voice synthesis device, and an electronic device. Background Art
[0002] With the continuous development of artificial intelligence technology, research and applications of artificial intelligence technology have been carried out in multiple fields. Natural Language Processing (NLP) and speech processing are important directions in artificial intelligence technology. For example, a text is synthesized into a synthesized speech through a voice synthesis model, and then the synthesized speech can be played to a user.
[0003] In some application scenarios, users pay more and more attention to the usage experience of intelligent software and hardware. In the process of intelligent interaction, voice interaction plays an important role, and voice synthesis technology is an important technology indispensable in the interaction process. Although users understand that the voice emitted by the machine is not the voice of a real person, in order to enrich the experience during the interaction, the synthesized voice should be as close as possible to the voice of a real person in terms of timbre and sound quality. Summary of the Invention
[0004] A voice synthesis method, a voice synthesis device, and an electronic device provided by an embodiment of the present disclosure can provide voice synthesis with a decoding speed and effect meeting requirements and are applicable to scenarios with a lightweight deployment structure.
[0005] According to one aspect of the present disclosure, a voice synthesis method is provided. The method includes: obtaining an acoustic feature corresponding to a voice to be synthesized, where the acoustic feature includes pitch, and the pitch refers to the height of the voice, representing the frequency and wavelength of the voice; processing the acoustic feature to generate a first audio signal; and correcting the first audio signal to obtain a corrected voice signal.
[0006] According to another aspect of the present disclosure, a voice synthesis device is provided. The device includes: an obtaining unit configured to obtain an acoustic feature corresponding to a voice to be synthesized, where the acoustic feature includes pitch, and the pitch refers to the height of the voice, representing the frequency and wavelength of the voice; a generating unit configured to process the acoustic feature to generate a first audio signal; and a correcting unit configured to correct the first audio signal to obtain a corrected voice signal.
[0007] The present disclosure first processes pitch as an acoustic feature to generate a first audio signal, and then corrects the first audio signal to obtain a corrected speech signal. Since the acoustic feature is processed through the above two steps respectively with an additional correction process for the audio signal, a higher-quality synthesized speech signal is obtained. The present disclosure avoids directly converting the acoustic feature into high-quality speech audio at one time using a complex neural network, has a smaller number of parameters, and can be applied to scenarios that require a lightweight deployment structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present disclosure are disclosed. In the drawings:
[0009] Figure 1 FIG. shows a schematic diagram of an example application scenario in which various methods described herein can be implemented according to an exemplary embodiment of the present disclosure;
[0010] Figure 2 FIG. shows a flowchart of a speech synthesis method according to an exemplary embodiment of the present disclosure;
[0011] Figure 3 FIG. shows a schematic structural diagram of a speech synthesis apparatus according to an exemplary embodiment of the present disclosure;
[0012] Figure 4A FIG. shows a flowchart of a speech synthesis method according to another exemplary embodiment of the present disclosure;
[0013] Figure 4B FIG. shows a schematic structural diagram of a speech synthesis apparatus according to another exemplary embodiment of the present disclosure;
[0014] Figure 5 FIG. shows a schematic structural diagram of a speech generation apparatus according to another exemplary embodiment of the present disclosure;
[0015] Figure 6 FIG. shows a schematic structural diagram of a speech correction apparatus according to another exemplary embodiment of the present disclosure;
[0016] Figure 7 FIG. shows a decoding diagram of a speech correction apparatus according to another exemplary embodiment of the present disclosure;
[0017] Figure 8 FIG. shows a schematic diagram of a speech synthesis apparatus according to another exemplary embodiment of the present disclosure;
[0018] Figure 9 FIG. shows a block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0020] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0021] As used herein, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.
[0022] It should be noted that the modifications of "one" and "a plurality" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more". The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] In the prior art, speech synthesis mainly includes three processing stages:
[0024] The first stage is the front-end processing stage, that is, converting the text corresponding to the speech to be synthesized into a vector that can be recognized by an acoustic model through processing. This front-end processing stage may include processes such as text normalization, disambiguation of polyphonic characters, and conversion to phonemes; among them, the acoustic model needs to perform model training on some texts, and / or different speech synthesis models perform model training on the speech in different languages spoken by different pronunciation objects respectively.
[0025] The second stage is the acoustic processing stage, which mainly converts the information processed by the acoustic module in the front-end processing stage into acoustic features. The technology in this stage is the key technology that determines the content and timbre of the generated audio. Among them, due to the different timbres of different pronunciation objects, after the model training is completed, the text corresponding to the speech to be synthesized will be synthesized according to the trained speech synthesis model, so that synthesized speech in different languages and with different timbres can be obtained.
[0026] The third stage is the acoustic decoding processing stage, which mainly includes the vocoder for decoding processing. Specifically, the vocoder is used to convert the acoustic features obtained in the above acoustic processing stage into the final required speech audio. This processing process has a great impact on the quality of the sound quality and will also affect the timbre to a certain extent.
[0027] To complete a speech synthesis system with excellent performance, the above front-end processing, acoustic processing, and acoustic decoding processing parts all need to well implement the functions of each module. Currently, the existing acoustic decoding processing stages, such as WaveNet, WaveRNN, MelGAN, HiFi-GAN, etc. based on neural networks, all convert acoustic features into speech audio at one time. The decoding speeds of WaveNet and WaveRNN are relatively slow and it is difficult to meet the synthesis speed requirements in actual applications; although MelGAN and HiFi-GAN have good decoding speeds and effects, their structures are too complex and the number of parameters is relatively large, which is not good enough for some scenarios that require lightweight deployment structures.
[0028] The speech synthesis processing method and speech synthesis processing device provided in this application can be applied to an application environment as Figure 1 shown. Among them, the electronic device 102 communicates with the server 104 through the network. This speech synthesis method can be executed by the electronic device 102 or the server 104, or jointly executed by the electronic device 102 and the server 104. Here, taking the execution of this speech synthesis method by the electronic device 102 as an example: The electronic device 102 obtains the text corresponding to the speech to be synthesized locally, or obtains the text corresponding to the speech to be synthesized from the server 104, and obtains the corresponding phoneme sequence based on this text; performs timbre processing on the phoneme sequence through the speech synthesis model to obtain acoustic features including target timbre information; among them, the speech synthesis model is trained by the electronic device 102 or the server 104 based on the phoneme sequence of the target text and the target speech of at least one target language, and is deployed on the electronic device 102; the target speech corresponds to the target text and is a speech with a target timbre generated according to the acoustic features extracted from speech samples with different timbres; performs speech synthesis on the acoustic features through the speech synthesis model to obtain synthesized speech in at least one target language and with the target timbre.
[0029] In addition, taking the execution of the speech synthesis method by the server 104 as an example for illustration: The server 104 respectively extracts acoustic features from speech samples of at least one target language to obtain training acoustic features; generates target speech of at least one target language with a target timbre based on the training acoustic features; when obtaining a training phoneme sequence from the target text corresponding to the target speech, performs timbre processing on the training phoneme sequence through a speech synthesis model to obtain training acoustic features including target timbre information; performs speech synthesis on the training acoustic features through the speech synthesis model to obtain predicted speech of at least one target language; the predicted speech has the target timbre corresponding to the target timbre information; adjusts the network parameters in the speech synthesis model based on the loss value between the predicted speech and the target speech, and then deploys the trained speech synthesis model to the electronic device 102, so that the electronic device 102 can implement the steps of the above speech synthesis method. In addition, the trained speech synthesis model can also be deployed to the server 104 so that the server 104 can also implement the steps of the above speech synthesis method.
[0030] Among them, the electronic device 102 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. In addition, the electronic device 102 can also be a smart vehicle-mounted device, and this smart vehicle-mounted device can use the phoneme sequence of the text to perform speech synthesis to obtain synthesized speech with a target timbre, thereby realizing voice interaction with the user.
[0031] The server 104 can be an independent physical server or a service node in a blockchain system. The service nodes in this blockchain system form a peer-to-peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In addition, the server 104 can also be a server cluster composed of multiple physical servers, and can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. A service end of a requirements management system can be installed on the server 104, and through this service end, it can interact with the terminal 104.
[0032] The electronic device 102 and the server 104 can be connected through communication connection methods such as Bluetooth, USB (Universal Serial Bus), or network, and this application does not make any restrictions here.
[0033] In one embodiment, asFigure 2 As shown in Figure 2 , a voice synthesis method is provided. This method is executed by an electronic device, which can be the Figure 1 electronic device 102 or the server 104 in Figure 1 , that is, this method can be executed by the electronic device 102 or the server 104. In this embodiment, taking the electronic device as the electronic device 102 as an example for illustration, the method includes the following steps:
[0034] S201, obtaining the acoustic features corresponding to the voice to be synthesized.
[0035] Among them, the acoustic features can be input acoustic feature parameters, including pitch (pitch, abbreviated as p). Pitch refers to the height of the voice, indicating the frequency and wavelength of the voice. Optionally, the acoustic feature parameters further include Mel spectrogram (Melspectrogram, m) and loudness (loudness, abbreviated as l), which will be introduced in detail in the following examples.
[0036] S202, processing the acoustic features and generating a first audio signal.
[0037] Specifically, the above acoustic features, such as pitch, or Mel spectrogram, pitch, and loudness, can be hierarchically processed by a voice generation device to synthesize the first audio signal.
[0038] S203, correcting the first audio signal to obtain a corrected voice signal.
[0039] Specifically, a voice correction device can be used to further correct the first audio signal by using a deep neural network to generate a synthesized voice signal with better sound quality.
[0040] In one embodiment, as Figure 3 shown in Figure 3 , a voice synthesis device 30 is provided, such as a vocoder for voice synthesis, which can be located in Figure 1 the electronic device 102 or the server 104 in Figure 1 . The voice synthesis device 30 can execute Figure 2 the steps S201 to S203 in Figure 2 . Combining Figure 2 with Figure 3 specifically described, the voice synthesis device 30 mainly consists of two parts:
[0041] The first part is the voice generation device 303, which can be, for example, a pre-vocoder. It is used to execute steps S201 and S202. Specifically, it receives the acoustic features 301 and generates the first audio signal 304;
[0042] The second part is the voice correction device 305, which can be, for example, a fixer-vocoder, and is used to execute step S203, specifically to correct the received first audio signal 304 and obtain the corrected voice signal 306.
[0043] The voice synthesis device 30 can be some simple voice synthesizers. Under the background of smaller network parameters, the voice synthesis device 30 is used to generate audio.
[0044] The acoustic-related parameter 301 is the input acoustic feature parameter, including pitch. Optionally, the acoustic feature parameter includes Mel spectrum, pitch, and loudness, which will be introduced in detail in the following examples.
[0045] The above Figure 2 The provided voice synthesis method and Figure 3 The provided voice synthesis device first performs a pre-vocoder process on the input acoustic features, such as the pitch in the acoustic feature parameters, to generate a first audio signal, and then corrects the generated first audio signal to obtain a higher-quality synthesized voice signal as the final voice. Because the audio correction process is added in the above two separate processes for the acoustic feature parameters, a higher-quality synthesized voice signal is obtained.
[0046] This embodiment avoids directly converting acoustic features into high-quality voice audio using a complex neural network at one time, has a small number of parameters, and can be applied to scenarios that require a lightweight deployment structure.
[0047] In another embodiment, as Figure 4A shown, another voice synthesis method is provided, which is executed by an electronic device. The electronic device can be Figure 1 the electronic device 102 or the server 104 in , that is, this method can be executed by the electronic device 102 or the server 104. In this embodiment, taking the electronic device as the electronic device 102 as an example for illustration, it includes the following steps:
[0048] S401, obtain the acoustic features corresponding to the voice to be synthesized.
[0049] The acoustic features can be the input acoustic feature parameters, including pitch p. Pitch refers to the height of the voice, indicating the frequency and wavelength of the voice. Optionally, the acoustic feature parameters also include Mel spectrum m and loudness l, which will be introduced in detail in the following examples.
[0050] S402, process the Mel spectrum, pitch, and loudness to obtain a first hidden state.
[0051] Specifically, the Mel spectrum, pitch, and loudness among the above acoustic feature parameters are subjected to first upsampling using the first copy sampling method to obtain a first hidden state, where the first copy sampling method will be described in detail in the following embodiments.
[0052] S403. Sample the Mel spectrum, pitch, and loudness respectively to generate the sampled Mel spectrum, pitch, and loudness.
[0053] Specifically, use the second copy sampling method to perform second upsampling on the Mel spectrum among the above acoustic feature parameters to obtain the sampled Mel spectrum signal; use the first interpolation sampling method to perform third upsampling on the pitch among the above acoustic feature parameters to obtain the sampled pitch signal; use the second interpolation sampling method to perform fourth upsampling on the loudness among the above acoustic feature parameters to obtain the sampled loudness signal.
[0054] S404. Process the first hidden state, the sampled Mel spectrum, pitch, and loudness to obtain conditional features.
[0055] Specifically, splice the first hidden state obtained after sampling in step S402 with the sampled Mel spectrum, pitch, and loudness obtained after sampling in step S403 to obtain conditional features.
[0056] S405. Generate the sine function of each of the N harmonic oscillators, and sum up the sine functions of each of the N harmonic oscillators to obtain a first audio signal.
[0057] Specifically, according to the conditional features, generate the sine function of each of the N harmonic oscillators, where the sine function includes amplitude and frequency, and N is a natural number; sum up the sine functions of each of the N harmonic oscillators to obtain a first audio signal.
[0058] S406. Correct the synthesized first audio signal through a voice correction device to obtain a corrected voice signal.
[0059] Specifically, the voice correction device corrects the audio sample values synthesized by the voice generation device, removes noises, reverberations, electronic sounds and other noises in the audio to improve the pronunciation clarity of the audio, and at the same time repairs the missing parts of the audio, so that the audio synthesized from the voice correction device is high-quality and high-fidelity voice.
[0060] In another embodiment, as Figure 4B shown, a voice synthesis device 40 is provided, such as a vocoder for voice synthesis, which can be located in Figure 1 the electronic device 102 or the server 104. The voice synthesis device 40 can executeFigure 4A Steps S401 to S405 in Figure 4A Combined with Figure 4B Specifically described, the speech synthesis device 40 mainly consists of two parts:
[0061] The first part is the speech generation device 43, which can be, for example, a pre-vocoder, and is used to execute steps S401 and S406, specifically including:
[0062] The first copy sampling module 431 is used to receive the acoustic features corresponding to the speech to be synthesized. The acoustic features include Mel spectrogram, pitch, and loudness, and process the Mel spectrogram, pitch, and loudness to obtain the first hidden state 441. Specifically, the first copy sampling module 431 performs first upsampling processing on the Mel spectrogram, pitch, and loudness in the above acoustic features by using the first copy sampling method to obtain the first hidden state, where the first copy sampling method will be described in detail in the following embodiments;
[0063] The second copy sampling module 432 is used to perform second upsampling processing on the Mel spectrogram in the above acoustic feature parameters by using the second copy sampling method to obtain the sampled Mel spectrogram signal;
[0064] The first interpolation sampling module 433 is used to perform third upsampling processing on the pitch in the above acoustic feature parameters by using the first interpolation sampling method to obtain the sampled pitch signal;
[0065] The second interpolation sampling module 434 is used to perform fourth upsampling processing on the loudness in the above acoustic feature parameters by using the second interpolation sampling method to obtain the sampled loudness signal;
[0066] The conditional feature generation module 435 is used to process the first hidden state, the sampled Mel spectrogram, the sampled pitch, and the sampled loudness to obtain the conditional feature.
[0067] Specifically, the first hidden state obtained after the first copy sampling module 431 performs sampling processing, the sampled Mel spectrogram obtained after the second copy sampling module 432 performs sampling processing, the sampled pitch obtained after the first interpolation sampling module 433 performs sampling processing, and the sampled loudness obtained after the second interpolation sampling module 434 performs sampling processing are concatenated to obtain the conditional feature.
[0068] The first audio signal generation module 436 is used to generate the sine function of each harmonic oscillator in the N harmonic oscillators, and sum the sine functions of each harmonic oscillator in the N harmonic oscillators to obtain the first audio signal 442.
[0069] Specifically, according to this conditional feature, a sine function of each harmonic oscillator among N harmonic oscillators is generated, where the sine function includes an amplitude and a frequency, and N is a natural number; the sine functions of each harmonic oscillator among the N harmonic oscillators are summed to obtain a first audio signal 442.
[0070] The speech synthesis device 40 further includes a second part, a speech correction device 45, which is similar to the speech correction device 305 in Figure 3 and can be, for example, a fixer-vocoder, which is used to perform step S405, specifically to perform correction processing on the received first audio signal 442 and synthesize a corrected speech signal 406. Among them, the speech correction device 45 can use a deep neural network, such as a convolutional neural network, to achieve the correction of audio sampling values.
[0071] Among them, the speech synthesis device 40 can be some simple speech synthesizers, and the speech synthesis device 40 is used to generate audio under the background of smaller network parameters.
[0072] The above Figure 4A provided speech synthesis method and Figure 4B the provided speech synthesis device, because for the input acoustic features, such as Mel spectrogram, pitch, and loudness, a pre-vocoder process is first performed to obtain a first audio signal, where the pre-vocoder process is to perform sampling processing on the input Mel spectrogram, pitch, and loudness to obtain a first hidden parameter, perform sampling processing on the Mel spectrogram, pitch, and loudness respectively to obtain the sampled Mel spectrogram, sampled pitch, and sampled loudness, and then splice the first hidden parameter, the sampled Mel spectrogram, the sampled pitch, and the sampled loudness to obtain a conditional feature, and then generate a sine function of each harmonic oscillator among N harmonic oscillators through N harmonic oscillators, sum the sine functions of each harmonic oscillator among the harmonic oscillators to obtain a first audio signal. Then, correction processing is performed on the obtained first audio signal to obtain a corrected speech signal, that is, the final output audio. In these two steps of processing, not only is the correction processing of the audio increased, but also the audio sampling signal is optimized in the first step, so higher-quality audio output can be obtained.
[0073] In another embodiment, as Figure 5 shown, a speech generation device 50 is provided, which can be, for example, a pre-vocoder. In this embodiment, the function of the speech generation device 50 is similar to that of the speech generation devices 303 and Figure 3 the speech generation device 43 in Figure 4B above, and all are used to generate audio signals. However, the processing method of the speech generation device 50 in this embodiment is specifically described in combination with a fully connected layer structure.
[0074] Among them, Figure 5 FIG. Figure 5 is a schematic diagram of the specific processing flow of the voice generation device 50 in this embodiment. Those skilled in the art know that the voice generation device 50 includes multiple modules for processing acoustic feature parameters. The multiple modules can refer to the module structure diagram of Figure 4B which will not be elaborated here. In this embodiment, the voice generation device 50 can be located alone in Figure 1 the electronic device 102 or the server 104 in. If the voice generation device 50 and the voice correction device in this embodiment are processed in the electronic device 102 and the server 104 respectively, it can reduce the network burden and make the processing in the electronic device 102 and the server 104 faster, and thus obtain higher-quality audio.
[0075] First, different sampling processes are performed on acoustic feature parameters, such as Mel spectrogram m, pitch p, and loudness l.
[0076] Specifically, the Mel spectrogram m is input into the first fully connected layer 511, the pitch p is input into the second fully connected layer 512, and the loudness l is input into the third fully connected layer 513. Among them, the first fully connected layer 511, the second fully connected layer 512, and the third fully connected layer 513 can all be a three-layer fully connected layer structure. Among them, each fully connected layer has three parameters that affect the model: the total number of layers (length) of the fully connected layer, the number of neurons (width) of a single fully connected layer, and the activation function. The first fully connected layer 511, the second fully connected layer 512, and the third fully connected layer 513 with a three-layer fully connected layer structure can better solve the non-linear problem. For example, it has a better effect in solving the non-linear problem in the implementation of vocoder decoding. Further, after being processed by each fully connected layer, the voice signal needs to be normalized and processed by an activation function (LeakyReLU).
[0077] As Figure 5 shown, the specific processing method of the voice generation device 50 in this embodiment is:
[0078] After the first fully connected layer 511 processes the input Mel spectrogram m, it is then normalized and processed by an activation function to obtain an upsampled Mel spectrogram m1;
[0079] After the second fully connected layer 512 processes the input pitch p, it is then normalized and processed by an activation function to obtain an upsampled pitch p1;
[0080] After the third fully connected layer 513 processes the loudness l, it is then normalized and processed by an activation function to obtain an upsampled loudness l1;
[0081] Perform splicing processing on the upsampled Mel spectrum m1, upsampled pitch p1, and upsampled loudness l1 and input them into the gated recurrent unit (GRU) network 52 to obtain the second hidden state h2;
[0082] The first copy sampling module 53 performs first upsampling on the hidden state h2 using the first copy sampling method to obtain the first hidden state h1.
[0083] Among them, the first copy sampling method is to copy each element in the second hidden state h2 (for example, the first upsampling multiple - 1) times: specifically, the second hidden state h2 = [0.5, 0.3, 0.2], when the first upsampling multiple is 2, after performing first upsampling through the first copy sampling method, the obtained first hidden state h1 = [0.5, 0.5, 0.3, 0.3, 0.2, 0.2].
[0084] As Figure 5 shown, the further processing method of the voice generation device 50 in this embodiment is:
[0085] Input the Mel spectrum m into the second copy sampling module 54, and the second copy sampling module 54 uses the second copy sampling method to perform second upsampling processing on the Mel spectrum m to obtain the upsampled Mel spectrum m2; among them, the second copy sampling method is to copy each element in the hidden Mel spectrum m (the second upsampling multiple - 1) times. Specifically, the Mel spectrum m = [0.5, 0.3, 0.2], when the second upsampling multiple is 2, after performing second upsampling through the second copy sampling method, the obtained second sampling signal m2 = [0.5, 0.5, 0.3, 0.3, 0.2, 0.2]. The first copy sampling method and the second copy sampling method can be the same or different, and the first upsampling multiple and the second upsampling multiple can be the same or different.
[0086] Input the pitch p into the first interpolation sampling module 55, and the first interpolation sampling module 55 uses the first interpolation sampling method to perform third upsampling processing on the pitch p to obtain the upsampled pitch p2. Specifically, the first interpolation sampling method is to insert (the third upsampling multiple - 1)*k random values into the elements of the pitch p, where k refers to the number of elements in the pitch p. For example, the random value can be the mean of two adjacent elements in the pitch p, or the value of one of the elements, such as the value of the first element or the last element. For example, the pitch p = [100, 200, 300], the third upsampling multiple is 2, and the random value is the mean, then k = 3, and p2 = [100, 150, 200, 250, 300, 300] is obtained.
[0087] The loudness l is input into the second interpolation sampling module 56. The second interpolation sampling module 56 performs a fourth upsampling process on the loudness l using the second interpolation sampling method to obtain the upsampled loudness l2. Here, the second interpolation sampling method means inserting (the fourth upsampling multiple - 1)*j random values into the elements of the loudness l, where j refers to the number of elements in the loudness l. For example, the random value can be the mean of two adjacent elements in the loudness l, or the value of one of the elements, such as the value of the first element or the last element. For illustration, if the loudness l = [-1, -2, -3], the fourth upsampling multiple is 2, and the random value is the mean, then j = 3, and the obtained upsampled loudness l2 = [-1, -1.5, -2, -2.5, -3, -3]. The first interpolation sampling method and the second interpolation sampling method can be the same or different, and the third upsampling multiple and the fourth upsampling multiple can be the same or different.
[0088] The first hidden state h1, the upsampled Mel spectrum m2, the upsampled pitch p2, and the upsampled loudness l2 obtained above are concatenated and input into the fourth fully connected layer 57. After processing, the conditional feature f is obtained. The fourth fully connected layer 57 can be a four - layer fully connected layer structure. The fourth fully connected layer 57 with a four - layer fully connected layer structure can better solve the non - linear problem. For example, it has a better effect in solving the non - linear problem in realizing the vocoder decoding. Among them, different input signals enter the fourth fully connected layer for processing respectively, and then the signals are normalized and processed by the activation function and output to obtain the conditional intermediate feature f.
[0089] The obtained conditional feature f is input into the fifth fully connected layer 58. The processed audio signal is input into N harmonic oscillators to generate the sine functions of each harmonic oscillator in the N harmonic oscillators, where N is a natural number. The fifth fully connected layer 58 is a one - layer fully connected layer structure. The input conditional feature f enters the fifth fully connected layer 58 and outputs N harmonic oscillation signals: harmonic oscillation signal 1, harmonic oscillation signal 2,..., harmonic oscillation signal N. Each harmonic oscillation signal is composed of an amplitude (A1, A2,..., AN), a frequency (F1, F2,..., FN), and an initial phase (Φ1, Φ2,..., ΦN). The values of the N harmonic oscillators at time t are summed up to obtain the audio at time t.
[0090] According to Figure 5 In the present embodiment shown, the number of parameters of the audio obtained by the voice generation device 50 through the above - mentioned processing method is less than 1MB. Therefore, it greatly reduces the processing time and complexity of calculating the number of parameters, not only reducing the burden of network processing, but also accelerating the audio generation speed.
[0091] As Figure 6As shown, a voice correction device 60 is provided, which can be a corrector for example. The corrector corrects the audio sample values synthesized by the voice generation device, removes noises such as noise, reverberation, and electroacoustic sound in the audio, so as to improve the pronunciation clarity of the audio, and makes the audio synthesized from the voice correction device have a high-quality and high-fidelity voice effect. Among them, the voice correction device can be implemented using a deep neural network (for example, a convolutional neural network). In this embodiment, the function of the voice correction device 60 is the same as that of the Figure 3 voice correction device 304 and Figure 4B voice correction device 44, both of which are used to synthesize high-quality and high-fidelity voices. However, the voice correction device 60 in this embodiment combines an encoder and a decoder to specifically describe the processing method of the audio correction device.
[0092] Among them, Figure 6 is a schematic diagram of the specific processing flow of the voice generation device 60 in this embodiment. Those skilled in the art know that the voice correction device 60 includes multiple modules to process audio signals. In this embodiment, the voice correction device 60 can be located alone in the Figure 1 electronic device 102 or server 104 in. If the voice generation device and the voice correction device 60 in this embodiment are processed in the electronic device 102 and the server 104 respectively, it can reduce the network burden and make the processing in the electronic device 102 and the server 104 faster, and thus obtain higher-quality audio.
[0093] As Figure 6 shown, the specific way for the voice generation device 60 to process the first audio signal synthesized by the voice generation device 50 in the above embodiment is:
[0094] The acoustic feature extraction module 61 extracts acoustic features from the first audio signal synthesized by the voice generation device 50 to obtain Fourier transform features. Specifically, the acoustic feature extraction module 61 performs Fourier transform on the first audio signal (including audio sample values), and extracts acoustic features from the Fourier-transformed audio signal to obtain the Fourier transform features.
[0095] The encoding device 62 (encoder) is used to perform downsampling processing on the Fourier transform features to obtain a downsampled intermediate feature vector. The encoding device 62 is composed of M layers of convolutional networks. The extracted acoustic features are processed by each layer of convolutional network and then followed by normalization and activation function processing, where M is a natural number.
[0096] The recurrent neural (LSTM) network 63 is used to process the downsampled intermediate feature vector to obtain temporal information.
[0097] The decoding device 64 (decoder) is used to perform upsampling processing and inverse Fourier transform processing on the timing information to obtain a corrected speech signal.
[0098] Specifically, the decoding device 64 receives the timing information and performs upsampling processing to obtain clean short-time Fourier transform (STFT) features through training and model learning. Among them, the training and model learning include the authenticity discrimination between the obtained speech and the target speech, and the improvement of the similarity between the obtained speech and the target speech. Further, the decoding device 64 is mainly composed of an M-layer transposed convolutional network. The input of each layer of the transposed convolutional network in the decoding device 64 is the output of the previous layer concatenated with the output of the corresponding structure of the encoding device 62.
[0099] Among them, clean short-time Fourier transform features are obtained through training and model learning. In the specific training and model learning, this embodiment uses an adversarial training strategy to ensure that the obtained audio is closer to the target audio. The discriminator used in the adversarial training is a multi-period discriminator (Multi-Period Discriminator). The role of the discriminator is to separately send the speech obtained after the above upsampling processing and the target audio into the discriminator for authenticity discrimination, and then improve the similarity between the obtained speech and the target speech through training and model learning to obtain the clean short-time Fourier transform speech features. The goal of model learning is to make the discriminator unable to distinguish the authenticity of the audio by improving the similarity between the obtained speech and the target audio, thereby reflecting to a certain extent that the obtained speech is extremely close to the target audio.
[0100] Please refer to Figure 7 , which is a schematic diagram of the processing flow of the encoding device 62, the recurrent neural network 63, and the decoding device 64. For example, if M is 3, the convolutional layer network structure of the encoding device 62 is e1, e2, e3; the transposed convolutional layer network structure of the decoding device 64 is d1, d2, d3, then the input of d1 is the output of the recurrent neural network 63 concatenated with the output of e3, the input of d2 is the output of d1 concatenated with the output of e2, and the input of d3 is the output of d2 concatenated with the output of e1.
[0101] Among them, the number of parameters of the speech correction device 60 in this embodiment is also about 1MB. Among them, high-quality speech signals are generated at high speed through the two low-parameter structures of the speech generation device 50 and the speech correction device 60.
[0102] Combined with Figure 5 the speech generation device in Figure 6Regarding the voice correction device therein, each processing module in the voice generation device and the voice correction device has a lightweight structure. The voice generation device can decode audible audio at an extremely fast speed, and the voice correction device then corrects the audio and converts it into high-quality audio output. Further, the voice generation device and the voice correction device can be separated in different networks or electronic devices to respectively implement the functions of their respective modules. Compared with the prior art method of obtaining audio using a complex network with a large number of parameters, the embodiments of the present disclosure have the advantages of simple and fast network processing, flexible configuration and lighter hardware structure, and higher-quality voice after correction.
[0103] The exemplary embodiments of the present disclosure further provide a voice synthesis device. Figure 8 The schematic diagram of a voice synthesis device 800 according to another exemplary embodiment of the present disclosure is shown. As Figure 8 shown, the device 800 includes: an acquisition unit 801, configured to acquire the acoustic features corresponding to the voice to be synthesized, wherein the acoustic features include pitch, and pitch refers to the height of the voice, indicating the frequency and wavelength of the voice;
[0104] a generation unit 802, configured to process the acoustic features to generate a first audio signal;
[0105] a correction unit 803, configured to correct the first audio signal to obtain a corrected voice signal.
[0106] The exemplary embodiments of the present disclosure further provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiments of the present disclosure.
[0107] The exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0108] The exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiments of the present disclosure.
[0109] Refer to Figure 9, a block diagram of an electronic device 900 that can serve as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0110] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0111] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information into the electronic device 900. The input unit 906 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 907 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 904 can include, but is not limited to, magnetic disks, optical disks. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0112] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above. For example, in some embodiments, methods S202 and S203 can be implemented as computer software programs that are tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to execute methods S202 and S203 in any other suitable way (e.g., by means of firmware).
[0113] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) that provides machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.
[0116] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0117] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or in a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0118] A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A voice synthesis method, comprising: Obtaining acoustic features corresponding to the voice to be synthesized, wherein the acoustic features include pitch, Mel spectrum, and loudness, and the pitch refers to the height of the voice, representing the frequency and wavelength of the voice; Performing hierarchical processing on the acoustic features to generate a first audio signal; Modifying the first audio signal to obtain a modified voice signal; Wherein, the performing hierarchical processing on the acoustic features to generate a first audio signal includes: Processing the Mel spectrum, pitch, and loudness to obtain a first hidden state; Sampling the Mel spectrum, pitch, and loudness respectively to generate a sampled Mel spectrum, pitch, and loudness; Processing the first hidden state, the sampled Mel spectrum, pitch, and loudness to obtain a conditional feature; Generating a sine function of each harmonic oscillator in N harmonic oscillators according to the conditional feature, wherein the sine function includes an amplitude and a frequency, and N is a natural number; Summing the sine functions of each harmonic oscillator in the N harmonic oscillators to obtain the first audio signal.
2. The method according to claim 1, wherein, Performing hierarchical processing on the Mel spectrum, pitch, and loudness to obtain a first hidden state, further including: Inputting the Mel spectrum, pitch, and loudness into a first fully connected layer, a second fully connected layer, and a third fully connected layer respectively to obtain an upsampled Mel spectrum, an upsampled pitch, and an upsampled loudness; Concatenating the obtained upsampled Mel spectrum, upsampled pitch, and upsampled loudness, and inputting them into a gated recurrent unit network for processing to obtain a second hidden state; Performing a first upsampling process on the second hidden state by using a first replication sampling method, to obtain the first hidden state, wherein the first replication sampling method includes replicating each element in the second hidden state (C1 - 1) times, and C1 represents the first upsampling multiple.
3. The method according to claim 1 or 2, wherein The respectively sampling the Mel spectrum, pitch, and loudness to generate a sampled Mel spectrum, pitch, and loudness includes: Performing a second upsampling process on the Mel spectrum by using a second replication sampling method to obtain a sampled Mel spectrum signal, wherein the second replication sampling method includes replicating each element in the Mel spectrum (C2 - 1) times, and C2 represents the second upsampling multiple; Performing a third upsampling process on the pitch by using a first interpolation sampling method to obtain a sampled pitch signal, wherein the first interpolation sampling method includes inserting K random values into the elements of the pitch, and K = (C3 - 1) × k, C3 represents the third upsampling multiple, and k represents the number of elements in the pitch; Performing a fourth upsampling process on the loudness by using a second interpolation sampling method to obtain a sampled loudness signal, wherein the second interpolation sampling method includes inserting J random values into the elements of the loudness, and J = (C4 - 1) × j, C4 represents the fourth upsampling multiple, and j represents the number of elements in the loudness.
4. The method according to claim 1 or 2, wherein The processing the first hidden state, the sampled Mel spectrum, pitch, and loudness to obtain a conditional feature includes: Concatenate the first hidden state, the sampled Mel spectrogram signal, the sampled pitch signal, and the sampled loudness signal, and input them into the fourth fully connected layer for processing to obtain the conditional feature. And, generating a sine function for each of the N harmonic oscillators according to the conditional feature includes: Input the conditional feature into the fifth fully connected layer for processing to obtain a sine function for each of the N harmonic oscillators.
5. The method according to any one of claims 1-2, wherein, The correcting the first audio signal to obtain a corrected speech signal includes: Extract acoustic features from the first audio signal to obtain Fourier transform features; Perform downsampling processing on the Fourier transform features to obtain a downsampled intermediate feature vector; Process the downsampled intermediate feature vector using a recurrent neural network to obtain temporal information; Perform upsampling processing and inverse Fourier transform processing on the temporal information to obtain the corrected speech signal.
6. The method according to claim 5, wherein, The performing downsampling processing on the Fourier transform features to obtain a downsampled intermediate feature vector includes: Use an encoding device to perform downsampling processing on the Fourier transform features to obtain the downsampled intermediate feature vector, where the encoding device consists of M convolutional networks, and each convolutional network is followed by normalization and an activation function, and M is a natural number.
7. The method according to claim 5, wherein, The performing upsampling processing and inverse Fourier transform processing on the temporal information to obtain the corrected speech signal includes: Use a decoding device to upsample the temporal information to obtain clean short-time Fourier transform features; Perform inverse Fourier transform processing on the short-time Fourier transform features to obtain the corrected speech signal, where the decoding device includes M transposed convolutional networks, and each transposed convolutional network is followed by normalization and an activation function, and where the input of each layer of the decoding device is the output of the previous layer concatenated with the output of the corresponding network structure of the encoder device.
8. A speech synthesis device, comprising: An acquisition unit, configured to acquire acoustic features corresponding to the speech to be synthesized, where the acoustic features include pitch, Mel spectrogram, and loudness, and the pitch refers to the height of the speech, indicating the frequency and wavelength of the speech; A generation unit, configured to perform hierarchical processing on the acoustic features to generate a first audio signal; A correction unit, configured to correct the first audio signal to obtain a corrected speech signal; Among them, the generating unit is used to perform hierarchical processing on the acoustic features to generate a first audio signal, including: processing the Mel spectrogram, pitch, and loudness to obtain a first hidden state; respectively sampling the Mel spectrogram, pitch, and loudness to generate the sampled Mel spectrogram, pitch, and loudness; processing the first hidden state, the sampled Mel spectrogram, pitch, and loudness to obtain a conditional feature; generating a sine function of each of the N harmonic oscillators according to the conditional feature, where the sine function includes an amplitude and a frequency, and the N is a natural number; adding the sine functions of each of the N harmonic oscillators to obtain the first audio signal.
9. An electronic device, comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice synthesis device
CN107039033A