Audio synthesis method, terminal device and computer-readable storage medium

By extracting and adjusting the fundamental frequency and spectrum envelope of the audio, a synthetic audio with a spectrum envelope consistent with the original audio is solved, and a higher quality audio synthesis is achieved.

CN114038474BActive Publication Date: 2025-05-27TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111562100.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-05-27
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

During the resampling process, existing audio synthesis technology will lead to loss of signal decimation and errors in interpolation, resulting in large differences in the sound quality of the synthetic audio from the original audio and low quality.

Method used

By extracting the first fundamental frequency and spectrum envelope of the audio to be synthesized, the first fundamental frequency is adjusted to obtain the second fundamental frequency, and the synthetic audio is generated based on the second fundamental frequency and spectrum envelope, so that the spectrum envelope of the synthetic audio is consistent with the audio to be synthesized.

Benefits of technology

While changing the pitch, maintain the consistency of tone, avoid errors caused by signal decimation or interpolation, and improve the sound quality of the synthetic audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038474B_ABST
    Figure CN114038474B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses an audio synthesis method, a terminal device and a computer-readable storage medium, wherein the method comprises: obtaining the audio to be synthesized; extracting a first fundamental frequency and a spectrum envelope from the audio to be synthesized, wherein the first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectrum envelope is used to indicate the timbre of the audio to be synthesized; adjusting the first fundamental frequency to obtain a second fundamental frequency; obtaining synthesized audio according to the second fundamental frequency and the spectrum envelope, wherein the spectrum envelope of the synthesized audio is consistent with the spectrum envelope of the audio to be synthesized. The present application can be applied to audio processing fields such as vocal tuning and vocal synthesis, while changing the pitch, maintaining the timbre characteristics. Compared with the existing resampling technology, the present application avoids the errors caused by signal extraction or interpolation, and improves the sound quality of the output audio while achieving the purpose of tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to an audio synthesis method, a terminal device, and a computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence technology and audio processing technology, it has gradually become possible to meet diverse audio synthesis requirements. For example, in different field requirements, the pitch (tone) of the original audio can be changed or retained to achieve the purpose of pitch correction. Currently, resampling the signal of the original audio can achieve the auditory output of variable speed and pitch. On this basis, by adding a suitable variable speed module, the effect of changing pitch without changing speed can be achieved. However, when resampling the signal, it is inevitable to introduce errors caused by signal extraction loss and interpolation estimation error, which will make the auditory sense of the synthesized audio quite different from that of the original audio in terms of sound quality, and the quality of the synthesized audio is not high. Summary of the Invention

[0003] Embodiments of the present application provide an audio synthesis method, a terminal device, and a computer-readable storage medium, which can improve the sound quality of the synthesized audio.

[0004] In a first aspect, embodiments of the present application provide an audio synthesis method, which includes: obtaining an audio to be synthesized; extracting a first fundamental frequency and a spectral envelope from the audio to be synthesized, where the first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectral envelope is used to indicate the timbre of the audio to be synthesized; adjusting the first fundamental frequency to obtain a second fundamental frequency; and obtaining a synthesized audio according to the second fundamental frequency and the spectral envelope, where the spectral envelope of the synthesized audio is the same as that of the audio to be synthesized. Based on the method described in the first aspect, while changing the fundamental frequency (pitch) of the output audio, the spectral envelope (timbre) of the output audio can be kept the same as that of the input. Compared with the existing resampling technology, the error caused by signal extraction or interpolation is avoided, and the sound quality of the synthesized audio is improved while achieving the purpose of pitch correction.

[0005] In a possible implementation manner, obtaining a synthesized audio according to the second fundamental frequency and the spectral envelope specifically includes: calling a trained first residual network model and a trained second residual network model to process the spectral envelope to obtain a first result and a second result; obtaining a third result according to the first result and the embedding vector of the second fundamental frequency; superimposing the second result and the third result to obtain an acoustic feature; and calling a trained audio synthesis model to process the acoustic feature to obtain a synthesized audio.

[0006] In a possible implementation manner, the trained audio synthesis model is composed of a convolutional layer and N transposed convolutional layers connected in sequence, where N is a positive integer greater than 1. The trained audio synthesis model is called to process the acoustic features to obtain the synthesized audio, which specifically includes: calling the convolutional layer to perform convolutional processing on the acoustic features to obtain a fourth result; calling the N transposed convolutional layers to perform transposed convolutional processing on the fourth result to obtain N transposed convolutional results; linearly interpolating and transforming M of the N transposed convolutional results and then stacking them layer by layer to obtain the synthesized audio; where M is a positive integer less than N.

[0007] In a possible implementation manner, the product of the reduction multiples of the N channels corresponding to the N transposed convolutional layers is the same as the dimensional component of the acoustic features.

[0008] In a possible implementation manner, the method further includes: determining a first training loss value based on the training sample set and the synthesized audio of the training sample set; calling a discriminant model to perform true / false discrimination on the synthesized audio of the training sample set to determine a second training loss value; training the parameters in the first residual network model, the second residual network model, the audio synthesis model, and the discriminant model according to the first training loss value and the second training loss value to obtain the trained first residual network model, the trained second residual network model, and the trained audio synthesis model.

[0009] In a possible implementation manner, the discriminant model includes an average pooling layer and a discriminant layer. The discriminant layer is composed of a convolutional layer, a max pooling layer, and two convolutional layers connected in sequence; the average pooling layer includes a first average pooling layer and a second average pooling layer, and the discriminant layer includes a first discriminant layer, a second discriminant layer, and a third discriminant layer; calling the discriminant model to perform true / false discrimination on the synthesized audio of the training sample set to determine a second training loss value, which specifically includes: calling the first discriminant layer to perform true / false discrimination on the synthesized audio of the training sample set to obtain a first discrimination result; calling the first average pooling layer to perform average pooling processing on the synthesized audio of the training sample set to obtain a first average pooling result; calling the second discriminant layer to perform true / false discrimination on the first average pooling result to obtain a second discrimination result; calling the second average pooling layer to perform average pooling processing on the first average pooling result to obtain a second average pooling result; calling the third discriminant layer to perform true / false discrimination on the second average pooling result to obtain a third discrimination result; determining the second training loss value according to the first discrimination result, the second discrimination result, and the third discrimination result.

[0010] In a possible implementation manner, determining the first training loss value based on the training sample set and the synthesized audio of the training sample set includes: determining the first Mel spectrogram of the training sample set and the second Mel spectrogram of the synthesized audio of the training sample set; determining the first training loss value as the minimum mean square error between the first Mel spectrogram and the second Mel spectrogram.

[0011] In a possible implementation, the method further includes: obtaining an original sample set, where the original sample set includes a plurality of dry audio samples, and each dry audio sample in the plurality of dry audio samples is represented by a pitch distribution vector; determining the sampling probability of each dry audio sample based on the pitch distribution vector of each dry audio sample, where the sampling probability is used to make the pitch of the sampled training sample set evenly distributed; and sampling the original sample set based on the sampling probability of each dry audio sample to obtain a training sample set.

[0012] In a possible implementation, the synthesized audio of the training sample set is obtained based on the fundamental frequency and spectral envelope of the training sample set.

[0013] In a second aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor; the above-mentioned memory is used to store a computer program; the above-mentioned processor is used to call the computer program from the memory, so that the terminal device executes any one of the methods in the first aspect above.

[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-readable instructions are stored. When the computer-readable instructions run on the terminal device in the second aspect above, the terminal device is enabled to execute any one of the methods in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a computer program or computer program product, including code or instructions. When the code or instructions run on a computer, the computer is enabled to execute any one of the methods in the first aspect above. Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below.

[0017] Figure 1 It is a schematic diagram of the spectrum of a voice provided by an embodiment of the present application;

[0018] Figure 2 It is a schematic flowchart of an audio synthesis method provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic flowchart of the application stage of an audio synthesis method provided by an embodiment of the present application;

[0020] Figure 4 It is a schematic diagram of a filtered sine signal provided by an embodiment of the present application;

[0021] Figure 5 It is a schematic diagram of the architecture of a feature splicing network provided by an embodiment of the present application;

[0022] Figure 6 It is a schematic diagram of the architecture of a residual network model provided by an embodiment of the present application;

[0023] Figure 7 It is a schematic diagram of the architecture of an audio synthesis model provided by an embodiment of the present application;

[0024] Figure 8 It is a schematic flowchart of a training method for an audio processing model provided by an embodiment of the present application;

[0025] Figure 9 It is a schematic diagram of the architecture of a discriminant model provided by an embodiment of the present application;

[0026] Figure 10 It is a schematic diagram of the structure of an audio synthesis device provided by an embodiment of the present application;

[0027] Figure 11 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0029] In the specification, claims and drawings of the present application, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0030] To better understand the solution of the present application, the following introduces the professional terms of the present application first:

[0031] ①. Pitch: Speech is composed of the superposition of simple harmonic vibrations of many frequencies. In the spectrogram of speech, the peak corresponding to the vibration with the lowest frequency is the fundamental tone, and the rest are overtones. The pitch is the frequency corresponding to the fundamental tone, and can also be called the fundamental frequency or tone. What is usually referred to as "out of tune" means that the pitch of the singer does not match the pitch of the note.

[0032] ②. Spectral envelope: When the sound wave generated by vocal cord vibration passes through the vocal tract composed of the oral cavity, nasal cavity, etc., resonance will occur. As a result of resonance, certain regions of the spectrum will be strengthened to form peaks. There are multiple peaks on the speech spectrum, and the heights of each peak on the spectrum are different. The ratio of the heights of these peaks determines the timbre of a person. If these peak values are connected by a smooth curve, it is the spectral envelope. Figure 1 It is a schematic diagram of the spectrum of a kind of speech. In this diagram, the frequency value corresponding to the first peak is the fundamental frequency, and the black smooth curve connecting each vertex is the spectral envelope.

[0033] ③. Periodic information and aperiodic information: Speech consists of periodic signals and aperiodic signals. The spectrum of a periodic signal is a discrete spectrum, and the peak values of each discrete frequency point can form a spectral envelope. The spectrum of an aperiodic signal is a continuous spectrum and has no spectral envelope. Therefore, only by combining the information in periodic signals and aperiodic signals can the original signal be perfectly synthesized.

[0034] ④. Mel spectrum: The audible speech frequency range of the human ear is 20 - 20 kHz, but the human ear does not have a linear perception relationship with the spectrum on the Hz scale. For example: When the listener gets used to a tone of 1000 Hz, if the frequency of the tone is increased to 2000 Hz, the listener's ear can only perceive that the frequency has increased a little bit, rather than doubling. Therefore, a Mel-scale filter bank is used to transform the linear spectrum group into a Mel spectrum group and convert the linear spectrum scale into a Mel spectrum scale to simulate the linear perception relationship of the human ear to frequency. That is to say, under the Mel spectrum scale, if the Mel spectra of two pieces of speech differ by a factor of two, the tones perceived by the human ear will also differ by approximately a factor of two.

[0035] ⑤. Source-filter model: This model regards sound as composed of an excitation and a corresponding filter. The excitation is equivalent to the vocal cords of the vocal structure, and the filter is equivalent to the human vocal tract and resonance cavity. The sound source excitation generates a voiced signal generated by a periodic pulse sequence and a silent signal generated by white noise excitation. Among them, the voiced signal is a quasi-periodic signal and contains the above-mentioned periodic information, and the silent signal is an aperiodic signal and includes the above-mentioned aperiodic information.

[0036] See Figure 2 , which is a schematic flowchart of an audio synthesis method provided by an embodiment of the present application. This process can be implemented by an audio synthesis device. The audio synthesis device can be a terminal device, or a device in the terminal device, or a device that can be used in matching with the terminal device. Exemplarily, the terminal device can be a device with data processing functions and input / output functions, including but not limited to devices such as computers, smartphones, and tablets. Among them, the audio synthesis device includes an audio processing model, and the audio processing model includes a residual network model and an audio synthesis model. Through the audio processing model, a synthesized audio with higher sound quality can be synthesized.

[0037] Figure 2 The process shown includes a training phase and an application phase. In the training phase, the audio signal is a training sample set of audio. The training sample set is successively subjected to feature extraction and feature combination to obtain acoustic features. Then, the acoustic features are input into an audio synthesis model to obtain synthesized audio, and the synthesized audio is input into a discriminant model for true / false discrimination. Additionally, the training sample set of audio and the synthesized audio are processed to extract Mel spectrograms. Based on the obtained Mel spectrograms, the minimum mean square error between the two can be determined. Based on the minimum mean square error and the results of true / false discrimination, the residual network model (applied to Figure 2 the feature combination phase), the audio synthesis model, and the discriminant model can be trained to determine the processing parameters in the above models.

[0038] In the application phase, the audio signal is the audio to be synthesized. The audio to be synthesized is subjected to feature extraction to obtain the fundamental frequency and spectral envelope. Then, by adjusting the fundamental frequency and combining the adjusted fundamental frequency with the spectral envelope (the residual network model involved in this phase has been determined with processing parameters after training) to obtain acoustic features. Finally, the acoustic features are input into the audio synthesis model (the audio synthesis model involved in this phase has been determined with processing parameters after training) to obtain synthesized audio.

[0039] Through Figure 2 the process shown, the model parameters involved in the audio synthesis method can be trained, and finally a trained model can be obtained. This model is used to adjust the fundamental frequency of newly input audio while keeping the spectral envelope (timbre) of the synthesized audio consistent with the spectral envelope of the newly input audio. Compared with the existing resampling technology, it avoids the errors caused by signal decimation or interpolation, improves the quality of the synthesized audio while achieving the purpose of pitch correction. The specific implementation manner of the above application phase will be described in detail in combination with the following Figure 3 embodiment, and the specific implementation manner of the training phase will be described in detail in combination with the following Figure 8 embodiment.

[0040] See Figure 3 , which is a schematic flowchart of the application phase of an audio synthesis method provided by an embodiment of the present application. This method is applied to the above audio synthesis device and includes steps S301 to S304, where:

[0041] S301. Obtain the audio to be synthesized.

[0042] In the embodiments of the present application, the audio to be synthesized is a pure dry audio without background audio. Exemplarily, when the present application is applied to audio processing of a music software, the audio to be synthesized can be a vocal singing audio without accompaniment or music. In a possible implementation manner, the method for obtaining the audio to be synthesized can be: obtaining the dry audio recorded by the user in real time through a module with an input function (exemplarily, a microphone, etc.) in the above audio synthesis device, calling the dry audio stored in the above audio synthesis device, or receiving the dry audio sent from other devices, etc., and the present application does not limit this.

[0043] It should be noted that when the embodiments of the present application are applied to specific products and technologies, before obtaining relevant data such as the audio to be synthesized, user permission or consent needs to be obtained, and the collection, use, and processing of these relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0044] S302. Extract a first fundamental frequency and a spectral envelope from the audio to be synthesized. The first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectral envelope is used to indicate the timbre of the audio to be synthesized.

[0045] The purpose of this step is to extract the audio features of the audio to be synthesized, and these audio features are related to the vocal characteristics of the audio to be synthesized.

[0046] Among them, the value of the first fundamental frequency is the lowest vibration frequency of the audio to be synthesized. The larger this value is, the higher the pitch of the audio to be synthesized. When the audio to be synthesized is a periodic signal, it is composed of the waveform corresponding to the first fundamental frequency and the second harmonic, third harmonic, etc. Due to the reciprocal relationship between frequency and time, the waveform corresponding to the first fundamental frequency is the waveform with the longest period among the above waveforms, and the value of the first fundamental frequency is the reciprocal of the period of this waveform.

[0047] For the above extraction of the first fundamental frequency from the audio to be synthesized, in a possible implementation manner, it includes: filtering the audio to be synthesized using low-pass filters with different cut-off frequencies, calculating the candidate fundamental frequencies and their credibility in each filtered signal (sinusoidal signal), and selecting the candidate fundamental frequency with the highest credibility as the first fundamental frequency. Exemplarily, as Figure 4 shown, it is a filtered sinusoidal signal. This sinusoidal signal includes intervals (1), (2), (3), and (4), and the sizes of the four intervals are basically equal. The candidate fundamental frequency is the reciprocal of the average value of the four intervals, and the credibility of the candidate fundamental frequency is the standard deviation of the four intervals. If the standard deviation of the four intervals is smaller, it means that the lengths of the four intervals do not differ much, and the credibility of this candidate fundamental frequency is higher.

[0048] For the above extraction of the first fundamental frequency from the audio to be synthesized, in another possible implementation, it includes: when the audio to be synthesized has periodicity, the autocorrelation function of the audio to be synthesized also has periodicity, and the period is the same as that of the audio to be synthesized. At integer multiples of the period of the periodic signal, the autocorrelation function of the signal can reach the maximum value. Therefore, the starting time of the signal can be ignored, and only the shift distance between two points where the maximum autocorrelation function values are obtained needs to be estimated. This shift distance is also the period of the signal (the reciprocal of the fundamental frequency).

[0049] Since audio signals are all non-stationary signals, when processing the audio to be synthesized, the short-time autocorrelation function is used to intercept the audio to be synthesized with a short-time window and perform short-time autocorrelation calculation. Exemplarily, the audio to be synthesized can be sampled and framed first, and the autocorrelation function is calculated for each audio frame. The expression of the autocorrelation function is shown in Formula 1-1, where s(n) represents the sampling value of the audio frame at the nth point, w(m) represents the window function, N represents the width of the window function, and τ represents the shift distance. By finding the τ corresponding to the maximum value of the autocorrelation function obtained from the expression, the magnitude of the first fundamental frequency can be estimated.

[0050]

[0051] It should be noted that the extraction of the first fundamental frequency from the audio to be synthesized can also be implemented by other methods, and this application does not limit this.

[0052] In the frequency-amplitude diagram of the speech signal, the resonance peaks of each frequency are connected by a smooth curve, and this smooth curve is the spectral envelope. For the above extraction of the spectral envelope from the audio to be synthesized, in one possible implementation, the audio to be synthesized can be windowed first, and the power spectrum after windowing is calculated; secondly, the power spectrum is smoothed, and the cepstrum of the power spectrum (inverse Fourier transform) is obtained. Since the image of the low-frequency part of the cepstrum of the power spectrum is roughly the same as the spectral envelope, the cepstrum is input into a low-pass filter for filtering to obtain the spectral envelope of the audio to be synthesized. It should be noted that the extraction of the spectral envelope from the audio to be synthesized can also be implemented by other methods, and this application does not limit this.

[0053] In addition, it can be understood that the above first fundamental frequency and spectral envelope can only be used to reflect the audio characteristics contained in the audio to be synthesized when it is a periodic signal. When the audio to be synthesized also includes aperiodic information, the above first fundamental frequency, spectral envelope or the waveform of the audio to be synthesized can be combined to extract other audio characteristics.

[0054] S303. Adjust the first fundamental frequency to obtain the second fundamental frequency.

[0055] The first fundamental frequency obtained through the above steps contains the original pitch information of the audio to be synthesized. However, in specific application scenarios, users need to adjust the first fundamental frequency to achieve the purpose of changing the pitch. Exemplarily, in a Karaoke scenario, the first fundamental frequency comes from the user's dry singing audio, and the value of this first fundamental frequency is inconsistent with the pitch of the original song (or accompaniment). Then, the user can adjust the value of the first fundamental frequency according to the difference between the first fundamental frequency and the original song (or accompaniment) to obtain a second fundamental frequency, which can achieve the purpose of correcting the singing voice. Or, in some scenarios (for example, in application scenarios that require voice disguise for security and confidentiality), users need to change a male voice to a female voice or to a voice with other pitches. Since the pitch of male voices is generally higher than that of female voices, the value of the first fundamental frequency can be appropriately increased to obtain a second fundamental frequency similar to that of a female voice. Or, in daily life, the pitch of a mobile voice assistant, a question-and-answer robot, an e-reader, and a virtual diva can be changed through the above steps.

[0056] S304. Obtain a synthesized audio according to the second fundamental frequency and the spectral envelope, and the spectral envelope of the synthesized audio is consistent with the spectral envelope of the audio to be synthesized.

[0057] In a possible implementation manner, step S304 specifically includes: calling a trained first residual network model and a trained second residual network model to process the spectral envelope to obtain a first result and a second result; obtaining a third result according to the embedding vector of the first result and the second fundamental frequency; superimposing the second result and the third result to obtain an acoustic feature; calling a trained audio synthesis model to process the acoustic feature to obtain a synthesized audio.

[0058] In this possible implementation manner, by performing feature splicing on the second fundamental frequency and the spectral envelope obtained through the above processing, an acoustic feature is obtained and used as the input of the trained audio synthesis model. Specifically, Figure 5 is a schematic diagram of the architecture of a feature splicing network provided in this application. Among them, the spectral envelope is processed by two trained residual networks to obtain a first result and a second result. Exemplarily, the schematic diagrams of the two residual networks are as Figure 6 shown, including a residual part 601 and a direct mapping part 602. That is to say, the relationship between the input and output of the residual network model can be represented by formula 1-2.

[0059] x′ = x + f(x, W) (1-2)

[0060] Where: f(x, W) represents the residual part, which is generally composed of two or three convolution operations, and W represents all network parameters in the residual part. Exemplarily, Figure 6The residual part shown includes two convolutional operations with a convolutional kernel size of 3*3. Batch normalization (BN) is used in the two convolutional operations. Since when training deep neural networks such as the residual network model, except for the input data of the input layer (the input data of the input layer will be pre-normalized), the distribution of the input data of each subsequent layer of the network will be affected by the parameter tuning of the previous network, resulting in a decrease in the training speed. Therefore, BN processing can normalize all the input data of the network to a normal distribution with a mean of 0 and a variance of 1, so that a larger learning rate can be used when training the entire network, greatly improving the training speed. It can be understood that to avoid the influence of BN processing on the features learned by the previous layer of the network, reconstruction parameters can be introduced to restore the features learned by a certain original layer. This BN processing is also used between the convolutional kernel and the rectified linear unit in the residual part 601. In addition, the residual part uses an activation function to increase the non-linear mapping ability of the model. Exemplarily, Figure 6 the residual part in uses the rectified linear unit (relu) as the activation function. The relu function makes a part of the output in the network be 0, thereby improving the sparsity of the network and alleviating the occurrence of the overfitting problem. It should be noted that for different training environments (such as different input data), the choice of the activation function may not be the same. This application only takes the relu function as an example for illustration, and does not limit the specific usage of the activation function.

[0061] In the above expression, x represents the direct mapping part 602. In a possible implementation, if the number of output feature maps of the direct mapping part and the residual part is different, a convolutional operation can be used on the output of the direct mapping part to perform dimensionality increase or dimensionality reduction.

[0062] It can be understood that the architectures of the first trained residual network model and the second trained residual network model are generally the same (see the example in Figure 6 ). The parameters in the first trained residual network model and the second trained residual network model can be the same, partially the same, or completely different. The spectral envelope respectively obtains a first result and a second result through these two residual networks. The second result contains the non-linear information of the audio. And the first result is used to multiply with the embedding vector of the fundamental frequency to obtain a third result, and the third result is used to indicate the linear information in the audio. Exemplarily, a fundamental frequency vector table needs to be pre-constructed for the embedding vector of the fundamental frequency. Each fundamental frequency can find a vector of a specified dimension in the fundamental frequency vector table according to its value. Each dimension corresponds to a frequency feature of the fundamental frequency. Exemplarily, the dimension of the fundamental frequency vector table can be 1200*1025, where each fundamental frequency can find a vector with a dimension of 1025 corresponding to its own dimension from the fundamental frequency vector table.

[0063] In a possible implementation, the trained audio synthesis model described above is composed of a convolutional layer and N transposed convolutional layers connected in sequence, where N is a positive integer greater than 1. When the trained audio synthesis model is called to process acoustic features to obtain synthesized audio, it specifically includes: calling the convolutional layer to perform convolutional processing on the acoustic features to obtain a fourth result; calling the N transposed convolutional layers to perform transposed convolutional processing on the fourth result to obtain N transposed convolutional results; linearly interpolating and transforming M of the N transposed convolutional results and then stacking them layer by layer to obtain the synthesized audio; where M is a positive integer less than N.

[0064] Exemplarily, Figure 7 FIG. [FIGURE NUMBER] is a schematic diagram of the architecture of a trained audio synthesis model. Among them, the acoustic features obtained through the above processing are input into a convolutional layer, and the result after convolution is successively upsampled and amplified by 5 transposed convolutions on the convolutional features. It can be understood that the upsampling method in deep learning is mainly applied in the field of computer vision to restore low-resolution images to high-resolution images. When this method is applied to this application, it can further extract and amplify the acoustic features in the audio to achieve the effect that the sound quality of the synthesized audio is higher than that of the original audio. Since the interpolation in traditional upsampling methods (such as nearest neighbor interpolation, bilinear interpolation, etc.) is similar to the manually established feature engineering, the neural network on which the audio synthesis model is based will not learn this during training, while using the above transposed convolution has learnable parameters and can perform upsampling optimally.

[0065] In addition, each transposed convolution layer will amplify the time dimension and shrink the channel dimension. The reason for shrinking the channel dimension is that as the number of layers increases and the size of the feature maps extracted by each layer shrinks, each feature map contains more features. Therefore, the number of channels is shrunk to reduce the number of feature maps output by each layer.

[0066] In a possible implementation, the product of the N shrinking or amplifying multiples is the same as the dimensional component of the acoustic features. Exemplarily, the input acoustic feature dimension is (T, 512). After 5 transposed convolutional layers, the amplification or reduction multiples of the transposed convolution are [8, 4, 4, 2, 2] respectively, then the dimension of the output synthesized audio is (T * 512, 1). Where T is the number of frames of the acoustic features, and 512 represents the number of samples contained in one frame. For example, if the audio to be synthesized is sampled to 32 kHz and each frame of audio data is 16 ms, then the number of samples contained in one frame is 32 * 16 = 512. It should be noted that there may be other values for the number N of the above transposed convolutional layers and the amplification or reduction multiples of the N transposed convolutional layers, and this application does not limit this.

[0067] After selecting the transposed convolution results of M out of the above N transposed convolutions and performing linear interpolation transformation, they are stacked layer by layer to obtain the synthesized audio; M is a positive integer less than N. For example, in Figure 7 In the audio synthesis model shown, the transposed convolution results of the second, third, fourth, and fifth transposed convolution layers are selected from the transposed convolution results of the 5-layer transposed convolution layer. The transposed convolution result of the second transposed convolution layer is linearly interpolated and then stacked with the transposed result of the third transposed convolution layer to obtain the first stacked result. The first stacked result is linearly interpolated and then stacked with the transposed convolution result of the fourth transposed convolution layer to obtain the second stacked result. The second stacked result is linearly interpolated and then stacked with the transposed convolution result of the fifth transposed convolution layer to obtain the third stacked result, which is the synthesized audio.

[0068] Among them, extracting the results (multiple feature maps) of multiple transposed convolutions can improve the quality of synthesizing audio through feature maps. And since the sizes of the feature maps increase layer by layer, linear interpolation transformation operations are required before stacking layer by layer to obtain the synthesized audio.

[0069] It should be noted that the hyperparameters of the convolution kernels (such as the size of the convolution kernel (usually odd numbers 1, 3, 5, 7), the number, the stride, etc.) used in the above trained audio synthesis model depend on the actual application scenario, and this application does not limit them.

[0070] Based on Figure 3 the described embodiments, a neural network can be used to perform non-linear modeling on the features extracted from the spectral envelope and the fundamental frequency. While changing the fundamental frequency, ensure that the spectral envelope of the output audio is consistent with the spectral envelope of the input audio (maintaining the timbre characteristics). Compared with the existing resampling techniques, this method avoids the errors caused by signal decimation or interpolation, and improves the sound quality of the output audio while achieving the purpose of pitch correction.

[0071] In the above content, a method for audio synthesis of the audio to be synthesized using the audio synthesis device of the embodiments of this application is introduced. The following content will further introduce the training method of the audio processing model in this audio synthesis device.

[0072] See Figure 8 , which is a schematic flowchart of a training method for an audio processing model provided by an embodiment of this application. The audio processing model includes a first residual network model, a second residual network model, an audio synthesis model, and a discriminant model. These models can be neural network models. The training method of the audio processing model is applied to the above audio synthesis device, including steps S801 to S803, where:

[0073] S801. Determine the first training loss value based on the training sample set and the synthesized audio of the training sample set.

[0074] Among them, before step S801, it is also necessary to obtain a training sample set. The methods for obtaining the training sample set include: obtaining an original sample set, where the original sample set includes multiple dry sound audio samples, and each dry sound audio sample in the multiple dry sound audio samples is represented by a pitch distribution vector; determining the sampling probability of each dry sound audio sample based on the pitch distribution vector of each dry sound audio sample, and the sampling probability is used to make the pitch of the sampled training sample set evenly distributed; sampling the original sample set based on the sampling probability of each dry sound audio sample to obtain the training sample set. Among them, when the embodiments of the present application are applied to specific products and technologies, before obtaining relevant data such as the training sample set, user permission or consent needs to be obtained, and the collection, use, and processing of these relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. The sampling probability can be understood as the probability value of each dry sound audio sample being used as a training sample. Generally speaking, the more the number of dry sound audio samples in a certain pitch distribution range, the smaller the sampling probability of the dry sound audio samples corresponding to this pitch distribution range. That is to say, the possibility of a certain dry sound audio sample in this pitch distribution range being used as a training sample is smaller. Based on this method, it can be ensured that the pitch distribution of the samples included in the training sample set for training the audio synthesis method is balanced, thereby making the training effect of the audio processing model better and the generalization ability higher.

[0075] Exemplarily, the vocal frequency of a human is between 60 Hz and 1200 Hz. In order to enable the finally trained audio processing model to have good synthesis ability for the audio in this frequency band and improve the synthesized audio, the training sample set needs to cover 60 Hz to 1200 Hz, and the pitch of the training sample set needs to be evenly distributed in this frequency band. For example, the original sample set includes L dry audio samples. The acquisition method of each dry audio sample can refer to the method for acquiring the audio to be synthesized introduced in the above step S301, which will not be elaborated here. Analyze each dry audio sample to obtain the pitch distribution vector n corresponding to each dry audio sample. Exemplarily, the dimension of the pitch distribution vector n can be 1200 (generally speaking, the choice of dimension is based on the maximum value of the frequency to be covered). Among them, the pitch distribution vector n corresponding to each dry audio sample is used to represent the number of frames of the dry audio sample at each pitch, or understood as the distribution situation at each pitch. Exemplarily, the total duration of a certain dry audio sample is 1600 milliseconds, and the duration of each frame in the dry audio sample is 16 milliseconds. Then the 1600-millisecond dry audio sample includes 1000 frames. If there are 0 frames distributed at 60 Hz among these 1000 frames, the value of the dimension of the pitch distribution vector n of this dry audio sample corresponding to 60 Hz is 0; if there are 10 frames distributed at 300 Hz among these 1000 frames, the value of the dimension of the pitch distribution vector n of this dry audio sample corresponding to 300 Hz is 10, etc. The corresponding values for other pitches will not be exemplified one by one here. Therefore, for m dry audio samples, a pitch distribution matrix A of L*n can be obtained. In order to achieve pitch distribution balance, it is necessary to solve the sample sampling probability corresponding to the L dry audio samples according to the pitch distribution matrix A. The calculation method of this sample sampling probability can refer to formula 1-3:

[0076] A T X = B, s.t. X i > 0

[0077] B i = sum(A) / L (1-3)

[0078] Among them, A is a pitch distribution matrix of L*n, X is the sample sampling probability corresponding to the audio sample distribution, B is a vector with dimension n*1, and each element of the B vector is equal. Through the above expression, the value of X that makes each vector in B equal (that is to say, the pitch distribution is uniform) can be obtained. The value of each element in X is the sampling probability of each dry audio sample in the L dry audio samples.

[0079] It can be understood that in the specific implementation process, for the dry audio samples with a higher sampling probability, it can be achieved by duplicating the dry audio samples. Specifically, if the training sample set requires 1000 dry audio samples, and the original sample set includes 500 dry audio samples, and the sampling probabilities of the 500 dry audio samples are [0.005, 0.004, 0.002,...], then the number of dry audio samples to be obtained from the original sample set for the 1000 dry audio samples corresponds to [5, 4, 2,...]. Then, it is necessary to duplicate the first dry audio sample 5 times and add it to the training sample set, duplicate the second dry audio sample 4 times and add it to the training sample set, duplicate the third dry audio sample 2 times and add it to the training sample set, and so on. There can be other ways to achieve the sampling probability, such as duplicating the identifiers of the dry audio samples (such as file names, etc.), and the embodiments of the present application do not limit this.

[0080] Among them, the synthesized audio of the above training sample set is obtained based on the fundamental frequency and spectral envelope of the training sample set. Specifically, assume that the training sample set includes X dry audio samples, and the X dry audio samples can be divided into S batches, with each batch containing Y dry audio samples. Each time, a batch of dry audio samples is input into the audio processing model for training. For each of the Y dry audio samples among the Y dry audio samples, the fundamental frequency and spectral envelope are extracted from the dry audio sample. The fundamental frequency is used to indicate the pitch of the dry audio sample, and the spectral envelope is used to indicate the timbre of the dry audio sample. The Y pairs (fundamental frequency, spectral envelope) of the Y dry audio samples are combined in terms of features to obtain the acoustic features of the Y dry audio samples. The acoustic features of the Y dry audio samples are input into the audio synthesis model to generate Y synthesized audio. Among them, the way of feature combination can refer to the above Figure 5 and Figure 6 for the detailed description, and the way of synthesizing the audio can refer to the above Figure 7 for the detailed description, which will not be elaborated here.

[0081] Among them, for the Y synthesized audio of the above Y dry audio samples, Mel spectrograms are extracted for each dry audio sample and its corresponding synthesized audio. Specifically, Mel spectrograms are extracted for each audio frame among the multiple audio frames included in each dry audio sample and each synthesized audio to obtain the first Mel spectrogram and the second Mel spectrogram. Among them, the first Mel spectrogram or the second Mel spectrogram includes a continuous plurality of Mel spectral data between the start frame and the end frame.

[0082] Exemplarily, the first audio is a dry audio sample or a synthesized audio of a dry audio sample. Taking the first audio as an example, the specific process of Mel spectrum extraction will be described next. First, it is necessary to perform frame division on the first audio to obtain multiple audio frames of the first audio, and then extract the Mel spectrum of each audio frame. This Mel spectrum is close to the human auditory perception, which is beneficial for the user to determine the error between the synthesized audio and the dry audio. In the embodiment of the present application, the Mel spectrum can be obtained by performing a Fourier transform on each audio frame and then inputting the linear spectrum obtained by the Fourier transform of each audio frame into a Mel filter bank. Here, taking the Mel spectrum diagram (i.e., Mel spectrum) output by the Mel filter bank as an example of a 128×R-dimensional Mel spectrum diagram, where 128 represents the number of Mel filter banks and R represents the frame length of the audio. Then, from the multiple Mel spectra corresponding to the multiple audio frames included in the first audio, a continuous plurality of Mel spectrum data between the start frame and the end frame are obtained as the Mel spectrum of the first audio.

[0083] After obtaining the first Mel spectrum of Y dry audio samples and the second Mel spectrum of Y synthesized audios based on the above processing, the loss function (the first training loss value) of the audio synthesis model is determined according to Formula 1-4:

[0084] loss g =MES(Mel(y),Mel(y′)) (1-4) Where Mel(y), Mel( ′ ) respectively represent the first Mel spectrum and the second Mel spectrum obtained by performing Mel spectrum extraction on a dry audio sample y and the synthesized audio y′ corresponding to a dry audio sample. MES represents the mean square error calculation, that is, the Euclidean distance between the first Mel spectrum of Y dry audio samples and the second Mel spectrum of Y synthesized audios. By using the above formula to calculate the Euclidean distance between the first Mel spectrum and the second Mel spectrum, the audio features of the synthesized audio corresponding to the dry audio sample are made to approach each other, that is, they have similar fundamental frequencies and spectral envelopes, and the sound quality of the synthesized audio is improved. When the value of the loss function continuously shrinks to a preset threshold or tends to be stable during training, it represents the end of the training phase. Each time during training, the value of the loss function is determined by the loss functions of the Y dry audio samples and their corresponding Y synthesized audios in the current batch.

[0085] It can be understood that the distance calculated by the above loss function between the first Mel spectrum and the second Mel spectrum can be not only the Euclidean distance, but also the Manhattan distance, the Chebyshev distance, etc., which is specifically determined according to the actual application scenario, and the present application does not limit this here.

[0086] S802. Invoke the discriminant model to perform true / false discrimination on the synthesized audios in the training sample set, and determine the second training loss value.

[0087] See Figure 9 , which is a schematic diagram of the architecture of a discrimination model provided by this application. Among them, the discrimination model includes an average pooling layer and a discrimination layer. The discrimination layer is composed of a convolutional layer, a max pooling layer, and two convolutional layers connected in sequence; the average pooling layer includes a first average pooling layer and a second average pooling layer, and the discrimination layer includes a first discrimination layer, a second discrimination layer, and a third discrimination layer; call the discrimination model to discriminate the authenticity of the synthesized audio of the training sample set, and determine the second training loss value, specifically including: calling the first discrimination layer to discriminate the authenticity of the synthesized audio of the training sample set to obtain a first discrimination result; calling the first average pooling layer to perform average pooling processing on the synthesized audio of the training sample set to obtain a first average pooling result; calling the second discrimination layer to discriminate the authenticity of the first average pooling result to obtain a second discrimination result; calling the second average pooling layer to perform average pooling processing on the first average pooling result to obtain a second average pooling result; calling the third discrimination layer to discriminate the authenticity of the second average pooling result to obtain a third discrimination result; determine the second training loss value according to the first discrimination result, the second discrimination result, and the third discrimination result.

[0088] Although the synthesized audio obtained through the above step S801 is close to the dry audio sample in the overall spectrogram, the sound quality optimization ability in the high-frequency part is poor (the sound quality optimization ability in the low-frequency part is good), resulting in high-frequency blurring and reducing the authenticity of the synthesized audio (exemplarily, the human ear can feel that it is an artificially synthesized audio, and there may even be an obvious electroacoustic effect). Therefore, this application adds a discrimination model in the training stage to discriminate the authenticity of the synthesized audio, thereby improving the authenticity of the synthesized audio from multiple scales.

[0089] Specifically, take a dry audio sample and its corresponding synthesized audio as an example for illustration. It can be understood that the dry audio sample is real audio data, so the label is true, and the synthesized audio is artificially synthesized audio data, so the label is false. Input the dry audio sample and the synthesized audio with labels into the discrimination layer, and the discrimination layer will learn the features in the dry audio sample and the synthesized audio (the discrimination layer contains three convolutional layers, and uses an average pooling layer to reduce the computational amount of parameters and prevent overfitting during training), and obtain the probability that the audio input into the discrimination model is true or false. Input this probability into the following cross-entropy loss function (Formula 1-5) to calculate the loss (the second training loss value) of the discrimination model:

[0090]

[0091] Among them, Y represents the number of dry audio samples or the number of synthesized audio input in the current batch; y i is the label of the audio input into the discrimination model. When the label is true, y i takes the value of 1. When the label is false, yi takes the value of 0; p i is the probability that the discrimination model judges the input audio to the discrimination model as true, 1 - i is the probability that the discrimination model judges the input audio to the discrimination model as false. The loss of each training of the discrimination model is calculated using the above formula, so as to promote the discrimination model to improve the probability of judging true audio as true and false audio as false. That is to say, the discrimination model can accurately distinguish the true and false of the input audio as much as possible. When the value of this loss function continuously shrinks to a preset threshold or stabilizes with training, it represents the end of the training phase.

[0092] S803. According to the first training loss value and the second training loss value, train the parameters in the first residual network model, the second residual network model, the audio synthesis model, and the discrimination model to obtain the trained first residual network model, the trained second residual network model, and the trained audio synthesis model.

[0093] In a possible implementation manner, the parameters in the entire audio processing model can be updated according to the first training loss value and the second training loss value obtained above, and then the audio processing model after multiple updates and removing the discrimination model is used for Figure 3 the embodiment shown to synthesize audio with higher sound quality.

[0094] Among them, as can be seen from the above loss function, the training objective of the audio synthesis model is that the audio features of the dry audio sample and its corresponding synthesized audio are close to each other, that is, they have similar fundamental frequencies and spectral envelopes; and the objective of the discrimination model is to make the probability of judging the dry audio sample as true and the synthesized audio as false as close to 1 as possible; that is to say, the training objectives of the audio synthesis model and the discrimination model are mutually adversarial. Therefore, when training the entire audio processing model, alternating iterative training is adopted (exemplarily, first train the discrimination model once, then train the audio synthesis model once, then train the discrimination model, and so on).

[0095] Specifically, for the first batch of training data containing Y dry audio samples, first, pass this batch of training data through the above-mentioned processing in sequence (when processing for the first time, the parameters in the audio processing model are the initial default values) to obtain the corresponding Y synthesized audios. At this time, the audio processing model has a true and false data set that can be input into the discriminant model. The true data set is the dry audio samples, and the false data set is the synthesized audios. Input the true and false data sets into the discriminant model for true and false discrimination and obtain the second training loss value. Fix the parameters of the audio synthesis model, and update the parameters of the remaining models (two residual network models and the discriminant model) according to the second training loss value. Input the second batch of training data containing Y dry audio samples into the audio processing model to obtain Y synthesized audios, and obtain the first training loss value. Fix the parameters of the discriminant model, and update the parameters of the remaining models according to the first training loss value. After the update is completed, input the next batch of training data again, and fix the parameters of the audio synthesis model, and update the parameters of the remaining models except the audio synthesis model. Continuously repeat the above process until the first training loss value and the second training loss value tend to be stable or reach the preset number of training times, and finally determine the parameters in the audio processing model.

[0096] It should be noted that when updating the parameters in the audio processing model according to the first training loss value or the second training loss value, the gradient backpropagation algorithm is used and according to the preset learning rate, the change values of the parameters in each layer of the network in each model are obtained in sequence. Moreover, the above-mentioned method of alternating iterative training can also train the discriminant model multiple times, and then train the audio synthesis model once, or train the audio synthesis model multiple times, and then train the discriminant model once. The setting of the training process is determined according to the specific implementation scenario, and the present application does not limit this.

[0097] Based on Figure 8 the described embodiments, a discriminant model can be added during the training process to train the entire audio processing model and update the parameter values in the audio processing model. The first residual network model, the second residual network model, and the audio synthesis model after training can be used for Figure 3 the embodiments shown to synthesize audio data with higher sound quality.

[0098] See Figure 10 , which is a schematic structural diagram of an audio synthesis device provided by an embodiment of the present application. This audio device includes an acquisition unit 1001 and a processing unit 1002. Among them:

[0099] The acquisition unit 1001 is used to acquire the audio to be synthesized;

[0100] The processing unit 1002 is used to extract the first fundamental frequency and spectral envelope from the audio to be synthesized. The first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectral envelope is used to indicate the timbre of the audio to be synthesized;

[0101] The processing unit 1002 is further configured to adjust the first fundamental frequency to obtain a second fundamental frequency;

[0102] The processing unit 1002 is further configured to obtain a synthesized audio according to the second fundamental frequency and the spectral envelope, and the spectral envelope of the synthesized audio is consistent with the spectral envelope of the audio to be synthesized.

[0103] In a possible implementation manner, when the processing unit 1002 is configured to obtain the synthesized audio according to the second fundamental frequency and the spectral envelope, it specifically includes:

[0104] Invoking a first residual network model and a second residual network model to process the spectral envelope to obtain a first result and a second result;

[0105] Obtaining a third result according to the embedding vector of the first result and the second fundamental frequency;

[0106] Superposing the second result and the third result to obtain an acoustic feature;

[0107] Invoking an audio synthesis model to process the acoustic feature to obtain a synthesized audio.

[0108] In a possible implementation manner, the above audio synthesis model is composed of a convolutional layer and N transposed convolutional layers connected in sequence, where N is a positive integer greater than 1; when the processing unit 1002 is configured to invoke the audio synthesis model to process the acoustic feature to obtain a synthesized audio, it specifically includes:

[0109] Invoking the convolutional layer to perform convolutional processing on the acoustic feature to obtain a fourth result;

[0110] Invoking N transposed convolutional layers to perform transposed convolutional processing on the fourth result to obtain N transposed convolutional results;

[0111] Performing linear interpolation transformation on M of the N transposed convolutional results and then stacking them layer by layer to obtain a synthesized audio; M is a positive integer less than N.

[0112] In a possible implementation manner, the product of the reduction multiples of the N channels corresponding to the above N transposed convolutional layers is the same as the dimensional component of the acoustic feature.

[0113] In a possible implementation manner, the processing unit 1002 is further configured to determine a first training loss value based on the training sample set and the synthesized audio of the training sample set;

[0114] Invoking a discriminant model to perform true / false discrimination on the synthesized audio of the training sample set to determine a second training loss value;

[0115] Train the parameters in the first residual network model, the second residual network model, the audio synthesis model, and the discriminative model according to the first training loss value and the second training loss value, so as to obtain the trained first residual network model, the trained second residual network model, and the trained audio synthesis model.

[0116] In a possible implementation, the discriminative model includes an average pooling layer and a discriminative layer. The discriminative layer is sequentially composed of a convolutional layer, a max pooling layer, and two convolutional layers; the average pooling layer includes a first average pooling layer and a second average pooling layer, and the discriminative layer includes a first discriminative layer, a second discriminative layer, and a third discriminative layer. When the processing unit 1002 is used to call the discriminative model to determine the true or false of the synthesized audio in the training sample set and determine the second training loss value, it specifically includes:

[0117] Call the first discriminative layer to determine the true or false of the synthesized audio in the training sample set to obtain a first discriminative result;

[0118] Call the first average pooling layer to perform average pooling on the synthesized audio in the training sample set to obtain a first average pooling result; call the second discriminative layer to determine the true or false of the first average pooling result to obtain a second discriminative result;

[0119] Call the second average pooling layer to perform average pooling on the first average pooling result to obtain a second average pooling result; call the third discriminative layer to determine the true or false of the second average pooling result to obtain a third discriminative result;

[0120] Determine the second training loss value according to the first discriminative result, the second discriminative result, and the third discriminative result.

[0121] In a possible implementation, when the processing unit 1002 determines the first training loss value based on the training sample set and the synthesized audio of the training sample set, it specifically includes:

[0122] Determine the first Mel spectrogram of the training sample set and the second Mel spectrogram of the synthesized audio of the training sample set;

[0123] Determine that the first training loss value is the minimum mean square error between the first Mel spectrogram and the second Mel spectrogram.

[0124] In a possible implementation, the acquisition unit 1001 is further configured to acquire an original sample set, and the original sample set includes multiple dry audio samples, and each dry audio sample in the multiple dry audio samples is represented by a pitch distribution vector;

[0125] The processing unit 1002 is further configured to determine the sampling probability of each dry audio sample based on the pitch distribution vector of each dry audio sample, and the sampling probability is used to make the pitch of the sampled training sample set evenly distributed;

[0126] Sampling the original sample set based on the sampling probability of each dry audio sample to obtain a training sample set.

[0127] In a possible implementation, the processing unit 1002 is further configured to obtain the synthesized audio of the training sample set based on the fundamental frequency and spectral envelope of the training sample set.

[0128] It should be noted that the functions of the various unit modules of the audio synthesis device according to the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.

[0129] See Figure 11 , which is a schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device 11 may include: one or more processors 1101, a memory 1102, and a transceiver 1103. The above-mentioned processors 1101, memory 1102, and transceiver 1103 are connected through a bus 1104. The memory 1102 is used to store a computer program, and the computer program includes program instructions. The processors 1101 and transceiver 1103 are used to execute the program instructions stored in the memory 702 and perform the following operations:

[0130] Obtain the audio to be synthesized;

[0131] Extract the first fundamental frequency and spectral envelope from the audio to be synthesized, where the first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectral envelope is used to indicate the timbre of the audio to be synthesized;

[0132] Adjust the first fundamental frequency to obtain a second fundamental frequency;

[0133] Obtain the synthesized audio according to the second fundamental frequency and spectral envelope, and the spectral envelope of the synthesized audio is consistent with the spectral envelope of the audio to be synthesized.

[0134] It should be understood that in some feasible embodiments, the above-mentioned processor 1101 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory 1102 may include a read-only memory and a random access memory, and provide instructions and data to the processor 501. A part of the memory 1102 may also include a non-volatile random access memory. For example, the memory 1102 may also store information about the device type.

[0135] In specific implementation, the above-mentioned terminal device may execute, through each of its built-in functional modules, the implementation manners provided in each of the steps as described above Figure 3 or Figure 8 in the above, and for the specific implementation manners, reference may be made to the implementation manners provided in each of the above steps, which will not be elaborated here.

[0136] The embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores computer-readable instructions executed by the aforementioned audio synthesis device, and the computer-readable instructions include program instructions. When the processor executes the above program instructions, it can execute the Figure 3 , Figure 9 methods in the corresponding embodiments, and therefore, details will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the program instructions may be deployed on one computer device, or executed on multiple computer devices located at one place, or alternatively, executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network may form a blockchain system.

[0137] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the method in the corresponding embodiment as described above. Therefore, it will not be elaborated here. Figure 3 , Figure 8 The method in the corresponding embodiment, and thus will not be repeated here.

[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0139] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An audio synthesis method, characterized in that, the method includes: Obtain the audio to be synthesized; Extract a first fundamental frequency and a spectral envelope from the audio to be synthesized, where the first fundamental frequency is used to indicate the pitch of the audio to be synthesized, and the spectral envelope is used to indicate the timbre of the audio to be synthesized; Adjust the first fundamental frequency to obtain a second fundamental frequency; Obtain a synthesized audio according to the second fundamental frequency and the spectral envelope, where the spectral envelope of the synthesized audio is consistent with the spectral envelope of the audio to be synthesized; The obtaining the synthesized audio according to the second fundamental frequency and the spectral envelope includes: Call the trained first residual network model and the trained second residual network model to process the spectral envelope to obtain a first result and a second result; Obtain a third result according to the embedding vector of the first result and the second fundamental frequency; Superimpose the second result and the third result to obtain an acoustic feature; Call the trained audio synthesis model to process the acoustic feature to obtain the synthesized audio.

2. The method according to claim 1, characterized in that, The trained audio synthesis model is composed of a convolutional layer and N transposed convolutional layers connected in sequence, where N is a positive integer greater than 1. The calling the trained audio synthesis model to process the acoustic feature to obtain the synthesized audio includes: Call the convolutional layer to perform convolutional processing on the acoustic feature to obtain a fourth result; Call the N transposed convolutional layers to perform transposed convolutional processing on the fourth result to obtain N transposed convolutional results; Perform linear interpolation transformation on M of the N transposed convolutional results and then layer-by-layer superimpose them to obtain the synthesized audio; M is a positive integer less than N.

3. The method according to claim 2, characterized in that, The product of the reduction multiples of the N channels corresponding to the N transposed convolutional layers is the same as the dimensional component of the acoustic feature.

4. The method according to claim 1, characterized in that, The method further includes: Determine a first training loss value based on the training sample set and the synthesized audio of the training sample set; Call the discriminant model to perform true / false discrimination on the synthesized audio of the training sample set to determine a second training loss value; According to the first training loss value and the second training loss value, train the parameters in the first residual network model, the second residual network model, the audio synthesis model and the discriminant model to obtain the trained first residual network model, the trained second residual network model and the trained audio synthesis model.

5. The method according to claim 4, characterized in that, The discriminant model includes an average pooling layer and a discriminant layer. The discriminant layer is composed of a convolutional layer, a max pooling layer and two convolutional layers connected in sequence; the average pooling layer includes a first average pooling layer and a second average pooling layer, and the discriminant layer includes a first discriminant layer, a second discriminant layer and a third discriminant layer; The calling the discriminant model to perform true / false discrimination on the synthesized audio of the training sample set to determine a second training loss value specifically includes: Call the first discrimination layer to perform true / false discrimination on the synthesized audio of the training sample set, and obtain a first discrimination result; Call the first average pooling layer to perform average pooling on the synthesized audio of the training sample set to obtain a first average pooling result; call the second discrimination layer to perform true / false discrimination on the first average pooling result to obtain a second discrimination result; Call the second average pooling layer to perform average pooling on the first average pooling result to obtain a second average pooling result; call the third discrimination layer to perform true / false discrimination on the second average pooling result to obtain a third discrimination result; Determine the second training loss value according to the first discrimination result, the second discrimination result, and the third discrimination result.

6. The method according to claim 4, wherein, the determining the first training loss value based on the training sample set and the synthesized audio of the training sample set includes: determining a first Mel spectrogram of the training sample set and a second Mel spectrogram of the synthesized audio of the training sample set; determining the first training loss value as the mean squared error between the first Mel spectrogram and the second Mel spectrogram.

7. The method according to claim 4, wherein, the method further includes: obtaining an original sample set, where the original sample set includes a plurality of dry audio samples, and each dry audio sample in the plurality of dry audio samples is represented by a pitch distribution vector; determining a sampling probability for each dry audio sample based on the pitch distribution vector of each dry audio sample, where the sampling probability is used to make the pitch of the sampled training sample set uniformly distributed; sampling the original sample set based on the sampling probability of each dry audio sample to obtain the training sample set.

8. The method according to claim 6, wherein, the synthesized audio of the training sample set is obtained based on the fundamental frequency and spectral envelope of the training sample set.

9. A terminal device, wherein, the terminal device includes a memory and a processor; the memory is used to store a computer program; the processor is used to call the computer program from the memory, so that the terminal device executes the method according to any one of claims 1-8.

10. A computer-readable storage medium, wherein, the computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions run on a terminal device, the terminal device is caused to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Audio processing method and device

    CN111916093A