Normalizing flow-based audio generation method and apparatus, device, and storage medium

By using a denoising decoder to process the periodic noise generated by the prior network in the standardized stream model, the problem of low audio quality in speech generation is solved, and high-quality audio generation is achieved.

WO2025129817A1PCT designated stage expired Publication Date: 2025-06-26PING AN TECH (SHENZHEN) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/079024
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-02-28
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When using a standardized stream model for speech generation, periodic operations using an affine coupling layer for dimension decomposition result in significant periodic noise in the generated audio, resulting in lower audio quality.

Method used

By randomly sampling the standard Gaussian distribution, the first variable vector is obtained, and then input it to the prior network for inverse transformation, obtain the first hidden variable vector, and then decode the first hidden variable vector into audio data through a denoising decoder.

Benefits of technology

Effectively neutralize periodic noise generated in prior network compression operations, generate clear and noise-free high-quality audio data, and improve the quality of audio generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024079024_26062025_PF_FP_ABST
    Figure CN2024079024_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A normalizing flow-based audio generation method and apparatus, a computer device, and a storage medium. The method comprises: performing random sampling on a standard Gaussian distribution to obtain a first variable vector (S101); inputting the first variable vector into a prior network in an audio generation model for inverse transformation so as to obtain a first hidden variable vector (S102); and inputting the first hidden variable vector into a noise reduction decoder in the audio generation model for decoding so as to obtain audio data (S103). Therefore, although the prior network needs to perform compression operation on the first variable vector to obtain the first hidden variable vector with periodic noise, the noise can be effectively neutralized by means of the noise reduction decoder, and finally, clear and noiseless high-quality audio data is generated, thereby achieving the purpose of improving the audio quality.
Need to check novelty before this filing date? Find Prior Art

Description

Audio generation method, device, equipment and storage medium based on standardized stream

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 202311750998.2, and invention name “Audio generation method, device, equipment and storage medium based on standardized stream”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, and storage medium for generating audio based on a standardized stream. Background Art

[0003] In recent years, with the continuous development of artificial intelligence technology, more and more intelligent scenarios require the application of audio generation technology.

[0004] Currently, common generative model frameworks for audio generation include Generative Adversarial Networks (GANs) and Normalizing Flow (NF) models. Because affine coupling layers are often used in NF models, dimensional partitioning of the data is required during computation. Therefore, when processing low-dimensional data, especially one-dimensional audio waveform points, dimensional compression is required: by forming an n-dimensional vector from every n points, the original data of dimension 1 and length T is compressed into a sequence of dimension n and length T / n. Technical issues

[0005] The inventors realized that when applying the normalized flow model for speech generation, the periodic operation of using the affine coupling layer for dimensional decomposition will cause the generated audio to have obvious periodic noise, resulting in lower quality of the generated audio.

[0006] Therefore, the audio generated by the speech generation model based on the standardized flow model has the problem of low quality.

[0007] Summary of the Invention

[0008] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for generating audio based on a standardized stream to solve the problem of low quality of audio generated by a speech generation model based on a standardized stream model.

[0009] A method for generating audio based on a standardized stream, the method comprising:

[0010] Randomly sample the standard Gaussian distribution to obtain the first variable vector;

[0011] Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector;

[0012] The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

[0013] An audio generation device based on a standardized stream, the device comprising:

[0014] A sampling unit, configured to perform random sampling on a standard Gaussian distribution to obtain a first variable vector;

[0015] an inverse transformation unit, configured to input the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector;

[0016] A decoding unit is used to input the first latent variable vector into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

[0017] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer-readable instructions:

[0018] Randomly sample the standard Gaussian distribution to obtain the first variable vector;

[0019] Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector;

[0020] The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

[0021] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the following steps:

[0022] Randomly sample the standard Gaussian distribution to obtain the first variable vector;

[0023] Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector;

[0024] The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

[0025] Technical Effects

[0026] In the embodiment of the present application, a first variable vector is obtained by sampling from a standard Gaussian distribution. The first variable vector is then inversely transformed using a priori network to obtain a first latent variable vector. This latent variable vector is then decoded into audio data using a noise reduction decoder. As can be seen, although the priori network requires a compression operation on the first variable vector, resulting in a first latent variable vector containing periodic noise, the noise reduction decoder effectively neutralizes this noise, ultimately generating clear, noise-free, high-quality audio data, thereby improving audio quality.

[0027] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below, and other features and advantages of the disclosure will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] FIG1 is a flowchart of an implementation of a method for generating audio based on a standardized stream in an embodiment of the present application;

[0030] FIG2 is a schematic structural diagram of a noise reduction encoder according to an embodiment of the present application;

[0031] FIG3 is a schematic diagram of a structure of a priori network in one embodiment of the present application;

[0032] FIG4 is a flowchart of a partial implementation of a method for generating audio based on a standardized stream in an embodiment of the present application;

[0033] FIG5 is a flowchart of a partial implementation of a method for generating audio based on a standardized stream in an embodiment of the present application;

[0034] FIG6 is a flowchart of a partial implementation of a method for generating audio based on a standardized stream in an embodiment of the present application;

[0035] FIG7 is a flowchart of a partial implementation of a method for generating audio based on a standardized stream in an embodiment of the present application;

[0036] FIG8 is a schematic structural diagram of an audio generation device based on a standardized stream in an embodiment of the present application;

[0037] FIG9 is a schematic structural diagram of a computer device in one embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0039] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0040] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0041] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0042] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0043] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0044] This application discloses a method, apparatus, computer device, and storage medium for audio generation based on a standardized stream. A first variable vector is obtained by sampling from a standard Gaussian distribution. The first variable vector is then inversely transformed using a priori network to obtain a first latent variable vector. This latent variable vector is then decoded into audio data using a noise reduction decoder. While the priori network requires compression of the first variable vector to obtain a first latent variable vector containing periodic noise, the noise reduction decoder effectively neutralizes this noise, ultimately generating clear, noise-free, high-quality audio data, thereby improving audio quality.

[0045] As shown in FIG1 , a schematic diagram of a method for generating audio based on a standardized stream disclosed in an embodiment of the present application is provided. The method is applicable to electronic devices with data processing capabilities, such as mobile phones, tablet computers, personal computers, and servers. The method in this embodiment may specifically include the following steps:

[0046] S101: Randomly sample the standard Gaussian distribution to obtain a first variable vector.

[0047] Specifically, the standard Gaussian distribution in this embodiment refers to a Gaussian distribution with a mean of 0 and a variance of 1, wherein the method of randomly sampling the standard Gaussian distribution includes but is not limited to the inverse transform method, rejection sampling method, importance sampling and resampling (Importance Sampling, Sampling-Importance-Resampling), Markov Monte Carlo sampling method, etc.

[0048] S102: Input the first variable vector into the prior network in the audio generation model for inverse transformation to obtain a first latent variable vector.

[0049] Specifically, the priori network in this embodiment is designed using a volume-invariant flow structure. Therefore, after the first variable vector is input into the priori network, a controllable first latent variable vector can be obtained.

[0050] The specific structure of the prior network in this embodiment can be shown in Figure 2. The prior network based on the standardized flow model is a fully reversible flow network. After the first variable vector is input into the prior network, a series of operations such as squeezing, adding coincidences, and fast inversion are performed on the input first variable vector to ultimately obtain the first latent variable vector.

[0051] S103: Input the first latent variable vector into the noise reduction decoder in the audio generation model for decoding to obtain audio data.

[0052] The normalized flow model is a completely reversible network structure. Therefore, the noise reduction decoder based on the normalized flow model is also completely reversible. That is, after being encoded by the encoder based on the normalized flow model, it can be perfectly reconstructed by the decoder based on the normalized flow model.

[0053] It can be understood that the encoder and decoder based on the standardized flow model can be the same model network, that is, the encoder of the standardized flow model can also be used as a decoder, and the encoding and decoding are completely reversible.

[0054] Specifically, if the a priori network in this embodiment causes periodic noise in the finally generated first variable vector during the squeezing process of the first variable vector, the encoder based on the standardized stream will neutralize the periodic noise that may exist in the first variable vector when decoding and restoring the first variable vector, and finally generate clear and noise-free audio data.

[0055] The model structure of the encoder or decoder in this embodiment is shown in Figure 3. The prior network based on the standardized flow model is a fully reversible flow network. The encoder and decoder based on the standardized flow can be implemented based on the same network structure. That is, the encoder can encode L audio into R latent variables and decode R latent variables into L audio.

[0056] This application discloses a method for audio generation based on a standardized stream. This method uses standard Gaussian distribution sampling to obtain a first variable vector, then uses a priori network to perform an inverse transform on the first variable vector to obtain a first latent variable vector. This first latent variable vector is then decoded into audio data using a noise reduction decoder. While the priori network requires compression of the first variable vector to obtain a first latent variable vector containing periodic noise, the noise reduction decoder effectively neutralizes this noise, ultimately generating clear, noise-free, high-quality audio data, thereby improving audio quality.

[0057] In one implementation, the a priori network in this embodiment can be trained through the following steps, as shown in FIG4 :

[0058] S401: Obtain sample audio required for training.

[0059] The sample audio includes but is not limited to music, songs, or conversations, and may be sourced from on-site recordings, audio database extraction, etc. In this embodiment, the source and specific type of the sample audio are not limited.

[0060] S402: Add random Gaussian noise to the sample audio to obtain a noisy audio sample.

[0061] When the sample audio is determined, a Gaussian noise that conforms to the Gaussian distribution is randomly generated, and the sample audio is noised according to the Gaussian noise to obtain a noise audio sample.

[0062] For example, taking the sample noise x as an example, a Gaussian noise ε is obtained by random sampling, and the noise x is added to obtain x+βε, where β is the noise weight, and x+βε is the noise audio sample after the sample noise is added.

[0063] S403: Inputting the noise audio sample into the first network structure and the second network structure of the audio generation model to be trained respectively for training to obtain a trained audio generation model.

[0064] It should be understood that the audio generation model in this embodiment can be divided into two parts, namely the first network structure and the second network structure. The first network structure and the second network structure are trained separately. When the training of the first network structure and the second network structure is completed, a trained audio generation model is obtained.

[0065] In the specific implementation based on FIG. 4 , step S403 in this embodiment can be implemented by the following steps, as shown in FIG. 5 :

[0066] S501: Input a noise audio sample into a first network structure to obtain a second variable vector.

[0067] Specifically, the first network structure in this embodiment includes a noise reduction encoder and a priori network. The second variable vector is obtained through the following steps based on the noise reduction encoder and the priori network, as shown in FIG6 :

[0068] S601: Input the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector.

[0069] The noisy audio sample is input into the noise reduction encoder so that the noise reduction encoder encodes the noisy audio sample to obtain a second latent variable vector between the noise reduction encoder and the prior network to be trained.

[0070] Specifically, the noise reduction encoder in this embodiment includes but is not limited to a noise reduction autoencoder (DAE: Denoising AutoEncoder). The noisy audio sample is input into the noise reduction autoencoder so that the noise reduction autoencoder performs noise reduction processing on the noisy audio sample while encoding the noisy audio sample to obtain a second latent variable vector. It should be noted that the structure of the noise reduction autoencoder in this embodiment is designed using a bidirectional autoregressive flow. Compared with the noise reduction encoder designed with a non-autoregressive flow, the noise reduction encoder obtained by using the bidirectional autoregressive flow design does not need to perform compression operations on the noisy audio sample, thereby avoiding the possibility of generating periodic noise.

[0071] S602: Input the second latent variable vector into the priori network to be trained for forward transformation to obtain a second variable vector.

[0072] The second variable vector conforms to the standard Gaussian distribution N(0, 1).

[0073] Specifically, the prior network in this embodiment employs a volume-invariant flow structure, such that the second latent variable vector is input into the prior network to be trained and subjected to a forward transformation, resulting in a second variable vector having the same volume as a standard Gaussian distribution. This volume-invariant flow structure allows the prior network to restrict the range of variation in the volume of the second latent variable vector during the forward transformation, thereby improving the convergence efficiency of the prior network to be trained and enhancing both the training efficiency and effectiveness of the prior network to be trained.

[0074] S502: Calculate a first loss value based on the second variable vector.

[0075] The first network model includes a priori network and a noise reduction encoder. Based on the priori network, the noise reduction encoder and the second variable vector, a first loss value of the first network model is calculated.

[0076] Specifically, the first loss value in this embodiment includes but is not limited to the maximum likelihood estimation. The formula for calculating the likelihood estimation can be as follows: logp(x+βε)=log(p(z p )+logdet|J(F)|+logdet|J(G)|

[0077] Where x+βε represents the noise audio sample, z p represents the second variable vector, det|J(F)| represents the Jacobian matrix determinant of the denoising encoder, det|J(G)| represents the Jacobian matrix determinant of the prior network, logp(x+βε) represents the likelihood estimate, where the maximum likelihood estimate represents the first loss value.

[0078] S503: Input the noisy audio sample into the second network structure to obtain restored audio.

[0079] Specifically, the first network structure in this embodiment includes a noise reduction encoder and a noise reduction decoder. Based on the noise reduction encoder and the noise reduction decoder, the second variable vector is obtained through the following steps, as shown in FIG7 :

[0080] S701: Input the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector.

[0081] The noisy audio sample is input into the noise reduction encoder so that the noise reduction encoder encodes the noisy audio sample to obtain a second latent variable vector between the noise reduction encoder and the prior network to be trained.

[0082] Specifically, the noise reduction encoder in this embodiment includes but is not limited to a noise reduction autoencoder (DAE: Denoising AutoEncoder). The noisy audio sample is input into the noise reduction autoencoder so that the noise reduction autoencoder performs noise reduction processing on the noisy audio sample while encoding the noisy audio sample to obtain a second latent variable vector. It should be noted that the structure of the noise reduction autoencoder in this embodiment is designed using a bidirectional autoregressive flow. Compared with the noise reduction encoder designed with a non-autoregressive flow, the noise reduction encoder obtained by using the bidirectional autoregressive flow design does not need to perform compression operations on the noisy audio sample, thereby avoiding the possibility of generating periodic noise.

[0083] S702: Input the second latent variable vector into the noise reduction decoder to be trained for decoding to obtain restored audio.

[0084] It should be understood that the noise reduction encoder encodes the noise audio vector into a second latent variable vector, and the noise reduction decoder decodes the second latent variable vector into restored audio.

[0085] In a specific implementation, the noise reduction decoder in this embodiment can adopt the structure of a feedforward convolutional neural network and a UNet structure, which not only ensures the decoding effect of the decoder but also reduces the computational complexity of the decoding process while ensuring that the noise reduction decoder is completely reversible.

[0086] S504: Calculate a second loss value based on the restored audio and the sample audio.

[0087] The restored audio and the sample audio are input into a second loss value calculation formula to obtain a second loss value.

[0088] The second loss value calculation formula is as follows:

[0089] Among them, x represents the sample audio, Indicates restoring audio. Represents the weight, L reb Represents the second loss value.

[0090] S505: When the first loss value and the second loss value meet the model training conditions, a trained audio generation model is obtained.

[0091] A total loss value is calculated based on the first loss value and the second loss value. Based on the total loss value, it is determined whether the audio generation model is fully trained. If the audio generation model is fully trained, a trained audio generation model is obtained.

[0092] Specifically, in this embodiment, whether the audio generation model has been fully trained can be determined by determining whether the total loss value is less than a preset loss value threshold. If the total loss value is less than the preset loss value threshold, the audio generation model is trained. If the total loss value is greater than the preset loss value threshold, the audio generation model is trained based on the audio sample until the total loss value is less than the preset loss value threshold. The audio generation model is trained. It should be noted that the above is only one way to determine whether the audio generation model has been fully trained. It can also be determined whether the audio generation model has been fully trained by determining whether the total loss value tends to be stable. Therefore, this embodiment does not limit the specific method of determining whether the audio generation model has been fully trained based on the total loss value.

[0093] It should be noted that in this embodiment, after obtaining the trained audio generative model, the noise reduction encoder in the model is discarded, while the prior network and noise reduction decoder in the model are retained. During the inference phase, the first variable vector is input from the prior network to the audio generative model, and the audio generative model is output from the noise reduction decoder to obtain audio data.

[0094] In addition, the audio data in this embodiment is only a piece of random audio data. If a specified audio segment needs to be generated, an audio generation condition can be added to the input first variable vector, where the audio generation condition includes but is not limited to audio generation duration, audio generation content, etc. Thus, according to the added audio generation condition, the specified audio data is finally output.

[0095] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0096] FIG8 is a schematic diagram of the structure of an audio generation device based on a standardized stream disclosed in an embodiment of the present application. The device is suitable for electronic devices with data processing capabilities, such as mobile phones, tablet computers, personal computers, and servers.

[0097] Specifically, the device in this embodiment may include the following units:

[0098] A sampling unit 801 is configured to perform random sampling on a standard Gaussian distribution to obtain a first variable vector;

[0099] An inverse transformation unit 802 is configured to input the first variable vector into a priori network in the audio generation model for inverse transformation to obtain a first latent variable vector;

[0100] The decoding unit 803 is configured to input the first latent variable vector into a noise reduction decoder in the audio generation model for decoding to obtain audio data.

[0101] This application discloses an audio generation device based on a standardized stream. This device obtains a first variable vector by sampling from a standard Gaussian distribution. This first variable vector is then inversely transformed using a priori network to obtain a first latent variable vector. This latent variable vector is then decoded into audio data using a noise reduction decoder. While the priori network requires compression of the first variable vector to produce a first latent variable vector containing periodic noise, the noise reduction decoder effectively neutralizes this noise, ultimately generating clear, noise-free, high-quality audio data and improving audio quality.

[0102] In one implementation, the audio generation model is trained as follows:

[0103] Get the sample audio required for training;

[0104] Add random Gaussian noise to the sample audio to obtain a noisy audio sample;

[0105] The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained to obtain a trained audio generation model.

[0106] In one implementation, the noise audio sample is input into the first network structure and the second network structure of the audio generation model to be trained for training, to obtain a trained audio generation model, including:

[0107] Inputting the noise audio sample into the first network structure to obtain a second variable vector;

[0108] Calculate the first loss value based on the second variable vector;

[0109] Input the noisy audio sample into the second network structure to obtain restored audio;

[0110] Calculate a second loss value based on the restored audio and the sample audio;

[0111] When the first loss value and the second loss value meet the model training conditions, a trained audio generation model is obtained.

[0112] In one implementation, the first network structure includes a noise reduction encoder and a priori network;

[0113] Input the noise audio sample into the first network structure to obtain the second variable vector, including:

[0114] Input the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector;

[0115] The second latent variable vector is input into the prior network to be trained for positive transformation to obtain the second variable vector.

[0116] In one implementation, the second network structure includes a noise reduction encoder and a noise reduction decoder;

[0117] Input the noisy audio sample into the second network structure to obtain the restored audio, including:

[0118] Input the noisy audio sample into the noise reduction encoder to obtain the second latent variable vector;

[0119] The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain the restored audio.

[0120] In one implementation, the noise reduction encoder adopts a bidirectional autoregressive flow model structure.

[0121] In one implementation, the prior network adopts a volume-invariant flow structure.

[0122] Regarding the specific limitations of the audio generation device based on standardized streams, please refer to the relevant limitations of the audio generation method based on standardized streams above, which will not be repeated here. The various modules in the above-mentioned audio generation device based on standardized streams can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0123] In one implementation, an embodiment of the present application discloses a computer device, which may be a server, and its internal structure diagram may be shown in Figure 9. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer-readable instructions, and a database. The storage medium includes a non-volatile readable storage medium and a volatile readable storage medium. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer-readable instructions are executed by the processor, a method for generating audio based on a standardized stream is implemented.

[0124] In one embodiment, a computer device is provided, comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the following steps are implemented:

[0125] Randomly sample the standard Gaussian distribution to obtain the first variable vector;

[0126] Inputting the first variable vector into the prior network in the audio generation model for inverse transformation to obtain a first latent variable vector;

[0127] The first latent variable vector is input into the noise reduction decoder in the audio generation model for decoding to obtain audio data.

[0128] In one implementation, the audio generation model is trained as follows:

[0129] Get the sample audio required for training;

[0130] Add random Gaussian noise to the sample audio to obtain a noisy audio sample;

[0131] The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained to obtain a trained audio generation model.

[0132] In one implementation, the noise audio sample is input into the first network structure and the second network structure of the audio generation model to be trained for training, to obtain a trained audio generation model, including:

[0133] Inputting the noise audio sample into the first network structure to obtain a second variable vector;

[0134] Calculate the first loss value based on the second variable vector;

[0135] Input the noisy audio sample into the second network structure to obtain restored audio;

[0136] Calculate a second loss value based on the restored audio and the sample audio;

[0137] When the first loss value and the second loss value meet the model training conditions, a trained audio generation model is obtained.

[0138] In one implementation, the first network structure includes a noise reduction encoder and a priori network;

[0139] Input the noise audio sample into the first network structure to obtain the second variable vector, including:

[0140] Input the noisy audio sample into the noise reduction encoder to obtain the second latent variable vector;

[0141] The second latent variable vector is input into the prior network to be trained for positive transformation to obtain the second variable vector.

[0142] In one implementation, the second network structure includes a noise reduction encoder and a noise reduction decoder;

[0143] Input the noisy audio sample into the second network structure to obtain the restored audio, including:

[0144] Input the noisy audio sample into the noise reduction encoder to obtain the second latent variable vector;

[0145] The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain the restored audio.

[0146] In one implementation, the noise reduction encoder adopts a bidirectional autoregressive flow model structure.

[0147] In one implementation, the prior network adopts a volume-invariant flow structure.

[0148] In one implementation, an embodiment of the present application discloses a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor in a computer device, the computer device is enabled to perform the steps of any embodiment of a method for generating audio based on a standardized stream as disclosed in this application. The computer-readable storage medium can be either non-volatile or volatile.

[0149] In one embodiment, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the following steps are implemented:

[0150] Randomly sample the standard Gaussian distribution to obtain the first variable vector;

[0151] Inputting the first variable vector into the prior network in the audio generation model for inverse transformation to obtain a first latent variable vector;

[0152] The first latent variable vector is input into the noise reduction decoder in the audio generation model for decoding to obtain audio data.

[0153] In one implementation, the audio generation model is trained as follows:

[0154] Get the sample audio required for training;

[0155] Add random Gaussian noise to the sample audio to obtain a noisy audio sample;

[0156] The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained to obtain a trained audio generation model.

[0157] In one implementation, the noise audio sample is input into the first network structure and the second network structure of the audio generation model to be trained for training, to obtain a trained audio generation model, including:

[0158] Inputting the noise audio sample into the first network structure to obtain a second variable vector;

[0159] Calculate the first loss value based on the second variable vector;

[0160] Input the noisy audio sample into the second network structure to obtain restored audio;

[0161] Calculate a second loss value based on the restored audio and the sample audio;

[0162] When the first loss value and the second loss value meet the model training conditions, a trained audio generation model is obtained.

[0163] In one implementation, the first network structure includes a noise reduction encoder and a priori network;

[0164] Input the noise audio sample into the first network structure to obtain the second variable vector, including:

[0165] Input the noisy audio sample into the noise reduction encoder to obtain the second latent variable vector;

[0166] The second latent variable vector is input into the prior network to be trained for positive transformation to obtain the second variable vector.

[0167] In one implementation, the second network structure includes a noise reduction encoder and a noise reduction decoder;

[0168] Input the noisy audio sample into the second network structure to obtain the restored audio, including:

[0169] Input the noisy audio sample into the noise reduction encoder to obtain the second latent variable vector;

[0170] The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain the restored audio.

[0171] In one implementation, the noise reduction encoder adopts a bidirectional autoregressive flow model structure.

[0172] In one implementation, the prior network adopts a volume-invariant flow structure.

[0173] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When the computer-readable instructions are executed, they may include processes such as the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0174] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0175] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for generating audio based on a standardized stream, wherein: The method comprises: Randomly sample the standard Gaussian distribution to obtain the first variable vector; Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector; The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

2. The method for generating audio based on a standardized stream according to claim 1, wherein: The audio generation model is trained in the following way: Get the sample audio required for training; Adding random Gaussian noise to the sample audio to obtain a noisy audio sample; The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained for training, so as to obtain the trained audio generation model.

3. The method for generating audio based on a standardized stream as claimed in claim 2, wherein: The step of inputting the noise audio sample into the first network structure and the second network structure of the audio generation model to be trained respectively for training to obtain the trained audio generation model comprises: Inputting the noise audio sample into the first network structure to obtain a second variable vector; Calculate a first loss value according to the second variable vector; Inputting the noise audio sample into the second network structure to obtain restored audio; Calculating a second loss value according to the restored audio and the sample audio; When the first loss value and the second loss value meet the model training conditions, the trained audio generation model is obtained.

4. The method for generating audio based on a standardized stream as claimed in claim 3, wherein: The first network structure includes a noise reduction encoder and the prior network; The step of inputting the noise audio sample into the first network structure to obtain a second variable vector comprises: Inputting the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the priori network to be trained for positive transformation to obtain a second variable vector.

5. The method for generating audio based on a standardized stream as claimed in claim 3, wherein: The second network structure includes a noise reduction encoder and the noise reduction decoder; The step of inputting the noise audio sample into the second network structure to obtain restored audio comprises: Inputting the noisy audio sample into a noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain restored audio.

6. The method for generating audio based on a standardized stream according to any one of claims 1 to 5, wherein: The noise reduction encoder adopts a bidirectional autoregressive flow model structure.

7. The method for generating audio based on a standardized stream according to any one of claims 1 to 5, wherein: The prior network adopts a volume-invariant flow structure.

8. An audio generation device based on a standardized stream, wherein: The device comprises: A sampling unit, used for randomly sampling a standard Gaussian distribution to obtain a first variable vector; An inverse transformation unit, used for inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector; A decoding unit is used to input the first latent variable vector into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein: When the processor executes the computer-readable instructions, the following steps are implemented: Randomly sample the standard Gaussian distribution to obtain the first variable vector; Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector; The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

10. The computer device of claim 9, wherein: The audio generation model is trained in the following way: Get the sample audio required for training; Adding random Gaussian noise to the sample audio to obtain a noisy audio sample; The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained for training, so as to obtain the trained audio generation model.

11. The computer device of claim 10, wherein: The step of inputting the noise audio sample into the first network structure and the second network structure of the audio generation model to be trained respectively for training to obtain the trained audio generation model comprises: Inputting the noise audio sample into the first network structure to obtain a second variable vector; Calculate a first loss value according to the second variable vector; Inputting the noise audio sample into the second network structure to obtain restored audio; Calculating a second loss value according to the restored audio and the sample audio; When the first loss value and the second loss value meet the model training conditions, the trained audio generation model is obtained.

12. The computer device of claim 11, wherein: The first network structure includes a noise reduction encoder and the prior network; The step of inputting the noise audio sample into the first network structure to obtain a second variable vector comprises: Inputting the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the priori network to be trained for positive transformation to obtain a second variable vector.

13. The computer device of claim 11, wherein: The second network structure includes a noise reduction encoder and the noise reduction decoder; The step of inputting the noise audio sample into the second network structure to obtain restored audio comprises: Inputting the noisy audio sample into a noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain restored audio.

14. The computer device according to any one of claims 9 to 13, wherein: The noise reduction encoder adopts a bidirectional autoregressive flow model structure.

15. The computer device according to any one of claims 9 to 13, wherein: The prior network adopts a volume-invariant flow structure.

16. A computer-readable storage medium storing computer-readable instructions, wherein: When the computer readable instructions are executed by a processor, the following steps are implemented: Randomly sample the standard Gaussian distribution to obtain the first variable vector; Inputting the first variable vector into a priori network in an audio generation model for inverse transformation to obtain a first latent variable vector; The first latent variable vector is input into a noise reduction decoder in an audio generation model for decoding to obtain audio data.

17. The readable storage medium according to claim 16, wherein: The audio generation model is trained in the following way: Get the sample audio required for training; Adding random Gaussian noise to the sample audio to obtain a noisy audio sample; The noise audio samples are respectively input into the first network structure and the second network structure of the audio generation model to be trained for training, so as to obtain the trained audio generation model.

18. The readable storage medium according to claim 17, wherein: The step of inputting the noise audio sample into the first network structure and the second network structure of the audio generation model to be trained respectively for training to obtain the trained audio generation model comprises: Inputting the noise audio sample into the first network structure to obtain a second variable vector; Calculate a first loss value according to the second variable vector; Inputting the noise audio sample into the second network structure to obtain restored audio; Calculating a second loss value according to the restored audio and the sample audio; When the first loss value and the second loss value meet the model training conditions, the trained audio generation model is obtained.

19. The readable storage medium according to claim 18, wherein: The first network structure includes a noise reduction encoder and the prior network; The step of inputting the noise audio sample into the first network structure to obtain a second variable vector comprises: Inputting the noisy audio sample into the noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the priori network to be trained for positive transformation to obtain a second variable vector.

20. The readable storage medium of claim 18, wherein: The second network structure includes a noise reduction encoder and the noise reduction decoder; The step of inputting the noise audio sample into the second network structure to obtain restored audio comprises: Inputting the noisy audio sample into a noise reduction encoder to obtain a second latent variable vector; The second latent variable vector is input into the noise reduction decoder to be trained for decoding to obtain restored audio.

Citation Information

Patent Citations

  • Voice noise method and system for data enhancement

    CN110211575A

  • Bidirectional computer-aided pronunciation training method and device adopting standardized flow

    CN115966109A

  • End-to-end speech synthesis method, device, equipment and medium

    CN116469375A

  • Audio confrontation sample generation method and device, equipment and storage medium

    CN116580694A

  • Apparatus for providing processed audio signal, method for providing processed audio signal, apparatus for providing neural network parameters and method for providing neural network parameters

    CN116648747A