Audio codec training method, audio processing method and device
By setting divergence threshold and joint training model parameters in audio codec training, the divergence loss collapse problem is solved, the diversity of audio characterization and reconstruction fidelity are improved, and more audio application scenarios are adapted to more audio application scenarios.
Patent Information
- Application Number
- CN202510863785.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing audio codecs have divergence loss collapse problems during training, resulting in reduced audio characterization diversity and insufficient reconstruction fidelity, and unstable training.
By setting a preset divergence threshold during the audio codec training process, the divergence distance between the latent spatial distribution and the standard Gaussian distribution is constrained, and end-to-end joint training is carried out in combination with the preset audio encoding model, variational autoencoder and audio decoding model, the model parameters are adjusted using divergence loss information and other loss information to avoid premature convergence of divergence loss.
It improves the diversity and robustness of audio characterization, ensures the stability of the training process, improves the fidelity of audio reconstruction, and enhances the audio encoding compression rate, reduces the amount of data transmission, and adapts to more audio application scenarios.
Smart Images

Figure CN120356476A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio processing, and in particular, to a training method for an audio codec, an audio processing method, and an apparatus. Background Art
[0002] With the wide development of audio processing technologies, audio-based interactions and the like have attracted much attention, such as text-to-speech (TTS), generating non-speech audio (TTA) based on text descriptions, audio transmission, etc. Among them, audio representation is a core task in audio processing. In related technologies, an audio codec is generally used to perform encoding and decoding processing on audio, such as a variational autoencoder. However, when using divergence loss in training, the existing processing methods of divergence loss will have a collapse problem, resulting in the degradation of the latent space, the loss of the diversity of audio representation, and thus the low fidelity of audio reconstruction and the instability of training. Summary of the Invention
[0003] The present disclosure provides a training method for an audio codec, an audio processing method, and an apparatus to at least improve the diversity and fidelity of audio representation. The technical solutions of the present disclosure are as follows: According to a first aspect of an embodiment of the present disclosure, a training method for an audio codec is provided. The preset audio codec to be trained includes a preset audio encoding model, a preset variational autoencoder, and a preset audio decoding model. The method includes: Obtaining sample audio features at the current training step; Inputting the sample audio features obtained by performing audio encoding processing on the sample audio features in the preset audio encoding model into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoding features and a latent space distribution; Inputting the sample variational encoding features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and obtaining first loss information based on the reconstructed audio features and the sample audio features; Determining the divergence distance between the latent space distribution and a standard Gaussian distribution; and obtaining divergence loss information based on the divergence distance and a preset divergence threshold when the divergence distance is greater than the preset divergence threshold; Adjusting the model parameters of the preset audio codec according to the first loss information and the divergence loss information; After the model parameters are adjusted, returning to the step of obtaining sample audio features at the current training step until the training iteration end condition is satisfied, and using the preset audio codec when the training iteration end condition is satisfied as the target audio codec.
[0004] According to a second aspect of the embodiments of the present disclosure, there is provided an audio processing method, including: Obtain a target text to be processed; and map the target text to an audio feature space to obtain intermediate audio features; Input the intermediate audio features into a target audio decoding model for audio decoding processing to obtain target audio features; the target audio decoding model is the preset audio decoding model in the target audio codec obtained by the training method according to any one of the above first aspects when the training iteration end condition is satisfied; Generate a target audio corresponding to the target text based on the target audio features.
[0005] According to a third aspect of the embodiments of the present disclosure, there is provided an audio processing method, including: An audio sending end inputs audio features of an audio to be transmitted into a target audio encoding model to obtain first encoded audio features; and inputs the first encoded audio features into a target variational autoencoder to output second encoded audio features; and sends the second encoded audio features to an audio receiving end; The audio receiving end inputs the received second encoded audio features into a target audio decoding model for audio decoding processing to obtain decoded audio features; Based on the decoded audio features, obtain a target reconstructed audio corresponding to the audio to be transmitted; Wherein, the target audio encoding model, the target variational autoencoder, and the target audio decoding model are respectively the preset audio encoding model, the preset variational autoencoder, and the preset audio decoding model in the target audio codec obtained by the training method according to any one of the above first aspects when the training iteration end condition is satisfied.
[0006] According to a fourth aspect of the embodiments of the present disclosure, there is provided a training device for an audio codec. The preset audio codec to be trained includes a preset audio encoding model, a preset variational autoencoder, and a preset audio decoding model; the device includes: A sample audio feature acquisition module, configured to acquire sample audio features at the current training step; An encoding module, configured to input the sample audio features into the preset audio encoding model for audio encoding processing to obtain sample audio encoded features, and input the sample audio encoded features into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoded features and a latent space distribution; A decoding module, configured to input the sample variational encoded features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and obtain first loss information based on the reconstructed audio features and the sample audio features; A divergence loss acquisition module, configured to determine the divergence distance between the latent space distribution and the standard Gaussian distribution; and in the case that the divergence distance is greater than a preset divergence threshold, obtain divergence loss information based on the divergence distance and the preset divergence threshold; A model parameter adjustment module, configured to adjust the model parameters of the preset audio codec according to the first loss information and the divergence loss information; A target audio codec acquisition module, configured to, after the model parameter adjustment, return to the step of acquiring the sample audio features at the current training step, until the training iteration end condition is satisfied, and use the preset audio codec when the training iteration end condition is satisfied as the target audio codec.
[0007] In a possible implementation manner, the apparatus further includes: A first model parameter adjustment module, configured to, in the case that the current training step is less than or equal to a first divergence training threshold, adjust the model parameters of the preset audio codec based on the first loss information.
[0008] In a possible implementation manner, the model parameter adjustment module includes: A second model parameter adjustment module, configured to, in the case that the current training step is greater than the first divergence training threshold and less than or equal to a first adversarial training threshold, obtain second loss information according to the first loss information and the divergence loss information; and adjust the model parameters of the preset audio codec based on the second loss information.
[0009] In a possible implementation manner, the second model parameter adjustment module includes: A dynamic divergence loss weight determination unit, configured to determine a dynamic divergence loss weight corresponding to the divergence loss information according to the current training step, the first divergence training threshold, a second divergence training threshold, and a preset divergence loss weight; the first divergence training threshold is less than the second divergence training threshold; A first divergence loss acquisition unit, configured to obtain first target divergence loss information based on the dynamic divergence loss weight and the divergence loss information; A second loss information acquisition unit, configured to use the sum of the first loss information and the first target divergence loss information as the second loss information.
[0010] In a possible implementation manner, the preset audio codec further includes a preset discriminator; the apparatus further includes: A discrimination result acquisition module, configured to input the reconstructed audio features and the sample audio features into the preset discriminator for discrimination processing to obtain a discrimination result; A discriminator loss and generator loss acquisition module, configured to obtain discriminator loss information and generator loss information based on the discrimination result; The model parameter adjustment module includes: A third model parameter adjustment module, configured to obtain third loss information according to the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and perform model parameter adjustment on the preset audio codec based on the third loss information.
[0011] In a possible implementation, the third model parameter adjustment module includes: A second divergence loss acquisition unit, configured to obtain second target divergence loss information based on a preset divergence loss weight and the divergence loss information when the current training step number is greater than a first adversarial training threshold and less than or equal to a second adversarial training threshold; An adversarial loss determination unit, configured to determine adversarial loss information based on the discriminator loss information and the generator loss information; A dynamic adversarial loss weight determination unit, configured to determine a dynamic adversarial loss weight corresponding to the adversarial loss information according to the current training step number, the first adversarial training threshold, the second adversarial training threshold, and a preset adversarial loss weight; A first adversarial loss acquisition unit, configured to obtain first adversarial loss information based on the dynamic adversarial loss weight and the adversarial loss information; A first model parameter adjustment unit, configured to obtain the first target loss information according to the first loss information, the second target divergence loss information, and the first adversarial loss information, and perform model parameter adjustment on the preset audio codec based on the first target loss information.
[0012] In a possible implementation, the third loss information includes second target loss information; the third model parameter adjustment module includes: A second divergence loss acquisition unit, configured to obtain second target divergence loss information based on a preset divergence loss weight and the divergence loss information when the current training step number is greater than the second adversarial training threshold; A second adversarial loss acquisition unit, configured to determine adversarial loss information based on the discriminator loss information and the generator loss information; and obtain second adversarial loss information based on a preset adversarial loss weight and the adversarial loss information; A second model parameter adjustment unit, configured to obtain the second target loss information according to the first loss information, the second target divergence loss information, and the second adversarial loss information, and perform model parameter adjustment on the preset audio decoding model and the preset discriminator based on the second target loss information.
[0013] In a possible implementation, the sample audio corresponding to the sample audio feature is an audio with a high sampling rate.
[0014] According to a fifth aspect of the embodiments of the present disclosure, there is provided an audio processing apparatus, including: A feature mapping module, configured to obtain a target text to be processed and map the target text to an audio feature space to obtain an intermediate audio feature; An audio decoding module, configured to input the intermediate audio feature into a target audio decoding model for audio decoding processing to obtain a target audio feature; the target audio decoding model is the preset audio decoding model in the target audio codec obtained by the training method according to any one of the first aspects above when the training iteration end condition is satisfied; An audio generation module, configured to generate a target audio corresponding to the target text based on the target audio feature.
[0015] According to a sixth aspect of the embodiments of the present disclosure, there is provided an audio processing apparatus, including: An audio sending end, configured to input an audio feature of an audio to be transmitted into a target audio encoding model to obtain a first encoded audio feature; and input the first encoded audio feature into a target variational autoencoder to output a second encoded audio feature; and send the second encoded audio feature to an audio receiving end; An audio receiving end, configured to input the received second encoded audio feature into a target audio decoding model for audio decoding processing to obtain a decoded audio feature; and obtain a target reconstructed audio corresponding to the audio to be transmitted based on the decoded audio feature; Wherein, the target audio encoding model, the target variational autoencoder, and the target audio decoding model are respectively the preset audio encoding model, the preset variational autoencoder, and the preset audio decoding model in the target audio codec obtained by the training method according to any one of the first aspects above when the training iteration end condition is satisfied.
[0016] According to a seventh aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the instructions to implement the method according to any one of the first aspect described above, or to implement the method according to the second aspect described above, or to implement the method according to the third aspect described above.
[0017] According to an eighth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method according to any one of the first aspect of the embodiments of the present disclosure, or the method according to the second aspect described above, or the method according to the third aspect described above. According to a ninth aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer instructions, when the computer instructions are executed by a processor, enabling the computer to execute the method according to any one of the first aspect of the embodiments of the present disclosure, or the method according to the second aspect described above, or the method according to the third aspect described above.
[0018] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: By obtaining sample audio features at the current training step; inputting the sample audio features into the preset audio encoding model for audio encoding processing to obtain sample audio encoding features, inputting the sample audio encoding features into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoding features and latent space distribution; inputting the sample variational encoding features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and based on the reconstructed audio features and the sample audio features, obtaining first loss information; determining the divergence distance between the latent space distribution and the standard Gaussian distribution; and in the case where the divergence distance is greater than a preset divergence threshold, obtaining divergence loss information based on the divergence distance and the preset divergence threshold; adjusting model parameters of the preset audio codec according to the first loss information and the divergence loss information; after the model parameters are adjusted, returning to the step of obtaining sample audio features at the current training step until a training iteration end condition is satisfied, and using the preset audio codec when the training iteration end condition is satisfied as the target audio codec.
[0019] Among them, by setting a preset divergence threshold to constrain the divergence distance between the latent space distribution and the standard Gaussian distribution, only when the divergence distance is greater than the preset divergence threshold, the divergence loss information is obtained based on the divergence distance and the preset divergence threshold. This enables the divergence loss to effectively avoid premature convergence to zero, thereby avoiding the collapse problem of the divergence loss, which can enhance the diversity of audio representations, have better robustness and noise resistance, and the training process is relatively stable. Based on this, the trained target audio codec can achieve high fidelity in audio reconstruction. In addition, the audio coding compression ratio in the target audio codec is also improved, greatly reducing the amount of audio transmission data and improving the audio transmission efficiency. Moreover, by using a preset audio coding model for audio coding before the preset variational autoencoder, the audio features are relatively rich, which can further avoid the collapse problem of the divergence loss corresponding to the preset variational autoencoder.
[0020] And by setting that the preset audio codec to be trained includes a preset audio coding model, a preset variational autoencoder, and a preset audio decoding model, it is possible to learn audio features through the preset audio coding model, learn the audio feature distribution based on the preset variational autoencoder, and realize the end-to-end joint training of audio encoding and decoding in combination with the preset audio decoding model, making the audio features learned by audio encoding and audio decoding more flexible and diverse, and capable of adapting to more audio application scenarios.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0022] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0023] Figure 1 is a schematic diagram of an application environment shown according to an exemplary embodiment.
[0024] Figure 2 is a flowchart of a method for training an audio codec shown according to an exemplary embodiment.
[0025] Figure 3 is a schematic diagram of a training architecture of an audio codec shown according to an exemplary embodiment.
[0026] Figure 4 is a schematic diagram of the process of staged training of an audio codec shown according to an exemplary embodiment.
[0027] Figure 5It is a schematic diagram of audio processing shown according to an exemplary embodiment.
[0028] Figure 6 It is another schematic diagram of audio processing shown according to an exemplary embodiment.
[0029] Figure 7 It is a block diagram of a training device for an audio codec shown according to an exemplary embodiment.
[0030] Figure 8 It is a block diagram of an electronic device for audio codec training or audio processing shown according to an exemplary embodiment.
[0031] Figure 9 It is a block diagram of an electronic device for audio codec training or audio processing shown based on an exemplary embodiment. Detailed implementation manners
[0032] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0034] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0035] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely applied in multiple fields. The solutions provided in the embodiments of the present application involve technologies such as machine learning / deep learning, and are specifically described through the following embodiments.
[0036] Please refer to Figure 1 , Figure 1It is a schematic diagram of an application environment shown according to an exemplary embodiment. As Figure 1 shown, the application environment may include server 01 and terminal 02.
[0037] In an alternative embodiment, server 01 may be used for the training process or audio processing of an audio codec. Specifically, server 01 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0038] In an alternative embodiment, terminal 02 may be combined with server 01 to trigger the above training, or display the training results, trigger audio generation, etc. Specifically, terminal 02 may include, but is not limited to, electronic devices such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. Optionally, the operating systems running on the electronic devices may include, but are not limited to, Android, IOS, Linux, Windows, etc.
[0039] In addition, it should be noted that Figure 1 what is shown is only an application environment of the audio codec training method provided by the present disclosure.
[0040] In the embodiments of the present specification, the above server 01 and terminal 02 may be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.
[0041] It should be noted that the following shows a possible step sequence, and actually does not limit that it must be strictly in this order. Some steps may be executed in parallel without dependence. The user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to data for display, training data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0042] Before introducing the method embodiments provided by the present application, a brief introduction to the application scenarios, related terms or nouns that may be involved in the method embodiments of the present application is given first, so as to facilitate the understanding of those skilled in the art of the present application.
[0043] TTS (Text-To-Speech): Text-to-speech synthesis is one of the core technologies of intelligent voice interaction. By converting the received text sequence into a natural and realistic speech waveform, it is fed back to the user. The text-to-speech synthesis technology directly affects the actual use effect of human-computer interaction. The text-to-speech synthesis technology involves multiple disciplines such as speech signal processing, pattern recognition, natural language processing, acoustics, and linguistics, and is an essential key technology in the field of information processing.
[0044] TTA (Text-to-Audio): It is a generative task that generates non-speech audio (such as environmental sound effects, music, scene sounds, etc.) based on text descriptions. Its core goal is to convert natural language input into high-fidelity and semantically matching audio signals through an algorithm model. Different from TTS, TTA focuses more on the synthesis of multi-element and time-sequential sound effects, which is more challenging.
[0045] VAE (Variational Autoencoder): A variational autoencoder transforms real samples into an ideal data distribution through an encoder network. The data distribution is passed to a decoder network to obtain generated samples. Further variational processing is performed on the autoencoder so that the output results of the variational autoencoder can correspond to the mean and variance of the target distribution. It should be noted that in this disclosure, the preset variational encoder and the target variational encoder are both VAEs, and the only difference is the VAE before and after training.
[0046] BERT (Bidirectional Encoder Representations from Transformers): A pre-trained language representation model based on the Transformer architecture that can be used for semantic recognition of text.
[0047] GAN: Generative Adversarial Network, which is a deep learning model that can generate new data similar to the training data through adversarial training of two neural networks. These two neural networks are the Generator and the Discriminator respectively.
[0048] Figure 2 It is a flowchart of a training method for an audio codec shown according to an exemplary embodiment. As Figure 2 shown, it may include the following steps.
[0049] In step S201, sample audio features at the current training step are obtained.
[0050] In the embodiments of this specification, the sample audio features may refer to the acoustic features corresponding to the sample audio, such as Mel spectrogram features, sampling points, spectral energy distribution extracted by Fourier transform, etc. This disclosure does not make any limitations in this regard.
[0051] In one example, the sample audio corresponding to the above sample audio features is an audio with a high sampling rate, such as a sample audio of 44.1 kHz. By setting the audio with a high sampling rate as the sample audio, the trained target audio codec can perform high-quality audio encoding and decoding processing.
[0052] The current training step number can refer to the current number of training steps being performed. The training step number can refer to the number of training iterations. Each training step number can correspond to a batch of training samples, such as a batch of sample audio features. After each batch of training samples undergoes an inference process based on the preset audio codec to be trained, the model parameters can be adjusted once. After completing one model parameter adjustment, the current training step number can be incremented by 1 as the current training step number to enter the next round of training iteration process. Exemplarily, the initial value of the training step number can be 0.
[0053] In one possible implementation, the sample audio at the current training step number can be obtained, and the Mel spectrogram features can be extracted from the sample audio as the sample audio features.
[0054] In step S203, the sample audio coding features obtained by inputting the sample audio features into the preset audio coding model for audio coding processing are input into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain the sample variational coding features and the latent space distribution.
[0055] Exemplarily, the preset audio codec to be trained can include a preset audio coding model, a preset variational autoencoder, and a preset audio decoding model, as Figure 3 shown. Among them, the preset audio coding model can refer to the audio coding model to be trained, such as a convolutional neural network (CNN), a recurrent neural network (RNN), etc. The preset variational autoencoder can refer to the VAE to be trained. The preset audio decoding model can refer to the audio decoding model to be trained, such as an audio decoding model based on CNN, an audio decoding model based on Transformer, etc. The present disclosure does not limit these.
[0056] Referring to Figure 3 , in one possible implementation, the sample audio features (for example, represented by Figure 3 in ) can be input into the preset audio coding model for audio coding processing to obtain the sample audio coding features (for example, represented by A in Figure 3 ). As an example, the sample audio coding feature A can be the encoding and compression result of , such as a sample audio coding vector.
[0057] Further, the sample audio encoding feature A can be input into a preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain a sample variational encoding feature (e.g., Figure 3 the Z output by the preset variational autoencoder in ) and a latent space distribution (e.g., denoted as
[0058] Exemplarily, the preset VAE (i.e., the preset variational autoencoder) can consist of three parts: an encoder, which is used to map the input data to the probability distribution of the latent space, i.e., the above-mentioned latent space distribution; latent space sampling, which is used to sample a latent variable from the probability distribution; and a decoder, which is used to reconstruct the latent variable into output data. For example, based on the following formula, the reparameterization method in latent space sampling can be used to obtain Z.
[0059] where Z can be a latent variable, i.e., the sample variational encoding feature. is the mean of the latent space distribution (e.g., denoted as ). is the standard deviation of the latent space distribution. ⊙ refers to element-wise multiplication, also known as the Hadamard product, i.e., the elements at the corresponding positions of two tensors of the same dimension are multiplied respectively. Here, it can be realized that and are multiplied element-wise to scale the randomness of the sampling.
[0060] can be an auxiliary random variable, following the standard normal distribution , where I is the identity matrix, meaning that the elements in each dimension independently follow a normal distribution with a mean of 0 and a variance of I . represents the standard multivariate normal distribution, with a mean of the zero vector and a covariance matrix of the identity matrix I , indicating that the random variables in each dimension are independent of each other and follow the standard normal distribution.
[0061] In step S205, the sample variational encoding feature is input into a preset audio decoding model for decoding processing to obtain a reconstructed audio feature; and based on the reconstructed audio feature and the sample audio feature, first loss information is obtained.
[0062] In one possible implementation, referring to Figure 3 , the sample variational encoding feature Z can be input into a preset audio decoding model for decoding processing, e.g., for mel spectrogram feature reconstruction processing, to obtain a reconstructed audio feature (e.g.,Figure 3 in )
[0063] Exemplarily, the reconstructed audio feature may correspond to the sample audio feature. For example, if the sample audio feature is a Mel spectrogram feature, the reconstructed audio feature may be a Mel spectrogram feature; if the sample audio feature is a sampling point, correspondingly, the reconstructed audio feature may be a sampling point.
[0064] Further, based on the reconstructed audio feature and the sample audio feature, first loss information can be obtained (for example, it can be expressed as L 1). Exemplarily, the first loss information can be calculated based on the mean squared error (MSE) loss function L 1, as shown in the following formula.
[0065]
[0066]
[0067] where refers to the sample audio feature, refers to the reconstructed audio feature.
[0068] In step S207, determine the divergence distance between the latent space distribution and the standard Gaussian distribution; and when the divergence distance is greater than a preset divergence threshold, based on the divergence distance and the preset divergence threshold, obtain divergence loss information.
[0069] In the embodiments of this specification, the preset divergence threshold can be set as a hyperparameter during the model training process, as a marginal constraint, to be used to constrain the gradient backpropagation process to stop before the divergence distance reaches 0, thereby avoiding the divergence collapse problem.
[0070] In a possible implementation manner, the divergence distance between the latent space distribution and the standard Gaussian distribution can be determined based on a preset divergence algorithm. And when the divergence distance is greater than the preset divergence threshold, based on the divergence distance and the preset divergence threshold, divergence loss information can be obtained. For example, the difference obtained by subtracting the preset divergence threshold from the divergence distance can be used as the divergence loss information. Exemplarily, the preset divergence algorithm may include the KL divergence algorithm, etc.
[0071] In one example, taking the KL divergence as an example, the function of step S207 can be implemented through the following formula, so that the divergence distance can stop calculating the gradient when it reaches the preset divergence threshold (represented by ).
[0072]
[0073] where max() is used to take the maximum value among them; refers to the divergence loss information; refers to the divergence distance (or KL divergence) between the latent space distribution and the standard Gaussian distribution, which is used to measure the difference between two probability distributions.
[0074] In step S209, according to the first loss information and the divergence loss information, the model parameters of the preset audio codec are adjusted.
[0075] In a possible implementation, the first loss information and the divergence loss information can be added up to obtain the total loss, so that the model parameters of the preset audio codec can be adjusted based on this total loss. For example, when the total loss is greater than a preset loss threshold (such as 0), the gradient can be calculated based on this total loss, so that the gradient backpropagation method can be used to adjust the model parameters of the preset audio codec. Specifically, the model parameters of the preset audio encoding model, the model parameters of the preset VAE, and the model parameters of the preset audio decoding model can all be adjusted.
[0076] In step S211, after the model parameters are adjusted, return to the step of obtaining the sample audio features at the current training step until the training iteration end condition is met, and use the preset audio codec when the training iteration end condition is met as the target audio codec.
[0077] In a possible implementation, after the model parameters are adjusted, the step of obtaining the sample audio features at the current training step can be returned to perform the next round of iterative training. For example, adding 1 to the current training step as the current training step of the new round of iteration. Thus, the updated sample audio features at the current training step can be obtained, and the above steps are repeated until the training iteration end condition is met, so that the preset audio codec when the training iteration end condition is met can be used as the target audio codec. Specifically, the preset audio encoding model when the training iteration end condition is met can be used as the trained target audio encoding model, the preset variational autoencoder when the training iteration end condition is met can be used as the trained target variational autoencoder, and the preset audio decoding model when the training iteration end condition is met can be used as the trained target audio decoding model.
[0078] Exemplarily, the training iteration end condition can include but is not limited to a preset loss threshold (such as 0), and the present disclosure does not limit this.
[0079] By setting a preset divergence threshold to constrain the divergence distance between the latent space distribution and the standard Gaussian distribution, divergence loss information is obtained based on the divergence distance and the preset divergence threshold only when the divergence distance is greater than the preset divergence threshold. This enables the divergence loss to effectively avoid premature convergence to zero, thereby preventing the collapse problem of the divergence loss, enhancing the diversity of audio representations, having better robustness and noise resistance, and ensuring a relatively stable training process. Based on this, the trained target audio codec can achieve high-fidelity reconstructed audio in audio reconstruction. In addition, the audio coding compression rate in the target audio codec is also improved, significantly reducing the amount of audio transmission data and enhancing the audio transmission efficiency. Moreover, by performing audio coding using a preset audio coding model before the preset variational autoencoder, the audio features are relatively rich, further avoiding the collapse problem of the divergence loss corresponding to the preset variational autoencoder.
[0080] And by setting that the preset audio codec to be trained includes a preset audio coding model, a preset variational autoencoder, and a preset audio decoding model, audio features can be learned through the preset audio coding model, the audio feature distribution can be learned based on the preset variational autoencoder, and end-to-end joint training of audio encoding and decoding can be achieved in combination with the preset audio decoding model, making the audio features learned by audio encoding and audio decoding more flexible and diverse and capable of adapting to more audio application scenarios.
[0081] Referring to Figure 3 , in an optional implementation, the above-mentioned preset audio codec may further include a preset discriminator, which may be the discriminator in a pre-trained generative adversarial network (GAN), and the present disclosure makes no limitation thereto. Based on this, the method may further include: inputting the reconstructed audio features and the sample audio features into the preset discriminator for discrimination processing to obtain a discrimination result; thereby, discriminator loss information and generator loss information can be obtained based on the discrimination result. As an example, the discriminator loss information can be calculated by the following formula and the generator loss information .
[0082]
[0083]
[0084]
[0085] where E represents the expectation of calculating data; represents the discrimination result of the preset discriminator on (which can be 1 or 0; for example, 1 can represent the real audio and 0 can represent the reconstructed audio); can represent gradient penalty; can represent an adversarial loss; , can be a preset coefficient, which is not limited in the present disclosure. max() is used to take the maximum value among them; refers to the sample audio feature, refers to the reconstructed audio feature; represents the discrimination result of the preset discriminator on .
[0086] Correspondingly, the above step S209 can be replaced by: obtaining third loss information according to the first loss information, divergence loss information, discriminator loss information, and generator loss information, and based on the third loss information, adjusting the model parameters of the preset audio codec, which may refer to adjusting the model parameters of the preset audio encoding model, preset VAE, preset audio decoding model, and preset discriminator. Exemplarily, the sum of the first loss information, divergence loss information, discriminator loss information, and generator loss information can be used as the third loss information, which is not limited in the present disclosure. For the content of model parameter adjustment, reference can be made to the corresponding introduction content above, which will not be elaborated here.
[0087] By combining the adversarial loss to train the preset audio codec, high-frequency detail features can be effectively supplemented, so that the reconstructed audio feature can be closer to the real audio feature, realizing high-fidelity audio decoding.
[0088] Before introducing the phased model training, some identifiers are introduced for the convenience of subsequent content description. Referring to Figure 4 , for example, the current training step can be represented as ; the first divergence training threshold can be represented as ; the second divergence training threshold can be represented as ; the first adversarial training threshold can be represented as ; the second adversarial training threshold can be represented as . Among them, the initial value of step can be 0. Exemplarily, 、 、 and can be the preset training step thresholds in the phased model training, which can be used to divide different training stages.
[0089] The first divergence training threshold can be used to indicate the starting training step corresponding to using the divergence loss information to participate in model parameter adjustment, that is, when it is greater than the first divergence training threshold, the divergence loss information is used to participate in the adjustment of model parameters.
[0090] Exemplarily, the second divergence training threshold may be greater than the first divergence training threshold and may be less than the first adversarial training threshold. For example, < < < 。 When the second divergence training threshold can be used to indicate that divergence loss information participates in the adjustment of model parameters, the following preset divergence loss weights are dynamically adjusted corresponding to the range of training steps.
[0091] The first adversarial training threshold can be used to indicate the starting training step corresponding to using adversarial loss information to participate in the adjustment of model parameters. That is, adversarial loss information is used to participate in the adjustment of model parameters only when it is greater than the first adversarial training threshold.
[0092] The second adversarial training threshold can be used to indicate the starting training step corresponding to using adversarial loss information to participate in the adjustment of model parameters and only adjusting the model parameters of the preset audio decoding model and the preset discriminator. That is, when it is greater than the second adversarial training threshold, the fine-tuning stage of the preset audio decoding model will be entered.
[0093] As an example, 、 、 And It can be obtained by pre-statistics based on the number of model training iterations (such as the total number of iterations from the start to the end of model training). For example, if the statistically obtained number of model training iterations is generally m times, then it can be set as: , that is, one-fourth of m; , that is, three-tenths of m; , that is, one-third of m; , that is, one-half of m.
[0094] The preset divergence loss weight can be expressed as , to characterize the upper limit of the divergence loss weight. Correspondingly, the dynamic divergence loss information can be expressed as . Exemplarily, the initial value can be 0.
[0095] The preset adversarial loss weight can be expressed as , to characterize the upper limit of the dynamic adversarial loss weight. Correspondingly, the dynamic adversarial loss weight can be expressed as , the initial value can be 0.
[0096] Refer to Figure 4, in a possible implementation, after the step of obtaining the first loss information, the method may further include: when the current training step is less than or equal to the first divergence training threshold (for example Figure 4 shown in ), which can be regarded as training stage one), based on the first loss information, adjust the parameters (or model parameters) of the preset audio codec. The first loss information can represent the audio reconstruction ability of the model. Setting to first adjust the parameters of the preset audio codec based on this first loss information can enable the preset audio codec to have basic audio reconstruction ability, and can reduce the computational complexity of the loss information and improve the training efficiency of the audio codec.
[0097] Referring to Figure 4 , in a possible implementation, the regularization enhancement stage can be entered. For example, the model parameters of the preset audio codec can be adjusted in combination with the divergence loss information. The regularization enhancement stage can refer to the situation where the current training step is greater than the first divergence training threshold and less than or equal to the first adversarial training threshold, that is, as Figure 4 shown in , which can be regarded as training stage two. Based on this, the above adjustment of the model parameters of the preset audio codec according to the first loss information and the divergence loss information can include: When the current training step is greater than the first divergence training threshold and less than or equal to the first adversarial training threshold, the second loss information can be obtained according to the first loss information and the divergence loss information. For example, the sum of the first loss information and the divergence loss information can be used as the second loss information. Further, based on this second loss information, the model parameters of the preset audio codec can be adjusted. By gradually introducing the divergence loss information of VAE, the collapse of VAE can be avoided, the robustness of model training can be improved, and a preset divergence threshold is set in the divergence loss information to constrain premature convergence to zero, avoiding the degradation problem of the latent space, thereby improving the diversity of audio representations, so that the trained target audio codec can obtain high-quality audio in downstream audio tasks.
[0098] In an example, the above obtaining the second loss information according to the first loss information and the divergence loss information can include: Determine the dynamic divergence loss weight corresponding to the divergence loss information according to the current training step, the first divergence training threshold, the second divergence training threshold, and the preset divergence loss weight; the first divergence training threshold is less than the second divergence training threshold. Exemplarily, the dynamic divergence loss weight can be obtained through the following formula .
[0099]
[0100] Further, based on the dynamic divergence loss weight and the divergence loss information, the first target divergence loss information can be obtained; for example, the product of the dynamic divergence loss weight and the divergence loss information can be used as the first target divergence loss information.
[0101] The sum of the first loss information and the first target divergence loss information is used as the second loss information. That is, the total loss of the first loss information and the first target divergence loss information can be used as the second loss information. Exemplarily, denoting the second loss information as L2, L2 can be obtained through the following formula.
[0102]
[0103] where, refers to the first target divergence loss information; refers to the divergence loss information; refers to the first loss information.
[0104] By setting the dynamic divergence loss weight, as the number of training steps increases, the weight of the divergence loss information in the total loss can be increased, so that the stability of model training can be controlled when the number of training steps increases.
[0105] In a possible implementation, the above third loss information may include the first target loss information. Correspondingly, the above obtaining the third loss information according to the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and adjusting the model parameters of the preset audio codec based on the third loss information may include: When the current number of training steps is greater than the first adversarial training threshold and less than or equal to the second adversarial training threshold, that is, as shown in Figure 4 the situation can be regarded as the third training stage. The second target divergence loss information can be obtained based on the preset divergence loss weight and the divergence loss information. For example, the product of the preset divergence loss weight and the divergence loss information can be used as the second target divergence loss information.
[0106] And, the adversarial loss information can be determined based on the discriminator loss information and the generator loss information. For example, the total loss of the discriminator loss information and the generator loss information can be used as the adversarial loss information.
[0107] Further, the dynamic adversarial loss weight corresponding to the adversarial loss information can be determined according to the current number of training steps, the first adversarial training threshold, the second adversarial training threshold, and the preset adversarial loss weight. Exemplarily, the dynamic adversarial loss weight can be obtained through the following formula .
[0108]
[0109] Based on the dynamic adversarial loss weight and the adversarial loss information, the first adversarial loss information is obtained. For example, the product of the dynamic adversarial loss weight and the adversarial loss information can be used as the first adversarial loss information.
[0110] According to the first loss information, the second target divergence loss information, and the first adversarial loss information, the first target loss information is obtained. For example, the sum of the first loss information, the second target divergence loss information, and the first adversarial loss information, i.e., the total loss, can be used as the first target loss information. And based on the first target loss information, the model parameters of the preset audio codec are adjusted. Here, it can refer to adjusting the model parameters of the preset audio coding model, the preset VAE, the preset audio decoding model, and the preset discriminator. Exemplarily, L Let 31 denote the first target loss information, which can be obtained through the following formula L 31.
[0111]
[0112] where, refers to the first loss information; denotes the second target divergence loss information; refers to the first adversarial loss information; for other parameters, please refer to the corresponding introduction above and will not be elaborated here.
[0113] By introducing the adversarial loss information, high-frequency details can be effectively supplemented, enabling subsequent audio reconstruction to be closer to the real audio and improving the quality of audio processing.
[0114] In a possible implementation, the above third loss information may include the second target loss information. Correspondingly, the above obtaining the third loss information according to the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and adjusting the model parameters of the preset audio codec based on the third loss information, may include: When the current training step number is greater than the second adversarial training threshold, i.e., in the case as Figure 4 shown in this case, it can be regarded as training stage four. The second target divergence loss information can be obtained based on the preset divergence loss weight and the divergence loss information. For example, the product of the preset divergence loss weight and the divergence loss information can be used as the second target divergence loss information.
[0115] Further, based on the discriminator loss information and the generator loss information, adversarial loss information is determined; and based on a preset adversarial loss weight and the adversarial loss information, second adversarial loss information is obtained. For example, the sum of the discriminator loss information and the generator loss information can be used as the adversarial loss information. Based on this, the product of the preset adversarial loss weight and the adversarial loss information can be used as the second adversarial loss information.
[0116] Moreover, based on the first loss information, the second objective divergence loss information, and the second adversarial loss information, second objective loss information can be obtained, and based on the second objective loss information, the model parameters of the preset audio decoding model and the preset discriminator are adjusted.
[0117] Exemplarily, the sum of the first loss information, the second objective divergence loss information, and the second adversarial loss information can be used as the second objective loss information. For example, L Let 32 represent the second objective loss information, which can be obtained through the following formula L 32.
[0118]
[0119] where refers to the first loss information; represents the second objective divergence loss information; refers to the second adversarial loss information; for other parameters, please refer to the corresponding introduction above and will not be elaborated here.
[0120] By introducing the adversarial loss information and for the preset audio decoding model and the preset discriminator, the preset audio decoding model can be further fine-tuned, so that subsequent audio reconstruction can be closer to the real audio and the quality of audio processing can be improved.
[0121] Refer to Figure 4 , it should be noted that after each model parameter adjustment, it can be determined whether the training iteration end condition is met. If the training iteration end condition is met, the training can be ended. If the training iteration end condition is not met, step can be updated by adding 1.
[0122] In a possible implementation manner, the present disclosure also provides an audio processing method, which may include: Obtain the target text to be processed; and the target text can be mapped to the audio feature space to obtain intermediate audio features. Exemplarily, the target text can be the text to be processed, such as lines, scene description text, etc., and the present disclosure does not limit this. As an example, the target text can be input into the BERT model to obtain the text semantic representation, and then this semantic representation can be input into the GAN network to map the text semantic representation to the audio feature space to obtain intermediate audio features, such as intermediate audio vectors.
[0123] Furthermore, the intermediate audio features can be input into the target audio decoding model for audio decoding processing to obtain the target audio features. Here, the target audio decoding model is the preset audio decoding model in the target audio codec obtained by the above training method when the training iteration end condition is satisfied. Exemplarily, the target audio features can correspond to the reconstructed audio features in the above training process, and can include but are not limited to Mel spectrum features, sampling points, etc.
[0124] And, based on the target audio features, the target audio corresponding to the target text can be generated. The target audio can include speech and non-speech, and the non-speech can include environmental sound effects, music, scene sounds, etc. Optionally, the target audio can also be stored.
[0125] In one example, referring to Figure 5 , if the target audio features are Mel spectrum features, a vocoder can be used to generate the target audio corresponding to the target text from the target audio features.
[0126] In another example, if the target audio features are sampling points, the vocoder can be omitted and the waveform corresponding to the sampling points can be used as the target audio.
[0127] In an application example, for example, during the process of generating an audio-visual video, for example, the image frames of a silent video and the lines corresponding to each image frame can be obtained, and then the lines can be used as the Figure 5 target text in, and through the above audio processing, the target audio corresponding to the lines can be obtained, so that the lines corresponding to each image frame, the target audio corresponding to each image frame, and each image frame can be combined into an audio-visual video, and based on the generation from text to audio, the generation from a silent video to an audio-visual video can be realized.
[0128] By using the target audio decoding model obtained by the above training method for the text-to-audio generation process, the audio reconstruction quality based on the text is higher, effectively optimizing the effect of the audio generation task, thereby making the video processing scenario more extensive.
[0129] In one possible implementation manner, the present disclosure also provides an audio processing method, and this method can include: The audio sending end inputs the audio features of the audio to be transmitted into the target audio encoding model to obtain the first encoded audio features; and inputs the first encoded audio features into the target variational autoencoder to output the second encoded audio features; and sends the second encoded audio features to the audio receiving end; Correspondingly, the audio receiving end can input the received second encoded audio features into the target audio decoding model for audio decoding processing to obtain decoded audio features.
[0130] The above processing process can refer to the corresponding steps in the above training process and will not be elaborated here.
[0131] Further, based on the decoded audio features, the target reconstructed audio corresponding to the audio to be transmitted is obtained. Here, the process of generating the target audio based on the target audio features can be referred to and will not be elaborated here either.
[0132] Among them, the target audio encoding model, the target variational autoencoder, and the target audio decoding model are successively the preset audio encoding model, the preset variational autoencoder, and the preset audio decoding model in the target audio codec obtained by the above training method when the training iteration end condition is met.
[0133] By using the target audio codec obtained by the above training method for audio encoding transmission and decoding processing, high compression rate can be achieved in audio encoding, thereby reducing the data transmission volume and improving the audio transmission efficiency; and high-fidelity reconstruction of the audio can be obtained at the audio receiving end based on the target audio decoding model, so that both the audio transmission efficiency and the decoding fidelity are relatively high.
[0134] Figure 7 It is a block diagram of a training device for an audio codec shown according to an exemplary embodiment. The preset audio codec to be trained may include a preset audio encoding model, a preset variational autoencoder, and a preset audio decoding model. Refer to Figure 7 , the device may include: A sample audio feature acquisition module 701, configured to acquire sample audio features at the current training step; An encoding module 703, configured to input the sample audio encoding features obtained by performing audio encoding processing on the sample audio features into the preset audio encoding model into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoding features and a latent space distribution; A decoding module 705, configured to input the sample variational encoding features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and obtain first loss information based on the reconstructed audio features and the sample audio features; The divergence loss acquisition module 707 is used to determine the divergence distance between the latent space distribution and the standard Gaussian distribution; and in the case that the divergence distance is greater than a preset divergence threshold, based on the divergence distance and the preset divergence threshold, obtain divergence loss information; The model parameter adjustment module 709 is used to adjust the model parameters of the preset audio codec according to the first loss information and the divergence loss information; The target audio codec acquisition module 711 is used to, after the model parameters are adjusted, return to the step of acquiring the sample audio features at the current training step number until the training iteration end condition is met, and use the preset audio codec when the training iteration end condition is met as the target audio codec.
[0135] By setting a preset divergence threshold to constrain the divergence distance between the latent space distribution and the standard Gaussian distribution, in the case that the divergence distance is greater than the preset divergence threshold, the divergence loss information is obtained based on the divergence distance and the preset divergence threshold. This enables the divergence loss to effectively avoid premature convergence to zero, thereby avoiding the collapse problem of the divergence loss, which can enhance the diversity of audio representations, have better robustness and noise resistance, and the training process is relatively stable. Based on this, the target audio codec obtained through training can achieve high fidelity in the reconstructed audio during audio reconstruction. And the audio coding compression rate in the target audio codec is also improved, greatly reducing the amount of audio transmission data and improving the audio transmission efficiency.
[0136] And by setting the preset audio codec to be trained to include a preset audio coding model, a preset variational autoencoder, and a preset audio decoding model, the audio features can be learned through the preset audio coding model, the audio feature distribution can be learned based on the preset variational autoencoder, and the end-to-end joint training of audio encoding and decoding can be realized in combination with the preset audio decoding model, making the audio features learned by audio encoding and audio decoding more flexible and diverse, and capable of adapting to more audio application scenarios.
[0137] In a possible implementation manner, the device further includes: The first model parameter adjustment module is used to, in the case that the current training step number is less than or equal to the first divergence training threshold, adjust the model parameters of the preset audio codec based on the first loss information.
[0138] In a possible implementation manner, the model parameter adjustment module 709 may include: A second model parameter adjustment module, configured to obtain second loss information according to the first loss information and the divergence loss information when the current training step is greater than a first divergence training threshold and less than or equal to a first adversarial training threshold; and adjust the model parameters of the preset audio codec based on the second loss information.
[0139] In a possible implementation manner, the second model parameter adjustment module includes: A dynamic divergence loss weight determination unit, configured to determine a dynamic divergence loss weight corresponding to the divergence loss information according to the current training step, the first divergence training threshold, a second divergence training threshold, and a preset divergence loss weight; the first divergence training threshold is less than the second divergence training threshold; A first divergence loss acquisition unit, configured to obtain first target divergence loss information based on the dynamic divergence loss weight and the divergence loss information; A second loss information acquisition unit, configured to use the sum of the first loss information and the first target divergence loss information as the second loss information.
[0140] In a possible implementation manner, the preset audio codec further includes a preset discriminator; the apparatus further includes: A discrimination result acquisition module, configured to input the reconstructed audio feature and the sample audio feature into the preset discriminator for discrimination processing to obtain a discrimination result; A discriminator loss and generator loss acquisition module, configured to obtain discriminator loss information and generator loss information based on the discrimination result; The model parameter adjustment module 709 may include: A third model parameter adjustment module, configured to obtain third loss information according to the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and adjust the model parameters of the preset audio codec based on the third loss information.
[0141] In a possible implementation manner, the third model parameter adjustment module includes: A second divergence loss acquisition unit, configured to obtain second target divergence loss information based on a preset divergence loss weight and the divergence loss information when the current training step is greater than the first adversarial training threshold and less than or equal to a second adversarial training threshold; An adversarial loss determination unit, configured to determine adversarial loss information based on the discriminator loss information and the generator loss information; A dynamic adversarial loss weight determination unit, configured to determine a dynamic adversarial loss weight corresponding to the adversarial loss information according to a current training step, the first adversarial training threshold, the second adversarial training threshold, and a preset adversarial loss weight; A first adversarial loss obtaining unit, configured to obtain first adversarial loss information based on the dynamic adversarial loss weight and the adversarial loss information; A first model parameter adjustment unit, configured to obtain first target loss information according to the first loss information, the second target divergence loss information, and the first adversarial loss information, and perform model parameter adjustment on the preset audio codec based on the first target loss information.
[0142] In a possible implementation manner, the third loss information includes second target loss information; the third model parameter adjustment module includes: A second divergence loss obtaining unit, configured to obtain second target divergence loss information based on a preset divergence loss weight and the divergence loss information when the current training step is greater than the second adversarial training threshold; A second adversarial loss obtaining unit, configured to determine adversarial loss information based on the discriminator loss information and the generator loss information; and obtain second adversarial loss information based on a preset adversarial loss weight and the adversarial loss information; A second model parameter adjustment unit, configured to obtain second target loss information according to the first loss information, the second target divergence loss information, and the second adversarial loss information, and perform model parameter adjustment on the preset audio decoding model and the preset discriminator based on the second target loss information.
[0143] In a possible implementation manner, the sample audio corresponding to the sample audio feature is an audio with a high sampling rate.
[0144] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0145] An embodiment of the present disclosure further provides an audio processing device, which may include: A feature mapping module, configured to obtain a target text to be processed and map the target text to an audio feature space to obtain intermediate audio features; An audio decoding module, configured to input the intermediate audio feature into a target audio decoding model for audio decoding processing to obtain a target audio feature; the target audio decoding model is the preset audio decoding model in the target audio codec obtained based on the training method described in any one of the above first aspects when the training iteration end condition is satisfied; An audio generation module, configured to generate a target audio corresponding to the target text based on the target audio feature.
[0146] An embodiment of the present disclosure further provides an audio processing device, which may include: An audio sending end, configured to input an audio feature of an audio to be transmitted into a target audio encoding model to obtain a first encoded audio feature; and input the first encoded audio feature into a target variational autoencoder to output a second encoded audio feature; and send the second encoded audio feature to an audio receiving end; An audio receiving end, configured to input the received second encoded audio feature into a target audio decoding model for audio decoding processing to obtain a decoded audio feature; and obtain a target reconstructed audio corresponding to the audio to be transmitted based on the decoded audio feature; Wherein, the target audio encoding model, the target variational autoencoder, and the target audio decoding model are sequentially the preset audio encoding model, the preset variational autoencoder, and the preset audio decoding model in the target audio codec obtained based on the training method described in any one of the above first aspects when the training iteration end condition is satisfied.
[0147] Figure 8 is a block diagram of an electronic device for audio codec training or audio processing shown according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as shown in Figure 8 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Wherein, the processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program stored in the non-volatile storage medium to run. The network interface of the electronic device is configured to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a method for audio codec training or audio processing. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad, or a mouse, etc. Those skilled in the art can understand that Figure 8 the structure shown in Figure 8 is only a block diagram of some structures related to the present disclosure solution, and does not constitute a limitation on the electronic device to which the present disclosure solution is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0148] Figure 9 is a block diagram of an electronic device for audio codec training or audio processing shown based on an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as shown in Figure 9 . The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for audio codec training or audio processing. Those skilled in the art can understand that Figure 9 the structure shown in Figure 9 is only a block diagram of some structures related to the present disclosure solution, and does not constitute a limitation on the electronic device to which the present disclosure solution is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the audio codec training or audio processing method as in the embodiments of the present disclosure.
[0149] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the audio codec training or audio processing method in the embodiments of the present disclosure. The computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0150] In an exemplary embodiment, a computer program product containing instructions is further provided. When it runs on a computer, the computer executes the audio codec training or audio processing method in the embodiments of the present disclosure.
[0151] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0152] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are to be considered as illustrative only, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0153] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for an audio codec, characterized in that, The preset audio codec to be trained includes a preset audio encoding model, a preset variational autoencoder, and a preset audio decoding model; the method includes: Obtain the sample audio features at the current training step. Input the sample audio features obtained by performing audio encoding processing on the sample audio features in the preset audio encoding model into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoding features and a latent space distribution. Input the sample variational encoding features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and based on the reconstructed audio features and the sample audio features, obtain first loss information. Determine the divergence distance between the latent space distribution and the standard Gaussian distribution; and in the case where the divergence distance is greater than a preset divergence threshold, obtain divergence loss information based on the divergence distance and the preset divergence threshold. Adjust the model parameters of the preset audio codec according to the first loss information and the divergence loss information. After the model parameters are adjusted, return to the step of obtaining the sample audio features at the current training step until the training iteration end condition is satisfied, and use the preset audio codec when the training iteration end condition is satisfied as the target audio codec.
2. The method according to claim 1, characterized in that After the step of obtaining the first loss information, the method further includes: In the case where the current training step is less than or equal to the first divergence training threshold, adjust the model parameters of the preset audio codec based on the first loss information.
3. The method according to claim 1, wherein The adjusting the model parameters of the preset audio codec according to the first loss information and the divergence loss information includes: In the case where the current training step is greater than the first divergence training threshold and less than or equal to the first adversarial training threshold, obtain second loss information according to the first loss information and the divergence loss information; and adjust the model parameters of the preset audio codec based on the second loss information.
4. The method according to claim 3, characterized in that, The obtaining the second loss information according to the first loss information and the divergence loss information includes: Determine a dynamic divergence loss weight corresponding to the divergence loss information according to the current training step, the first divergence training threshold, the second divergence training threshold, and a preset divergence loss weight; the first divergence training threshold is less than the second divergence training threshold. Obtain first target divergence loss information based on the dynamic divergence loss weight and the divergence loss information. Use the sum of the first loss information and the first target divergence loss information as the second loss information.
5. The method according to claim 1, wherein The preset audio codec further includes a preset discriminator; the method further includes: Input the reconstructed audio features and the sample audio features into the preset discriminator for discrimination processing to obtain a discrimination result. Obtain discriminator loss information and generator loss information based on the discrimination result. The adjusting the model parameters of the preset audio codec according to the first loss information and the divergence loss information includes: Based on the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, third loss information is obtained, and based on the third loss information, the model parameters of the preset audio codec are adjusted.
6. The method according to claim 5, wherein The third loss information includes first target loss information; the obtaining of the third loss information based on the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and the adjusting of the model parameters of the preset audio codec based on the third loss information includes: When the current training step is greater than the first adversarial training threshold and less than or equal to the second adversarial training threshold, based on a preset divergence loss weight and the divergence loss information, second target divergence loss information is obtained; Based on the discriminator loss information and the generator loss information, adversarial loss information is determined; According to the current training step, the first adversarial training threshold, the second adversarial training threshold, and a preset adversarial loss weight, a dynamic adversarial loss weight corresponding to the adversarial loss information is determined; Based on the dynamic adversarial loss weight and the adversarial loss information, first adversarial loss information is obtained; According to the first loss information, the second target divergence loss information, and the first adversarial loss information, the first target loss information is obtained, and based on the first target loss information, the model parameters of the preset audio codec are adjusted.
7. The method according to claim 5, wherein The third loss information includes second target loss information; the obtaining of the third loss information based on the first loss information, the divergence loss information, the discriminator loss information, and the generator loss information, and the adjusting of the model parameters of the preset audio codec based on the third loss information includes: When the current training step is greater than the second adversarial training threshold, based on a preset divergence loss weight and the divergence loss information, second target divergence loss information is obtained; Based on the discriminator loss information and the generator loss information, adversarial loss information is determined; and based on a preset adversarial loss weight and the adversarial loss information, second adversarial loss information is obtained; According to the first loss information, the second target divergence loss information, and the second adversarial loss information, the second target loss information is obtained, and based on the second target loss information, the model parameters of the preset audio decoding model and the preset discriminator are adjusted.
8. The method according to claim 1, characterized in that, The sample audio corresponding to the sample audio feature is an audio with a high sampling rate.
9. An audio processing method, characterized in that, The method includes: Obtaining a target text to be processed; and mapping the target text to an audio feature space to obtain intermediate audio features; Inputting the intermediate audio features into a target audio decoding model for audio decoding processing to obtain target audio features; the target audio decoding model is the preset audio decoding model in the target audio codec obtained by the training method according to any one of claims 1 to 8 when the training iteration end condition is satisfied; Based on the target audio features, a target audio corresponding to the target text is generated.
10. An audio processing method, characterized in that, The method includes: The audio sending end inputs the audio features of the audio to be transmitted into the target audio encoding model to obtain the first encoded audio features; and inputs the first encoded audio features into the target variational autoencoder to output the second encoded audio features; and sends the second encoded audio features to the audio receiving end; The audio receiving end inputs the received second encoded audio features into the target audio decoding model for audio decoding processing to obtain decoded audio features; Based on the decoded audio features, the target reconstructed audio corresponding to the audio to be transmitted is obtained; Wherein, the target audio encoding model, the target variational autoencoder, and the target audio decoding model are respectively the target audio codec obtained by the training method according to any one of claims 1 to 8, the preset audio encoding model, the preset variational autoencoder, and the preset audio decoding model when the training iteration end condition is satisfied.
11. A training device for an audio codec, characterized in that, The preset audio codec to be trained includes a preset audio encoding model, a preset variational autoencoder, and a preset audio decoding model; the device includes: A sample audio feature acquisition module, configured to acquire sample audio features at the current training step; An encoding module, configured to input the sample audio features obtained by performing audio encoding processing on the sample audio features into the preset audio encoding model into the preset variational autoencoder for latent space distribution prediction and feature sampling processing to obtain sample variational encoded features and a latent space distribution; A decoding module, configured to input the sample variational encoded features into the preset audio decoding model for decoding processing to obtain reconstructed audio features; and based on the reconstructed audio features and the sample audio features, obtain first loss information; A divergence loss acquisition module, configured to determine the divergence distance between the latent space distribution and the standard Gaussian distribution; and in the case where the divergence distance is greater than a preset divergence threshold, obtain divergence loss information based on the divergence distance and the preset divergence threshold; A model parameter adjustment module, configured to adjust the model parameters of the preset audio codec according to the first loss information and the divergence loss information; A target audio codec acquisition module, configured to, after the model parameters are adjusted, return to the step of acquiring the sample audio features at the current training step until the training iteration end condition is satisfied, and use the preset audio codec when the training iteration end condition is satisfied as the target audio codec.
12. An audio processing device, characterized in that, Includes: A feature mapping module, configured to acquire the target text to be processed and map the target text to the audio feature space to obtain intermediate audio features; An audio decoding module, configured to input the intermediate audio features into the target audio decoding model for audio decoding processing to obtain target audio features; the target audio decoding model is the preset audio decoding model in the target audio codec obtained by the training method according to any one of claims 1 to 8 when the training iteration end condition is satisfied; An audio generation module, configured to generate the target audio corresponding to the target text based on the target audio features.
13. An audio processing system, characterized in that, Includes: An audio sending end, which is used to input the audio features of the audio to be transmitted into a target audio coding model to obtain a first coded audio feature; and input the first coded audio feature into a target variational autoencoder to output a second coded audio feature; and send the second coded audio feature to an audio receiving end; An audio receiving end, which is used to input the received second coded audio feature into a target audio decoding model for audio decoding processing to obtain a decoded audio feature; and based on the decoded audio feature, obtain a target reconstructed audio corresponding to the audio to be transmitted; wherein the target audio coding model, the target variational autoencoder, and the target audio decoding model are respectively the target audio codec obtained by the training method according to any one of claims 1 to 8, the preset audio coding model, the preset variational autoencoder, and the preset audio decoding model when the training iteration end condition is satisfied.
14. An electronic device, characterized in that, Comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the training method according to any one of claims 1 to 8, or to implement the audio processing method according to claim 9 or 10.
15. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method according to any one of claims 1 to 8, or the electronic device is enabled to execute the audio processing method according to claim 9 or 10.
16. A computer program product, characterized in that, Comprising computer instructions, which when executed by a processor, cause the computer to execute the training method according to any one of claims 1 to 8, or cause the computer to execute the audio processing method according to claim 9 or 10.
Citation Information
Patent Citations
Multimodal fusion audio generation method and device based on diffusion model
CN116884391A
Model training method and device, speech recognition method and device, equipment and storage medium
CN117437904A
Audio signal processing method and device, computer equipment and storage medium
CN119314499A
Systems and methods for parallel wave generation in end-to-end text-to-speech
US20190180732A1
Data compression using jointly trained encoder, decoder, and prior neural networks
US20210004677A1
Cited By
Training method and device of audio coding and decoding model
CN120766693A
A method and apparatus for training an audio codec model
CN120766693B