Audio processing method, apparatus, system, and storage medium

By using an audio conversion model to encode the audio to be processed into a target discrete coded signal and then decode it into the target audio, the problem of audio distortion at low bit rates is solved, high-precision audio conversion and true audio restoration are achieved, the operation process is simplified and hardware costs are reduced.

CN119252265BActive Publication Date: 2026-03-20PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing audio encoding and decoding technologies suffer from severe audio distortion at low bit rates, affecting the listening experience and rendering them impractical.

Method used

An audio conversion model is adopted, in which the audio to be processed is encoded into the target continuous coded signal by the encoder, and then quantized into the target discrete coded signal based on the preset codebook vector. The decoder then decodes it into the target audio. The audio conversion model is trained with reconstruction loss, codebook loss and encoding/decoding preservation loss as constraints.

Benefits of technology

It improves audio conversion accuracy, reduces audio distortion, enhances the auditory effect of output audio, simplifies the operation process, and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119252265B_ABST
    Figure CN119252265B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio processing method, device, system and storage medium, and relates to the technical field of data processing, which at least solves the problem of serious distortion of output audio in the related art. The method comprises: inputting the audio to be processed into an audio conversion model to perform the following operations using the audio conversion model: calling an encoder to encode the audio to be processed into a target continuous coding signal; based on a preset codebook, the target continuous coding signal is vector quantized into a target discrete coding signal; calling a decoder to decode the target discrete coding signal into a target audio composed of a continuous target output coding signal based on the preset codebook; wherein the preset codebook comprises an association mapping relationship between the continuous coding signal and the discrete coding signal, and the audio conversion model is trained with the reconstruction loss, the codebook loss and the encoding and decoding preservation loss of the encoder and the decoder as the constraint target; and outputting the target audio obtained after the audio conversion model completes the operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to an audio processing method, device, system and storage medium. BACKGROUND

[0002] With the development demand of voice intelligence in various service fields such as intelligent medical treatment and financial field, audio processing has become an increasingly growing application trend, and audio coding technology has important application value in audio digital signal processing. In order to ensure that the audio data can be efficiently stored, transmitted and played, in the current audio coding process, the audio signal is usually converted from analog to digital, and the data size is reduced through compression and encoding, and then recovered to analog signal through decoding and reconstruction, so as to realize the conversion of audio signal to analog signal output. Since the above-mentioned analog signal is obtained based on the signal compression processing technology of reducing data size through compression and encoding, some frequency bands are filtered with a small number of bits, so that the output audio is seriously distorted, especially in the case of extremely low bit number (for example, 16 kilobits per second or less), the above-mentioned audio coding mode will cause the output audio to be obviously distorted, greatly affecting the hearing effect, and not having practical value. SUMMARY

[0003] The present application provides an audio processing method, device, system and storage medium to at least solve the problem of serious distortion of the output audio in the related art. The technical solution of the present application is as follows:

[0004] According to a first aspect of the embodiment of the present application, an audio processing method is provided, which comprises: inputting the audio to be processed into an audio conversion model to perform the following operations by using the audio conversion model: calling an encoder to encode the audio to be processed into a target continuous coding signal; based on a preset codebook, vector quantizing the target continuous coding signal into a target discrete coding signal; calling a decoder to decode the target discrete coding signal into a target audio composed of continuous target output coding signals based on the preset codebook; wherein the preset codebook comprises an associated mapping relationship between the continuous coding signal and the discrete coding signal, and the audio conversion model is trained with a reconstruction loss, a codebook loss and a coding and decoding keeping loss of the encoder and the decoder as constraint targets; and outputting the target audio obtained after the audio conversion model completes the operation.

[0005] In an implementation manner, the reconstruction loss represents an audio signal loss between the input sample audio inputting the audio conversion model and the output audio corresponding to the output of the audio conversion model; the codebook loss represents an encoding loss between the continuous encoding signal output by the encoder and the discrete encoding signal processed based on the preset codebook vector quantization, in a case that the first network parameter of the encoder and the second network parameter of the decoder are determined; and the preservation loss represents an encoding loss between the continuous encoding signal output by the encoder and the discrete encoding signal processed based on the preset codebook vector quantization, in a case that the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoding signal and the discrete encoding signal in the preset codebook are determined.

[0006] In another implementation manner, the constraint target includes a first constraint target and a second constraint target; and before the audio to be processed is input to the audio conversion model, the method further includes: alternately performing the following training processes until the preset codebook meeting the first constraint target and the encoder under the first network parameter and the decoder under the second network parameter meeting the second constraint target are obtained: training the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoding signal and the discrete encoding signal in the preset codebook, with the first constraint target that the reconstruction loss is less than a first loss threshold and the codebook loss is less than a second loss threshold; and training the first network parameter of the encoder and the second network parameter of the decoder, with the second constraint target that the reconstruction loss is less than the first loss threshold and the preservation loss is less than a third loss threshold.

[0007] In another implementation manner, training the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoding signal and the discrete encoding signal in the preset codebook, with the first constraint target that the reconstruction loss is less than the first loss threshold and the codebook loss is less than the second loss threshold, includes: when the codebook loss is greater than or equal to the second loss threshold, determining, from each code word, a target code word between the discrete encoding signal and the continuous encoding signal with a signal difference greater than or equal to a first preset difference; determining a first gradient representing a change trend of the code word value of the target code word, based on a linear relationship represented between the code word value corresponding to the target code word and adjacent code words; and adjusting the code word value of the target code word according to the first gradient.

[0008] In another implementation manner, adjusting the code word value of the target code word according to the first gradient includes: when the first gradient is greater than or equal to 0, reducing the code word value of the target code word; and when the first gradient is less than 0, increasing the code word value of the target code word.

[0009] In another implementation, the first network parameter of the encoder and the second network parameter of the decoder are trained with the second constraint target that the reconstruction loss is less than the first loss threshold and the preservation loss is less than the third loss threshold, including: determining a third gradient of the discrete encoded signal after the vector quantization processing of the continuous encoded signal vector according to a second gradient of a variation trend between the continuous encoded signals output by the encoder; when the preservation loss is greater than or equal to the third loss threshold, if a signal loss between the continuous encoded signal and the corresponding sample continuous encoded signal is greater than a fourth loss threshold, adjusting the first network parameter of the encoder according to a mapping relationship between the second gradient and the first network parameter; and if a signal loss between the discrete encoded signal and the corresponding sample discrete encoded signal is greater than a fifth loss threshold, adjusting the second network parameter of the decoder according to a mapping relationship between the third gradient and the second network parameter.

[0010] In another implementation, the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoded signal and the discrete encoded signal in the preset codebook are trained with the first constraint target that the reconstruction loss is less than the first loss threshold and the codebook loss is less than the second loss threshold, including: adjusting the code word values corresponding to each code word in the preset codebook when the reconstruction loss is greater than or equal to the first loss threshold and / or the codebook loss is greater than or equal to the second loss threshold; and the first network parameter of the encoder and the second network parameter of the decoder are trained with the second constraint target that the reconstruction loss is less than the first loss threshold and the preservation loss is less than the third loss threshold, including: adjusting the first network parameter and the second network parameter when the reconstruction loss is greater than or equal to the first loss threshold and / or the preservation loss is greater than or equal to the third loss threshold.

[0011] According to a second aspect of the embodiments of the present application, an audio processing apparatus is provided, which comprises: an input unit configured to input an audio to be processed into an audio conversion model, so as to perform the following operations by using the audio conversion model: calling an encoder to encode the audio to be processed into a target continuous encoded signal; vector quantizing the target continuous encoded signal into a target discrete encoded signal based on a preset codebook; calling a decoder to decode the target discrete encoded signal into a target audio composed of continuous target output encoded signals based on the preset codebook; wherein the preset codebook comprises an associated mapping relationship between the continuous encoded signal and the discrete encoded signal, and the audio conversion model is trained with the constraint targets of the reconstruction loss, the codebook loss and the codec preservation loss of the encoder and the decoder; and an output unit configured to output the target audio obtained after the audio conversion model completes the operations.

[0012] According to a third aspect of the embodiments of the present application, an audio processing system is provided, which comprises an encoder, a decoder and a preset codebook, and the system is provided with an audio conversion model, and the system is configured to perform the audio processing method of the first aspect and any possible implementation manner thereof.

[0013] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement the audio processing method according to the first aspect and any possible implementation thereof.

[0014] According to a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores instructions, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the audio processing method according to the first aspect and any possible implementation thereof.

[0015] According to a sixth aspect of the embodiments of the present application, a computer program product is provided, and the computer program product comprises computer instructions, when the computer instructions are run on an electronic device, the electronic device executes the audio processing method according to the first aspect and any possible implementation thereof.

[0016] The embodiments of the present application provide at least the following beneficial effects: directly converting the to-be-processed audio into the output target audio according to the audio conversion model, simple implementation, easy operation, no need for additional signal processing modules, and reduction of hardware cost and operation process. In the process of directly converting the to-be-processed audio into the target audio by using the audio conversion model, the target continuous coding signal vector in the encoder is quantized into a target discrete coding signal based on a preset codebook, so as to ensure that the audio signals of each frequency band in the original audio can be converted into the corresponding target discrete coding signal, thereby ensuring that the target discrete coding signal input into the decoder can more completely represent the original audio information, so that the decoder can better restore the original audio information according to the preset codebook, thereby ensuring that the output audio is more realistic. At the same time, the audio conversion model is formed by constraining the target training in three dimensions of reconstruction loss, codebook loss, and coding and decoding maintenance loss of the encoder and the decoder, thereby ensuring the accuracy of the encoder, the decoder, and the preset codebook, and improving the audio conversion accuracy of the audio conversion model.

[0017] The above audio signal conversion process directly converts the output to-be-processed audio into the target audio based on the audio conversion model, rather than relying on signal processing in the related art to filter and compress the signal, which can better restore and retain the audio information of the original audio, thereby reducing the distortion of the output audio and improving the auditory effect of the output audio.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure and, without in any way intending to limit the present disclosure, do not constitute an improper limitation thereof.

[0020] Figure 1 is a schematic diagram of an audio processing system according to an exemplary embodiment;

[0021] Figure 2 is a flowchart of an audio processing method according to an exemplary embodiment;

[0022] Figure 3 is a schematic diagram of an audio processing process according to an exemplary embodiment;

[0023] Figure 4 is a block diagram of an audio processing device according to an exemplary embodiment;

[0024] Figure 5 is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0025] In order to make the ordinary person in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings.

[0026] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0027] Before the audio processing method provided by the embodiments of the present application is described in detail, the application scenarios and implementation environments involved in the embodiments of the present application are briefly introduced.

[0028] First, the application scenarios involved in the present application are briefly introduced.

[0029] With the development of voice intelligence in various service fields such as intelligent medical care and finance, audio processing has become an increasingly growing application trend, and audio coding technology has important application value in audio digital signal processing. In order to ensure that audio data can be efficiently stored, transmitted and played, in the current audio coding process, the audio signal is usually converted from analog to digital, and the data size is reduced through compression and encoding, and then restored to analog signal through decoding and reconstruction, so as to realize the conversion of audio signal to analog signal output. Since the above-mentioned analog signal is obtained based on the signal compression processing technology of reducing data size through compression and encoding, some frequency bands with less bit number signals are filtered, resulting in serious distortion of the output audio, especially in the case of extremely low bit number (for example, 16 kilobits per second or less), the above-mentioned audio coding method will cause the output audio to have obvious distortion, which greatly affects the hearing effect and has no practical value.

[0030] Generally speaking, there is much redundant information in audio signal, and it is high and not practical to directly transmit the original audio signal. The traditional signal processing method generally encodes the frequency band less sensitive to human ear using less bit number through modeling of human ear and psychoacoustics, so as to reduce the total bit number of signal.

[0031] It is found through research that in the case of extremely low bit number, the current audio coding method will cause the audio to have obvious distortion, which greatly affects the hearing effect and has no practical value. With the development of deep learning technology, data-driven audio signal processing method has become possible. Vector quantized varitional autoencoder (VQ-VAE) can encode data blocks into fixed discrete codes by quantizing the intermediate hidden variables of VAE (i.e. generation model), and restore the discrete codes to data through decoder.

[0032] To solve the above problems, the present application provides an audio processing method, which directly converts the to-be-processed audio into the output target audio according to an audio conversion model, is simple and easy to operate, does not require an additional signal processing module, and reduces hardware cost and operation process. In the process of directly converting the to-be-processed audio into the target audio by using the audio conversion model, the target continuous coding signal vector in the encoder is quantized into a target discrete coding signal based on a preset codebook, so as to ensure that the audio signal of each frequency band in the original audio can be converted into the corresponding target discrete coding signal, thereby ensuring that the target discrete coding signal input into the decoder can more completely represent the original audio information, so that the decoder can better restore the original audio information according to the preset codebook, thereby ensuring that the output audio is more real. At the same time, the audio conversion model is formed by constraining the target training in three dimensions of reconstruction loss, codebook loss, and coding and decoding maintenance loss of the encoder and the decoder, thereby ensuring the accuracy of the encoder, the decoder, and the preset codebook, and improving the audio conversion accuracy of the audio conversion model.

[0033] Secondly, the implementation architecture related to the present application is briefly introduced below.

[0034] Figure 1 FIG. 1 is a schematic diagram of an audio processing system 10 provided by the present disclosure. As shown in FIG. 1, the audio processing system 10 includes an encoder 101 and a decoder 102, and is provided with a preset codebook. Figure 1

[0035] In some embodiments, the encoder 101 and the decoder 102 can be connected through a wired network or a wireless network. The preset codebook can be referred to as a codebook or a coding book.

[0036] In some embodiments, the audio processing system 10 can be arranged in a terminal device in a service platform such as a consultation platform or a financial platform.

[0037] Specifically, in the process of transmitting the signal, the sender of the voice signal or the audio signal needs a trained encoder and a preset codebook, and the receiver needs a preset codebook and a decoder paired with the encoder. The voice sender can encode each coding signal in the codebook into an integer, i.e., the coding signal can be represented by coding, so that a 2 k ​A codebook of codes (i.e., codewords) can be encoded by k bits per frame. Further, for an input 16 kHz sample rate speech audio signal, the frame length is 10 milliseconds per frame by down-sampling the encoder, and 4096 discrete codes are used in the codebook, thus the bit rate in the proposed codec model is 1.2 kbps, which is much lower than the bit rate of related signal processing algorithms. The receiver queries the codebook according to the bits of the received discrete code signal to obtain the corresponding code of each frame, and inputs the decoder to obtain the final restored speech audio signal.

[0038] In some other embodiments, the terminal device can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) \ virtual reality (VR) device, or the like, which can install and use a content community application (such as Kuaishou), and the specific form of the terminal device is not specially limited in the present disclosure. It can interact with the user through one or more ways such as a keyboard, a touchpad, a touch screen, a remote controller, voice interaction, or a handwriting device.

[0039] In addition, the number and type of terminal devices are not limited in the present application.

[0040] The audio processing method provided in the embodiments of the present application can be applied to the audio processing system in the implementation architecture shown in the foregoing Figure 1 For ease of understanding, the audio processing method provided in the present application is specifically introduced below in combination with the drawings.

[0041] Figure 2 is a flowchart of an audio processing method according to an exemplary embodiment, as shown in Figure 2 The audio processing method includes the following steps.

[0042] S21, input the audio to be processed to an audio conversion model, to perform the following operations by using the audio conversion model: calling an encoder to encode the audio to be processed into a target continuous code signal; based on a preset codebook, vector quantizing the target continuous code signal into a target discrete code signal; calling a decoder to decode the target discrete code signal into a target audio composed of continuous target output code signals based on the preset codebook.

[0043] The preset codebook includes an associated mapping relationship between the continuous encoded signal and the discrete encoded signal, and the audio conversion model is trained with a reconstruction loss, a codebook loss, and a codec preservation loss of the encoder and the decoder as constraint targets.

[0044] The reconstruction loss represents an audio signal loss between an input sample audio input into the audio conversion model and an output audio output by the audio conversion model.

[0045] The codebook loss represents an encoding loss between the continuous encoded signal output by the encoder and the discrete encoded signal processed by the preset codebook vector quantization under a condition that the first network parameter of the encoder and the second network parameter of the decoder are determined.

[0046] The preservation loss represents an encoding loss between the continuous encoded signal output by the encoder and the discrete encoded signal processed by the preset codebook vector quantization under a condition that code word values corresponding to each code word representing the associated mapping relationship between the continuous encoded signal and the discrete encoded signal in the preset codebook are determined.

[0047] The audio conversion model can be a VQ-VAE model based on a VAE generation model and modified based on vector quantization, i.e., a vector quantization generation model.

[0048] The above vector quantization is a technique for discretizing continuous data. Specifically, it achieves data compression and feature extraction by mapping input data to a set of discrete code vectors (codebook). The set of discrete code vectors is the discrete encoded signal in this application.

[0049] In an embodiment, in order to extract and compress features by reducing the dimensionality of data while retaining important information, the above-mentioned encoder is a down-sampling encoder. The down-sampling encoder is used to extract global features of the speech audio, and then these features are passed to the up-sampling layer (decoder) for pixel-level classification.

[0050] The down-sampling encoder is usually stacked by multiple convolutional layers and pooling layers. The convolutional layer performs convolution operation on the input audio to be processed by sliding the convolution kernel, thereby extracting local features. After each pooling layer, the size of the feature map is reduced, and the number of channels (the number of feature maps) is usually increased. This hierarchical structure enables the network to gradually capture higher-level features.

[0051] The down-sampling encoder samples the audio to be processed according to the sampling points following the Gaussian distribution to mark the discrete encoded signal corresponding to the continuous encoded signal.

[0052] In an embodiment, as shown in Figure 3 the audio processing process is described as follows for the audio to be processed being speech.

[0053] In the model application process, the input voice is passed through an encoder based on a convolution model, which includes stacked convolution layers and multi-step convolution layers (strided convolution) for downsampling, to obtain a continuous encoding after downsampling. The nearest discrete encoding to the continuous encoding is found in the codebook, and the discrete encoding is input into a decoder based on convolution layers and transposed convolution to obtain a restored voice signal. The codebook includes code 1, code 2, code 3, …, code k, etc.

[0054] In the model training process, since the process from continuous encoding to discrete encoding is not differentiable, the gradient at the continuous encoding needs to be copied to the discrete encoding. A combination of three loss functions is used in the training process. First, the reconstruction loss of the input audio and the output audio. The sampling points of the voice audio signal are generally Gaussian distributed, so the MSE (Mean Squared Error) loss function is used to determine the reconstruction loss. At the same time, the codebook loss and the preservation loss are also used.

[0055] Specifically, the codebook loss is the MSE loss between the continuous encoding and the discrete encoding of the codebook trained by the fixed codec layer, and there is a linear relationship between the values of the code words in the codebook. The preservation loss is the MSE loss between the continuous encoding and the discrete encoding of the codec trained by the fixed codebook. These two losses can maintain the stability of the codebook during training while training the codebook and the codec.

[0056] In the above implementation steps, in the process of directly converting the to-be-processed audio into the target audio by using the audio conversion model, the target continuous encoding signal vector in the encoder is quantized into a target discrete encoding signal based on the preset codebook, so as to ensure that the audio signals of each frequency band in the original audio can be converted into the corresponding target discrete encoding signal, thereby ensuring that the target discrete encoding signal input into the decoder can more completely represent the original audio information, so that the decoder can better restore the original audio information according to the preset codebook, thereby ensuring that the output audio is more realistic.

[0057] At the same time, the audio conversion model is trained by constraining the target training in three dimensions of the reconstruction loss, the codebook loss, and the codec preservation loss of the encoder and the decoder, thereby ensuring the accuracy of the encoder, the decoder, and the preset codebook, and thereby improving the audio conversion accuracy of the audio conversion model.

[0058] S22, the target audio obtained after the output audio conversion model completes the operation.

[0059] The above directly converts the to-be-processed audio into the target audio according to the audio conversion model.

[0060] Through the above embodiment, the audio signal conversion process is based on the audio conversion model to directly convert the output to-be-processed audio into target audio, rather than relying on signal processing in the related art to filter and compress the signal, which can better restore and retain the audio information of the original audio, thereby reducing the distortion of the output audio and improving the auditory effect of the output audio.

[0061] As a refinement and extension of the foregoing embodiment, in order to fully describe the specific implementation process of the present embodiment, the present embodiment provides another audio processing method.

[0062] In one embodiment, the constraint target includes a first constraint target and a second constraint target. Based on this, before step S21, the audio conversion model is trained, and the following training steps are alternately performed until the preset codebook that meets the first constraint target and the encoder under the first network parameter and the decoder under the second network parameter that meet the second constraint target are obtained.

[0063] First, the first constraint target is that the reconstruction loss is less than a first loss threshold and the codebook loss is less than a second loss threshold, and the codebook values corresponding to each code word in the preset codebook that represents the association mapping relationship between the continuous coded signal and the discrete coded signal are trained.

[0064] Second, the second constraint target is that the reconstruction loss is less than the first loss threshold and the preservation loss is less than a third loss threshold, and the first network parameter of the encoder and the second network parameter of the decoder are trained.

[0065] Through the above implementation steps, the encoder and the decoder that are trained are included in the audio conversion model, and the preset codebook that is trained is included in the audio conversion model, so as to ensure the accuracy of the audio conversion model.

[0066] As one embodiment, the specific steps of training the codebook values corresponding to each code word in the preset codebook are as follows.

[0067] First, when the codebook loss is greater than or equal to the second loss threshold, the target code word whose signal difference between the discrete coded signal and the continuous coded signal is greater than or equal to a first preset difference is determined from each code word.

[0068] Second, based on the linear relationship represented by the codebook values corresponding to the target code word and the adjacent code words, a first gradient representing the change trend of the codebook values of the target code word is determined.

[0069] It can be understood that based on the linear relationship represented by the codebook values corresponding to each code word, a first gradient representing the change trend between the codebook values of each code word is determined.

[0070] Thirdly, according to the first gradient, the code word value of the target code word is adjusted correspondingly.

[0071] In some embodiments, the first gradient is represented by a derivative obtained by differentiating a curve of the linear relationship.

[0072] Specifically, the above-mentioned adjustment of the code word value of the target code word according to the first gradient includes the following two cases.

[0073] Firstly, when the first gradient is greater than or equal to 0, the code word value of the target code word is reduced.

[0074] Specifically, when the first gradient is greater than or equal to 0, it indicates that the code word value is too large, and the code word value of the target code word can be reduced according to a preset reduction step.

[0075] Secondly, when the first gradient is less than 0, the code word value of the target code word is increased.

[0076] Specifically, when the first gradient is less than 0, it indicates that the code word value is too small, and the code word value of the target code word can be increased according to a preset increase step.

[0077] As an embodiment, the training steps of the first network parameter of the encoder and the second network parameter of the decoder are as follows.

[0078] Firstly, according to the second gradient of the change trend between the continuous encoding signals output by the encoder, the third gradient of the discrete encoding signals after the vector quantization processing of the continuous encoding signals is determined.

[0079] Because in the training process, the process from the continuous encoding signal to the discrete encoding signal is not differentiable, the gradient of the continuous encoding signal is mapped and copied to the corresponding discrete encoding signal to ensure that the training process is.

[0080] Secondly, when the keeping loss is greater than or equal to the third loss threshold, if the signal loss between the continuous encoding signal and the corresponding sample continuous encoding signal is greater than the fourth loss threshold, the first network parameter of the encoder is adjusted according to the mapping relationship between the second gradient and the first network parameter; and if the signal loss between the discrete encoding signal and the corresponding sample discrete encoding signal is greater than the fifth loss threshold, the second network parameter of the decoder is adjusted according to the mapping relationship between the third gradient and the second network parameter.

[0081] In some embodiments, the first constraint target is that the reconstruction loss is less than a first loss threshold and the codebook loss is less than a second loss threshold, and the code word values corresponding to each code word in the preset codebook are trained, which specifically includes the following steps: when the reconstruction loss is greater than or equal to the first loss threshold and / or the codebook loss is greater than or equal to the second loss threshold, the code word values corresponding to each code word in the preset codebook are adjusted.

[0082] The second constraint target is that the reconstruction loss is less than a first loss threshold and the preservation loss is less than a third loss threshold, and the first network parameter of the encoder and the second network parameter of the decoder are trained, which specifically includes: when the reconstruction loss is greater than or equal to the first loss threshold and / or the preservation loss is greater than or equal to the third loss threshold, the first network parameter and the second network parameter are adjusted.

[0083] To implement the above functions, the audio processing apparatus includes hardware structures and / or software modules corresponding to each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application of the technical solution and the design constraints. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered beyond the scope of the present application.

[0084] The embodiments of the present disclosure also provide an audio processing apparatus as shown in Figure 4 The apparatus includes an input unit 401, an output unit 402, and a training unit 403.

[0085] The input unit 401 is configured to input a to-be-processed audio into an audio conversion model to perform the following operations by using the audio conversion model: calling an encoder to encode the to-be-processed audio into a target continuous encoding signal; based on a preset codebook, vector quantizing the target continuous encoding signal into a target discrete encoding signal; calling a decoder to decode the target discrete encoding signal into a target audio composed of a continuous target output encoding signal based on the preset codebook; and the preset codebook includes an associated mapping relationship between a continuous encoding signal and a discrete encoding signal, and the audio conversion model is trained with a reconstruction loss, a codebook loss, and a preservation loss of encoding and decoding of the encoder and the decoder as constraint targets.

[0086] The output unit 402 is configured to output the target audio obtained after the audio conversion model completes the operation.

[0087] In an implementation, the reconstruction loss represents an audio signal loss between the input sample audio inputting the audio conversion model and the output audio outputted by the audio conversion model; the codebook loss represents an encoding loss between the continuous encoding signal outputted by the encoder and the discrete encoding signal processed based on the preset codebook vector quantization, in a case that the first network parameter of the encoder and the second network parameter of the decoder are determined; and the preservation loss represents an encoding loss between the continuous encoding signal outputted by the encoder and the discrete encoding signal processed based on the preset codebook vector quantization, in a case that the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoding signal and the discrete encoding signal in the preset codebook are determined.

[0088] In another implementation, the constraint target includes a first constraint target and a second constraint target; and the training unit 403 is further configured to alternately execute the following training processes until the preset codebook satisfying the first constraint target and the encoder with the first network parameter and the decoder with the second network parameter satisfying the second constraint target are obtained: training the code word values corresponding to each code word representing the associated mapping relationship between the continuous encoding signal and the discrete encoding signal in the preset codebook, with the first constraint target that the reconstruction loss is less than a first loss threshold and the codebook loss is less than a second loss threshold; and training the first network parameter of the encoder and the second network parameter of the decoder, with the second constraint target that the reconstruction loss is less than the first loss threshold and the preservation loss is less than a third loss threshold.

[0089] In another implementation, the training unit 403 is specifically configured to: when the codebook loss is greater than or equal to the second loss threshold, determine, from each code word, a target code word with a signal difference between the discrete encoding signal and the continuous encoding signal greater than or equal to a first preset difference; determine a first gradient representing a change trend of the code word value of the target code word based on a linear relationship represented between the code word value of the target code word and adjacent code words; and adjust the code word value of the target code word according to the first gradient.

[0090] In another implementation, adjusting the code word value of the target code word according to the first gradient includes: when the first gradient is greater than or equal to 0, reducing the code word value of the target code word; and when the first gradient is less than 0, increasing the code word value of the target code word.

[0091] In another implementation, the training unit 403 is specifically configured to: determine a third gradient of the discrete coded signals after the vector quantization processing of the continuous coded signals according to a second gradient of the variation trend between the continuous coded signals output by the encoder; when the signal loss of the continuous coded signals and the corresponding sample continuous coded signals is greater than a fourth loss threshold value while the maintaining loss is greater than or equal to the third loss threshold value, adjust the first network parameters of the encoder according to the mapping relationship between the second gradient and the first network parameters; and when the signal loss of the discrete coded signals and the corresponding sample discrete coded signals is greater than a fifth loss threshold value, adjust the second network parameters of the decoder according to the mapping relationship between the third gradient and the second network parameters.

[0092] In another implementation, the training unit 403 is specifically configured to: adjust the code word values corresponding to each code word in the preset codebook when the reconstruction loss is greater than or equal to the first loss threshold value and / or the codebook loss is greater than or equal to the second loss threshold value; and train the first network parameters of the encoder and the second network parameters of the decoder with the second constraint target that the reconstruction loss is less than the first loss threshold value and the maintaining loss is less than the third loss threshold value, including: adjusting the first network parameters and the second network parameters when the reconstruction loss is greater than or equal to the first loss threshold value and / or the maintaining loss is greater than or equal to the third loss threshold value.

[0093] As to the apparatus in the above-mentioned embodiments, the specific manner in which each unit module performs the operation has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0094] Figure 5 is a schematic diagram of an electronic device provided by the present application. As Figure 5 The electronic device 50 can include at least one processor 501 and a memory 503 for storing processor-executable instructions. The processor 501 is configured to execute the instructions in the memory 503 to implement the audio processing method in the following embodiments.

[0095] In addition, the electronic device 50 can further include a communication bus 502, at least one communication interface 504, an input device 506, and an output device 505.

[0096] The processor 501 can be a central processing unit (CPU), a micro processing unit, an ASIC, or one or more integrated circuits for controlling the execution of programs of the present application.

[0097] The communication bus 502 can include a path for transmitting information between the above-mentioned components.

[0098] The communication interface 504, using any transceiver-like mechanism, is used to communicate with other devices or communication networks, such as an Ethernet network, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0099] The input device 506 is used to receive input signals and the output device 505 is used to output signals.

[0100] The memory 503 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing instructions or data that can be accessed by a computer, but not limited to. The memory can exist independently, connected to the processing unit through a bus. The memory can also be integrated with the processing unit.

[0101] The memory 503 is configured to store instructions for implementing the solutions of the present application, and the processor 501 is configured to control the execution of the instructions stored in the memory 503. The processor 501 is configured to execute the instructions stored in the memory 503, thereby realizing the functions in the methods of the present application.

[0102] In a specific implementation, as an example, the processor 501 can include one or more CPUs, such as the CPU0 and the CPU1 in the Figure 5 In a specific implementation, as an example, the processor 501 can include one or more CPUs, such as the CPU0 and the CPU1 in the

[0103] In a specific implementation, as an example, the electronic device 50 can include multiple processors, such as the processor 501 and the processor 507 in the Figure 5 In a specific implementation, as an example, the electronic device 50 can include multiple processors, such as the processor 501 and the processor 507 in the

[0104] The electronic device is, for example, Figure 5The apparatus shown includes: a processor 501 and a memory 503 for storing instructions executable by the processor 501; wherein the processor 501 is configured to execute the executable instructions to implement the audio processing method of any of the possible implementation manners described above. And can achieve the same technical effects, to avoid repetition, here will not repeat.

[0105] The embodiments of the present application also provide a computer readable storage medium, when the instructions in the computer readable storage medium are executed by the processor of the audio processing apparatus or the electronic device, the audio processing apparatus or the electronic device can execute the audio processing method of any of the possible implementation manners described above. And can achieve the same technical effects, to avoid repetition, here will not repeat.

[0106] The embodiments of the present application also provide a computer program product, including a computer program or instructions, the computer program or instructions are executed by the processor to execute the audio processing method of any of the possible implementation manners described above. And can achieve the same technical effects, to avoid repetition, here will not repeat.

[0107] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The specification and examples given are intended as illustrative only and not limiting of the true scope and spirit of the application. What is intended to be claimed is set forth in the following claims.

[0108] It should be understood that the application is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the appended claims.

Claims

1. An audio processing method, characterized in that, The method includes: The audio to be processed is input into an audio conversion model, which performs the following operations: an encoder is invoked to encode the audio into a target continuous coded signal; based on a preset codebook, the target continuous coded signal vector is quantized into a target discrete coded signal; and a decoder is invoked to decode the target discrete coded signal into target audio composed of continuous target output coded signals based on the preset codebook. The preset codebook includes the correlation mapping relationship between continuous and discrete coded signals, and the audio conversion model is trained with reconstruction loss, codebook loss, and the encoding / decoding preservation loss of the encoder and decoder as constraints. Output the target audio obtained by the audio conversion model after completing the operation; The reconstruction loss characterizes the audio signal loss between the input sample audio of the audio conversion model and the corresponding output audio of the audio conversion model; the codebook loss characterizes the coding loss between the continuous coded signal output by the encoder and the discrete coded signal quantized based on the preset codebook vector, given that the first network parameters of the encoder and the second network parameters of the decoder are determined; the preservation loss characterizes the coding loss between the continuous coded signal output by the encoder and the discrete coded signal quantized based on the preset codebook vector, given that the codeword values ​​corresponding to each codeword representing the correlation mapping relationship between the continuous coded signal and the discrete coded signal in the preset codebook are determined.

2. The audio processing method according to claim 1, characterized in that, The constraint objectives include a first constraint objective and a second constraint objective; before inputting the audio to be processed into the audio conversion model, the method further includes: The following training process is performed alternately until the preset encoder that satisfies the first constraint objective and the encoder under the first network parameters and the decoder under the second network parameters that satisfy the second constraint objective are obtained: With the reconstruction loss being less than a first loss threshold and the codebook loss being less than a second loss threshold as the first constraint objectives, the codeword values ​​corresponding to each codeword in the preset codebook that represent the correlation mapping relationship between continuous and discrete coded signals are trained. Furthermore, the first network parameters of the encoder and the second network parameters of the decoder are trained with the reconstruction loss being less than the first loss threshold and the retention loss being less than the third loss threshold as the second constraint objectives.

3. The audio processing method according to claim 2, characterized in that, The step of training the codeword values ​​corresponding to each codeword in the preset encoding book, which represents the correlation mapping relationship between continuous and discrete encoded signals, with the reconstruction loss being less than a first loss threshold and the codebook loss being less than a second loss threshold as the first constraint objective, includes: When the codebook loss is greater than or equal to the second loss threshold, a target codeword is determined from each codeword whose signal difference between the discrete coded signal and the continuous coded signal is greater than or equal to a first preset difference. Based on the linear relationship between the codeword value corresponding to the target codeword and adjacent codewords, a first gradient representing the changing trend of the codeword value of the target codeword is determined. The codeword value of the target codeword is adjusted accordingly based on the first gradient.

4. The audio processing method according to claim 3, characterized in that, The step of adjusting the codeword value of the target codeword according to the first gradient includes: When the first gradient is greater than or equal to 0, the codeword value of the target codeword is reduced; When the first gradient is less than 0, the codeword value of the target codeword is increased.

5. The audio processing method according to claim 2, characterized in that, The step of training the first network parameters of the encoder and the second network parameters of the decoder, with the reconstruction loss being less than a first loss threshold and the retention loss being less than a third loss threshold as the second constraint objectives, includes: The third gradient of the discrete encoded signal after vector quantization of the continuous encoded signal is determined according to the second gradient of the changing trend between the continuous encoded signals output by the encoder. When the retention loss is greater than or equal to the third loss threshold, if the signal loss between the continuous encoded signal and the corresponding sample continuous encoded signal is greater than the fourth loss threshold, the first network parameters of the encoder are adjusted according to the mapping relationship between the second gradient and the first network parameters; and if the signal loss between the discrete encoded signal and the corresponding sample discrete encoded signal is greater than the fifth loss threshold, the second network parameters of the decoder are adjusted according to the mapping relationship between the third gradient and the second network parameters.

6. The audio processing method according to any one of claims 2 to 5, characterized in that, The step of training the codeword values ​​corresponding to each codeword in the preset encoding book, which represent the correlation mapping relationship between continuous and discrete encoded signals, with the reconstruction loss being less than a first loss threshold and the codebook loss being less than a second loss threshold as the first constraint target, includes: adjusting the codeword values ​​corresponding to each codeword in the preset encoding book when the reconstruction loss is greater than or equal to the first loss threshold and / or the codebook loss is greater than or equal to the second loss threshold; The step of training the first network parameters of the encoder and the second network parameters of the decoder, with the reconstruction loss being less than the first loss threshold and the retention loss being less than the third loss threshold as the second constraint objectives, includes: If the reconstruction loss is greater than or equal to the first loss threshold and / or the retention loss is greater than or equal to the third loss threshold, the first network parameters and the second network parameters are adjusted.

7. An audio processing device, characterized in that, The device includes: An input unit is used to input the audio to be processed into an audio conversion model, which then performs the following operations: invoking an encoder to encode the audio into a target continuous coded signal; quantizing the target continuous coded signal vector into a target discrete coded signal based on a preset codebook; and invoking a decoder to decode the target discrete coded signal into target audio composed of continuous target output coded signals based on the preset codebook. The preset codebook includes the correlation mapping relationship between continuous and discrete coded signals, and the audio conversion model is trained with reconstruction loss, codebook loss, and the encoding / decoding preservation loss of the encoder and decoder as constraints. The output unit is used to output the target audio obtained by the audio conversion model after completing the operation; The reconstruction loss characterizes the audio signal loss between the input sample audio of the audio conversion model and the corresponding output audio of the audio conversion model; the codebook loss characterizes the coding loss between the continuous coded signal output by the encoder and the discrete coded signal quantized based on the preset codebook vector, given that the first network parameters of the encoder and the second network parameters of the decoder are determined; the preservation loss characterizes the coding loss between the continuous coded signal output by the encoder and the discrete coded signal quantized based on the preset codebook vector, given that the codeword values ​​corresponding to each codeword representing the correlation mapping relationship between the continuous coded signal and the discrete coded signal in the preset codebook are determined.

8. An audio processing system, characterized in that, The system includes an encoder, a decoder, and a preset codec, and the system is equipped with an audio conversion model. The system is configured to perform the audio processing method as described in any one of claims 1-6.

9. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the audio processing method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Voice conversion method, voice conversion device, electronic equipment and storage medium

    CN115206333A

  • Voice conversion method and device, equipment and medium

    CN118072749A