Audio coding method, electronic device, and program product
By using a pre-trained codec neural network model for audio encoding and decoding, the problem of poor speech quality at ultra-low bit rates is solved, achieving efficient audio compression and high-quality audio restoration, which is suitable for communication environments with limited bandwidth.
Patent Information
- Application Number
- CN202511270796.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing audio encoding and decoding methods cannot effectively maintain speech quality at ultra-low bit rates, resulting in distorted speaker timbre or unclear speech, which affects the voice communication experience.
By employing a pre-trained encoding and decoding neural network model, and through local acoustic feature extraction, intra-frame dependency modeling, and inter-frame dependency modeling, combined with multi-layer quantization processing, efficient feature extraction and compression are achieved to meet the communication needs of different bit rates.
Achieving high-quality audio reproduction at ultra-low bitrates is suitable for bandwidth-constrained communication scenarios, improving the quality and efficiency of voice communication.
Smart Images

Figure CN120783775B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, and in particular to an audio coding method, an electronic device and a program product. BACKGROUND
[0002] In modern real-time communication systems, there are often scenarios where communication signals are weak, such as underground spaces, tunnels, elevator rooms, ring-shaped building interiors, remote areas, and post-disaster rescue sites. In such a weak network communication environment, audio coding technology needs to support narrowband transmission at an ultra-low code rate. However, existing audio coding methods cannot effectively maintain speech quality at an ultra-low code rate. The reconstructed audio after decoding performs poorly in timbre restoration and semantic clarity, often exhibiting speaker timbre distortion or unclear statements, which affects the speech communication experience.
[0003] Therefore, how to realize high-quality audio restoration in an ultra-low code rate communication environment is a problem that needs to be solved at present. SUMMARY
[0004] The present application provides an audio coding method, an electronic device and a program product to solve the problem that existing audio coding methods cannot realize high-quality audio restoration in an ultra-low code rate communication environment.
[0005] The present application provides an audio coding method, comprising:
[0006] Obtaining audio data to be processed, and converting the audio data to be processed into a spectrum graph;
[0007] Inputting the spectrum graph into an encoder of a pre-trained coding and decoding neural network model, sequentially performing local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling on the spectrum graph, and obtaining an encoding vector output by the encoder;
[0008] Inputting the encoding vector into a residual vector quantizer of the coding and decoding neural network model, performing multi-layer quantization processing on the encoding vector, and obtaining a codebook index output by the residual vector quantizer;
[0009] Inputting the codebook index into a decoder of the coding and decoding neural network model, performing decoding processing on the codebook index, and obtaining a reconstructed spectrum graph output by the decoder.
[0010] According to the audio coding method provided by the present application, the spectrum graph is input into the encoder of the pre-trained coding and decoding neural network model, and the spectrum graph is sequentially subjected to local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling, and an encoding vector output by the encoder is obtained, comprising:
[0011] The spectral graph is convoluted by using an encoding convolution layer of the encoder to obtain a local feature vector;
[0012] The intra-frame dependency of the local feature vector is modeled by using a first transformer layer of the encoder to obtain a global feature vector;
[0013] The inter-frame dependency of the global feature vector is modeled by using a first gated recurrent unit (GRU) layer of the encoder to obtain the encoding vector.
[0014] According to the audio encoding and decoding method provided by the application, the encoding convolution layer includes a first convolution layer, a second convolution layer and a third convolution layer, and the convolution kernel of the first convolution layer, the second convolution layer and the third convolution layer has a size of 1 in the time dimension.
[0015] According to the audio encoding and decoding method provided by the application, the step length of the first convolution layer and the second convolution layer in the frequency dimension is 2.
[0016] According to the audio encoding and decoding method provided by the application, the residual vector quantizer includes N residual quantization layers, and N≥3;
[0017] The encoding vector is input into the residual vector quantizer of the encoding and decoding neural network model to perform multi-layer quantization processing on the encoding vector to obtain a codebook index output by the residual vector quantizer, including:
[0018] The encoding vector is input into a first residual quantization layer to obtain a first residual vector and a first codebook index output by the first residual quantization layer;
[0019] The first residual vector is input into a second residual quantization layer to obtain a second residual vector and a second codebook index output by the second residual quantization layer;
[0020] Similarly, the N-1 residual vector is input into an N residual quantization layer to obtain an N residual vector and an N codebook index output by the N residual quantization layer;
[0021] The codebook index includes the first codebook index, the second codebook index to the N codebook index.
[0022] According to the audio encoding and decoding method provided by the application, the encoding vector is input into a first residual quantization layer to obtain a first residual vector and a first codebook index output by the first residual quantization layer, including:
[0023] The encoding vector is factorized to obtain a reduced dimension encoding vector;
[0024] retrieving a first codebook vector most similar to the reduced dimensionality encoded vector in a first codebook of the first residual quantization layer;
[0025] calculating a residual between the reduced dimensionality encoded vector and the first codebook vector to obtain a first residual vector;
[0026] determining the first codebook index according to a position of the first codebook vector in the first codebook.
[0027] According to the audio coding method provided by the present application, the retrieving a first codebook vector most similar to the reduced dimensionality encoded vector in a first codebook of the first residual quantization layer comprises:
[0028] performing L2 normalization processing on each codebook vector in the first codebook to obtain a normalized codebook vector;
[0029] performing L2 normalization processing on the reduced dimensionality encoded vector to obtain a normalized reduced dimensionality encoded vector;
[0030] performing cosine similarity calculation on the normalized reduced dimensionality encoded vector and the normalized codebook vector, and determining the first codebook vector most similar to the reduced dimensionality encoded vector according to the calculation result.
[0031] According to the audio coding method provided by the present application, the inputting the encoded vector into a residual vector quantizer of the coding and decoding neural network model, performing multi-layer quantization processing on the encoded vector to obtain a codebook index output by the residual vector quantizer comprises:
[0032] obtaining a target number of quantization layers;
[0033] inputting the encoded vector into the residual vector quantizer of the coding and decoding neural network model, performing multi-layer quantization processing on the encoded vector according to the target number of quantization layers to obtain the codebook index output by the residual vector quantizer.
[0034] According to the audio coding method provided by the present application, the inputting the codebook index into a decoder of the coding and decoding neural network model, performing decoding processing on the codebook index to obtain a reconstructed spectrum output by the decoder comprises:
[0035] querying a reconstructed codebook vector corresponding to the codebook index;
[0036] reconstructing an inter-frame dependency relationship according to the reconstructed codebook vector by using a second GRU layer of the decoder to obtain a reconstructed global feature vector;
[0037] reconstructing, by using a second Transformer layer of the decoder, an intra-frame dependency relationship according to the reconstructed global feature vector, to obtain a reconstructed local feature vector;
[0038] performing, by using a decoding convolution layer of the decoder, transposed convolution processing on the reconstructed local feature vector, to obtain the reconstructed spectrogram.
[0039] According to the audio coding method provided by the application, before the codebook index is input into the decoder of the coding neural network model and decoded to obtain the reconstructed spectrogram output by the decoder, the method further comprises:
[0040] determining a target decoder;
[0041] The codebook index is input into the decoder of the coding neural network model and decoded to obtain the reconstructed spectrogram output by the decoder, comprising:
[0042] The codebook index is input into the target decoder and decoded to obtain the reconstructed spectrogram output by the target decoder.
[0043] According to the audio coding method provided by the application, the coding neural network model is trained in the following manner:
[0044] obtaining sample audio data and converting the sample audio data into a sample spectrogram;
[0045] training an initial coding neural network model based on the sample spectrogram and the sample audio data to obtain the pre-trained coding neural network model.
[0046] According to the audio coding method provided by the application, when training the initial coding neural network model, in each training iteration, a random number of codebooks are randomly discarded to regularize the initial residual vector quantizer, wherein the random number is randomly determined according to a preset range.
[0047] According to the audio coding method provided by the application, the initial coding neural network model is trained based on the sample spectrogram and the sample audio data to obtain the pre-trained coding neural network model, comprising:
[0048] training an initial coding neural network model based on the sample spectrogram and the sample audio data to obtain a first coding neural network model;
[0049] A second codec neural network model is constructed based on the first codec neural network model. The second codec neural network model has the same architecture as the first codec neural network model. The number of channels in each layer of the decoder in the second codec neural network model is less than the number of channels in the corresponding layer of the decoder in the first codec neural network model.
[0050] The first codec neural network model is used as the teacher model, and the second codec neural network model is used as the student model. The student model is trained based on the teacher model to obtain the pre-trained codec neural network model.
[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio encoding / decoding method as described above.
[0052] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the audio encoding / decoding method as described in any of the preceding claims.
[0053] The audio encoding / decoding method, electronic device, and program product provided by this invention converts the audio data to be processed into a spectrogram. Leveraging the spectrogram's ability to simultaneously capture the time-varying characteristics and spectral structure of speech signals, and its superior alignment with the human auditory system's perception mechanism, this lays the foundation for more efficient feature extraction and compression. The spectrogram is then input into the encoder of a pre-trained encoding / decoding neural network model. Local acoustic feature extraction, intra-frame dependency modeling, and inter-frame dependency modeling are performed sequentially on the spectrogram to comprehensively capture key perceptual information from the highly redundant spectrogram across multiple dimensions and scales, extracting a high-information-density encoding vector containing core acoustic features. Next, the encoding vector is input into the residual vector quantizer of the encoding / decoding neural network model. Through multi-layer quantization processing, the encoding vector is converted into a codebook index, achieving ultra-high compression rates. Furthermore, the multi-layer quantization structure of the residual vector quantizer possesses inherent scalability, allowing it to flexibly adapt to communication requirements at different bit rates, making it particularly suitable for bandwidth-constrained communication scenarios. Finally, the codebook index is input into the decoder of the encoding / decoding neural network model, and the reconstructed spectrogram is obtained through decoding processing to restore the audio data. Because the decoder, encoder, and residual vector quantizer are trained together, the decoder can reconstruct the spectrogram with high quality, which in turn facilitates the reconstruction of high-quality audio data. In summary, this invention can achieve high-quality audio reconstruction in ultra-low bitrate communication environments, while also achieving efficient compression. Attached Figure Description
[0054] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those of ordinary skill in the art without any creative effort.
[0055] Figure 1 is one of the flowcharts of the audio coding method provided by the present application.
[0056] Figure 2 is one of the flowcharts of the audio coding method provided by the present application.
[0057] Figure 3 is one of the flowcharts of the audio coding method provided by the present application.
[0058] Figure 4 is one of the flowcharts of the audio coding method provided by the present application.
[0059] Figure 5 is the architecture diagram of the coding neural network model provided by the present application.
[0060] Figure 6 is the structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0061] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those of ordinary skill in the art without any creative effort.
[0062] The present application provides an audio coding method, an electronic device and a program product, which will be described below in conjunction with the accompanying drawings. Figures 1-6
[0063] Figure 1 is one of the flowcharts of the audio coding method provided by the present application, as shown in Figure 1 , the audio coding method comprises the following steps.
[0064] In step S110, the audio data to be processed is obtained, and the audio data to be processed is converted into a spectrum graph.
[0065] In this embodiment, the audio data to be processed is first obtained, and then the audio data to be processed is converted into a spectrum graph.
[0066] In real-time communication scenarios, when acquiring the audio data to be processed, a sampling rate of 16 kHz is usually adopted, which achieves a good balance between computational resource consumption and voice quality. Too low a sampling rate will introduce significant distortion, affecting voice intelligibility; while too high a sampling rate, although it can improve audio fidelity, will significantly increase the computational overhead in the encoding and decoding process, which is not conducive to the engineering deployment and real-time requirements of the system.
[0067] In real-time communication scenarios, the audio data to be processed can be converted into a spectrogram through a Short-Time Fourier Transform (STFT). Specifically, the audio data to be processed can be frame-processed according to a preset window length and a preset window shift to obtain audio frames; the preset window length is 15 ms to 25 ms, and the preset window shift is 25% to 50% of the preset window length; the audio frames are sequentially windowed to obtain windowed audio frames; the windowed audio frames are subjected to a fast Fourier transform to obtain frequency domain signals; and the frequency domain signals are converted to obtain a spectrogram corresponding to the audio data to be processed. The specific execution process can be referred to in the following embodiments, which will not be described here.
[0068] The audio data to be processed is converted into a spectrogram because the spectrogram, a time-frequency domain representation, has more advantages in voice processing tasks, and can capture both the time-varying characteristics and the spectral structure of the voice signal. At the same time, this dual representation is not only more suitable for the analysis of non-stationary signals, but also more consistent with the perception mechanism of the human auditory system, thus helping to achieve more efficient feature extraction and compression in the subsequent neural network structure.
[0069] In step S120, the spectrogram is input into an encoder of a pre-trained encoding and decoding neural network model, and the spectrogram is sequentially subjected to local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling to obtain an encoding vector output by the encoder.
[0070] The encoding and decoding neural network model is trained based on sample audio data. The encoding and decoding neural network model mainly includes an encoder, a residual vector quantizer (RVQ) and a decoder. The structure of the encoding and decoding neural network model can refer to Figure 5 .
[0071] First, the spectrogram is input into the encoder, and the encoder sequentially performs local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling on the spectrogram to obtain an encoding vector output by the encoder.
[0072] The encoder comprises an encoding convolutional layer, a Transformer layer and a GRU (Gated Recurrent Unit) layer.
[0073] Further, the convolution kernel of the encoding convolutional layer has a size of 1 in the time dimension, so as to ensure that the model has the characteristics of causality and low delay, thereby eliminating the need to cache the convolutional state of the historical time step in the streaming process, thereby significantly reducing the memory occupation. The Transformer layer is used to model the intra-frame dependency, and the self-attention mechanism of the Transformer can model the global dependency in parallel, effectively breaking through the bottleneck of the traditional RNN (Recurrent Neural Network) structure in sequential processing and remote dependency decay. In order to maintain causality in the entire model structure, the inter-frame dependency is modeled by using a unidirectional GRU layer. By using the above-mentioned dual-path architecture of the Transformer layer + GRU layer to model the intra-frame and inter-frame dependencies of the signal respectively, the audio encoding effect can be improved.
[0074] In step S130, the encoding vector is input into the residual vector quantizer of the codec neural network model, and the encoding vector is subjected to multi-layer quantization processing to obtain the codebook index output by the residual vector quantizer.
[0075] The residual vector quantizer adopts a residual vector quantization structure.
[0076] After obtaining the encoding vector, the encoding vector is input into the residual vector quantizer, and the encoding vector is subjected to multi-layer quantization processing to obtain the codebook index output by the residual vector quantizer.
[0077] In an embodiment, the residual vector quantizer comprises N residual quantization layers, N≥3, the encoding vector is first input into a first residual quantization layer to obtain a first residual vector and a first codebook index output by the first residual quantization layer, then the first residual vector is input into a second residual quantization layer to obtain a second residual vector and a second codebook index output by the second residual quantization layer, and so on. The (N-1)th residual vector is input into an Nth residual quantization layer to obtain an Nth residual vector and an Nth codebook index output by the Nth residual quantization layer. The codebook index comprises the first codebook index, the second codebook index, and the Nth codebook index.
[0078] In another embodiment, the target quantization layer number can be obtained first, wherein the target quantization layer number n≤N, the encoding vector is input into the residual vector quantizer of the codec neural network model, and the encoding vector is subjected to multi-layer quantization processing according to the target quantization layer number n to obtain the codebook index output by the residual vector quantizer. Correspondingly, the codebook index comprises the first codebook index to the n-th codebook index.
[0079] Step S140, input the codebook index into the decoder of the coding and decoding neural network model, decode the codebook index to obtain the reconstructed spectrum graph output by the decoder.
[0080] The structure of the decoder is basically symmetrical to that of the encoder, and includes GRU layers, a Transformer layer and a decoding convolution layer connected in sequence.
[0081] The codebook index is input into the decoder of the coding and decoding neural network model, the corresponding reconstructed codebook vector is queried first, and then the codebook vector is sequentially reconstructed for inter-frame dependency, intra-frame dependency and transposed convolution processing to obtain the reconstructed spectrum graph output by the decoder.
[0082] Further, after obtaining the reconstructed spectrum graph, audio reconstruction can be performed according to the reconstructed spectrum graph to obtain reconstructed audio data. In the audio reconstruction, an inverse short-time Fourier transform (ISTFT) can be used.
[0083] The audio coding method provided by the embodiment of the present application converts the audio data to be processed into a spectrum graph, which can capture the time-varying characteristics and spectral structure of the speech signal at the same time, and is more consistent with the perception mechanism of the human auditory system, which can lay a foundation for more efficient feature extraction and compression. Then, the spectrum graph is input into the encoder of the pre-trained coding and decoding neural network model, and the spectrum graph is sequentially subjected to local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling to comprehensively capture key perceptual information in the highly redundant spectrum graph from multiple dimensions and multiple scales, and extract a high-information-density coding vector containing core acoustic features. Next, the coding vector is input into the residual vector quantizer of the coding and decoding neural network model, and the coding vector is converted into a codebook index through multi-layer quantization processing, which can realize ultra-high compression ratio, and the multi-layer quantization structure of the residual vector quantizer has natural scalability, which can flexibly adapt to different communication requirements of different code rates, especially for bandwidth-limited communication scenarios. Finally, the codebook index is input into the decoder of the coding and decoding neural network model, and the reconstructed spectrum graph is obtained through decoding processing, which is used to restore the audio data. Since the decoder is trained together with the encoder and the residual vector quantizer, the decoder can restore the spectrum graph with high quality, which is conducive to restoring high-quality audio data. In summary, the embodiment of the present application can realize high-quality audio restoration in an ultra-low code rate communication environment, and can also realize efficient compression.
[0084] Based on any of the above embodiments, Figure 2 is a second flowchart of the audio coding method provided by the present application, as Figure 2As shown, step S120 includes step S121, step S122 and step S123.
[0085] In step S121, the spectral graph is subjected to convolution processing by using the encoding convolution layer of the encoder to obtain a local feature vector.
[0086] In this embodiment, the encoder includes an encoding convolution layer, a Transformer layer (denoted as a first Transformer layer) and a GRU layer (denoted as a first GRU layer).
[0087] The spectral graph is input into the encoding convolution layer to perform convolution processing on the spectral graph to obtain a local feature vector output by the encoding convolution layer.
[0088] Further, the encoding convolution layer includes a first convolution layer, a second convolution layer and a third convolution layer, and the convolution kernel of the first convolution layer, the second convolution layer and the third convolution layer has a size of 1 in the time dimension. By fixing the size of the convolution kernel in the time dimension to 1, the model can be ensured to have the characteristics of causality and low delay, so that the convolution state of the historical time step does not need to be cached in the streaming processing process, thereby significantly reducing the memory occupation.
[0089] Further, the number of channels of the first convolution layer, the second convolution layer and the third convolution layer is 32, 64 and 64 respectively, and the size of the convolution kernel is 1x3.
[0090] Further, the step length of the first convolution layer and the second convolution layer in the frequency dimension is 2. The first two convolution layers use convolution with a step length of 2 in the frequency dimension, which can effectively compress the spectral resolution while improving the calculation efficiency and expanding the receptive field of the convolution kernel.
[0091] In an embodiment, the strides of the first convolution layer, the second convolution layer and the third convolution layer are set to 1x2, 1x2 and 1x1 respectively.
[0092] Further, after the convolution processing on the spectral graph, a preset activation function can be used to perform nonlinear activation on the convolution result to obtain the local feature vector. The preset activation function can be a PReLU (Parametric Rectified Linear Unit) activation function, a ReLU (Rectified Linear Unit) activation function. It should be noted that, compared with the ReLU activation function, the use of the PReLU activation function can improve the performance and training stability of the model. In addition, in the embodiment of the present application, no normalization operation is introduced after the convolution operation.
[0093] Step S122, modeling the intra-frame dependency of the local feature vector by using a first transformer layer of the encoder to obtain a global feature vector.
[0094] The local feature vector is input into the first transformer layer to model the intra-frame dependency of the local feature vector, and a global feature vector output by the first transformer layer is obtained.
[0095] The self-attention mechanism of the transformer can model the global dependency in parallel, which can effectively break through the bottleneck of the traditional RNN (Recurrent Neural Network) structure in sequential processing and remote dependency decay.
[0096] Step S123, modeling the inter-frame dependency of the global feature vector by using a first gated recurrent unit (GRU) layer of the encoder to obtain the encoding vector.
[0097] The global feature vector is input into the first GRU layer to model the inter-frame dependency of the global feature vector, and an encoding vector output by the first GRU layer is obtained.
[0098] By using the unidirectional GRU layer to model the inter-frame dependency, the time sequence continuity of the audio signal can be captured, and due to the causal processing characteristic, the entire encoder does not need to wait for future information, can run in a streaming mode, and the processing delay is greatly reduced.
[0099] The audio encoding and decoding method provided in the embodiment of the application models the intra-frame and inter-frame dependency of the signal by using the first transformer layer and the first GRU layer, respectively, and the above-mentioned dual-path architecture can improve the audio encoding effect.
[0100] Based on any one of the above embodiments, the residual vector quantizer includes N residual quantization layers, and N is greater than or equal to 3, Figure 3 is a third flowchart of the audio encoding and decoding method provided in the application, as shown in Figure 3 The step S130 includes a step S131, a step S132 and a step S133.
[0101] Step S131, inputting the encoding vector into a first residual quantization layer to obtain a first residual vector and a first codebook index output by the first residual quantization layer.
[0102] In this embodiment, the residual vector quantizer adopts a residual vector quantization structure, including N residual quantization layers, where N≥3, and further, N≤9. Each residual quantization layer corresponds to a codebook. In this embodiment, all residual quantization layers are used completely to maximize the performance of the model.
[0103] The encoding vector is input into the first residual quantization layer, and in the first codebook of the first residual quantization layer, the codebook vector most similar to the encoding vector is searched, which is recorded as the first codebook vector.
[0104] The residual between the encoding vector and the first codebook vector is calculated to obtain the first residual vector. First residual vector = encoding vector - first codebook vector.
[0105] At the same time, the first codebook index is determined according to the position of the first codebook vector in the first codebook.
[0106] Step S132, input the first residual vector into the second residual quantization layer to obtain the second residual vector and the second codebook index output by the second residual quantization layer.
[0107] Then, the first residual vector is input into the second residual quantization layer, and in the second codebook of the second residual quantization layer, the codebook vector most similar to the first residual vector is searched, which is recorded as the second codebook vector.
[0108] The residual between the first residual vector and the second codebook vector is calculated to obtain the second residual vector. Second residual vector = first residual vector - second codebook vector.
[0109] At the same time, the second codebook index is determined according to the position of the second codebook vector in the second codebook.
[0110] Step S133, by analogy, input the N-1 residual vector into the N residual quantization layer to obtain the N residual vector and the N codebook index output by the N residual quantization layer.
[0111] The codebook index includes the first codebook index, the second codebook index to the N codebook index.
[0112] By analogy, repeat the above process until the N residual quantization layer, input the N-1 residual vector into the N residual quantization layer, and in the N codebook of the N residual quantization layer, the codebook vector most similar to the N-1 residual vector is searched, which is recorded as the N codebook vector.
[0113] The residual between the N-1 residual vector and the N codebook vector is calculated to obtain the N residual vector. N residual vector = N-1 residual vector - N codebook vector.
[0114] Meanwhile, according to the position of the Nth codebook vector in the Nth codebook, the Nth codebook index is determined.
[0115] The codebook indexes include a first codebook index, a second codebook index,..., an (N-1)th codebook index, and an Nth codebook index.
[0116] The audio coding method provided by the embodiment of the application introduces a residual vector quantizer including multiple residual quantization layers in the audio coding process, and through the residual vector quantizer, the complex details of the audio signal can be better captured in the multiple quantization processes, and the quantization error can be reduced, so that the high sound quality can be maintained at a low bit rate.
[0117] Based on any of the above embodiments, the step S131 includes a step S1311, a step S1312, a step S1313, and a step S1314.
[0118] In the step S1311, the encoding vector is factorized to obtain a reduced dimension encoding vector.
[0119] In the embodiment, the factorization technology is a vector quantization technology, which can reduce the dimension of a high-dimensional vector to obtain a low-dimensional vector.
[0120] The encoding vector is factorized to obtain a reduced dimension encoding vector, which is denoted as a reduced dimension encoding vector. For example, the dimension of the encoding vector is 124, which is reduced to 8 through factorization.
[0121] In the step S1312, a first codebook vector most similar to the reduced dimension encoding vector is searched in the first codebook of the first residual quantization layer.
[0122] Then, a first codebook vector most similar to the reduced dimension encoding vector is searched in the first codebook of the first residual quantization layer.
[0123] It should be understood that the dimension of the first codebook is set to be the same as that of the reduced dimension encoding vector.
[0124] In an embodiment, for any reduced dimension encoding vector, the Euclidean distance between the reduced dimension encoding vector and each codebook vector in the first codebook can be directly calculated, and the codebook vector corresponding to the minimum Euclidean distance is determined as the first codebook vector.
[0125] In another embodiment, the L2 normalization processing is first performed on the reduced dimension encoding vector and each codebook vector in the first codebook, respectively, and then for any normalized reduced dimension encoding vector, the cosine similarity between the normalized reduced dimension encoding vector and each normalized codebook vector in the first codebook is calculated, and the codebook vector corresponding to the minimum cosine similarity is determined as the first codebook vector.
[0126] Step S1313, a residual error vector is obtained by calculating the residual error between the dimension-reduced encoding vector and the first codebook vector.
[0127] Step S1314, the first codebook index is determined according to the position of the first codebook vector in the first codebook.
[0128] Then, the residual error between the dimension-reduced encoding vector and the first codebook vector is calculated to obtain the first residual error vector, and the first codebook index is determined according to the position of the first codebook vector in the first codebook.
[0129] It should be understood that the factorization method is also used when other codebook vectors (such as the second codebook vector, …, the Nth codebook vector) are determined subsequently.
[0130] The audio coding method provided by the embodiment of the application can limit the search process of the codebook vector in the low-dimensional space by using the above factorization strategy, so that the search of the codebook vector and the codebook index can be implemented efficiently, and the corresponding high-dimensional codebook vector can be searched based on the codebook index, so that better reconstruction quality can be obtained.
[0131] Based on any of the above embodiments, the step S1312 includes steps S13121, S13122 and S13123.
[0132] Step S13121, each codebook vector in the first codebook is subjected to L2 normalization processing to obtain a normalized codebook vector.
[0133] Step S13122, the dimension-reduced encoding vector is subjected to L2 normalization processing to obtain a normalized dimension-reduced encoding vector.
[0134] Step S13123, the normalized dimension-reduced encoding vector and the normalized codebook vector are subjected to cosine similarity calculation, and the first codebook vector most similar to the dimension-reduced encoding vector is determined according to the calculation result.
[0135] It should be noted that the execution order of step S13121 and step S13122 is not limited.
[0136] In this embodiment, the L2 normalization processing is to adjust the L2 norm (i.e. the length or modulus of the vector) of the vector to 1, so that the Euclidean distance between the vectors is only related to the included angle (i.e. the cosine similarity) between them.
[0137] The L2 normalization processing is performed on each codebook vector in the first codebook to obtain a normalized codebook vector, and the L2 normalization processing is performed on the dimension-reduced encoding vector to obtain a normalized dimension-reduced encoding vector. Then, for any normalized dimension-reduced encoding vector, the cosine similarity between the normalized dimension-reduced encoding vector and each normalized codebook vector is calculated, and the codebook vector corresponding to the minimum cosine similarity is determined as the first codebook vector most similar to the dimension-reduced encoding vector.
[0138] It should be understood that, in subsequent determination of other codebook vectors (such as a second codebook vector, …, an Nth codebook vector), the L2 normalization processing is further combined on the basis of the factorization processing, and the specific processing process is similar to the above process.
[0139] The audio encoding and decoding method provided by the embodiment of the application can convert the distance metric from the Euclidean distance to the cosine similarity when searching for the most similar codebook vector through the L2 normalization processing, can simplify the similarity calculation, and thus helps to improve the utilization efficiency of the codebook. In addition, it should be noted that the L2 normalization processing can help to improve the stability of the model training in the model training process.
[0140] Based on any of the above embodiments, the step S130 further includes a step S134 and a step S135.
[0141] In the step S134, the target quantization layer number is obtained.
[0142] In an embodiment, the target quantization layer number is pre-set.
[0143] In another embodiment, the target quantization layer number can be determined according to the network bandwidth. The higher the network bandwidth is, the greater the target quantization layer number is. For example, when the network bandwidth is high, the target quantization layer number is determined as N, that is, all the N residual network layers are used to transmit 9 kbps data to obtain the highest quality reconstruction; when the network bandwidth is poor, the target quantization layer number is determined as 1 / 3N~1 / 2N (upward integer), and 3 kbps data is transmitted to obtain an acceptable basic quality reconstruction; and when the network bandwidth is medium, the target quantization layer number is determined as 1 / 2N~2 / 3N (upward integer), and 5 kbps data is transmitted to obtain a quality and bit rate between the above two.
[0144] In the step S135, the encoding vector is input into the residual vector quantizer of the encoding and decoding neural network model, the encoding vector is subjected to multi-layer quantization processing according to the target quantization layer number, and the codebook index output by the residual vector quantizer is obtained.
[0145] Assuming that the target quantization layer number is n, n≤N, the encoding vector is first input to the first residual quantization layer to obtain the first residual vector output by the first residual quantization layer and the first codebook index, then the first residual vector is input to the second residual quantization layer to obtain the second residual vector output by the second residual quantization layer and the second codebook index, and so on. The (n-1)th residual vector is input to the nth residual quantization layer to obtain the nth residual vector output by the nth residual quantization layer and the nth codebook index. The codebook index includes the first codebook index, the second codebook index, and the nth codebook index.
[0146] Further, when input to the residual quantization layer for processing, the above factorization strategy and L2 normalization processing can be adopted, and the specific processing process can refer to the above embodiments, which will not be repeated here.
[0147] The audio encoding and decoding method provided by the embodiment of the application adjusts the number of residual quantization layers, thereby realizing flexible adaptation to different network bandwidths and meeting different encoding requirements.
[0148] Based on any of the above embodiments, Figure 4 is a fourth flowchart of the audio encoding and decoding method provided by the application, as shown in Figure 4 The step S140 further includes steps S141, S142, S143, and S144.
[0149] In step S141, the codebook index corresponding reconstruction codebook vector is queried.
[0150] The structure of the decoder is basically symmetrical to that of the encoder, including a GRU layer (denoted as a second GRU layer), a Transformer layer (denoted as a second Transformer layer), and a decoding convolutional layer connected in sequence. The above-mentioned dual-path architecture of the Transformer layer+GRU layer can respectively reconstruct the intra-frame and inter-frame dependencies, thereby improving the audio decoding effect.
[0151] In an embodiment, in the process of performing multi-layer quantization processing on the encoding vector to obtain the codebook index, if the encoding vector is not factorized, after the codebook index is obtained, the codebook vector corresponding to the codebook index is first queried in the codebook according to the codebook index, denoted as a reconstruction codebook vector.
[0152] In another implementation, in the process of multi-layer quantization of the encoding vector to obtain the codebook index, if the encoding vector is factorized, after the codebook index is obtained, the codebook vector corresponding to the codebook index is first queried in the codebook according to the codebook index, and then the dimension is raised to the dimension of the original encoding vector to obtain the reconstructed codebook vector.
[0153] In step S142, the inter-frame dependency is reconstructed according to the reconstructed codebook vector by using the second GRU layer of the decoder to obtain a reconstructed global feature vector.
[0154] The codebook vector is input into the second GRU layer of the decoder, the reconstructed codebook vector is taken as the initial hidden state, and the global feature vector is generated in an autoregressive manner, denoted as a reconstructed global feature vector, to realize the reconstruction of the inter-frame dependency, and thus the reconstructed global feature vector output by the second GRU layer is obtained.
[0155] In step S143, the intra-frame dependency is reconstructed according to the reconstructed global feature vector by using the second Transformer layer of the decoder to obtain a reconstructed local feature vector.
[0156] The reconstructed global feature vector is input into the second Transformer layer of the decoder to perform the intra-frame dependency reconstruction on the reconstructed global feature vector, and thus the reconstructed local feature vector output by the second Transformer layer is obtained.
[0157] In step S144, the decoding convolution layer of the decoder is used to perform transposed convolution processing on the reconstructed local feature vector to obtain the reconstructed spectrogram.
[0158] The reconstructed local feature vector is input into the decoding convolution layer of the decoder to perform transposed convolution processing on the reconstructed local feature vector, and thus the reconstructed spectrogram output by the decoding convolution layer is obtained.
[0159] Further, the decoding convolution layer includes a first transposed convolution layer, a second transposed convolution layer and a third transposed convolution layer, and the convolution kernel of the first transposed convolution layer, the second transposed convolution layer and the third transposed convolution layer has a size of 1 in the time dimension.
[0160] Further, the number of channels of the first transposed convolution layer, the second transposed convolution layer and the third transposed convolution layer is 64, 64 and 32 in sequence, and the convolution kernel size is 1x3.
[0161] Further, the strides of the first transposed convolution layer, the second transposed convolution layer and the third transposed convolution layer are set to 1x1, 1x2 and 1x2 respectively.
[0162] The audio coding method provided by the embodiment of the application firstly utilizes the second GRU layer to reconstruct the interframe dependency relationship according to the codebook vector, thereby preferentially restoring the macro timing structure of the audio signal, and then utilizes the second Transformer layer to reconstruct the intraframe dependency relationship of each frame, since the self-attention mechanism of the Transformer can capture long-distance dependency between features, the complex key spectral details inside each frame can be accurately restored, and then the reconstructed spectral graph is obtained by utilizing the decoding convolution layer, so as to be used for reconstructing the audio. By preferentially reconstructing the interframe timing relationship and then finely restoring the intraframe spectral structure, the balance problem between time fluency and detail fidelity in audio decoding can be solved, and the quality of the reconstructed audio after decoding is high.
[0163] Further, it needs to be noted that, in an embodiment, the number of channels of each layer of the decoder and the number of channels of the corresponding layer of the encoder are designed to be symmetrical, that is, the number of channels of each layer of the decoder and the number of channels of the corresponding layer of the encoder are consistent. That is, the number of channels of the second GRU layer is the same as the number of channels of the first GRU layer, the number of channels of the second Transformer layer is the same as the number of channels of the first Transformer layer, and the number of channels of the decoding convolution layer is the same as the number of channels of the encoding convolution layer. Specifically, the number of channels of the first transposed convolution layer is the same as the number of channels of the third convolution layer, the number of channels of the second transposed convolution layer is the same as the number of channels of the second convolution layer, and the number of channels of the third transposed convolution layer is the same as the number of channels of the first convolution layer.
[0164] In another embodiment, the number of channels of each layer of the decoder and the number of channels of the corresponding layer of the encoder are designed to be asymmetrical. Specifically, the number of channels of the second GRU layer is less than the number of channels of the first GRU layer, the number of channels of the second Transformer layer is less than the number of channels of the first Transformer layer, and the number of channels of the decoding convolution layer is less than the number of channels of the encoding convolution layer.
[0165] This is because in actual applications, the audio coding process often presents an asymmetrical characteristic. In particular, in the actual application of a conference scenario, the audio coding process presents obvious asymmetrical characteristics: usually only the local audio needs to be encoded once, but multiple remote audios need to be decoded in real time at the same time. This characteristic causes the computational load of the decoder to be significantly higher than that of the encoder, so that the computational efficiency of the decoder becomes a key factor restricting the practical application of the coding and decoding neural network model. Therefore, when training the coding and decoding neural network model, on the basis of the symmetrical coding and decoding architecture obtained by training, the asymmetrical coding and decoding architecture is further trained, the number of channels of each layer of the decoder is reduced, the computational complexity and the parameter amount are reduced, and the lightweight coding and decoding neural network model obtained by training can be applied to terminal devices with limited computing resources. When the audio data to be processed is coded and decoded by using the lightweight coding and decoding neural network model, the consumption of computing resources can be reduced, so that the lightweight coding and decoding neural network model can stably operate in a complex communication scenario.
[0166] Specifically, the channel numbers of the second GRU layer, the second Transformer layer and the decoding convolution layer can be reduced to 20%-80% of the channel numbers of the first GRU layer, the first Transformer layer and the encoding convolution layer, respectively.
[0167] Further, the encoding and decoding neural network model can include multiple decoders, and different decoders have the same structure, that is, each includes a second GRU layer, a second Transformer layer and a decoding convolution layer, and the only difference is that the channel numbers of the layers are different.
[0168] Therefore, before the step S130, there is also a step of determining a target decoder, and at this time, the step S130 includes inputting the codebook index into the target decoder and decoding the codebook index to obtain a reconstructed spectrum graph output by the target decoder.
[0169] When determining the target decoder, the determination can be made according to the computing capability of the terminal device, the application scenario, the network bandwidth and the like, and the specific rules are not limited here. After determining the target decoder, the codebook index is input into the target decoder, and the codebook index is decoded to obtain a reconstructed spectrum graph output by the target decoder. The specific processing process can refer to the above embodiments, and will not be repeated here.
[0170] For example, the decoder can include three decoders, the channel numbers of the layers of the first decoder are the same as the channel numbers of the corresponding layers of the encoder, the channel numbers of the layers of the second decoder are reduced to 50% of the channel numbers of the layers of the first decoder, and the channel numbers of the layers of the third decoder are reduced to 25% of the channel numbers of the layers of the first decoder. On a high-end terminal device, a high-quality decoder with full channel numbers (i.e., the first decoder described above) can be enabled, and in a resource-constrained terminal device, a lightweight version with channel numbers compressed to 50% (i.e., the second decoder described above) is switched, and in an extremely low-power consumption scenario, an extremely simple decoder with channel numbers compressed to 25% (i.e., the third decoder described above) can be switched.
[0171] The audio encoding and decoding method provided by the embodiment of the application can switch different decoders according to the computing capability, power consumption limitation or application scenario requirement of different terminal devices, can significantly reduce the computing overhead while maintaining good audio reconstruction quality.
[0172] Based on any of the above embodiments, the step of converting the to-be-processed audio data into a spectrum graph includes steps S111, S112, S113 and S114.
[0173] In step S111, the audio data to be processed is frame-processed according to a preset window length and a preset window shift to obtain audio frames; the preset window length is 15 ms to 25 ms, and the preset window shift is 25% to 50% of the preset window length.
[0174] In a real-time communication scenario, the audio data to be processed can be converted into a spectrum graph through STFT.
[0175] First, the audio data to be processed is frame-processed according to a preset window length and a preset window shift to obtain audio frames. The preset window length ranges from 15 ms to 25 ms. This parameter range can achieve a good trade-off between time resolution and frequency resolution.
[0176] For example, when the preset window length is 20 ms, and the audio in the real-time communication scenario usually adopts a sampling rate of 16 kHz, the window length in sampling points is 320 sampling points according to the above sampling rate, that is, each audio frame contains 320 sampling points. When the preset window shift is 50% of the preset window length, the corresponding window shift is 160 sampling points. That is, the first audio frame with a length of 320 sampling points is intercepted from the starting point of the audio data to be processed; then, the second audio frame with a length of 320 sampling points is intercepted from the 160th sampling point; and so on, until the end of the audio data to be processed. In this way, multiple audio frames that overlap with each other are obtained.
[0177] In step S112, each audio frame is sequentially windowed to obtain a windowed audio frame.
[0178] In order to reduce the spectral leakage effect caused by subsequent Fourier transform, a window function needs to be applied to each audio frame for windowing to obtain a windowed audio frame. Specifically, a Hanning window can be used.
[0179] In step S113, the windowed audio frame is subjected to fast Fourier transform to obtain a frequency domain signal.
[0180] Each windowed audio frame is subjected to fast Fourier transform (FFT) to convert it from the time domain to the frequency domain to obtain a frequency domain signal.
[0181] Further, in order to improve the calculation efficiency of FFT and obtain a smoother spectrum, an integer power of 2 greater than or equal to the window length is usually selected as the number of points of FFT.
[0182] In step S114, the frequency domain signal is converted to obtain a spectrum graph corresponding to the audio data to be processed.
[0183] According to the frequency domain signal, a spectrum graph corresponding to the to-be-processed audio data is converted. The X-axis of the spectrum graph represents time, and the Y-axis represents frequency.
[0184] The audio coding method provided in the embodiment of the present application can convert the to-be-processed audio data into a spectrum graph that can intuitively reflect the change of the frequency component of the to-be-processed audio data over time. The spectrum graph has clear time resolution and frequency resolution, and such a time-frequency domain representation form has more advantages in speech processing tasks, and can capture the time-varying characteristics and spectral structure of the speech signal at the same time. At the same time, such a dual representation is more suitable for the analysis of non-stationary signals, and is more in line with the perception mechanism of the human auditory system, thereby helping to realize more efficient feature extraction and compression in the subsequent neural network structure.
[0185] Based on any of the above embodiments, the training process of the coding and decoding neural network model comprises steps S10 and S20.
[0186] In step S10, sample audio data is obtained, and the sample audio data is converted into a sample spectrum graph.
[0187] In step S20, based on the sample spectrum graph and the sample audio data, an initial coding and decoding neural network model is trained to obtain the pre-trained coding and decoding neural network model.
[0188] In this embodiment, an initial coding and decoding neural network model is first constructed, which includes an encoder (denoted as an initial encoder), a residual vector quantizer (denoted as an initial residual vector quantizer), and a decoder (denoted as an initial decoder). The structures of the initial encoder and the initial decoder are basically the same, the initial encoder includes an encoding convolutional layer (denoted as an initial encoding convolutional layer), a Transformer layer (denoted as a first initial Transformer layer), and a GRU layer (denoted as a first initial GRU layer) connected in sequence, and the initial encoding convolutional layer includes three convolutional layers, which are denoted as a first initial convolutional layer, a second initial convolutional layer, and a third initial convolutional layer. The initial decoder includes a GRU layer (denoted as a second initial GRU layer), a Transformer layer (denoted as a second initial Transformer layer), and a decoding convolutional layer (denoted as an initial decoding convolutional layer) connected in sequence, and the initial decoding convolutional layer includes three transposed convolutional layers, which are denoted as a first initial transposed convolutional layer, a second initial transposed convolutional layer, and a third initial transposed convolutional layer.
[0189] The number of channels of each layer of the initial encoder and the number of channels of the corresponding layer of the initial decoder can be set to be the same, specifically, the number of channels of the first initial Transformer layer and the second initial Transformer layer is the same, the number of channels of the first initial GRU layer and the second initial GRU layer is the same, the number of channels of the first initial convolution layer and the third initial transposed convolution layer is the same, the number of channels of the second initial convolution layer and the second initial transposed convolution layer is the same, and the number of channels of the third initial convolution layer and the first initial transposed convolution layer is the same.
[0190] The sample audio data can include, but is not limited to, single-person voice, single-person voice with noise, single-person voice with reverberation, single-person mixed audio (i.e., single-person voice with noise and reverberation), and multi-person mixed audio.
[0191] After obtaining the sample audio data, the sample audio data is converted into a sample spectrogram, and then the initial codec neural network model is trained based on the sample spectrogram and the sample audio data to obtain the pre-trained codec neural network model.
[0192] In an embodiment, during the training process, the sample spectrogram can be input into the initial codec neural network model to obtain a sample reconstructed spectrogram, and then sample reconstructed audio data is obtained based on the sample reconstructed spectrogram, the spectral reconstruction loss is calculated based on the sample spectrogram and the sample reconstructed spectrogram, the time-domain waveform loss is calculated based on the sample audio data and the sample reconstructed audio data, the total loss is calculated according to the spectral reconstruction loss and the time-domain waveform loss, and the parameters of the initial codec neural network model are updated through back propagation according to the total loss to train the initial codec neural network model until the initial codec neural network model converges, and the final codec neural network model is obtained.
[0193] In another embodiment, in addition to the time-domain waveform loss and the spectral reconstruction loss, the quantization error loss and / or the generative adversarial loss can be further combined when calculating the total loss to train the initial codec neural network model until the initial codec neural network model converges, and the final codec neural network model is obtained.
[0194] The audio codec method provided by the embodiment of the present application can be used for audio codec processing by using the above trainable codec neural network model.
[0195] Based on any of the above embodiments, when training the initial codec neural network model, a random number of codebooks are randomly discarded in each training iteration to regularize the initial residual vector quantizer, wherein the random number is randomly determined according to a preset range.
[0196] In the embodiment, to prevent overfitting or collaborative adaptation between residual quantization layers of the residual vector quantizer, a random dropping strategy can be used to regularize the initial residual vector quantizer, which can force each residual quantization layer to learn more general and robust features, and finally improve the generalization ability and robustness of the entire coding and decoding neural network model on unknown audio data.
[0197] Specifically, in each training iteration, a random number of codebooks are randomly dropped, where the random number is randomly determined according to a preset range. When dropping the codebooks, they are also randomly selected. Dropping the codebooks means dropping / skipping the corresponding residual quantization layer.
[0198] When N = 9, there are 9 codebooks, and the preset range can be [0, 6]. In each training iteration, [0, 6] codebooks are randomly dropped, for example, 1 or 2 codebooks are randomly dropped. For example, when 1 codebook is randomly dropped, any one of the first to ninth residual quantization layers can be randomly dropped. In this way, not only the robustness of the residual vector quantizer is enhanced, but also the coding and decoding neural network model obtained by training can support variable bit rate transmission in the range of 3-9 kbps.
[0199] The audio coding method provided by the embodiment of the application can enhance the robustness of the residual vector quantizer by using the random dropping strategy to regularize the initial residual vector quantizer, and also enables the coding and decoding neural network model obtained by training to support variable bit rate transmission in different ranges. Compared with the traditional baseline residual vector quantization method, the embodiment of the application performs better in terms of code rate efficiency under the same configuration.
[0200] Based on any of the above embodiments, the loss function used in the training process of the coding and decoding neural network model is determined based on multiple kinds of spectral reconstruction loss, time-domain waveform loss, quantization error loss, and generative adversarial loss; wherein the spectral reconstruction loss is determined based on a sample spectrum graph and a sample reconstructed spectrum graph; and / or the time-domain waveform loss is determined based on sample audio data and sample reconstructed audio data; and / or the quantization error loss is determined based on codebook loss and commitment loss; and / or the generative adversarial loss is determined based on adversarial loss and feature matching loss.
[0201] The loss function used in the training process of the coding and decoding neural network model is determined based on multiple kinds of spectral reconstruction loss, time-domain waveform loss, quantization error loss, and generative adversarial loss.
[0202] The spectral reconstruction loss is determined based on a sample spectrum graph and a sample reconstructed spectrum graph.
[0203] The spectrum reconstruction loss is composed of three parts: Mel-spectrum magnitude loss, log-amplitude spectrum loss, and complex spectrum loss. Spectrum reconstruction loss = Mel-spectrum magnitude loss + log-amplitude spectrum loss + complex spectrum loss.
[0204] wherein Mel-spectrum magnitude loss is determined based on the Mel-spectrum magnitude corresponding to the sample spectrum and the reconstructed Mel-spectrum magnitude corresponding to the sample reconstructed spectrum, based on the Mel-frequency scale, emphasizing the spectral region related to human ear perception. Its calculation formula is as follows:
[0205] ;
[0206] wherein, Mel-spectrum magnitude loss is represented by T, which represents the total number of time frames, M t Mel-spectrum magnitude is represented by reconstructed Mel-spectrum magnitude is represented by || ||1, which represents the L1 norm calculation.
[0207] Log-amplitude spectrum loss, which is determined based on the sample spectrum and the sample reconstructed spectrum, applies a logarithmic transformation on the spectral amplitude, enhancing the sensitivity to weak signals. Its calculation formula is as follows:
[0208] ;
[0209] wherein, Log-amplitude spectrum loss is represented by sample spectrum of time frame t is represented by sample reconstructed spectrum of time frame t is represented by amplitude operation is represented by pre-set constant is represented by, which can be used to maintain numerical stability.
[0210] Complex spectrum loss, which is determined based on the sample spectrum and the sample reconstructed spectrum, directly measures the reconstruction error in the complex spectrum space to capture phase and amplitude information. Its calculation formula is as follows:
[0211] ;
[0212] wherein, Complex spectrum loss is represented by operator for extracting real part is represented by operator for extracting imaginary part is represented by
[0213] Further, time-domain waveform loss is determined based on sample audio data and sample reconstructed audio data.
[0214] The time-domain waveform loss directly measures the reconstruction error in the waveform domain to ensure the accuracy of the restored audio signal in the time dimension. The time-domain waveform loss is calculated using the L1 norm, which can effectively capture multiple key features of waveform fidelity. The calculation formula is as follows:
[0215]
[0216]
[0217] Further, the quantization error loss is determined based on the codebook loss and the commitment loss. The calculation formula is as follows:
[0218]
[0219]
[0220] The quantization error loss contains two components: codebook loss and commitment loss, both of which use the stop gradient operation but have different application targets. The goal of the codebook loss is to keep the quantized representation as close as possible to the original feature space, which is achieved by minimizing the distance between the embedding vector and its corresponding quantized vector. This process updates the embedding vector by applying a gradient, while keeping the quantized vector unchanged, to stabilize the training process. The calculation formula is as follows:
[0221]
[0222]
[0223] The commitment loss is used to maintain the diversity of the codebook and prevent codebook collapse (i.e., all inputs are mapped to the same code word). The calculation formula is as follows:
[0224] .
[0225] Further, the generative adversarial loss is determined based on the generative adversarial loss and the feature matching loss.
[0226] The generative adversarial loss is generated by multiple resolution spectrum, which is calculated by multiple spectrum resolution levels to capture local and global spectrum features at the same time. The multi-resolution strategy calculates the spectrum by using different window sizes in the short-time Fourier transform, thereby generating representations with different time-frequency resolutions. The window and step size are usually selected as powers of 2 to cover the feature extraction range from fine to coarse. The generative adversarial loss makes the reconstructed audio closer to the real audio distribution, significantly improving the overall naturalness and listening quality of the generated audio.
[0227] The calculation formula of the generative adversarial loss is as follows:
[0228] ;
[0229] wherein, denotes the generative adversarial loss, denotes the adversarial loss, denotes the feature matching loss, denotes the third preset weight coefficient.
[0230] The calculation formula of the adversarial loss is as follows:
[0231] ;
[0232] wherein, R denotes the resolution layer number, and E denotes the expected value, is the discriminator at resolution r, is the target spectrum graph at resolution r, is the reconstructed spectrum graph generated at resolution r.
[0233] The calculation formula of the feature matching loss is as follows:
[0234] ;
[0235] wherein, is the number of layers in the discriminator denotes the feature mapping of the i-th layer in the discriminator
[0236] The audio coding method provided by the embodiment of the application introduces multiple loss functions in the training process of the coding neural network model: spectrum reconstruction loss, time domain waveform loss, quantization error loss, and generative adversarial loss. Through the above loss functions, the fidelity and perceptual quality of audio reconstruction can be improved at the same time.
[0237] Based on any of the above embodiments, the step S20 comprises a step S21, a step S22 and a step S23.
[0238] In the step S21, the initial codec neural network model is trained based on the sample spectrogram and the sample audio data to obtain a first codec neural network model.
[0239] In practical applications, the audio codec process often presents an asymmetric characteristic. In particular, in the practical application of a conference scenario, the audio codec process presents a significant asymmetric characteristic: usually only the local audio needs to be encoded once, but multiple remote audios need to be decoded simultaneously in real time. This characteristic causes the computational load of the decoder to be significantly higher than that of the encoder, so that the computational efficiency of the decoder becomes a key factor restricting the practical application of the neural network codec. Therefore, in the present embodiment, when training the codec neural network model, on the basis of the symmetric codec architecture obtained by training, the asymmetric codec architecture is further trained to reduce the number of channels of each layer of the decoder, so as to reduce the computational complexity and the parameter quantity.
[0240] The initial codec neural network model is trained based on the sample spectrogram and the sample audio data to obtain a first codec neural network model. Specifically, the sample spectrogram can be input into the initial codec neural network model to obtain a sample reconstructed spectrogram, and then the sample reconstructed audio data is obtained based on the sample reconstructed spectrogram. Then, based on the sample spectrogram, the sample reconstructed spectrogram, the sample audio data, and the sample reconstructed audio data, a total loss is determined, and the parameters of the initial codec neural network model are updated through back propagation according to the total loss to train the initial codec neural network model, until the initial codec neural network model converges, and a final codec neural network model is obtained. The total loss is determined based on multiple of the spectral reconstruction loss, the time-domain waveform loss, the quantization error loss, and the generative adversarial loss.
[0241] In the step S22, a second codec neural network model is constructed according to the first codec neural network model, the architecture of the second codec neural network model is the same as that of the first codec neural network model, and the number of channels of each layer of the decoder in the second codec neural network model is less than that of the corresponding layer of the decoder in the first codec neural network model.
[0242] According to the first coding-decoding neural network model, a second coding-decoding neural network model is constructed, the second coding-decoding neural network model has the same architecture as the first coding-decoding neural network model, and the only difference is that the number of channels of each layer of the decoder in the second coding-decoding neural network model is less than the number of channels of the corresponding layer of the decoder in the first coding-decoding neural network model. Specifically, the number of channels of each layer of the decoder in the second coding-decoding neural network model can be reduced to 20%-80% of the number of channels of the corresponding layer of the decoder in the first coding-decoding neural network model.
[0243] In step S23, the first coding-decoding neural network model is used as a teacher model, the second coding-decoding neural network model is used as a student model, the student model is trained according to the teacher model, and the pre-trained coding-decoding neural network model is obtained.
[0244] The first coding-decoding neural network model is used as a teacher model, the second coding-decoding neural network model is used as a student model, the student model is trained according to the teacher model by a knowledge distillation technology, and a pre-trained coding-decoding neural network model is obtained.
[0245] In the distillation process, the problem of mismatching of the number of channels between the teacher model and the student model can be solved by introducing an adaptive layer with a 1x1 convolution kernel. After channel adaptation, a distillation loss can be added to the original loss function. That is, the loss function spectrum includes multiple types of reconstruction loss, time domain waveform loss, quantization error loss and generative adversarial loss, and also includes distillation loss. The distillation loss can use an L2 loss function. Through the L2 loss function, the features of each layer of the decoder can be distilled layer by layer. The calculation formula of the distillation loss is as follows:
[0246] ;
[0247] wherein, represents the distillation loss, I represents the total number of feature layers, i represents the feature layer index, represents the feature output of the i-th layer of the teacher model, represents the feature output of the i-th layer of the student model.
[0248] It should be noted that since the structure of the encoder and the residual vector quantizer of the teacher model and the student model is completely the same, the parameters of the encoder and the residual vector quantizer of the student model are directly inherited from the teacher model and remain fixed, and only the decoder part is trained and optimized. Through the above-mentioned manner, the following three important advantages can be brought: (1) coding-decoding compatibility, i.e., the decoders of the teacher and student models can process the same encoded stream; (2) flexible deployment, i.e., the full-amount decoder (high quality) or the light-weight decoder (high efficiency) can be selected according to the computing resources; (3) seamless switching, i.e., zero-delay switching between different decoders can be realized.
[0249] In an embodiment, the trained student model is directly used as the pre-trained codec neural network model. In this case, the encoder of the codec neural network model is lightweight and can be applied to a terminal with limited computing resources while ensuring good audio quality.
[0250] In another embodiment, a codec neural network model including multiple decoder branches is constructed as the pre-trained codec neural network model according to the teacher model and the trained student model. That is, the codec neural network model includes an encoder, a residual vector quantizer, and multiple decoders, wherein the encoder and the residual vector quantizer can be derived from the teacher model, and the multiple decoders are derived from the decoder of the trained student model in addition to being derived from the teacher model. The decoders derived from the trained student model are actually asymmetric decoders, can share the same encoder structure, and are all compatible with the teacher model, thereby realizing multi-level deployment and dynamic adaptation. In this way, not only the flexibility and adaptability of the system are improved, but also a new idea is provided for a unified audio coding standard for a heterogeneous device environment in the future. In this case, seamless switching to a lightweight decoder can be realized when the computing resources are limited, while ensuring good audio quality.
[0251] The audio codec method provided by the embodiments of the present application can significantly reduce the computing overhead while maintaining good audio reconstruction quality, is suitable for terminal devices with limited computing resources, and can stably operate in complex communication scenarios.
[0252] Based on any of the above embodiments, the step S10 includes steps S11, S12, S13, and S14.
[0253] In step S11, a single speaker voice, background sound, and room impulse response are obtained.
[0254] Considering the existing audio codec model, it is unstable when facing diversified audio content (such as music, different language speech, or multi-person conversation), lacks good scene adaptability and robustness, and is difficult to meet the diversity requirements in actual applications. Therefore, in the present embodiment, sample audio data is constructed by using a data enhancement technique for model training, which can improve the generalization ability of the model in various real scenarios.
[0255] The single speaker voice refers to a "clean" voice signal emitted by a speaker without mixing any other sound (such as background noise, music, or other speakers). The single speaker voice can include multiple languages. The single speaker voice can be derived from a public data set, for example, the LibriSpeech data set.
[0256] Background sound refers to all other sounds existing in the environment, for example, noise (e.g., street environment sound, office air conditioner sound, computer fan sound, keyboard tapping sound, car horn sound, etc.), music sound, etc. The background sound can be derived from a public dataset, for example, the MUSAN dataset.
[0257] The room impulse response (RIR) describes the acoustic propagation characteristics of sound in a closed room under certain conditions, which can be generated by measuring the reflection, refraction and scattering process of the room to the sound signal. The room impulse response can be derived from a public database, for example, the Aachen Impulse Response (AIR) database, which contains real RIR signals collected in different rooms (such as conference rooms, staircases, offices) and different positions. The room impulse response can also be generated by a toolkit (for example, Pyroomacoustics (a Python room acoustic simulation library)), specifically, by setting different room sizes, wall reflection coefficients and sound source / microphone positions to dynamically generate simulated RIR signals to further increase data diversity.
[0258] Step S12, superimposing part of the single voice speech in the single voice speech and the background sound to obtain a single voice noisy audio.
[0259] Step S13, convolving the single voice noisy audio and the room impulse response to generate a single voice mixed audio.
[0260] A certain proportion of single voice speech (i.e., part of the single voice speech) is superimposed with the background sound to obtain a single voice noisy audio. When superimposing, superimposition can be performed according to a random signal-to-noise ratio, which is randomly selected from a preset signal-to-noise ratio range. The preset signal-to-noise ratio range is [-5dB, 25dB]. By randomly selecting the signal-to-noise ratio, different noise levels can be simulated, so that the model trained can cope with various complex noise conditions.
[0261] Then, the single voice noisy audio is convolved with the room impulse response to generate a single voice mixed audio. Convolution is a mathematical operation. In audio processing, convolving the single voice noisy audio with the room impulse response is equivalent to playing the single voice noisy audio in the room represented by the room impulse response, and recording the single voice mixed audio. The single voice mixed audio further superimposes reverberation on the basis of the superimposed noise. That is, the single voice mixed audio is a single voice audio with noise and reverberation.
[0262] It should be understood that the room impulse responses include multiple room impulse responses, and when convolution is performed on the room impulse responses, one room impulse response can be randomly selected for convolution to simulate the acoustic environment of various spaces in the real world.
[0263] At step S14, the multi-person mixed audio is generated according to the partial single-person mixed audio and the partial single-person mixed audio.
[0264] The sample audio data includes at least two of the other single-person speech than the partial single-person speech, the other single-person mixed audio than the partial single-person mixed audio, and the multi-person mixed audio.
[0265] The partial single-person mixed audio is mixed with a random number of single-person mixed audios to generate the multi-person mixed audio, for example, two mixed audios or three mixed audios, to simulate the situation in which multiple persons speak at the same time.
[0266] Further, two or three of the other single-person speech than the partial single-person speech, the other single-person mixed audio than the partial single-person mixed audio, and the multi-person mixed audio can be selected as the sample audio data. It should be understood that, when the sample audio data includes the other single-person speech than the partial single-person speech, the other single-person mixed audio than the partial single-person mixed audio, and the multi-person mixed audio, the training effect of the model is better than that of only including two of them.
[0267] The audio coding method provided by the embodiment of the present application can efficiently and automatically generate large-scale, diversified, and rich-acoustic-characteristics sample audio data in the above manner, which covers various complex environments from quiet to noisy, from no reverberation to strong reverberation, and from single-person speech to multi-person conversation, and provides a solid data foundation for training a coding and decoding neural network model with stronger robustness and better performance. The initial coding and decoding neural network model is trained by using the above sample audio data, so that the coding and decoding capability of the model for audio data in a complex environment can be improved.
[0268] Figure 6 An example of an entity structure diagram of an electronic device is shown in FIG. 6. Figure 6 As shown in FIG. 6, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 can communicate with each other through the communications bus 640. The processor 610 can invoke the logic instructions in the memory 630 to execute the above-provided audio coding method.
[0269] Moreover, the logic instructions in the memory 630 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0270] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the audio coding method provided above.
[0271] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e. they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0272] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary hardware resources, and of course can also be realized by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0273] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio encoding and decoding method, characterized in that, include: Acquire the audio data to be processed and convert it into a spectrogram; The spectrogram is input into the encoder of a pre-trained encoding / decoding neural network model. Local acoustic features are extracted from the spectrogram to obtain local feature vectors. Intra-frame dependency modeling is then performed on the local feature vectors to obtain global feature vectors. Inter-frame dependency modeling is then performed on the global feature vectors to obtain the encoded vector output by the encoder. The encoded vector is input into the residual vector quantizer of the encoding / decoding neural network model, and the encoded vector is subjected to multi-level quantization processing to obtain the codebook index output by the residual vector quantizer; wherein, the codebook index includes multiple codebook indices corresponding to the number of levels obtained after multi-level quantization processing; The codebook index is input into the decoder of the encoding / decoding neural network model, and the codebook index is decoded to obtain the reconstructed spectrogram output by the decoder.
2. The audio encoding and decoding method according to claim 1, characterized in that, The process involves inputting the spectrogram into the encoder of a pre-trained encoding / decoding neural network model, extracting local acoustic features from the spectrogram to obtain local feature vectors, performing intra-frame dependency modeling on the local feature vectors to obtain global feature vectors, and then performing inter-frame dependency modeling on the global feature vectors to obtain the encoded vector output by the encoder, including: The encoder's coding convolutional layer is used to perform convolution processing on the spectrogram to obtain local feature vectors; The intra-frame dependencies of the local feature vectors are modeled using the first Transformer layer of the encoder to obtain the global feature vectors. The inter-frame dependencies of the global feature vector are modeled using the first gated recurrent unit (GRU) layer of the encoder to obtain the encoded vector.
3. The audio encoding and decoding method according to claim 2, characterized in that, The encoding convolutional layer includes a first convolutional layer, a second convolutional layer, and a third convolutional layer, wherein the convolutional kernels of the first convolutional layer, the second convolutional layer, and the third convolutional layer all have a size of 1 in the time dimension.
4. The audio encoding / decoding method according to claim 3, characterized in that, The stride of both the first and second convolutional layers is 2 in the frequency dimension.
5. The audio encoding / decoding method according to claim 1, characterized in that, The residual vector quantizer includes N residual quantization layers, where N ≥ 3; The step of inputting the encoded vector into the residual vector quantizer of the encoding / decoding neural network model, performing multi-level quantization on the encoded vector, and obtaining the codebook index output by the residual vector quantizer includes: The encoded vector is input to the first residual quantization layer to obtain the first residual vector and the first codebook index output by the first residual quantization layer; The first residual vector is input to the second residual quantization layer to obtain the second residual vector and the second codebook index output by the second residual quantization layer; Similarly, the (N-1)th residual vector is input to the Nth residual quantization layer to obtain the Nth residual vector and the Nth codebook index output by the Nth residual quantization layer; The codebook index includes the first codebook index, the second codebook index, and so on up to the Nth codebook index.
6. The audio encoding and decoding method according to claim 5, characterized in that, The step of inputting the encoded vector into the first residual quantization layer to obtain the first residual vector and the first codebook index output by the first residual quantization layer includes: Factorize the encoded vector to obtain a dimension-reduced encoded vector; In the first codebook of the first residual quantization layer, the first codebook vector most similar to the dimensionality reduction coding vector is retrieved; Calculate the residual between the dimensionality-reduced encoding vector and the first codebook vector to obtain the first residual vector; The first codebook index is determined based on the position of the first codebook vector in the first codebook.
7. The audio encoding / decoding method according to claim 6, characterized in that, The step of retrieving the first codebook vector most similar to the dimensionality-reduced encoding vector in the first codebook of the first residual quantization layer includes: L2 normalization is performed on each codebook vector in the first codebook to obtain the normalized codebook vector; The dimensionality-reduced encoding vector is subjected to L2 normalization to obtain the normalized dimensionality-reduced encoding vector; The cosine similarity between the normalized dimensionality-reduced encoding vector and the normalized codebook vector is calculated, and the first codebook vector most similar to the dimensionality-reduced encoding vector is determined based on the calculation result.
8. The audio encoding and decoding method according to claim 1, characterized in that, The step of inputting the encoded vector into the residual vector quantizer of the encoding / decoding neural network model, performing multi-level quantization on the encoded vector, and obtaining the codebook index output by the residual vector quantizer includes: Obtain the target quantization layer number; The encoded vector is input into the residual vector quantizer of the encoding / decoding neural network model, and the encoded vector is subjected to multi-level quantization processing according to the target quantization layer number to obtain the codebook index output by the residual vector quantizer.
9. The audio encoding and decoding method according to claim 1, characterized in that, The step of inputting the codebook index into the decoder of the encoding / decoding neural network model, decoding the codebook index, and obtaining the reconstructed spectrogram output by the decoder includes: Query the reconstructed codebook vector corresponding to the codebook index; Using the second GRU layer of the decoder, the inter-frame dependencies are reconstructed based on the reconstructed codebook vector to obtain the reconstructed global feature vector; Using the second Transformer layer of the decoder, the intra-frame dependencies are reconstructed based on the reconstructed global feature vector to obtain the reconstructed local feature vector; The reconstructed local feature vectors are transposed and convolved using the decoding convolutional layer of the decoder to obtain the reconstructed spectrogram.
10. The audio encoding and decoding method according to claim 1, characterized in that, The decoder includes multiple decoders. Before inputting the codebook index into the decoder of the encoding / decoding neural network model, decoding the codebook index, and obtaining the reconstructed spectrogram output by the decoder, the method further includes: Determine the target decoder; The step of inputting the codebook index into the decoder of the encoding / decoding neural network model, decoding the codebook index, and obtaining the reconstructed spectrogram output by the decoder includes: The codebook index is input into the target decoder, and the codebook index is decoded to obtain the reconstructed spectrogram output by the target decoder.
11. The audio encoding and decoding method according to any one of claims 1 to 10, characterized in that, The encoding / decoding neural network model is trained in the following manner: Acquire sample audio data and convert the sample audio data into a sample spectrogram; Based on the sample spectrogram and the sample audio data, the initial codec neural network model is trained to obtain the pre-trained codec neural network model.
12. The audio encoding and decoding method according to claim 11, characterized in that, During the training of the initial codec neural network model, in each training iteration, a random number of codebooks are randomly discarded to regularize the initial residual vector quantizer, wherein the random number is randomly determined according to a preset range.
13. The audio encoding / decoding method according to claim 11, characterized in that, The step of training the initial codec neural network model based on the sample spectrogram and the sample audio data to obtain the pre-trained codec neural network model includes: Based on the sample spectrogram and the sample audio data, the initial codec neural network model is trained to obtain the first codec neural network model; A second codec neural network model is constructed based on the first codec neural network model. The second codec neural network model has the same architecture as the first codec neural network model. The number of channels in each layer of the decoder in the second codec neural network model is less than the number of channels in the corresponding layer of the decoder in the first codec neural network model. The first codec neural network model is used as the teacher model, and the second codec neural network model is used as the student model. The student model is trained based on the teacher model to obtain the pre-trained codec neural network model.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the audio encoding / decoding method as described in any one of claims 1 to 13.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio encoding / decoding method as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Audio coding and decoding method and device based on neural network, equipment and storage medium
CN119152863A
Audio coding and decoding method, device, equipment and medium
CN120412605A