Audio coding and decoding processing method and device, electronic equipment and storage medium

By using a multi-level encoding and decoding network in audio encoding and decoding, and using residual signals for multiple encoding and decoding, the problems of complexity and inefficiency of different compression code rate algorithms in the prior art are solved, and efficient audio encoding and decoding is achieved.

CN120148532APending Publication Date: 2025-06-13PANORAMIC SOUND (BEIJING) INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311714578.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, compression algorithms with different compression code rates need to be used to encode and code audio code streams, resulting in incomplexity and inefficiency.

Method used

By obtaining the input signal and decoded signal of each level of the codec network in the multi-level codec network, determining the residual signal, and re-encoding it as the input signal of the next level codec network, forming a multi-level codec stream. Multi-stage encoding and decoding audio signals are used to multiple encoding and decode audio signals to improve compression efficiency and quality.

Benefits of technology

Residual recoding method based on deep learning is realized, and audio codec streams are efficiently encoded and decoded, reducing redundant information and distortion errors, and improving audio compression efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148532A_ABST
    Figure CN120148532A_ABST
Patent Text Reader

Abstract

The invention provides an audio coding and decoding processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an input signal and a decoding signal of each stage of coding and decoding network in the multi-stage coding and decoding network; determining a residual signal of each level of coding and decoding network according to the input signal and the decoding signal of each level of coding and decoding network; taking the residual signal of each level of coding and decoding network as an input signal of the next level of coding and decoding network to carry out coding again so as to obtain a multi-level coding code stream of the multi-level coding and decoding network; and performing decoding processing on each level of coding code stream corresponding to each level of decoder network in the multi-level coding and decoding network by adopting each level of decoder network in the multi-level coding and decoding network, and adding each level of decoding code stream of each level of decoder network to obtain a decoding code stream of the multi-level coding and decoding network. The technical problem that compression algorithms with different compression code rates need to be adopted to encode and decode audio code streams in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to audio processing technologies, and in particular to an audio encoding and decoding processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, during the encoding and decoding process of audio-visual bitstreams, it is usually necessary to configure multiple compression bitrates according to different transmission / storage conditions, so as to use a compression algorithm with a higher bitrate when the conditions are good and a compression algorithm with a lower bitrate when the conditions are poor.

[0003] Moreover, compression algorithms based on deep learning all adopt the method of retraining with different bitrates, that is, designing and training a deep learning network according to one bitrate, and several deep learning networks need to be trained for several bitrates.

[0004] It can be seen that there are technical problems in the related art of encoding and decoding audio bitstreams using compression algorithms with different compression bitrates. Summary of the Invention

[0005] The present application provides an audio encoding and decoding processing method, apparatus, electronic device, and storage medium, which are used to solve the technical problem in the related art that it is necessary to use compression algorithms with different compression bitrates to encode and decode audio bitstreams, and to achieve the technical effect of encoding and decoding audio bitstreams in a manner of performing residual re-encoding based on deep learning.

[0006] On the one hand, the present application provides an audio encoding and decoding processing method, and the method includes:

[0007] Obtaining the input signal and decoded signal of each encoding and decoding network in a multi-level encoding and decoding network;

[0008] Determining the residual signal of each encoding and decoding network according to the input signal and decoded signal of each encoding and decoding network;

[0009] Using the residual signal of each encoding and decoding network as the input signal of the next-level encoding and decoding network for re-encoding to obtain the multi-level encoded bitstream of the multi-level encoding and decoding network;

[0010] Respectively using each decoder network in the multi-level encoding and decoding network to perform decoding processing on each-level encoded bitstream corresponding to each decoder network in the multi-level encoded bitstream, and adding the decoded bitstreams of each level of each decoder network obtained to obtain the decoded bitstream of the multi-level encoding and decoding network.

[0011] An optional implementation method is to use the residual signal of each level of the codec network as the input signal of the next-level codec network for re-encoding to obtain the multi-level encoded bitstream of the multi-level codec network, including:

[0012] Use the next-level codec network to perform dimensionality reduction processing on the residual signal to obtain the multi-dimensional vector of the next-level codec network;

[0013] Use the quantization network in the multi-level codec network to retrieve the codeword corresponding to the multi-dimensional vector in the codebook, so that the index value of the codeword is used to replace the multi-dimensional vector in subsequent processing to obtain the multi-level encoded bitstream of the multi-level codec network.

[0014] An optional implementation method is to perform decoding processing on each level of encoded bitstream corresponding to each level of decoder network in the multi-level encoded bitstream to obtain each level of decoded bitstream corresponding to each level of decoder network, including:

[0015] Use the inverse quantization network in the multi-level codec network to retrieve the corresponding multi-dimensional vector from the codebook according to the index value of the codeword;

[0016] Use the multi-level decoder network in the multi-level codec network to perform decoding processing on each level of encoded bitstream corresponding to each level of decoder network, so that each performs dimensionality increase processing on the multi-dimensional vector to obtain each level of decoded bitstream corresponding to each level of decoder network.

[0017] An optional implementation method, before using the quantization network in the multi-level codec network to retrieve the codeword corresponding to the multi-dimensional vector, the method further includes:

[0018] If the multi-dimensional vector does not match the codebook tensor of the quantization network, use the mapping network in the multi-level codec network to process the multi-dimensional vector to match the codebook tensor of the quantization network.

[0019] On the other hand, the present application provides an audio codec processing device, and the device includes:

[0020] An acquisition module, configured to acquire the input signal and the decoded signal of each level of codec network in the multi-level codec network;

[0021] A determination module, configured to determine the residual signal of each level of codec network according to the input signal and the decoded signal of each level of codec network;

[0022] An encoding processing module, configured to use the residual signal of each level of the encoding and decoding network as the input signal of the next-level encoding and decoding network for re-encoding, so as to obtain the multi-level encoded bitstreams of the multi-level encoding and decoding network;

[0023] A decoding processing module, configured to respectively use each decoder network in the multi-level encoding and decoding network to perform decoding processing on each-level encoded bitstream corresponding to each decoder network in the multi-level encoded bitstream, and add the decoded bitstreams of each level of the obtained decoder network to obtain the decoded bitstream of the multi-level encoding and decoding network.

[0024] An optional implementation manner, the encoding processing module includes:

[0025] A first processing unit, configured to use the next-level encoding and decoding network to perform dimensionality reduction processing on the residual signal to obtain a multi-dimensional vector of the next-level encoding and decoding network;

[0026] A second processing unit, configured to use the quantization network in the multi-level encoding and decoding network to retrieve the codeword corresponding to the multi-dimensional vector in the code table, so as to use the index value of the codeword to replace the multi-dimensional vector in subsequent processing to obtain the multi-level encoded bitstreams of the multi-level encoding and decoding network.

[0027] An optional implementation manner, the decoding processing module includes:

[0028] A third processing unit, configured to use the inverse quantization network in the multi-level encoding and decoding network to retrieve the corresponding multi-dimensional vector from the code table according to the index value of the codeword;

[0029] A fourth processing unit, configured to use the multi-level decoder network in the multi-level encoding and decoding network to perform decoding processing on each-level encoded bitstream corresponding to each decoder network in the multi-level encoded bitstream, so as to respectively perform dimensionality increase processing on the multi-dimensional vector to obtain each-level decoded bitstream corresponding to each decoder network.

[0030] On the other hand, the present application provides an electronic device, including: a processor, and a memory connected to the above-mentioned processor; the above-mentioned memory stores computer execution instructions; the above-mentioned processor executes the computer execution instructions stored in the above-mentioned memory to implement the method as described in any one of the above.

[0031] On the other hand, the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method as described in any one of the above.

[0032] On the other hand, the present application provides a computer program product, including a computer program which, when executed by a processor, implements any of the above-mentioned methods.

[0033] The audio encoding and decoding processing method, device, electronic device and storage medium provided by the present application obtain the input signal and decoding signal of each level of the multi-level encoding and decoding network; determine the residual signal of each level of the above-mentioned encoding and decoding network according to the input signal and decoding signal of each level of the above-mentioned encoding and decoding network; use the residual signal of each level of the above-mentioned encoding and decoding network as the input signal of the next level of the encoding and decoding network for re-encoding to obtain the multi-level encoding code stream of the above-mentioned multi-level encoding and decoding network; respectively use each level of decoder network in the above-mentioned multi-level encoding and decoding network to perform decoding processing on each level of encoding code stream corresponding to each level of decoder network in the above-mentioned multi-level encoding code stream, and add the decoding code stream of each level of the obtained decoder network to obtain the decoding code stream of the above-mentioned multi-level encoding and decoding network. It can solve the technical problem in the related art that different compression algorithms with different compression rates are required to encode and decode the audio code stream, and achieve the technical effect of encoding and decoding the audio code stream by means of residual re-encoding based on deep learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0035] Figure 1 is a schematic flowchart of an audio encoding and decoding processing method provided by an embodiment of the present application;

[0036] Figure 2 is a schematic encoding diagram of an optional three-level encoding and decoding network provided by an embodiment of the present application;

[0037] Figure 3 is a schematic decoding diagram of an optional three-level decoding network provided by an embodiment of the present application;

[0038] Figure 4 is a schematic model construction diagram of an optional encoding and decoding network provided by an embodiment of the present application;

[0039] Figure 5 is a schematic decoding diagram of an optional two-level decoding network provided by an embodiment of the present application;

[0040] Figure 6 is a structural block diagram of an audio encoding and decoding processing device provided by an embodiment of the present application;

[0041] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0042] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of Specific Embodiments

[0043] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0044] The audio encoding and decoding processing method provided by the present application aims to solve the above technical problems in the prior art. The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the drawings.

[0045] Figure 1 is a schematic flowchart of an audio encoding and decoding processing method provided by an embodiment of the present application, as Figure 1 shown, the method includes:

[0046] S101, obtaining the input signal and the decoded signal of each level of the multi-level encoding and decoding network;

[0047] S102, determining the residual signal of each level of the above-mentioned encoding and decoding network according to the input signal and the decoded signal of each level of the above-mentioned encoding and decoding network;

[0048] S103, using the residual signal of each level of the above-mentioned encoding and decoding network as the input signal of the next level of the encoding and decoding network for re-encoding to obtain the multi-level encoded bitstreams of the above-mentioned multi-level encoding and decoding network;

[0049] S104, respectively using each level of the decoder network in the above-mentioned multi-level encoding and decoding network to decode each level of the encoded bitstream corresponding to each level of the decoder network in the above-mentioned multi-level encoded bitstream, and adding the decoded bitstreams of each level of the above-mentioned decoder network to obtain the decoded bitstream of the above-mentioned multi-level encoding and decoding network.

[0050] The audio encoding and decoding processing method provided by the embodiments of the present application is specifically a multi-level encoding method based on deep learning residual re-encoding. A multi-level encoded bitstream is formed by using the difference (residual signal) between the input signal and the decoded signal of each level of the encoding and decoding network as the input signal of the next-level encoder for re-encoding. The multi-level encoding and decoding network is used to encode and decode the audio signal multiple times to improve the compression efficiency and quality of the audio. By calculating the residual signal of each level of the encoding and decoding network and using it as the input signal of the next-level encoding and decoding network, this method realizes the step-by-step dimensionality reduction and dimensionality increase processing of the audio signal, reducing redundant information and distortion errors.

[0051] For example, in applications such as pre-encoded live recordings, the original bitstream is pre-encoded into multiple levels: first-level encoded bitstream + second-level encoded bitstream + third-level encoded bitstream +... + N-level encoded bitstream, and then a part or all of them are dynamically transmitted according to the actual situation (such as network conditions or user requirements).

[0052] For another example, in applications such as voice communication and live broadcast, multi-level encoding can be performed in real time according to network conditions or user requirements: first-level encoded bitstream + multi-level encoded bitstream for dynamic transmission.

[0053] Figure 2 It is a schematic diagram of an optional multi-level (N-level, N >= 1, shown with N = 3 as an example) encoding and decoding network. As Figure 2 shown, assuming that the audio bitstream is encoded into: total encoded bitstream = first-level encoded bitstream + second-level encoded bitstream +... + N-level encoded bitstream, then one of the following combinations is used according to the actual situation during encoding / transmission / decoding:

[0054] (1) First-level encoded bitstream;

[0055] (2) First-level encoded bitstream + second-level encoded bitstream;

[0056] (3) First-level encoded bitstream + second-level encoded bitstream + third-level encoded bitstream;

[0057] (4) And so on;

[0058] (5) First-level encoded bitstream + second-level encoded bitstream + third-level encoded bitstream... + N-level encoded bitstream.

[0059] As Figure 2As shown, in the multi-level encoding and decoding network, the first-level encoding and decoding network: the first-level encoder network + the first-level decoder network, is used to train the encoding and decoding network of the input audio; the second-level encoding and decoding network: the second-level encoder network + the second-level decoder network, is used to train the encoding and decoding network of the first difference (i.e., the input audio minus the first-level decoded bitstream); the third-level encoding and decoding network: the third-level encoder network + the third-level decoder network, is used to train the encoding and decoding network of the second difference (i.e., the first difference minus the second-level decoded bitstream); the N-level encoding and decoding network: the N-level encoder network + the N-level decoder network, is used to train the encoding and decoding network of the N - 1th difference (i.e., the N - 2th difference minus the N - 1th decoded bitstream).

[0060] During training, the first-level encoding and decoding network corresponding to the first-level encoded bitstream can be trained first; then the second-level encoding and decoding network corresponding to the second-level encoded bitstream is trained. At this time, the previously trained first-level encoding and decoding network is used, and only the second-level encoding and decoding network is trained; then the third-level encoding and decoding network corresponding to the third-level encoded bitstream is trained. At this time, the previously trained first-level encoding and decoding network and the second-level encoding and decoding network are used, and only the third-level encoding and decoding network is trained; and so on. Finally, the N-level encoding and decoding network corresponding to the N-level encoded bitstream is trained. At this time, the previously trained first-level encoding and decoding network + the second-level encoding and decoding network +... + the N - 1th-level encoding and decoding network are used, and only the N-level decoder network is trained.

[0061] Adopting the embodiment of the present application, in the encoding process, the input audio bitstream passes through the first-level encoder network to obtain the first-level encoded bitstream; the first-level encoded bitstream passes through the first-level decoder network to obtain the first-level decoded bitstream; then the input audio bitstream is subtracted from the first-level decoded bitstream to obtain the first difference, and the first difference passes through the second-level encoder network to obtain the second-level encoded bitstream; the second-level encoded bitstream passes through the second-level decoder network to obtain the second-level decoded bitstream; the first difference is subtracted from the second-level decoded bitstream to obtain the second difference; the second difference passes through the third-level encoder network to obtain the third-level encoded bitstream; the first-level encoded bitstream, the second-level encoded bitstream, and the third-level encoded bitstream output in the above steps are combined to be the multi-level encoded bitstream of the multi-level encoding and decoding network.

[0062] In the decoding process, taking the decoding schematic diagram of the third-level encoded bitstream as shown in Figure 3 as an example, the first-level encoded bitstream is taken out from the encoded bitstream, and the first-level encoded bitstream passes through the first-level decoder network to obtain the first-level decoded bitstream; the second-level encoded bitstream is taken out from the encoded bitstream, and the second-level encoded bitstream passes through the second-level decoder network to obtain the second-level decoded bitstream; the third-level encoded bitstream is taken out from the encoded bitstream, and the third-level encoded bitstream passes through the third-level decoder network to obtain the third-level decoded bitstream; the first-level decoded bitstream, the second-level decoded bitstream, and the third-level decoded bitstream obtained in the above steps are added together to obtain the decoded bitstream of the multi-level encoding and decoding network.

[0063] As shown in Figure 4As shown in the figure, in the embodiments of the present application, the model construction of the multi-level encoding and decoding network is divided into three processes, namely, the training process, the encoding process, and the decoding process:

[0064] The training process, taking Figure 1 the three-level deep learning network shown in the figure as an example. Among them, as Figure 4 shown on the left, each level of the encoder network and as Figure 4 shown on the right, the decoder network both adopt deep learning networks with the same structure (in practice, different levels can adopt different encoder-decoder networks).

[0065] In the embodiments of the present application, the network composition can be but is not limited to: the training sample is an audio with a single-channel 48kHz sampling rate; the encoder network, the decoder network, and the mapping network are all composed of convolutional networks; the purpose of the mapping network is to make the tensor output by the encoder network match the codebook tensor of the quantization network; it is also possible to use a codebook tensor that matches the output tensor of the encoder network. In this case, the mapping network can be not used; the time-domain downsampling multiples of the encoder sub-networks 1, 2, 3, and 4 are 3, 4, 5, and 5 respectively; the time-domain upsampling multiple of the decoder network is the same as that of the corresponding encoder network; the codebook of the quantization network is (1024, 64), that is, it is composed of 1024 vectors with a length of 64; the encoded bitstream is the index value of the codeword corresponding to the output vector of the quantization network.

[0066] In the embodiments of the present application, the role of the mapping network is to transform the output of the encoder network to be equal to the length of the codeword (the length is 64), and the inverse quantization network is the process of finding the corresponding codeword from the codebook using the index of the codeword;

[0067] As an alternative embodiment, the training process: First, train the one-level encoding and decoding network as Figure 4 shown in the figure: perform signal processing on the input signal and truncate it into a segment of (1, 9600), that is, a single channel with 9600 sample points; successively downsample it to a multi-dimensional vector of (64, 3200), (128, 800), (256, 160), and (512, 32) through the first-level encoder sub-networks 1, 2, 3, and 4; (512, 32) is transformed into (64, 32) through the mapping network, that is, 32 vectors with a length of 64; find the index of the nearest codeword from the codebook in the first-level quantization network, and find 32 indexes for 32 vectors with a length of 64; each index is composed of 10 bits (10 bits cover the index range of the codebook 1024).

[0068] After that, transform 32 vectors of length 64 into 32 10-bit vectors = 320 bits, and output 32 10-bit indexes; find the corresponding vectors (inverse quantization) from the codebook according to the indexes to obtain vectors (64, 32); output to the first-level decoder network, and sequentially increase the dimension to (256, 160), (128, 800), (64, 3200), (1, 9600); the first-level decoder network outputs a vector with 9600 mono-sample points (i.e., the first-level decoded bitstream).

[0069] In the embodiment of the present application, the loss function Loss of the above training process includes two parts Loss1 + Loss2:

[0070] Loss function 1: The loss between the input vector xi (1, 9600) and the output vector yi (1, 9600)

[0071] Loss function 2: The loss between the vector ei (64, 32) in the first-level codebook corresponding to the actual coding vector gi (64, 32) and the actual vector gi

[0072] The above loss function is the objective function, and the gradient descent method is used to adjust the model parameters (the first-level encoding and decoding network), that is, the model parameters are adjusted as follows:

[0073]

[0074] Among them, θ represents the model parameters to be adjusted, and Loss represents the reconstruction loss, is the partial derivative symbol.

[0075] After that, train the second-level encoding and decoding network (using an encoding and decoding network with the same structure), and solidify the weights of the first-level encoding and decoding network as the weights after training in the training network. Figure 4 the same structure of the encoding and decoding network)

[0076] For example, the input signal can be processed, truncated into (1, 9600), that is, a monophonic segment with 9600 sample points; a vector of (1, 9600), that is, the first-level decoded bitstream, is generated through a first-level encoding and decoding network; the input signal (1, 9600) is subtracted from the first-level decoded bitstream (1, 9600) to obtain a first difference (1, 9600), and the first difference is used as the input bitstream of the second-level encoding and decoding network; the first difference is successively reduced in dimension to multi-dimensional vectors of (64, 3200), (128, 800), (256, 160), and (512, 32) through the second-level encoder sub-networks 1, 2, 3, and 4; it is transformed into (64, 32) through a mapping network, that is, 32 vectors with a length of 64; the index of the nearest codeword is found from the code table in the second-level quantization network, and 32 indexes are found for 32 vectors with a length of 64; each index consists of 10 bits (the 10 bits cover the index range of the code table); 32 vectors with a length of 64 are transformed into 32 10-bit vectors = 320 bits; 32 10-bit indexes are output; the corresponding vectors (inverse quantization) are found from the code table according to the indexes to obtain vectors (64, 32); it is output to the second-level decoder network and successively increased in dimension to (256, 160), (128, 800), (64, 3200), and (1, 9600); the second-level decoder network outputs a vector with 9600 monophonic sample points (that is, the second-level decoded bitstream). Other encoder networks are analogized using the above training method.

[0077] Optionally, the loss function Loss in the above training process includes Loss1 + Loss2:

[0078] Loss function 1: The L1 loss between the first difference xi(1, 9600) and the second-level decoded bitstream yi(1, 9600):

[0079]

[0080] Loss function 2: The L1 loss between the vector ai(64, 32) in the second-level code table corresponding to the actual encoded vector bi(64, 32) and the actual vector bi:

[0081]

[0082] This loss function Loss is the objective function, and the gradient descent method is used to adjust the model parameters (the second-level encoding and decoding network), that is, the model parameters are adjusted using the following formula:

[0083]

[0084] Among them, θ represents the model parameters to be adjusted, and Loss represents the reconstruction loss. is the partial derivative symbol.

[0085] An optional implementation manner is to use the residual signal of each level of the above-mentioned encoding and decoding network as the input signal of the next level of the encoding and decoding network for re-encoding to obtain the multi-level encoded bitstream of the above-mentioned multi-level encoding and decoding network, including:

[0086] Use the next-level encoding and decoding network above to perform dimensionality reduction processing on the above residual signal to obtain the multi-dimensional vector of the next-level encoding and decoding network;

[0087] Use the quantization network in the above multi-level encoding and decoding network to retrieve the codeword corresponding to the above multi-dimensional vector in the codebook, so that the index value of the above codeword is used to replace the above multi-dimensional vector in subsequent processing to obtain the multi-level encoded bitstream of the above multi-level encoding and decoding network.

[0088] In the embodiment of the present application, a quantization network is introduced into the multi-level encoding and decoding network to perform codebook retrieval and index value replacement on the multi-dimensional vector, further reducing the bit rate and storage space of the audio.

[0089] An optional implementation manner is that before using the quantization network in the above multi-level encoding and decoding network to retrieve the codeword corresponding to the above multi-dimensional vector, the method further includes: if the above multi-dimensional vector does not match the codebook tensor of the quantization network, use the mapping network in the above multi-level encoding and decoding network to process the above multi-dimensional vector to match the codebook tensor of the quantization network.

[0090] In the embodiment of the present application, by using a quantization network and an inverse quantization network to perform codebook retrieval and index value replacement on the multi-dimensional vector, the bit rate and storage space of the audio are further reduced. The method also processes the unmatched multi-dimensional vector by using a mapping network, enhancing the adaptability and compatibility of the algorithm.

[0091] For example, taking the first-level encoded bitstream + the second-level encoded bitstream as an example, in an optional encoding process, the sample length is 96,000 sample points for mono (1, 96000); (1, 96000) passes through the first-level encoder network and is successively reduced in dimension to (64, 32000), (128, 8000), (256, 1600), (512, 320). The multi-dimensional vector (512, 320) generated in the above steps is transformed to (64, 320) through the mapping network.

[0092] After that, for 320 vectors with a length of 64, the indices of the closest codewords are found from the codebook in the first-level quantization network respectively, and 320 10-bit indices are output, which is the first-level encoded bitstream; the 320 10-bit indices are used to find the corresponding vectors (first-level inverse quantization) from the first-level codebook, obtaining vectors (64, 320); the multi-dimensional vectors (64, 320) obtained from the above steps are successively upsampled to (256, 1600), (128, 8000), (64, 32000), (1, 96000) through the first-level decoder network, obtaining new multi-dimensional vectors (1, 96000), which is the first-level decoded bitstream; the input bitstream (1, 96000) is subtracted from the first-level decoded bitstream (1, 96000) to obtain the first difference (1, 96000); the first difference (1, 96000) is successively downsampled to (64, 32000), (128, 8000), (256, 1600), (512, 320) through the second-level encoder network. After the multi-dimensional vectors (512, 320) generated from the above steps are transformed to (64, 320) through the mapping network; then, for 320 vectors with a length of 64, the index values of the closest codewords are found from the codebook in the second-level quantization network respectively, and 320 10-bit index values are output as the second-level encoded bitstream; finally, the first-level encoded bitstream 320 * 10-bit and the second-level encoded bitstream 320 * 10-bit output in the above embodiments are obtained, which is the second-level encoded bitstream of the second-level encoding and decoding network.

[0093] An optional implementation manner is to perform decoding processing on each level of encoded bitstream corresponding to each level of decoder network in the above multi-level encoded bitstream to obtain each level of decoded bitstream corresponding to each level of decoder network, including:

[0094] Using the inverse quantization network in the above multi-level encoding and decoding network, the corresponding multi-dimensional vectors are retrieved from the codebook according to the index values of the codewords;

[0095] Using the multi-level decoder network in the above multi-level encoding and decoding network, decoding processing is performed on each level of encoded bitstream corresponding to each level of decoder network, so as to perform upsampling processing on the multi-dimensional vectors respectively, and obtain each level of decoded bitstream corresponding to each level of decoder network.

[0096] In the embodiment of the present application, an inverse quantization network is introduced into the multi-level encoding and decoding network, the corresponding multi-dimensional vectors are retrieved from the codebook according to the index values, and the multi-level decoder network is used to perform upsampling processing on the multi-dimensional vectors to restore the original audio signal.

[0097] Taking the first-level encoded bitstream + the second-level encoded bitstream as an example, an optional decoding process provided in the embodiment of the present application can be as follows: Figure 5 As shown below:

[0098] For example, the input is a first-level encoded bitstream with a coding bitstream of 320 * 10 bits and a second-level encoded bitstream of 320 * 10 bits; 320 vectors (64, 320) are obtained through inverse quantization according to the 320 indexes of the first-level encoded bitstream, that is, the first-level inverse quantization; then the vectors (64, 320) are successively upsampled to (256, 1600), (128, 8000), (64, 32000), (1, 96000) through the first-level decoder sub-network, and (1, 96000) is the first-level decoded bitstream; 320 vectors (64, 320) are obtained through inverse quantization according to the 320 indexes of the second-level encoded bitstream, that is, the second-level inverse quantization.

[0099] Then, the (64, 320) obtained in the above steps is successively upsampled to (256, 1600), (128, 8000), (64, 32000), (1, 96000) through the second-level decoder sub-network, and (1, 96000) is the second-level decoded bitstream; then, the first-level decoded bitstream and the second-level decoded bitstream are added sample by sample to obtain the second-level decoded bitstream of the second-level encoding and decoding network.

[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0101] According to one or more embodiments of the present application, an audio encoding and decoding processing device is provided. Figure 6 This is a structural block diagram of an audio encoding and decoding processing device provided by an embodiment of the present application, as Figure 6 shown. The above device includes:

[0102] An acquisition module 601, configured to acquire an input signal and a decoded signal of each level of the multi-level encoding and decoding network;

[0103] A determination module 602, configured to determine a residual signal of each level of the above encoding and decoding network according to the input signal and the decoded signal of each level of the above encoding and decoding network;

[0104] An encoding processing module 603, configured to use the residual signal of each level of the above encoding and decoding network as an input signal of the next-level encoding and decoding network for re-encoding to obtain multi-level encoded bitstreams of the above multi-level encoding and decoding network;

[0105] A decoding processing module 604, configured to separately perform decoding processing on each level of encoded bitstream corresponding to each level of decoder network in the above multi-level encoding and decoding network by using each level of decoder network in the above multi-level encoding and decoding network, and add the decoded bitstreams of each level of decoder network obtained to obtain the decoded bitstream of the above multi-level encoding and decoding network.

[0106] According to one or more embodiments of the present application, the above encoding processing module includes:

[0107] A first processing unit, configured to perform dimensionality reduction processing on the above residual signal by using the above next-level encoding and decoding network to obtain a multi-dimensional vector of the above next-level encoding and decoding network;

[0108] A second processing unit, configured to retrieve, by using the quantization network in the above multi-level encoding and decoding network, a codeword corresponding to the above multi-dimensional vector in a code table, so that in subsequent processing, the index value of the above codeword is used to replace the above multi-dimensional vector to obtain the multi-level encoded bitstream of the above multi-level encoding and decoding network.

[0109] According to one or more embodiments of the present application, the above decoding processing module includes:

[0110] A third processing unit, configured to retrieve, by using the inverse quantization network in the above multi-level encoding and decoding network, the corresponding above multi-dimensional vector from the above code table according to the index value of the above codeword;

[0111] A fourth processing unit, configured to perform decoding processing on each level of encoded bitstream corresponding to each level of decoder network in the above multi-level encoded bitstream by using the multi-level decoder network in the above multi-level encoding and decoding network, so as to perform dimensionality increase processing on the above multi-dimensional vector respectively to obtain each level of decoded bitstream corresponding to each level of decoder network respectively.

[0112] In an exemplary embodiment, the embodiment of the present application further provides an electronic device, including: a processor, and a memory connected to the above processor;

[0113] The above memory stores computer-executable instructions;

[0114] The above processor executes the computer-executable instructions stored in the above memory to implement the method as described in any one of the above.

[0115] In an exemplary embodiment, the embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the above.

[0116] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program which, when executed by a processor, implements any of the above methods.

[0117] To implement the above embodiment, the embodiment of the present application further provides an electronic device. Referring to Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiment of the present application. The electronic device 700 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, messaging devices, game consoles, medical devices, fitness devices, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiment of the present application.

[0118] As Figure 7 shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0119] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7An electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0120] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present application are performed.

[0121] It should be noted that the above-mentioned computer-readable medium in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0122] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.

[0123] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0124] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0126] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0127] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.

[0128] In the context of this application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0129] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include known common knowledge or conventional technical means in this technical field that are not disclosed in this application. The specification and examples are only to be considered as exemplary, and the true scope and spirit of this application are pointed out by the following claims.

[0130] It should be understood that this application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.

Claims

1. An audio encoding and decoding processing method, characterized in that, the method includes: obtaining the input signal and the decoded signal of each level of the multi-level encoding and decoding network; determining the residual signal of each level of the encoding and decoding network according to the input signal and the decoded signal of each level of the encoding and decoding network; using the residual signal of each level of the encoding and decoding network as the input signal of the next level of the encoding and decoding network for re-encoding to obtain the multi-level encoded bitstream of the multi-level encoding and decoding network; respectively using each level of the decoder network in the multi-level encoding and decoding network to perform decoding processing on each level of the encoded bitstream corresponding to each level of the decoder network in the multi-level encoded bitstream, and adding the decoded bitstreams of each level of the decoder network obtained to obtain the decoded bitstream of the multi-level encoding and decoding network.

2. The method according to claim 1, characterized in that, using the residual signal of each level of the encoding and decoding network as the input signal of the next level of the encoding and decoding network for re-encoding to obtain the multi-level encoded bitstream of the multi-level encoding and decoding network, including: performing dimensionality reduction processing on the residual signal by using the next level of the encoding and decoding network to obtain the multi-dimensional vector of the next level of the encoding and decoding network; using the quantization network in the multi-level encoding and decoding network to retrieve the codeword corresponding to the multi-dimensional vector in the code table, so that the index value of the codeword is used to replace the multi-dimensional vector in subsequent processing to obtain the multi-level encoded bitstream of the multi-level encoding and decoding network.

3. The method according to claim 2, characterized in that, performing decoding processing on each level of the encoded bitstream corresponding to each level of the decoder network in the multi-level encoded bitstream to obtain each level of the decoded bitstream corresponding to each level of the decoder network, including: using the inverse quantization network in the multi-level encoding and decoding network to retrieve the corresponding multi-dimensional vector from the code table according to the index value of the codeword; using the multi-level decoder network in the multi-level encoding and decoding network to perform decoding processing on each level of the encoded bitstream corresponding to each level of the decoder network in the multi-level encoded bitstream, so that each performs dimensionality increase processing on the multi-dimensional vector to obtain each level of the decoded bitstream corresponding to each level of the decoder network.

4. The method according to claim 2, characterized in that, before using the quantization network in the multi-level encoding and decoding network to retrieve the codeword corresponding to the multi-dimensional vector, the method further includes: if the multi-dimensional vector does not match the code table tensor of the quantization network, using the mapping network in the multi-level encoding and decoding network to process the multi-dimensional vector to match the code table tensor of the quantization network.

5. An audio encoding and decoding processing device, characterized in that, the device includes: an obtaining module, configured to obtain the input signal and the decoded signal of each level of the multi-level encoding and decoding network; a determining module, configured to determine the residual signal of each level of the encoding and decoding network according to the input signal and the decoded signal of each level of the encoding and decoding network; An encoding processing module, configured to use the residual signal of each level of the codec network as the input signal of the next-level codec network for re-encoding, so as to obtain the multi-level encoded bitstreams of the multi-level codec network; A decoding processing module, configured to respectively use each level of decoder network in the multi-level codec network to perform decoding processing on each level of encoded bitstream in the multi-level encoded bitstream corresponding to each level of decoder network, and add the decoded bitstreams of each level of decoder network obtained, so as to obtain the decoded bitstream of the multi-level codec network.

6. The apparatus according to claim 5, wherein, the encoding processing module includes: A first processing unit, configured to use the next-level codec network to perform dimensionality reduction processing on the residual signal to obtain a multi-dimensional vector of the next-level codec network; A second processing unit, configured to use the quantization network in the multi-level codec network to retrieve the codeword corresponding to the multi-dimensional vector in the code table, so that in subsequent processing, the index value of the codeword is used to replace the multi-dimensional vector, so as to obtain the multi-level encoded bitstreams of the multi-level codec network.

7. The apparatus according to claim 6, wherein, the decoding processing module includes: A third processing unit, configured to use the inverse quantization network in the multi-level codec network to retrieve the corresponding multi-dimensional vector from the code table according to the index value of the codeword; A fourth processing unit, configured to use the multi-level decoder network in the multi-level codec network to perform decoding processing on each level of encoded bitstream in the multi-level encoded bitstream corresponding to each level of decoder network, so as to respectively perform dimensionality increase processing on the multi-dimensional vector to obtain the decoded bitstreams of each level corresponding to each level of decoder network.

8. An electronic device, wherein, it includes: a processor, and a memory connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 4.

9. A computer-readable storage medium, wherein, computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 4.

10. A computer program product, wherein, it includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Technique for encoding / decoding of codebook indices for quantized MDCT spectrum in scalable speech and audio codecs

    CN101849258A

  • Audio coding and decoding method and device, storage medium and computer equipment

    CN116504254A

  • Encoding and decoding device for sequential reproduction of picture signal

    JP1988310292A

  • Hierarchical encoding method and hierarchical decoding method for sound signal

    JP2004301954A

  • KR20210058731A