Audio decoding method, decoding device, computer equipment, program product and medium

By using the first convolution operation and feature recombination operation in the audio decoding process, instead of the deconvolution operation, the problem of high computational complexity of deconvolution operation in the prior art is solved, and a more efficient decoding process is achieved.

CN119943068APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311477636.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the existing end-to-end audio encoding and decoding scheme based on deep learning, the calculation complexity of the deconvolution operation is high, resulting in the calculation complexity of the decoding device.

Method used

By using the first convolution operation and feature recombination operation in the decoding process instead of the deconvolution operation, upsampling and feature recombination of the encoded data are realized, thereby obtaining the decoded data.

Benefits of technology

Effectively reduce the computational complexity during the decoding process, improve the decoding efficiency, and reduce the number of deconvolution operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943068A_ABST
    Figure CN119943068A_ABST
Patent Text Reader

Abstract

The invention discloses an audio decoding method, a decoding device, computer equipment, a computer program product and a non-volatile computer readable storage medium. The method comprises the following steps: carrying out nonlinear transformation operation on coded data to obtain decoded data; the nonlinear transformation operation comprises an up-sampling operation, the up-sampling operation comprises a first convolution operation and a feature recombination operation, the first convolution operation is used for increasing the number of channels of an input first feature vector to obtain an intermediate feature vector, and the feature recombination operation is used for recombining a plurality of channels of the intermediate feature vector to obtain a second feature vector; the decoding data is determined according to the second feature vector. Therefore, the first convolution operation and the feature recombination operation are utilized to achieve the effect achieved by the deconvolution operation in the prior art, so that the number of times of the deconvolution operation is reduced, the decoding efficiency is improved, and the calculation complexity in the decoding process is effectively controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of coding and decoding technology, and more specifically, to an audio decoding method, a decoding device, a computer device, a computer program product, and a non-volatile computer-readable storage medium. Background Art

[0002] In recent years, deep learning schemes have been widely used in signal processing technologies of different dimensions (such as audio, image, and video). In an end-to-end audio codec scheme based on deep learning, the encoding device can convert the original audio signal into encoded data. The decoding device obtains the transmitted information according to the encoded data, and then decodes the encoded data to obtain the final reconstructed audio signal. However, the current scheme uses a series of deconvolution layers in the decoding device to implement upsampling processing, and the computational complexity of deconvolution is high, resulting in a large computational complexity of the decoding device. Summary of the invention

[0003] The embodiments of the present application provide an audio decoding method, a decoding device, a computer device, a computer program product and a non-volatile computer-readable storage medium, which utilize a first convolution operation and a feature recombination operation to achieve the effect achieved by the deconvolution operation in the prior art, so as to effectively control the computational complexity in the decoding process.

[0004] The audio decoding method of the embodiment of the present application includes performing a nonlinear transformation operation on the encoded data to obtain decoded data; the nonlinear transformation operation includes an upsampling operation, the upsampling operation includes a first convolution operation and a feature recombination operation, the first convolution operation is used to increase the number of channels of the input first feature vector to obtain an intermediate feature vector, the feature recombination operation is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector, and the decoded data is determined based on the second feature vector.

[0005] In some embodiments, the nonlinear transformation operation is performed on the encoded data to obtain the decoded data, including: performing a second convolution operation on the encoded data to extract the first feature vector; performing the upsampling operation on the first feature vector to obtain the second feature vector; and performing a third convolution operation on the second feature vector to obtain the decoded data.

[0006] In some embodiments, the nonlinear transformation operation on the encoded data to obtain decoded data also includes: performing a second convolution operation on the encoded data to extract the first feature vector; performing a first nonlinear mapping operation on the first feature vector; performing the upsampling operation on the first feature vector after nonlinear mapping to obtain the second feature vector; performing a second nonlinear mapping operation on the second feature vector; and performing a third convolution operation on the second feature vector after nonlinear mapping to obtain the decoded data.

[0007] In some embodiments, performing the upsampling operation on the first feature vector to obtain the second feature vector includes: performing the upsampling operation on the first feature vector multiple times to obtain the second feature vector, the number of upsampling operations is determined according to the number of upsampling layers at the decoding end, the first feature vector is the input feature vector of the first upsampling operation, the second feature vector is the output feature vector of the last upsampling operation, and the input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.

[0008] In some embodiments, the number of channels of the intermediate feature vector is N times the number of channels of the first feature vector, and the feature recombination operation is used to recombine the N channels into one channel, where N is an integer.

[0009] In some embodiments, reorganizing the multiple channels of the intermediate feature vector includes: arranging the features corresponding to the positions in the multiple channels in sequence to reorganize them into one channel, and the number of features of the reorganized channel is equal to the sum of the number of features of the multiple channels before the reorganization.

[0010] In some embodiments, the number of upsampling operations and the convolution kernel size of the first convolution operation are determined according to the content complexity of the encoded data.

[0011] In some embodiments, the number and the convolution kernel size are positively correlated with the content complexity, and the content complexity is determined according to at least one of the sampling rate, the number of channels, and the number of sound source types of the encoded data.

[0012] In some embodiments, the upsampling operation also includes a nonlinear mapping operation, the intermediate feature vector includes a first intermediate feature vector and a second intermediate feature vector, the nonlinear mapping operation is used to perform nonlinear mapping on the first intermediate feature vector according to a preset activation function to obtain the second intermediate feature vector, and the feature recombination operation is used to recombine multiple channels of the second intermediate feature vector to obtain the second feature vector.

[0013] The decoding device of the embodiment of the present application includes a nonlinear transformation module, which is used to perform a nonlinear transformation operation on the encoded data to obtain decoded data. The nonlinear transformation module includes an upsampling module, and the upsampling module includes a first convolution unit and a feature recombination unit, the first convolution unit is used to increase the number of channels of the input first feature vector to obtain an intermediate feature vector, and the feature recombination unit is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector, and the decoded data is determined according to the second feature vector.

[0014] In some embodiments, the nonlinear transformation module further includes a first convolution module and a second convolution module. The first convolution module is used to perform a second convolution operation on the encoded data to extract the first feature vector. The upsampling module is used to perform the upsampling operation on the first feature vector to obtain the second feature vector. The second convolution module is used to perform a third convolution operation on the second feature vector to obtain the decoded data.

[0015] In some embodiments, the upsampling module includes multiple upsampling modules, which are connected in sequence, the first feature vector is the input feature vector of the first upsampling module, the second feature vector is the output feature vector of the last upsampling module, and the output feature vector of the previous upsampling module serves as the input feature vector of the next upsampling module.

[0016] In some embodiments, the upsampling module further includes an activation unit, the intermediate feature vector includes a first intermediate feature vector and a second intermediate feature vector, the activation unit is used to perform nonlinear mapping on the first intermediate feature vector according to a preset activation function to obtain the second intermediate feature vector, and the feature recombination unit is used to recombine multiple channels of the second intermediate feature vector to obtain the second feature vector.

[0017] In some embodiments, the upsampling module further includes one or more residual units, and the residual units are used to perform a fourth convolution operation on the output feature vector of the feature recombination unit to obtain the output feature vector of the upsampling module.

[0018] The computer device of an embodiment of the present application includes a processor, a memory and a computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the audio decoding method described in any of the above embodiments.

[0019] The computer program product of the embodiments of the present application includes a computer program, and when the computer program is executed by a processor, the audio decoding method described in any of the above embodiments is performed.

[0020] The non-volatile computer-readable storage medium of the embodiment of the present application includes a computer program. When the computer program is executed by a processor, the processor executes the audio decoding method described in any of the above embodiments.

[0021] The audio decoding method, decoding device, computer equipment, computer program product and non-volatile computer-readable storage medium of the embodiment of the present application perform an up-sampling operation on the encoded data through the first convolution operation and the feature recombination operation in the nonlinear transformation operation to increase the number of channels of the input first feature vector, and recombine multiple channels of the intermediate feature vector to obtain the second feature vector, and then obtain the encoded data. In this way, the present application uses the first convolution operation and the feature recombination operation to achieve the effect achieved by the deconvolution operation in the prior art, that is, to obtain the second feature vector. And the computational complexity of the first convolution operation and the feature recombination operation is less than the computational complexity of the deconvolution operation, so the audio decoding method of the present application can reduce the number of deconvolution operations, improve the decoding efficiency and effectively reduce the computational complexity in the decoding process.

[0022] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0024] Figure 1 It is a schematic diagram of an application scenario of an audio decoding method in some embodiments of the present application;

[0025] Figure 2 This is an example diagram of the basic process of the encoding and decoding scheme of certain implementation methods of the present application;

[0026] Figure 3 is a flowchart of an audio decoding method in some implementation modes of the present application;

[0027] Figure 4 is a schematic diagram of a scenario of an audio decoding method in some implementation modes of the present application;

[0028] Figure 5 is a schematic diagram of a scenario of an audio decoding method in some implementation modes of the present application;

[0029] Figure 6 is a flowchart of an audio decoding method in some implementation modes of the present application;

[0030] Figure 7is a flowchart of an audio decoding method in some implementation modes of the present application;

[0031] Figure 8 is a schematic diagram of the structure of a computer device of certain embodiments of the present application;

[0032] Fig. 9 It is a schematic diagram of the connection status of a non-volatile computer-readable storage medium and a processor in certain embodiments of the present application. DETAILED DESCRIPTION

[0033] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.

[0034] To facilitate understanding of this application, the terms appearing in this application are explained below:

[0035] Encoding and decoding: The audio encoding process is to compress the audio into smaller data, and the decoding process is to restore the smaller data to audio. The encoded smaller data is used for network transmission and occupies less bandwidth.

[0036] Sampling rate: The sampling rate describes the number of data contained in a unit of time (1 second). For example, a 16k sampling rate contains 16,000 sampling points, and each sampling point corresponds to a short integer.

[0037] Codebook: A collection of multiple vectors. The encoding device and the decoding device both store the same codebook.

[0038] Quantization: Find the closest vector in the codebook for the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.

[0039] Quantizer: The quantizer is responsible for quantization and updating the vectors in the codebook.

[0040] Figure 1 A schematic diagram schematically shows an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied.

[0041] like Figure 1 As shown, the system architecture includes a plurality of terminal devices, which can communicate with each other through, for example, a network. For example, the system architecture may include a first terminal device 101 and a second terminal device 102 interconnected through a network. Figure 1 In the embodiment of the present invention, the first terminal device 101 and the second terminal device 102 perform unidirectional data transmission.

[0042] For example, the first terminal device 101 can encode audio and video data (such as an audio and video data stream collected by the terminal device) for transmission to the second terminal device 102 via the network. The encoded audio and video data is transmitted in the form of one or more encoded audio and video streams. The second terminal device 102 can receive the encoded audio and video data from the network, decode the encoded audio and video data to restore the audio and video data, and play or display content based on the restored audio and video data.

[0043] Figure 2 The following diagram shows an example of the basic process of an end-to-end audio encoding and decoding solution. During encoding, the encoder at the data transmitter first encodes the input audio signal to generate a binary code stream, and then the data transmitter sends the binary code stream to the data receiver. After receiving the binary code stream, the data receiver decodes the binary code stream through the decoder to obtain a reconstructed audio signal.

[0044] In one embodiment of the present application, the system architecture may include a third terminal device 103 and a fourth terminal device 104 that perform bidirectional transmission of encoded audio and video data, which may occur, for example, during an audio and video conference. For bidirectional data transmission, each of the third terminal device 103 and the fourth terminal device 104 may encode audio and video data (e.g., an audio and video data stream collected by the terminal device) for transmission to the other terminal device of the third terminal device 103 and the fourth terminal device 104 through a network. Each of the third terminal device 103 and the fourth terminal device 104 may also receive the encoded audio and video data transmitted by the other terminal device of the third terminal device 103 and the fourth terminal device 104, and may decode the encoded audio and video data to restore the audio and video data, and play or display content according to the restored audio and video data.

[0045] exist Figure 1 In the embodiment of the present invention, the first terminal device 101, the second terminal device 102, the third terminal device 103 and the fourth terminal device 104 may be servers, personal computers and smart phones, but the principles disclosed in the present application may not be limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or dedicated audio and video conferencing equipment. The network represents any number of networks that transmit encoded audio and video data between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the fourth terminal device 104, including, for example, wired and / or wireless communication networks. The communication network may exchange data in circuit switching and / or packet switching channels. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purpose of the present application, unless explained below, the architecture and topology of the network may be irrelevant to the operation disclosed in the present application.

[0046] The audio decoding method of this application will be described in detail below:

[0047] See also Figure 3 and Figure 4 , the present application embodiment provides an audio decoding method, the audio decoding method comprising:

[0048] Step 011: performing a nonlinear transformation operation on the encoded data to obtain decoded data;

[0049] The nonlinear transformation operation includes an upsampling operation, and the upsampling operation includes a first convolution operation and a feature recombination operation. The first convolution operation is used to increase the number of channels of the input first feature vector to obtain an intermediate feature vector. The feature recombination operation is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector. The decoded data is determined based on the second feature vector.

[0050] Specifically, the data transmitting end can encode and compress the audio data through the encoding device 200. For example, the encoding device 200 may include a third convolution module 210, a downsampling module 220, and a fourth convolution module 230. During encoding, the encoding device 200 will first use the third convolution module 210, the downsampling module 220, and the fourth convolution module 230 to perform a nonlinear transformation on the input original data. The nonlinear transformation operation at this time includes a downsampling operation and a convolution operation to complete feature extraction, increase the number of channels of the original data, and reduce the dimension of the original data. The encoding device 200 may also include a first linear transformation module 240 and a quantization module 250 (for example, a quantizer). After the fourth convolution module 230 completes the convolution operation, a coding feature vector can be obtained, and the first linear transformation module 240 is used to perform a linear transformation operation on the coding feature vector. Then, the feature vector obtained after the linear transformation operation is input into the quantization module 250 for quantization (such as using a residual-based vector quantizer) operation, and a binary code stream is generated according to the quantization result, thereby obtaining the coded data. For example, the feature vector obtained after the linear transformation operation can be input into the quantizer, so that the vector index corresponding to each encoded feature vector can be queried in the codebook. The vector index can then be transmitted to the data receiving end, which decodes the vector index through the decoding device 100 to obtain the restored data.

[0051] The decoding device 100 may include an inverse quantization module 20, a nonlinear transformation module 10 and a second linear transformation module 30. The inverse quantization module 20 may perform an inverse quantization operation on the encoded data. The second linear transformation module 30 may perform a linear transformation operation on the encoded data after the inverse quantization operation. The nonlinear transformation module 10 may perform a nonlinear transformation operation on the encoded data. The nonlinear transformation module 10 includes an upsampling module 11, and the upsampling module 11 includes a first convolution unit 111 and a feature recombination unit 112.

[0052] After receiving the encoded data, the data receiving end can use the inverse quantization module 20 to perform inverse quantization. For example, after receiving the vector index transmitted by the network, the codebook feature vector corresponding to the vector index can be first queried by the quantizer in the codebook. Feature extraction is then performed based on the result of the inverse quantization to obtain the first feature vector corresponding to the encoded data. Next, the decoding device 100 inputs the first feature vector into the upsampling module 11. At this time, the first convolution unit 111 can perform a first convolution operation on the first feature vector, increase the number of channels of the first feature vector, and reduce the dimension of the first feature vector to obtain an intermediate feature vector. For example, the first convolution unit 111 performs a 1*1 convolution on the first feature vector to obtain an intermediate feature vector whose number of channels is twice the number of channels of the first feature vector. Next, the feature recombination unit 112 can perform a feature recombination operation on the intermediate feature vector to recombine multiple channels of the intermediate feature vector into one or more channels. For example Figure 5 As shown, the intermediate feature vector is divided into 4 channels, namely channel A, channel B, channel C and channel D. In one embodiment, the features of channel A and channel B can be reorganized into channel E, and the features of channel C and channel D can be reorganized into channel F. In another embodiment, the features of the first half of channel A and the features of the first half of channel B are reorganized into the first half of channel E, and the features of the second half of channel C and the features of the second half of channel D are reorganized into the second half of channel E; the features of the first half of channel C and the features of the first half of channel D are reorganized into the first half of channel F, and the features of the second half of channel A and the features of the second half of channel B are reorganized into the second half of channel F. In this way, a second feature vector containing features of multiple channels can be obtained. Among them, the number of channels of the intermediate feature vector is N times the number of channels of the first feature vector, and the feature reorganization operation is used to reorganize N channels into one channel, and N is an integer, for example, N is 2, 3, 6 or 9, which can be specifically set according to the reorganization requirements of the intermediate feature vector. Finally, the decoding device 100 may determine the decoded data according to the second eigenvector, for example, determine the decoded data according to the result of performing a convolution operation on the second eigenvector.

[0053] In the prior art, a deconvolution operation is usually used to obtain the second eigenvector, while the present application uses a first convolution operation and a feature reorganization operation to obtain the second eigenvector, that is, the present application uses the first convolution operation and the feature reorganization operation to perform the function of the deconvolution operation, thereby using the first convolution operation and the feature reorganization operation to replace part of the deconvolution operation in the decoding process. The computational complexity of the convolution operation is less than the computational complexity of the deconvolution operation, and the computational complexity of the feature reorganization operation is also less than the computational complexity of the deconvolution operation, so that the computational complexity of the audio decoding method of the present application is less than the computational complexity of the audio decoding method of the prior art.

[0054] The audio decoding method of the embodiment of the present application performs an up-sampling operation on the encoded data through the first convolution operation and the feature recombination operation in the nonlinear transformation operation to increase the number of channels of the input first feature vector, and recombine multiple channels of the intermediate feature vector to obtain the second feature vector, and then obtain the encoded data. In this way, the present application uses the first convolution operation and the feature recombination operation to achieve the effect achieved by the deconvolution operation in the prior art, that is, to obtain the second feature vector. And the computational complexity of the first convolution operation and the feature recombination operation is less than the computational complexity of the deconvolution operation, so the audio decoding method of the present application can reduce the number of deconvolution operations, improve decoding efficiency and effectively reduce the computational complexity in the decoding process.

[0055] See also Figure 4 and Figure 6 In some implementations, step 011: performing a nonlinear transformation operation on the encoded data to obtain decoded data, includes:

[0056] Step 0111: performing a second convolution operation on the encoded data to extract a first eigenvector;

[0057] Step 0112: performing an upsampling operation on the first eigenvector to obtain a second eigenvector;

[0058] Step 0113: Perform a third convolution operation on the second eigenvector to obtain decoded data.

[0059] Specifically, after receiving the original data, the encoding device 200 can perform data sampling on the original data to be encoded, and can obtain the original feature vector. The original feature vector is input into the third convolution module 210, and the third feature vector can be obtained after convolution processing. Then, the encoding device 200 can input the third feature vector into the downsampling module 220, so as to use the downsampling module 220 to perform feature extraction, reduce the dimension of the third feature vector, and increase the number of channels of the third feature vector, thereby obtaining a fourth feature vector. Then, the fourth convolution module 230 performs convolution processing on the fourth feature vector to obtain a coded feature vector, thereby obtaining coded data, and completing the encoding of the original data.

[0060] Correspondingly, the nonlinear transformation module 10 also includes a first convolution module 12 and a second convolution module 13. After acquiring the encoded data, the first convolution module 12 performs a second convolution operation on the encoded data to extract features and obtain a first feature vector. Next, the decoding device 100 may input the first feature vector into the upsampling module 11 to perform an upsampling operation on the first feature vector, upsample and extract features from the first feature vector, and thereby obtain a second feature vector. Then, the decoding device 100 inputs the second feature vector into the second convolution module 13, performs a third convolution operation on the second feature vector, and extracts features from the second feature vector. Finally, the decoding device 100 may obtain decoded data based on the second feature vector after feature extraction.

[0061] For example, in the input stage of the encoding device 200, after receiving the original data of 24 kilohertz (khz), the original data to be encoded can be sampled to obtain an original feature vector with a channel number of 1 and a dimension of 19200; the original feature vector is input to the third convolution module 210, and after convolution processing, a third feature vector with a channel number of 32 and a dimension of 19200 can be obtained. In some optional embodiments, in order to improve the encoding efficiency, the encoding device 200 can simultaneously encode a batch of feature vectors with a number of B. In the downsampling stage of the encoding device 200, the downsampling module 220 can be used to process the third feature vector to obtain a fourth feature vector with a channel number of 512 and a dimension of 50. In the output stage of the encoding device 200, the fourth convolution module 230 performs convolution processing on the fourth feature vector to obtain a coded feature vector with a channel number of K and a dimension of 50. Wherein, K is a preset vector quantization dimension, for example, it can be 32. Next, the encoded feature vector is input into the quantizer, and a binary code stream can be generated according to the quantization result to obtain 75 Hz encoded data, thereby completing the encoding of the original signal.

[0062] After the decoding device 100 obtains the encoded data, at the input stage of the decoding device 100, the codebook in the inverse quantization module 20 is used to obtain the codebook feature vector corresponding to the encoded data, and then the codebook feature vector is input to the first convolution module 12 to perform a second convolution operation on the encoded data, thereby obtaining a first feature vector with a channel number of 512 and a dimension of 50. In the upsampling stage of the decoding device 100, the upsampling module 11 can be used to process the first feature vector to obtain a second feature vector with a channel number of 32 and a dimension of 19200. In the output stage of the decoding device 100, the second convolution module 13 performs convolution processing on the second feature vector to restore the decoded data of 24khz, a channel number of 1, and a dimension of 19200.

[0063] In this way, the decoding device 100 can use the first convolution module 12, the upsampling module 11 and the second convolution module 13 to convert the encoded data into the first feature vector, the second feature vector and the decoded data in sequence, so as to accurately obtain the decoded data according to the encoded data. At the same time, the upsampling module 11 uses the first convolution unit 111 and the feature recombination unit 112 to replace the deconvolution operation, so that the computational complexity in the decoding process is effectively reduced.

[0064] Optionally, an activation function, such as relu, leaky-relu or attention module, is a function added to an artificial neural network to help the network learn complex patterns in the data. The activation function can be used to add a nonlinear operation (i.e., nonlinear mapping operation) after the convolution operation to make the output of the neural network more complex and more expressive.

[0065] Therefore, after performing the second convolution operation on the encoded data, the first nonlinear mapping operation can also be performed on the first feature vector, that is, the first nonlinear mapping operation is performed on the first feature vector using the activation function, and then the first feature vector after nonlinear mapping is upsampled to obtain the second feature vector. The upsampling operation also includes the convolution operation, so the second nonlinear mapping operation can also be performed on the second feature vector, and the third convolution operation can be performed on the second feature vector after nonlinear mapping to obtain the decoded data. In this way, the decoding device 100 can use the activation function to increase the nonlinear factor in the nonlinear operation, so that the decoding device 100 can express more complex features.

[0066] See also Figure 4 and Figure 7 In some embodiments, step 0112: performing an upsampling operation on the first feature vector to obtain a second feature vector includes:

[0067] Step 01121: Perform multiple upsampling operations on the first feature vector to obtain a second feature vector. The number of upsampling operations is determined according to the number of upsampling layers at the decoding end. The first feature vector is the input feature vector of the first upsampling operation, and the second feature vector is the output feature vector of the last upsampling operation. The input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.

[0068] Specifically, the number of downsampling modules 220 of the encoding device 200 may be multiple, and the multiple downsampling modules 220 are connected in sequence to perform multiple downsampling on the feature vector, thereby reducing the dimension of the third feature vector multiple times, increasing the number of channels of the third feature vector, and then obtaining a fourth feature vector whose dimension and number of channels meet the preset requirements. Among them, the third feature vector is the input feature vector of the first downsampling operation, the fourth feature vector is the output feature vector of the last downsampling operation, and the input feature vector of each downsampling operation is the output feature vector of the next upsampling operation.

[0069] Correspondingly, the decoding device 100 at the decoding end may have multiple upsampling modules 11, and the multiple upsampling modules 11 are connected in sequence. In the case of performing an upsampling operation on the first feature vector, the decoding device 100 may perform multiple upsampling operations on the first feature vector, wherein the number of upsampling operations is determined according to the number of upsampling layers at the decoding end, that is, the number of upsampling operations is determined according to the number of upsampling modules 11 at the decoding end, so as to achieve multiple dimensionality increases and multiple reductions in the number of channels. Among them, the first feature vector is the input feature vector of the first upsampling operation, the second feature vector is the output feature vector of the last upsampling operation, and the input feature vector of each upsampling operation is the output feature vector of the last upsampling operation. In this way, the first feature vector can be upsampled multiple times to obtain a second feature vector having the same dimension as the original data.

[0070] For example, the third eigenvector has 32 channels and 19200 dimensions. The first downsampling module 220 reduces the dimension of the third eigenvector to 1 / 2, and obtains a fifth eigenvector with 64 channels and 9600 dimensions; the second downsampling module 220 reduces the dimension of the fifth eigenvector to 1 / 4, and obtains a sixth eigenvector with 128 channels and 2400 dimensions; the third downsampling module 220 reduces the dimension of the sixth eigenvector to 1 / 6, and obtains a seventh eigenvector with 256 channels and 400 dimensions; the fourth downsampling module 220 reduces the dimension of the seventh eigenvector to 1 / 8, and obtains a fourth eigenvector with 512 channels and 50 dimensions. Then the encoding device 200 can generate encoding data according to the fourth eigenvector.

[0071] The number of channels of the first eigenvector is 512 and the dimension is 50. After the first convolution module 12 determines the first eigenvector according to the encoded data, the first upsampling module 11 increases the dimension of the first eigenvector to 8 times, and obtains the eighth eigenvector with the number of channels of 256 and the dimension of 400; the second upsampling module 11 increases the dimension of the eighth eigenvector to 6 times, and obtains the ninth eigenvector with the number of channels of 128 and the dimension of 2400; the third upsampling module 11 increases the dimension of the ninth eigenvector to 4 times, and obtains the tenth eigenvector with the number of channels of 64 and the dimension of 9600; the fourth upsampling module 11 increases the dimension of the tenth eigenvector to 2 times, and obtains the second eigenvector with the number of channels of 32 and the dimension of 19200.

[0072] Of course, the decoding device 100 may also set multiple first convolution units 111 and feature recombination units 112 in one upsampling module 11 to complete multiple upsampling operations in one upsampling module 11.

[0073] In this way, the decoding device 100 can perform multiple upsampling operations on the first feature vector to gradually increase the dimension of the first feature vector to the original dimension of the original data received by the encoding device 200, and gradually reduce the number of channels of the first feature vector to the original number of channels of the original data received by the encoding device 200, thereby accurately reconstructing the data based on the encoded data.

[0074] See also Figure 4 In some embodiments, reorganizing multiple channels of the intermediate feature vector includes: arranging the features corresponding to the positions in the multiple channels in sequence to reorganize them into one channel, and the number of features of the reorganized channel is equal to the sum of the number of features of the multiple channels before the reorganization.

[0075] Specifically, when reorganizing multiple channels of the intermediate feature vector, the feature reorganization unit 112 may arrange the features corresponding to the positions in the multiple channels in sequence, that is, arrange the first features in the multiple channels together, arrange the second features in the multiple channels together, and so on. Among them, when the feature reorganization unit 112 arranges the features corresponding to the positions in the multiple channels in sequence, the order of arrangement may be arranged from large to small according to the sequence number of the channel, or from small to large, or may be arranged at random. For example, if there are 4 channels, the 4 channels may be sorted in sequence. When arranging the first features of the 4 channels together, the feature reorganization unit 112 may obtain features from the 4 channels in sequence according to the order of 1234, or according to the order of 4321, or according to the order of 2134, and then arrange them. In this way, the feature reorganization unit 112 may reorganize the multiple channels into one channel according to the corresponding arrangement order, and the number of features of the reorganized channel is equal to the sum of the number of features before the reorganization, so as to reduce the number of channels of the intermediate feature vector while keeping the number of features of the intermediate feature vector unchanged.

[0076] See also Figure 4 In some embodiments, the number of upsampling operations and the convolution kernel size of the first convolution operation are determined according to the content complexity of the encoded data.

[0077] Specifically, convolution kernels are a common tool in machine learning and computer vision for performing convolution operations on data such as images, audio, and video. A convolution kernel is a two-dimensional matrix that performs element-by-element product operations with the original data and adds the results to obtain a new value. The size of the convolution kernel can be various, such as 1x1, 3x3, 5x5, or 7x7. The size of the convolution kernel can be adjusted as needed to better capture the features in the data. In general, the larger the convolution kernel, the larger the receptive field, the more image information is seen, and the better the global features obtained. However, large convolution kernels will lead to a surge in the amount of calculation and reduced computing performance.

[0078] Therefore, the convolution kernel size of the first convolution operation can be determined according to the content complexity of the encoded data, and the content complexity and the size of the convolution kernel are positively correlated. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, so the convolution kernel is larger. The lower the content complexity, the smaller the convolution kernel, so that the amount of calculation can be ensured while ensuring the convolution effect.

[0079] Similarly, the more times the upsampling operation is performed, the more information is obtained, and the better the global features are. However, too many upsampling operations will result in excessive computation. The number of upsampling operations can also be determined according to the content complexity of the encoded data, and the content complexity is positively correlated with the number of upsampling operations. The higher the content complexity, the more information is required for convolution while ensuring the convolution effect, and the more upsampling operations are performed. The lower the content complexity, the smaller the convolution kernel, so that the amount of computation can be ensured while ensuring the convolution effect.

[0080] Among them, the content complexity can be determined according to at least one of the sampling rate, the number of channels, and the number of sound source types of the encoded data. It can be understood that the higher the sampling rate, the more channels, the more sound source types, and the higher the content complexity. The content complexity can be determined according to one of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, the content complexity is determined according to the sampling rate. Alternatively, the content complexity can be determined according to any two of the sampling rate, the number of channels, and the number of sound source types of the encoded data, for example, according to the sampling rate and the number of channels, or according to the number of channels and the number of sound source types. Alternatively, the content complexity can be determined according to the sampling rate, the number of channels, and the number of sound source types of the encoded data.

[0081] In this way, the appropriate number of upsampling operations and the appropriate size of the convolution kernel of the first convolution operation can be determined according to the content complexity of the encoded data, so as to ensure the convolution effect on the one hand and ensure that the amount of calculation is not too large on the other hand, thereby ensuring the computing performance during the decoding process.

[0082] See also Figure 4 In some embodiments, the upsampling operation also includes a nonlinear mapping operation, the intermediate feature vector includes a first intermediate feature vector and a second intermediate feature vector, the nonlinear mapping operation is used to perform nonlinear mapping on the first intermediate feature vector according to a preset activation function to obtain a second intermediate feature vector, and the feature recombination operation is used to recombine multiple channels of the second intermediate feature vector to obtain a second feature vector.

[0083] Specifically, the upsampling module 11 further includes an activation unit 113. The first convolution unit 111 processes the input first feature vector using a first convolution operation, and after increasing the number of channels of the first feature vector, a first intermediate feature vector can be obtained. Then the activation unit 113 can use a nonlinear mapping operation to perform nonlinear mapping on the first intermediate feature vector according to the activation function to obtain a second intermediate feature vector. Finally, the feature recombination unit 112 can use a feature recombination operation to recombine multiple channels of the second intermediate feature vector into one or more channels to obtain a second feature vector.

[0084] In this way, the decoding device 100 can use the activation function to increase the nonlinear factors in the upsampling operation, so that the decoding device 100 can express more complex features.

[0085] See also Figure 4 In some embodiments, the upsampling module 11 further includes one or more residual units 114, and the residual unit 114 is used to perform a fourth convolution operation on the output feature vector of the feature recombination unit 112 to obtain the output feature vector of the upsampling module 11.

[0086] Specifically, the residual unit 114 is a component widely used in the field of deep learning, and is often used to enhance the expressiveness of neural networks and improve performance. Its function is to extract useful features from input data and then output the results. The residual unit 114 can be used to enhance the expressiveness of deep neural networks, reduce the impact of problems such as gradient disappearance and gradient explosion, and effectively alleviate overfitting problems.

[0087] Therefore, the upsampling module 11 may further include one or more residual units 114, and the residual unit 114 may perform a fourth convolution operation on the output feature vectors of the feature recombination unit 112, such as the second feature vector, the eighth feature vector, the ninth feature vector, and the tenth feature vector, to extract features from the output feature vectors, and use the output feature vectors after feature extraction as the output feature vectors of the upsampling module 11. In this way, by using the residual unit 114, the decoding device 100 can construct a deeper neural network and achieve a higher degree of abstraction, while reducing network training time and consumption of computing resources.

[0088] See also Figure 4 In order to better implement the audio decoding method of the embodiment of the present application, the embodiment of the present application also provides a decoding device 100. The decoding device 100 may include a nonlinear transformation module 10. The nonlinear transformation module 10 is used to perform a nonlinear transformation operation on the encoded data to obtain decoded data. The nonlinear transformation module 10 includes an upsampling module 11, and the upsampling module 11 includes a first convolution unit 111 and a feature recombination unit 112. The first convolution unit 111 is used to increase the number of channels of the first feature vector to obtain an intermediate feature vector, and the feature recombination unit 112 is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector, and the decoded data is determined according to the second feature vector.

[0089] The decoding device 100 is described above from the perspective of a functional module in conjunction with the accompanying drawings. The functional module can be implemented in hardware form, can be implemented in software form, or can be implemented in combination with hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form of the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware encoding processor to execute, or can be executed by a combination of hardware and software modules in the encoding processor. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0090] See also Figure 8 The computer device 300 of the embodiment of the present application includes a processor 310, a memory 320 and a computer program, wherein the computer program is stored in the memory 320 and executed by the processor 310, and the computer program includes instructions for executing the audio decoding method of any of the above embodiments.

[0091] Optionally, the computer device 300 can be any device with image processing capabilities, such as a server or a terminal device (such as a mobile phone, a tablet computer, a display device, a laptop computer, a smart watch, a head-mounted display device, a game console, etc.).

[0092] like Figure 8 As shown, the processor 310 included in the computer device 300 is a central processing unit (CPU), and the memory 320 includes a read-only memory 321 (ROM) and a random access memory 322 (RAM). The central processor can perform various appropriate actions and processes according to the program stored in the read-only memory 321 or the program loaded from the storage part 380 to the random access memory 322. In the random access memory 322, various programs and data required for system operation are also stored. The central processor, the read-only memory 321 and the random access memory 322 are connected to each other through a bus 330. The input / output interface 340 (Input / Output interface, i.e., I / O interface) is also connected to the bus 330.

[0093] The following components are connected to the input / output interface 340: an input section 350 including a keyboard, a mouse, etc.; an output section 360 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 370 including a hard disk, etc.; and a communication section 380 including a network interface card such as a LAN card, a modem, etc. The communication section 380 performs communication processing via a network such as the Internet. A drive 390 is also connected to the input / output interface 340 as needed. A removable medium 391, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 390 as needed so that a computer program read therefrom is installed into the storage section 370 as needed.

[0094] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 380, and / or installed from the removable medium 391. When the computer program is executed by the central processing unit, various functions defined in the system of the present application are executed.

[0095] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a processor, various functions defined in the system of the present application are executed.

[0096] The computer program product of the embodiments of the present application includes a computer program, and when the computer program is executed by the processor 310, the audio decoding method of any of the above embodiments is executed.

[0097] See also Fig. 9 The embodiment of the present application also provides a computer-readable storage medium 400 on which a computer program 410 is stored. When the computer program 410 is executed by a processor 420 (for example, the processor 310 of the above-mentioned computer device 300), the steps of the audio decoding method of any of the above-mentioned embodiments are implemented. For the sake of brevity, they are not repeated here.

[0098] In the description of this specification, the descriptions with reference to the terms "certain embodiments", "in an example", "exemplarily", etc., mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0099] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0100] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An audio decoding method, characterized in that: The method comprises: Performing a nonlinear transformation operation on the encoded data to obtain decoded data; The nonlinear transformation operation includes an upsampling operation, and the upsampling operation includes a first convolution operation and a feature recombination operation. The first convolution operation is used to increase the number of channels of the input first feature vector to obtain an intermediate feature vector. The feature recombination operation is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector. The decoded data is determined based on the second feature vector.

2. The audio decoding method according to claim 1, characterized in that: The performing a nonlinear transformation operation on the encoded data to obtain decoded data includes: Performing a second convolution operation on the encoded data to extract the first feature vector; Performing the upsampling operation on the first feature vector to obtain the second feature vector; A third convolution operation is performed on the second feature vector to obtain the decoded data.

3. The audio decoding method according to claim 1, characterized in that: The performing of a nonlinear transformation operation on the encoded data to obtain decoded data further includes: performing a second convolution operation on the encoded data to extract the first feature vector; Performing a first nonlinear mapping operation on the first eigenvector; Performing the upsampling operation on the first eigenvector after nonlinear mapping to obtain the second eigenvector; performing a second nonlinear mapping operation on the second eigenvector; A third convolution operation is performed on the second feature vector after nonlinear mapping to obtain the decoded data.

4. The audio decoding method according to claim 2, characterized in that: The performing the upsampling operation on the first feature vector to obtain the second feature vector includes: The upsampling operation is performed multiple times on the first feature vector to obtain the second feature vector, the number of upsampling operations is determined according to the number of upsampling layers at the decoding end, the first feature vector is the input feature vector of the first upsampling operation, the second feature vector is the output feature vector of the last upsampling operation, and the input feature vector of each upsampling operation is the output feature vector of the last upsampling operation.

5. The audio decoding method according to claim 1, characterized in that: The number of channels of the intermediate feature vector is N times the number of channels of the first feature vector, and the feature recombination operation is used to recombine the N channels into one channel, where N is an integer.

6. The audio decoding method according to any one of claims 1 to 5, characterized in that: The reorganizing of the multiple channels of the intermediate feature vector comprises: The features corresponding to the positions in the multiple channels are arranged in sequence to be reorganized into one channel, and the number of features of the reorganized channel is equal to the total number of features of the multiple channels before the reorganization.

7. The audio decoding method according to claim 1, characterized in that: The number of upsampling operations and the convolution kernel size of the first convolution operation are determined according to the content complexity of the encoded data.

8. The audio decoding method according to claim 7, characterized in that: The number of times and the convolution kernel size are positively correlated with the content complexity, and the content complexity is determined according to at least one of the sampling rate, the number of channels, and the number of sound source types of the encoded data.

9. The audio decoding method according to claim 1, characterized in that: The upsampling operation also includes a third nonlinear mapping operation, the intermediate feature vector includes a first intermediate feature vector and a second intermediate feature vector, the third nonlinear mapping operation is used to perform nonlinear mapping on the first intermediate feature vector according to a preset activation function to obtain the second intermediate feature vector, and the feature recombination operation is used to recombine multiple channels of the second intermediate feature vector to obtain the second feature vector.

10. A decoding device, characterized in that: include: A non-linear transformation module, used for performing a non-linear transformation operation on the encoded data to obtain decoded data; The nonlinear transformation module includes an upsampling module, which includes a first convolution unit and a feature recombination unit. The first convolution unit is used to increase the number of channels of the input first feature vector to obtain an intermediate feature vector. The feature recombination unit is used to recombine multiple channels of the intermediate feature vector to obtain a second feature vector. The decoded data is determined based on the second feature vector.

11. The decoding device according to claim 10, characterized in that: The nonlinear transformation module comprises: A first convolution module, configured to perform a second convolution operation on the encoded data to extract the first feature vector; The upsampling module is used to perform the upsampling operation on the first feature vector to obtain the second feature vector; The second convolution module is used to perform a third convolution operation on the second feature vector to obtain the decoded data.

12. The decoding device according to claim 10, characterized in that: The upsampling modules include multiple ones, and the multiple upsampling modules are connected in sequence. The first feature vector is the input feature vector of the first upsampling module, the second feature vector is the output feature vector of the last upsampling module, and the output feature vector of the previous upsampling module is used as the input feature vector of the next upsampling module.

13. The decoding device according to claim 10, characterized in that: The upsampling module also includes an activation unit, the intermediate feature vector includes a first intermediate feature vector and a second intermediate feature vector, the activation unit is used to perform nonlinear mapping on the first intermediate feature vector according to a preset activation function to obtain the second intermediate feature vector, and the feature recombination unit is used to recombine multiple channels of the second intermediate feature vector to obtain the second feature vector.

14. The decoding device according to any one of claims 10 to 13, characterized in that: The up-sampling module further includes one or more residual units, and the residual units are used to perform a fourth convolution operation on the output feature vector of the feature recombination unit to obtain the output feature vector of the up-sampling module.

15. A computer device, characterized in that: include: Processor, memory; and A computer program, wherein the computer program is stored in the memory and executed by the processor, and the computer program includes instructions for executing the audio decoding method according to any one of claims 1 to 9.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the audio decoding method according to any one of claims 1 to 9 is implemented.

17. A non-volatile computer-readable storage medium containing a computer program, wherein when the computer program is executed by a processor, the processor executes the audio decoding method according to any one of claims 1 to 9.