Audio encoding and decoding method, device and equipment

By introducing different number and structures of feature processing modules into the audio encoding and codec network, the problem of low audio encoding and codec efficiency in the prior art is solved, more efficient audio data encoding and codec is achieved, and flexible network adjustment scheme is provided.

CN119943065APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311467514.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing end-to-end audio encoding and decoding schemes have low audio encoding and decoding efficiency due to the limitations of network structure.

Method used

By introducing feature processing modules with different numbers, structures and sampling multiples into the encoding network and the decoding network, it is ensured that the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and upsampling and feature transformation processing are performed on the decoding end to recover the audio data.

Benefits of technology

It reduces the complexity of audio encoding and decoding, improves the efficiency of audio encoding and decoding, and provides flexibility to adjust the network structure at the decoding end to improve performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943065A_ABST
    Figure CN119943065A_ABST
Patent Text Reader

Abstract

The invention provides an audio encoding and decoding method, device and equipment, which can be applied to the fields of artificial intelligence, audio encoding and decoding and the like, and comprises the following steps: obtaining a quantization result corresponding to audio data by decoding a code stream; and performing inverse quantization on the quantization result to obtain a reconstruction coding vector of the audio data, and finally performing at least one of up-sampling and feature transformation on the reconstruction coding vector through a decoding network to obtain a reconstruction value of the audio data. Wherein the total up-sampling multiple of the decoding network is consistent with the total down-sampling multiple of the coding network, and the coding feature processing modules included in the coding network are different from the decoding feature processing modules included in the decoding network in at least one of the number, the network structure and the sampling rate. Therefore, the network structures of the coding network and the decoding network can be respectively adjusted according to actual needs, so that the coding and decoding efficiency of the audio data is improved, and the coding and decoding complexity is effectively controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to an audio encoding and decoding method, apparatus, and device. Background Art

[0002] With the rapid development of deep learning technology, deep learning technology has been widely used in the processing technology of signals of different dimensions, such as audio, image and video. Taking audio signals as an example, in an end-to-end audio codec solution based on deep learning, the encoder maps the audio signal into a coding vector through the coding network, and further generates the corresponding binary code stream file through quantization technology. The decoding end obtains the quantization result by reading the binary code stream file, and dequantizes the quantization result through the dequantization technology to obtain the reconstructed coding vector, and then uses the reconstructed coding vector as the input of the decoding network to decode and obtain the final reconstructed audio signal.

[0003] However, the current end-to-end audio coding and decoding solutions are limited by the network structure, resulting in low audio coding and decoding efficiency. Summary of the invention

[0004] The present application provides an audio coding and decoding method, apparatus, and device, which can reduce the complexity of audio coding and decoding and improve the efficiency of audio coding and decoding.

[0005] In a first aspect, the present application provides an audio decoding method, comprising:

[0006] Decoding the bitstream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data, wherein the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through a coding network;

[0007] De-quantizing the quantization result to obtain a reconstructed coding vector of the audio data;

[0008] The reconstructed coding vector is subjected to at least one of upsampling and feature transformation processing through a decoding network to obtain a reconstructed value of the audio data, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples, the decoding feature processing module is used to perform at least one of upsampling and feature transformation processing on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation processing on the input feature vector.

[0009] In a second aspect, the present application provides an audio encoding method, comprising:

[0010] Get the audio data to be encoded;

[0011] The audio data is subjected to at least one of downsampling and feature transformation by a coding network to obtain a coding vector of the audio data, wherein the total downsampling multiple of the coding network is consistent with the total upsampling multiple of the decoding network, and the coding feature processing modules included in the coding network are different from the decoding feature processing modules included in the decoding network in at least one of the number, network structure, and sampling rate, the coding feature processing modules are used to perform at least one of downsampling and feature transformation on the input feature vector, and the decoding feature processing modules are used to perform at least one of upsampling and feature transformation on the input feature vector;

[0012] The coding vector of the audio data is quantized to obtain a quantization result, and the quantization result is encoded to obtain a code stream.

[0013] In a third aspect, the present application provides an audio decoding device, comprising:

[0014] A decoding unit, configured to decode a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data, wherein the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through a coding network;

[0015] A dequantization unit, used for dequantizing the quantization result to obtain a reconstructed coding vector of the audio data;

[0016] A reconstruction unit, used to perform at least one of upsampling and feature transformation processing on the reconstructed coding vector through a decoding network to obtain a reconstructed value of the audio data, wherein the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples, the decoding feature processing module is used to perform at least one of upsampling and feature transformation processing on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation processing on the input feature vector.

[0017] In some embodiments, the decoding network includes multiple decoding feature processing modules, at least one of the multiple decoding feature processing modules is used to upsample the input feature vector; the reconstruction unit is specifically used to upsample and feature transform the reconstructed coding vector through the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0018] In some embodiments, the reconstruction unit is specifically used to perform at least one of upsampling and feature transformation on the i-1th feature vector through the i-th decoding feature processing module to obtain the i-th feature vector, where i is a positive integer. If i is equal to 1, the i-1th feature vector is the reconstructed coding vector; if i is greater than 1, the i-1th feature vector is the feature vector output by the i-1th decoding feature processing module; perform at least one of upsampling and feature transformation on the i-th feature vector through the i+1th decoding feature processing module, repeat the execution, and determine the reconstructed value of the audio data based on the feature vector output by the last decoding feature processing module.

[0019] In some embodiments, the i-th decoding feature processing module includes R deconvolution layers, where R is a positive integer; a reconstruction unit, specifically used to perform at least one of upsampling and feature transformation on the i-1-th feature vector through the R deconvolution layers to obtain the i-th feature vector.

[0020] In some embodiments, at least one of the R deconvolution layers is an upsampling deconvolution layer; the reconstruction unit is specifically used to upsample and transform the i-1th feature vector through the R deconvolution layers to obtain the i-th feature vector.

[0021] In some embodiments, the i-th decoding feature processing module also includes an activation layer and a reconstruction unit, which is specifically used to upsample and linearly transform the i-1-th feature vector through the R deconvolution layers to obtain an upsampled feature vector; and perform a nonlinear transformation on the upsampled feature vector through the activation layer to obtain the i-th feature vector.

[0022] In some embodiments, one of the R deconvolution layers is an upsampling deconvolution layer, and the other layers are non-upsampling deconvolution layers.

[0023] In some embodiments, the decoding feature processing module includes at least one residual unit.

[0024] In some embodiments, the decoding network includes an input layer and a reconstruction unit, which is specifically used to process the reconstructed coding vector through the input layer to obtain a first feature vector; through the multiple decoding feature processing modules, the first feature vector is upsampled and feature transformed to obtain the reconstructed value of the audio data.

[0025] In some embodiments, the decoding network includes an output layer and a reconstruction unit, which is specifically used to upsample and perform feature transformation processing on the reconstructed coding vector through the multiple decoding feature processing modules to obtain a second feature vector; and process the second feature vector through the output layer to obtain a reconstructed value of the audio data.

[0026] In a fourth aspect, the present application provides an audio encoding device, including:

[0027] An acquisition unit, used for acquiring audio data to be encoded;

[0028] a transform unit, configured to perform at least one of downsampling and feature transformation on the audio data through an encoding network to obtain an encoding vector of the audio data, wherein a total downsampling multiple of the encoding network is consistent with a total upsampling multiple of the decoding network, and an encoding feature processing module included in the encoding network is different from a decoding feature processing module included in the decoding network in at least one of quantity, network structure, and sampling rate, the encoding feature processing module is configured to perform at least one of downsampling and feature transformation on an input feature vector, and the decoding feature processing module is configured to perform at least one of upsampling and feature transformation on an input feature vector;

[0029] The encoding unit is used to quantize the encoding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a code stream.

[0030] In some embodiments, the encoding network includes multiple encoding feature processing modules, at least one of the multiple encoding feature processing modules is used to downsample the input feature vector, and P is a positive integer; the transformation unit is specifically used to downsample and feature transform the audio data through the multiple encoding feature processing modules to obtain the encoding vector of the audio data, and P is a positive integer.

[0031] In some embodiments, the transformation unit is specifically used to perform at least one of downsampling and feature transformation on the i-1th feature vector through the i-th coding feature processing module to obtain the i-th feature vector, where i is a positive integer. If i is equal to 1, the i-1th feature vector is the audio data; if i is greater than 1, the i-1th feature vector is the feature vector output by the i-1th coding feature processing module; perform at least one of downsampling and feature transformation on the i-th feature vector through the i+1th coding feature processing module, repeat the execution, and determine the coding vector of the audio data based on the feature vector output by the last coding feature processing module.

[0032] In some embodiments, the i-th encoding feature processing module includes T convolutional layers, where T is a positive integer; a transformation unit, specifically used to perform at least one of downsampling and feature transformation on the i-1-th feature vector through the T convolutional layers to obtain the i-th feature vector.

[0033] In some embodiments, at least one of the T convolutional layers is a downsampling convolutional layer; the transformation unit is specifically used to downsample and linearly transform the i-1th feature vector through the T convolutional layers to obtain the i-th feature vector.

[0034] In some embodiments, the i-th encoding feature processing module also includes an activation layer and a transformation unit, which is specifically used to downsample and linearly transform the i-1-th feature vector through the T convolutional layers to obtain a downsampled feature vector; and perform a nonlinear transformation on the downsampled feature vector through the activation layer to obtain the i-th feature vector.

[0035] In some embodiments, one of the T convolutional layers is a downsampling convolutional layer, and the other layers are non-downsampling convolutional layers.

[0036] In some embodiments, the encoding feature processing module includes at least one residual unit.

[0037] In some embodiments, the encoding network includes an input layer and a transformation unit, which is specifically used to process the audio data through the input layer to obtain a first feature vector before downsampling and feature transformation processing the audio data through the multiple encoding feature processing modules to obtain the encoding vector of the audio data; and then downsampling and feature transformation processing the first feature vector through the multiple encoding feature processing modules to obtain the encoding vector.

[0038] In some embodiments, the encoding network includes an output layer and a transformation unit, which is specifically used to downsample and transform the audio data through the multiple encoding feature processing modules to obtain a second feature vector; and process the second feature vector through the output layer to obtain the encoding vector.

[0039] In a fifth aspect, an electronic device is provided, comprising a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method in any one of the first to second aspects or their implementations.

[0040] In a sixth aspect, a chip is provided for implementing the method in any one of the first to second aspects or their respective implementations. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method in any one of the first to second aspects or their respective implementations.

[0041] In a seventh aspect, a computer-readable storage medium is provided for storing a computer program, wherein the computer program enables a computer to execute the method in any one of the first to second aspects above or in each of their implementations.

[0042] In an eighth aspect, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions enable a computer to execute the method in any one of the first to second aspects above or in each of their implementations.

[0043] In a ninth aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or in each of their implementations.

[0044] In summary, the present application obtains the quantization result corresponding to the audio data by decoding the code stream, and the quantization result is obtained by quantizing the coding vector of the audio data, and the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through the coding network. Then, the decoding end dequantizes the quantization result to obtain the reconstructed coding vector of the audio data, and finally performs at least one of upsampling and feature transformation on the reconstructed coding vector through the decoding network to obtain the reconstructed value of the audio data. Among them, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the encoding feature processing module included in the encoding network is different from the decoding feature processing module included in the decoding network in at least one of the number, network structure, and sampling rate (for example, not mirror symmetric), wherein the encoding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector, and the decoding feature processing module is used to perform at least one of upsampling and feature transformation on the input feature vector. In this way, in some cases, in order to reduce the decoding complexity, the number of decoding feature processing modules in the decoding network can be reduced. In some cases, if the decoding performance needs to be improved, the number of decoding feature processing modules in the decoding network can be increased, etc., thereby improving the encoding and decoding efficiency of the audio data and effectively controlling the complexity of the encoding and decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 A schematic block diagram of an audio codec system according to an embodiment of the present application;

[0047] Figure 2 A schematic diagram of an end-to-end audio codec system based on deep learning involved in an embodiment of the present application;

[0048] Figure 3 A network structure block diagram of a codec constructed based on a convolutional neural network in one embodiment of the present application;

[0049] Figure 4 A schematic diagram of a residual-based vector quantizer according to an embodiment of the present application;

[0050] Figure 5 A flowchart of an audio decoding method provided in an embodiment of the present application;

[0051] Figure 6 It is a schematic diagram of dequantization;

[0052] Figures 7 to 10 A schematic diagram of a decoding network involved in an embodiment of the present application;

[0053] Figures 11A to 11D A schematic diagram of a decoding feature processing module involved in an embodiment of the present application;

[0054] FIG. 12A to FIG. 12C Another schematic diagram of a decoding network involved in an embodiment of the present application;

[0055] Fig.13 A schematic diagram of a flow chart of an audio encoding method provided in an embodiment of the present application;

[0056] Figures 14 to 17 A schematic diagram of a coding network involved in an embodiment of the present application;

[0057] Figures 18A to 18D A schematic diagram of a coding feature processing module involved in an embodiment of the present application;

[0058] FIG. 19A to FIG. 19C Another schematic diagram of a coding network involved in an embodiment of the present application;

[0059] Fig. 20 It is a quantitative schematic diagram;

[0060] Fig.21 is a schematic block diagram of an audio decoding device provided by an embodiment of the present application;

[0061] Fig. 22 is a schematic block diagram of an audio encoding device provided by an embodiment of the present application;

[0062] Fig.23 It is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0064] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In an embodiment of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units does not have to be limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "multiple" refers to two or more than two.

[0065] The technical solution proposed in this application can be applied to technical fields such as artificial intelligence and audio coding and decoding, and can be used to reduce the complexity of audio coding and decoding while ensuring the performance of audio and video coding and decoding, thereby improving the efficiency of audio coding and decoding.

[0066] The following is an introduction to the relevant concepts involved in the embodiments of the present application.

[0067] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0068] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0070] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0071] The embodiments of this application mainly introduce the application of artificial intelligence technology in audio coding and decoding technology.

[0072] Audio encoding and decoding: The audio encoding process is to compress the audio into smaller data, and the decoding process is to restore the smaller data to audio. The encoded smaller data is used for network transmission and occupies less bandwidth.

[0073] Audio sampling rate: The audio sampling rate describes the number of data contained in a unit of time (1 second). For example, an 8k sampling rate contains 8000 sampling points, and each sampling point corresponds to a short integer.

[0074] Codebook: A collection of multiple vectors. The encoder and decoder both store the same codebook.

[0075] Quantization: Find the closest vector in the codebook for the input vector, return it as a replacement for the input vector, and return the corresponding codebook index position.

[0076] Quantizer: The quantizer is responsible for quantization and updating the vectors in the codebook.

[0077] Audio frame: Indicates the minimum duration of voice for a single transmission in the network.

[0078] Short Time Fourier Transform: STFT. It divides a long signal into several shorter signals of equal length, and then calculates the Fourier transform of each shorter segment. It is usually used to describe the changes in the frequency domain and time domain, and is an important tool in time-frequency analysis.

[0079] The audio coding method provided in the embodiments of the present application can be applied to the field of audio coding, the field of hardware audio coding, the field of dedicated circuit video coding, the field of real-time audio coding, etc. For example, the scheme of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard. Alternatively, the scheme of the present application can be combined with other exclusive or industry standards. It should be understood that the technology of the present application is not limited to any specific coding standard or technology.

[0080] The audio coding and decoding method provided in the embodiments of the present application can be applied to any end-to-end audio coding and decoding solution based on deep learning.

[0081] To facilitate understanding, first combine Figure 1 The audio codec system involved in the embodiment of the present application is introduced.

[0082] Figure 1 This is a schematic block diagram of an audio codec system involved in an embodiment of the present application. It should be noted that: Figure 1 This is just an example. The audio codec system of the embodiment of the present application includes but is not limited to Figure 1 As shown. Figure 1As shown, the audio codec system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compression) the audio data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded audio data.

[0083] The encoding device 110 of the embodiment of the present application can be understood as a device with an audio encoding function, and the decoding device 120 can be understood as a device with an audio decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, playback devices, digital media players, audio game consoles, car computers, etc.

[0084] In some embodiments, the encoding device 110 may transmit the encoded audio data (eg, a code stream) to the decoding device 120 via the channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded audio data from the encoding device 110 to the decoding device 120.

[0085] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded audio data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded audio data according to the communication standard and transmit the modulated audio data to the decoding device 120. The communication medium includes a wireless communication medium, such as a radio frequency spectrum, and optionally, the communication medium may also include a wired communication medium, such as one or more physical transmission lines.

[0086] In another example, the channel 130 includes a storage medium, which can store the audio data encoded by the encoding device 110. The storage medium includes a variety of locally accessible data storage media, such as optical disks, DVDs, flash memories, etc. In this example, the decoding device 120 can obtain the encoded audio data from the storage medium.

[0087] In another example, the channel 130 may include a storage server that can store the audio data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded audio data from the storage server. Optionally, the storage server can store the encoded audio data and can transmit the encoded audio data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.

[0088] In some embodiments, the encoding device 110 includes an audio encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.

[0089] In some embodiments, the encoding device 110 may further include an audio source 111 in addition to the audio encoder 112 and the input interface 113 .

[0090] The audio source 111 may include at least one of an audio acquisition device (eg, a microphone), an audio archive, an audio input interface, and a computer voice system, wherein the audio input interface is used to receive audio data from an audio content provider, and the computer voice system is used to generate audio data.

[0091] The audio encoder 112 encodes the audio data from the audio source 111 to generate a bitstream. The bitstream contains the encoding information of the audio data in the form of a bitstream. The encoding information may include the encoded audio data and associated data. The associated data may include quantization parameters and other syntax structures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.

[0092] The audio encoder 112 transmits the encoded audio data directly to the decoding device 120 via the output interface 113. The encoded audio data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.

[0093] In some embodiments, the decoding device 120 includes an input interface 121 and an audio decoder 122 .

[0094] In some embodiments, the decoding device 120 may include a playback device 123 in addition to the input interface 121 and the audio decoder 122 .

[0095] The input interface 121 includes a receiver and / or a modem. The input interface 121 can receive the encoded audio data through the channel 130 .

[0096] The audio decoder 122 is used to decode the encoded audio data to obtain decoded audio data, and transmit the decoded audio data to the playback device 123 .

[0097] The playback device 123 plays the decoded audio data. The playback device 123 may be integrated with the decoding device 120 or may be external to the decoding device 120. The playback device 123 may include a variety of playback devices.

[0098] also, Figure 1 This is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1For example, the technology of the present application can also be applied to single-sided audio encoding or single-sided audio decoding.

[0099] Figure 2 Schematic diagram of an end-to-end audio codec system based on deep learning involved in an embodiment of the present application. Figure 2 As shown, the audio codec system of the embodiment of the present application includes: an encoding network 210, a quantization module 211, an inverse quantization module 212 and a decoding network 213.

[0100] During encoding, the encoding end (also called the transmitting end) will first input the input audio data into the encoding network 210 for nonlinear transformation to obtain the encoding vector (also called embedded sequence or hidden variable, etc.) of the input audio data. Then, the encoding vector of the audio data is quantized by the quantization module 211 to obtain the quantization result of the encoding vector. For example, a residual-based vector quantizer is used to select the corresponding quantization parameter according to the target bit rate. Finally, the quantized encoding vector is encoded and converted into a binary code stream.

[0101] During decoding, the decoding end (also called the receiving end) first recovers the quantization result of the coding vector from the bit stream, and then further recovers the coding vector through the inverse quantization module 212, and inputs it into the decoding network 213 for nonlinear transformation to obtain reconstructed audio data.

[0102] Figure 3 This is a block diagram of the network structure of a codec built based on a convolutional neural network in one embodiment of the present application.

[0103] like Figure 3 As shown, the network structure of the codec includes an encoding network 310 and a decoding network 320, wherein the encoding network 310 can be implemented as software as shown in FIG. Figure 1 The audio and video encoding device 110 and the decoding network 320 shown in FIG. 3 can be implemented as software as shown in FIG. Figure 1 The audio and video decoding device 120 is shown. In some embodiments, the encoding network 310 is also referred to as an encoder 310 , and the decoding network 320 is also referred to as a decoder 320 .

[0104] At the data transmission end, the audio data may be encoded and compressed through the encoding network 310. In an embodiment of the present application, the encoding network 310 may include an input layer 311, one or more encoding modules 312 and an output layer 313.

[0105] Exemplarily, the input layer 311 and the output layer 313 may be convolutional layers constructed based on a one-dimensional convolution kernel, and multiple (e.g., 4) encoding modules (EncoderBlock) 312 are sequentially connected between the input layer 311 and the output layer 313. Each encoding module 312 includes multiple residual (ResidualUnit) modules, and each residual module includes multiple convolutional layers.

[0106] For example, at the input stage of the encoder, the original audio data to be encoded is sampled to obtain a vector with c channels and w dimensions; the vector is input to the input layer 311, and after convolution processing, a feature vector with 32c channels and w dimensions can be obtained. In some optional implementations, in order to improve encoding efficiency, the encoding network 310 can encode a batch of audio vectors at the same time.

[0107] In the downsampling stage of the encoder, the first encoding module reduces the vector dimension to 1 / 2 and increases the number of channels by 2 times, obtaining a feature vector with 64 channels and 1 / 2w dimension; the second encoding module reduces the vector dimension to 1 / 4 and increases the number of channels by 2 times, obtaining a feature vector with 128 channels and 1 / 8w dimension; the third encoding module reduces the vector dimension to 1 / 5 and increases the number of channels by 2 times, obtaining a feature vector with 256 channels and 1 / 40w dimension; the fourth encoding module reduces the vector dimension to 1 / 8 and increases the number of channels by 2 times, obtaining a feature vector with 512 channels and 1 / 320w dimension.

[0108] In the output stage of the encoder, the output layer 313 performs convolution processing on the feature vector with a channel number of 512c and a dimension of 1 / 320w to obtain a coding vector with a channel number of 1 and a dimension of K.

[0109] The coding vector is input to the quantizer 330, and the codebook index corresponding to the coding vector can be queried in the codebook, and the codebook index is encoded to obtain a binary code stream, which is then sent to the data receiving end.

[0110] The data receiving end decodes the received binary code stream to obtain a codebook index, and performs inverse quantization based on the codebook index to obtain a reconstructed coding vector. Finally, the reconstructed coding vector is decoded through a decoding network 320 to obtain restored audio data.

[0111] In one embodiment of the present application, the decoding network 320 may include an input layer 321, one or more decoding modules 322, and an output layer 323. Each decoding module 322 includes a plurality of residual units, and each residual unit includes a plurality of convolutional layers.

[0112] After the data receiving end decodes the code stream to obtain the codebook index, the data receiving end may first query the codebook vector corresponding to the codebook index in the codebook through the quantizer 320, and then obtain the encoding vector reconstructed by the audio data based on the codebook vector. For example, the reconstructed encoding vector may be a vector with a channel number of 1 and a dimension of K. In some optional implementations, in order to improve decoding efficiency, the data receiving end may decode a batch of codebook vectors at the same time.

[0113] In the input stage of the decoder, the reconstructed coding vector is input to the input layer 321, and after convolution processing, a feature vector with a channel number of 512c and a dimension of 1 / 320w can be obtained.

[0114] In the decoding stage of the decoder, the first decoding module increases the vector dimension to 8 times and reduces the number of channels by 2 times, obtaining a feature vector with 256 channels and 1 / 40w dimension; the second decoding module increases the vector dimension to 5 times and reduces the number of channels by 2 times, obtaining a feature vector with 128 channels and 1 / 8w dimension; the third decoding module increases the vector dimension to 4 times and reduces the number of channels by 2 times, obtaining a feature vector with 64 channels and 1 / 2w dimension; the fourth decoding module increases the vector dimension to 2 times and reduces the number of channels by 2 times, obtaining a feature vector with 32 channels and w dimension.

[0115] In the output stage of the decoder, the output layer 323 performs convolution processing on the feature vector with 32 channels and a dimension of w, and restores the reconstructed audio data with 1 channel and a dimension of w.

[0116] In some embodiments, in order to improve the audio coding effect, when quantizing the coding vector of the audio data, a residual-based vector quantizer is used, that is, the above-mentioned Figure 3 The quantizer 330 in is a residual-based vector quantizer.

[0117] In one example, if Figure 4 As shown in the figure, the residual-based vector quantizer contains N vector quantization layers (VQ), each of which maintains a codebook of length L. The length of each codeword is M, that is, a codebook includes L codebook vectors, and the length of each codebook vector is M. Its specific operation is: the encoding vector x∈R S×D As input (S is the number of frames, D is the dimension of the feature). Assume that the feature of the i-th frame of x is recorded as y 0 ∈R 1×D , will first pass through the first vector quantization layer, and query the codebook corresponding to the first vector quantization layer with y 0 ∈R 1×DThe codebook vector closest to the nearest distance obtains the corresponding vector quantization result And the corresponding codebook index index_i0, and calculate the quantized residual As the input of the next vector quantization layer. Then iterate N-1 times the same quantization and residual calculation operations until all N vector quantization layers are traversed, and output the codebook index index_ij corresponding to each layer, j = 1, ... N, and write it into the bit stream after binarization (such as 7 can be represented as 0 ... 111, the total number of bits is log2 (L)). The quantization process ends after processing all frames of x, and outputs the bit stream to be transmitted.

[0118] The decoder decodes the bitstream and reads the corresponding codebook index for each frame. ij ,i=1,…S,j=1,…N, recover the codebook vector from the codebook. Specifically, for the i-th frame, assume that the information recovered at each layer is Then sum the results of each layer to get the reconstructed coding vector of the i-th frame

[0119] By the above Figure 3 It can be seen that in the current end-to-end audio codec solution, the network structures of the encoding network and the decoding network are strictly mirror-symmetric, that is, the decoding network is the mirror-symmetric of the encoding network. When the network structure of the encoding network is complex, the corresponding network structure of the decoding network is also complex, which increases the complexity of the encoding and decoding of audio data exponentially, making the audio encoding and decoding efficiency low. In addition, since the current decoding network is determined by the network structure of the encoding network, it is impossible to complete the operation of improving the audio decoding effect at the decoding end, resulting in unsatisfactory audio decoding effect.

[0120] In order to solve the above technical problems, the embodiment of the present application proposes a codec network structure, and the network models of the encoding network and the decoding network are relatively independent. Specifically, the decoding end decodes the code stream to obtain the quantization result corresponding to the audio data, and the quantization result is obtained by quantizing the encoding vector of the audio data, and the encoding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through the encoding network. Then, the quantization result is inversely quantized to obtain the reconstructed encoding vector of the audio data, and finally the reconstructed encoding vector is subjected to at least one of upsampling and feature transformation through the decoding network to obtain the reconstructed value of the audio data. Among them, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the encoding feature processing module included in the encoding network is different from the decoding feature processing module included in the decoding network in at least one of the number, network structure, and sampling rate, wherein the encoding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector, and the decoding feature processing module is used to perform at least one of upsampling and feature transformation on the input feature vector. In this way, in some cases, in order to reduce the decoding complexity, the number of decoding feature processing modules in the decoding network can be reduced. In some cases, if the decoding performance needs to be improved, the number of decoding feature processing modules in the decoding network can be increased, etc., thereby improving the encoding and decoding efficiency of the audio data and effectively controlling the complexity of the encoding and decoding.

[0121] The technical solutions of the embodiments of the present application are described in detail below through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0122] First, taking the decoding end as an example, the audio decoding method provided in the embodiment of the present application is introduced.

[0123] Figure 5 1 is a flowchart of an audio decoding method provided in an embodiment of the present application. The execution subject of the embodiment of the present application is a device with an audio decoding function, such as an audio decoding device. In some embodiments, the audio decoding device can be Figure 1 For the convenience of description, the present application embodiment is described by taking the execution subject as a decoding device as an example.

[0124] like Figure 5 As shown, the audio decoding method of the embodiment of the present application includes:

[0125] S101, decoding a bit stream to obtain a quantization result corresponding to the audio data.

[0126] The quantization result is obtained by quantizing the coding vector of the audio data, and the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through a coding network.

[0127] In the embodiment of the present application, the audio data to be decoded may be a segment of audio data of any length.

[0128] In one example, the audio data to be decoded includes one or more audio frames. In one example, the audio data may also include an incomplete audio frame, such as 1 / 2 of an audio frame or 1 / 4 of an audio frame.

[0129] In the embodiment of the present application, the audio frame can be understood as a data segment with a specified time length obtained after framing and windowing the original audio data.

[0130] The embodiment of the present application does not limit the specific method of obtaining the original audio data.

[0131] In some examples, the original audio data may be speech captured by a terminal.

[0132] In some examples, the original audio data may be a sound signal collected in an Internet voice call or video call scenario.

[0133] In some examples, the original audio data may be a sound signal collected in a live broadcast scenario, a sound signal collected in an online singing scenario, or a sound signal collected in a voice broadcast scenario.

[0134] In some examples, the original audio data may be audio data acquired from a storage resource, such as stored voice, music, video, etc.

[0135] In a possible implementation, when dividing the original audio data into audio frames, the embodiment of the present application may set a preset time length for division, for example, dividing every 10 ms of original audio in the original audio data into an audio frame.

[0136] In order to store and transmit audio data over long distances, the acquired original audio data needs to be encoded to reduce the size of the audio data, thereby reducing the storage space of the audio data or reducing the traffic bandwidth consumed by long-distance transmission.

[0137] In the process of audio encoding, the acquired audio data is firstly subjected to nonlinear transformation, such as upsampling and feature transformation of the audio data through a coding network to obtain the coding vector of the audio data. Then, the coding vector of the audio data is quantized by a quantizer, and finally the quantized coding vector is encoded to obtain a bit stream.

[0138] After the corresponding decoding end obtains the bit stream, it decodes the bit stream to obtain the quantization result corresponding to the audio data.

[0139] S102: Dequantize the quantization result to obtain a reconstructed coding vector of the audio data.

[0140] The decoding end decodes the received bit stream, obtains the quantization result corresponding to the audio data, and then dequantizes the quantization result to obtain the reconstructed coding vector of the audio.

[0141] In some embodiments, if the quantizer used in the embodiment of the present application is Figure 4 When the residual-based vector quantizer is used as shown, the quantization result corresponding to the audio data includes the codebook index corresponding to each quantization layer of the quantizer and the residual vector corresponding to the last quantization layer. At this time, when the decoder performs inverse quantization on the quantization result, it first searches the corresponding codebook based on the codebook index corresponding to each quantization layer to obtain the corresponding vector, and then accumulates the vectors obtained by the search with the residual vector corresponding to the last quantization layer to obtain the reconstructed coding vector of the audio data.

[0142] For example, Figure 6 As shown, assuming that the residual-based vector quantizer includes three quantization layers, the decoder decodes the bitstream and obtains the residual vector corresponding to the last quantization layer of the quantizer. And the codebook index index0, index1 and index2 corresponding to each of the three quantization layers. In this way, the decoder first queries the codebook vector corresponding to the codebook index index2 in the codebook corresponding to the third quantization layer, and then determines the codebook vector as the reconstructed value y3' of the output vector of the third vector layer, and compares the vector y3' with the residual vector The sum of the values ​​is determined as the residual vector The reconstruction value Next, the codebook vector corresponding to the codebook index index1 is queried in the codebook corresponding to the second quantization layer, and the codebook vector is determined as the reconstructed value y2' of the output vector of the second vector layer, and the output vector y2' is compared with the reconstructed value of the residual vector The sum of the values ​​is determined as the residual vector The reconstruction value Next, the codebook vector corresponding to the codebook index index0 is queried in the codebook corresponding to the first quantization layer, and the codebook vector is determined as the reconstructed value y1' of the output vector of the first vector layer, and the vector y1' is compared with the reconstructed value of the residual vector The sum of the values ​​is determined as the reconstructed coding vector y0′ of the audio data.

[0143] In some embodiments, the above Figure 6 The number of quantization layers included in the residual-based vector quantizer shown can be adaptively selected based on the statistical value of effective audio information of the audio data to be processed.

[0144] After the decoding end obtains the reconstructed coding vector of the audio data based on the above steps, step S103 is executed.

[0145] S103. Perform at least one of upsampling and feature transformation on the reconstructed coding vector through a decoding network to obtain a reconstructed value of the audio data.

[0146] The total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing module included in the decoding network is different from the encoding feature processing module included in the encoding network in at least one of the number, network structure, and sampling rate. The difference here includes asymmetry, for example, the decoding feature processing module and the encoding feature processing module are not mirror-image objects, and the upsampling multiple of the decoding feature processing module and the downsampling multiple of the encoding feature processing module are not mirror-image symmetric.

[0147] The coding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector. For example, at least one coding feature processing module only has a feature transformation function but does not have a downsampling function. For another example, at least one coding feature processing module has both a feature transformation function and a downsampling function.

[0148] The decoding feature processing module is used to perform at least one of upsampling and feature transformation on the input feature vector. For example, at least one decoding feature processing module only has a feature transformation function and does not have an upsampling function. For another example, at least one decoding feature processing module has both a feature transformation function and an upsampling function.

[0149] In the embodiment of the present application, feature transformation can be understood as converting feature information from one set of data to another set of data. The feature transformation in the embodiment of the present application includes at least one of linear transformation (such as convolution operation) and nonlinear transformation (such as activation operation).

[0150] In an embodiment of the present application, after the decoding end obtains the reconstructed coding vector of the audio data based on the above steps, the reconstructed coding vector is input into the decoding network for at least one of upsampling and feature transformation to obtain the reconstructed value of the audio data.

[0151] The current decoding network structure is a mirror image of the encoding network structure, for example Figure 3As shown, the encoding network includes 1 input layer, 4 encoding feature processing modules and 1 output layer, and the corresponding decoding network also includes 1 input layer, 4 encoding feature processing modules and 1 output layer, and the decoding feature processing module in the decoding network is consistent with the network structure of the encoding feature processing module in the encoding network. In this way, when the network structure of the encoding network is complex, the network structure of the corresponding decoding network is also complex. For example, when the encoding network includes 8 encoding feature processing modules, the decoding network also includes 8 decoding feature processing modules, which increases the complexity of the encoding and decoding of the audio data exponentially, making the audio encoding and decoding efficiency low. In addition, since the current decoding network is determined by the network structure of the encoding network, it is impossible to complete the operation of improving the audio decoding effect at the decoding end, resulting in unsatisfactory audio decoding effect.

[0152] In order to solve the above technical problems, in the embodiment of the present application, the network models of the encoding network and the decoding network are independent, so that the network structures of the encoding network and the decoding network can be adjusted according to actual needs. For example, in order to reduce the decoding complexity, the decoding feature processing module in the decoding network can be reduced. For another example, if the decoding performance needs to be improved, the number of decoding feature processing modules in the decoding network can be increased, etc., thereby improving the encoding and decoding efficiency of the audio data and effectively controlling the complexity of the encoding and decoding.

[0153] That is to say, in the embodiment of the present application, there is no restriction on whether the encoding network and the decoding network are mirror-symmetric, as long as the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network.

[0154] The network structure of the decoding network in the embodiment of the present application may include at least the following situations or a combination of the following situations:

[0155] In the first case, the encoding network includes one input layer, while the decoding network does not include an input layer, or the input layer included in the decoding network is inconsistent with the network structure of the input layer included in the encoding network. For example, the input layer included in the encoding network includes one convolutional layer, while the input layer of the decoding network of the embodiment of the present application may include two or more deconvolutional layers. For another example, the input layer included in the encoding network includes one convolutional layer, while the input layer of the decoding network of the embodiment of the present application may include a residual unit, etc.

[0156] In the second case, the encoding network includes one output layer, while the decoding network does not include an output layer, or the output layer included in the decoding network is inconsistent with the network structure of the output layer included in the encoding network. For example, the output layer included in the encoding network includes one convolutional layer, while the output layer of the decoding network of the embodiment of the present application may include two or more deconvolutional layers. For another example, the output layer included in the encoding network includes one convolutional layer, while the output layer of the decoding network of the embodiment of the present application may include a residual unit, etc.

[0157] In the third case, the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples.

[0158] In an example of the third case, the number of decoding feature processing modules included in the decoding network of the embodiment of the present application is different from the number of encoding feature processing modules included in the encoding network. For example, the encoding network includes n encoding feature processing modules, and the decoding network includes m decoding feature processing modules, and m is not equal to n.

[0159] In another example of the third case, the network structure of the decoding feature processing module included in the decoding network of the embodiment of the present application is not mirror-symmetrical with the network of the encoding feature processing module included in the encoding network. Exemplarily, the encoding feature processing module of the embodiment of the present application includes multiple residual units, and the decoding feature processing module of the present application includes multiple deconvolution layers. Exemplarily, the hierarchy of the decoding feature processing module of the embodiment of the present application is not mirror-symmetrical with the hierarchy of the encoding feature processing module, for example, the encoding feature processing module includes 4 convolution layers, and the decoding feature processing module includes 3 deconvolution layers.

[0160] In another example of the third case, the network structure of the decoding feature processing module included in the decoding network of the embodiment of the present application is not mirror-symmetric with the sampling multiple of the encoding feature processing module included in the encoding network.

[0161] In one example, the number of decoding feature processing modules included in the decoding network is the same as the number of encoding feature processing modules included in the encoding network, but the upsampling multiples of the decoding feature processing modules included in the decoding network are not mirror-symmetric with the downsampling multiples of the encoding feature processing modules included in the encoding network. For example, the encoding network includes 4 encoding feature processing modules, and the downsampling multiples of the 4 encoding feature processing modules are [8, 6, 4, 2] respectively, and the upsampling multiples of the 4 decoding feature processing modules included in the decoding network can be any group of upsampling multiples except [2, 4, 6, 8]. For example, the upsampling multiples of the 4 decoding feature processing modules included in the decoding network of the embodiment of the present application can be [2, 6, 4, 8], [2, 8, 4, 6], etc.

[0162] In one example, the number of decoding feature processing modules included in the decoding network is different from the number of encoding feature processing modules included in the encoding network, and the upsampling multiples of the decoding feature processing modules included in the decoding network are not mirror-symmetric with the downsampling multiples of the encoding feature processing modules included in the encoding network. For example, the encoding network includes 4 encoding feature processing modules, and the downsampling multiples of these 4 encoding feature processing modules are [8, 6, 4, 2] respectively, while the upsampling multiples of the decoding network including 3 decoding feature processing modules can be [8, 6, 8], etc.

[0163] In an embodiment of the present application, the decoding network is used to perform at least one of upsampling and feature transformation on the reconstructed coding vector of the audio data to restore the audio data.

[0164] That is to say, the decoding network of the embodiment of the present application has upsampling and feature transformation functions.

[0165] The feature transformation includes linear transformation and / or nonlinear transformation.

[0166] In the embodiment of the present application, the decoding network has an upsampling function, which can be understood as at least one module included in the decoding network having an upsampling function. Similarly, the decoding network has a feature transformation function, which can be understood as at least one module included in the decoding network having a feature transformation function.

[0167] In some embodiments, some modules in the decoding network have upsampling functions, and some modules have feature transformation functions.

[0168] In some embodiments, some modules in the decoding network have an upsampling function, some modules have a feature conversion function, and some modules have both an upsampling function and a feature conversion function. Alternatively, some modules in the decoding network have an upsampling function, and some modules have both an upsampling function and a feature conversion function. Alternatively, some modules in the decoding network have a feature conversion function, and some modules have both an upsampling function and a feature conversion function.

[0169] That is to say, in the embodiment of the present application, there is no restriction on which modules in the decoding network have the upsampling function and which modules have the feature transformation function, and they can be set according to actual needs.

[0170] In some embodiments, Figure 7 As shown, the decoding network proposed in the embodiment of the present application includes multiple decoding feature processing modules.

[0171] The number of decoding feature processing modules included in the decoding network of the embodiment of the present application may be the same as or different from the number of encoding feature processing modules included in the encoding network, and the embodiment of the present application does not impose any limitation on this.

[0172] In one example, in order to reduce decoding complexity, the number of decoding feature processing modules in the decoding network can be reduced. For example, the number of decoding feature processing modules included in the decoding network is set to be smaller than the number of encoding feature processing modules included in the encoding network.

[0173] In one example, if the decoding performance needs to be improved, the number of decoding feature processing modules in the decoding network can be increased. For example, the number of decoding feature processing modules included in the decoding network is set to be greater than the number of encoding feature processing modules included in the encoding network.

[0174] The decoding feature processing module of the embodiment of the present application is used to perform at least one of upsampling and feature transformation on the input feature vector.

[0175] In some embodiments, each of the multiple decoding feature processing modules included in the decoding network does not have an upsampling function, but completes a feature transformation function. At this time, the decoding network may also include other upsampling modules, such as at least one upsampling deconvolution layer. Exemplarily, some or all of the upsampling deconvolution layers in the at least one upsampling deconvolution layer may be continuously or discontinuously interspersed between at least two encoding feature processing modules in the multiple decoding feature processing modules. Alternatively, the three upsampling deconvolution layers may be arranged before the multiple decoding feature processing modules, or after the multiple decoding feature processing modules. For example, assuming that the decoding network includes three upsampling deconvolution layers, the three upsampling deconvolution layers may be continuously interspersed between two decoding feature processing modules in the multiple decoding feature processing modules. Or the three upsampling deconvolution layers may be interspersed before or after different decoding feature processing modules in the multiple decoding feature processing modules, for example, one or more upsampling deconvolution layers may be interspersed before or after a decoding feature processing module. Alternatively, the three upsampling deconvolution layers are arranged before the multiple decoding feature processing modules, or the three upsampling deconvolution layers are arranged after the multiple decoding feature processing modules. The embodiment of the present application does not limit the specific distribution of the decoding feature processing modules and the upsampling modules in the decoding network. At this time, the decoding end can input the reconstructed coding vector of the audio data determined above into the decoding network, and the multiple decoding feature processing modules in the decoding network implement feature conversion processing, and the decoding feature processing module in the decoding network implements upsampling processing, and finally the decoding network outputs the reconstructed value of the audio data.

[0176] In some embodiments, at least one decoding feature processing module among the multiple decoding feature processing modules included in the decoding network is used to perform upsampling processing on the input feature vector. In this case, the above S103 includes the following step S103-A:

[0177] S103-A. Upsample and transform the reconstructed coding vector through multiple decoding feature processing modules to obtain a reconstructed value of the audio data.

[0178] In this implementation, at least one of the multiple decoding feature processing modules included in the decoding network has an upsampling function. In this way, the decoding end can perform upsampling and feature transformation processing on the reconstructed coding vector of the above-determined audio data through the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0179] In this implementation, the function of each decoding feature processing module in the multiple decoding feature processing modules may be the same or different.

[0180] In one example, each of the multiple decoding feature processing modules has upsampling and feature transformation processing functions. At this time, the decoding end can perform upsampling and feature transformation processing on the reconstructed coding vector of the above-determined audio data through each of the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0181] In one example, some of the above-mentioned multiple decoding feature processing modules have upsampling and feature transformation processing functions, and some of the decoding feature processing modules only have feature transformation functions but do not have upsampling functions. At this time, the decoding end can perform upsampling and feature transformation processing on the reconstructed coding vector of the audio data through the decoding feature processing modules with upsampling and feature transformation functions among the multiple decoding feature processing modules, and perform feature transformation processing on the reconstructed coding vector of the audio data through the decoding feature processing modules with only feature transformation functions but not upsampling functions among the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0182] The embodiment of the present application does not limit the specific distribution structure of the above-mentioned multiple decoding feature processing modules.

[0183] In some embodiments, the above-mentioned multiple decoding feature processing modules can be connected in parallel. As an example, Figure 8As shown, the decoding network of the embodiment of the present application includes an input layer, three decoding feature processing modules and an output layer, wherein the three decoding feature processing modules are connected in parallel. In this way, the decoding end processes the reconstructed coding vector of the audio data through the input layer to obtain a feature vector 1, and then, the feature vector 1 is respectively input into the three decoding feature processing modules for processing, such as upsampling and / or feature transformation processing, to obtain a feature vector output by each of the three decoding feature processing modules. Then, the feature vectors respectively output by the three decoding feature processing modules are fused, such as by addition or multiplication, to obtain a fused feature vector 2, and finally, the fused feature vector 2 is input into the output layer for processing to obtain the reconstructed value of the audio data. Optionally, the decoding network may not include Figure 8 The input layer and / or output layer shown.

[0184] In some embodiments, some of the above-mentioned multiple decoding feature processing modules are connected in parallel, and some of the decoding feature processing modules are connected in series. Fig. 9 As shown, the decoding network of the embodiment of the present application includes an input layer, three decoding feature processing modules and an output layer, wherein two of the three decoding feature processing modules are connected in parallel and then connected in series with another decoding feature processing module. In this way, the decoding end processes the reconstructed coding vector of the audio data through the input layer to obtain a feature vector 1, and then, the feature vector 1 is respectively input into the two decoding feature processing modules connected in parallel for processing, such as upsampling and / or feature transformation processing, to obtain a feature vector output by each of the two decoding feature processing modules connected in parallel. Then, the feature vectors respectively output by the two decoding feature processing modules connected in parallel are fused, such as addition or multiplication, to obtain a fused feature vector 3, and the fused feature vector 3 is input into the decoding feature processing modules connected in series for upsampling and / or feature transformation processing to obtain a feature vector 4, and finally, the feature vector 4 is input into the output layer for processing to obtain the reconstructed value of the audio data. Optionally, the decoding network may not include Figure 8 The input layer and / or output layer shown.

[0185] In some embodiments, the above-mentioned multiple decoding feature processing modules are connected in series. In this case, the decoding end S103-A includes the following steps S103-A1 and S103-A2:

[0186] S103-A1, performing at least one of upsampling and feature transformation on the i-1th feature vector through the i-th decoding feature processing module to obtain the i-th feature vector, where i is a positive integer. If i is equal to 1, the i-1th feature vector is a reconstructed coding vector. If i is greater than 1, the i-1th feature vector is a feature vector output by the i-1th decoding feature processing module;

[0187] S103-A2, perform at least one of upsampling and feature transformation on the i-th feature vector through the i+1-th decoding feature processing module, and repeat the process to determine the reconstructed value of the audio data based on the feature vector output by the last decoding feature processing module.

[0188] As an example, Fig.10 As shown, the decoding network of the embodiment of the present application includes an input layer, three decoding feature processing modules and an output layer, wherein the three decoding feature processing modules are connected in series. In this way, the decoding end processes the reconstructed coding vector of the audio data through the input layer to obtain feature vector 1, and then, the feature vector 1 is respectively input into the first decoding feature processing module for processing, such as upsampling and / or feature transformation processing, to obtain feature vector 5. Then, feature vector 5 is input into the second decoding feature processing module for upsampling and / or feature transformation processing to obtain feature vector 6. Then, feature vector 6 is input into the third decoding feature processing module for upsampling and / or feature transformation processing to obtain feature vector 7. Finally, feature vector 7 is input into the output layer for processing to obtain the reconstructed value of the audio data. Optionally, the decoding network may not include Figure 8 The input layer and / or output layer shown.

[0189] In one example, all the decoding feature processing modules of the above-mentioned multiple decoding feature processing modules have upsampling and feature transformation processing functions. At this time, the decoding end can upsample and feature transform the reconstructed coding vector of the audio data determined above through each of the multiple decoding feature processing modules to obtain the reconstructed value of the audio data. For example, if multiple decoding feature processing modules are connected in series, the decoding end inputs the reconstructed coding vector of the audio data into the first decoding feature processing module among the multiple decoding feature processing modules for upsampling and feature transformation processing, and obtains the upsampled and feature transformed feature vector output by the first decoding feature processing module. Then, the feature vector output by the first decoding feature processing module is input into the second decoding feature processing module for upsampling and feature transformation processing, and so on, to obtain the upsampled and feature transformed feature vector output by the last decoding feature processing module, and then the reconstructed value of the audio data is obtained based on the feature vector.

[0190] In one example, some of the above-mentioned multiple decoding feature processing modules have upsampling and feature transformation processing functions, and some decoding feature processing modules only have feature transformation functions but do not have upsampling functions. At this time, the decoding end can perform upsampling and feature transformation processing on the reconstructed code vector of the audio data through some decoding feature processing modules, and perform feature transformation processing on the reconstructed code vector of the audio data through some decoding feature processing modules to obtain the reconstructed value of the audio data. For example, if multiple decoding feature processing modules are connected in series, assuming that the first decoding feature processing module among the multiple decoding feature processing modules has specific upsampling and feature transformation functions, and the second decoding module has feature transformation functions but does not have upsampling functions, then the decoding end inputs the reconstructed code vector of the audio data into the first decoding feature processing module among the multiple decoding feature processing modules for upsampling and feature transformation processing, and obtains the upsampling and feature transformed feature vector output by the first decoding feature processing module. Then, the feature vector output by the first decoding feature processing module is input into the second decoding feature processing module for feature transformation processing, and so on, to obtain the feature vector output by the last decoding feature processing module, and then obtain the reconstructed value of the audio data based on the feature vector.

[0191] In some embodiments, when only multiple decoding feature processing modules in the decoding network of the embodiment of the present application have an upsampling function, and other modules do not have an upsampling function, the total upsampling multiple of the decoding network is equal to the total upsampling multiples of the multiple decoding feature processing modules, which is equal to the total upsampling multiples of the encoding network.

[0192] For example, suppose that the decoding network includes multiple decoding feature processing modules, and these multiple decoding feature processing modules are upsampled d times. Suppose that the multiples of these d upsamplings are M1, M2, ..., Md respectively, and the total upsampling multiples of the multiple decoding feature processing modules are M1*M2*...Md. Suppose that the encoding network includes multiple encoding feature processing modules, and these multiple encoding feature processing modules are downsampled e times. Suppose that the multiples of these e downsamplings are N1, N2, ..., Ne respectively, and the total downsampling multiples of the multiple encoding feature processing modules are N1*N2*...*Ne. In the embodiment of the present application, it is sufficient to ensure that the total upsampling multiples M1*M2*...Md of the multiple decoding feature processing modules are equal to the total downsampling multiples N1*N2*...*Ne of the multiple encoding feature processing modules, and for a certain Ni or Mj, there is no direct corresponding relationship between their values.

[0193] In some embodiments, in addition to the multiple decoding feature processing modules having upsampling functions, the decoding network of the embodiment of the present application may also include other upsampling modules. For example, the embodiment of the present application sets one or more upsampling deconvolution layers between, before or after the multiple decoding feature processing modules. In this way, the total upsampling multiple of the decoding network is the product of the total upsampling multiples of the multiple decoding feature processing modules and the total upsampling multiples of the upsampling modules.

[0194] As can be seen from the above, in the embodiment of the present application, the network structures of multiple decoding feature processing modules can be the same or different, and the embodiment of the present application does not limit this. For example, some of the multiple decoding feature processing modules have an upsampling function, while some decoding feature processing modules do not have an upsampling function.

[0195] In some embodiments, Fig.11A As shown, the i-th decoding feature processing module of the embodiment of the present application includes R deconvolution layers, where R is a positive integer. At this time, the i-th decoding feature processing module in the above S103-A1 performs at least one of upsampling and feature transformation on the i-1-th feature vector to obtain the i-th feature vector, including the following steps S103-A11:

[0196] S103-A11. Perform at least one of upsampling and feature transformation on the i-1th feature vector through R deconvolution layers to obtain the i-th feature vector.

[0197] The embodiment of the present application includes R deconvolution layers for the ith decoding feature processing module, so that the decoding end can upsample and / or perform feature transformation processing on the i-1th feature vector output by the ith decoding feature processing module through the R deconvolution layers in the ith decoding feature processing module to obtain the ith feature vector. It should be noted that the embodiment of the present application does not limit the change of the number of channels of the input feature vector by the R deconvolution layers. For example, each of the R deconvolution layers can reduce the number of channels of the input feature vector by 2 times (of course, it can be other multiples). For another example, the change multiples of the number of channels of the input feature vector by each of the R deconvolution layers can be the same or different, and the present application does not limit this.

[0198] In the embodiment of the present application, the number of deconvolution layers included in different decoding feature processing modules may be the same or different. Here, an example in which the i-th decoding feature processing module includes R deconvolution layers is used for illustration.

[0199] The embodiment of the present application does not limit the specific functions of the R deconvolution layers included in the i-th decoding feature processing module.

[0200] In some embodiments, none of the above R deconvolution layers have an upsampling function, that is, the R deconvolution layers of the embodiment of the present application have a feature transformation function. For example, for each of the R deconvolution layers, the deconvolution layer performs a linear transformation on the input feature vector, but does not perform upsampling on the input feature vector. For example, the moving step size of the R deconvolution layers is equal to 1, and only the input feature vector is linearly transformed without changing the dimension. Optionally, the deconvolution kernel sizes of the R deconvolution layers may be the same. Optionally, the deconvolution kernel sizes of the R deconvolution layers may not be exactly the same.

[0201] In some embodiments, at least one of the R deconvolution layers is an upsampling deconvolution layer. In this case, performing at least one of upsampling and feature transformation on the i-1th feature vector through the R deconvolution layers in S103-A11 to obtain the i-th feature vector includes the following steps S103-A111:

[0202] S103-A111, upsample and transform the i-1th feature vector through R deconvolution layers to obtain the i-th feature vector.

[0203] In this implementation, at least one of the R deconvolution layers included in the i-th decoding feature processing module has an upsampling function, and the embodiment of the present application refers to the deconvolution layer with an upsampling function as an upsampling deconvolution layer. At this time, when the decoding end inputs the i-1th feature vector into the i-th decoding feature processing module, the R deconvolution layers in the i-th decoding feature processing module perform upsampling and feature transformation processing on the i-1th feature vector to obtain the i-th feature vector.

[0204] In a possible implementation, each of the R deconvolution layers has upsampling and feature transformation functions. For example, Fig. 11B As shown, R=3, the upsampling multiple stride (abbreviated as S) of the first deconvolution layer among the three deconvolution layers is A, the upsampling multiple of the second deconvolution layer is B, and the upsampling multiple of the third deconvolution layer is C. In this way, the decoder inputs the i-1th feature vector into the first deconvolution layer for upsampling, the upsampling multiple is A, and performs a linear transformation to obtain feature vector 1, then inputs feature vector 1 into the second deconvolution layer for upsampling, the upsampling multiple is B, and performs a linear transformation to obtain feature vector 2, and finally, inputs feature vector 2 into the third deconvolution layer for upsampling, the upsampling multiple is C, and performs a linear transformation to obtain the i-th feature vector.

[0205] In a possible implementation, some of the R deconvolution layers have upsampling and feature transformation functions, and some of the deconvolution layers do not have upsampling functions. Exemplarily, one of the R deconvolution layers has an upsampling function, and the other deconvolution layers do not have an upsampling function. That is, one of the R deconvolution layers is an upsampling deconvolution layer, and the other layers are non-upsampling deconvolution layers.

[0206] For example, Fig. 11C As shown, R=3, the upsampling multiple of the first deconvolution layer among the three deconvolution layers is A, and the second and third deconvolution layers do not have upsampling functions (i.e., the upsampling multiples are 1). In this way, the decoder inputs the i-1th feature vector into the first deconvolution layer for upsampling, the upsampling multiple is A, and performs linear transformation to obtain feature vector 1, then inputs feature vector 1 into the second deconvolution layer for linear transformation to obtain feature vector 3, and finally, inputs feature vector 3 into the third deconvolution layer for linear transformation to obtain the i-th feature vector.

[0207] In the embodiment of the present application, the connection method of the R deconvolution layers is as follows Fig. 11B and Fig. 11C In addition to the series connection shown, parallel connection or mixed connection can also be used. For example, some of the R deconvolution layers are connected in parallel and then connected in series with other deconvolution layers. In other words, the embodiment of the present application does not limit the specific connection method of the R deconvolution layers. At the same time, the embodiment of the present application does not limit the selection of parameters such as the deconvolution kernel size, the number of deconvolution kernels, the moving step size, and the padding number of the R deconvolution layers.

[0208] In some embodiments, the above-mentioned i-th decoding feature processing module may further include an activation layer in addition to R deconvolution layers. In this case, the above-mentioned upsampling and feature transformation processing of the i-1th feature vector through the R deconvolution layers to obtain the i-th feature vector includes: upsampling and linear transformation processing of the i-1th feature vector through the R deconvolution layers to obtain the upsampled feature vector; and nonlinear transformation of the upsampled feature vector through the activation layer to obtain the i-th feature vector.

[0209] In the embodiments of the present application, Fig.11D As shown, the i-th decoding feature processing module includes R deconvolution layers and an activation layer, and the activation layer is used to perform nonlinear transformation on the input data.

[0210] In one example, at least one of the R deconvolution layers is an upsampling deconvolution layer. At this time, when the decoding end inputs the i-1th feature vector into the i-th decoding feature processing module, the R deconvolution layers in the i-th decoding feature processing module upsample and linearly transform the i-1th feature vector to obtain the upsampled feature vector. Then, the upsampled feature vector is input into the activation layer for nonlinear transformation to obtain the i-th feature vector.

[0211] The embodiment of the present application does not limit the type of activation function included in the activation layer. Exemplarily, the activation function can be: relu function, sigmoid function, tanh function, leaky relu function, elu activation function, etc.

[0212] In some embodiments, at least one decoding feature processing module of the embodiments of the present application may include one or more residual units.

[0213] In some embodiments, Fig. 12A As shown, the decoding network of the embodiment of the present application includes an input layer in addition to the above-mentioned multiple decoding feature processing modules. At this time, the decoding end upsamples and performs feature transformation processing on the reconstructed coding vector through multiple decoding feature processing modules. Before obtaining the reconstructed value of the audio data, the reconstructed coding vector is first processed through the input layer to obtain a first feature vector. Then, the first feature vector is upsampled and performed feature transformation processing through multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0214] The embodiment of the present application does not limit the specific network structure of the input layer. For example, the input layer includes one or more deconvolution layers.

[0215] In some embodiments, the above-mentioned input layer does not have an upsampling function, that is, the input layer includes a deconvolution layer that does not have an upsampling function.

[0216] In some embodiments, the input layer has an upsampling function, that is, the input layer includes a deconvolution layer having an upsampling function. In this case, the total upsampling multiple of the decoding network includes the upsampling multiple of the input layer and the upsampling multiples of the multiple decoding feature processing modules.

[0217] In some embodiments, Fig. 12B As shown, the decoding network of the embodiment of the present application includes an output layer. At this time, the decoding end upsamples and performs feature transformation processing on the reconstructed coding vector through multiple decoding feature processing modules to obtain a reconstructed value of the audio data, including: upsampling and performing feature transformation processing on the reconstructed coding vector through multiple decoding feature processing modules to obtain a second feature vector; processing the second feature vector through the output layer to obtain a reconstructed value of the audio data.

[0218] The embodiment of the present application does not limit the specific network structure of the output layer. For example, the output layer includes one or more deconvolution layers.

[0219] The output layer of the embodiment of the present application is used to restore the convolved feature vector to the size of the original audio data. Assuming that the dimension of the audio data to be encoded is w and the number of channels is c, the dimension of the output feature of the output layer is w and the number of channels is c.

[0220] In some embodiments, the output layer does not have an upsampling function, that is, the output layer includes a deconvolution layer that does not have an upsampling function.

[0221] In some embodiments, the output layer has an upsampling function, that is, the output layer includes a deconvolution layer having an upsampling function. In this case, the total upsampling multiple of the decoding network includes the upsampling multiple of the input layer and the upsampling multiples of the multiple decoding feature processing modules.

[0222] In some embodiments, Fig. 12C As shown, the decoding network of the embodiment of the present application includes an input layer, multiple decoding networks and an output layer. At this time, the decoding end performs upsampling and feature transformation processing on the reconstructed coding vector through the decoding network to obtain the reconstructed value of the audio data, including: the decoding end processes the reconstructed coding vector of the audio data through the input layer to obtain a first feature vector, and then processes the first feature vector through multiple decoding feature processing modules to obtain a second feature vector. Finally, the second feature vector is processed through the output layer to obtain the reconstructed value of the audio data.

[0223] In some embodiments, the input layer and / or the output layer have an upsampling function, that is, the input layer and / or the output layer include a deconvolution layer having an upsampling function. In this case, the total upsampling multiple of the decoding network includes the upsampling multiple of the input layer, the upsampling multiples of the multiple decoding feature processing modules, and the upsampling multiple of the output layer.

[0224] In the audio decoding method provided by the embodiment of the present application, the decoding end obtains the quantization result corresponding to the audio data by decoding the code stream, and the quantization result is obtained by quantizing the coding vector of the audio data, and the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through the coding network. Then, the decoding end dequantizes the quantization result to obtain the reconstructed coding vector of the audio data, and finally performs at least one of upsampling and feature transformation on the reconstructed coding vector through the decoding network to obtain the reconstructed value of the audio data. Among them, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing module included in the decoding network is different from the encoding feature processing module included in the encoding network in at least one of the number, network structure, and sampling multiples. The decoding feature processing module is used to perform at least one of upsampling and feature transformation on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector. In this way, the network structures of the encoding network and the decoding network can be adjusted according to actual needs. For example, in order to reduce the decoding complexity, the decoding feature processing modules in the decoding network can be reduced. For another example, if the decoding performance needs to be improved, the number of decoding feature processing modules in the decoding network can be increased, etc., thereby improving the encoding and decoding efficiency of the audio data and effectively controlling the complexity of encoding and decoding.

[0225] The above describes the audio decoding method involved in the embodiment of the present application. The following describes the audio encoding method provided in the embodiment of the present application by taking the encoding end as an example.

[0226] Fig.13 1 is a flow chart of an audio coding method provided by an embodiment of the present application. The execution subject of the embodiment of the present application may be a device with a specific audio coding function, such as an audio coding device. In some embodiments, the audio coding device may be Figure 1 For ease of description, the embodiment of the present application is described by taking the execution subject as an encoding device as an example.

[0227] like Fig.13 As shown, the audio encoding method of the embodiment of the present application includes the following steps:

[0228] S201: Obtain audio data to be encoded.

[0229] In the embodiment of the present application, the audio data to be encoded may be a segment of audio data of any length.

[0230] In one example, the audio data to be encoded includes one or more audio frames. In one example, the audio data may also include an incomplete audio frame, such as 1 / 2 of an audio frame or 1 / 4 of an audio frame.

[0231] In the embodiment of the present application, the audio frame can be understood as a data segment with a specified time length obtained after framing and windowing the original audio data.

[0232] The embodiment of the present application does not limit the specific method of obtaining the original audio data.

[0233] In some examples, the original audio data may be speech captured by a terminal.

[0234] In some examples, the original audio data may be a sound signal collected in an Internet voice call or video call scenario.

[0235] In some examples, the original audio data may be a sound signal collected in a live broadcast scenario, a sound signal collected in an online singing scenario, or a sound signal collected in a voice broadcast scenario.

[0236] In some examples, the original audio data may be audio data acquired from a storage resource, such as stored voice, music, video, etc.

[0237] In a possible implementation, when dividing the original audio data into audio frames, the embodiment of the present application may set a preset time length for division, for example, dividing every 10 ms of original audio in the original audio data into an audio frame.

[0238] In order to store and transmit audio data over long distances, the acquired original audio data needs to be encoded to reduce the size of the audio data, thereby reducing the storage space of the audio data or reducing the traffic bandwidth consumed by long-distance transmission.

[0239] S202. Perform at least one of downsampling and feature transformation on the audio data through a coding network to obtain a coding vector of the audio data.

[0240] Among them, the total downsampling multiple of the encoding network is consistent with the total downsampling multiple of the decoding network, and the encoding feature processing modules included in the encoding network and the decoding feature processing modules included in the decoding network are different in at least one of the number, network structure, and sampling rate.

[0241] The decoding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector. For example, at least one decoding feature processing module only has a feature transformation function and does not have a downsampling function. For another example, at least one decoding feature processing module has both a feature transformation function and a downsampling function.

[0242] The coding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector. For example, at least one coding feature processing module only has a feature transformation function and does not have a downsampling function. For another example, at least one coding feature processing module has both a feature transformation function and a downsampling function.

[0243] In the embodiment of the present application, feature transformation can be understood as converting feature information from one set of data to another set of data. The feature transformation in the embodiment of the present application includes at least one of linear transformation (such as convolution operation) and nonlinear transformation (such as activation operation).

[0244] In an embodiment of the present application, after obtaining the audio data to be encoded based on the above steps, the encoding end inputs the audio data to be encoded into the encoding network for at least one of downsampling and feature transformation to obtain the encoding vector of the audio data.

[0245] The current encoding network structure is a mirror image of the decoding network structure, for example Figure 3 As shown, the encoding network includes 1 input layer, 4 encoding feature processing modules and 1 output layer, and the corresponding decoding network also includes 1 input layer, 4 encoding feature processing modules and 1 output layer, and the decoding feature processing module in the decoding network has the same network structure as the encoding feature processing module in the encoding network. In this way, when the network structure of the decoding network is complex, the network structure of the corresponding encoding network is also complex. For example, when the decoding network includes 8 decoding feature processing modules, the encoding network also includes 8 encoding feature processing modules, which doubles the complexity of audio data encoding and decoding, making the audio encoding and decoding efficiency low.

[0246] In order to solve the above technical problems, in an embodiment of the present application, the network models of the encoding network and the decoding network are independent, so that the network structures of the encoding network and the decoding network can be adjusted according to actual needs. For example, in order to reduce the encoding complexity, the encoding feature processing modules in the encoding network can be reduced. For another example, if it is necessary to improve the encoding performance, the number of encoding feature processing modules in the encoding network can be increased, etc., thereby improving the encoding efficiency of the audio data and effectively controlling the complexity of encoding and decoding.

[0247] That is to say, in the embodiment of the present application, there is no restriction on whether the encoding network and the decoding network are mirror-symmetric, as long as the total downsampling multiple of the encoding network is consistent with the total upsampling multiple of the decoding network.

[0248] The network structure of the coding network in the embodiment of the present application may include at least the following situations or a combination of the following situations:

[0249] In the first case, the decoding network includes one input layer, while the encoding network does not include an input layer, or the network structure of the input layer included in the encoding network is inconsistent with that of the input layer included in the decoding network. For example, the input layer included in the decoding network includes one convolutional layer, while the input layer of the encoding network of the embodiment of the present application may include two or more deconvolutional layers. For another example, the input layer included in the decoding network includes one convolutional layer, while the input layer of the encoding network of the embodiment of the present application may include a residual unit, etc.

[0250] In the second case, the decoding network includes one output layer, while the encoding network does not include an output layer, or the output layer included in the encoding network is inconsistent with the network structure of the output layer included in the decoding network. For example, the output layer included in the decoding network includes one convolutional layer, while the output layer of the encoding network of the embodiment of the present application may include two or more deconvolutional layers. For another example, the output layer included in the decoding network includes one convolutional layer, while the output layer of the encoding network of the embodiment of the present application may include a residual unit, etc.

[0251] In a third case, the encoding feature processing modules included in the encoding network are different from the decoding feature processing modules included in the decoding network in at least one of the number, network structure, and sampling multiples.

[0252] In an example of the third case, the number of encoding feature processing modules included in the encoding network of the embodiment of the present application is different from the number of decoding feature processing modules included in the decoding network. For example, the decoding network includes m decoding feature processing modules, and the encoding network includes n encoding feature processing modules, and n is not equal to m.

[0253] In another example of the third case, the network structure of the encoding feature processing module included in the encoding network of the embodiment of the present application is not mirror-symmetrical with the network structure of the decoding feature processing module included in the decoding network. Exemplarily, the decoding feature processing module of the embodiment of the present application includes multiple residual units, and the encoding feature processing module of the present application includes multiple convolutional layers. Exemplarily, the hierarchy of the encoding feature processing module of the embodiment of the present application is not mirror-symmetrical with the hierarchy of the decoding feature processing module, for example, the decoding feature processing module includes 4 convolutional layers, and the encoding feature processing module includes 3 convolutional layers.

[0254] In another example of the third case, the network structure of the encoding feature processing module included in the encoding network of the embodiment of the present application is not mirror-symmetric with the sampling multiples of the decoding feature processing module included in the decoding network.

[0255] In one example, the number of encoding feature processing modules included in the encoding network is the same as the number of decoding feature processing modules included in the decoding network, but the downsampling multiples of the encoding feature processing modules included in the encoding network are not mirror-symmetric with the upsampling multiples of the decoding feature processing modules included in the decoding network. For example, the decoding network includes 4 decoding feature processing modules, and the upsampling multiples of the 4 decoding feature processing modules are [2, 4, 6, 8] respectively, and the downsampling multiples of the 4 encoding feature processing modules included in the encoding network can be any group of upsampling multiples except [8, 6, 4, 2]. For example, the downsampling multiples of the 4 encoding feature processing modules included in the encoding network of the embodiment of the present application can be [8, 4, 6, 2], [8, 8, 3, 2], etc.

[0256] In one example, the number of encoding feature processing modules included in the encoding network is different from the number of decoding feature processing modules included in the decoding network, and the downsampling multiples of the encoding feature processing modules included in the encoding network are not mirror-symmetric with the upsampling multiples of the decoding feature processing modules included in the decoding network. For example, the decoding network includes 4 decoding feature processing modules, and the upsampling multiples of these 4 decoding feature processing modules are [2, 4, 6, 8] respectively, while the downsampling multiples of the encoding network including 3 encoding feature processing modules may be [8, 6, 8], etc.

[0257] In an embodiment of the present application, the encoding network is used to perform at least one of downsampling and feature transformation on the audio data to obtain an encoding vector of the audio data.

[0258] That is to say, the encoding network of the embodiment of the present application has downsampling and feature transformation functions.

[0259] The feature transformation includes linear transformation and / or nonlinear transformation.

[0260] In the embodiment of the present application, the coding network has a downsampling function, which can be understood as at least one module included in the coding network having a downsampling function. Similarly, the coding network has a feature transformation function, which can be understood as at least one module included in the coding network having a feature transformation function.

[0261] In some embodiments, some modules in the encoding network have a downsampling function, and some modules have a feature transformation function. Among the modules with the feature transformation function, some modules have a linear transformation function, and some modules have a nonlinear transformation function.

[0262] In some embodiments, some modules in the coding network have a downsampling function, some modules have a feature conversion function, and some modules have both a downsampling function and a feature conversion function. Alternatively, some modules in the coding network have a downsampling function, and some modules have both a downsampling function and a feature conversion function. Alternatively, some modules in the coding network have a feature conversion function, and some modules have both a downsampling function and a feature conversion function.

[0263] That is to say, in the embodiment of the present application, there is no restriction on which modules in the encoding network have the downsampling function and which modules have the feature transformation function, and they can be set according to actual needs.

[0264] In some embodiments, Fig.14 As shown, the encoding network proposed in the embodiment of the present application includes multiple encoding feature processing modules.

[0265] The number of encoding feature processing modules included in the encoding network of the embodiment of the present application may be the same as or different from the number of decoding feature processing modules included in the decoding network, and the embodiment of the present application does not impose any restrictions on this.

[0266] In one example, in order to reduce the encoding complexity, the number of encoding feature processing modules in the encoding network can be reduced. For example, the number of encoding feature processing modules included in the encoding network is set to be smaller than the number of decoding feature processing modules included in the decoding network.

[0267] In one example, if the encoding performance needs to be improved, the number of encoding feature processing modules in the encoding network can be increased. For example, the number of encoding feature processing modules included in the encoding network is set to be greater than the number of decoding feature processing modules included in the decoding network.

[0268] The encoding feature processing module of the embodiment of the present application is used to perform at least one of downsampling and feature transformation on the input feature vector.

[0269] In some embodiments, each of the multiple coding feature processing modules does not have a downsampling function, but completes a feature transformation function. At this point, the coding network may also include other downsampling modules, such as at least one downsampling convolution layer (e.g., a pooling layer). Exemplarily, some or all of the downsampling convolution layers in the at least one downsampling convolution layer may be continuously or discontinuously interspersed between at least two coding feature processing modules in the multiple coding feature processing modules. Alternatively, the three downsampling convolution layers may be arranged before the multiple coding feature processing modules, or after the multiple coding feature processing modules. For example, assuming that the coding network includes three downsampling convolution layers, the three downsampling convolution layers may be continuously interspersed between two coding feature processing modules in the multiple coding feature processing modules. Or the three downsampling convolution layers may be interspersed before or after different coding feature processing modules in the multiple coding feature processing modules, such as interspersing one or more downsampling convolution layers before or after a coding feature processing module. Alternatively, the three downsampling convolutional layers are arranged before a plurality of coding feature processing modules, or the three downsampling convolutional layers are arranged after a plurality of coding feature processing modules. The embodiment of the present application does not limit the specific distribution of coding feature processing modules and downsampling modules in the coding network. At this time, the coding end can input the above-obtained audio data into the coding network, the plurality of coding feature processing modules in the coding network implement feature transformation processing, the downsampling module in the coding network implements downsampling processing, and finally the coding network outputs the coding vector of the audio data.

[0270] In some embodiments, at least one encoding feature processing module among the multiple encoding feature processing modules included in the encoding network is used to downsample the input feature vector. In this case, the above S202 includes the following step S202-A:

[0271] S202-A. Downsample and perform feature transformation processing on the audio data through multiple coding feature processing modules to obtain a coding vector of the audio data.

[0272] In this implementation, at least one of the multiple encoding feature processing modules included in the encoding network has a downsampling function. In this way, the encoding end can perform downsampling and feature transformation processing on the above-mentioned acquired audio data through these multiple encoding feature processing modules to obtain the encoding vector of the audio data. For example, the encoding end can perform downsampling and feature transformation processing on the above-mentioned acquired audio data with a dimension of w and a number of channels of c through these multiple encoding feature processing modules to obtain an encoding vector with a dimension of k and a number of channels of 1.

[0273] In this implementation, the function of each encoding feature processing module in the multiple encoding feature processing modules may be the same or different.

[0274] In one example, each of the plurality of coding feature processing modules has downsampling and feature transformation processing functions. At this time, the encoding end can perform downsampling and feature transformation processing on the acquired audio data through each of the plurality of coding feature processing modules to obtain a coding vector of the audio data.

[0275] In one example, some of the above-mentioned multiple coding feature processing modules have downsampling and feature transformation processing functions, and some of the coding feature processing modules only have feature transformation functions but do not have downsampling functions. At this time, the encoding end can perform downsampling and feature transformation processing on the audio data volume through the coding feature processing modules with downsampling and feature transformation functions among the multiple coding feature processing modules, and perform feature transformation processing on the audio data through the coding feature processing modules with only feature transformation functions but not downsampling functions among the multiple coding feature processing modules to obtain the coding vector of the audio data.

[0276] The embodiment of the present application does not limit the specific distribution structure of the above-mentioned multiple coding feature processing modules.

[0277] In some embodiments, the above-mentioned multiple coding feature processing modules can be connected in parallel. As an example, Fig.15 As shown, the encoding network of the embodiment of the present application includes an input layer, three encoding feature processing modules and an output layer, wherein the three encoding feature processing modules are connected in parallel. In this way, the encoding end processes the audio data through the input layer to obtain a feature vector 1, and then inputs the feature vector 1 into the three encoding feature processing modules for processing, such as downsampling and / or feature transformation processing, to obtain a feature vector output by each of the three encoding feature processing modules. Then, the feature vectors output by the three encoding feature processing modules are fused, such as by addition or multiplication, to obtain a fused feature vector 2, and finally the fused feature vector 2 is input into the output layer for processing to obtain the encoding vector of the audio data. Optionally, the encoding network may not include Fig.15 The input layer and / or output layer shown.

[0278] In some embodiments, some of the plurality of coding feature processing modules are connected in parallel, and some of the coding feature processing modules are connected in series. Fig.16As shown, the encoding network of the embodiment of the present application includes an input layer, three encoding feature processing modules and an output layer, wherein two of the three encoding feature processing modules are connected in parallel and then connected in series with another encoding feature processing module. In this way, the encoding end processes the audio data through the input layer to obtain a feature vector 1, and then, the feature vector 1 is respectively input into the two encoding feature processing modules connected in parallel for processing, such as downsampling and / or feature transformation processing, to obtain a feature vector output by each of the two encoding feature processing modules connected in parallel. Then, the feature vectors respectively output by the two encoding feature processing modules connected in parallel are fused, such as by addition or multiplication, to obtain a fused feature vector 3, and the fused feature vector 3 is input into the encoding feature processing module connected in series for downsampling and / or feature transformation processing to obtain a feature vector 4, and finally, the feature vector 4 is input into the output layer for processing to obtain the encoding vector of the audio data. Optionally, the encoding network may not include Fig.16 The input layer and / or output layer shown.

[0279] In some embodiments, the above-mentioned multiple coding feature processing modules are connected in series. In this case, the coding end S202-A includes the following steps S202-A1 and S202-A2:

[0280] S202-A1, performing at least one of downsampling and feature transformation on the i-1th feature vector through the i-th coding feature processing module to obtain the i-th feature vector, where i is a positive integer. If i is equal to 1, the i-1th feature vector is audio data, and if i is greater than 1, the i-1th feature vector is a feature vector output by the i-1th coding feature processing module;

[0281] S202-A2, perform at least one of downsampling and feature transformation on the i-th feature vector through the i+1-th coding feature processing module, repeat the process, and determine the coding vector of the audio data based on the feature vector output by the last coding feature processing module.

[0282] As an example, Fig.17As shown, the encoding network of the embodiment of the present application includes an input layer, three encoding feature processing modules and an output layer, wherein the three encoding feature processing modules are connected in series. In this way, the encoding end processes the audio data through the input layer to obtain feature vector 1, and then, the feature vector 1 is respectively input into the first encoding feature processing module for processing, such as downsampling and / or feature transformation processing, to obtain feature vector 5. Then, feature vector 5 is input into the second encoding feature processing module for downsampling and / or feature transformation processing to obtain feature vector 6. Then, feature vector 6 is input into the third encoding feature processing module for downsampling and / or feature transformation processing to obtain feature vector 7. Finally, feature vector 7 is input into the output layer for processing to obtain the encoding vector of the audio data. Optionally, the encoding network may not include Fig.17 The input layer and / or output layer shown.

[0283] In one example, all the encoding feature processing modules of the above-mentioned multiple encoding feature processing modules have downsampling and feature transformation processing functions. At this time, the encoding end can downsample and feature transform the above-mentioned audio data through each of the multiple encoding feature processing modules to obtain the encoding vector of the audio data. For example, if multiple encoding feature processing modules are connected in series, the encoding end inputs the audio data into the first encoding feature processing module among the multiple encoding feature processing modules for downsampling and feature transformation processing, and obtains the feature vector output by the first encoding feature processing module after downsampling and feature transformation. Then, the feature vector output by the first encoding feature processing module is input into the second encoding feature processing module for downsampling and feature transformation processing, and so on, to obtain the feature vector output by the last encoding feature processing module after downsampling and feature transformation, and then obtain the encoding vector of the audio data based on the feature vector.

[0284] In one example, some of the above-mentioned multiple coding feature processing modules have downsampling and feature transformation processing functions, and some of the coding feature processing modules only have feature transformation functions but do not have downsampling functions. At this time, the encoding end can perform downsampling and feature transformation processing on the audio data through some of the coding feature processing modules, and perform feature transformation processing on the audio data through some of the coding feature processing modules to obtain the encoding vector of the audio data. For example, if multiple coding feature processing modules are connected in series, assuming that the first coding feature processing module among the multiple coding feature processing modules has specific downsampling and feature transformation functions, and the second solution module has feature transformation functions but does not have downsampling functions, then the encoding end inputs the audio data into the first coding feature processing module among the multiple coding feature processing modules for downsampling and feature transformation processing, and obtains the feature vector after downsampling and feature transformation output by the first coding feature processing module. Then, the feature vector output by the first coding feature processing module is input into the second coding feature processing module for feature transformation processing, and so on, to obtain the feature vector output by the last coding feature processing module, and then obtain the encoding vector of the audio data based on the feature vector.

[0285] In some embodiments, when only multiple encoding feature processing modules in the encoding network of the embodiment of the present application have a downsampling function, and other modules do not have a downsampling function, the total downsampling multiple of the encoding network is equal to the total downsampling multiples of the multiple encoding feature processing modules, which is equal to the total upsampling multiples of the decoding network.

[0286] For example, assuming that the encoding network includes multiple encoding feature processing modules, and these multiple encoding feature processing modules are downsampled e times, assuming that the multiples of these e downsamplings are N1, N2, ..., Ne respectively, and the total downsampling multiples of the multiple encoding feature processing modules are N1*N2*...*Ne. Assuming that the decoding network includes multiple decoding feature processing modules, and these multiple decoding feature processing modules are upsampled d times, assuming that the multiples of these d upsamplings are M1, M2, ..., Md respectively, and the total upsampling multiples of the multiple decoding feature processing modules are M1*M2*...Md. In the embodiment of the present application, it is sufficient to ensure that the total downsampling multiples N1, N2, ..., Ne of the multiple encoding feature processing modules are equal to the total downsampling multiples N1*N2*...*Ne of the multiple decoding feature processing modules, and for a certain Ni or Mj, there is no direct corresponding relationship between their values.

[0287] In some embodiments, in addition to the multiple coding feature processing modules having the downsampling function, the coding network of the embodiment of the present application may also include other downsampling modules. For example, the embodiment of the present application sets one or more downsampling convolution layers between, before or after the multiple coding feature processing modules. In this way, the total downsampling multiple of the coding network is the product of the total downsampling multiples of the multiple coding feature processing modules and the total downsampling multiples of the downsampling modules.

[0288] As can be seen from the above, in the embodiment of the present application, the network structures of multiple coding feature processing modules can be the same or different, and the embodiment of the present application does not limit this. For example, some of the multiple coding feature processing modules have a downsampling function, while some coding feature processing modules do not have a downsampling function.

[0289] In some embodiments, Fig.18A As shown, the i-th coding feature processing module of the embodiment of the present application includes T convolution layers, where T is a positive integer. At this time, the i-th coding feature processing module in the above S202-A1 performs at least one of downsampling and feature transformation on the i-1-th feature vector to obtain the i-th feature vector, including the following steps of S202-A11:

[0290] S202-A11. Perform at least one of downsampling and feature transformation on the i-1th feature vector through T convolutional layers to obtain the i-th feature vector.

[0291] The embodiment of the present application includes T convolutional layers for the ith coding feature processing module. In this way, the encoding end can downsample and / or perform feature transformation processing on the i-1th feature vector output by the ith coding feature processing module through the T convolutional layers in the ith coding feature processing module to obtain the ith feature vector. It should be noted that the embodiment of the present application does not limit the change of the number of channels of the input feature vector by these T convolutional layers. For example, each of the T convolutional layers can reduce the number of channels of the input feature vector by 2 times (of course, it can be other multiples). For another example, the change multiples of the number of channels of the input feature vector by each of the T convolutional layers can be the same or different, and the present application does not limit this.

[0292] The embodiment of the present application does not limit the specific functions of the T convolutional layers included in the i-th encoding feature processing module.

[0293] In some embodiments, none of the T convolutional layers mentioned above have a downsampling function, that is, the T convolutional layers of the embodiment of the present application have a feature transformation function. For example, for each of the T convolutional layers, the convolutional layer performs a linear transformation on the input feature vector, but does not perform a downsampling process on the input feature vector. For example, the moving step size of the T convolutional layers is equal to 1, and only a linear transformation is performed on the input feature vector without changing the dimension. Optionally, the convolution kernel sizes of the T convolutional layers may be the same. Optionally, the convolution kernel sizes of the T convolutional layers may not be exactly the same.

[0294] In some embodiments, at least one of the T convolutional layers is a downsampling convolutional layer. In this case, performing at least one of downsampling and feature transformation on the i-1th feature vector through the T convolutional layers in S202-A11 to obtain the i-th feature vector includes the following steps of S202-A111:

[0295] S202-A111. Downsample and transform the i-1th feature vector through T convolutional layers to obtain the i-th feature vector.

[0296] In this implementation, at least one of the T convolutional layers included in the i-th coding feature processing module has a downsampling function, and the embodiment of the present application refers to the convolutional layer with a downsampling function as a downsampling convolutional layer. At this time, when the encoder inputs the i-1th feature vector into the i-th coding feature processing module, the T convolutional layers in the i-th coding feature processing module downsample and perform feature transformation processing on the i-1th feature vector to obtain the i-th feature vector.

[0297] In a possible implementation, each of the T convolutional layers has downsampling and feature transformation functions. For example, Fig.18B As shown, R=3, the downsampling multiple of the first convolutional layer among the three convolutional layers is A, the downsampling multiple of the second convolutional layer is B, and the downsampling multiple of the third convolutional layer is C. In this way, the encoder inputs the i-1th feature vector into the first convolutional layer for downsampling, the downsampling multiple S is A, and performs a linear transformation to obtain feature vector 1. Then, feature vector 1 is input into the second convolutional layer for downsampling, the downsampling multiple is B, and a linear transformation is performed to obtain feature vector 2. Finally, feature vector 2 is input into the third convolutional layer for downsampling, the downsampling multiple is C, and a linear transformation is performed to obtain the i-th feature vector.

[0298] In a possible implementation, some of the T convolutional layers have downsampling and feature transformation functions, and some of the convolutional layers do not have downsampling functions. Exemplarily, one of the T convolutional layers has a downsampling function, and the other convolutional layers do not have a downsampling function. That is, one of the T convolutional layers is a downsampling convolutional layer, and the other layers are non-downsampling convolutional layers.

[0299] For example, Fig. 18C As shown, R=3, the downsampling multiple S of the first convolution layer among the three convolution layers is A, and the second and third convolution layers do not have the downsampling function, that is, the downsampling multiple is 1. In this way, the encoder inputs the i-1th feature vector into the first convolution layer for downsampling, the downsampling multiple S is A, and performs linear transformation to obtain feature vector 1, then inputs feature vector 1 into the second convolution layer for linear transformation to obtain feature vector 3, and finally, inputs feature vector 3 into the third convolution layer for linear transformation to obtain the i-th feature vector.

[0300] In the embodiment of the present application, the connection mode of the T convolutional layers, in addition to the series connection shown above, can also be connected in parallel, or mixed connection, for example, some of the T convolutional layers are connected in parallel and then connected in series with other convolutional layers. In other words, the embodiment of the present application does not limit the specific connection mode of the T convolutional layers, and at the same time, the embodiment of the present application does not limit the selection of parameters such as the convolution kernel size, the number of convolution kernels, the moving step size, and the padding number of the T convolutional layers.

[0301] In some embodiments, the above-mentioned i-th encoding feature processing module may further include an activation layer in addition to T convolutional layers. In this case, the above-mentioned downsampling and feature transformation processing of the i-1th feature vector through T convolutional layers to obtain the i-th feature vector includes: downsampling and linear transformation processing of the i-1th feature vector through T convolutional layers to obtain a downsampled feature vector; and nonlinear transformation of the downsampled feature vector through the activation layer to obtain the i-th feature vector.

[0302] In the embodiments of the present application, Fig.18D As shown, the i-th encoding feature processing module includes T convolutional layers and an activation layer, and the activation layer is used to perform nonlinear transformation on the input data.

[0303] In one example, at least one of the T convolutional layers is a downsampling convolutional layer. At this time, when the encoder inputs the i-1th feature vector into the i-th encoding feature processing module, the T convolutional layers in the i-th encoding feature processing module downsample and linearly transform the i-1th feature vector to obtain a downsampled feature vector. Then, the downsampled feature vector is input into the activation layer for nonlinear transformation to obtain the i-th feature vector.

[0304] The embodiment of the present application does not limit the type of activation function included in the activation layer. Exemplarily, the activation function can be: relu function, sigmoid function, tanh function, leaky relu function, elu activation function, etc.

[0305] In some embodiments, at least one encoding feature processing module of the embodiments of the present application may include one or more residual units.

[0306] In some embodiments, Fig.19A As shown, the encoding network of the embodiment of the present application includes an input layer in addition to the above-mentioned multiple encoding feature processing modules. At this time, the encoding end downsamples and performs feature transformation processing on the audio data through multiple encoding feature processing modules. Before obtaining the encoding vector of the audio data, the audio data is first processed through the input layer to obtain a first feature vector. Then, the first feature vector is downsampled and feature transformed through multiple encoding feature processing modules to obtain the encoding vector of the audio data.

[0307] The embodiment of the present application does not limit the specific network structure of the input layer. For example, the input layer includes one or more convolutional layers.

[0308] In some embodiments, the input layer does not have a downsampling function, that is, the input layer including the convolution layer does not have a downsampling function.

[0309] In some embodiments, the input layer has a downsampling function, that is, the input layer includes a convolutional layer with a downsampling function. In this case, the total downsampling multiple of the encoding network includes the downsampling multiple of the input layer and the downsampling multiples of the multiple encoding feature processing modules.

[0310] In some embodiments, Fig.19B As shown, the encoding network of the embodiment of the present application includes an output layer. At this time, the encoding end downsamples and performs feature transformation processing on the audio data through multiple encoding feature processing modules to obtain the encoding vector of the audio data, including: downsampling and performing feature transformation processing on the audio data through multiple encoding feature processing modules to obtain a second feature vector; processing the second feature vector through the output layer to obtain the encoding vector of the audio data.

[0311] The embodiment of the present application does not limit the specific network structure of the output layer. For example, the output layer includes one or more convolutional layers.

[0312] The output layer of the embodiment of the present application is used to convert the convolved feature vector into a coding vector of a preset dimension (for example, k) and a preset number of channels (for example, 1). Assuming that the dimension of the audio data to be decoded is w and the number of channels is c, the dimension of the coding vector output by the output layer is k and the number of channels is 1.

[0313] In some embodiments, the output layer does not have a downsampling function, that is, the output layer includes a convolutional layer that does not have a downsampling function.

[0314] In some embodiments, the output layer has a downsampling function, that is, the output layer includes a convolutional layer having a downsampling function. In this case, the total downsampling multiple of the encoding network includes the downsampling multiple of the input layer and the downsampling multiples of the multiple encoding feature processing modules.

[0315] In some embodiments, Fig.19C As shown, the coding network of the embodiment of the present application includes an input layer, multiple coding networks and an output layer. At this time, the coding end downsamples and performs feature transformation processing on the audio data through the coding network to obtain the coding vector of the audio data, including: the coding end processes the audio data through the input layer to obtain a first feature vector, and then processes the first feature vector through multiple coding feature processing modules to obtain a second feature vector. Finally, the second feature vector is processed through the output layer to obtain the coding vector of the audio data.

[0316] In some embodiments, the input layer and / or the output layer have a downsampling function, that is, the input layer and / or the output layer include a convolutional layer with a downsampling function. In this case, the total downsampling multiple of the encoding network includes the downsampling multiple of the input layer, the downsampling multiples of the multiple encoding feature processing modules, and the downsampling multiple of the output layer.

[0317] After the encoding end generates the encoding vector of the audio data based on the above steps, it executes the following step S203.

[0318] S203, quantize the coding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a code stream.

[0319] The embodiment of the present application does not limit the specific manner in which the encoding end quantizes the encoding vector of the audio data.

[0320] In some embodiments, the encoding end samples a preset encoder, quantizes the encoding vector of the audio data, and obtains a quantization result of the audio data.

[0321] In some embodiments, when the encoding end quantizes the encoding vector of the audio data, it can determine the effective audio information statistics of the audio data, and then select the target quantizer of the audio data based on the effective audio information statistics of the audio data, and then use the target quantizer to quantize the encoding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a bit stream.

[0322] In some embodiments, if the quantizer used in the embodiment of the present application is as follows Figure 4 When the residual-based vector quantizer is used as shown, the encoding end inputs the encoding vector of the audio data into the first quantization layer of the target quantizer for quantization, and obtains the first quantization vector and the codebook subscript corresponding to the first quantization layer; based on the encoding vector of the audio data and the first quantization vector, a first residual vector is obtained; the first residual vector is input into the second quantization layer of the target quantizer for quantization, and it is iterated repeatedly to obtain the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook subscript corresponding to each quantization layer in the target quantizer; the residual vector corresponding to the last quantization layer of the target quantizer and the codebook subscript corresponding to each quantization layer in the target quantizer are encoded to obtain a code stream.

[0323] For example, Fig. 20 As shown in FIG. 1 , assuming that the target quantizer includes three quantization layers, the encoder inputs the encoding vector y0 of the i-th audio frame into the first quantization layer of the target quantizer for quantization, searches for a codebook vector matching the encoding vector y0 in the codebook vectors included in the codebook 1 corresponding to the first quantization layer, and uses it as the output vector y1 of the first quantization layer. At the same time, the codebook index index0 of the vector y1 in the codebook 1 corresponding to the first quantization layer is recorded. The difference between the encoding vector y0 and the vector y1 is used as the residual vector Then the residual vector The input is quantized in the second quantization layer of the target quantizer, and the codebook vector included in the codebook 2 corresponding to the second quantization layer is searched for the residual vector The matched codebook vector is used as the output vector y2 of the second quantization layer. At the same time, the codebook index index1 of the vector y2 in the codebook 2 corresponding to the second quantization layer is recorded. Then, the residual vector The difference with vector y2 is used as the residual vector Then the residual vector The input is quantized in the third quantization layer of the target quantizer, and the codebook vector included in the codebook 3 corresponding to the third quantization layer is searched for the residual vector The matched codebook vector is used as the output vector y3 of the third quantization layer. At the same time, the codebook index index2 of the vector y3 in the codebook 3 corresponding to the third quantization layer is recorded.

[0324] Next, the residual vector The difference with vector y3 is used as the residual vector

[0325] As can be seen from the above, in the embodiment of the present application, the target quantizer is used to quantize the coding vector of the audio data, and the final quantization result obtained includes at least the residual vector corresponding to the last quantization layer of the target quantizer, and the codebook index corresponding to each quantization layer. Fig. 20 As shown, the final quantization result includes the residual vector corresponding to the last quantization layer of the target quantizer And the codebook indexes index0, index1 and index2 corresponding to each quantization layer.

[0326] Next, the encoder encodes the above quantization results to obtain a bitstream. For example, the encoder encodes the residual vector corresponding to the last quantization layer of the target quantizer And the codebook indexes index0, index1 and index2 corresponding to each quantization layer are encoded to form a binary code stream.

[0327] The audio encoding method provided in the embodiment of the present application, the encoding end obtains the audio data to be encoded, and downsamples and performs feature transformation processing on the audio data through the encoding network to obtain the encoding vector of the audio data. Finally, the encoding end quantizes the encoding vector of the audio data to obtain a quantization result, and encodes the quantization result to obtain a code stream. The total downsampling multiple of the encoding network is consistent with the total upsampling multiple of the decoding network, and the encoding network and the decoding network are not completely mirror-symmetrical, so that the network structures of the encoding network and the decoding network can be adjusted according to actual needs. For example, in order to reduce the encoding complexity, the encoding feature processing module in the encoding network can be reduced. For another example, if it is necessary to improve the encoding performance, the number of encoding feature processing modules in the encoding network can be increased, thereby improving the encoding efficiency of the audio data and effectively controlling the encoding complexity.

[0328] Combination of the above Figures 5 to 20 , describes in detail the audio encoding and decoding method embodiment of the present application, and the following is combined with Figure 21 to Figure 22 , describe in detail the device embodiments of the present application.

[0329] Fig.21 1 is a schematic block diagram of an audio decoding apparatus provided in an embodiment of the present application. The apparatus 10 can be applied to a decoding device.

[0330] like Fig.21 As shown, the audio decoding device 10 includes:

[0331] A decoding unit 11 is used to decode the bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data, and the coding vector is obtained by downsampling and feature transforming the audio data through a coding network;

[0332] A dequantization unit 12, configured to dequantize the quantization result to obtain a reconstructed coding vector of the audio data;

[0333] The reconstruction unit 13 is used to perform at least one of upsampling and feature transformation on the reconstructed coding vector through a decoding network to obtain a reconstructed value of the audio data, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples. The decoding feature processing module is used to perform at least one of upsampling and feature transformation on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation on the input feature vector.

[0334] In some embodiments, the decoding network includes multiple decoding feature processing modules, and at least one of the multiple decoding feature processing modules is used to upsample the input feature vector; the reconstruction unit 13 is specifically used to perform at least one of upsampling and feature transformation on the reconstructed coding vector through the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

[0335] In some embodiments, the reconstruction unit 13 is specifically used to perform at least one of upsampling and feature transformation on the i-1th feature vector through the i-th decoding feature processing module to obtain the i-th feature vector, where i is a positive integer less than or equal to Q, and if i is equal to 1, then the i-1th feature vector is the reconstructed coding vector, and if i is greater than 1, then the i-1th feature vector is the feature vector output by the i-1th decoding feature processing module; perform at least one of upsampling and feature transformation on the i-th feature vector through the i+1th decoding feature processing module, repeat the execution, and determine the reconstructed value of the audio data based on the feature vector output by the last decoding feature processing module.

[0336] In some embodiments, the i-th decoding feature processing module includes R deconvolution layers, where R is a positive integer; the reconstruction unit 13 is specifically used to perform at least one of upsampling and feature transformation on the i-1-th feature vector through the R deconvolution layers to obtain the i-th feature vector.

[0337] In some embodiments, at least one of the R deconvolution layers is an upsampling deconvolution layer; the reconstruction unit 13 is specifically used to upsample and transform the i-1th feature vector through the R deconvolution layers to obtain the i-th feature vector.

[0338] In some embodiments, the i-th decoding feature processing module also includes an activation layer, a reconstruction unit 13, which is specifically used to upsample and linearly transform the i-1-th feature vector through the R deconvolution layers to obtain an upsampled feature vector; and perform a nonlinear transformation on the upsampled feature vector through the activation layer to obtain the i-th feature vector.

[0339] In some embodiments, one of the R deconvolution layers is an upsampling deconvolution layer, and the other layers are non-upsampling deconvolution layers.

[0340] In some embodiments, the decoding feature processing module includes at least one residual unit.

[0341] In some embodiments, the decoding network includes an input layer, a reconstruction unit 13, which is specifically used to process the reconstructed coding vector through the input layer to obtain a first feature vector; through the multiple decoding feature processing modules, the first feature vector is upsampled and feature transformed to obtain the reconstructed value of the audio data.

[0342] In some embodiments, the decoding network includes an output layer, a reconstruction unit 13, which is specifically used to upsample and feature transform the reconstructed coding vector through the multiple decoding feature processing modules to obtain a second feature vector; and process the second feature vector through the output layer to obtain the reconstructed value of the audio data.

[0343] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Fig.21 The device shown can execute the above-mentioned audio decoding method embodiment, and the above-mentioned and other operations and / or functions of each module in the device are respectively for realizing the above-mentioned method embodiment, which will not be described in detail here for the sake of brevity.

[0344] Fig. 22 20 is a schematic block diagram of an audio encoding device provided in an embodiment of the present application. The device 20 can be applied to an encoding device.

[0345] like Fig. 22 As shown, the audio encoding device 20 includes:

[0346] An acquisition unit 21, used to acquire audio data to be encoded;

[0347] A transformation unit 22, configured to perform at least one of downsampling and feature transformation on the audio data through an encoding network to obtain an encoding vector of the audio data, wherein the total downsampling multiple of the encoding network is consistent with the total upsampling multiple of the decoding network, and the encoding feature processing modules included in the encoding network are different from the decoding feature processing modules included in the decoding network in at least one of the number, network structure, and sampling rate, the encoding feature processing modules are configured to perform at least one of downsampling and feature transformation on the input feature vector, and the decoding feature processing modules are configured to perform at least one of upsampling and feature transformation on the input feature vector;

[0348] The encoding unit 23 is used to quantize the encoding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a code stream.

[0349] In some embodiments, the encoding network includes multiple encoding feature processing modules, at least one of the multiple encoding feature processing modules is used to downsample the input feature vector; the transformation unit 22 is specifically used to downsample and feature transform the audio data through the multiple encoding feature processing modules to obtain the encoding vector of the audio data.

[0350] In some embodiments, the transformation unit 23 is specifically used to perform at least one of downsampling and feature transformation on the i-1th feature vector through the i-th coding feature processing module to obtain the i-th feature vector, where i is a positive integer. If i is equal to 1, the i-1th feature vector is the audio data; if i is greater than 1, the i-1th feature vector is the feature vector output by the i-1th coding feature processing module; perform at least one of downsampling and feature transformation on the i-th feature vector through the i+1th coding feature processing module, repeat the execution, and determine the coding vector of the audio data based on the feature vector output by the last coding feature processing module.

[0351] In some embodiments, the i-th encoding feature processing module includes T convolutional layers, where T is a positive integer; the transformation unit 23 is specifically used to perform at least one of downsampling and feature transformation on the i-1-th feature vector through the T convolutional layers to obtain the i-th feature vector.

[0352] In some embodiments, at least one of the T convolutional layers is a downsampling convolutional layer; the transformation unit 23 is specifically used to downsample and linearly transform the i-1th feature vector through the T convolutional layers to obtain the i-th feature vector.

[0353] In some embodiments, the i-th encoding feature processing module also includes an activation layer, a transformation unit 23, which is specifically used to downsample and linearly transform the i-1-th feature vector through the T convolutional layers to obtain a downsampled feature vector; and perform a nonlinear transformation on the downsampled feature vector through the activation layer to obtain the i-th feature vector.

[0354] In some embodiments, one of the T convolutional layers is a downsampling convolutional layer, and the other layers are non-downsampling convolutional layers.

[0355] In some embodiments, the encoding feature processing module includes at least one residual unit.

[0356] In some embodiments, the encoding network includes an input layer and a transformation unit 23, which is specifically used to process the audio data through the input layer to obtain a first feature vector before downsampling and feature transformation processing are performed on the audio data through the multiple encoding feature processing modules to obtain the encoding vector of the audio data; and then downsampling and feature transformation processing are performed on the first feature vector through the multiple encoding feature processing modules to obtain the encoding vector.

[0357] In some embodiments, the encoding network includes an output layer, a transformation unit 23, which is specifically used to downsample and transform the audio data through the multiple encoding feature processing modules to obtain a second feature vector; and process the second feature vector through the output layer to obtain the encoding vector.

[0358] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Fig. 22 The device shown can execute the above-mentioned audio encoding method embodiment, and the above-mentioned and other operations and / or functions of each module in the device are respectively for realizing the above-mentioned method embodiment, which will not be described in detail here for the sake of brevity.

[0359] The above describes the device of the embodiment of the present application from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software modules in the decoding processor to perform. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0360] Fig.23 is a schematic block diagram of an electronic device provided in an embodiment of the present application, Fig.23 The electronic device may be the above-mentioned encoding device or the decoding device.

[0361] like Fig.23 As shown, the electronic device 30 may include:

[0362] The memory 31 and the processor 32, the memory 31 is used to store the computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiment of the present application.

[0363] For example, the processor 32 may be configured to execute the steps in the method 200 according to the instructions in the computer program 33 .

[0364] In some embodiments of the present application, the processor 32 may include but is not limited to:

[0365] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0366] In some embodiments of the present application, the memory 31 includes but is not limited to:

[0367] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0368] In some embodiments of the present application, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording pages provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 33 in the electronic device.

[0369] like Fig.23 As shown, the electronic device 30 may further include:

[0370] The transceiver 34 may be connected to the processor 32 or the memory 31 .

[0371] The processor 32 may control the transceiver 34 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.

[0372] It should be understood that the various components in the computing device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0373] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.

[0374] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.

[0375] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a solid state drive (solid state disk, SSD)), etc.

[0376] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0377] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0378] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0379] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An audio decoding method, characterized in that: include: Decoding the bitstream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data, wherein the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through a coding network; De-quantizing the quantization result to obtain a reconstructed coding vector of the audio data; The reconstructed coding vector is subjected to at least one of upsampling and feature transformation processing through a decoding network to obtain a reconstructed value of the audio data, the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples, the decoding feature processing module is used to perform at least one of upsampling and feature transformation processing on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation processing on the input feature vector.

2. The method according to claim 1, characterized in that: The decoding network comprises a plurality of decoding feature processing modules, at least one of the plurality of decoding feature processing modules is used to perform upsampling processing on an input feature vector; The step of performing at least one of upsampling and feature transformation on the reconstructed coding vector through a decoding network to obtain a reconstructed value of the audio data includes: The reconstructed coding vector is upsampled and feature transformed through the multiple decoding feature processing modules to obtain the reconstructed value of the audio data.

3. The method according to claim 2, characterized in that The step of upsampling and feature transforming the reconstructed coding vector by the multiple decoding feature processing modules to obtain the reconstructed value of the audio data includes: Performing at least one of upsampling and feature transformation on the i-1th feature vector through the i-th decoding feature processing module to obtain the i-th feature vector, wherein i is a positive integer, and if i is equal to 1, the i-1th feature vector is the reconstructed coding vector, and if i is greater than 1, the i-1th feature vector is the feature vector output by the i-1th decoding feature processing module; The i+1th decoding feature processing module performs at least one of upsampling and feature transformation on the i-th feature vector, and the process is repeated. Based on the feature vector output by the last decoding feature processing module, the reconstructed value of the audio data is determined.

4. The method according to claim 3, characterized in that The i-th decoding feature processing module includes R deconvolution layers, where R is a positive integer; The step of performing at least one of upsampling and feature transformation on the i-1th feature vector by the i-th decoding feature processing module to obtain the i-th feature vector comprises: The i-th feature vector is obtained by performing at least one of upsampling and feature transformation on the i-1th feature vector through the R deconvolution layers.

5. The method according to claim 4, characterized in that At least one deconvolution layer among the R deconvolution layers is an upsampling deconvolution layer; The step of performing at least one of upsampling and feature transformation on the i-1th feature vector through the R deconvolution layers to obtain the i-th feature vector comprises: The i-1th feature vector is up-sampled and feature transformed by the R deconvolution layers to obtain the i-th feature vector.

6. The method according to claim 5, characterized in that The i-th decoding feature processing module further includes an activation layer, and the i-th feature vector is upsampled and feature transformed by the R deconvolution layers to obtain the i-th feature vector, including: Performing upsampling and linear transformation processing on the i-1th feature vector through the R deconvolution layers to obtain an upsampled feature vector; The upsampled feature vector is nonlinearly transformed through the activation layer to obtain the i-th feature vector.

7. The method according to claim 5, characterized in that One of the R deconvolution layers is an upsampling deconvolution layer, and the other layers are non-upsampling deconvolution layers.

8. The method according to any one of claims 2 to 7, characterized in that: The decoding feature processing module includes at least one residual unit.

9. The method according to any one of claims 2 to 7, characterized in that: The decoding network includes an input layer. Before the reconstructed coding vector is upsampled and feature transformed by the multiple decoding feature processing modules to obtain the reconstructed value of the audio data, the method further includes: Processing the reconstructed coding vector through an input layer to obtain a first eigenvector; The step of upsampling and feature transforming the reconstructed coding vector by the multiple decoding feature processing modules to obtain the reconstructed value of the audio data includes: The first feature vector is upsampled and feature transformed by the multiple decoding feature processing modules to obtain a reconstructed value of the audio data.

10. The method according to any one of claims 2 to 7, characterized in that: The decoding network includes an output layer, and the reconstructed coding vector is upsampled and feature transformed by the multiple decoding feature processing modules to obtain the reconstructed value of the audio data, including: The reconstructed coding vector is subjected to upsampling and feature transformation processing by the multiple decoding feature processing modules to obtain a second feature vector; The second feature vector is processed through the output layer to obtain a reconstructed value of the audio data.

11. An audio encoding method, characterized in that: include: Get the audio data to be encoded; The audio data is subjected to at least one of downsampling and feature transformation by a coding network to obtain a coding vector of the audio data, wherein the total downsampling multiple of the coding network is consistent with the total upsampling multiple of the decoding network, and the coding feature processing modules included in the coding network are different from the decoding feature processing modules included in the decoding network in at least one of the number, network structure, and sampling rate, the coding feature processing modules are used to perform at least one of downsampling and feature transformation on the input feature vector, and the decoding feature processing modules are used to perform at least one of upsampling and feature transformation on the input feature vector; The coding vector of the audio data is quantized to obtain a quantization result, and the quantization result is encoded to obtain a code stream.

12. The method according to claim 11, characterized in that The step of performing at least one of downsampling and feature transformation on the audio data through a coding network to obtain a coding vector of the audio data includes: Performing at least one of downsampling and feature transformation on the i-1th feature vector through the i-th coding feature processing module to obtain the i-th feature vector, wherein i is a positive integer, and if i is equal to 1, the i-1th feature vector is the audio data, and if i is greater than 1, the i-1th feature vector is the feature vector output by the i-1th coding feature processing module; The i+1th coding feature processing module performs at least one of downsampling and feature transformation on the i-th feature vector, and the process is repeated to determine the coding vector of the audio data based on the feature vector output by the last coding feature processing module.

13. The method according to claim 12, characterized in that The i-th encoding feature processing module includes T convolutional layers, where T is a positive integer; The step of performing at least one of downsampling and feature transformation on the i-1th feature vector by the i-th encoding feature processing module to obtain the i-th feature vector comprises: The i-th feature vector is obtained by performing at least one of downsampling and feature transformation on the i-1th feature vector through the T convolutional layers.

14. The method according to claim 13, characterized in that At least one convolutional layer among the T convolutional layers is a downsampling convolutional layer; The step of performing at least one of downsampling and feature transformation on the i-1th feature vector through the T convolutional layers to obtain the i-th feature vector comprises: The i-1th feature vector is downsampled and linearly transformed through the T convolutional layers to obtain the i-th feature vector.

15. The method according to claim 13 or 14, characterized in that The i-th encoding feature processing module further includes an activation layer, and the i-1-th feature vector is downsampled and feature transformed by the T convolutional layers to obtain the i-th feature vector, including: Downsampling and linear transformation processing are performed on the i-1th feature vector through the T convolutional layers to obtain a downsampled feature vector; The downsampled feature vector is nonlinearly transformed through the activation layer to obtain the i-th feature vector.

16. The method according to any one of claims 11 to 14, characterized in that: The encoding network includes an input layer, and performing at least one of downsampling and feature transformation processing on the audio data through the encoding network to obtain an encoding vector of the audio data includes: Processing the audio data through an input layer to obtain a first feature vector; The first feature vector is downsampled and feature transformed through multiple encoding feature processing modules to obtain the encoding vector.

17. The method according to any one of claims 11 to 14, characterized in that: The encoding network includes an output layer, and performing at least one of downsampling and feature transformation processing on the audio data through the encoding network to obtain an encoding vector of the audio data includes: The audio data is downsampled and feature transformed by multiple encoding feature processing modules to obtain a second feature vector; The second feature vector is processed by the output layer to obtain the encoding vector.

18. An audio decoding device, characterized in that: include: A decoding unit, configured to decode a bit stream to obtain a quantization result corresponding to the audio data, wherein the quantization result is obtained by quantizing a coding vector of the audio data, wherein the coding vector is obtained by performing at least one of downsampling and feature transformation on the audio data through a coding network; A dequantization unit, used for dequantizing the quantization result to obtain a reconstructed coding vector of the audio data; A reconstruction unit, used to perform at least one of upsampling and feature transformation processing on the reconstructed coding vector through a decoding network to obtain a reconstructed value of the audio data, wherein the total upsampling multiple of the decoding network is consistent with the total downsampling multiple of the encoding network, and the decoding feature processing modules included in the decoding network are different from the encoding feature processing modules included in the encoding network in at least one of the number, network structure, and sampling multiples, the decoding feature processing module is used to perform at least one of upsampling and feature transformation processing on the input feature vector, and the encoding feature processing module is used to perform at least one of downsampling and feature transformation processing on the input feature vector.

19. An audio encoding device, characterized in that: include: An acquisition unit, used for acquiring audio data to be encoded; a transform unit, configured to perform at least one of downsampling and feature transformation on the audio data through an encoding network to obtain an encoding vector of the audio data, wherein a total downsampling multiple of the encoding network is consistent with a total upsampling multiple of the decoding network, and an encoding feature processing module included in the encoding network is different from a decoding feature processing module included in the decoding network in at least one of quantity, network structure, and sampling rate, the encoding feature processing module is configured to perform at least one of downsampling and feature transformation on an input feature vector, and the decoding feature processing module is configured to perform at least one of upsampling and feature transformation on an input feature vector; The encoding unit is used to quantize the encoding vector of the audio data to obtain a quantization result, and encode the quantization result to obtain a code stream.

20. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 10 or 11 to 17.