Audio encoding method and apparatus, audio decoding method and apparatus, and readable storage medium

By adopting multiple encoding modes and code rate modes in audio encoding and decoding, the problem of single audio decoding quality in the prior art is solved, and the output of reconstructed audio signals of different quality levels is realized to meet the diverse needs of users.

WO2025113123A1PCT designated stage expired Publication Date: 2025-06-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2024/130151
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-06
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing audio decoding technologies cannot meet users' needs for reconstructing audio signals of different quality levels, resulting in a single audio decoding quality.

Method used

By adopting multiple encoding modes and code rate modes in the audio encoding and decoding process, the encoding characteristics of the audio signal are extracted, and the signal encoding and decoding is performed according to the target encoding mode and code rate mode, reconstructed audio signals of different quality levels are generated.

Benefits of technology

It realizes that while ensuring the audio decoding efficiency, it outputs reconstructed audio signals of different quality levels, improves the diversified quality of audio signals, and meets the actual application needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130151_05062025_PF_FP_ABST
    Figure CN2024130151_05062025_PF_FP_ABST
Patent Text Reader

Abstract

An audio encoding method and apparatus, an audio decoding method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The audio decoding method comprises: acquiring an audio code stream encapsulation, wherein the audio code stream encapsulation comprises an audio code stream, the audio code stream is obtained by performing audio encoding on an audio signal on the basis of a target encoding mode and a target code rate mode, the target encoding mode is acquired from among a plurality of candidate encoding modes, and the target code rate mode is acquired from among a plurality of candidate code rate modes; in response to a decoding request for the audio code stream encapsulation, acquiring the target encoding mode and the target code rate mode from a frame header comprised in the audio code stream encapsulation; performing signal decoding on the audio code stream on the basis of the target encoding mode and the target code rate mode, so as to obtain an encoding feature estimation value corresponding to the audio code stream; and performing reconstruction on the encoding feature estimation value on the basis of the target encoding mode, so as to obtain a reconstructed audio signal corresponding to the audio code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Audio encoding method, audio decoding method, device, and readable storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The embodiments of this application are based on the Chinese patent application with application number 202311614893.4 and application date November 29, 2023, and claim the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into the embodiments of this application as a reference. Technical Field

[0003] The present application relates to artificial intelligence technology, and in particular to an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0004] Audio coding and decoding technology is a key application in the field of artificial intelligence and a core technology for communication services, including remote audio and video calls. Simply put, speech coding technology aims to transmit as much voice information as possible using less network bandwidth. From the perspective of Shannon's information theory, speech coding is a form of source coding. The goal of source coding is to compress the data volume as much as possible at the encoding end, removing redundancy, while enabling lossless (or near-lossless) recovery at the decoding end.

[0005] In the related art, the quality of the audio generated by decoding at the decoding end is relatively simple and cannot meet user needs.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can output reconstructed audio signals of different quality levels while ensuring the efficiency of audio decoding.

[0008] The technical solution of the embodiment of the present application is implemented as follows:

[0009] The present application provides an audio encoding method, which is applied to an electronic device and includes:

[0010] In response to an encoding request for an audio signal, obtaining a target encoding mode for the audio signal from a plurality of encoding modes, and obtaining a target bit rate mode for the audio signal from a plurality of bit rate modes;

[0011] extracting the coding features of the audio signal from the audio signal using the target coding mode;

[0012] Performing signal encoding on the coding feature using the target bit rate mode to obtain an audio code stream of the audio signal;

[0013] Determining a frame header based on the target coding mode and the target bit rate mode;

[0014] An audio code stream encapsulation of the audio signal is generated based on the audio code stream and the frame header.

[0015] The present invention provides an audio decoding method for an electronic device, including:

[0016] Obtaining an audio stream package, wherein the audio stream package includes an audio stream, the audio stream being obtained by encoding an audio signal using a target coding mode and a target bit rate mode, wherein the target coding mode is obtained from a plurality of candidate coding modes, and the target bit rate mode is obtained from a plurality of candidate bit rate modes;

[0017] In response to a decoding request for the audio stream encapsulation, obtaining the target coding mode and the target bit rate mode from a frame header included in the audio stream encapsulation;

[0018] Decoding the audio code stream using the target coding mode and the target bit rate mode to obtain an estimated coding feature value corresponding to the audio code stream;

[0019] The coding feature estimation value is reconstructed through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

[0020] The present invention provides an audio encoding device, including:

[0021] a second acquisition module configured to, in response to an encoding request for an audio signal, acquire a target encoding mode for the audio signal from a plurality of encoding modes, and acquire a target bit rate mode for the audio signal from a plurality of bit rate modes;

[0022] an extraction module configured to extract the coding features of the audio signal from the audio signal using the target coding mode;

[0023] a signal encoding module configured to perform signal encoding on the encoding feature using the target bit rate mode to obtain an audio code stream of the audio signal;

[0024] A construction module configured to determine a frame header based on the target coding mode and the target bit rate mode;

[0025] The generating module is configured to generate an audio code stream encapsulation of the audio signal based on the audio code stream and the frame header.

[0026] The present invention provides an audio decoding device, including:

[0027] a first acquisition module configured to acquire an audio stream package, wherein the audio stream package includes an audio stream, and the audio stream is obtained by audio encoding an audio signal using a target coding mode and a target bit rate mode, wherein the target coding mode is obtained from a plurality of candidate coding modes, and the target bit rate mode is obtained from a plurality of candidate bit rate modes;

[0028] In response to a decoding request for the audio stream encapsulation, obtaining the target coding mode and the target bit rate mode from a frame header included in the audio stream encapsulation;

[0029] a signal decoding module configured to perform signal decoding on the audio stream using the target coding mode and the target bit rate mode to obtain a coding feature estimation value corresponding to the audio stream;

[0030] The reconstruction module is configured to reconstruct the coding feature estimation value through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

[0031] An embodiment of the present application provides an electronic device, comprising:

[0032] a memory for storing computer-executable instructions;

[0033] The processor is used to implement the audio encoding method or audio decoding method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0034] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implements the audio encoding method or audio decoding method provided in the embodiment of the present application.

[0035] An embodiment of the present application provides a computer program product, including computer-executable instructions, which, when executed by a processor, implement the audio encoding method or audio decoding method provided in the embodiment of the present application.

[0036] The embodiments of the present application have the following beneficial effects:

[0037] By decoding the audio code stream using different coding modes and bit rate modes, we can obtain coding feature estimation values ​​of different precisions. These coding feature estimation values ​​of different precisions are then reconstructed to obtain reconstructed audio signals of different quality levels, thereby improving the diversity of the quality of the reconstructed audio signals to meet the actual application needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] FIG1 is a schematic diagram showing a comparison of spectrums at different bit rates provided in an embodiment of the present application;

[0039] FIG2 is a schematic diagram of the architecture of an audio codec system provided in an embodiment of the present application;

[0040] FIG3A is a first structural diagram of an electronic device provided in an embodiment of the present application;

[0041] FIG3B is a second structural diagram of an electronic device provided in an embodiment of the present application;

[0042] FIG4A is a schematic diagram of a first flow chart of an audio encoding method provided in an embodiment of the present application;

[0043] FIG4B is a schematic diagram of a second flow chart of the audio encoding method provided in an embodiment of the present application;

[0044] FIG4C is a schematic diagram of a third flow chart of the audio encoding method provided in an embodiment of the present application;

[0045] FIG4D is a schematic diagram of a fourth flow chart of the audio encoding method provided in an embodiment of the present application;

[0046] FIG4E is a schematic diagram of a fifth flow chart of the audio encoding method provided in an embodiment of the present application;

[0047] FIG5A is a schematic diagram of a first flow chart of an audio decoding method provided in an embodiment of the present application;

[0048] FIG5B is a schematic diagram of a second flow chart of the audio decoding method provided in an embodiment of the present application;

[0049] FIG5C is a schematic diagram of a third flow chart of the audio decoding method provided in an embodiment of the present application;

[0050] FIG6A is a schematic diagram of a channel without group convolution provided in an embodiment of the present application;

[0051] FIG6B is a schematic diagram of a channel using grouped convolution according to an embodiment of the present application;

[0052] FIG6C is a schematic diagram of a voice communication link provided in an embodiment of the present application;

[0053] 7 is a schematic diagram of a flow chart of a multi-mode multi-rate speech encoding method according to an embodiment of the present application;

[0054] FIG8 is a schematic diagram of a filter bank provided in an embodiment of the present application;

[0055] FIG9A is a schematic diagram of a common convolutional network provided in an embodiment of the present application;

[0056] FIG9B is a schematic diagram of a dilated convolutional network provided in an embodiment of the present application;

[0057] FIG10 is a schematic diagram of frequency band extension provided in an embodiment of the present application;

[0058] FIG11 is a schematic diagram of a third neural network provided in an embodiment of the present application;

[0059] FIG12A is a schematic diagram of a residual block structure used in a coding block provided in an embodiment of the present application;

[0060] FIG12B is a schematic diagram of the residual unit structure provided in an embodiment of the present application;

[0061] FIG13A is a schematic diagram of the structure of code stream encapsulation in a broadband coding mode provided by an embodiment of the present application;

[0062] FIG13B is a schematic structural diagram of a code stream encapsulation that does not include residuals and flatness edge information in an ultra-wideband coding mode provided by an embodiment of the present application;

[0063] FIG13C is a schematic diagram showing the structure of a code stream encapsulation including a residual and excluding flatness edge information in an ultra-wideband coding mode according to an embodiment of the present application;

[0064] FIG13D is a schematic diagram showing the structure of a code stream encapsulation that does not include residuals but includes flatness edge information in an ultra-wideband coding mode according to an embodiment of the present application;

[0065] FIG13E is a schematic diagram showing the structure of a code stream encapsulation including residual and flatness edge information in an ultra-wideband coding mode provided by an embodiment of the present application;

[0066] FIG14 is a schematic diagram of a first neural network provided in an embodiment of the present application;

[0067] FIG15A is a schematic structural diagram of a code stream encapsulation in a wideband stereo coding mode provided by an embodiment of the present application;

[0068] FIG15B is a schematic structural diagram of a bitstream encapsulation that does not include residuals and flatness edge information in an ultra-wideband stereo coding mode provided by an embodiment of the present application;

[0069] FIG15C is a schematic diagram showing the structure of a bitstream encapsulation including a residual but not flatness edge information in an ultra-wideband stereo coding mode provided by an embodiment of the present application;

[0070] FIG15D is a schematic diagram showing the structure of a bitstream encapsulation that does not include residuals but includes flatness side information in an ultra-wideband stereo coding mode provided by an embodiment of the present application;

[0071] FIG15E is a schematic structural diagram of a code stream encapsulation including residual and flatness edge information in an ultra-wideband stereo coding mode provided by an embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0073] In the following description, the terms "first\second" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0074] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0075] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0077] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0078] 1) Neural Network (NN): A mathematical model that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing. This network relies on the complexity of the system to process information by adjusting the connections between its numerous nodes.

[0079] 2) Deep Learning (DL): A new research direction in the field of machine learning (ML), deep learning studies the inherent patterns and representational hierarchies of sample data. The information gained from this learning process is highly helpful in interpreting data such as text, images, and sounds. Its ultimate goal is to enable machines to have human-like analytical learning capabilities and to recognize data such as text, images, and sounds.

[0080] 3) Quantization: This is the process of approximating a signal's continuous values ​​(or a large number of discrete values) to a finite number (or a small number) of discrete values. Quantization includes vector quantization (VQ) and scalar quantization.

[0081] Vector quantization is an effective lossy compression technique, based on Shannon's rate-distortion theory. The basic principle of vector quantization is to replace the input vector with the index of the codeword in the code table that best matches it (also known as the quantization value) for transmission and storage, while decoding requires only a simple table lookup. For example, several scalar data items form a vector space, which is then divided into several small regions. During quantization, the vectors falling into the small regions are replaced with the corresponding indices of the input vectors.

[0082] Scalar quantization is the quantization of scalars, i.e. one-dimensional vector quantization, which divides the dynamic range into several small intervals, each of which has a representative value (i.e., index). When the input signal falls into a certain interval, the input signal is quantized to the representative value.

[0083] 4) Entropy Coding: This is a lossless coding method that uses the entropy principle to prevent any information loss during the encoding process. It is also a key module in lossy coding and is located at the end of the encoder. Entropy coding includes Shannon coding, Huffman coding, Exp-Golomb coding, and arithmetic coding.

[0084] 5) Quadrature Mirror Filters (QMF): This is a filter pair that combines analysis and synthesis. The QMF analysis filter decomposes the subband signal to reduce the signal bandwidth so that each subband signal can be processed smoothly through its own channel. The QMF synthesis filter synthesizes the subband signals recovered by the decoder, for example, through zero-value interpolation and bandpass filtering to reconstruct the original audio signal.

[0085] 6) Stream encapsulation: refers to packaging the encoded stream into a specific format for easy storage or network transmission. During the encapsulation process, the encoded and compressed stream is placed in a file according to a certain format to form a container. This container not only contains the stream, but may also contain some metadata, such as encoding type, bit rate, frame rate and other information. The encapsulation format (also known as the container format) can be regarded as a shell of the stream (audio stream, video stream, etc.), and the stream is stored in the file of this encapsulation format. The choice of encapsulation format depends on the specific application scenario and requirements. Different encapsulation formats support different encoding formats. For example, MP4, Flash video format (FLV, Adobe Flash Video), etc. are all common encapsulation formats.

[0086] 7) Wideband: This term reflects the resolution or sampling rate of an audio signal. According to standards organizations such as the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the 3rd Generation Partnership Project (3GPP), wideband has a sampling rate of 16,000 Hz and an effective bandwidth of up to 8,000 Hz. Wideband transmission rates are typically above 1.54 Mbps (megabits per second).

[0087] 8) Ultra-wideband: This term reflects the resolution or sampling rate of an audio signal. According to standards organizations such as ITU-T and 3GPP, ultra-wideband has a sampling rate of 32,000 Hz and an effective bandwidth of up to 16,000 Hz. Ultra-wideband is a wireless communication technology. Generally speaking, ultra-wideband refers to an audio signal with a bandwidth of at least 500 MHz, or a ratio of the audio signal's bandwidth to its center frequency exceeding 20%.

[0088] Before specifically introducing the audio encoding method and audio decoding method provided in the embodiments of the present application, the QMF filter bank, the dilated convolutional network, and the frequency band extension are first introduced.

[0089] The QMF filter bank is a filter pair that includes analysis and synthesis. For the QMF analysis filter, the input signal with a sampling rate of Fs can be decomposed into two signals with a sampling rate of Fs / 2, representing the QMF low-pass signal and the QMF high-pass signal respectively. The spectral response of the low-pass part H_Low(z) and the high-pass part H_High(z) of the QMF filter is shown in Figure 8. Based on the relevant theoretical knowledge of the QMF analysis filter bank, the correlation between the coefficients of the low-pass filter and the high-pass filter can be easily described, as shown in formula (1):

[0090] h High (k) = -1 k h Low (k) (1)

[0091] Among them, h Low (k) represents the coefficient of low-pass filtering, h High (k) represents the coefficient of high-pass filtering.

[0092] Similarly, according to QMF related theory, the QMF synthesis filter bank can be described based on the QMF analysis filter bank H_Low(z) and H_High(z), as shown in formula (2).

[0093] G Low (z)=H Low (z)

[0094] G High (z)=(-1)*H High (z) (2)

[0095] Among them, G Low (z) represents the recovered low-pass signal, G High (z) represents the recovered high-pass signal.

[0096] The low-pass and high-pass signals recovered by the decoding end are synthesized by the QMF synthesis filter bank, so that the reconstructed signal (ie, the synthesized signal) with the sampling rate Fs corresponding to the input signal can be restored.

[0097] Referring to Figures 9A and 9B, Figure 9A is a schematic diagram of a normal convolution (e.g., causal convolution) network provided in an embodiment of the present application, and Figure 9B is a schematic diagram of a dilated convolution network provided in an embodiment of the present application. Relative to ordinary convolution networks, dilated convolution can increase the receptive field while keeping the size of the feature map unchanged, and can also avoid errors caused by upsampling and downsampling. Although the convolution kernel sizes (Kernel Size) shown in Figures 9A and 9B are both 3×3; however, the receptive field 901 of the ordinary convolution shown in Figure 9A is only 3, while the receptive field 902 of the dilated convolution shown in Figure 9B reaches 5. That is to say, for a convolution kernel of size 3×3, the receptive field of the ordinary convolution shown in Figure 9A is 3, and the dilation rate (the number of intervals between points in the convolution kernel) is 1; while the receptive field of the dilated convolution shown in Figure 9B is 5, and the dilation rate is 2.

[0098] The convolution kernel can also be moved on a plane similar to Figure 9A or Figure 9B. This involves the concept of stride rate. For example, each time the convolution kernel shifts 1 square, the corresponding stride rate is 1.

[0099] There's also the concept of convolution channels, which refers to the number of convolution kernel parameters used to perform the convolution analysis. In theory, a greater number of channels provides a more comprehensive signal analysis and higher accuracy; however, a higher number of channels also increases complexity. For example, a 1×320 tensor can be convolved using 24 channels, resulting in a 24×320 tensor as the output.

[0100] It should be noted that the size of the dilated convolution kernel (for example, for speech signals, the size of the convolution kernel can be set to 1×3), the expansion rate, the shift rate and the number of channels can be defined by yourself according to actual application needs. The embodiments of this application do not make specific restrictions on this.

[0101] Figure 10 shows a schematic diagram of frequency band extension (or band replication). First, the wideband signal is reconstructed, then copied onto an ultra-wideband signal, and finally shaped based on the ultra-wideband envelope. The frequency domain implementation scheme shown in Figure 10 specifically includes: 1) implementing a core layer encoding at a low sampling rate; 2) selectively copying the low-frequency spectrum to the high-frequency portion; and 3) applying a gain to the copied high-frequency spectrum based on pre-recorded boundary information (such as describing the energy correlation between high and low frequencies). With a bit rate of only 1-2 kbps, the sampling rate can be doubled.

[0102] Speech coding technology aims to transmit as much voice information as possible while minimizing network bandwidth. Speech codecs can achieve compression ratios of over 10 times. This means that 10MB of speech data can be compressed by the codec to only 1MB, significantly reducing the bandwidth required to transmit the information. For example, for a wideband speech signal with a sampling rate of 16,000Hz, using a 16-bit sampling depth (the level of detail in recording speech intensity), the bit rate (the amount of data transmitted per unit time) of the uncompressed version is 256kbps. Using speech coding technology, even with lossy encoding, the quality of the reconstructed speech signal within a bit rate range of 10-20kbps can be close to that of the uncompressed version, and may even be perceived as indistinguishable. For services requiring even higher sampling rates, such as ultra-wideband speech at 32,000Hz, the bit rate must be at least 30kbps.

[0103] To ensure smooth communication within communication systems, the industry deploys standard voice codec protocols, such as those from international and domestic standards organizations like ITU-T, 3GPP, IETF, AVS, and CCSA, as well as standards like G.711, G.722, the AMR series, EVS, and OPUS. Figure 1 shows a schematic diagram comparing spectra at different bit rates, demonstrating the relationship between compression bit rate and quality. Curve 101 shows the spectrum of the original speech (i.e., the uncompressed signal); Curve 102 shows the spectrum of the OPUS encoder at a bit rate of 20 kbps; and Curve 103 shows the spectrum of the OPUS encoder at a bit rate of 6 kbps. As shown in Figure 1, as the encoding bit rate increases, the compressed signal becomes closer to the original signal.

[0104] The principles of speech coding are roughly as follows: speech coding can directly encode speech waveform samples sample by sample; or, based on the principle of human vocalization, relevant low-dimensional features are extracted, the encoding end encodes the features, and the decoding end reconstructs the speech signal based on these parameters.

[0105] The above coding principles are all derived from voice signal modeling, that is, compression methods based on signal processing, which cannot guarantee the encoding quality of audio. In order to improve coding efficiency while ensuring voice quality, embodiments of the present application provide an audio encoding method, an audio decoding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The following describes exemplary applications of electronic devices provided by embodiments of the present application. The electronic devices provided by embodiments of the present application can be implemented as terminal devices, or as servers, or implemented in collaboration with terminal devices and servers. The following describes an example of an electronic device implemented as a terminal device.

[0106] For example, see Figure 2, which is a schematic diagram of the architecture of the audio encoding and decoding system 10 provided in an embodiment of the present application. The audio decoding system 10 includes: a server 200, a network 300, a terminal device 400 (i.e., an encoding end) and a terminal device 500 (i.e., a decoding end), wherein the network 300 can be a local area network, a wide area network, or a combination of the two.

[0107] In some embodiments, a client 410 runs on the terminal device 400. The client 410 can be various types of clients, such as an instant messaging client, a web conferencing client, a live broadcast client, a browser, etc. In response to an audio collection instruction triggered by a sender (such as the initiator of a web conferencing session, a host, or the initiator of a voice call), the client 410 calls the microphone of the terminal device 400 to collect audio signals, and performs audio encoding processing on the collected audio signals to obtain an audio stream.

[0108] For example, the client 410 calls the audio encoding method provided in an embodiment of the present application to encode the collected audio signal, that is, obtains a target encoding mode for the audio signal from multiple encoding modes, and obtains a target bit rate mode for the audio signal from multiple bit rate modes; through the target encoding mode, extracts the encoding features of the audio signal from the audio signal; through the target bit rate mode, encodes the encoding features of the audio signal to obtain an audio code stream of the audio signal; determines a frame header based on the target encoding mode and the target bit rate mode; and generates an audio code stream encapsulation of the audio signal based on the audio code stream and the frame header.

[0109] The client 410 can send the audio code stream package to the server 200 via the network 300, so that the server 200 sends the audio code stream package to the terminal device 500 associated with the recipient (such as a participant, audience, or recipient of a voice call in a web conference).

[0110] After receiving the audio code stream package sent by the server 200, the client 510 (such as an instant messaging client, a web conferencing client, a live broadcast client, a browser, etc.) running on the terminal device 500 can perform audio decoding processing on the audio code stream package to obtain a reconstructed audio signal, thereby realizing audio communication.

[0111] For example, the client 510 calls the audio decoding method provided in an embodiment of the present application to decode the received audio stream encapsulation, that is, obtains the target coding mode and the target bit rate mode from the frame header included in the audio stream encapsulation; wherein, the audio stream included in the audio stream encapsulation is obtained by audio encoding the audio signal through the target coding mode and the target bit rate mode, the target coding mode is obtained from multiple coding modes, and the target bit rate mode is obtained from multiple bit rate modes; through the target coding mode and the target bit rate mode, the audio stream is signal decoded to obtain the coding feature estimation value corresponding to the audio stream; through the target coding mode, the coding feature estimation value corresponding to the audio stream is reconstructed to obtain the reconstructed audio signal corresponding to the audio stream.

[0112] For example, the server 200 shown in Figure 2 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices 400 and 500 shown in Figure 2 can be smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, car terminals, etc., but are not limited to these. The terminal devices (such as terminal devices 400 and terminal devices 500) and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0113] In some embodiments, the terminal device or server 200 can also implement the audio encoding method or audio decoding method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a live broadcast APP, a web conferencing APP, or an instant messaging APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module or plug-in in any form.

[0114] 3A , which is a schematic diagram of the structure of an electronic device 500 provided in an embodiment of the present application. Taking the electronic device 500 as a terminal device as an example, the electronic device 500 shown in FIG3A includes: at least one processor 520, a memory 550, at least one network interface 530, and a user interface 540. The various components in the electronic device 500 are coupled together via a bus system 550. It will be understood that the bus system 550 is used to implement connection and communication between these components. In addition to including a data bus, the bus system 550 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in FIG3A , various buses are labeled as the bus system 550.

[0115] The processor 520 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0116] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 520.

[0117] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0118] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0119] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0120] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 530. Exemplary network interfaces 530 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).

[0121] In some embodiments, the audio encoding device provided in the embodiments of the present application can be implemented in software. Figure 3A shows an audio encoding device 555 stored in the memory 550, which can be software in the form of programs and plug-ins, including the following software modules: a second acquisition module 5551, an extraction module 5552, a signal encoding module 5553, a construction module 5554 and a generation module 5555, wherein the second acquisition module 5551, the extraction module 5552, the signal encoding module 5553, the construction module 5554 and the generation module 5555 are used to implement the audio encoding function. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented.

[0122] Referring to FIG. 3B , FIG. 3B is a schematic diagram of the structure of an electronic device 600 provided in an embodiment of the present application. For example, the electronic device 600 is a terminal device. The electronic device 600 shown in FIG. 3B includes at least one processor 620, a memory 650, at least one network interface 630, and a user interface 640. The various components in the electronic device 600 are coupled together via a bus system 650. The memory 650 includes an operating system 651 and a network communication module 652. It should be noted that the functions of the structure in FIG. 3B are similar to those of the structure in FIG. 3A . The audio encoding device provided in an embodiment of the present application can be implemented in software. FIG. 3B shows an audio decoding device 655 stored in the memory 650. This device can be software in the form of a program or plug-in, and includes the following software modules: a first acquisition module 6551, a signal decoding module 6552, and a reconstruction module 6553. The first acquisition module 6551, the signal decoding module 6552, and the reconstruction module 6553 are used to implement audio decoding functionality. These modules are logically connected and can be arbitrarily combined or further separated depending on the functionality implemented.

[0123] As previously mentioned, the audio encoding method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to FIG. 4A , FIG. 4A is a schematic flow chart of the audio encoding method provided in the embodiments of the present application. The audio encoding function is implemented by the audio encoding method, and the following description is provided in conjunction with steps 101 to 105 shown in FIG. 4A .

[0124] Before describing the following steps, the audio encoding method provided in the embodiment of the present application is explained. The audio encoding method provided in the embodiment of the present application includes two encoding methods (i.e., the signal encoding method in the signal processing technology and the feature encoding (i.e., feature extraction) in the artificial intelligence technology. The feature encoding in the artificial intelligence technology is detailed in step 102 below, and the signal encoding method in the signal processing technology is detailed in step 103 below.

[0125] In step 101, in response to an encoding request for an audio signal, a target encoding mode for the audio signal is acquired from a plurality of encoding modes, and a target bit rate mode for the audio signal is acquired from a plurality of bit rate modes.

[0126] The encoding request is used to instruct the audio signal to be encoded. The encoding request includes encoding configuration information, such as the sampling rate of the audio signal, the type of audio signal (e.g., wideband signal, ultra-wideband signal), etc. The multiple encoding modes include wideband encoding mode and ultra-wideband encoding mode. The wideband encoding mode is used to indicate that the audio signal is encoded with wideband, indicating that the audio signal is a wideband signal. Compared with ultra-wideband signals, the sampling rate of wideband signals is low, so there is no need to decompose the audio signal into sub-bands. The ultra-wideband encoding mode is used to indicate that the audio signal is encoded with ultra-wideband, indicating that the audio signal is an ultra-wideband signal. Compared with wideband signals, the sampling rate of ultra-wideband signals is high, and the audio signal can be decomposed into sub-bands and then encoded. The rate mode is used to indicate the use of a specified code table and corresponding code rate for signal encoding or decoding. Different rate modes correspond to different code tables and code rates. During the configuration stage, different encoding modes and rate modes can be configured for different audio signals.

[0127] As an example of obtaining an audio signal, the encoding end responds to the audio collection instruction triggered by the sender (such as the initiator of the online conference, the host, the initiator of the voice call, etc.), and calls the microphone of the terminal device of the encoding end to collect the audio signal to obtain the audio signal (also called the input signal).

[0128] As an example of obtaining an encoding request for an audio signal, the encoding end selects a target encoding mode for the audio signal from multiple encoding modes in response to the sender's selection operation for multiple encoding models. It can also select a target bit rate mode for the audio signal from multiple bit rate modes in response to the sender's selection operation for multiple bit rate modes, and generate an encoding request for the audio signal based on the target encoding mode and the target bit rate mode, thereby configuring different encoding modes and bit rate modes for different audio signals according to user needs.

[0129] As another example of obtaining an encoding request for an audio signal, the encoding end may pre-store a configuration table, wherein the configuration table stores a correspondence between different configuration information and different encoding modes, as well as a correspondence between different configuration information and different bit rate modes. After the encoding end obtains the audio signal, the encoding request for the audio signal is generated based on the configuration information of the audio signal, so as to subsequently obtain the configuration information of the audio signal from the encoding request, and query the configuration table based on the configuration information, determine the encoding mode corresponding to the queried configuration information of the audio signal as the target encoding mode, and determine the bit rate mode corresponding to the queried configuration information of the audio signal as the target bit rate mode.

[0130] In step 102, encoding features of the audio signal are extracted from the audio signal using a target encoding mode.

[0131] For example, when the target coding mode is the wideband coding mode, the coding features of the audio signal are extracted from the audio signal based on the wideband coding mode; when the target coding mode is the ultra-wideband coding mode, the coding features of the audio signal are extracted from the audio signal based on the ultra-wideband coding mode.

[0132] As shown in Figure 4B, Figure 4B is a flow chart of the audio encoding method provided in an embodiment of the present application. When the target encoding mode is the broadband encoding mode, Figure 4B shows that step 102 in Figure 4A can be implemented through steps 1021A-1022A: in step 1021A, the first neural network is called based on the broadband encoding mode; in step 1022A, the encoding features of the audio signal are extracted from the audio signal through the first neural network.

[0133] Here, when the target coding mode is a wideband coding mode, it indicates that the audio signal is a wideband signal (low-frequency signal), and artificial intelligence technology can be used to directly extract the coding features of the audio signal, so that a corresponding audio code stream can be subsequently generated based on the coding features of the audio signal. It should be noted that the embodiments of the present application are not limited to the structure of the first neural network. The first neural network can be a convolutional neural network, a deep neural network, etc.

[0134] As shown in FIG4C , FIG4C is a schematic flow chart of an audio encoding method provided in an embodiment of the present application. FIG4C shows that step 1022A in FIG4B can be implemented through steps 10221A to 10222A:

[0135] In step 10221A, feature extraction is performed on the audio signal to obtain audio features of the audio signal.

[0136] Here, the embodiment of the present application can call a first neural network (NN) based on the audio signal, and extract audio features from the audio signal through the first neural network, so as to subsequently continue to perform feature refinement based on the audio features. It should be noted that the embodiment of the present application is not limited to the structure of the first NN. The first NN can be a convolutional neural network, a deep neural network, etc.

[0137] In step 10222A, residual processing is performed on the audio features using at least one residual unit included in the first neural network to obtain coding features of the audio signal.

[0138] Here, the first neural network includes 4 encoding blocks, and each encoding block includes 4 or 5 residual units.

[0139] In neural network models, residual units refer to a special structure used to build residual networks (Residual Networks, ResNets). Residual units are designed to address the problems of vanishing and exploding gradients during deep neural network training, and to help the network better learn features. Residual units introduce skip connections, which directly add the input to the output rather than simply passing it between layers. This skip connection allows the network to learn the residual function—the difference between the input and output—rather than directly learning the mapping relationship. This design makes the network easier to optimize and also helps alleviate the vanishing gradient problem.

[0140] Here, by performing residual processing on the audio features on the encoding side, based on the characteristics of residual processing, while ensuring comprehensive learning of the audio features, it is also possible to better utilize the shallow feature information of the audio features and avoid missing the shallow feature information of the audio features.

[0141] Based on the characteristics of the residual unit, the residual processing in step 10222A is used to calculate the residual of the audio feature on the encoding side, and determine the residual of the audio feature as the encoding feature for subsequent signal encoding. For example, the residual of the audio feature is obtained by adding the audio feature to the output of the residual unit, that is, the audio feature is used as the input of the residual unit. After the audio feature is processed by the residual unit, the output of the residual unit is obtained, and through the jump connection characteristics of the residual unit, the input of the residual unit and the output of the residual unit are added to obtain the residual of the audio feature.

[0142] As shown in FIG4D , FIG4D is a flow chart of an audio encoding method provided in an embodiment of the present application. When the target encoding mode is an ultra-wideband encoding mode, FIG4D shows that step 102 in FIG4A can be implemented through steps 1021B to 1024B:

[0143] In step 1021B, the audio signal is decomposed into sub-bands to obtain a low-frequency sub-band signal and a high-frequency sub-band signal of the audio signal.

[0144] Here, when the target coding mode is the ultra-wideband coding mode, it indicates that the audio signal is an ultra-wideband signal. The ultra-wideband signal can be decomposed into sub-bands to obtain low-frequency sub-band signals and high-frequency sub-band signals of the audio signal.

[0145] It should be noted that the embodiment of the present application does not limit the frequency bands of the low-frequency sub-band signal and the high-frequency sub-band signal. That is, the low-frequency sub-band signal and the high-frequency sub-band signal obtained by decomposition can be two sub-band signals obtained by evenly dividing the frequency band of the audio signal, or can be two sub-band signals obtained by unevenly dividing the frequency band of the audio signal. For example, if the effective bandwidth of the audio signal x(n) is 0-16kHz, then the low-frequency sub-band signal x(n) is 0-16kHz. LB (n) and high frequency sub-band signal x HB The effective bandwidth of (n) is 0-8kHz and 8-16kHz respectively, and the low-frequency sub-band signal x LB (n) and high frequency sub-band signal x HB The effective bandwidth of (n) can also be 0-6 kHz and 6-16 kHz, respectively. In addition, the embodiment of the present application does not limit the number of divided frequency bands. That is, the frequency band of the audio signal can be divided evenly or unevenly to obtain two sub-band signals, or the frequency band of the audio signal can be divided evenly or unevenly to obtain more than two sub-band signals, for example, 3, 4, or more sub-band signals.

[0146] The audio signal includes a low-frequency part and a high-frequency part. The low-frequency signal (i.e., the low-frequency sub-band signal) is based on the characteristics of the audio signal and is separated from the audio signal of a specific sampling rate through a filter. The high-frequency signal (i.e., the high-frequency sub-band signal) is separated from the audio signal of a specific sampling rate. For example, if the effective bandwidth of the audio signal x(n) is 0-16kHz, then the effective bandwidth of the low-frequency signal is 0-8kHz, and the effective bandwidth of the high-frequency signal x(n) is 0-16kHz. HB The effective bandwidth of (n) may be 6-16 kHz. In addition, the embodiment of the present application does not limit the frequency band division of the audio signal. For example, the audio signal may be divided evenly or unevenly to obtain uniform low-frequency signals and high-frequency signals.

[0147] As an example of subband decomposition, the audio signal is decomposed into low-frequency subband signals x by a QMF analysis filter. LB (n) and high frequency sub-band signal x HB (n) Since the low-frequency sub-band signal has a greater impact on audio coding than the high-frequency sub-band signal, differential signal processing can be performed on the low-frequency sub-band signal and the high-frequency sub-band signal subsequently.

[0148] In some embodiments, step 1021B can be implemented by: sampling the audio signal to obtain a sampled signal, wherein the sampled signal includes multiple sample points obtained by sampling; low-pass filtering the sampled signal to obtain a low-pass filtered signal; down-sampling the low-pass filtered signal to obtain a low-frequency sub-band signal of the audio signal; high-pass filtering the sampled signal to obtain a high-pass filtered signal; down-sampling the high-pass filtered signal to obtain a high-frequency sub-band signal of the audio signal.

[0149] Here, the audio signal is a continuous analog signal, the sampling signal is a discrete digital signal, and the sampling point is a sampling value obtained by sampling the audio signal.

[0150] In the field of digital signal processing, downsampling is used to reduce the sampling rate of an audio signal to reduce data volume, reduce system complexity, or adapt to specific application requirements. The downsampling factor can be a multiple of 2, such as 2, 4, or 8.

[0151] As an example, take an audio signal with a sampling rate of Fs = 32000 Hz as an input signal, sample the audio signal to obtain a sampled signal x(n) consisting of 640 sample points. Call the analysis filter (2 channels) in the QMF filter bank, perform low-pass filtering on the sampled signal to obtain a low-pass filtered signal, perform high-pass filtering on the sampled signal to obtain a high-pass filtered signal, and downsample the low-pass filtered signal to obtain a low-frequency subband signal x of the audio signal. LB (n), downsample the high-pass filtered signal to obtain the high-frequency sub-band signal x of the audio signal HB (n). Low frequency subband signal x LB (n) and high frequency sub-band signal x HB The effective bandwidth of (n) is 0-8kHz and 8-16kHz respectively, and the low-frequency sub-band signal x LB (n) and high frequency sub-band signal x HB The number of sample points of (n) is 320.

[0152] It should be noted that a QMF filter bank is a filter pair that combines analysis and synthesis. The QMF analysis filter decomposes an input signal with a sampling rate of Fs into two signals with a sampling rate of Fs / 2, representing a QMF low-pass signal and a QMF high-pass signal, respectively. The low-pass and high-pass signals recovered at the decoder are then combined through a QMF synthesis filter to recover a reconstructed signal with the sampling rate Fs corresponding to the input signal.

[0153] In the field of digital signal processing, the embodiments of the present application first filter the audio signal through a filter (such as a low-frequency filter, a high-pass filter) to remove high-frequency components and aliasing interference in the audio signal to ensure that the downsampled audio signal does not lose necessary information; then, through downsampling processing, a sampling point is retained at every certain sampling point in the filtered audio signal, thereby reducing the sampling rate of the audio signal.

[0154] In step 1022B, the low-frequency features of the low-frequency sub-band signal are extracted from the low-frequency sub-band signal through the second neural network.

[0155] It should be noted that step 1022B is similar to step 10221A, except that step 10221A processes the audio signal, while step 1022B processes the low-frequency subband signal. The first neural network and the second neural network have similar structures. For example, the second neural network, like the first neural network, includes four coding blocks, each of which includes four or five residual units.

[0156] In some embodiments, step 1022B may be implemented through steps 10221B to 10222B:

[0157] In step 10221B, feature extraction is performed on the low-frequency sub-band signal to obtain low-frequency features of the low-frequency sub-band signal.

[0158] Here, the embodiment of the present application can call a second neural network (NN) based on the low-frequency subband signal, and extract low-frequency features from the low-frequency subband signal through the second neural network, so as to continue to perform feature refinement based on important low-frequency features. It should be noted that the embodiment of the present application is not limited to the structure of the second NN, and the second NN can be a convolutional neural network, a deep neural network, etc.

[0159] In some embodiments, step 1022B can be implemented in the following manner: performing causal convolution processing on the low-frequency subband signal of the audio signal through a second neural network to obtain causal convolution features; and performing pooling processing on the causal convolution features to obtain low-frequency coding features of the low-frequency subband signal.

[0160] In the field of audio codecs, neural network (NN) operations such as causal convolution and pooling play a crucial role in processing audio signals and extracting features. In audio codecs, causal convolution can be used to extract local features from audio signals. By applying a convolution kernel (a learnable filter), convolution can be performed on the time dimension of the audio signal to capture patterns and resonances. Causal convolution can extract both time-domain and frequency-domain features from audio signals, which can be used for tasks such as noise reduction, feature extraction, and signal separation. Pooling reduces the time dimension of audio signals, thereby reducing data complexity and computational effort. Pooling samples a local region of the input signal and aggregates the information within that region, such as the maximum or average value, to produce a more compact feature representation. In audio codecs, pooling can help improve network robustness and generalization, reducing the risk of overfitting. In audio codecs, operations such as convolution and pooling can be used to implement tasks such as feature extraction, encoding, and decoding of audio signals by constructing appropriate neural network structures. These operations help improve the efficiency and quality of audio signal processing and expand the application scope of audio codec technology in fields such as audio processing, speech recognition, and music generation.

[0161] For example, referring to the network structure diagram of the second NN shown in Figure 11, the second NN includes a causal convolution layer and a preprocessing layer. First, a 16-channel causal convolution layer is called to expand the input tensor (i.e., the low-frequency subband signal) to a 16×320 causal convolution feature; then, the 16×320 causal convolution feature is preprocessed by the preprocessing layer. For example, after performing a convolution operation on the 16×320 causal convolution feature, a pooling process with a factor of 2 is performed, and the activation function can be a parametric rectified linear unit (PReLU, Parametric Rectified Linear Unit) to generate a 16×160 tensor (i.e., low-frequency coding feature).

[0162] In step 10222B, residual processing is performed on the low-frequency coding features using at least one residual unit included in the second neural network to obtain low-frequency features of the low-frequency sub-band signal.

[0163] Here, by performing residual processing on the low-frequency coding features of the low-frequency sub-band signal, based on the characteristics of residual processing, while ensuring comprehensive learning of the low-frequency coding features, it is also possible to better utilize the shallow feature information of the low-frequency coding features and avoid missing the shallow feature information of the low-frequency coding features.

[0164] In some embodiments, step 10222B can be implemented through steps 102221B-102222B.

[0165] In step 102221B, feature residual processing is performed on the low-frequency features by at least one residual unit included in the second neural network to obtain residual features of the low-frequency sub-band signal.

[0166] The feature residual processing in step 102221B is used to calculate the residual of the low-frequency feature, and determine the residual of the low-frequency feature as the residual feature of the low-frequency sub-band signal for subsequent feature encoding.

[0167] In some embodiments, when the at least one residual unit is a single residual unit, step 102221B may be implemented by performing a single residual processing on the low-frequency features using the single residual unit to obtain the residual features of the low-frequency sub-band signal. The single residual processing of the single residual unit is used to calculate a single residual corresponding to the low-frequency sub-band signal on the encoding side.

[0168] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 102221B can be implemented in the following manner: performing a residual processing on the low-frequency coding feature through the first residual unit of the plurality of cascaded residual units, wherein the residual processing of the first residual unit is used to calculate the residual of the low-frequency coding feature, and the residual of the low-frequency coding feature is determined as the residual result of the first residual unit; the residual result output by the first residual unit is output to the subsequent cascaded residual unit, and the residual processing and the output of the residual result are continued through the subsequent cascaded residual unit, wherein the residual processing of the subsequent cascaded residual unit is used to calculate the residual of the residual result input to the subsequent cascaded residual unit; and the residual result output by the last residual unit is used as the residual feature of the low-frequency subband signal.

[0169] For example, the second neural network includes 4 coding blocks, each coding block includes 4 or 5 residual units. As shown in Figure 12A, when at least one residual unit for feature residual has 5 cascaded residual units, the first residual unit performs a residual process on the low-frequency coding feature and outputs the residual result output by the first residual unit to the second residual unit; the second residual unit performs a residual process on the residual result output by the first residual unit and outputs the residual result output by the second residual unit to the third residual unit; the third residual unit performs a residual process on the residual result output by the second residual unit and outputs the residual result output by the third residual unit to the fourth residual unit; the fourth residual unit performs a residual process on the residual result output by the third residual unit and outputs the residual result output by the fourth residual unit to the fifth residual unit; the fifth residual unit performs a residual process on the residual result output by the fourth residual unit to obtain the residual feature of the low-frequency subband signal.

[0170] The embodiments of the present application are not limited to the number of coding blocks in the neural network, and can be any positive integer such as 2, 3, 4, 5, etc. The embodiments of the present application are also not limited to the number of residual units in the coding block, and can be any positive integer such as 2, 3, 4, 5, 6, etc. The number of residual units in multiple coding blocks can be the same or different, for example, one coding block contains 4 residual units and another coding block contains 5 residual units.

[0171] In some embodiments, the processing process of the residual unit is as follows: the kth residual unit of multiple cascaded residual units performs the following processing: convolution processing is performed on the input of the kth residual unit to obtain the convolution result of the kth residual unit; the convolution result of the kth residual unit is added to the input of the kth residual unit to obtain the residual result output by the kth residual unit, wherein k is a positive integer that increases in sequence, 1≤k≤J, J is the number of residual units, when k is 1, the input of the kth residual unit is the low-frequency coding feature, and when k is not 1, the input of the kth residual unit is the residual result output by the k-1th residual unit of the residual feature.

[0172] Continuing with the above embodiment, each residual unit includes a dilated convolution operator; performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit can be achieved in the following manner: performing the following processing on the kth residual unit of multiple cascaded residual units: performing dilated convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit. That is, performing dilated convolution processing on the low-frequency coding features through the dilated convolution operator included in the first residual unit to obtain the dilated convolution result of the first residual unit. Performing the following processing on the jth residual unit of multiple cascaded residual units: performing dilated convolution processing on the residual result output by the j-1th residual unit through the dilated convolution operator included in the jth residual unit to obtain the dilated convolution result of the jth residual unit, wherein j is a positive integer that increases successively, 1<j≤J, and J is the number of residual units. It should be noted that each residual unit contains a dilated convolution operator with a specified dilation rate. Using a dilated convolution operator with a progressive dilation rate is equivalent to using different receptive fields to extract input features at different resolutions, which can better perform a relatively comprehensive analysis of the data. After each residual unit is convolved with the dilated convolution operator, it is added to the shallow features from the jump connection (i.e., the input of each residual unit), thereby directly utilizing the shallow feature information, allowing the network to fully utilize the shallow feature information during the learning process.

[0173] Continuing with the above embodiment, each residual unit includes not only a dilated convolution operator but also at least one causal convolution operator. After performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit, causal convolution processing is performed on the obtained dilated convolution result using the at least one causal convolution operator included in the kth residual unit, and the obtained causal convolution result is used as the convolution result of the kth residual unit. That is, causal convolution processing is performed on the dilated convolution result of the first residual unit using the at least one causal convolution operator included in the first residual unit, and the obtained causal convolution result is used as the convolution result output by the first residual unit. The dilated convolution operator included in the j-th residual unit is used to perform dilated convolution processing on the residual result output by the j-1-th residual unit. After obtaining the dilated convolution result of the j-th residual unit, the dilated convolution result of the j-th residual unit is causally convolved by at least one causal convolution operator included in the j-th residual unit, and the causal convolution result of the j-th residual unit is used as the convolution result of the j-th residual unit. It should be noted that each residual unit also includes at least one causal convolution operator, which continues to extract local information of the features input to the causal convolution operator through the causal convolution operator.

[0174] In neural network models, causal convolution is a special type of convolution when processing time series data (audio signals are a type of time series data). It ensures that the output of the neural network depends only on the current and previous time steps, thereby maintaining temporal causality. In practical applications, causal convolution can adjust the size of the convolution kernel to ensure that the convolution kernel does not span the area before the current time step. This can effectively capture long-term dependencies in the time series while avoiding the problem of vanishing or exploding gradients caused by confusion caused by future information. Causal convolution is particularly important in fields such as natural language processing, speech recognition, and time series prediction because it follows the temporal order of the data, avoids confusion with past information, and can effectively process and predict long time series data. In tasks such as speech recognition and time series prediction, causal convolution has demonstrated superior performance due to its ability to maintain temporal order.

[0175] In some embodiments, grouped convolution can be applied to the convolution operator of the residual unit (including the dilated convolution operator and the causal convolution operator). Grouped convolution involves dividing the input channels into multiple groups for convolution operations, with only the input channels and output channels within each group being associated. It should be noted that when the input channels are divided into multiple groups, the corresponding output channels are also divided into multiple groups, i.e., the number of input channel groups is the same as the number of output channel groups. This ensures that after the convolution is performed within the group, only the input channels and output channels within each group are associated. Here, assume that the input channels of a feature input to a convolution operator are 4 and the output channels are 4. If the number of groups is 1, each input channel is associated with 4 output channels. If the number of groups is 2, the 4 input channels are first divided into two groups, 0-1 and 2-3. Within each group, the input channels are associated with the output channels within the group. For example, input channels 0-1 in the first group are associated with output channels 0-1, and input channels 2-3 in the second group are associated with output channels 2-3. As shown in Figure 6A, when the grouped convolution scheme is not used, each input channel is associated with four output channels; as shown in Figure 6B, when the grouped convolution scheme is not used, the 0th output channel is only associated with the 0th-1st input channels, not with the 2nd-3rd input channels, and the 2nd output channel is only associated with the 2nd-3rd input channels, not with the 0th-1st input channels. This comparison shows that the introduction of grouped convolution can prevent any input channel from being associated with all output channels, reduce the number of connections, and reduce complexity.

[0176] For example, when grouped convolution is applied to the dilated convolution operator included in the residual unit, dilated convolution processing is performed on the low-frequency coding features, which can be achieved by: grouping the input channels of the low-frequency coding features to obtain multiple groups, wherein each group includes the first elements corresponding to at least two channels in the low-frequency coding features; and performing dilated convolution processing on the first elements in each group. When grouped convolution is applied to the causal convolution operator included in the residual unit, causal convolution processing is performed on the obtained dilated convolution result, which can be achieved by: grouping the input channels of the dilated convolution result to obtain multiple groups, wherein each group includes the second elements corresponding to at least two channels in the dilated convolution result; and performing causal convolution processing on the second element in each group.

[0177] For example, grouped convolution can be applied to the convolution operator of the residual unit (including the dilated convolution operator and the causal convolution operator). Grouped convolution divides the input channels into multiple groups for convolution operations, and only the input channels and output channels within each group are associated. It should be noted that when the input channels are divided into multiple groups, the corresponding output channels are also divided into multiple groups. That is, the number of input channel groups is the same as the number of output channel groups. This ensures that after the convolution within a group, only the input channels and output channels within each group are associated. Here, assume that the input channel input to a feature of a convolution operator has 4 input channels and 4 output channels. If the number of groups is 1, each input channel is associated with 4 output channels. If the number of groups is 2, the 4 input channels are first divided into two groups, 0-1 and 2-3. Within each group, the input channels are associated with the output channels within the group. For example, input channels 0-1 in the first group are associated with output channels 0-1, and input channels 2-3 in the second group are associated with output channels 2-3. As shown in Figure 6A, when the grouped convolution scheme is not used, each input channel is associated with four output channels; as shown in Figure 6B, when the grouped convolution scheme is not used, the 0th output channel is only associated with the 0th-1st input channels, not with the 2nd-3rd input channels, and the 2nd output channel is only associated with the 2nd-3rd input channels, not with the 0th-1st input channels. This comparison shows that the introduction of grouped convolution can prevent any input channel from being associated with all output channels, reduce the number of connections, and reduce complexity.

[0178] Following step 102221B, in step 102222B, feature encoding processing is performed on the residual features to obtain low-frequency features of the low-frequency sub-band signal.

[0179] Here, feature coding is performed on the residual features to obtain low-frequency features of the low-frequency sub-band signal, so that signal coding can be performed based on the low-frequency features to obtain a low-frequency code stream of the audio signal.

[0180] In some embodiments, step 102222B can be implemented by: performing convolution processing on the residual features to obtain convolution features, wherein the number of channels of the convolution features is greater than the number of channels of the residual features; and performing pooling processing on the convolution features to obtain low-frequency features of the low-frequency sub-band signal.

[0181] For example, the second NN is called based on the low-frequency subband signal. After processing by the residual unit in the second NN, the residual features are obtained. The residual features are then convolved through the convolution layer in the second NN to increase the number of channels of the residual features. Finally, the convolution features are pooled through the pooling layer in the second NN to obtain the low-frequency features of the low-frequency subband signal. Of course, the second NN may also include a causal convolution layer, which performs causal convolution on the low-frequency features to obtain the low-frequency features after causal convolution. The low-frequency features of the low-frequency subband signal after causal convolution are then signal encoded to obtain the low-frequency code stream of the audio signal.

[0182] For example, as shown in Figure 11, after the second NN is called based on the low-frequency subband signal, the low-frequency coding features (the 16×160 tensor obtained after preprocessing in Figure 11) are obtained through the second NN. The second NN includes four cascaded coding blocks with different downsampling factors (Down_factor). Each coding block contains a residual block (including at least one residual unit), a convolutional layer, and a pooling layer. Each residual block includes five residual units (Residual Units) based on dilated convolution (the input and output feature dimensions of the residual unit do not change); a convolutional layer is used to double the number of input channels, and the activation function can be PReLU to ensure data volume and avoid data loss; the pooling layer is a pooling operation including a Down_factor to complete downsampling and achieve data compression. Here, the Down_factor of the four coding blocks is set to 2, 4, 4, and 5, respectively. Therefore, the number of output channels of the four coding blocks is set to 32, 64, 128, and 256, respectively. After being processed by the four encoding blocks, the input 16×160 tensor is converted into 32×80, 64×20, 128×5, and 256×1 tensors respectively. For example, the residual block in the first encoding block performs residual processing on the low-frequency encoding feature (i.e., the 16×160 tensor), and the residual result output by the residual block in the first encoding block is output to the feature encoding block in the first encoding block. After being processed by the feature encoding block in the first encoding block (including a convolution layer and a pooling layer), the encoding result of the feature encoding block in the first encoding block (i.e., a 32×80 tensor) is obtained, and the encoding result of the feature encoding block in the first encoding block (i.e., a 32×80 tensor) is output to the second encoding block; the residual block in the second encoding block is used to process the first The encoding result of the feature coding block in the coding block (i.e., a 32×80 tensor) is subjected to residual processing, and the residual result output by the residual block in the second coding block is output to the feature coding block in the second coding block. After processing by the feature coding block in the second coding block (including a convolution layer and a pooling layer), the encoding result of the feature coding block in the second coding block (i.e., a 64×20 tensor) is obtained, and the encoding result of the feature coding block in the first coding block (i.e., a 64×20 tensor) is output to the third coding block; the above processing is performed in sequence, and the output of the last coding block is used as a low-frequency feature. Among them, the embodiment of the present application is not limited to the number of coding blocks, and can be any positive integer such as 2, 3, 4, 5, etc.

[0183] In step 1023B, high frequency analysis is performed on the high frequency sub-band signal to obtain high frequency features of the high frequency sub-band signal.

[0184] Because low-frequency subband signals have a greater impact on audio coding than high-frequency subband signals, differential signal processing is performed on low-frequency and high-frequency subband signals, resulting in a lower dimension of the high-frequency features than the low-frequency features. For example, the dimension of the low-frequency features is 56, while the dimension of the high-frequency features is 8. High-frequency analysis is used to reduce the dimensionality of the high-frequency subband signals, thereby achieving data compression. High-frequency coding features are features that characterize high-frequency subband signals, and their dimension is smaller than that of the high-frequency subband signals.

[0185] In some embodiments, a fifth neural network may be called to extract high-frequency features of the high-frequency sub-band signal from the high-frequency sub-band signal. The fifth neural network has a structure similar to that of the second neural network. The dimension of the high-frequency features is smaller than the dimension of the low-frequency features.

[0186] In some embodiments, high-frequency analysis of a high-frequency subband signal to obtain high-frequency features of the high-frequency subband signal can be performed by: framing the high-frequency subband signal to obtain multiple subframes of the high-frequency subband signal; performing frequency band expansion processing on each subframe to obtain a subband spectrum envelope for each subframe; and using the subband spectrum envelopes corresponding to the multiple subframes as the high-frequency features of the high-frequency subband signal. The number of subframes can be 2, 4, 6, etc., and the embodiments of the present application are not limited to the number of subframes.

[0187] Here, since the high-frequency sub-band signal is relatively less important to quality than the low-frequency sub-band signal, another method can be used to compress the high-frequency sub-band signal, namely, band expansion (recovering the broadband speech signal from the band-limited narrowband speech signal) to quickly compress the high-frequency sub-band signal and extract the high-frequency features of the high-frequency sub-band signal.

[0188] In some embodiments, frequency band expansion processing is performed on each subframe to obtain the subband spectrum envelope of each subframe, which can be achieved in the following way: frequency domain transformation is performed based on multiple sample points included in the subframe to obtain transformation coefficients corresponding to the multiple sample points; the transformation coefficients corresponding to the multiple sample points are divided into multiple subbands; the transformation coefficients included in each subband are averaged to obtain the average energy corresponding to each subband, and the average energy is used as the subband spectrum envelope corresponding to each subband.

[0189] It should be noted that the frequency domain transformation methods in the embodiments of the present application include modified discrete cosine transform (MDCT), discrete cosine transform (DCT), fast Fourier transform (FFT), etc., and the embodiments of the present application are not limited to the frequency domain transformation method. The averaging processing in the embodiments of the present application includes arithmetic mean and geometric mean, and the embodiments of the present application are not limited to the averaging processing method.

[0190] In some embodiments, frequency domain transform is performed based on multiple sample points included in a subframe to obtain transform coefficients corresponding to the multiple sample points respectively. This can be achieved in the following manner: performing the following processing on each subframe: obtaining an adjacent subframe adjacent to the subframe; based on the multiple sample points included in the adjacent subframe and the multiple sample points included in the subframe, performing discrete cosine transform processing on the multiple sample points included in the subframe to obtain transform coefficients corresponding to the multiple sample points included in the subframe respectively.

[0191] As an example, for a high frequency subband signal x comprising 320 points HB (n), the 320-point high-frequency sub-band signal x HB (n), divided into two subframes (the first subframe and the second subframe), each subframe includes 160 points. For any subframe, that is, the high-frequency subband signal including 160 points, the modified discrete cosine transform (MDCT) is called to generate 160-point MDCT coefficients (that is, the transform coefficients corresponding to the multiple sample points included in the subframe). Specifically, if there is a 50% overlap, the first subframe of the n+1th frame can be merged (spliced) with the second subframe of the nth frame, and the 320-point MDCT is calculated to obtain the 160-point MDCT coefficients of the first subframe of the n+1th frame; the second subframe of the n+1th frame can be merged with the first subframe of the n+1th frame, and the 320-point MDCT is calculated to obtain the 160-point MDCT coefficients of the second subframe of the n+1th frame.

[0192] In some embodiments, the process of performing geometric averaging on the transform coefficients included in each sub-band is as follows: determining the sum of the squares of the transform coefficients corresponding to the sample points included in each sub-band; and determining the ratio of the sum of the squares to the number of sample points included in the sub-band to obtain the average energy corresponding to each sub-band.

[0193] Following step 1023B, in step 1024B, the low-frequency features and the high-frequency features are determined as coding features of the audio signal.

[0194] Following step 102 , in step 103 , the coding feature is signal-coded using a target bit rate mode to obtain an audio code stream of the audio signal.

[0195] Here, the rate mode is used to encode the signal using a specified code table and a corresponding code rate. Therefore, the coding characteristics of the audio signal are encoded using the code table and code rate corresponding to the target rate mode to obtain an audio code stream of the audio signal. In the field of digital signal processing, step 103 can be implemented as follows: the coding characteristics are encoded based on the digital signal using the code table and code rate corresponding to the target rate mode to obtain an audio code stream of the audio signal.

[0196] In some embodiments, when the target coding mode is a broadband coding mode, the target bit rate mode is used to indicate the use of a target code table to perform signal encoding on the coding features of the audio signal; step 103 can be implemented in the following manner: using the target code table, quantizing the coding features of the audio signal to obtain a quantized value of the coding feature; and entropy encoding the quantized value using the target bit rate corresponding to the target code table to obtain an audio code stream of the audio signal.

[0197] As shown in Figure 4E, Figure 4E is a flow chart of the audio coding method provided in an embodiment of the present application. When the target coding mode is the ultra-wideband coding mode, the target bit rate mode includes an indication of using a first code table to perform signal encoding on the low-frequency features and using a second code table to perform signal encoding on the high-frequency features; Figure 4E shows that step 103 in Figure 4A can be implemented through steps 1031-1033: in step 1031, the low-frequency features are quantized using the first code table to obtain quantized values ​​of the low-frequency features, and the quantized values ​​of the low-frequency features are entropy encoded using the bit rate corresponding to the first code table to obtain a low-frequency code stream of the low-frequency sub-band signal; in step 1032, the high-frequency features are quantized using the second code table to obtain quantized values ​​of the high-frequency features, and the quantized values ​​of the high-frequency features are entropy encoded using the bit rate corresponding to the second code table to obtain a high-frequency code stream of the high-frequency sub-band signal; in step 1033, an audio code stream of the audio signal is constructed based on the low-frequency code stream and the high-frequency code stream.

[0198] The degree of quantization (i.e., quantization accuracy) corresponding to different code tables may be different. For example, the quantization accuracy of the first code table is greater than the quantization accuracy of the second code table. Since the value range of each dimension in the low-frequency features is [-1, 1], the interval [-1, 1] can be divided into 11 equal parts to form a first code table containing 11 elements, and the low-frequency features are quantized using the first code table. Since the value range of each dimension in the high-frequency features is [-1, 1], the interval [-1, 1] can be divided into 5 equal parts to form a second code table containing 5 elements, and the low- and high-frequency features are quantized using the second code table.

[0199] Here, for the low-frequency feature F of the low-frequency subband signal LB (n) and the high-frequency coding feature F of the high-frequency sub-band signal HB (n) Scalar quantization (each component is quantized separately) and entropy coding methods can be performed. In addition, the embodiments of the present application do not limit the technical combination of vector quantization (combining multiple adjacent components into a vector for joint quantization) and entropy coding.

[0200] It should be noted that when the target bit rate mode only includes an instruction to use the first code table to encode the low-frequency features and the second code table to encode the high-frequency features, the low-frequency code stream and the high-frequency code stream are combined to obtain the audio code stream of the audio signal.

[0201] In some embodiments, the target bit rate mode also includes an instruction to use a third code table to perform signal encoding on the residual features of the high-frequency features; before step 1033, the subband corresponding to the subframe of the high-frequency subband signal is determined, and the subband is divided into multiple subsets; the transform coefficients included in each subset are averaged to obtain the average energy corresponding to each subset, and the average energy is used as the envelope value corresponding to each subset; in the quantized value of the high-frequency feature, the quantized value of the subband spectrum envelope of the subband corresponding to the subset is determined, and the difference between the envelope value corresponding to the subset and the quantized value of the subband spectrum envelope is used as the first residual value; based on the first residual value, the residual features of the high-frequency feature are determined; the residual features of the high-frequency feature are quantized using the third code table to obtain the quantized value of the residual features of the high-frequency feature, and the quantized value of the residual features of the high-frequency feature is entropy encoded using the code rate corresponding to the third code table to obtain a residual code stream of the residual feature; correspondingly, step 1033 can be implemented by using the low-frequency code stream, the high-frequency code stream and the residual code stream as the audio code stream of the audio signal. The quantization degree of the third code table may be different from the quantization degrees of the first code table and the second code table.

[0202] In some embodiments, when the number of residual values ​​is N, N is a positive integer greater than 1, and the residual features of the high-frequency features are determined based on the first residual value, which can be achieved in the following manner: determining the quantization value of the nth residual value, and determining the sum of the quantization value of the nth residual value and the quantization value of the subband spectrum envelope of the corresponding subband; taking the difference between the envelope value corresponding to the subset and the sum as the n+1th residual value; determining the N residual values ​​as the residual features of the high-frequency features; wherein n is a positive integer that increases successively, 1≤n≤N.

[0203] Assume that the number of residual values ​​is 2 for example, and the residual value is determined in the following way:

[0204] 1) For each 10ms subframe, the subframe is divided into 4 subbands, and each subband is further evenly divided into 2 subsets.

[0205] For example, the set of envelope values ​​of four sub-bands is {e1, e2, e3, e4}.

[0206] 2) For each subset of each subband, calculate the envelope value of the subset.

[0207] For example, the set of envelope values ​​of all subsets is {e 1a ,e 1b ,e 2a ,e 2b ,e 3a ,e 3b ,e 4a ,e 4b}.

[0208] 3) Subtract the corresponding high-frequency feature vector quantization value (ie, the quantization value of the high-frequency feature) from the envelope value of each subset to obtain the first high-frequency feature vector residual (referred to as the first residual value).

[0209] For example, the high-frequency feature vector quantization value is Then the first high-frequency eigenvector residual is

[0210] 4) The first high-frequency feature residual is subjected to similar quantization and entropy coding to obtain a first high-frequency feature vector residual quantized value (hereinafter referred to as the quantized value of the first residual value).

[0211] Here, considering that the dynamic range of the residual value is much lower than the original value, only 2 bits need to be allocated when scalar quantizing each residual. In this way, a reference bit rate of 2*16*50=1.6kbps is required (similarly, because it is non-uniformly distributed entropy coding, the actual bit rate is generally less than 1.6kbps) to obtain the residual quantization value of the first high-frequency eigenvector.

[0212] 5) Subtract the sum of the corresponding high-frequency feature vector quantization value and the first high-frequency feature vector residual quantization value from the envelope value of each subset to obtain the second high-frequency feature vector residual (referred to as the second residual value).

[0213] In step 104, a frame header is determined based on the target coding mode and the target bit rate mode.

[0214] Here, an empty frame header can be constructed first, and the empty frame header includes a coding mode bit and a rate mode bit, and the target coding mode is written into the coding mode bit included in the empty frame header, and the target rate mode is written into the rate mode bit included in the empty frame header to determine the frame header. The embodiment of the present application is not limited to the number of bits in the frame header, nor is it limited to the position of the target coding mode and the target rate mode in the frame header. For example, the target coding mode can be written into the first bit of the frame header or the last bit of the frame header, and is not limited to the number of bits occupied by the coding mode bit and the rate mode bit in the frame header.

[0215] In step 105, an audio code stream package of the audio signal is generated based on the audio code stream and the frame header.

[0216] For example, after obtaining the audio code stream and the frame header, the audio code stream and the frame header are packaged into a file according to a specific format to obtain an audio code stream package, which includes the audio code stream and the frame header.

[0217] In some embodiments, before step 105 , the flatness side information of the high frequency sub-band signal is determined; therefore, step 105 can be implemented by combining the audio code stream, the frame header, and the flatness side information to obtain an audio code stream encapsulation of the audio signal.

[0218] For example, the frame header is placed at the beginning of the audio stream encapsulation, and the audio stream and flatness side information are placed after the frame header. The embodiments of the present application are not limited to the order of the audio stream and flatness side information. In some embodiments, the target coding mode and target bit rate mode can also be placed at other locations in the audio stream encapsulation.

[0219] The flatness of a high-frequency subband signal refers to the uniform distribution of the amplitudes of each frequency component within the frequency domain. Flatness is an important metric for measuring high-frequency subband signal quality. Flatness side information provides additional information about the flatness of a high-frequency subband signal when processing it. During frequency-domain analysis of a high-frequency subband signal, this flatness side information includes the amplitude distribution of the high-frequency subband signal at each frequency point. This information is used to determine whether the high-frequency subband signal remains consistent within a specific frequency range or whether it exhibits abrupt peaks. During signal equalization, this flatness side information helps identify dips and peaks within the high-frequency subband signal, enabling appropriate gain adjustments to be made to flatten the overall frequency response and improve sound quality. Flatness side information can also be used to analyze signal distortion. If the amplitude of a high-frequency subband signal at certain frequencies differs significantly from that at other frequencies, this can cause distortion, impacting sound quality. In short, the flatness edge information of high-frequency sub-band signals is of great significance for signal preprocessing, analysis, equalization and post-processing, and helps to improve the accuracy and effect of audio signal processing.

[0220] In some embodiments, determining the flatness side information of the high-frequency subband signal can be achieved by: dividing the transform coefficients included in the subframe of the high-frequency subband signal into multiple blocks; determining a first flatness of each block, and determining a second flatness of the low-frequency specified frequency band of the audio signal; when the first flatness is less than the second flatness, or the first flatness is less than a flatness threshold, setting the flatness side information to a first value, wherein the first value indicates that flattening processing is required during audio decoding; when the first flatness is greater than or equal to the second flatness, and the first flatness is greater than or equal to the flatness threshold, setting the flatness side information to a second value, wherein the second value indicates that flattening processing is not required during audio decoding.

[0221] The low-frequency designated band is a low-frequency band set according to actual application requirements, and the flatness threshold is also a threshold set according to actual application requirements.

[0222] For example, here the first value may be 1 and the second value may be 0; the first value may also be 0 and the second value may be 1.

[0223] In some embodiments, the frame header further includes at least one channel bit, and the channel bit is used to indicate whether a mono encoding method or a stereo encoding method is used to encode the audio signal. When the audio signal is a stereo signal obtained by downmixing the input signal of the left channel and the input signal of the right channel, the channel bit in the frame header is set to a first value (e.g., 1), and the first value indicates that the audio signal is encoded using the stereo encoding method. The input signal of the left channel and the input signal of the right channel are subjected to parametric stereo encoding to obtain a stereo feature vector, and then an audio code stream encapsulation of the audio signal is generated based on the stereo feature vector, the audio code stream, and the frame header; when the audio signal is a mono input signal, the channel bit in the frame header is set to a second value (e.g., 0), and the second value indicates that the audio signal is encoded using the mono encoding method, and an audio code stream encapsulation of the audio signal is generated based on the audio code stream and the frame header.

[0224] As previously mentioned, the audio decoding method provided in the embodiments of the present application can be implemented by various types of electronic devices. Referring to Figure 5A , Figure 5A is a schematic flow chart of the audio decoding method provided in the embodiments of the present application. The audio decoding function is implemented by the audio decoding method. The audio decoding method is the inverse process of the audio encoding method described above, and will be described below in conjunction with the steps shown in Figure 5A .

[0225] Before describing the following steps, the audio decoding method provided in the embodiment of the present application is explained. The audio decoding method provided in the embodiment of the present application includes two decoding methods (i.e., a signal decoding method in signal processing technology and a feature decoding (i.e., a reconstruction method) in artificial intelligence technology. The signal decoding method in signal processing technology is detailed in step 202 below, and the feature decoding in artificial intelligence technology is detailed in step 203 below.

[0226] In step 200, an audio stream package is obtained.

[0227] The audio code stream encapsulation includes an audio code stream, which is obtained by audio encoding an audio signal through a target coding mode and a target bit rate mode. The target coding mode is obtained from multiple candidate coding modes, and the target bit rate mode is obtained from multiple candidate bit rate modes.

[0228] As an example, after the audio code stream package is obtained by encoding using the audio encoding method shown in FIG4A , the audio code stream package is transmitted to a decoding end, and the decoding end receives the audio code stream package.

[0229] In step 201, in response to a decoding request for an audio stream encapsulation, a target coding mode and a target bit rate mode are obtained from a frame header included in the audio stream encapsulation.

[0230] Here, the decoding request is used to instruct to perform audio decoding processing on the audio code stream.

[0231] As an example, after the audio code stream package is encoded by the audio encoding method shown in Figure 4A, the audio code stream package is transmitted to the decoding end. After the decoding end receives the audio code stream package, it performs audio decoding processing on the audio code stream package. First, the target encoding mode and target bit rate mode are obtained from the frame header included in the audio code stream package, for example, the target encoding mode is obtained from the encoding mode bit included in the frame header, and the target bit rate mode is obtained from the bit rate mode bit included in the frame header.

[0232] In step 202, the audio code stream is decoded using the target coding mode and the target bit rate mode to obtain a coding feature estimation value corresponding to the audio code stream.

[0233] Since signal decoding is the inverse of signal encoding, the audio stream is decoded using the decoding mode and target bitrate mode corresponding to the target coding mode to obtain the corresponding coding feature estimate. For example, if the target coding mode is wideband, the audio stream is decoded using the wideband decoding mode to obtain the corresponding coding feature estimate. If the target coding mode is ultra-wideband, the audio stream is decoded using the ultra-wideband decoding mode to obtain the corresponding coding feature estimate.

[0234] It should be noted that signal decoding is the inverse process of signal encoding. Since the process of decoding the received code stream at the decoding end is the inverse process of the encoding process at the encoding end, the value generated in the decoding process is an estimate relative to the value in the encoding process. For example, the data value representing the coding feature generated in the decoding process is an estimate relative to the coding feature in the encoding process. In other words, the former is not necessarily equivalent to the data value of the original coding feature. For example, there may be subtle differences between the two due to the encoding and decoding operation process, and there are situations where decoding cannot completely restore the data value of the original coding feature. Therefore, the former is also called "coding feature estimate value".

[0235] In some embodiments, the target coding mode is a broadband coding mode, and the target bit rate mode is used to indicate the use of a target code table to perform signal encoding on the coding features of the audio signal; therefore, step 202 can be implemented in the following manner: using the target bit rate corresponding to the target code table, entropy decoding is performed on the audio bit stream to obtain a quantization value corresponding to the audio bit stream; using the target code table, inverse quantization processing is performed on the quantization value corresponding to the audio bit stream to obtain an estimated value of the coding features corresponding to the audio bit stream.

[0236] Inverse quantization is achieved by querying a quantization table, which is a mapping table generated by quantization during the encoding process. For example, entropy decoding is first performed on the received bitstream, and then an estimate of the feature vector is obtained by querying a quantization table (i.e., inverse quantization, which is a mapping table generated by quantization during the encoding process). This is the estimated encoding feature value corresponding to the audio bitstream.

[0237] Referring to FIG5B , FIG5B is a flow chart of the audio decoding method provided in an embodiment of the present application, wherein the target coding mode is an ultra-wideband coding mode, and the target bit rate mode includes an instruction to use a first code table to perform signal encoding on the low-frequency characteristics of the audio signal, and to use a second code table to perform signal encoding on the high-frequency characteristics of the audio signal, and the audio code stream includes a low-frequency code stream and a high-frequency code stream; therefore, FIG5B shows that step 202 can be implemented through steps 2021-2024: in step 2021, a low-frequency code stream and a high-frequency code stream are obtained from the audio code stream through the target coding mode; in step 2022, a low-frequency code stream and a high-frequency code stream are obtained from the audio code stream through the first code table The low-frequency code stream is entropy decoded according to the bit rate corresponding to the second code table to obtain the quantization value corresponding to the low-frequency code stream, and the quantization value corresponding to the low-frequency code stream is inversely quantized using the first code table to obtain the low-frequency feature estimation value corresponding to the low-frequency code stream; in step 2023, the high-frequency code stream is entropy decoded according to the bit rate corresponding to the second code table to obtain the quantization value corresponding to the high-frequency code stream, and the quantization value corresponding to the high-frequency code stream is inversely quantized using the second code table to obtain the high-frequency feature estimation value corresponding to the high-frequency code stream; in step 2024, the coding feature estimation value corresponding to the audio code stream is determined based on the low-frequency feature estimation value and the high-frequency feature estimation value.

[0238] The data values ​​representing low-frequency features generated during the decoding process are estimates relative to the low-frequency features generated during the encoding process. That is, the former are not necessarily equivalent to the data values ​​of the original low-frequency features. For example, there may be subtle differences between the two due to the encoding and decoding processes, and there may be situations where decoding cannot completely restore the data values ​​of the original low-frequency features. Therefore, the former are also called "low-frequency feature estimates." Similarly, the high-frequency feature estimates generated during the decoding process are also estimates relative to the high-frequency features generated during the encoding process.

[0239] For example, for a low-frequency code stream, first perform entropy decoding on the low-frequency code stream using the code rate corresponding to the first code table to obtain the quantization value corresponding to the low-frequency code stream, and then obtain the low-frequency feature estimation value F′ corresponding to the low-frequency code stream by looking up the first code table. LB (n); For the high-frequency code stream, first perform entropy decoding on the high-frequency code stream using the code rate corresponding to the second code table to obtain the quantization value corresponding to the high-frequency code stream, and then obtain the high-frequency feature estimation value F′ corresponding to the high-frequency code stream by looking up the second code table. HB(n). It should be noted that when the target bit rate mode only includes instructions for using the first code table to encode the low-frequency features of the audio signal and using the second code table to encode the high-frequency features of the audio signal, the low-frequency feature estimation value and the high-frequency feature estimation value are determined as the coding feature estimation value corresponding to the audio bit stream, wherein the high-frequency feature estimation value is used to participate in high-frequency reconstruction to generate a high-frequency subband signal estimation value.

[0240] In some embodiments, the audio bitstream also includes a residual bitstream, and the target bitrate mode further includes an instruction to use a third code table to perform signal encoding on the residual features of the high-frequency features corresponding to the audio bitstream; before step 2024, entropy decoding is performed on the residual bitstream using the bitrate corresponding to the third code table to obtain a quantization value corresponding to the residual bitstream, and the quantization value corresponding to the residual bitstream is inversely quantized using the third code table to obtain a residual feature estimate corresponding to the residual bitstream. Accordingly, step 2024 can be implemented by summing the high-frequency feature estimate and the residual feature estimate to determine the final high-frequency feature estimate; and determining the low-frequency feature estimate and the final high-frequency feature estimate as the coding feature estimate corresponding to the audio bitstream.

[0241] The ith value in the high-frequency feature estimate is added to the 2i-1th value and the 2ith value in the residual feature estimate to obtain the 2i-1th value and the 2ith value in the final high-frequency feature estimate. For example, the high-frequency feature estimate is F′ HB (n)=[e′1,e′2,…,e′8], the residual feature estimate is F′ C (n) = [e′ 1C ,e′ 2C ,…,e′ 16C ], then the final estimated value of the high-frequency feature is F′ FHB (n) = [e′1 + e′ 1C ,e′1+e′ 2C ,…,e′8+e′ 16C ].

[0242] In step 203, the estimated coding feature value is reconstructed using the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

[0243] It should be noted that the reconstruction process is the inverse of the extraction process. Because decoding the received bitstream at the decoder is the inverse of encoding at the encoder, the values ​​generated during the decoding process are estimates relative to the values ​​during the encoding process. For example, the reconstructed audio signal generated during the decoding process is an estimate relative to the audio signal during the encoding process. This means that the reconstructed audio signal is not necessarily identical to the original audio signal. Subtle differences may exist between the two due to the encoding and decoding operations, and decoding may not be able to fully restore the original audio signal.

[0244] In some embodiments, the target coding mode is a broadband coding mode, and step 203 can be implemented by calling a third neural network based on the broadband coding mode; reconstructing the coding feature estimation value corresponding to the audio code stream through the third neural network to obtain a reconstructed audio signal corresponding to the audio code stream.

[0245] It should be noted that the reconstruction process through the third neural network is the inverse process of the extraction process through the first neural network.

[0246] As an example, reconstructing the coding feature estimation value corresponding to the audio code stream to obtain the reconstructed audio signal corresponding to the audio code stream can be achieved in the following manner: utilizing at least one residual unit included in the third neural network to perform residual processing on the coding feature estimation value corresponding to the audio code stream to obtain the audio feature estimation value corresponding to the audio code stream; performing feature reconstruction on the audio feature estimation value corresponding to the audio code stream to obtain the reconstructed audio signal corresponding to the audio code stream.

[0247] The third neural network includes four decoding blocks, each of which includes four or five residual units. The embodiments of the present application are not limited to the number of decoding blocks, and can be any positive integer such as 2, 3, 4, or 5. The embodiments of the present application are also not limited to the number of residual units in a decoding block, and can be any positive integer such as 2, 3, 4, 5, or 6.

[0248] Based on the characteristics of the residual unit, at least one residual unit included in the third neural network is used to perform residual processing on the estimated value of the coding feature corresponding to the audio code stream for calculating the residual of the coding feature on the decoding side. For example, the residual of the coding feature is obtained by adding the coding feature to the output of the residual unit on the decoding side, that is, the coding feature is used as the input of the residual unit, and after the coding feature is processed by the residual unit on the decoding side, the output of the residual unit is obtained, and through the jump connection characteristics of the residual unit, the input of the residual unit on the decoding side and the output of the residual unit are added to obtain the residual of the coding feature.

[0249] 5C , which is a flow chart of an audio decoding method provided in an embodiment of the present application, wherein the target coding mode is an ultra-wideband coding mode, and FIG5C shows that step 203 can be implemented through steps 2031 to 2033:

[0250] In step 2031, a fourth neural network is used to perform feature reconstruction on the low-frequency feature estimate included in the coding feature estimate to obtain a low-frequency subband signal estimate corresponding to the audio stream. The structure of the fourth neural network is similar to that of the third neural network.

[0251] In some embodiments, step 2031 can be implemented through steps 20311-20312.

[0252] In step 20311, residual processing is performed on the low-frequency feature estimate using at least one residual unit included in the fourth neural network to obtain a low-frequency coding feature estimate. The fourth neural network includes four decoding blocks, each of which includes four or five residual units.

[0253] In some embodiments, step 20311 can be implemented through steps 203111-203112:

[0254] In step 203111, feature decoding processing is performed on the low-frequency feature estimation value to obtain residual features.

[0255] For example, feature decoding is the inverse process of feature encoding, and feature decoding processing is performed on the low-frequency feature estimation value to obtain the residual feature (an estimation value).

[0256] In some embodiments, step 203111 can be implemented by: performing convolution processing on the low-frequency features to obtain convolution features, wherein the number of channels of the convolution features is smaller than the number of channels of the low-frequency features; and performing upsampling processing on the convolution features to obtain residual features.

[0257] In the field of audio codecs, upsampling is used to increase the resolution of feature maps (i.e., convolutional features) to more accurately reconstruct audio signals. Upsampling involves interpolation or other forms of upsampling techniques to generate higher-precision feature maps, which helps to better restore the original details and characteristics of the audio signal during the decoding process. By using neural network techniques such as convolution, pooling, and upsampling in audio decoding, useful features can be effectively extracted, computational complexity can be reduced, and the original content of the audio signal can be more accurately reconstructed. These technologies are of great significance for improving the performance and efficiency of audio decoding and help promote the development and application of audio codec technology.

[0258] Of course, before step 203111, causal convolution can be performed on the low-frequency features to obtain the low-frequency features after causal convolution, and step 203111 can be executed based on the low-frequency features after causal convolution, that is, feature decoding processing is performed on the low-frequency features after causal convolution to obtain residual features.

[0259] In step 203112, feature residual processing is performed on the residual features through at least one residual unit included in the fourth neural network to obtain a low-frequency coding feature estimation value.

[0260] Here, residual processing is performed on the residual features to ensure that the residual features are fully learned while making better use of the shallow feature information of the residual features to avoid missing the shallow feature information. The residual features obtained by performing feature decoding processing on the low-frequency feature estimation value are estimated values ​​relative to the residual features on the encoding side. Since the encoding process and the decoding process are inverse processes to each other, the residual features obtained by performing feature decoding processing on the low-frequency feature estimation value are not the residual results obtained by residual calculation on the decoding side. After obtaining the residual features obtained by performing feature decoding processing on the low-frequency feature estimation value, residual calculation can be performed on the residual features obtained by performing feature decoding processing on the low-frequency feature estimation value.

[0261] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 203112 may be implemented in the following manner: performing a residual process on the residual feature through one residual unit to obtain a low-frequency coding feature estimation value.

[0262] In some embodiments, when at least one residual unit is a plurality of cascaded residual units, step 203112 can be implemented in the following manner: performing a residual processing on the residual feature through the first residual unit of the plurality of cascaded residual units; outputting the residual result output by the first residual unit to the subsequent cascaded residual unit, and continuing to perform a residual processing and output the residual result through the subsequent cascaded residual unit; and using the residual result output by the last residual unit as a low-frequency coding feature estimation value.

[0263] In some embodiments, the processing process of the residual unit is as follows: the kth residual unit of multiple cascaded residual units performs the following processing: convolution processing is performed on the input of the kth residual unit to obtain the convolution result of the kth residual unit; the convolution result of the kth residual unit is added to the input of the kth residual unit to obtain the residual result output by the kth residual unit, wherein k is a positive integer that increases in sequence, 1≤k≤J, J is the number of residual units, when k is 1, the input of the kth residual unit is the residual feature, when k is not 1, the input of the kth residual unit is the residual result output by the k-1th residual unit.

[0264] Continuing with the above embodiment, each residual unit includes a dilated convolution operator; performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit can be achieved in the following manner: performing the following processing on the kth residual unit of multiple cascaded residual units: performing dilated convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit. That is, performing dilated convolution processing on the residual features through the dilated convolution operator included in the first residual unit to obtain the dilated convolution result of the first residual unit. Performing the following processing on the jth residual unit of multiple cascaded residual units: performing dilated convolution processing on the residual result output by the j-1th residual unit through the dilated convolution operator included in the jth residual unit to obtain the dilated convolution result of the jth residual unit, wherein j is a positive integer that increases successively, 1<j≤J, and J is the number of residual units.

[0265] Continuing with the above embodiment, each residual unit includes not only a dilated convolution operator but also at least one causal convolution operator. After performing convolution processing on the input of the kth residual unit to obtain the convolution result of the kth residual unit, causal convolution processing is performed on the obtained dilated convolution result using the at least one causal convolution operator included in the kth residual unit, and the obtained causal convolution result is used as the convolution result of the kth residual unit. That is, causal convolution processing is performed on the dilated convolution result of the first residual unit using the at least one causal convolution operator included in the first residual unit, and the obtained causal convolution result is used as the convolution result output by the first residual unit. The residual result output by the j-1th residual unit is subjected to dilated convolution processing through the dilated convolution operator included in the jth residual unit. After obtaining the dilated convolution result of the jth residual unit, the dilated convolution result of the jth residual unit is subjected to causal convolution processing through at least one causal convolution operator included in the jth residual unit, and the causal convolution result of the jth residual unit is used as the convolution result of the jth residual unit.

[0266] In some embodiments, when grouped convolution is applied to the dilated convolution operator included in the residual unit, dilated convolution processing is performed on the low-frequency features, which can be achieved by: grouping the input channels of the residual features to obtain multiple groups, wherein each group includes the first elements corresponding to at least two channels in the residual features; and performing dilated convolution processing on the first elements in each group. When grouped convolution is applied to the causal convolution operator included in the residual unit, causal convolution processing is performed on the obtained dilated convolution result, which can be achieved by: grouping the input channels of the dilated convolution result to obtain multiple groups, wherein each group includes the second elements corresponding to at least two channels in the dilated convolution result; and performing causal convolution processing on the second element in each group.

[0267] In some embodiments, the fourth neural network for audio decoding includes multiple cascaded decoding blocks, each decoding block includes a feature decoding block and at least one residual unit; step 203111 can be implemented by step 203111A, and step 203112 can be implemented by step 203112A: in step 203111A, feature decoding processing is performed on the low-frequency features through the feature decoding blocks in the multiple cascaded decoding blocks to obtain residual features; correspondingly, in step 203112A, residual processing is performed on the residual features once through at least one residual unit in the multiple cascaded decoding blocks to obtain a low-frequency coding feature estimate.

[0268] In some embodiments, step 203111A can be implemented in the following manner: performing feature decoding processing on the low-frequency features through the feature decoding block in the first decoding block of multiple cascaded decoding blocks, and outputting the decoding result output by the feature decoding block in the first decoding block to at least one residual unit in the first decoding block; performing feature decoding processing on the residual result output by at least one residual unit in the i-1th decoding block through the feature decoding block in the i-th decoding block of multiple cascaded decoding blocks, and outputting the decoding result output by the feature decoding block in the i-th decoding block to at least one residual unit in the i-th decoding block; using the decoding result output by the feature decoding block in the last decoding block as the residual feature; wherein i is a positive integer that increases successively, 1<i≤I, and I is the number of decoding blocks. The decoding result output by the feature decoding block in the first decoding block is obtained by performing the following processing on the feature decoding block in the first decoding block of the multiple cascaded decoding blocks: performing convolution processing on the low-frequency features to obtain convolution features, wherein the number of channels of the convolution features is less than the number of channels of the low-frequency features; and performing upsampling processing on the convolution features to obtain the decoding result output by the feature decoding block in the first decoding block. The decoding result output by the feature decoding block in the i-th decoding block is obtained by performing the following processing on the feature decoding block in the i-th decoding block: performing convolution processing on the residual result output by at least one residual unit in the i-1-th decoding block to obtain convolution features, wherein the number of channels of the convolution features is less than the number of channels of the residual result output by the at least one residual unit; and performing upsampling processing on the convolution features to obtain the decoding result output by the feature decoding block in the i-th decoding block.

[0269] As an example, the fourth neural network includes 4 decoding blocks, each of which includes 4 or 5 residual units.

[0270] In some embodiments, step 203112A may be implemented by performing a residual process on the residual feature through at least one residual unit in the last decoding block of a plurality of cascaded decoding blocks to obtain a low-frequency coding feature estimation value.

[0271] In step 2032, high frequency reconstruction is performed on the high frequency feature estimation value included in the coding feature estimation value to obtain a high frequency subband signal estimation value corresponding to the audio code stream.

[0272] It should be noted that the feature reconstruction at the decoding end is the inverse process of the extraction process at the encoding end, and the high-frequency reconstruction at the decoding end is the inverse process of the high-frequency analysis at the encoding end.

[0273] In some embodiments, high-frequency reconstruction is performed on the high-frequency feature estimation value included in the coding feature to obtain a high-frequency sub-band signal estimation value corresponding to the audio code stream. This can be achieved by calling a sixth neural network to perform feature reconstruction on the high-frequency feature estimation value included in the coding feature to obtain a high-frequency sub-band signal estimation value corresponding to the audio code stream, wherein the structure of the sixth neural network is similar to that of the third neural network.

[0274] In some embodiments, step 2032 may be implemented through steps 20321 to 20324:

[0275] In step 20321, frequency domain transformation is performed on the first half of the sample points and the second half of the sample points included in the low-frequency subband signal estimation value to obtain first transformation coefficients corresponding to the first half of the sample points and second transformation coefficients corresponding to the second half of the sample points.

[0276] It should be noted that the frequency domain transformation methods of the embodiments of the present application include modified discrete cosine transform (MDCT), discrete cosine transform (DCT), fast Fourier transform (FFT), etc. The embodiments of the present application are not limited to the frequency domain transformation method.

[0277] In step 20322, based on the first transform coefficient, the high-frequency feature estimation value is subjected to inverse frequency band expansion processing to obtain a first high-frequency sub-band signal estimation value.

[0278] In some embodiments, step 20322 can be implemented by: performing spectrum replication processing on the transform coefficients of the second half of the first transform coefficients to obtain the first reference transform coefficients of the reference high-frequency sub-band signal; based on the first half of the sub-band spectrum envelope corresponding to the high-frequency feature estimation value, performing a gain on the first reference transform coefficient to obtain the gained first reference transform coefficient; performing an inverse frequency domain transform on the gained first reference transform coefficient to obtain the first high-frequency sub-band signal estimation value.

[0279] In some embodiments, based on the first half sub-band spectral envelope corresponding to the high-frequency feature estimation value, the first reference transformation coefficient is gained to obtain the gained first reference transformation coefficient, including: based on the first half sub-band spectral envelope, the first reference transformation coefficient of the reference high-frequency sub-band signal is divided into multiple first sub-bands; for any first sub-band among the multiple first sub-bands, the following processing is performed: determining the first average energy corresponding to the first sub-band in the first half sub-band spectral envelope, and determining the second average energy corresponding to the first sub-band; determining the gain factor based on the ratio of the first average energy to the second average energy; multiplying the gain factor by each first reference transformation coefficient included in the first sub-band to obtain the gained first reference transformation coefficient.

[0280] As an example, firstly, the low frequency subband signal estimation value x′ generated by the decoding end is LB (n), two 320-point MDCT transforms are also performed to generate two groups of 160-point MDCT coefficients (i.e., the first transform coefficients corresponding to the sample points in the first half and the second transform coefficients corresponding to the sample points in the second half).

[0281] Then, for the first group of 160 MDCT coefficients (i.e., the first transform coefficients corresponding to the first half of the sample points), x′ LB (n) The 160-point MDCT coefficients generated are copied to generate the MDCT coefficients for the high-frequency portion. Referring to the basic characteristics of the speech signal, the low-frequency portion has more harmonics and the high-frequency portion has fewer harmonics. Therefore, to avoid simple copying, which would cause the artificially generated high-frequency MDCT spectrum to contain too many harmonics, the last 80 points of the 160-point MDCT coefficients that the low-frequency subband relies on can be used as a master. The spectrum can be copied twice to generate reference values ​​for the 160-point MDCT coefficients of the high-frequency subband signal (i.e., the first reference transform coefficients of the reference high-frequency subband signal).

[0282] Next, the 8 sub-band spectral envelopes obtained previously are called (i.e., the 8 sub-band spectral envelopes obtained after querying the quantization table, i.e., the sub-band spectral envelopes corresponding to the high-frequency features). These 8 sub-band spectral envelopes correspond to 8 high-frequency sub-bands, and the reference values ​​of the MDCT coefficients of the generated 160-point reference high-frequency sub-band signal are divided into 8 reference high-frequency sub-bands (i.e., the first reference transform coefficients of the reference high-frequency sub-band signal are divided into multiple first sub-bands). Band-wise, based on a high-frequency sub-band and the corresponding reference high-frequency sub-band, the reference values ​​of the MDCT coefficients of the generated 320-point reference high-frequency sub-band signal are gained (multiplication in the frequency domain). For example, the gain factor is calculated based on the average energy of the high-frequency sub-band (i.e., the first average energy) and the average energy of the corresponding reference high-frequency sub-band (the second average energy), and the MDCT coefficient corresponding to each point in the corresponding reference high-frequency sub-band is multiplied by the gain factor to ensure that the energy of the virtually generated high-frequency MDCT coefficients at the decoding end is close to the original coefficient energy at the encoding end.

[0283] For example, assuming that the average energy of the reference high-frequency sub-band (i.e., the first sub-band divided by the reference values ​​of the 160-point MDCT coefficients of the generated high-frequency signal) is Y_L, and the average energy of the current high-frequency sub-band (i.e., the sub-band corresponding to the sub-band spectrum envelope decoded based on the bitstream) is Y_H, then a gain factor a = sqrt(Y_H / Y_L) is calculated, where sqrt() represents a square root calculation function, which is used to calculate the square root of (Y_H / Y_L). With the gain factor a, the MDCT coefficients of each point in the reference high-frequency sub-band are directly multiplied by a. The average energy of the MDCT coefficients (virtually generated) after gain is very close to the original one at the encoding end.

[0284] Finally, an inverse MDCT transform is called to generate an estimated value of the first subframe of the high frequency subband signal (ie, a first high frequency subband signal estimated value).

[0285] In step 20323, based on the second transform coefficient, the high-frequency feature estimation value is subjected to inverse frequency band expansion processing to obtain a second high-frequency sub-band signal estimation value.

[0286] Here, step 20323 is similar to step 20322, except that the processing object of step 20323 is the second transformation coefficient, and the processing object of step 20322 is the first transformation coefficient.

[0287] In some embodiments, step 20323 can be implemented by: performing spectrum replication processing on the transform coefficients of the second half of the second transform coefficients to obtain second reference transform coefficients of the reference high-frequency sub-band signal; performing a gain on the second reference transform coefficients based on the second half of the sub-band spectrum envelope corresponding to the high-frequency feature estimation value to obtain the gained second reference transform coefficients; performing an inverse frequency domain transform on the gained second reference transform coefficients to obtain a second high-frequency sub-band signal estimation value.

[0288] In some embodiments, based on the second half of the sub-band spectrum envelope corresponding to the high-frequency feature estimation value, the second reference transformation coefficient is gained to obtain the gained second reference transformation coefficient, including: based on the second half of the sub-band spectrum envelope, the second reference transformation coefficient of the reference high-frequency sub-band signal is divided into multiple second sub-bands; for any second sub-band among the multiple second sub-bands, the following processing is performed: determining the first average energy corresponding to the second sub-band in the second half of the sub-band spectrum envelope, and determining the second average energy corresponding to the second sub-band; determining a gain factor based on the ratio of the first average energy to the second average energy; multiplying the gain factor by each second reference transformation coefficient included in the second sub-band to obtain the gained second reference transformation coefficient.

[0289] In step 20324, the first high frequency sub-band signal estimation value and the second high frequency sub-band signal estimation value are combined to obtain a high frequency sub-band signal estimation value corresponding to the audio bit stream.

[0290] In step 2033, sub-band synthesis is performed on the low-frequency sub-band signal estimation value and the high-frequency sub-band signal estimation value to obtain a reconstructed audio signal corresponding to the audio code stream.

[0291] For example, subband synthesis is the inverse process of subband decomposition. The decoding end performs subband synthesis on the low-frequency subband signal estimation value and the high-frequency subband signal estimation value to restore the reconstructed audio signal, wherein the reconstructed audio signal is the restored reconstructed signal.

[0292] In some embodiments, sub-band synthesis is performed on the low-frequency sub-band signal estimation value and the high-frequency sub-band signal estimation value to obtain a reconstructed audio signal, including: up-sampling the low-frequency sub-band signal estimation value to obtain a low-pass filtered signal; up-sampling the high-frequency sub-band signal estimation value to obtain a high-frequency filtered signal; and filtering and synthesizing the low-pass filtered signal and the high-frequency filtered signal to obtain a reconstructed audio signal.

[0293] For example, after obtaining the low-frequency subband signal estimation value and the high-frequency subband signal estimation value, the low-frequency subband signal estimation value and the high-frequency subband signal estimation value are subjected to subband synthesis through a QMF synthesis filter to restore the reconstructed audio signal.

[0294] In some embodiments, the frame header further includes at least one channel bit, and the channel bit is used to indicate whether a mono decoding method or a stereo decoding method is used to decode the audio stream. When the channel bit in the frame header included in the audio stream encapsulation is a first value, it indicates that a stereo decoding method is required to decode the audio stream, that is, a synthetic stereo signal is reconstructed by using the decoded reconstructed audio signal and the stereo feature vector in the audio stream encapsulation; when the channel bit in the frame header included in the audio stream encapsulation is a second value, it indicates that a mono decoding method is required to decode the audio stream, that is, the decoded reconstructed audio signal is used as a mono reconstructed audio signal.

[0295] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0296] The embodiments of the present application can be applied to various audio scenarios, such as voice calls, instant messaging, etc. The following description will be made using a voice call as an example:

[0297] In related technologies, the principles of speech coding are roughly as follows: speech coding can directly encode speech waveform samples sample by sample; or, based on the principle of human vocalization, relevant low-dimensional features are extracted, the encoding end encodes the features, and the decoding end reconstructs the speech signal based on these parameters.

[0298] The above coding principles all come from speech signal modeling, that is, a compression method based on signal processing. In order to improve the coding quality while ensuring the efficiency of speech coding relative to the compression method based on signal processing. The embodiment of the present application provides a multi-mode multi-rate speech coding method (i.e., an audio coding method and an audio decoding method). Based on the characteristics of the audio signal, for the important part (low-frequency sub-band signal), after processing based on the neural network (NN, Neural Network) technology, a feature vector with a lower dimension than the input low-frequency sub-band signal is obtained. Among them, the neural network adopts an operation similar to "blocking" to reduce the complexity of the algorithm and improve the coding effect. For low-frequency feature vectors (i.e., the above-mentioned low-frequency features), code tables with different quantization precisions can be used for quantization and encoding based on the same low-frequency feature vector to achieve multi-rate coding and decoding effects. For high-frequency feature vectors (i.e., the above-mentioned high-frequency features), multi-level quantization and encoding are adopted to achieve multi-rate coding and decoding effects.

[0299] The embodiment of the present application can be applied to the voice communication link shown in Figure 6C. Taking the Voice over Internet Protocol (VoIP) conference system as an example, the voice codec technology involved in the embodiment of the present application is deployed in the encoding and decoding parts to solve the basic function of voice compression. The encoder is deployed on the uplink client 601, and the decoder is deployed on the downlink client 602. The voice is collected through the uplink client and pre-processed, enhanced, and encoded. The encoded code stream is transmitted to the downlink client 602 via the network. The downlink client 602 performs decoding, enhancement, and other processing to play back the decoded voice on the downlink client 602.

[0300] To ensure forward compatibility (i.e., compatibility between the new encoder and the existing encoder), a transcoder must be deployed in the system's backend (i.e., server) to ensure interoperability between the new encoder and the existing encoder. For example, if the transmitter (uplink client) is the new NN encoder and the receiver (downlink client) is the Public Switched Telephone Network (PSTN) (G.722), the backend must execute the NN decoder to generate the voice signal and then call the G.722 encoder to generate a specific bitstream to implement the transcoding function. This allows the receiver to correctly decode the specific bitstream.

[0301] The audio encoding method and audio decoding method provided in the embodiments of the present application will be described below in conjunction with the high-frequency part and the low-frequency part of the mono signal.

[0302] The following describes the audio encoding method and audio decoding method provided by the embodiment of the present application with reference to FIG7 :

[0303] The following processing is performed on the encoding side:

[0304] For the input audio signal x(n) of the nth frame, use the analysis filter to decompose it into the low-frequency subband signal x LB (n) and high frequency sub-band signal x HB (n).

[0305] For the low-frequency subband signal x LB (n), call the second NN to obtain the low-dimensional feature vector F LB (n), that is, low-frequency features, feature vector F LB The dimension of (n) is smaller than that of the low-frequency subband signal to reduce the amount of data. For example, for each frame x LB (n), call the neural network (encoding part) to generate a lower-dimensional feature vector F LB(n). The embodiments of the present application do not limit other NN structures, such as autoencoder, full-connection (FC) network, long short-term memory (LSTM) network, convolutional neural network (CNN) + LSTM, etc. Among them, the use of similar "blocking" operations within the neural network can reduce the complexity of the algorithm and improve the encoding effect. LB (n), using code tables with different quantization precisions for quantization and encoding to achieve multi-rate encoding and decoding effects.

[0306] For the high frequency subband signal x HB (n), considering that high frequencies are not as important to quality as low frequencies, the high frequency subband signal x HB (n) Other schemes can be used to extract the feature vector F HB (n). For example, the frequency band extension technology based on speech signal analysis can realize the generation of high-frequency sub-band signals with only a small number of bits; it can also use the same NN structure as the low-frequency sub-band signal or a more streamlined network (for example, the output feature vector is smaller than the low-frequency feature vector F LB (n) is smaller). For the high frequency feature vector F HB (n), adopts multi-level quantization and encoding to achieve multi-rate encoding and decoding effects.

[0307] The eigenvector corresponding to the subband signal (i.e. F LB (n) and F HB (n)) performs vector quantization or scalar quantization, and performs entropy coding on the quantized value, and transmits the encoded code stream (low-frequency code stream and high-frequency code stream) to the decoding end.

[0308] The following processing is performed on the decoding end:

[0309] Decode the code stream (including low-frequency code stream and high-frequency code stream) received by the decoding end to obtain the estimated value F′ of the low-frequency feature vector LB (n) and the estimated value of the high-frequency eigenvector F′ HB (n).

[0310] For the low-frequency part, the estimated value F′ based on the low-frequency eigenvector LB (n) Call the third NN to obtain the estimated value x′ of the low-frequency subband signal LB (n). Among them, the neural network adopts a similar "blocking" operation to reduce the complexity of the algorithm and improve the encoding effect. According to the actual codeword length in the code stream, code tables with different quantization precisions can be used to generate F' with different precisions. LB (n) and x′LB (n), to achieve multi-rate encoding and decoding effects.

[0311] For the high frequency part, the estimated value F′ based on the high frequency eigenvector HB (n) Call high-frequency reconstruction to generate an estimate x′ of the high-frequency subband signal HB (n). According to the actual codeword length in the code stream, code tables with different quantization precisions can be used to generate F′ with different precisions. HB (n) and x′ HB (n), to achieve multi-rate encoding and decoding effects.

[0312] Finally, the QMF synthesis filter is called to generate the reconstructed synthetic speech signal x′(n).

[0313] The audio encoding method and audio decoding method provided in the embodiments of the present application are described in detail below.

[0314] In some embodiments, a speech signal with a sampling rate of Fs = 32000 Hz is used as an example (it should be noted that the method provided in the embodiments of the present application is also applicable to scenarios with other sampling rates, including but not limited to: 8000 Hz, 32000 Hz, and 48000 Hz). At the same time, assuming that the frame length is set to 20 ms, then for Fs = 32000 Hz, each frame contains 640 sample points.

[0315] The following describes the encoding end and the decoding end in detail with reference to the flowchart shown in FIG7 .

[0316] The process of encoding the low-frequency and high-frequency parts of the mono signal is as follows:

[0317] For an audio signal (mono-channel signal) with a sampling rate of Fs=32000 Hz, the input signal of the nth frame includes 640 sample points, which are recorded as input signal x(n), that is, a mono-channel input signal.

[0318] Step 11: Call the QMF analysis filter to decompose the signal.

[0319] Call the QMF analysis filter (2-channel QMF) and perform downsampling to obtain two sub-band signals, namely the low-frequency sub-band signal x LB (n) and high frequency sub-band signal x HB (n). Low frequency subband signal x LB (n) and high frequency sub-band signal x HB The effective bandwidth of (n) is 0-8kHz and 8-16kHz respectively, and the low-frequency sub-band signal x LB (n) and high frequency sub-band signal x HB The number of sample points of (n) is 320.

[0320] Step 12: Based on the low-frequency subband signal, call the second NN.

[0321] Based on the low-frequency subband signal x LB (n), call the second NN to generate a lower-dimensional feature vector F LB (n). It should be noted that x LB The dimension of (n) is 320, F LB The dimension of (n) is 56. From the perspective of data volume, the second NN plays the role of "dimensionality reduction" and realizes the function of data compression. LB (n) dimension, or other dimensions smaller than x LB (n) dimension.

[0322] Referring to the network structure diagram of the second NN shown in FIG11 , the process of data compression performed by the second NN is described in detail below:

[0323] First, a 16-channel causal convolution is called to expand the input tensor (i.e., vector) to a 16×320 tensor.

[0324] Then, the 16×320 tensor is preprocessed. For example, after performing a convolution operation on the 16×320 tensor, a pooling operation with a factor of 2 is performed, and the activation function can be PReLU to generate a 16×160 tensor.

[0325] Next, four encoding blocks with different downsampling factors (Down_factor) are cascaded. Each encoding block contains a residual block, a convolutional layer, and a pooling layer. Each residual block includes five residual units (RUs) based on dilated convolution (the input and output feature dimensions of the RUs do not change). A convolutional layer is used to double the number of input channels, and the activation function can be PReLU to ensure data volume and avoid data loss. The pooling layer is a pooling operation with a Down_factor to complete downsampling and achieve data compression. Here, the Down_factor of the four encoding blocks is set to 2, 4, 4, and 5, respectively. Therefore, the number of output channels of the four encoding blocks is set to 32, 64, 128, and 256, respectively. After processing by the four encoding blocks, the input 16×160 tensor is converted into 32×80, 64×20, 128×5, and 256×1 tensors, respectively. The embodiment of the present application is not limited to the number of coding blocks, and can be any positive integer such as 2, 3, 4, 5, etc. In addition, the embodiment of the present application is not limited to the number of residual units in a coding block, and can be any positive integer such as 2, 3, 4, 5, 6, etc. The number of residual units in multiple coding blocks can be the same or different, for example, one coding block contains 4 residual units and another coding block contains 5 residual units.

[0326] Here we further introduce the residual unit. The residual unit refers to a module in a deep neural network. By introducing cross-layer connections in the neural network, the neural network is easier to optimize during the training process, avoiding problems such as gradient disappearance or gradient explosion. The core idea is to perform residual learning on the input within the module, that is, to bypass a part of the layer through a direct path and pass the input information directly to the output, so that the network can better utilize shallow feature information during the learning process. Figure 12A is a schematic diagram of the residual block structure used in the encoding block in the third NN. This residual block includes 4 residual units based on dilated convolution, and each residual unit contains a dilated convolution block with a specified dilation rate (Dilation rate), that is, each dilated convolution block contains a convolution operator with a specified dilation rate (such as Dilation rate = 3). In an embodiment of the present application, the use of dilated convolution blocks with 5 progressive dilation rates is equivalent to using different receptive fields to extract input features at different resolutions, which can better perform a relatively comprehensive analysis of the data. After the residual is processed by 5 dilation-rate-specified dilation-rate dilated convolution blocks, it is added to the input from the jump connection to obtain the output of the residual block, and the output is sent to the convolution layer connected to the residual block.

[0327] Here, any residual unit in Figure 12A is further described, as shown in Figure 12B. For any residual unit, it contains at least one dilated convolution with a specified expansion rate (used to expand the receptive field), and PReLU can be used as the activation function; in addition, one or more causal convolutions (used to extract local information) can be cascaded, and PReLU can be used as the activation function. The convolution kernel size of the dilated convolution with the above-mentioned specified expansion rate can be 3, 5, 7, 9, etc., and the convolution kernel size of the above-mentioned causal convolution can be 1, 3, etc. This embodiment of the present application does not limit the convolution kernel size of the dilated convolution or causal convolution with the above-mentioned specified expansion rate. In addition, the causal convolution or dilated convolution in the embodiment of the present application can also be implemented by other specific convolution units with similar or equivalent functions.

[0328] Furthermore, for residual units, a grouped convolution algorithm is introduced to reduce algorithmic complexity. Grouped convolution divides the input channels into multiple groups for convolution operations, where only the input channels within each group are associated with the output channels. Here, assume there are 16 input channels and 32 output channels. If the number of groups is 1, each input channel is associated with all 32 output channels. If the number of groups is 2, the 16 input channels are first divided into two groups, 0-7 and 8-15. Within each group, the input channels are associated with the output channels within that group. For example, input channels 0-7 in the first group are associated with output channels 0-15, while input channels 8-15 in the second group are associated with output channels 16-31. For example, output channel 0 is associated only with input channels 0-7, not with input channels 8-15. Output channel 25 is associated only with input channels 8-15, not with input channels 0-7. From this comparison, it can be seen that the introduction of grouped convolution can avoid the association between any input channel and all output channels, reduce the number of connections, and reduce complexity. Of course, since the larger the number of groups, the smaller the correlation between the input channel and the output channel, which will also affect the encoding effect, it is not the case that the larger the number of groups, the better. In the embodiment of the present application, the hole convolution contained in the 5 residual blocks corresponding to the 4 coding blocks can use different grouping number configurations. The specific grouping number configuration is shown in Table 1.

[0329] Table 1. Grouping configurations used by residual units in different coding blocks

[0330] Finally, after similar preprocessing of the causal convolution, the 256×1 tensor can output a 56-dimensional feature vector F LB (n). According to the second NN calculation, the 56-dimensional feature vector F LB Each value in (n) is between [-1,1].

[0331] Step 13: For the high frequency sub-band signal xHB (n) Perform high-frequency analysis.

[0332] The purpose of high frequency analysis is to extract the high frequency subband signal x HB (n) to generate a lower-dimensional feature vector F HB (n). The present application embodiment is not limited to F HB (n) dimension, or other dimensions smaller than x LB (n) dimension, but F HB The dimension of (n) must be smaller than F LB (n) dimension.

[0333] In some embodiments, referring to step 12, another NN structure similar to the second NN can be introduced to generate a low-dimensional feature vector. High-frequency subband signals are less critical to quality than low-frequency subband signals. Therefore, the NN structure for high-frequency subband signals does not need to be as complex as the second NN. The NN structure for high-frequency subband signals is similar to that of the second NN, but compared to the second NN structure, the NN structure for high-frequency subband signals has a significantly reduced number of channels.

[0334] The present embodiment proposes another method for compressing high-frequency sub-band signals, namely, frequency band expansion (recovering a broadband audio signal from a band-limited narrowband audio signal). The following describes in detail the application of frequency band expansion in the present embodiment:

[0335] For the high frequency subband signal x consisting of 320 points (ie, sample points) HB (n), the 320-point high-frequency sub-band signal x HB (n), divided into two subframes (the first subframe and the second subframe), each containing 160 points. Compared to the 640-point 50% overlap MDCT transform, the algorithm delay can be reduced by 10ms by using 320-point 50% overlap MDCT transform by framing. Therefore, for each 20ms frame, two 320-point MDCT transforms can be performed to extract features.

[0336] For any subframe, i.e., a high-frequency subband signal comprising 160 points, a modified discrete cosine transform (MDCT) is invoked to generate 160-point MDCT coefficients. Specifically, if there is a 50% overlap, the first subframe of the n+1th frame can be merged (concatenated) with the second subframe of the nth frame, and a 320-point MDCT can be calculated to obtain the 160-point MDCT coefficients for the first subframe of the n+1th frame. The second subframe of the n+1th frame can also be merged with the first subframe of the n+1th frame, and a 320-point MDCT can be calculated to obtain the 160-point MDCT coefficients for the second subframe of the n+1th frame.

[0337] For the 160-point MDCT coefficients of any one of the above two subframes, the following processing is performed in the same manner: the MDCT coefficients of the 160 points are divided into N subbands, where a subband is a group of multiple adjacent MDCT coefficients. The MDCT coefficients of the 160 points can be divided into 4 subbands. For example, the 160 points can be evenly distributed, that is, the number of points included in each subband is the same. Of course, the embodiment of the present application cannot perform non-uniform division of the 160 points, such as a subband with a lower frequency includes fewer MDCT coefficients (higher frequency resolution), and a subband with a higher frequency includes more MDCT coefficients (lower frequency resolution).

[0338] According to Nyquist's sampling theorem (to recover the original signal from a sampled signal without distortion, the sampling frequency must be greater than twice the original signal's highest frequency. If the sampling frequency is less than twice the highest frequency, the signal spectrum will exhibit aliasing, while if the sampling frequency is greater than twice the highest frequency, the signal spectrum will exhibit no aliasing), the 320-point MDCT coefficients above represent a spectrum between 8 and 16 kHz. However, ultra-wideband voice communication does not necessarily require a spectrum extending to 16 kHz. For example, if the spectrum is set to 14 kHz, only the MDCT coefficients of the first 240 points need to be considered, and the number of subbands can be controlled to 6.

[0339] For each sub-band, calculate the average energy of all MDCT coefficients in the current sub-band as the sub-band spectrum envelope (the spectrum envelope is a smooth curve passing through the main peak points of the spectrum). For example, the MDCT coefficients included in the current sub-band are x(n), n = 1, 2, ..., 40, then the average energy Y = ((x(1) 2 +x(2) 2 +…+x(40) 2 ) / 40). When the 160-point MDCT coefficients are divided into 4 sub-bands, 4 sub-band spectrum envelopes can be obtained. These 4 sub-band spectrum envelopes are the sub-band spectrum envelopes corresponding to the sub-bands. Where i=1 represents the first subframe, and i=2 represents the second subframe. Get the eigenvector F of the high-frequency subband signal HB (n).

[0340] In summary, using either of the two methods (NN structure and band expansion) described above, a 160-dimensional high-frequency subband signal can be output as a 4-dimensional feature vector. Therefore, for each 20ms frame, only 8 dimensions of data are needed to represent the high-frequency information, significantly improving coding efficiency.

[0341] Step 14: quantization encoding.

[0342] For the eigenvector F of the low-frequency subband signalLB (n) and the eigenvector F of the high-frequency subband signal HB (n), both scalar quantization (each component is quantized separately) and entropy coding methods can be used. In addition, the embodiments of the present application do not limit the technical combination of vector quantization (combining multiple adjacent components into a vector for joint quantization) and entropy coding.

[0343] Here for the low-frequency eigenvector F LB According to the above description, after the low-frequency sub-band signal is processed by the second NN, a 56-dimensional feature vector F is obtained. LB (n) The present application embodiment provides a method based on scalar quantization and entropy coding, including: 1) for F LB For each dimension in (n), the interval [-1, 1] is divided into 11 equal parts to form a codebook containing 11 elements, and each dimension value is quantized into one of the 11 elements; 2) According to the Shannon entropy theorem, for a codebook containing 11 elements and uniformly distributed, the entropy (average bits) is Therefore, the average bit rate corresponding to each frame of the low-frequency subband signal is 193.76 bits. 3) For a 20ms frame, there are 50 frames per second, resulting in an average bitrate of 9.69 kbps. Based on entropy coding theory, we can perform probability distribution statistics for each of the above dimensions and generate 56 code tables. Generally, each dimension has a non-uniform distribution, so the actual bitrate is around 9.69 kbps, or even less.

[0344] In addition, in order to realize the low-frequency feature vector F LB The goal of multi-rate encoding and decoding (n) can be implemented as follows:

[0345] The code table of the above implementation method, that is, the code table containing 11 elements in each dimension, is named low-frequency code table-1, and the reference bit rate is 9.69 kbps.

[0346] Similarly, for each dimension, the interval [-1, 1] is divided into 9 equal parts, resulting in an entropy of 3.17 and an average bitrate of 3.17 * 56 * 50 / 1000 = 8.88 kbps. This is named Low-Frequency Code Table-2, and the reference bitrate is 8.88 kbps. This provides at least two bitrate modes, each corresponding to different quality levels, achieving multi-rate encoding.

[0347] Similarly, for each dimension, the interval [-1, 1] is divided into 7 equal parts, with an average bit rate of 7.86 kbps, which is named low-frequency code table-3; the interval [-1, 1] is divided into 5 equal parts, with an average bit rate of 6.50 kbps, which is named low-frequency code table-4.

[0348] In this way, for the original 56-dimensional feature vector calculated for each frame of data, by configuring the quantization accuracy, an encoding mode of low-frequency features of at least 4 code rates is achieved. It should be noted that the embodiment of the present application does not limit other multi-rate construction methods and the number of code tables.

[0349] For each frame (including 2 subframes), the eigenvector F of the 8-dimensional high-frequency subband signal HB (n), a specific implementation of its quantization coding includes: 1) performing scalar quantization on each dimension separately, allocating 5 bits; thus, each frame is allocated 40 bits; 2) for a 20ms frame mode, there are 50 frames per second, so the average bit rate is 2kbps. If entropy coding (including but not limited to Huffman coding or range coding) is performed on the quantized values ​​of each dimension, the overall bit rate will be less than 2kbps. The above code table is defined as high-frequency code table-1, and the quantization value is defined as the high-frequency feature vector quantization value (i.e., the quantization value of the high-frequency feature).

[0350] In addition, in order to realize the high frequency feature vector F HB The multi-rate encoding and decoding goal of (n) can be further introduced into residual processing, which can be implemented as follows:

[0351] 1) For each 10ms subframe, the 4 subbands of the subframe are further evenly divided into 2 subsets.

[0352] For example, the set of envelope values ​​of four sub-bands is {e1, e2, e3, e4}.

[0353] 2) For each subset of each subband, calculate the envelope value of the subset.

[0354] For example, the set of envelope values ​​of all subsets is {e 1a ,e 1b ,e 2a ,e 2b ,e 3a ,e 3b ,e 4a ,e 4b}.

[0355] 3) Subtract the corresponding high-frequency eigenvector quantization value from the envelope value of each subset to obtain the first high-frequency eigenvector residual (referred to as the first residual value).

[0356] For example, the high-frequency feature vector quantization value is Then the first high-frequency eigenvector residual is

[0357] 4) The first high-frequency feature residual is subjected to similar quantization and entropy coding to obtain a first high-frequency feature vector residual quantized value (hereinafter referred to as the quantized value of the first residual value).

[0358] Here, considering that the dynamic range of the residual value is much lower than the original value, only 2 bits need to be allocated when scalar quantizing each residual. In this way, a reference bit rate of 2*16*50=1.6kbps is required (similarly, because it is non-uniformly distributed entropy coding, the actual bit rate is generally less than 1.6kbps) to obtain the residual quantization value of the first high-frequency eigenvector.

[0359] 5) Subtract the sum of the corresponding high-frequency feature vector quantization value and the first high-frequency feature vector residual quantization value from the envelope value of each subset to obtain the second high-frequency feature vector residual (referred to as the second residual value).

[0360] Here, similarly, the quantization and entropy coding of the second high-frequency eigenvector residual can be completed using a 1.6 kbps reference bit rate.

[0361] Similarly, as more residuals are calculated, it is possible to ensure that the high-frequency eigenvector reconstructed at the decoding end is infinitely close to the original high-frequency eigenvector. In addition, based on the presence or absence of high-frequency eigenvector residuals, or the number of high-frequency eigenvector residuals, a multi-rate coding effect is achieved. The embodiment of the present application can define the number of types of residual coding (no residual coding, first high-frequency eigenvector residual, second high-frequency eigenvector residual, third high-frequency eigenvector residual, etc.) according to the bit rate requirements, and is not limited here.

[0362] In addition, an embodiment of the present application provides a solution for additionally extracting side information at the encoding end, which is used to guide more refined processing at the decoding end to improve audio quality. In principle, the main function of frequency band expansion is to copy the low-frequency spectrum to the high frequency and then perform amplification. However, the high-frequency spectrum is flatter than the low-frequency spectrum and does not have so many harmonic components. If directly copied, the artificially generated high frequency will contain too many harmonic components. Therefore, an embodiment of the present application proposes to estimate specific side information at the encoding end and write it into the bitstream. Based on the above-mentioned flatness side information, the decoding end determines whether additional processing is required at the decoding end. It should be noted that the above-mentioned flatness side information will only prompt the decoding end whether the current frame needs additional processing. In a special case, the entire speech does not require additional processing.

[0363] The above side information estimation is implemented as follows at the encoder end:

[0364] 1) For each 10ms subframe, the 160-point MDCT coefficients are divided into two blocks. The division can be uniform, or the MDCT coefficients of the first two subbands can be used as the first block and the MDCT coefficients of the last two subbands can be used as the second block.

[0365] 2) For each block, calculate the power of each MDCT coefficient, i.e., p(i) = c(i) 2 .

[0366] 3) Calculate the arithmetic mean of all MDCT coefficients of the current block: where I represents the number of MDCT coefficients of the current block.

[0367] 4) Calculate the geometric mean of all MDCT coefficients of the current block: where In represents the logarithmic operation and I represents the number of MDCT coefficients of the current block.

[0368] 5) Calculate the first flatness of the current block Generally, sfp Hi is a value in [0, 1].

[0369] 6) Similarly, calculate the second flatness sfp of the MDCT coefficients in the specified low-frequency band Lo .

[0370] 7) When (sfp Hi [i] < sfp Lo [i]) or (sfp Hi [i] < SF_THD), Sfp_FLAG = 0; otherwise, Sfp_FLAG = 1. Here, SF_THD represents the set flatness threshold, and Sfp_FLAG represents the flatness side information.

[0371] Here, when Sfp_FLAG = 1, it means that the decoding end needs additional operations to avoid excessive harmonics artificially generated at high frequencies. According to the above operations, 4 additional bits are required for each frame, representing the flatness side information of two blocks in two 10 - ms sub - frames. The embodiments of the present application are not limited to the above extraction method, and other analysis methods can also be used to extract the above flatness side information to guide the decoding process at the decoding end later.

[0372] After quantization coding, a bitstream can be generated. According to experiments, in the range of 5 - 10 kbps, high - quality compression can be achieved for 16 - kHz wideband signals; in the range of 8 - 15 kbps, high - quality compression can be achieved for 32 - kHz ultra - wideband signals.

[0373] In addition, in order to achieve the above - mentioned goal of multi - mode and multi - rate encoding and decoding, a good bitstream structure needs to be designed. As shown in FIGS. 13A - 13E, the implementation form of the bitstream structure is as follows:

[0374] For each frame of code stream, it generally includes an 8-bit frame header. Among them, the 8-bit frame header contains at least a 1-bit coding mode bit. According to the 1-bit coding mode bit, it can be known whether to use wideband coding (the value of the 1-bit coding mode bit is 0) or ultra-wideband coding (the value of the 1-bit coding mode bit is 1), that is, the coding mode includes two types (wideband coding and ultra-wideband coding). In addition, it also includes a 2-bit coding rate bit. According to the 2-bit coding rate bit, it can be known which code rate mode is used in the corresponding wideband coding mode or ultra-wideband coding mode. It should be noted that the embodiments of the present application are not limited to the positions of the 1-bit coding mode bit and the 2-bit coding rate bit in the frame header. For example, the 1-bit coding mode bit can be located at the first position of the frame header, and the 2-bit coding rate bit can be located at the last two positions of the frame header.

[0375] As shown in Figure 13A, for wideband coding mode, the frame header is followed by a 56-dimensional low-frequency feature vector index. Because entropy coding is used, the code length varies. As described above, depending on the 2-bit rate mode bit, the encoded low-frequency bitstream is quantized using different low-frequency code tables and stored in the 56-dimensional low-frequency feature vector index.

[0376] The 2-bit coding rate in wideband coding mode is explained as follows:

[0377] 1) 00, indicating the use of low-frequency code table -4, with a reference average bit rate of 6.5kbps, which is also the lowest bit rate in wideband coding mode;

[0378] 2) 01, indicating the use of low-frequency code table-3, with a reference average bit rate of 7.86 kbps;

[0379] 3) 10, indicating the use of low-frequency code table-2, with a reference average bit rate of 8.88 kbps;

[0380] 4) 11, indicating the use of low-frequency code table -1, with a reference average bit rate of 9.69 kbps, which is also the highest bit rate in the wideband coding mode.

[0381] For ultra-wideband coding mode, the 2-bit rate mode bit is defined according to whether residual coding is used. The 2-bit code rate bit in ultra-wideband coding mode is interpreted as follows:

[0382] 1) 00 (excluding residual coding) indicates that the low-frequency part uses low-frequency code table-4. The bitstream encapsulation structure includes quantization coding using low-frequency code table-4, ultra-wideband coding of 4 spectral envelopes, and 4-bit flatness. The reference average bit rate is 6.5+2+0.2=8.7kbps.

[0383] 2) 01 (excluding residual coding) indicates that the low-frequency part uses low-frequency code table -1. The bitstream encapsulation structure includes quantization coding using low-frequency code table -1, ultra-wideband coding of 4 spectral envelopes, and 4-bit flatness. The reference average bit rate is 9.69 + 2 + 0.2 = 11.9 kbps.

[0384] 3) 10 (including first residual coding), indicating that the low-frequency part uses low-frequency code table -1. The code stream encapsulation structure includes quantization coding using low-frequency code table -1, ultra-wideband coding of 4 spectral envelopes, 8 first residuals, and 4-bit flatness. The reference average bit rate is 9.69 + 2 + 1.6 + 0.2 = 13.5 kbps;

[0385] 4) 11 (including the first residual coding and the second residual coding) indicates that the low-frequency part uses the low-frequency code table-1. The code stream encapsulation structure includes quantization coding using the low-frequency code table-1, ultra-wideband coding of 4 spectral envelopes, 8 first residuals, 8 second residuals, and 4-bit flatness. The reference average bit rate is 9.69+2+1.6+1.6+0.2=15.0kbps.

[0386] As shown in Figures 13A-13C , based on the 2-bit rate mode bit, the low-frequency feature vector index and the high-frequency feature vector index are placed immediately after the frame header. As shown in Figures 13D-13E , if the bitstream encapsulation structure includes 4-bit flatness edge information, the 4-bit edge information can be placed immediately after the 8-dimensional high-dimensional feature vector index.

[0387] In summary, when using an encoder for audio encoding, the user inputs the voice data to be encoded and the correct configuration parameters to the encoder. For example, when the input voice data is wideband, the sampling rate is 16000Hz, and when the input voice data is ultra-wideband, the sampling rate is 32000Hz. For ease of description, a wideband signal sampled at 16000Hz and an ultra-wideband signal sampled at 32000Hz are used as examples, but the embodiments of the present application are not limited to other combinations.

[0388] In addition, the user needs to configure the configuration parameters of the bit rate mode to be used, and the encoder can select the appropriate code table based on the configured bit rate mode. Generally, the higher the bit rate mode, the more bits the encoder will use for encoding, corresponding to higher reconstructed voice quality. As shown above, the embodiment of the present application provides four bit rate modes for both the wideband coding mode and the ultra-wideband coding mode, but the embodiment of the present application is not limited to the four bit rate modes.

[0389] For example, when the user inputs a broadband signal based on 16000Hz sampling, the encoder can use the broadband coding mode for encoding, directly use the second NN on the input broadband signal to extract low-frequency features, quantize and encode the extracted low-frequency features to obtain the code stream of the broadband signal; when the user inputs an ultra-wideband signal based on 32000Hz sampling, the encoder can use the ultra-wideband coding mode for encoding, first pass the input ultra-wideband signal through the analysis filter to separate the low-frequency sub-band signal and the high-frequency sub-band signal, and directly use the second NN on the low-frequency sub-band signal to extract the low-frequency feature F LB (n), for F LB (n) Quantization coding is performed to obtain low-frequency code stream, and high-frequency sub-band signals are subjected to high-frequency analysis to extract high-frequency features F HB (n), for F HB (n) Perform quantization encoding to obtain a high-frequency code stream. In addition, according to the code rate mode input at the encoding end, the encoder can select the corresponding code table for quantization encoding.

[0390] The decoding process of the low-frequency part and the high-frequency part is as follows:

[0391] Step 21: quantization decoding.

[0392] Quantization decoding is the inverse process of quantization encoding. For the received code stream (including high-frequency code stream and low-frequency code stream), entropy decoding is first performed, and the estimated value F′ of the feature vector of the low-frequency code stream is obtained by looking up the quantization table. LB (n) and the estimated value F′ of the eigenvector of the high-frequency code stream HB (n).

[0393] Referring to FIG13E , the configuration in which the ultra-wideband coding mode and the rate mode bit is 10 is taken as an example for the following description:

[0394] First, parse the frame header of the bitstream encapsulation. If the coding mode bit is 1, ultra-wideband coding is used. If the rate mode bit is 10, the wideband portion is encoded using low-frequency code table -1. After ultra-wideband encoding, the bitstream encapsulation contains an 8-dimensional high-frequency feature vector index and a 16-dimensional first high-frequency feature vector residual index.

[0395] For the broadband part, the low-frequency code stream is parsed by entropy coding and decoding technology, and the estimated value F′ of the 56-dimensional low-frequency feature vector can be obtained through the low-frequency code table-1 LB (n).

[0396] For the ultra-wideband part, first, based on the 5-bit code table, an 8-dimensional high-frequency feature vector estimation value is obtained for each frame; then, based on the 2-bit code table, a 16-dimensional first high-frequency feature vector residual estimation value is obtained for each frame. Through the corresponding 8-dimensional high-frequency feature vector estimation value (i.e., the above-mentioned high-frequency feature estimation value) and the 16-dimensional first high-frequency feature vector residual estimation value (i.e., the above-mentioned first residual feature estimation value), the 16-dimensional high-frequency feature vector final estimation value (i.e., the above-mentioned high-frequency feature final estimation value) can be obtained. The i-th value in the high-frequency feature vector estimation value is added to the 2i-1-th value and the 2i-th value in the residual estimation value of the first high-frequency feature vector, respectively, to obtain the 2i-1-th value and the 2i-th value in the final estimation value of the high-frequency feature vector. For example, the high-frequency feature estimation value is F′ HB (n) = [e′1, e′2, …, e′8], the residual estimate of the first high-frequency eigenvector is F′ C (n) = [e′ 1C ,e′ 2C ,…,e′ 16C ], then the final estimated value of the high-frequency feature is F′ FHB (n) = [e′1 + e′ 1C ,e′1+e′ 2C ,…,e′8+e′ 16C ].

[0397] Referring to FIG13D , the configuration in which the ultra-wideband coding mode and the rate mode bit is 01 is taken as an example for the following description:

[0398] First, parse the frame header of the bitstream encapsulation. If the coding mode bit is 1, it means that ultra-wideband coding is used. If the rate mode bit is 01, it means that the wideband part uses the low-frequency code table -1 during encoding. After the ultra-wideband part is encoded, the bitstream encapsulation contains an 8-dimensional high-frequency feature vector index.

[0399] For the broadband part, the low-frequency code stream is parsed by entropy coding and decoding technology, and the estimated value F′ of the 56-dimensional low-frequency feature vector can be obtained through the low-frequency code table-4. LB (n).

[0400] For the ultra-wideband part, first obtain an 8-dimensional high-frequency feature vector estimate for each frame based on the 5-bit code table, and then copy the 8-dimensional high-frequency feature vector estimate to obtain a 16-dimensional high-frequency feature vector final estimate (i.e., the above-mentioned high-frequency feature final estimate). For example, the high-frequency feature estimate is F′ HB (n)=[e′1,e′2,…,e′8], then the final estimated value of the high-frequency feature is F′ FHB (n)=[e′1,e′1,…,e′8,e′8].

[0401] Here, the configuration of ultra-wideband coding mode and the bit rate mode bit being 11 is used as an example for the following description:

[0402] First, parse the frame header of the bitstream encapsulation. If the coding mode bit is 1, ultra-wideband coding is used. If the rate mode bit is 11, the wideband portion is encoded using low-frequency code table -1. After ultra-wideband encoding, the bitstream encapsulation includes an 8-dimensional high-frequency feature vector index, a 16-dimensional first high-frequency feature vector residual index, and a 16-dimensional second high-frequency feature vector residual index.

[0403] For the broadband part, the low-frequency code stream is parsed by entropy coding and decoding technology, and the estimated value F′ of the 56-dimensional low-frequency feature vector can be obtained through the low-frequency code table-1 LB (n).

[0404] For the ultra-wideband part, first, based on the 5-bit code table, an 8-dimensional high-frequency feature vector estimate is obtained for each frame; then, based on the 2-bit code table, a 16-dimensional first high-frequency feature vector residual estimate and a 16-dimensional second high-frequency feature vector residual estimate are obtained for each frame. The final estimate of the 16-dimensional high-frequency feature vector (i.e., the final estimate of the high-frequency feature) can be obtained through the corresponding 8-dimensional high-frequency feature vector estimate (i.e., the high-frequency feature estimate), the 16-dimensional first high-frequency feature vector residual estimate (i.e., the first residual feature estimate), and the 16-dimensional second high-frequency feature vector residual estimate (i.e., the second residual feature estimate). The ith value in the high-frequency eigenvector estimate is added to the 2i-1th value in the first high-frequency eigenvector residual estimate and the 2i-1th value in the second high-frequency eigenvector residual estimate to obtain the 2i-1th value in the high-frequency eigenvector final estimate. The ith value in the high-frequency eigenvector estimate is added to the 2i-th value in the first high-frequency eigenvector residual estimate and the 2i-th value in the second high-frequency eigenvector residual estimate to obtain the 2i-th value in the high-frequency eigenvector final estimate. For example, the high-frequency eigenvector estimate is F′ HB (n) = [e′1, e′2, …, e′8], the residual estimate of the first high-frequency eigenvector is F′ C (n) = [e′ 1C ,e′ 2C ,…,e′ 16C ], the second high frequency eigenvector residual estimate is F′ B (n) = [e′ 1B ,e′ 2B ,…,e′ 16B ], then the final estimated value of the high-frequency feature is F′ FHB (n) = [e′1 + e′ 1C +e′ 1B ,e′1+e′ 2C +e′ 2B ,…,e′8+e′ 16C +e′ 16B ].

[0405] Step 22: Estimated value F′ based on the eigenvector of the low-frequency code stream LB (n), call the third NN.

[0406] First, based on the estimated value F′ of the eigenvector of the low-frequency code stream LB (n), call the third NN shown in Figure 14 to generate the estimated value x′ of the low-frequency subband signal LB (n). The third NN is similar to the second NN, such as causal convolution. The post-processing structure is similar to the pre-processing structure in the second NN. The decoding block structure is symmetrical with the encoding block on the encoding side. The encoding block on the encoding side first performs dilated convolution and then pooling to complete downsampling. The decoding block on the decoding side first performs pooling to complete upsampling and then performs dilated convolution. The specific process of the third NN is as follows:

[0407] First, a causal convolution is called, which can transform the input tensor F′ into LB (n), a tensor expanded from 56×1 to 256×1.

[0408] Next, cascade four decoding blocks with different upsampling factors (Up_factor). Each decoding block contains a convolution layer, an upsampling module, and a residual block, wherein a convolution layer is used to halve the number of input channels; an upsampling module contains a specific Up_factor for completing upsampling; and a residual block includes five residual units (Residual Unit) based on hole convolution. The Up_factor of the four decoding blocks is set to 5, 4, 4, and 2, respectively. Therefore, the number of output channels of the four decoding blocks is set to 128, 64, 32, and 16, respectively. After processing by the four decoding blocks, the 256×1 tensor is converted into tensors of 128×5, 64×20, 32×80, and 16×160, respectively. Among them, the embodiment of the present application is not limited to the number of decoding blocks, and can be any positive integer such as 2, 3, 4, or 5.

[0409] Here, for the upsampling module containing a specific Up_factor, a Repeat operation can be used to complete the upsampling operation by repeatedly filling, thereby saving complexity.

[0410] The configuration of the five residual units based on dilated convolution at the decoder is similar to that of the residual units at the encoder, including but not limited to the internal structure of the residual units, convolution kernel size, and dilation rate. Table 2 shows the number of groups used for dilated convolution in the decoder block. The decoder block uses a larger number of groups, 2, to associate more input and output channels, improving speech reconstruction quality.

[0411] Table 2. Grouping configurations used by residual units in different decoding blocks

[0412] Then, the 16×160 tensor output by the cascaded decoding block is post-processed. For example, a Repeat operation with a factor of 2 is performed on the 16×160 tensor output by the cascaded decoding block to complete upsampling, and then a convolution operation is performed and an activation function is used for the PReLU operation to generate a 16×320 tensor.

[0413] Finally, a causal convolution is called to convert the input 16×320 tensor into a 1×320 tensor to reconstruct the low-frequency subband signal.

[0414] Step 23: Estimation of the eigenvector F′ of the high frequency sub-band signal HB (n) Perform high-frequency reconstruction.

[0415] Similar to the high-frequency analysis at the encoding end, the high-frequency reconstruction in the embodiment of the present application includes two solutions.

[0416] The first implementation of high frequency reconstruction corresponds to the first implementation of high frequency analysis in the encoding end. The estimated value F′ based on the eigenvector of the high frequency subband signal HB (n), call the neural network to generate the estimated value x′ of the high-frequency subband signal HB (n).

[0417] The second implementation of high-frequency reconstruction corresponds to the band extension technique used in high-frequency analysis at the encoder end. Based on the 16 MDCT subband spectral envelopes decoded from the high-frequency bitstream (each 10ms subframe contains 8 subband spectral envelopes), i.e., the final estimate of the high-frequency feature vector, the following operations are performed:

[0418] First, the estimated value x′ of the low-frequency subband signal generated by the decoding end LB (n), two 320-point MDCT transforms are also performed to generate two groups of 160-point MDCT coefficients (i.e., MDCT coefficients of the low-frequency part, including the first group of 160-point MDCT coefficients and the second group of 160-point MDCT coefficients).

[0419] Then, x′ LB The 160-point MDCT coefficients generated by (n) are replicated to generate the MDCT coefficients for the high-frequency portion. Referring to the basic characteristics of speech signals, low-frequency portions have more harmonics, while high-frequency portions have fewer. Therefore, to avoid the artificially generated high-frequency MDCT spectrum containing too many harmonics due to simple replication, the last 80 points of the 160-point MDCT coefficients of the low-frequency subband can be used as a master. The spectrum is replicated twice to generate reference values ​​for the 160-point MDCT coefficients of the high-frequency subband signal.

[0420] It should be noted that if the code stream encapsulation contains flatness edge information, additional operations are required. When Sfp_FLAG == 0, no operation is performed. When Sfp_FLAG == 1, additional flattening operations are required. The flattening operation is implemented as follows:

[0421] Taking the MDCT coefficient c(i) at position i as an example, calculate the average power spectrum e of the neighborhood of c(i) ave (i), as shown in formula (3).

[0422] Here, c(i-1) represents the MDCT coefficient at the i-1th position, and c(i+1) represents the MDCT coefficient at the i+1th position.

[0423] When e ave When (i)==0, the MDCT coefficient c(i) at position i is not smoothed, i.e. c(i)=c(i); otherwise, the MDCT coefficient c(i) at position i is smoothed, i.e. c(i) is updated to the value of c(i) and e(i). ave (i), that is, c(i) = c(i) / e ave (i).

[0424] After the flattening operation, the 8 sub-band spectral envelopes obtained previously (i.e., the 8 sub-band spectral envelopes obtained by querying the quantization table) are called. These 8 sub-band spectral envelopes correspond to 8 high-frequency sub-bands, and the reference values ​​of the MDCT coefficients of the generated 160-point high-frequency sub-band signals are divided into 8 reference high-frequency sub-bands. Based on a high-frequency sub-band and the corresponding reference high-frequency sub-band, the reference values ​​of the MDCT coefficients of the generated 160-point high-frequency sub-band signals are gained (multiplication in the frequency domain). For example, the gain factor is calculated based on the average energy of the high-frequency sub-band and the average energy of the corresponding reference high-frequency sub-band, and the MDCT coefficient corresponding to each point in the corresponding reference high-frequency sub-band is multiplied by the gain factor to ensure that the energy of the virtually generated high-frequency MDCT coefficients by decoding is close to the original coefficient energy at the encoding end.

[0425] For example, assuming that the average energy of the reference high-frequency sub-band (i.e., the sub-band divided by the reference values ​​of the 160-point MDCT coefficients of the generated high-frequency signal) is Y_L, and the average energy of the current high-frequency sub-band (i.e., the sub-band corresponding to the sub-band spectrum envelope decoded based on the bitstream) is Y_H, then a gain factor a = sqrt(Y_H / Y_L) is calculated, where sqrt() represents the square root calculation function, which is used to calculate the square root of (Y_H / Y_L). With the gain factor a, the MDCT coefficients of each point in the reference high-frequency sub-band are directly multiplied by a. The average energy of the MDCT coefficients (virtually generated) after gain is very close to the original energy at the encoding end.

[0426] Finally, the inverse MDCT transform is called to generate the estimated value of the first subframe of the high-frequency subband signal (the estimated value of the subframe calculated by the MDCT coefficients of the first group of 160 points) and the estimated value of the second subframe of the high-frequency subband signal (the estimated value of the subframe calculated by the MDCT coefficients of the second group of 160 points). The estimated value of the high-frequency subband signal x′ is obtained by combining the estimated value of the first subframe and the estimated value of the second subframe. HB (n) Perform inverse MDCT transform on the 320-point MDCT coefficients after gain to generate 640-point estimated values. By overlapping, take the first 320 effective estimated values ​​as x′ HB (n).

[0427] Step 24: Synthesize the filter

[0428] The estimated value x′ of the low-frequency subband signal is obtained at the decoding end LB (n) and the estimated value x′ of the high frequency subband signal HB (n), it is only necessary to upsample and call the QMF synthesis filter to generate the 640-point reconstructed signal x′(n).

[0429] The embodiment of the present application can collect data and jointly train the relevant networks of the encoding and decoding ends to obtain optimal parameters. The user only needs to prepare the data and set the corresponding network structure. After the training is completed in the background, the trained model can be put into use.

[0430] In a voice communication system, communication modes include monophonic voice communication and stereophonic communication. Therefore, the embodiments of the present application are not limited to the above-mentioned monophonic voice communication codec technology, but can also be stereo codec technology. Among them, stereo codec technology can enable the encoder to receive signals from both left and right channels (i.e., input signals for the left channel and input signals for the right channel), so that when the user wears headphones to make a call, a certain left and right spatial sense can be achieved.

[0431] The audio encoding method and audio decoding method provided in the embodiments of the present application will be described below in conjunction with a stereo signal.

[0432] It should be noted that the stereo encoding method of the embodiment of the present application includes but is not limited to the following two methods: 1) the left and right channels are respectively encoded and decoded using the above-mentioned mono signal encoding and decoding technology, and the stereo signal is restored at the decoding end by combining the decoded single-channel reconstructed signal; 2) parametric stereo technology is used for encoding and decoding.

[0433] Among them, parametric stereo coding technology is a technology that effectively encodes a stereo signal by treating it as a mono signal plus a small amount of parameter overhead describing the stereo image information. At the encoding end, the input signals of the left and right channels (i.e., stereo signals) are downmixed into one signal, and the one signal is regarded as the above-mentioned input signal x(n), and then encoded using the above-mentioned mono signal encoding method of the embodiment of the present application. In addition, the characteristic information of the correlation between the left and right channels (i.e., stereo feature vector, generally 1-2kbps) is extracted by parametric stereo coding, and the stereo feature vector is used to mix the side information of the mono channel. At the decoding end, a downmix signal (i.e., the above-mentioned reconstructed signal) is obtained by decoding, and the stereo signals of the left and right channels are reconstructed by the decoded downmix signal and the stereo feature vector describing the correlation between the left and right channels.

[0434] Following the above embodiment, for the 8-bit frame header in the code stream encapsulation, 1-bit channel bit can be allocated, and the 1-bit stereo coding bit can be used to determine whether mono coding or stereo coding is used. The specific implementation is as follows:

[0435] On the encoding side, if the number of input channels is 1, the channel flag is set to 0, and the channel bit of the corresponding frame header is written to 0, indicating that the input signal is mono encoded, that is, the encoding method for the input signal x(n) is reused for encoding; if the number of input channels is 2, that is, stereo, the channel flag is set to 1, and the channel bit of the corresponding frame header is written to 1, indicating that the input signal is stereo encoded, that is, the parametric stereo encoding method introduced above is reused for encoding.

[0436] On the decoding side, the channel bit corresponding to the frame header is parsed. If the channel bit is 0, it indicates mono decoding, and the above decoding method is reused to decode the reconstructed signal; if the channel bit is 1, it indicates stereo decoding, and the parametric stereo decoding method introduced above is reused for decoding.

[0437] The code stream encapsulation shown in Figures 15A-15E carries a stereo feature vector on the basis of the code stream encapsulation shown in Figures 13A-13E above. At the decoding end, the stereo signals of the left and right channels are reconstructed through the decoded downmix signal and the stereo feature vector describing the correlation between the left and right channels.

[0438] In addition, for the 8-bit frame header in the code stream encapsulation, the channel bit may not be allocated. For example, if there is no channel bit, mono encoding is used by default.

[0439] Please note that the frame header in each embodiment of the present application is not limited to 8 bits, and can be more or less bits. In another embodiment, more bits can be allocated to the channel bit, such as 2 bits of channel bit, to represent the codec of more channels, such as 3.1 channels or 4.1 channels, etc.

[0440] In summary, the multi-mode multi-rate encoding and decoding method provided in the embodiment of the present application significantly improves the audio quality while ensuring acceptable complexity compared to signal processing solutions through the organic combination of signal decomposition, signal processing technology and deep neural network.

[0441] The audio encoding method or audio decoding method provided by the embodiments of the present application has been described so far in conjunction with the exemplary application and implementation of the terminal device provided by the embodiments of the present application. The embodiments of the present application also provide a method for processing a bit stream, where the bit stream is decoded based on the above-mentioned audio decoding method or generated according to the above-mentioned audio encoding method.

[0442] So far, the audio encoding method or audio decoding method provided by the embodiment of the present application has been described in combination with the exemplary application and implementation of the terminal device provided by the embodiment of the present application. The embodiment of the present application also provides an audio encoding device and an audio decoding device. In actual applications, the functional modules in the audio encoding device and the audio decoding device can be implemented by the hardware resources of the electronic device (such as terminal equipment, server or server cluster), such as computing resources such as processors, communication resources (such as for supporting various communication modes such as optical cables and cellular), and memories. Figure 3A shows an audio encoding device 555 stored in the memory 550 and Figure 3B shows an audio decoding device 655 stored in the memory 650, which can be software in the form of programs and plug-ins, for example, software modules designed in programming languages ​​such as C / C++ and Java, application software designed in programming languages ​​such as C / C++ and Java, or dedicated software modules in large software systems, application program interfaces, plug-ins, cloud services, etc. The following examples illustrate different implementation methods.

[0443] The audio encoding device 555 includes a series of modules, including a feature extraction module 5551, an encoding module 5552, and a signal encoding module 5553. The following further describes how the various modules in the audio encoding device 555 provided in the embodiment of the present application cooperate to implement audio encoding.

[0444] The second acquisition module 5551 is configured to, in response to an encoding request for an audio signal, obtain a target encoding mode for the audio signal from multiple encoding modes and obtain a target bit rate mode for the audio signal from multiple bit rate modes; the extraction module 5552 is configured to extract the encoding features of the audio signal from the audio signal through the target encoding mode; the signal encoding module 5553 is configured to perform signal encoding on the encoding features through the target bit rate mode to obtain an audio code stream of the audio signal; the construction module 5554 is configured to determine a frame header based on the target encoding mode and the target bit rate mode; the generation module 5555 is configured to generate an audio code stream encapsulation of the audio signal based on the audio code stream and the frame header.

[0445] In some embodiments, when the target coding mode is a wideband coding mode, the extraction module 5552 is further configured to call a first neural network based on the wideband coding mode; and extract the coding features of the audio signal from the audio signal through the first neural network.

[0446] In some embodiments, the extraction module 5552 is further configured to perform feature extraction on the audio signal to obtain audio features of the audio signal; and perform residual processing on the audio features using at least one residual unit included in the first neural network to obtain coding features of the audio signal.

[0447] In some embodiments, the first neural network includes 4 encoding blocks, and each encoding block includes 4 or 5 residual units.

[0448] In some embodiments, the target bit rate mode is used to indicate the use of a target code table to perform signal encoding on the coding features of the audio signal; the signal encoding module 5553 is also configured to quantize the coding features of the audio signal through the target code table to obtain a quantized value of the coding feature; and entropy encode the quantized value through the target code rate corresponding to the target code table to obtain an audio code stream of the audio signal.

[0449] In some embodiments, when the target coding mode is an ultra-wideband coding mode, the extraction module 5552 is further configured to perform sub-band decomposition on the audio signal to obtain a low-frequency sub-band signal and a high-frequency sub-band signal of the audio signal; extract the low-frequency features of the low-frequency sub-band signal from the low-frequency sub-band signal through a second neural network; perform high-frequency analysis processing on the high-frequency sub-band signal to obtain high-frequency features of the high-frequency sub-band signal; and determine the low-frequency features and the high-frequency features as coding features of the audio signal.

[0450] In some embodiments, the extraction module 5552 is further configured to frame the high-frequency sub-band signal to obtain multiple sub-frames of the high-frequency sub-band signal; perform frequency band expansion processing on each of the sub-frames to obtain the sub-band spectrum envelope of each of the sub-frames; and use the sub-band spectrum envelopes corresponding to the multiple sub-frames as the high-frequency features of the high-frequency sub-band signal.

[0451] In some embodiments, the extraction module 5552 is further configured to perform a frequency domain transform based on the multiple sample points included in the subframe to obtain the transform coefficients corresponding to the multiple sample points respectively; divide the transform coefficients corresponding to the multiple sample points respectively into multiple sub-bands; average the transform coefficients included in each of the sub-bands to obtain the average energy corresponding to each of the sub-bands, and use the average energy as the sub-band spectrum envelope corresponding to each of the sub-bands.

[0452] In some embodiments, the target bit rate mode includes an instruction to use a first code table to perform signal encoding on the low-frequency feature, and to use a second code table to perform signal encoding on the high-frequency feature; the signal encoding module 5553 is further configured to quantize the low-frequency feature through the first code table to obtain the quantized value of the low-frequency feature, and entropy encode the quantized value of the low-frequency feature through the code rate corresponding to the first code table to obtain the low-frequency code stream of the low-frequency sub-band signal; quantize the high-frequency feature through the second code table to obtain the quantized value of the high-frequency feature, and entropy encode the quantized value of the high-frequency feature through the code rate corresponding to the second code table to obtain the high-frequency code stream of the high-frequency sub-band signal; and construct the audio code stream of the audio signal based on the low-frequency code stream and the high-frequency code stream.

[0453] In some embodiments, the target bit rate mode also includes an instruction to use a third code table to perform signal encoding on the residual feature of the high-frequency feature; the signal encoding module 5553 is further configured to determine the sub-band corresponding to the sub-frame of the high-frequency sub-band signal, and divide the sub-band into multiple subsets; average the transform coefficients included in each of the subsets to obtain the average energy corresponding to each of the subsets, and use the average energy as the envelope value corresponding to each of the subsets; determine the quantized value of the sub-band spectrum envelope of the sub-band corresponding to the subset in the quantized value of the high-frequency feature, and use the quantized value of the sub-band spectrum envelope of the sub-band corresponding to the subset The difference between the envelope value and the quantized value of the sub-band spectrum envelope is used as the first residual value; based on the first residual value, the residual feature of the high-frequency feature is determined; the residual feature of the high-frequency feature is quantized through the third code table to obtain the quantized value of the residual feature of the high-frequency feature, and the quantized value of the residual feature of the high-frequency feature is entropy encoded through the code rate corresponding to the third code table to obtain the residual code stream of the residual feature; the construction module 5554 is also configured to use the low-frequency code stream, the high-frequency code stream and the residual code stream as the audio code stream of the audio signal.

[0454] In some embodiments, when the number of residual values ​​is N, N is a positive integer greater than 1, and the signal encoding module 5553 is further configured to determine the quantization value of the nth residual value, and determine the sum of the quantization value of the nth residual value and the quantization value of the subband spectrum envelope of the corresponding subband; the difference between the envelope value corresponding to the subset and the sum is used as the n+1th residual value; the N residual values ​​are determined as the residual features of the high-frequency features; wherein n is a positive integer that increases successively, 1≤n≤N.

[0455] In some embodiments, before generating the audio code stream encapsulation of the audio signal based on the audio code stream and the frame header, the signal encoding module 5553 is also configured to determine the flatness side information of the high-frequency sub-band signal; the generation module 5555 is also configured to combine the audio code stream, the frame header and the flatness side information to obtain the audio code stream encapsulation of the audio signal.

[0456] In some embodiments, the signal encoding module 5553 is further configured to divide the transform coefficients included in the subframe of the high-frequency subband signal into multiple blocks; determine a first flatness of each of the blocks, and determine a second flatness of the low-frequency specified frequency band of the audio signal; when the first flatness is less than the second flatness, or the first flatness is less than a flatness threshold, set the flatness side information to a first value, wherein the first value indicates that flattening processing is required during audio decoding; when the first flatness is greater than or equal to the second flatness, and the first flatness is greater than or equal to the flatness threshold, set the flatness side information to a second value, wherein the second value indicates that flattening processing is not required during audio decoding.

[0457] In some embodiments, the frame header further includes at least one channel bit, where the channel bit is used to indicate whether mono encoding or stereo encoding is used to encode the audio signal.

[0458] The audio decoding device 655 includes a series of modules, including a first acquisition module 6551, a signal decoding module 6552, and a reconstruction module 6553. The following further describes the solution of implementing audio encoding by cooperating with each module in the audio encoding device 555 provided in the embodiment of the present application.

[0459] The first acquisition module 6551 is configured to obtain an audio code stream encapsulation, wherein the audio code stream encapsulation includes an audio code stream, and the audio code stream is obtained by audio encoding an audio signal through a target coding mode and a target bit rate mode, wherein the target coding mode is obtained from multiple candidate coding modes, and the target bit rate mode is obtained from multiple candidate bit rate modes; in response to a decoding request for the audio code stream encapsulation, the target coding mode and the target bit rate mode are obtained from a frame header included in the audio code stream encapsulation; the signal decoding module 6552 is configured to perform signal decoding on the audio code stream through the target coding mode and the target bit rate mode to obtain a coding feature estimation value corresponding to the audio code stream; and the reconstruction module 6553 is configured to reconstruct the coding feature estimation value through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

[0460] In some embodiments, when the target coding mode is a wideband coding mode, the target bit rate mode is used to indicate the use of a target code table to perform signal encoding on the coding features of the audio signal; the signal decoding module 6552 is also configured to perform entropy decoding on the audio code stream using the target bit rate corresponding to the target code table to obtain a quantization value corresponding to the audio code stream; and perform inverse quantization processing on the quantization value corresponding to the audio code stream using the target code table to obtain an estimated coding feature value corresponding to the audio code stream.

[0461] In some embodiments, when the target coding mode is a wideband coding mode, the reconstruction module 6553 is further configured to call a third neural network based on the wideband coding mode; through the third neural network, the coding feature estimation value corresponding to the audio code stream is reconstructed to obtain a reconstructed audio signal corresponding to the audio code stream.

[0462] In some embodiments, the reconstruction module 6553 is further configured to perform residual processing on the coding feature estimation value corresponding to the audio code stream using at least one residual unit included in the third neural network to obtain the audio feature estimation value corresponding to the audio code stream; and perform feature reconstruction on the audio feature estimation value corresponding to the audio code stream to obtain a reconstructed audio signal corresponding to the audio code stream.

[0463] In some embodiments, the third neural network includes 4 decoding blocks, each of which includes 4 or 5 residual units.

[0464] In some embodiments, when the target coding mode is an ultra-wideband coding mode, the target bit rate mode includes an instruction to use a first code table to encode the low-frequency features of the audio signal, and to use a second code table to encode the high-frequency features of the audio signal, and the audio code stream includes a low-frequency code stream and a high-frequency code stream. The signal decoding module 6552 is further configured to obtain the low-frequency code stream and the high-frequency code stream from the audio code stream according to the target coding mode; perform entropy decoding on the low-frequency code stream according to the code rate corresponding to the first code table to obtain a quantization value corresponding to the low-frequency code stream, and perform inverse quantization processing on the quantization value corresponding to the low-frequency code stream according to the first code table to obtain a low-frequency feature estimation value corresponding to the low-frequency code stream; perform entropy decoding on the high-frequency code stream according to the code rate corresponding to the second code table to obtain a quantization value corresponding to the high-frequency code stream, and perform inverse quantization processing on the quantization value corresponding to the high-frequency code stream according to the second code table to obtain a high-frequency feature estimation value corresponding to the high-frequency code stream; and determine a coding feature estimation value corresponding to the audio code stream based on the low-frequency feature estimation value and the high-frequency feature estimation value.

[0465] In some embodiments, the audio code stream also includes a residual code stream, and the target bit rate mode also includes an instruction to use a third code table to perform signal encoding on the residual features of the high-frequency features corresponding to the audio code stream; the signal decoding module 6552 is further configured to perform entropy decoding on the residual code stream using the code rate corresponding to the third code table to obtain a quantization value corresponding to the residual code stream, and perform inverse quantization processing on the quantization value corresponding to the residual code stream using the third code table to obtain a residual feature estimation value corresponding to the residual code stream; the high-frequency feature estimation value and the residual feature estimation value are summed to determine the final estimation value of the high-frequency feature; the low-frequency feature estimation value and the final estimation value of the high-frequency feature are determined as the coding feature estimation value corresponding to the audio code stream.

[0466] In some embodiments, the reconstruction module 6553 is further configured to, when the target coding mode is ultra-wideband coding, perform feature reconstruction on the low-frequency feature estimation value included in the coding feature estimation value through a fourth neural network to obtain a low-frequency sub-band signal estimation value corresponding to the audio code stream; perform high-frequency reconstruction on the high-frequency feature estimation value included in the coding feature estimation value to obtain a high-frequency sub-band signal estimation value corresponding to the audio code stream; perform sub-band synthesis on the low-frequency sub-band signal estimation value and the high-frequency sub-band signal estimation value to obtain a reconstructed audio signal corresponding to the audio code stream.

[0467] In some embodiments, the reconstruction module 6553 is further configured to perform a frequency domain transform on the sample points of the first half and the sample points of the second half included in the low-frequency subband signal estimation value to obtain a first transform coefficient corresponding to the sample points of the first half and a second transform coefficient corresponding to the sample points of the second half; based on the first transform coefficient, the high-frequency feature estimation value is subjected to inverse processing of frequency band expansion to obtain a first high-frequency subband signal estimation value; based on the second transform coefficient, the high-frequency feature estimation value is subjected to inverse processing of frequency band expansion to obtain a second high-frequency subband signal estimation value; and the first high-frequency subband signal estimation value and the second high-frequency subband signal estimation value are combined to obtain a high-frequency subband signal estimation value corresponding to the audio code stream.

[0468] In some embodiments, the reconstruction module 6553 is also configured to perform spectrum replication processing on the transform coefficients of the second half of the first transform coefficients to obtain first reference transform coefficients of the reference high-frequency sub-band signal; based on the first half of the sub-band spectrum envelope corresponding to the high-frequency feature estimation value, perform a gain on the first reference transform coefficient to obtain the gained first reference transform coefficient; perform an inverse frequency domain transform on the gained first reference transform coefficient to obtain a first high-frequency sub-band signal estimation value.

[0469] In some embodiments, the audio code stream encapsulation includes flatness side information; after the transform coefficients of the second half of the first transform coefficients are subjected to spectrum replication processing to obtain the first reference transform coefficients of the reference high-frequency sub-band signal, the reconstruction module 6553 is further configured to determine the i-1th first reference transform coefficient, the i-th first reference transform coefficient, and the i+1th first reference transform coefficient when the flatness side information represents that flattening processing is required during audio decoding; determine the average power spectrum based on the i-1th first reference transform coefficient, the i-th first reference transform coefficient, and the i+1th first reference transform coefficient; and use the ratio of the i-th first reference transform coefficient to the average power spectrum as the new i-th first reference transform coefficient; wherein, 1<i<I, i is a positive integer, and I is the number of the first reference transform coefficients.

[0470] In some embodiments, the frame header further includes at least one channel bit, and the channel bit is used to indicate whether a mono decoding method or a stereo decoding method is used to perform audio decoding on the audio code stream.

[0471] The present invention provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the audio encoding method or audio decoding method described in the present invention.

[0472] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the audio encoding method or audio decoding method provided in an embodiment of the present application, for example, the audio encoding method shown in Figure 4A.

[0473] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.

[0474] In some embodiments, computer executable instructions (executable instructions for short) may be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as standalone programs or as modules, components, subroutines or other units suitable for use in a computing environment.

[0475] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0476] As an example, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0477] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0478] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An audio decoding method, applied to an electronic device, the method comprising: Obtaining an audio code stream package, wherein the audio code stream package includes an audio code stream, and the audio code stream is obtained by audio encoding an audio signal through a target coding mode and a target bit rate mode, wherein the target coding mode is obtained from a plurality of candidate coding modes, and the target bit rate mode is obtained from a plurality of candidate bit rate modes; In response to a decoding request for the audio code stream package, acquiring the target coding mode and the target bit rate mode from a frame header included in the audio code stream package; By using the target coding mode and the target bit rate mode, signal decoding is performed on the audio bit stream to obtain a coding feature estimation value corresponding to the audio bit stream; The coding feature estimation value is reconstructed through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

2. The method according to claim 1, wherein: When the target coding mode is a wideband coding mode, the target bit rate mode is used to indicate that a target code table is used to perform signal coding on the coding feature of the audio signal; The step of decoding the audio bitstream by the target coding mode and the target bitrate mode to obtain a coding feature estimation value corresponding to the audio bitstream includes: Performing entropy decoding on the audio bit stream according to the target bit rate corresponding to the target code table to obtain a quantization value corresponding to the audio bit stream; The quantized value is inversely quantized using the target code table to obtain a coding feature estimation value corresponding to the audio code stream.

3. The method according to claim 1 or 2, wherein: When the target coding mode is a wideband coding mode, reconstructing the coding feature estimation value through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream includes: Invoking a third neural network based on the broadband coding mode; The coding feature estimation value is reconstructed through the third neural network to obtain a reconstructed audio signal corresponding to the audio code stream.

4. The method according to claim 3, wherein: The reconstructing the estimated value of the coding feature to obtain a reconstructed audio signal corresponding to the audio code stream includes: Using at least one residual unit included in the third neural network, performing residual processing on the coding feature estimate value to obtain an audio feature estimate value corresponding to the audio code stream; Feature reconstruction is performed on the audio feature estimation value to obtain a reconstructed audio signal corresponding to the audio code stream.

5. The method according to claim 4, wherein: The third neural network includes 4 decoding blocks, each of which includes 4 or 5 residual units.

6. The method according to any one of claims 1 to 5, wherein: When the target coding mode is an ultra-wideband coding mode, the target bit rate mode includes an instruction to use a first code table to encode a low-frequency feature of the audio signal, and use a second code table to encode a high-frequency feature of the audio signal, and the audio code stream includes a low-frequency code stream and a high-frequency code stream; The step of decoding the audio bitstream by the target coding mode and the target bitrate mode to obtain a coding feature estimation value corresponding to the audio bitstream includes: Obtaining the low-frequency code stream and the high-frequency code stream from the audio code stream through the target coding mode; Performing entropy decoding on the low-frequency code stream using the code rate corresponding to the first code table to obtain a quantization value corresponding to the low-frequency code stream, and performing inverse quantization processing on the quantization value corresponding to the low-frequency code stream using the first code table to obtain a low-frequency feature estimation value corresponding to the low-frequency code stream; The high frequency code stream is entropy decoded according to the code rate corresponding to the second code table to obtain the high frequency code stream The quantized value corresponding to the high frequency code stream is obtained by inverse quantization processing on the quantized value corresponding to the high frequency code stream through the second code table to obtain a high frequency feature estimation value corresponding to the high frequency code stream; Based on the low-frequency feature estimation value and the high-frequency feature estimation value, a coding feature estimation value corresponding to the audio bitstream is determined.

7. The method according to claim 6, wherein: The audio code stream further includes a residual code stream, and the target bit rate mode further includes an instruction to use a third code table to perform signal encoding on residual features of high-frequency features corresponding to the audio code stream; The method further comprises: entropy decoding the residual code stream using the code rate corresponding to the third code table to obtain a quantization value corresponding to the residual code stream, and inverse quantizing the quantization value corresponding to the residual code stream using the third code table to obtain a residual feature estimation value corresponding to the residual code stream; and / or, The determining, based on the low-frequency feature estimation value and the high-frequency feature estimation value, the coding feature estimation value corresponding to the audio bitstream includes: The sum of the high-frequency feature estimation value and the residual feature estimation value is determined as a final high-frequency feature estimation value; The low-frequency feature estimation value and the high-frequency feature final estimation value are determined as the coding feature estimation values ​​corresponding to the audio bitstream.

8. The method according to any one of claims 1 to 7, wherein: The step of reconstructing the coding feature estimation value through the target coding mode to obtain a reconstructed audio signal corresponding to the audio bitstream includes: When the target coding mode is ultra-wideband coding, a low-frequency feature estimation value included in the coding feature estimation value is subjected to feature reconstruction by a fourth neural network to obtain a low-frequency subband signal estimation value corresponding to the audio bitstream; Performing high-frequency reconstruction on the high-frequency feature estimation value included in the coding feature estimation value to obtain a high-frequency subband signal estimation value corresponding to the audio bit stream; Sub-band synthesis is performed on the low-frequency sub-band signal estimation value and the high-frequency sub-band signal estimation value to obtain a reconstructed audio signal corresponding to the audio code stream.

9. The method according to claim 8, wherein: The high-frequency reconstruction of the high-frequency feature estimation value included in the coding feature to obtain the high-frequency sub-band signal estimation value corresponding to the audio code stream includes: Performing frequency domain transformation on the sample points of the first half and the sample points of the second half included in the low-frequency subband signal estimation value to obtain first transformation coefficients corresponding to the sample points of the first half and second transformation coefficients corresponding to the sample points of the second half; Based on the first transform coefficient, performing inverse processing of frequency band expansion on the high-frequency feature estimation value to obtain a first high-frequency sub-band signal estimation value; Based on the second transform coefficient, performing inverse processing of frequency band expansion on the high-frequency feature estimation value to obtain a second high-frequency sub-band signal estimation value; The first high frequency sub-band signal estimation value and the second high frequency sub-band signal estimation value are combined to obtain a high frequency sub-band signal estimation value corresponding to the audio code stream.

10. The method according to claim 9, wherein: The step of performing inverse processing of frequency band expansion on the high frequency feature estimation value based on the first transform coefficient to obtain a first high frequency subband signal estimation value comprises: Performing spectrum replication on the transform coefficients of the second half of the first transform coefficients to obtain first reference transform coefficients of a reference high frequency subband signal; Based on the first half of the sub-band spectrum envelope corresponding to the high-frequency feature estimation value, the first reference transform coefficient is amplified to obtain the amplified first reference transform coefficient; An inverse frequency domain transform is performed on the first reference transform coefficient after the gain to obtain a first high frequency subband signal estimation value.

11. The method according to claim 10, wherein: The audio code stream encapsulation includes flatness edge information; After performing spectrum replication on the transform coefficients of the second half of the first transform coefficients to obtain first reference transform coefficients of the reference high frequency sub-band signal, the method further includes: When the flatness side information represents that flattening processing is required during audio decoding, determining an i-1th first reference transform coefficient, an i-th first reference transform coefficient, and an i+1th first reference transform coefficient; Determine an average power spectrum based on the (i-1)th first reference transform coefficient, the (i)th first reference transform coefficient, and the (i+1)th first reference transform coefficient; Using a ratio of the i-th first reference transform coefficient to the average power spectrum as a new i-th first reference transform coefficient; Among them, 1<i<I, i is a positive integer, and I is the number of the first reference transformation coefficients.

12. The method according to any one of claims 1 to 11, wherein: The frame header further includes at least one channel bit, and the channel bit is used to indicate whether a mono decoding method or a stereo decoding method is used to perform audio decoding on the audio code stream.

13. An audio encoding method, applied to an electronic device, the method comprising: In response to an encoding request for an audio signal, acquiring a target encoding mode for the audio signal from a plurality of encoding modes, and acquiring a target bit rate mode for the audio signal from a plurality of bit rate modes; Extracting the coding feature of the audio signal from the audio signal by using the target coding mode; By using the target bit rate mode, signal encoding is performed on the coding feature to obtain an audio bit stream of the audio signal; Determining a frame header based on the target coding mode and the target bit rate mode; An audio code stream encapsulation of the audio signal is generated based on the audio code stream and the frame header.

14. The method according to claim 13, wherein: When the target coding mode is a wideband coding mode, extracting the coding feature of the audio signal from the audio signal by using the target coding mode includes: Invoking a first neural network based on the broadband coding mode; The encoding features of the audio signal are extracted from the audio signal through the first neural network.

15. The method according to claim 14, wherein: The step of extracting the coding feature of the audio signal from the audio signal comprises: Extracting features from the audio signal to obtain audio features of the audio signal; The audio feature is subjected to residual processing by utilizing at least one residual unit included in the first neural network to obtain a coding feature of the audio signal.

16. The method according to claim 15, wherein: The first neural network includes 4 encoding blocks, and each of the encoding blocks includes 4 or 5 residual units.

17. The method according to any one of claims 13 to 16, wherein: The target bit rate mode is used to indicate that a target code table is used to perform signal encoding on the coding feature of the audio signal; The step of performing signal encoding processing on the encoding feature through the target bit rate mode to obtain an audio bit stream of the audio signal includes: quantizing the coding feature using the target code table to obtain a quantized value of the coding feature; The quantized value is entropy encoded according to the target bit rate corresponding to the target code table to obtain an audio bit stream of the audio signal.

18. The method according to any one of claims 13 to 17, wherein: When the target coding mode is an ultra-wideband coding mode, extracting the coding feature of the audio signal from the audio signal by using the target coding mode includes: Performing sub-band decomposition on the audio signal to obtain a low-frequency sub-band signal and a high-frequency sub-band signal of the audio signal; extracting low-frequency features of the low-frequency sub-band signal from the low-frequency sub-band signal through a second neural network; Performing high-frequency analysis on the high-frequency sub-band signal to obtain high-frequency features of the high-frequency sub-band signal; The low-frequency feature and the high-frequency feature are determined as coding features of the audio signal.

19. The method according to claim 18, wherein: The performing high frequency analysis on the high frequency sub-band signal to obtain the high frequency features of the high frequency sub-band signal includes: Dividing the high frequency sub-band signal into frames to obtain a plurality of sub-frames of the high frequency sub-band signal; Performing frequency band extension processing on each of the subframes to obtain a subband spectrum envelope of each of the subframes; Sub-band spectrum envelopes corresponding to the multiple sub-frames are used as high-frequency features of the high-frequency sub-band signal.

20. The method according to claim 19, wherein: The performing frequency band extension processing on each of the subframes to obtain a subband spectrum envelope of each of the subframes includes: Performing frequency domain transformation based on multiple sample points included in the subframe to obtain transformation coefficients corresponding to the multiple sample points respectively; Dividing the transform coefficients corresponding to the plurality of sample points into a plurality of sub-bands; The transform coefficients included in each of the sub-bands are averaged to obtain average energy corresponding to each of the sub-bands, and the average energy is used as the sub-band spectrum envelope corresponding to each of the sub-bands.

21. The method according to claim 18, wherein: The target bit rate mode includes an instruction to use a first code table to encode the low-frequency feature and a second code table to encode the high-frequency feature; The step of encoding the encoding feature by the target bit rate mode to obtain an audio bit stream of the audio signal includes: quantizing the low-frequency feature using the first code table to obtain a quantized value of the low-frequency feature, and entropy encoding the quantized value of the low-frequency feature using a code rate corresponding to the first code table to obtain a low-frequency code stream of the low-frequency subband signal; quantizing the high-frequency feature using the second code table to obtain a quantized value of the high-frequency feature, and entropy encoding the quantized value of the high-frequency feature using a code rate corresponding to the second code table to obtain a high-frequency code stream of the high-frequency subband signal; An audio code stream of the audio signal is constructed based on the low-frequency code stream and the high-frequency code stream.

22. The method according to claim 21, wherein: The target bit rate mode further includes an instruction to use a third code table to perform signal encoding on the residual feature of the high-frequency feature; The method further comprises: Determine a subband corresponding to a subframe of the high frequency subband signal, and divide the subband into a plurality of subsets; Performing averaging processing on the transform coefficients included in each of the subsets to obtain average energy corresponding to each of the subsets, and using the average energy as the envelope value corresponding to each of the subsets; Determine the quantized value of the subband spectrum envelope of the subband corresponding to the subset in the quantized value of the high-frequency feature, and use the difference between the envelope value corresponding to the subset and the quantized value of the subband spectrum envelope as the first residual value; Determining a residual feature of the high-frequency feature based on the first residual value; quantizing the residual features of the high-frequency features through the third code table to obtain quantized values ​​of the residual features of the high-frequency features, and entropy encoding the quantized values ​​of the residual features of the high-frequency features through the code rate corresponding to the third code table to obtain a residual code stream of the residual features; and / or The step of constructing the audio code stream of the audio signal based on the low-frequency code stream and the high-frequency code stream includes: The low-frequency code stream, the high-frequency code stream and the residual code stream are used as the audio code stream of the audio signal.

23. The method according to claim 22, wherein: When the number of the residual values ​​is N, N is a positive integer greater than 1, the determining the residual feature of the high-frequency feature based on the first residual value includes: Determine a quantized value of an nth residual value, and determine a sum of the quantized value of the nth residual value and a quantized value of a sub-band spectral envelope of a corresponding sub-band; Taking the difference between the envelope value corresponding to the subset and the sum as the (n+1)th residual value; Determine the N residual values ​​as residual features of the high-frequency features; Wherein, n is a positive integer which increases successively, 1≤n≤N.

24. The method according to any one of claims 21 to 23, wherein: Before generating the audio code stream encapsulation of the audio signal based on the audio code stream and the frame header, the method further includes: Determining flatness edge information of the high frequency sub-band signal; The step of generating an audio code stream encapsulation of the audio signal based on the audio code stream and the frame header includes: The audio code stream, the frame header and the flatness edge information are combined to obtain an audio code stream encapsulation of the audio signal.

25. The method according to claim 24, wherein: The step of determining the flatness edge information of the high frequency sub-band signal comprises: Dividing the transform coefficients included in the subframe of the high frequency subband signal into a plurality of blocks; determining a first flatness of each of the blocks, and determining a second flatness of a low-frequency designated frequency band of the audio signal; When the first flatness is less than the second flatness, or the first flatness is less than a flatness threshold, setting the flatness edge information to a first value, wherein the first value indicates that flattening processing is required during audio decoding; When the first flatness is greater than or equal to the second flatness, and the first flatness is greater than or equal to the flatness threshold, the flatness edge information is set to a second value, wherein the second value indicates that flattening processing is not required during audio decoding.

26. The method according to any one of claims 13 to 25, wherein: The frame header further includes at least one channel bit, and the channel bit is used to indicate whether a mono encoding method or a stereo encoding method is used to perform audio encoding on the audio signal.

27. A method for processing a bit stream, wherein the bit stream is decoded based on the audio decoding method according to any one of claims 1 to 12, or is generated according to the audio encoding method according to any one of claims 13 to 26.

28. An audio decoding device, the device comprising: A first acquisition module is configured to acquire an audio code stream package, wherein the audio code stream package includes an audio code stream, and the audio code stream is obtained by audio encoding an audio signal through a target coding mode and a target bit rate mode, wherein the target coding mode is acquired from a plurality of candidate coding modes, and the target bit rate mode is acquired from a plurality of candidate bit rate modes; In response to a decoding request for the audio code stream package, acquiring the target coding mode and the target bit rate mode from a frame header included in the audio code stream package; A signal decoding module, configured to perform signal decoding on the audio bitstream according to the target coding mode and the target bitrate mode, and obtain a coding feature estimation value corresponding to the audio bitstream; The reconstruction module is configured to reconstruct the coding feature estimation value through the target coding mode to obtain a reconstructed audio signal corresponding to the audio code stream.

29. An audio encoding device, the device comprising: a second acquisition module, configured to, in response to an encoding request for an audio signal, acquire a target encoding mode for the audio signal from a plurality of encoding modes, and acquire a target bit rate mode for the audio signal from a plurality of bit rate modes; an extraction module, configured to extract the coding features of the audio signal from the audio signal through the target coding mode; A signal encoding module, configured to perform signal encoding on the encoding feature through the target bit rate mode to obtain an audio bit stream of the audio signal; A construction module configured to determine a frame header based on the target coding mode and the target bit rate mode; The generating module is configured to generate an audio code stream encapsulation of the audio signal based on the audio code stream and the frame header.

30. An electronic device, comprising: A memory for storing computer executable instructions; A processor, used to implement the audio encoding method described in any one of claims 1 to 12, or the audio decoding method described in any one of claims 13 to 26 when executing the computer executable instructions stored in the memory.

31. A computer-readable storage medium storing computer-executable instructions, which when executed by a processor implements the audio decoding method described in any one of claims 1 to 12, or the audio encoding method described in any one of claims 13 to 26.

Citation Information

Patent Citations

  • Stereo encoding and decoding method, a coder-decoder and encoding and decoding system

    CN101572088A

  • Variable-bit-rate encoder, variable-bit-rate decoder, variable-bit-rate encoding method and variable-bit-rate decoding method based on AMR (adaptive multi-rate)-NB (narrow band) voice signals

    CN104517612A

  • Audio coding and decoding method and device, equipment and storage medium

    CN114694664A

  • Audio coding method and device, equipment, storage medium and program product

    CN115116454A

  • Audio coding method, audio decoding method, audio coding device, audio decoding device and readable storage medium

    CN117476024A

Cited By

  • Voice coding and decoding method, device, equipment and medium

    CN121054006A

  • Audio encoding method and apparatus, audio decoding method and apparatus, and readable storage medium

    EP4679420A4