Multi-channel audio signal encoding and decoding method and device

By generating energy/amplitude balanced side information for channel pairs, the number of multi-channel audio encoding bits is reduced, saving bits for other functional modules, thereby improving the quality of the audio signal at the decoding end and the coding efficiency.

CN113948096BActive Publication Date: 2025-10-03HUAWEI TECH CO LTD

Patent Information

Application Number
CN202010699711.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-17
Publication Date
2025-10-03
Estimated Expiration
2040-07-17

AI Technical Summary

Technical Problem

How to reduce the number of bits required to encode multi-channel audio to improve the quality of the reconstructed signal at the decoding end.

Method used

By generating energy/amplitude balanced side information for channel pairs, the encoded bitstream carries the energy/amplitude balanced side information for K channel pairs, but does not carry the energy/amplitude balanced side information for unpaired channels. This reduces the number of bits for the energy/amplitude balanced side information in the encoded bitstream, saving bits that can be allocated to other functional modules of the encoder.

Benefits of technology

The quality of the audio signal reconstructed at the decoding end is improved, and the coding efficiency and coding effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113948096B_ABST
    Figure CN113948096B_ABST
Patent Text Reader

Abstract

The present application provides a multi-channel audio signal encoding and decoding method and apparatus. Embodiments of the present application can reduce the number of bits of multi-channel side information, thereby allocating the saved bits to other functional modules of the encoder to improve the quality of the audio signal reconstructed at the decoding end and improve the encoding quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to audio coding and decoding technology, and in particular to a method and device for coding and decoding multi-channel audio signals. Background Art

[0002] With the continuous development of multimedia technology, audio has been widely used in multimedia communications, consumer electronics, virtual reality, human-computer interaction, and other fields. Audio coding is one of the key technologies in multimedia technology. Audio coding removes redundant information from the original audio signal to compress the data volume for easier storage or transmission.

[0003] Multi-channel audio coding is the encoding of more than two channels. Common examples include 5.1-channel, 7.1-channel, 7.1.4-channel, and 22.2-channel. Multi-channel original audio signals are filtered, paired, stereo-processed, multi-channel side information generated, quantized, entropy-coded, and multiplexed to form a serial bit stream for easy transmission over channels or storage on digital media.

[0004] Among them, how to reduce the coding bits of multi-channel side information to improve the quality of the reconstructed signal at the decoding end has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The present application provides a multi-channel audio signal encoding and decoding method and apparatus, which are beneficial to improving the quality of encoded and decoded audio signals.

[0006] In a first aspect, an embodiment of the present application provides a multi-channel audio signal encoding method, which may include: obtaining audio signals of P channels of a current frame of a multi-channel audio signal, where P is a positive integer greater than 1, the P channels include K channel pairs, each channel pair includes two channels, K is a positive integer, and P is greater than or equal to K*2. Obtaining the energy / amplitude of each audio signal of the P channels. Based on the energy / amplitude of each audio signal of the P channels, generating energy / amplitude balanced side information of the K channel pairs. Encoding the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain a coded bitstream.

[0007] This implementation generates energy / amplitude balanced side information for channel pairs. The coded bitstream carries the energy / amplitude balanced side information for K channel pairs, but does not carry the energy / amplitude balanced side information for unpaired channels. This reduces the number of bits of energy / amplitude balanced side information in the coded bitstream, and reduces the number of bits of multi-channel side information. The saved bits can be allocated to other functional modules of the encoder to improve the quality of the audio signal reconstructed at the decoder and enhance the encoding quality.

[0008] For example, the saved bits can be used to encode multi-channel audio signals to reduce the compression rate of the data portion and improve the quality of the audio signal reconstructed at the decoding end.

[0009] In other words, the coded bitstream includes a control information portion and a data portion. The control information portion may include the aforementioned energy / amplitude balancing side information, and the data portion may include the aforementioned multi-channel audio signal. That is, the coded bitstream includes the multi-channel audio signal and the control information generated during the encoding process of the multi-channel audio signal. Embodiments of the present application can improve the quality of the audio signal reconstructed at the decoding end by reducing the number of bits occupied by the control information portion and increasing the number of bits occupied by the data portion.

[0010] It should be noted that the saved bits can also be used for other control information transmission, and the embodiments of the present application are not limited to the above examples.

[0011] In one possible design, the K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, the fixed-point energy / amplitude scaling ratio being a fixed-point value of an energy / amplitude scaling coefficient, the energy / amplitude scaling coefficient being obtained based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance, the energy / amplitude scaling identifier being used to identify whether the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

[0012] This implementation enables the decoding end to perform energy de-balancing through the fixed-point energy / amplitude scaling ratio and the energy / amplitude scaling identifier of the current channel pair to obtain a decoded signal.

[0013] By converting the floating-point energy / amplitude scaling coefficients into fixed-point energy / amplitude scaling coefficients, the bits occupied by the energy / amplitude equalization side information can be saved, thereby improving transmission efficiency.

[0014] In one possible design, the K channel pairs include the current channel pair, and generating energy / amplitude-balanced side information for the K channel pairs based on the energies / amplitudes of the audio signals of the P channels may include: determining the energies / amplitudes of the audio signals of the two channels of the current channel pair after energy / amplitude balance based on the energies / amplitudes of the audio signals of the two channels of the current channel pair before energy / amplitude balance. Generating the energy / amplitude-balanced side information for the current channel pair based on the energies / amplitudes of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energies / amplitudes of the audio signals of the two channels after energy / amplitude balance.

[0015] This implementation method, by performing energy / amplitude balancing on the two channels within a channel pair, can achieve that even for channel pairs with large energy differences, a large energy difference can still be maintained after energy / amplitude balancing, thereby meeting the encoding requirements of the channel pairs with large energy / amplitude in subsequent encoding processing, improving encoding efficiency and encoding effect, and thus improving the quality of the audio signal reconstructed at the decoding end.

[0016] In one possible design, the current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio of the first channel, a fixed-point energy / amplitude scaling ratio of the second channel, an energy / amplitude scaling identifier of the first channel, and an energy / amplitude scaling identifier of the second channel.

[0017] This implementation method uses fixed-point energy / amplitude scaling ratios and energy / amplitude scaling identifiers for each of the two channels in the current channel pair to enable the decoder to perform energy de-balancing, thereby further reducing the bits occupied by the energy / amplitude balanced side information of the current channel pair on the basis of the decoded signal.

[0018] In one possible design, generating side information of energy / amplitude balance of the current channel pair based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance can include: determining the energy / amplitude scaling coefficient of the qth channel and the energy / amplitude scaling identifier of the qth channel based on the energy / amplitude of the audio signal of the qth channel of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signal of the qth channel after energy / amplitude balance. Determining the fixed-point energy / amplitude scaling ratio of the qth channel based on the energy / amplitude scaling coefficient of the qth channel, where q is one or two.

[0019] In one possible design, determining the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the respective energies / amplitudes of the audio signals of the two channels of the current channel pair may include: determining the average energy / amplitude of the audio signals of the current channel pair based on the energies / amplitudes of the audio signals of the two channels of the current channel pair before energy / amplitude equalization; and determining the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the average energy / amplitude of the audio signals of the current channel pair.

[0020] This implementation method, by performing energy / amplitude balancing on the two channels within a channel pair, can achieve that even for channel pairs with large energy differences, a large energy difference can still be maintained after energy / amplitude balancing, thereby meeting the encoding requirements of the channel pairs with large energy / amplitude in subsequent encoding processing, thereby improving the quality of the audio signal reconstructed at the decoding end.

[0021] In one possible design, encoding the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain a coded bitstream may include encoding the energy / amplitude balanced side information of the K channel pairs, the channel pair indexes corresponding to the K and K channel pairs, and the audio signals of the P channels to obtain a coded bitstream.

[0022] In a second aspect, an embodiment of the present application provides a multi-channel audio signal decoding method, which may include: obtaining a code stream to be decoded. Demultiplexing the code stream to be decoded to obtain a current frame of the multi-channel audio signal to be decoded, the current frame including the number K of channel pairs, the channel pair indexes corresponding to the K channel pairs, and side information of energy / amplitude balance of the K channel pairs. Based on the channel pair indexes corresponding to the K channel pairs and the side information of energy / amplitude balance of the K channel pairs, the current frame of the multi-channel audio signal to be decoded is decoded to obtain a decoded signal of the current frame, where K is a positive integer and each channel pair includes two channels.

[0023] In one possible design, the K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, wherein the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance, and the energy / amplitude scaling identifier is used to identify whether the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

[0024] In one possible design, the K channel pairs include a current channel pair, and based on the channel pair indexes corresponding to the K channel pairs and the energy / amplitude balance side information of the K channel pairs, the current frame of the multi-channel audio signal to be decoded is decoded to obtain a decoded signal of the current frame. This may include: based on the channel pair index corresponding to the current channel pair, stereo decoding processing is performed on the current frame of the multi-channel audio signal to be decoded to obtain audio signals of two channels of the current channel pair of the current frame. Based on the energy / amplitude balance side information of the current channel pair, energy / amplitude de-balancing processing is performed on the audio signals of the two channels of the current channel pair to obtain decoded signals of the two channels of the current channel pair.

[0025] In one possible design, the current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio of the first channel, a fixed-point energy / amplitude scaling ratio of the second channel, an energy / amplitude scaling identifier of the first channel, and an energy / amplitude scaling identifier of the second channel.

[0026] The technical effects of the multi-channel audio signal decoding method can refer to the technical effects of the corresponding encoding method described above, which will not be repeated here.

[0027] In a third aspect, an embodiment of the present application provides an audio signal encoding device, which may be an audio encoder, or a chip or system-on-chip of an audio encoding device, or a functional module in an audio encoder for implementing the method of the first aspect or any possible design of the first aspect. The audio signal encoding device may implement the functions performed in the first aspect or each possible design of the first aspect, and the functions may be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, in one possible design, the audio signal encoding device may include: an acquisition module, an equalization side information generation module, and an encoding module.

[0028] In a fourth aspect, an embodiment of the present application provides an audio signal decoding device, which may be an audio decoder, or a chip or system-on-chip of an audio decoding device, or a functional module in an audio decoder for implementing the method of the second aspect or any possible design of the second aspect. The audio signal decoding device may implement the functions performed in the second aspect or each possible design of the second aspect, and the functions may be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, in one possible design, the audio signal decoding device may include: an acquisition module, a demultiplexing module, and a decoding module.

[0029] In a fifth aspect, an embodiment of the present application provides an audio signal encoding device, characterized in that it includes: a non-volatile memory and a processor coupled to each other, and the processor calls the program code stored in the memory to execute the above-mentioned first aspect or any possible design method of the above-mentioned first aspect.

[0030] In a sixth aspect, an embodiment of the present application provides an audio signal decoding device, characterized in that it includes: a non-volatile memory and a processor coupled to each other, and the processor calls the program code stored in the memory to execute the method of the above-mentioned second aspect or any possible design of the above-mentioned second aspect.

[0031] In a seventh aspect, an embodiment of the present application provides an audio signal encoding device, characterized in that it includes: an encoder, wherein the encoder is used to execute the method of the above-mentioned first aspect or any possible design of the above-mentioned first aspect.

[0032] In an eighth aspect, an embodiment of the present application provides an audio signal decoding device, characterized in that it includes: a decoder, wherein the decoder is used to execute the method of the second aspect or any possible design of the second aspect.

[0033] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium, characterized in that it includes an encoded code stream obtained according to the above-mentioned first aspect or any possible design method of the above-mentioned first aspect.

[0034] In the tenth aspect, an embodiment of the present application provides a computer-readable storage medium, including a computer program. When the computer program is executed on a computer, the computer executes the method described in any one of the first aspects above, or executes the method described in any one of the second aspects above.

[0035] In an eleventh aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by a computer, it is used to execute any one of the methods in the first aspect or any one of the methods in the second aspect.

[0036] In the twelfth aspect, the present application provides a chip comprising a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method as described in any one of the first aspects above, or to execute the method as described in any one of the second aspects above.

[0037] In the thirteenth aspect, the present application provides a coding and decoding device, which includes an encoder and a decoder, the encoder is used to execute the method of the above-mentioned first aspect or any possible design of the above-mentioned first aspect, and the decoder is used to execute the above-mentioned second aspect or any possible design of the above-mentioned second aspect.

[0038] The multi-channel audio signal encoding and decoding method and apparatus of the embodiments of the present application obtains the audio signals of P channels of the current frame of the multi-channel audio signal and the respective energies / amplitudes of the P channel audio signals, where the P channels include K channel pairs. Based on the respective energies / amplitudes of the audio signals of the P channels, side information for energy / amplitude balance of the K channel pairs is generated. Based on the side information for energy / amplitude balance of the K channel pairs, the audio signals of the P channels are encoded to obtain a coded bitstream. By generating the side information for energy / amplitude balance of the channel pairs, the coded bitstream carries the side information for energy / amplitude balance of the K channel pairs but does not carry the side information for energy / amplitude balance of the unpaired channels. This reduces the number of bits of the energy / amplitude balance side information in the coded bitstream, reduces the number of bits of the multi-channel side information, and allocates the saved bits to other functional modules of the encoder to improve the quality of the audio signal reconstructed at the decoding end and improve the encoding quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of an example of an audio encoding and decoding system in an embodiment of the present application;

[0040] Figure 2 This is a flowchart of a multi-channel audio signal encoding method according to an embodiment of the present application;

[0041] Figure 3 This is a flowchart of a multi-channel audio signal encoding method according to an embodiment of the present application;

[0042] Figure 4 A schematic diagram of the processing process of the encoding end in an embodiment of the present application;

[0043] Figure 5 Schematic diagram of the processing process of the multi-channel encoding processing unit according to an embodiment of the present application;

[0044] Figure 6 A schematic diagram of a process for writing multi-channel side information according to an embodiment of the present application;

[0045] Figure 7 This is a flowchart of a multi-channel audio signal decoding method according to an embodiment of the present application;

[0046] Figure 8 A schematic diagram of the processing process of the decoding end in an embodiment of the present application;

[0047] Figure 9 Schematic diagram of the processing process of the multi-channel decoding processing unit according to an embodiment of the present application;

[0048] Figure 10 This is a flowchart of multi-channel side information parsing according to an embodiment of the present application;

[0049] Figure 11 1 is a structural diagram of an audio signal encoding device 1100 according to an embodiment of the present application;

[0050] Figure 12 1 is a structural diagram of an audio signal encoding device 1200 according to an embodiment of the present application;

[0051] Figure 13 13 is a structural diagram of an audio signal decoding device 1300 according to an embodiment of the present application;

[0052] Figure 14 14 is a structural diagram of an audio signal decoding device 1400 according to an embodiment of the present application. DETAILED DESCRIPTION

[0053] The terms "first", "second", etc., as used in the embodiments of the present application, are only used for the purpose of distinguishing descriptions and are not to be understood as indicating or implying relative importance, nor as indicating or implying an order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, including a series of steps or units. Methods, systems, products, or devices are not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0054] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single, multiple, or some can be single and some can be multiple.

[0055] The following describes the system architecture used in the embodiments of this application. Figure 1 , Figure 1 The schematic block diagram of the audio encoding and decoding system 10 used in the embodiment of the present application is given as an example. Figure 1 As shown, audio encoding and decoding system 10 may include source device 12 and destination device 14. Source device 12 generates encoded audio data, and thus, source device 12 may be referred to as an audio encoding device. Destination device 14 may decode the encoded audio data generated by source device 12, and thus, destination device 14 may be referred to as an audio decoding device. Various implementations of source device 12, destination device 14, or both may include one or more processors and a memory coupled to the one or more processors. The memory may include, but is not limited to, RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures accessible by a computer, as described herein. Source device 12 and destination device 14 may include a variety of devices, including desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, televisions, speakers, digital media players, video game consoles, in-vehicle computers, any wearable devices, virtual reality (VR) devices, servers providing VR services, augmented reality (AR) devices, servers providing AR services, wireless communication devices, or the like.

[0056] Although Figure 1Source device 12 and destination device 14 are depicted as separate devices, but device embodiments may also include both source device 12 and destination device 14, or functionality of both, i.e., source device 12 or corresponding functionality and destination device 14 or corresponding functionality. In such embodiments, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or using separate hardware and / or software, or any combination thereof.

[0057] Source device 12 and destination device 14 may be communicatively connected via link 13, and destination device 14 may receive encoded audio data from source device 12 via link 13. Link 13 may include one or more media or devices capable of moving the encoded audio data from source device 12 to destination device 14. In one example, link 13 may include one or more communication media that enable source device 12 to transmit the encoded audio data directly to destination device 14 in real time. In this example, source device 12 may modulate the encoded audio data according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated audio data to destination device 14. The one or more communication media may include wireless and / or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from source device 12 to destination device 14.

[0058] The source device 12 includes an encoder 20. Optionally, the source device 12 may also include an audio source 16, a preprocessor 18, and a communication interface 22. In a specific implementation, the encoder 20, the audio source 16, the preprocessor 18, and the communication interface 22 may be hardware components in the source device 12 or software programs in the source device 12. They are described as follows:

[0059] The audio source 16 may include or may be any type of sound capture device, for example, for capturing real-world sounds, and / or any type of audio generation device. The audio source 16 may be a microphone for capturing sound or a memory for storing audio data. The audio source 16 may also include any type of interface (internal or external) for storing previously captured or generated audio data and / or acquiring or receiving audio data. When the audio source 16 is a microphone, the audio source 16 may be, for example, a local microphone or an integrated microphone integrated into the source device; when the audio source 16 is a memory, the audio source 16 may be, for example, a local memory or an integrated memory integrated into the source device. When the audio source 16 includes an interface, the interface may be, for example, an external interface for receiving audio data from an external audio source, such as an external sound capture device, such as a microphone, an external memory, or an external audio generation device. The interface may be any type of interface according to any proprietary or standardized interface protocol, such as a wired or wireless interface or an optical interface.

[0060] In the embodiment of the present application, the audio data transmitted from the audio source 16 to the pre-processor 18 may also be referred to as original audio data 17 .

[0061] The preprocessor 18 is configured to receive the original audio data 17 and perform preprocessing on the original audio data 17 to obtain preprocessed audio 19 or preprocessed audio data 19. For example, the preprocessing performed by the preprocessor 18 may include filtering or denoising.

[0062] The encoder 20 (or audio encoder 20 ) is used to receive the preprocessed audio data 19 and to execute the embodiments of the various encoding methods described below to implement the application of the audio signal encoding method described in this application on the encoding side.

[0063] The communication interface 22 may be configured to receive the encoded audio data 21 and transmit the encoded audio data 21 to the destination device 14 or any other device (e.g., a memory) via the link 13 for storage or direct reconstruction. The other device may be any device for decoding or storage. The communication interface 22 may, for example, be configured to encapsulate the encoded audio data 21 into a suitable format, such as data packets, for transmission over the link 13.

[0064] The destination device 14 includes a decoder 30. Optionally, the destination device 14 may also include a communication interface 28, an audio post-processor 32, and a speaker device 34. Each of these is described below:

[0065] Communication interface 28 may be configured to receive encoded audio data 21 from source device 12 or any other source, such as a storage device, such as an encoded audio data storage device. Communication interface 28 may be configured to transmit or receive encoded audio data 21 via link 13 between source device 12 and destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network, or any combination thereof, or any type of private or public network, or any combination thereof. Communication interface 28 may, for example, be configured to decapsulate data packets transmitted by communication interface 22 to obtain encoded audio data 21.

[0066] Both communication interface 28 and communication interface 22 may be configured as unidirectional or bidirectional communication interfaces and may be used, for example, to send and receive messages to establish connections, confirm and exchange any other information related to communication links and / or data transmission, such as encoded audio data transmission.

[0067] The decoder 30 (or decoder 30) is configured to receive the encoded audio data 21 and provide decoded audio data 31 or decoded audio 31. In some embodiments, the decoder 30 may be configured to execute the embodiments of the various decoding methods described below to implement the application of the audio signal decoding method described in this application on the decoding side.

[0068] The audio post-processor 32 is configured to perform post-processing on the decoded audio data 31 (also referred to as reconstructed audio data) to obtain post-processed audio data 33. The post-processing performed by the audio post-processor 32 may include, for example, rendering or any other processing, and may also be configured to transmit the post-processed audio data 33 to a speaker device 34.

[0069] The speaker device 34 is configured to receive the post-processed audio data 33 to play the audio to, for example, a user or viewer. The speaker device 34 may be or include any type of loudspeaker for presenting the reconstructed sound.

[0070] Although, Figure 1 Source device 12 and destination device 14 are depicted as separate devices, but device embodiments may also include both source device 12 and destination device 14, or functionality of both, i.e., source device 12 or corresponding functionality and destination device 14 or corresponding functionality. In such embodiments, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or using separate hardware and / or software, or any combination thereof.

[0071] It is obvious to those skilled in the art based on the description that the functionality or Figure 1The presence and (precise) division of functionality of the source device 12 and / or destination device 14 shown may vary depending on the actual device and application. The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or stationary device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a video camera, a desktop computer, a set-top box, a television, a camera, a car device, a stereo, a digital media player, an audio game console, an audio streaming device (such as a content service server or content distribution server), a broadcast receiver device, a broadcast transmitter device, smart glasses, a smart watch, etc., and may not use or use any type of operating system.

[0072] The encoder 20 and the decoder 30 may be implemented as any of a variety of suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the techniques are partially implemented in software, the device may store the software instructions in a suitable non-transitory computer-readable storage medium and may use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered one or more processors.

[0073] In some cases, Figure 1 The audio encoding and decoding system 10 shown in the figure is merely an example, and the technology of the present application can be applied to audio encoding settings (e.g., audio encoding or audio decoding) that do not necessarily include any data communication between the encoding and decoding devices. In other examples, data can be retrieved from local storage, streamed over a network, etc. The audio encoding device can encode data and store the data in a memory, and / or the audio decoding device can retrieve data from a memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to a memory and / or retrieve data from a memory and decode data.

[0074] The encoder may be a multi-channel encoder, for example, a stereo encoder, a 5.1-channel encoder, or a 7.1-channel encoder.

[0075] The above-mentioned audio data can also be called an audio signal. The audio signal in the embodiment of the present application refers to the input signal in the audio encoding device. The audio signal may include multiple frames. For example, the current frame may specifically refer to a certain frame in the audio signal. In the embodiment of the present application, the encoding and decoding of the current frame audio signal is used as an example. The previous frame or the next frame of the current frame in the audio signal can be encoded and decoded accordingly according to the encoding and decoding method of the current frame audio signal. The encoding and decoding process of the previous frame or the next frame of the current frame in the audio signal will not be described one by one. In addition, the audio signal in the embodiment of the present application may be a multi-channel audio signal, that is, including P channels. The embodiment of the present application is used to implement multi-channel audio signal encoding and decoding.

[0076] The above-mentioned encoder can implement the multi-channel audio signal encoding method of the embodiment of the present application to reduce the number of bits of multi-channel side information, thereby allocating the saved bits to other functional modules of the encoder to improve the quality of the reconstructed audio signal at the decoding end and improve the encoding quality. The specific implementation method can be referred to the detailed explanation of the following embodiment.

[0077] Figure 2 This is a flowchart of a multi-channel audio signal encoding method according to an embodiment of the present application. The execution subject of the embodiment of the present application may be the above-mentioned encoder, such as Figure 2 As shown, the method of this embodiment may include:

[0078] Step 201: Obtain audio signals of P channels of a current frame of a multi-channel audio signal and the energy / amplitude of each of the P channels, where the P channels include K channel pairs. The multi-channel signal can be a 5.1-channel signal (corresponding to P being 5+1=6), a 7.1-channel signal (corresponding to P being 7+1=8), or an 11.1-channel signal (corresponding to P being 11+1=12), etc.

[0079] Each channel pair includes two channels, P is a positive integer greater than 1, K is a positive integer, and P is greater than or equal to K*2.

[0080] In some embodiments, P=2K. By screening and pairing the multi-channel signal in the current frame of the multi-channel audio signal, K channel pairs can be obtained. The above-mentioned P channels include K channel pairs.

[0081] In some embodiments, P=2*K+Q, where Q is a positive integer. The P-channel audio signal also includes Q unpaired mono audio signals. Taking a 5.1-channel signal as an example, the 5.1 channel includes a left (L) channel, a right (R) channel, a center (C) channel, a low frequency effects (LFE) channel, a left surround (LS) channel, and a right surround (RS) channel. According to the multi-channel processing indicator (MultiProcFlag), the channels participating in the multi-channel processing are filtered out from the 5.1 channels. For example, the channels participating in the multi-channel processing include the L channel, the R channel, the C channel, the LS channel, and the RS channel. Pairing is performed among the channels participating in the multi-channel processing. For example, the L channel and the R channel are paired to form a first channel pair. The LS channel and the RS channel are paired to form a second channel pair. The LFE channel and the C channel are unpaired channels. That is, P=6, K=2, and Q=2. The P channels include a first channel pair, a second channel pair, and an unpaired LFE channel and a C channel.

[0082] Exemplarily, the channels participating in multi-channel processing may be paired by determining K channel pairs through multiple iterations, i.e., one channel pair is determined in each iteration. For example, in a first iteration, the inter-channel correlation value between any two channels among the P channels participating in the multi-channel processing is calculated. In the first iteration, the two channels with the highest inter-channel correlation values ​​are selected to form a channel pair. In a second iteration, the two channels with the highest inter-channel correlation values ​​among the remaining channels (excluding the paired channels among the P channels) are selected to form a channel pair. This process is repeated in this way, resulting in K channel pairs.

[0083] It should be noted that the embodiment of the present application may also adopt other pairing methods to determine K channel pairs, and the embodiment of the present application is not limited to the above exemplary description of pairing.

[0084] Step 202: Generate energy / amplitude balanced side information for K channel pairs based on the energy / amplitude of the audio signals of the P channels.

[0085] It should be noted that the "energy / amplitude" in the embodiments of the present application refers to energy or amplitude, and in the actual processing process, for the processing of a frame, if the energy is processed at the beginning, then the energy will be processed in the subsequent processing, or if the amplitude is processed at the beginning, then the amplitude will be processed in the subsequent processing.

[0086] For example, based on the energy of the audio signals in P channels, energy-balanced side information for K channel pairs is generated. That is, energy balancing is performed using the energy of the P channels to obtain energy-balanced side information. Alternatively, based on the amplitude of the audio signals in P channels, energy-balanced side information for K channel pairs is generated. That is, energy balancing is performed using the amplitude of the P channels to obtain energy-balanced side information. Alternatively, based on the amplitude of the audio signals in P channels, amplitude-balanced side information for K channel pairs is generated. That is, amplitude balancing is performed using the amplitude of the P channels to obtain amplitude-balanced side information.

[0087] Specifically, an embodiment of the present invention performs stereo encoding processing on a channel pair. In order to improve encoding efficiency and encoding effect, for example, before performing stereo encoding processing on the current channel pair, the energy / amplitude of the audio signals of the two channels of the current channel pair can be first energy / amplitude balanced to obtain the energy / amplitude of the two channels after energy / amplitude balance, and then subsequent stereo encoding processing is performed based on the energy / amplitude after energy / amplitude balance. In one embodiment, energy / amplitude balance can be based on the audio signals of the two channels of the current channel pair, but not based on the audio signals corresponding to other channel pairs and / or mono channels outside the current channel pair; in another embodiment, energy / amplitude balance can be based on the audio signals of the two channels of the current channel pair, and can also be further based on the audio signals corresponding to other channel pairs and / or mono channels.

[0088] The energy / amplitude balanced side information is used by the decoding end to perform energy / amplitude de-balancing to obtain a decoded signal.

[0089] In one implementation, the energy / amplitude balancing side information may include a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier. The fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, which is obtained based on the energy / amplitude before energy / amplitude balancing and the energy / amplitude after energy / amplitude balancing. The energy / amplitude scaling identifier is used to indicate whether the energy / amplitude after energy / amplitude balancing is amplified or reduced relative to the energy / amplitude before energy / amplitude balancing. The energy / amplitude scaling coefficient may be an energy / amplitude scaling coefficient, which is between (0, 1).

[0090] Taking a channel pair as an example, the side information of the energy / amplitude balance of the channel pair may include the fixed-point energy / amplitude scaling ratio and the energy / amplitude scaling identifier of the channel pair. Taking the channel pair including the first channel and the second channel as an example, the fixed-point energy / amplitude scaling ratio of the channel pair includes the fixed-point energy / amplitude scaling ratio of the first channel and the fixed-point energy / amplitude scaling ratio of the second channel, and the energy / amplitude scaling identifier of the channel pair includes the energy / amplitude scaling identifier of the first channel and the energy / amplitude scaling identifier of the second channel. Taking the first channel as an example, the fixed-point energy / amplitude scaling ratio of the first channel is the fixed-point value of the energy / amplitude scaling coefficient of the first channel, and the energy / amplitude scaling coefficient of the first channel is obtained based on the energy / amplitude of the audio signal of the first channel before energy / amplitude balance and the energy / amplitude of the audio signal of the first channel after energy / amplitude balance. The energy / amplitude scaling identifier of the first channel is obtained based on the energy / amplitude of the audio signal of the first channel before energy / amplitude balance and the energy / amplitude of the audio signal of the first channel after energy / amplitude balance. For example, the energy / amplitude scaling coefficient of the first channel is the smaller of the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization and the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization, divided by the larger of the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization and the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization. For example, if the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization is greater than the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization, then the energy / amplitude scaling coefficient of the first channel is the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization divided by the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization. When the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization is greater than the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization, the energy / amplitude scaling indicator of the first channel is 1. When the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization is greater than the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization, the energy / amplitude scaling flag of the first channel is 0. Of course, it is understandable that it is also possible to set the energy / amplitude scaling flag of the first channel to 0 when the energy / amplitude of the audio signal of the first channel before energy / amplitude equalization is greater than the energy / amplitude of the audio signal of the first channel after energy / amplitude equalization. The implementation principle is similar, and the embodiments of the present application are not limited to the above example.

[0091] The energy / amplitude scaling coefficient of the embodiment of the present application may also be referred to as a floating-point energy / amplitude scaling coefficient.

[0092] In another implementation, the side information of the energy / amplitude balance can include a fixed-point energy / amplitude scaling ratio. The fixed-point energy / amplitude scaling ratio is a fixed-point value of the energy / amplitude scaling coefficient, which is the ratio of the energy / amplitude before energy / amplitude balance to the energy / amplitude after energy / amplitude balance. That is, the energy / amplitude scaling coefficient is the energy / amplitude before energy / amplitude balance divided by the energy / amplitude after energy / amplitude balance. When the energy / amplitude scaling coefficient is less than 1, the decoding end can determine that the energy / amplitude after energy / amplitude balance is amplified relative to the energy / amplitude before energy / amplitude balance. When the energy / amplitude scaling coefficient is greater than 1, the decoding end can determine that the energy after energy / amplitude balance is reduced relative to the energy / amplitude before energy / amplitude balance. Of course, it is understandable that the energy / amplitude scaling coefficient can also be the energy / amplitude after energy / amplitude balance divided by the energy / amplitude before energy / amplitude balance. The implementation principle is similar, and the embodiments of the present application are not limited to the above examples. In this implementation, the energy / amplitude balancing side information may not include an energy / amplitude scaling identifier.

[0093] Step 203: Encode the audio signals of the P channels based on the energy / amplitude balanced side information of the K channel pairs to obtain a coded bitstream.

[0094] The energy / amplitude-balanced side information for the K channel pairs and the P channel audio signals are encoded to obtain a coded bitstream. Specifically, the energy / amplitude-balanced side information for the K channel pairs is written into the coded bitstream. In other words, the coded bitstream carries the energy / amplitude-balanced side information for the K channel pairs, but does not carry the energy / amplitude-balanced side information for unpaired channels. This reduces the number of bits of the energy / amplitude-balanced side information in the coded bitstream.

[0095] In some embodiments, the coded bitstream also carries the number of channel pairs and K channel pair indices for the current frame. The number of channel pairs and the K channel pair indices are used by the decoding end to perform stereo decoding, energy / amplitude de-equalization, and other processing. A channel pair index is used to indicate the two channels included in a channel pair. In other words, one implementation of step 203 is to encode the energy / amplitude balanced side information of the K channel pairs, the number of channel pairs, the K channel pair indices, and the audio signals of the P channels to obtain a coded bitstream. The number of channel pairs can be K. The K channel pair indices include the channel pair indices corresponding to each of the K channel pairs.

[0096] The order in which the number of channel pairs, the K channel pair indices, and the energy / amplitude balance side information for the K channel pairs are written into the encoded bitstream is to write the number of channel pairs first so that when a decoder decodes the received bitstream, it first obtains the number of channel pairs. The K channel pair indices and the energy / amplitude balance side information for the K channel pairs are then written.

[0097] It should also be noted that the number of channel pairs can be 0, meaning there are no paired channels. In this case, the number of channel pairs and the P-channel audio signal are encoded to obtain a coded bitstream. When the decoder decodes the received bitstream and first detects that the number of channel pairs is 0, it can directly decode the current frame of the multi-channel audio signal to be decoded without parsing for side information related to energy / amplitude balance.

[0098] Before obtaining the encoded bitstream, energy / amplitude equalization may be performed on coefficients in the current frame of the channel according to the fixed-point energy / amplitude scaling ratio and energy / amplitude scaling flag of the channel.

[0099] In this embodiment, P channels of a current frame of a multi-channel audio signal are obtained, where the P channels include K channel pairs. Based on the energy / amplitude of the audio signals in the P channels, energy / amplitude-balanced side information for the K channel pairs is generated. Based on the energy / amplitude-balanced side information for the K channel pairs, the P channel audio signals are encoded to obtain a coded bitstream. By generating the energy / amplitude-balanced side information for the channel pairs, the coded bitstream carries the energy / amplitude-balanced side information for the K channel pairs but does not carry the energy / amplitude-balanced side information for unpaired channels. This reduces the number of bits of the energy / amplitude-balanced side information in the coded bitstream, thus reducing the number of bits of the multi-channel side information. The saved bits can be allocated to other functional modules of the encoder, thereby improving the quality of the audio signal reconstructed at the decoder and thus improving the encoding quality.

[0100] Figure 3 This is a flowchart of a multi-channel audio signal encoding method according to an embodiment of the present application. The execution subject of the embodiment of the present application may be the above-mentioned encoder. Figure 2 A specific implementation method of the method described in the embodiment shown is as follows: Figure 3 As shown, the method of this embodiment may include:

[0101] Step 301: Obtain P channels of audio signals of a current frame of a multi-channel audio signal.

[0102] Step 302: Screen and pair the P channels of the current frame of the multi-channel audio signal to determine K channel pairs and K channel pair indexes.

[0103] The specific implementation of screening and pairing can be found in Figure 2Explanation of step 201 of the illustrated embodiment.

[0104] A channel pair index is used to indicate the two channels included in the channel pair. Different values ​​of the channel pair index correspond to different two channels. The correspondence between the channel pair index value and the two channels can be preset.

[0105] Taking a 5.1-channel signal as an example, for example, through screening and pairing, the L channel and the R channel are paired to form a first channel pair. The LS channel and the RS channel are paired to form a second channel pair. The LFE channel and the C channel are unpaired channels. That is, K = 2. The first channel pair index is used to indicate the pairing of the L channel and the R channel. For example, the value of the first channel pair index is 0. The second channel pair index is used to indicate the pairing of the LS channel and the RS channel. For example, the value of the second channel pair index is 9.

[0106] Step 303: Perform energy / amplitude equalization processing on the audio signals of the K channel pairs respectively to obtain the energy / amplitude-equalized audio signals of the K channel pairs and energy / amplitude-equalized side information of the K channel pairs.

[0107] Taking the energy / amplitude balancing of a channel pair as an example, one possible implementation involves performing energy / amplitude balancing at the channel pair level: determining the energy / amplitude of the audio signals of the two channels of the channel pair after energy / amplitude balancing based on the energy / amplitude of the audio signals of the two channels before energy / amplitude balancing. Based on the energy / amplitude of the audio signals of the two channels of the channel pair before energy / amplitude balancing and the energy / amplitude of the audio signals of the two channels after energy / amplitude balancing, energy / amplitude balancing side information for the current channel pair is generated, and the audio signals of the two channels after energy / amplitude balancing are obtained.

[0108] Specifically, determining the energy / amplitude of the audio signals of the two channels of the channel pair after energy / amplitude equalization can be performed by determining an average energy / amplitude of the audio signals of the channel pair based on the energy / amplitude of the audio signals of the two channels of the channel pair before energy / amplitude equalization, and determining the energy / amplitude of the audio signals of the two channels of the channel pair after energy / amplitude equalization based on the average energy / amplitude of the audio signals of the channel pair. For example, the energy / amplitude of the audio signals of the two channels of the channel pair after energy / amplitude equalization is equal, and both are the average energy / amplitude of the audio signals of the channel pair.

[0109] As described above, a channel pair may include a first channel and a second channel, and the side information of the energy / amplitude balance of the channel pair includes: a fixed-point energy / amplitude scaling ratio of the first channel, a fixed-point energy / amplitude scaling ratio of the second channel, an energy / amplitude scaling identifier of the first channel, and an energy / amplitude scaling identifier of the second channel.

[0110] In some embodiments, an energy / amplitude scaling coefficient of the qth channel of the channel pair can be determined based on the energy / amplitude of the audio signal of the qth channel before energy / amplitude equalization and the energy / amplitude of the audio signal of the qth channel after energy / amplitude equalization. A fixed-point energy / amplitude scaling ratio of the qth channel can be determined based on the energy / amplitude scaling coefficient of the qth channel. An energy / amplitude scaling identifier of the qth channel can be determined based on the energy / amplitude of the qth channel before energy / amplitude equalization and the energy / amplitude of the qth channel after energy / amplitude equalization. Wherein, q is one or two.

[0111] For example, the fixed-point energy / amplitude scaling ratio of the qth channel of a channel pair and the energy / amplitude scaling identifier of the qth channel may be determined according to the following formulas (1) to (3).

[0112] The fixed-point energy / amplitude scaling ratio of the qth channel is calculated according to formulas (1) and (2).

[0113] scaleInt_q=ceil((1<<M)×scaleF_q) (1)

[0114] scaleInt_q=clip(scaleInt_q, 1, 2 M -1) (2)

[0115] Where scaleInt_q is the fixed-point energy / amplitude scaling factor for the qth channel, scaleF_q is the floating-point energy / amplitude scaling factor for the qth channel, M is the number of bits used to convert the floating-point energy / amplitude scaling factor to the fixed-point energy / amplitude scaling factor, clip(x, a, b) is a bidirectional clamping function that clamps x to the range [a, b], where clip((x), (a), (b)) = max(a, min(b, (x))), where a ≤ b, and ceil(x) rounds x upwards. M can be any integer, for example, 4.

[0116] When energy_q>energy_qe, energyBigFlag_q is set to 1, when energy_q≤energy_qe, energyBigFlag_q is set to 0.

[0117] Among them, energy_q is the energy / amplitude before energy / amplitude equalization of the q-th channel, energy_qe is the energy / amplitude after energy / amplitude equalization of the q-th channel, and energyBigFlag_q is the energy / amplitude scaling flag of the q-th channel. energy_qe can be the average value of the energy / amplitude of the two channels of the channel pair.

[0118] The determination method of scaleF_q in the above formula (1): when energy_q > energy_qe, scaleF_q = energy_qe / energy_q; when energy_q ≤ energy_qe, scaleF_q = energy_q / energy_qe.

[0119] Among them, energy_q is the energy / amplitude before energy / amplitude equalization of the q-th channel, energy_q e is the energy / amplitude after energy / amplitude equalization of the q-th channel, and scaleF_q is the floating-point energy / amplitude scaling ratio coefficient of the q-th channel.

[0120] Among them, energy_q is determined by the following formula (3).

[0121]

[0122] Among them, sampleCoef(q, i) represents the i-th coefficient of the current frame of the q-th channel before energy / amplitude equalization, and N is the number of frequency domain coefficients of the current frame.

[0123] In the process of energy / amplitude equalization processing, the current frame of the q-th channel can be subjected to energy / amplitude equalization according to the fixed-point energy / amplitude scaling ratio of the q-th channel and the energy / amplitude scaling flag of the q-th channel, so as to obtain the audio signal after energy / amplitude equalization of the q-th channel.

[0124] For example, when energyBigFlag_q is 1, q e (i) = q(i) × scaleInt_q / (1 << M). When energyBigFlag_q is 0, q e (i) = q(i) × (1 << M) / scaleInt_q.

[0125] Among them, i is used to identify the coefficient of the current frame, q(i) is the i-th frequency domain coefficient of the current frame before energy / amplitude equalization, and q e (i) is the i-th frequency domain coefficient of the current frame after energy / amplitude equalization, and M is the number of fixed-point bits from the floating-point energy / amplitude scaling ratio coefficient to the fixed-point energy / amplitude scaling ratio.

[0126] Another possible implementation method is to perform energy / amplitude balancing processing with all channels, all channel pairs, or part of all channels as the granularity. For example, based on the energy / amplitude of each audio signal of the P channels before energy / amplitude balancing, the energy / amplitude average value of the audio signals of the P channels is determined, and based on the energy / amplitude average value of the audio signals of the P channels, the energy or amplitude of each audio signal of the two channels of a channel pair after energy / amplitude balancing is determined. For example, the energy / amplitude average value of the audio signals of the P channels can be used as the energy or amplitude of the audio signal of any one channel of a channel pair after energy / amplitude balancing. That is, the method for determining the energy or amplitude after energy / amplitude balancing is different from the above-mentioned possible implementation method, and other methods for determining the side information of energy / amplitude balancing can be the same. The specific implementation method can be found in the above description and will not be repeated here.

[0127] In the above embodiment, the side information of energy / amplitude equalization of the current channel pair includes a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the first channel, and a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the second channel. That is, for the current channel (the first channel or the second channel), the side information includes both the fixed-point energy / amplitude scaling ratio and the energy / amplitude scaling identifier. This is because when obtaining the energy / amplitude scaling ratio, the larger of the energy / amplitude of the current channel before energy / amplitude equalization and the energy / amplitude of the current channel after energy / amplitude equalization is divided by the smaller one, or the smaller one is divided by the larger one. Therefore, the obtained energy / amplitude scaling ratio is fixedly greater than or equal to 1, or the obtained energy / amplitude scaling ratio is fixedly less than or equal to 1. Therefore, simply using the energy / amplitude scaling ratio or the fixed-point energy / amplitude scaling ratio cannot determine whether the energy / amplitude after energy / amplitude equalization is greater than the energy / amplitude before energy / amplitude equalization. Therefore, the energy / amplitude scaling identifier is required to indicate.

[0128] In another embodiment of this aspect, the energy / amplitude of the current channel before energy / amplitude equalization and the energy / amplitude of the current channel after energy / amplitude equalization can be fixedly used, or the energy / amplitude of the previous channel after energy / amplitude equalization and the energy / amplitude of the current channel before energy / amplitude equalization can be fixedly used. In this way, there is no need to indicate it through an energy / amplitude scaling identifier. Accordingly, the side information of the current channel can include a fixed-point energy / amplitude scaling ratio, but does not need to include an energy / amplitude scaling identifier.

[0129] Step 304: Perform stereo processing on the energy / amplitude-equalized audio signals of the K channel pairs to obtain the stereo-processed audio signals of the K channel pairs and stereo side information of the K channel pairs.

[0130] Taking a channel pair as an example, stereo processing is performed on the energy / amplitude balanced audio signals of the two channels of the channel pair to obtain stereo processed audio signals of the two channels and generate stereo side information of the channel pair.

[0131] Step 305: Encode the stereo-processed audio signals of the K channel pairs, the energy / amplitude-balanced side information of the K channel pairs, the stereo side information of the K channel pairs, the K, K channel pair indexes, and the audio signals of the unpaired channels to obtain a coded bitstream.

[0132] Encode the stereo processed audio signals of the K channel pairs, the energy / amplitude balanced side information of the K channel pairs, the stereo side information of the K channel pairs, the number of channel pairs (K), the indexes of the K channel pairs, and the audio signals of the unpaired channels to obtain a coded bitstream for decoding and reconstructing the audio signal at a decoding end.

[0133] In this embodiment, audio signals of P channels of a current frame of a multi-channel audio signal are obtained, the P channels of the current frame of the multi-channel audio signal are screened and paired, K channel pairs and K channel pair indexes are determined, energy / amplitude equalization processing is performed on the audio signals of each of the K channel pairs, the energy / amplitude-equalized audio signals of each of the K channel pairs and energy / amplitude-equalized side information of each of the K channel pairs are obtained, stereo processing is performed on the energy / amplitude-equalized audio signals of each of the K channel pairs, the stereo-processed audio signals of each of the K channel pairs and the stereo side information of each of the K channel pairs are obtained, and the stereo-processed audio signals of the K channel pairs, the energy / amplitude-equalized side information of the K channel pairs, the stereo side information of the K channel pairs, K, the K channel pair indexes, and the audio signals of unpaired channels are encoded to obtain an encoded bitstream. By generating energy / amplitude balanced side information for channel pairs, the encoded bitstream carries the energy / amplitude balanced side information for K channel pairs, but does not carry the energy / amplitude balanced side information for unpaired channels. This reduces the number of bits of the energy / amplitude balanced side information in the encoded bitstream, and reduces the number of bits of multi-channel side information. The saved bits can be allocated to other functional modules of the encoder, thereby improving the quality of the audio signal reconstructed at the decoding end and improving the encoding quality.

[0134] The following embodiment takes a 5.1-channel signal as an example to schematically illustrate the multi-channel audio signal encoding method of the embodiment of the present application.

[0135] Figure 4 This is a schematic diagram of the processing process of the encoding end of the embodiment of the present application, such as Figure 4As shown, the encoding end may include a multi-channel encoding processing unit 401, a channel encoding unit 402, and a stream multiplexing interface 403. The encoding end may be the encoder described above.

[0136] The multi-channel encoding processing unit 401 is used to filter, pair, and stereo process the input signal, generate energy / amplitude balanced side information, and generate stereo side information. In this embodiment, the input signal is a 5.1 (L channel, R channel, C channel, LFE channel, LS channel, RS channel) signal.

[0137] In one example, the multi-channel encoding processing unit 401 pairs the L channel signal and the R channel signal to form a first channel pair, and obtains the center channel M1 channel signal and the side channel S1 channel signal through stereo processing, and pairs the LS channel signal and the RS channel signal to form a second channel pair, and obtains the center channel M2 channel signal and the side channel S2 channel signal through stereo processing. Figure 5 The embodiment shown.

[0138] The multi-channel encoding processing unit 401 outputs stereo-processed M1 channel signal, S1 channel signal, M2 channel signal, S2 channel signal, LFE channel signal and C channel signal that have not been stereo-processed, as well as energy / amplitude balanced side information, stereo side information and channel pair index.

[0139] The channel encoding unit 402 is configured to encode the stereo-processed M1 channel signal, S1 channel signal, M2 channel signal, S2 channel signal, the unstereo-processed LFE channel signal and C channel signal, as well as multi-channel side information, and output encoded channels E1-E6. This multi-channel side information may include energy / amplitude equalization side information, stereo side information, and channel pair indexes. It is understood that this multi-channel side information may also include bit allocation side information, entropy coding side information, etc., which is not specifically limited in this embodiment of the present application. The channel encoding unit 402 sends the encoded channels E1-E6 to the stream multiplexing interface 403.

[0140] The code stream multiplexing interface 403 multiplexes the six coded channels E1-E6 to form a serial bit stream (bitStream), namely, a coded code stream, to facilitate the transmission of the multi-channel audio signal in a channel or the storage in a digital medium.

[0141] Figure 5 This is a schematic diagram of the processing process of the multi-channel encoding processing unit of the embodiment of the present application, as shown in FIG. Figure 5As shown, the above-mentioned multi-channel encoding processing unit 401 may include a multi-channel screening unit 4011 and an iterative processing unit 4012, and the iterative processing unit 4012 may include a group pair decision unit 40121, a channel pair energy / amplitude balancing unit 40122, a channel pair energy / amplitude balancing unit 40123, a stereo processing unit 40124 and a stereo processing unit 40125.

[0142] The multi-channel screening unit 4011 screens out the channels participating in multi-channel processing from the 5.1 input channels (L channel, R channel, C channel, LS channel, RS channel, LFE channel) according to the multi-channel processing indicator (MultiProcFlag), including L channel, R channel, C channel, LS channel, and RS channel.

[0143] In the first iteration step, the pair determination unit 40121 in the iterative processing unit 4012 calculates the inter-channel correlation value between each pair of channels among the L channel, R channel, C channel, LS channel, and RS channel. In the first iteration step, the channel pair (L channel, R channel) with the highest inter-channel correlation value among the channels (L channel, R channel, C channel, LS channel, RS channel) is selected to form the first channel pair. The L channel and the R channel are energy / amplitude balanced by the channel pair energy / amplitude equalization unit 40122 to obtain the L channel. e Vocal channel and R e Channel. Stereo processing unit 40124 pairs L e Vocal channel and R e The channels are stereo processed to obtain the side information of the first channel pair and the center channel M1 and side channel S1 after stereo processing. The side information of the first channel pair includes the energy / amplitude balanced side information of the first channel pair, stereo side information and channel index. In the second iterative step, the channel pair (LS channel, RS channel) with the highest inter-channel correlation value among the channels (C channel, LS channel, RS channel) is selected to form the second channel pair. The LS channel and the RS channel are energy / amplitude balanced by the energy / amplitude equalization unit 40123 to obtain the LS channel. e Channel and RS e Channel. Stereo processing unit 40125 pairs of LS e Channel and RS e The channels are stereo-processed to obtain side information for the second channel pair and the stereo-processed center channel M2 and side channel S2. The side information for the second channel pair includes energy / amplitude balance side information for the second channel pair, stereo side information, and channel indexes. The side information for the first channel pair and the side information for the second channel pair constitute multi-channel side information.

[0144] The channel pair energy / amplitude balancing unit 40122 and the channel pair energy / amplitude balancing unit 40123 average the energy / amplitude of the input channel pairs to obtain energy / amplitude after energy / amplitude balance.

[0145] For example, the channel pair energy / amplitude balancing unit 40122 may determine the energy / amplitude after energy / amplitude balancing by using the following formula (4).

[0146] energy_avg_pair1=avg(energy_L,energy_R) (4)

[0147] The Avg(a1, a2) function outputs the average of the two parameters a1 and a2. energy_L is the frame energy / amplitude of the L channel before energy / amplitude equalization, energy_R is the frame energy / amplitude of the R channel before energy / amplitude equalization, and energy_avg_pair1 is the energy / amplitude of the first channel pair after energy / amplitude equalization.

[0148] Wherein, energy_L and energy_R can be determined by the above formula (3).

[0149] The channel pair energy / amplitude balancing unit 40123 can determine the energy / amplitude after energy / amplitude balancing by the following formula (4).

[0150] energy_avg_pair2=avg(energy_LS,energy_RS) (5)

[0151] The Avg(a1, a2) function outputs the average of the two parameters a1 and a2. energy_LS is the frame energy / amplitude of the LS channel before energy / amplitude equalization, energy_RS is the frame energy / amplitude of the RS channel before energy / amplitude equalization, and energy_avg_pair2 is the energy / amplitude of the second channel pair after energy / amplitude equalization.

[0152] During the energy / amplitude balancing process, energy / amplitude balancing side information for the first channel pair and energy / amplitude balancing side information for the second channel pair are generated, as described in the above embodiment. This energy / amplitude balancing side information for the first channel pair and the second channel pair is transmitted in the encoded bitstream to guide energy / amplitude balancing at the decoder.

[0153] The method for determining the side information of the energy / amplitude balance of the first channel pair is explained.

[0154] S01: Calculate the energy / amplitude energy_avg_pair1 of the first channel pair after being equalized by the channel pair energy / amplitude equalization unit 40122. The energy_avg_pair1 is determined using the above formula (4).

[0155] S02: Calculate a floating-point energy / amplitude scaling coefficient for the L channel of the first channel pair.

[0156] In one example, the floating-point energy / amplitude scaling factor for the L channel is scaleF_L. The floating-point energy / amplitude scaling factor is between (0, 1). If energy_L>energy_Le, scaleF_L=energy_Le / energy_L. Conversely, if energy_L≤energy_Le, scaleF_L=energy_L / energy_Le.

[0157] Where energy_Le is equal to energy_avg_pair1.

[0158] S03: Calculate a fixed-point energy / amplitude scaling ratio of the L channel of the first channel pair.

[0159] In one example, the fixed-point energy / amplitude scaling factor for the L channel is scaleInt_L. The number of fixed-point bits from the floating-point energy / amplitude scaling factor scaleF_L to the fixed-point energy / amplitude scaling factor scaleInt_L is a fixed value. The number of fixed-point bits determines the accuracy of the floating-point to fixed-point conversion, while also taking into account transmission efficiency (because side information takes up bits). Assuming the number of fixed-point bits is 4 (i.e., M = 4), the formula for calculating the fixed-point energy / amplitude scaling factor for the L channel is as follows:

[0160] scaleInt_L=ceil((1<<4)×scaleF_L)

[0161] scaleInt_L=clip(scaleInt_L, 1, 15)

[0162] clip((x), (a), (b)) = max(a, min(b, (x))), where a ≤ b. Ceil(x) rounds x upwards. clip(x, a, b) is a two-way clamping function that clamps x to the range [a, b].

[0163] S04: Calculate the energy / amplitude scaling flag of the L channel of the first channel pair.

[0164] In one example, the energy / amplitude scaling flag of the L channel is energyBigFlag_L. If energy_L>energy_Le, energyBigFlag_L is set to 1, otherwise if energy_L≤energy_Le, energyBigFlag_L is set to 0.

[0165] Perform energy / amplitude balancing on each coefficient in the current frame of the L channel as follows:

[0166] If energyBigFlag_L is 1, L e (i) = L(i) × scaleInt_L / (1<<4). Where i is used to identify the coefficient of the current frame, L(i) is the i-th frequency domain coefficient of the current frame before energy / amplitude equalization, and L e (i) is the i-th frequency domain coefficient of the current frame after energy / amplitude equalization. If energyBigFlag_L is 0, L e (i)=L(i)×(1<<4) / scaleInt_L.

[0167] Similar S01 to S04 operations can be performed on the R channel of the first channel pair to obtain the floating-point energy / amplitude scaling coefficient scaleF_R, the fixed-point energy / amplitude scaling ratio scaleInt_R, the energy / amplitude scaling flag energyBigFlag_R, and the current frame R after energy / amplitude equalization. e That is, replace L with R in the above S01 to S04.

[0168] Similar operations S01 to S04 can be performed on the LS channel of the second channel pair to obtain the floating-point energy / amplitude scaling factor scaleF_LS, the fixed-point energy / amplitude scaling factor scaleInt_LS, the energy / amplitude scaling flag energyBigFlag_LS, and the current frame LSe after energy / amplitude equalization. That is, replace L in the above S01 to S04 with LS.

[0169] Perform similar S01 to S04 operations on the second channel and the RS channel to obtain the floating-point energy / amplitude scaling coefficient scaleF_RS, the fixed-point energy / amplitude scaling ratio scaleInt_RS, the energy / amplitude scaling flag energyBigFlag_RS, and the current frame RS after energy / amplitude equalization. e .

[0170] Write multi-channel side information to the encoded bitstream, the multi-channel side information including the number of channel pairs, side information of energy / amplitude balance of the first channel pair, the first channel pair index, side information of energy / amplitude balance of the second channel pair, and the second channel pair index.

[0171] For example, the number of channel pairs is currPairCnt, the energy / amplitude balanced side information of the first channel pair and the energy / amplitude balanced side information of the second channel pair are two-dimensional arrays, and the first channel pair index and the second channel pair index are one-dimensional arrays. For example, the fixed-point energy / amplitude scaling ratios of the first channel pair are PairILDScale[0][0] and PairILDScale[0][1], the energy / amplitude scaling flags of the first channel pair are energyBigFlag[0][0] and energyBigFlag[0][1], the fixed-point energy / amplitude scaling ratios of the second channel pair are PairILDScale[1][0] and PairILDScale[1][1], and the energy / amplitude scaling flags of the second channel pair are energyBigFlag[1][0] and energyBigFlag[1][1]. The first channel pair index is PairIndex[0], and the second channel pair index is PairIndex[1].

[0172] The number of channel pairs currPairCnt may be of a fixed bit length, for example, may be composed of 4 bits, and may identify a maximum of 16 stereo pairs.

[0173] The values ​​of the channel pair index PairIndex[pair] are defined as shown in Table 1. The channel pair index can be variable-length coded for transmission in the coded bitstream to save bits and for audio signal recovery at the decoder. For example, PairIndex[0] = 0, indicating that the channel pair includes the R channel and the L channel.

[0174] Table 1 Channel pair index mapping table for 5 channels

[0175] 0(L) 1(R) 2(C) 3(RS) 4(RS) 0(L) 0 1 3 6 1(R) 2 4 7 2(C) 5 8 3(RS) 9 4(RS)

[0176] In this embodiment, PairILDScale[0][0]=scaleInt_L, PairILDScale[0][1]=scaleInt_R.

[0177] PairILDScale[1][0]=scaleInt_LS. PairILDScale[1][1]=scaleInt_RS.

[0178] energyBigFlag[0][0]=energyBigFlag_L. energyBigFlag[0][1]=energyBigFlag_R.

[0179] energyBigFlag[1][0]=energyBigFlag_LS. energyBigFlag[1][1]=energyBigFlag_RS.

[0180] PairIndex[0]=0 (L and R). PairIndex[1]=9 (LS and RS).

[0181] For example, the process of writing multi-channel side information into the code stream is as follows: Figure 6 As shown. Step 601, set the variable pair = 0, and write the number of channel pairs into the bitstream. For example, the number of channel pairs currPairCnt can be 4 bits. Step 602, determine whether pair is less than the number of channel pairs. If so, execute step 603, if not, end. Step 603, write the index of the i-th channel pair into the bitstream. i = pair + 1, for example, write PairIndex[0] into the bitstream. Step 604, write the fixed-point energy / amplitude scaling ratio of the i-th channel pair into the bitstream. For example, write PairILDScale[0][0] and PairILDScale[0][1] into the bitstream. PairILDScale[0][0] and PairILDScale[0][1] can each occupy 4 bits. Step 605, write the energy / amplitude scaling flag of the i-th channel pair into the bitstream. For example, write energyBigFlag[0][0] and energyBigFlag[0][1] into the bitstream. energyBigFlag[0][0] and energyBigFlag[0][1] can each occupy 1 bit. Step 606: Write the stereo side information of the i-th channel pair into the bitstream, with pair=pair+1, and return to step 602. After returning to step 602, write PairIndex[1], PairILDScale[1][0], PairILDScale[1][1], energyBigFlag[1][0], and energyBigFlag[1][1] into the bitstream until the end.

[0182] Figure 7 This is a flowchart of a multi-channel audio signal decoding method according to an embodiment of the present application. The execution subject of the embodiment of the present application may be the above-mentioned decoder, such as Figure 7 As shown, the method of this embodiment may include:

[0183] Step 701: Obtain a code stream to be decoded.

[0184] The code stream to be decoded may be the coded code stream obtained in the above coding method embodiment.

[0185] Step 702: Demultiplex the code stream to be decoded to obtain a current frame of the multi-channel audio signal to be decoded and the number of channel pairs included in the current frame.

[0186] Taking a 5.1-channel signal as an example, after demultiplexing the decoded bitstream, an M1 channel signal, an S1 channel signal, an M2 channel signal, an S2 channel signal, an LFE channel signal, and a C channel signal, as well as the number of channel pairs, are obtained.

[0187] Step 703: Determine whether the number of channel pairs is equal to 0. If so, execute step 704; if not, execute step 705.

[0188] Step 704: Decode the current frame of the multi-channel audio signal to be decoded to obtain a decoded signal of the current frame.

[0189] When the number of channel pairs is equal to 0, that is, all channels are not paired, the current frame of the multi-channel audio signal to be decoded can be decoded to obtain a decoded signal of the current frame.

[0190] Step 705: Parse the current frame to obtain K channel pair indices and energy / amplitude balance side information of the K channel pairs included in the current frame.

[0191] When the number of channel pairs is equal to K, the current frame can be further parsed to obtain other control information, such as K channel pair indices and side information of energy / amplitude balance of the K channel pairs of the current frame, so as to perform energy / amplitude de-balancing in the subsequent decoding process of the current frame of the multi-channel audio signal to be decoded to obtain a decoded signal of the current frame.

[0192] Step 706: Decode the current frame of the multi-channel audio signal to be decoded according to the K channel pair indices and the energy / amplitude balanced side information of the K channel pairs to obtain a decoded signal of the current frame.

[0193] Taking a 5.1-channel signal as an example, the M1, S1, M2, S2, LFE, and C channel signals are decoded to obtain the L, R, LS, RS, LFE, and C channel signals. During the decoding process, energy / amplitude de-equalization is performed based on the energy / amplitude balanced side information of the K channel pairs.

[0194] In some embodiments, the side information of energy / amplitude balance of a channel pair may include a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the channel pair. For detailed explanations, please refer to the explanations of the aforementioned encoding embodiment and will not be repeated here.

[0195] In this embodiment, the current frame of the multi-channel audio signal to be decoded and the number of channel pairs included in the current frame are obtained by demultiplexing the bitstream to be decoded. When the number of channel pairs is greater than 0, the current frame is further parsed to obtain K channel pair indices and energy / amplitude balance side information for the K channel pairs. Based on the K channel pair indices and the energy / amplitude balance side information for the K channel pairs, the current frame of the multi-channel audio signal to be decoded is decoded to obtain a decoded signal of the current frame. Because the bitstream sent by the encoder does not carry energy / amplitude balance side information for unpaired channels, the number of bits of energy / amplitude balance side information in the encoded bitstream can be reduced, thereby reducing the number of bits of multi-channel side information. The saved bits can be allocated to other functional modules of the encoder to improve the quality of the audio signal reconstructed by the decoder.

[0196] The following embodiment takes a 5.1-channel signal as an example to schematically illustrate the multi-channel audio signal decoding method of the embodiment of the present application.

[0197] Figure 8 This is a schematic diagram of the processing process of the decoding end of the embodiment of the present application, such as Figure 8 As shown, the decoding end may include a code stream demultiplexing interface 801, a channel decoding unit 802 and a multi-channel decoding processing unit 803. The decoding process of this embodiment is as described above. Figure 4 and Figure 5 The inverse of the encoding process of the illustrated embodiment.

[0198] The code stream demultiplexing interface 801 is used to demultiplex the code stream output by the encoding end to obtain six encoded audio channels E1-E6.

[0199] The channel decoding unit 802 is used to perform inverse entropy coding and inverse quantization on the coded channels E1-E6 to obtain a multi-channel signal, including the center channel M1 and the side channel S1 of the first channel pair, the center channel M2 and the side channel S2 of the second channel pair, and the unpaired C channel and LFE channel. The channel decoding unit 802 also decodes to obtain multi-channel side information. The multi-channel side information includes the above Figure 4 The illustrated embodiment includes side information generated during the channel encoding process (e.g., entropy coded side information), and side information generated during the multi-channel encoding process (e.g., energy / amplitude balancing side information of channel pairs).

[0200] The multi-channel decoding unit 803 performs multi-channel decoding on the center channel M1 and side channel S1 of the first channel pair, and the center channel M2 and side channel S2 of the second channel pair. Using multi-channel side information, the center channel M1 and side channel S1 of the first channel pair are decoded into the L channel and the R channel, and the center channel M2 and side channel S2 of the second channel pair are decoded into the LS channel and the RS channel. The L channel, R channel, LS channel, RS channel, unpaired C channel, and LFE channel constitute the output of the decoder.

[0201] Figure 9 This is a schematic diagram of the processing process of the multi-channel decoding processing unit of the embodiment of the present application, as shown in FIG. Figure 9 As shown, the multi-channel decoding processing unit 803 may include a multi-channel screening unit 8031 ​​and a multi-channel decoding processing submodule 8032. The multi-channel encoding processing submodule 8032 includes two stereo decoding boxes, an energy / amplitude de-equalization unit 8033 and an energy / amplitude de-equalization unit 8034.

[0202] The multi-channel screening unit 8031 ​​screens out the M1 channel, S1 channel, M2 channel, and S2 channel participating in multi-channel processing from the 5.1 input channels (M1 channel, S1 channel, C channel, M2 channel, S2 channel, and LFE channel) based on the number of channel pairs and the channel pair index in the multi-channel side information.

[0203] The stereo decoding box in the multi-channel decoding processing submodule 8032 is used to perform the following steps: instruct the stereo decoding box to decode the first channel pair (M1, S1) into L e Vocal channel and R e Channel. According to the stereo side information of the second channel pair, the stereo decoding box is instructed to decode the second channel pair (M2, S2) into LS e Channel and RS e Vocal channel.

[0204] The energy / amplitude de-equalization unit 8033 is configured to perform the following steps: guiding the first channel pair de-equalization unit according to the energy / amplitude side information of the first channel pair, e Vocal channel and R e The energy / amplitude of the channels is de-equalized to restore the L channel and the R channel. The energy / amplitude de-equalization unit 8034 is used to perform the following steps: instruct the first channel pair de-equalization unit to convert the LS e Channel, RS e The channels are restored to LS channel and RS channel.

[0205] The process of multi-channel side information decoding is explained. Figure 10This is a flowchart of a multi-channel side information analysis embodiment of the present application. This embodiment is the above Figure 6 The reverse process of the embodiment shown, such as Figure 10 As shown, step 701 parses the bitstream to obtain the number of channel pairs in the current frame. For example, the number of channel pairs, currPairCnt, occupies 4 bits in the bitstream. Step 702 determines whether the number of channel pairs in the current frame is zero. If so, the process ends; if not, the process proceeds to step 703. If the number of channel pairs, currPairCnt, is zero, it indicates that the current frame has not been paired, and therefore, there is no need to parse and obtain side information for energy / amplitude balance. If the number of channel pairs, currPairCnt, is not zero, the energy / amplitude balance side information for the first, ..., currPairCnt-th channel pairs is cyclically parsed. For example, the variable pair is set to 0. Then, subsequent steps 703 to 707 are executed. Step 703 determines whether pair is less than the number of channel pairs. If so, the process proceeds to step 704; if not, the process ends. Step 704 parses the index of the i-th channel pair from the bitstream. i = pair + 1. Step 705: Parse the fixed-point energy / amplitude scaling ratio of the i-th channel pair from the bitstream. For example, PairILDScale[pair][0] and PairILDScale[pair][1]. Step 706: Parse the energy / amplitude scaling flag of the i-th channel pair from the bitstream. For example, energyBigFlag[pair][0] and energyBigFlag[pair][1]. Step 707: Parse the stereo side information of the i-th channel pair from the bitstream, where pair = pair + 1. Return to step 703 and continue until all channel pair indices, fixed-point energy / amplitude scaling ratios, and energy / amplitude scaling flags are parsed.

[0206] The side information parsing process of the first channel pair and the second channel pair is described using the encoding end 5.1 (L, R, C, LFE, LS, RS) signal as an example.

[0207] The side information parsing process for the first channel pair is as follows: parse the 4-bit channel pair index PairIndex[0] from the bitstream and map it to the L channel and R channel according to the definition rules of the channel pair index. parse the L channel's fixed-point energy / amplitude scaling ratio PairILDScale[0][0] and the R channel's fixed-point energy / amplitude scaling ratio PairILDScale[0][1] from the bitstream. parse the L channel's energy / amplitude scaling flag energyBigFlag[0][0] and the R channel's energy / amplitude scaling flag energyBigFlag[0][1] from the bitstream. parse the stereo side information for the first channel pair from the bitstream. Side information parsing for the first channel pair is complete.

[0208] The side information parsing process for the second channel pair is as follows: parse the 4-bit channel pair index PairIndex[1] from the bitstream and map it to the LS channel and RS channel according to the definition rule of the channel pair index. Parse the fixed-point energy / amplitude scaling ratio PairILDScale[1][0] of the LS channel and the fixed-point energy / amplitude scaling ratio PairILDScale[1][1] of the RS channel from the bitstream. Parse the energy / amplitude scaling flag energyBigFlag[1][0] of the LS channel and the energyBigFlag[1][1] of the RS channel from the bitstream. Parse the stereo side information of the second channel pair from the bitstream. The side information parsing of the second channel pair is completed.

[0209] The energy / amplitude de-equalization unit 8033 is used to convert the L e Vocal channel and R e The process of energy / amplitude equalization of the channels is as follows:

[0210] The floating-point energy / amplitude scaling coefficient scaleF_L of the L channel is calculated based on the fixed-point energy / amplitude scaling factor PairILDScale[0][0] of the L channel and the energy / amplitude scaling flag energyBigFlag[0][0] of the L channel. If the energy / amplitude scaling flag energyBigFlag[0][0] of the L channel is 1, scaleF_L = (1<<4) / PairILDScale[0][0]; if the energy / amplitude scaling flag energyBigFlag[0][0] of the L channel is 0, scaleF_L = PairILDScale[0][0] / (1<<4).

[0211] The frequency domain coefficient of the L channel after energy / amplitude de-equalization is obtained according to the floating point energy / amplitude scaling coefficient scaleF_L of the L channel. e (i)×scaleF_L; where i is used to identify the coefficient of the current frame, L(i) is the i-th frequency domain coefficient of the current frame before energy / amplitude equalization, and L e (i) is the i-th frequency domain coefficient of the current frame after energy / amplitude equalization.

[0212] The floating-point energy / amplitude scaling coefficient scaleF_R of the R channel is calculated based on the fixed-point energy / amplitude scaling coefficient PairILDScale[0][1] of the R channel and the energy / amplitude scaling flag energyBigFlag[0][1] of the R channel. If the energy / amplitude scaling flag energyBigFlag[0][1] of the R channel is 1, scaleF_R = (1<<4) / PairILDScale[0][1]; if the energy / amplitude scaling flag energyBigFlag[0][1] of the R channel is 0, scaleF_R = PairILDScale[0][1] / (1<<4).

[0213] The frequency domain coefficient of the R channel after energy / amplitude de-equalization is obtained according to the floating point energy / amplitude scaling coefficient scaleF_R of the R channel. e (i)×scaleF_R; where i is used to identify the coefficient of the current frame, L(i) is the i-th frequency domain coefficient of the current frame before energy / amplitude equalization, and L e (i) is the i-th frequency domain coefficient of the current frame after energy / amplitude equalization.

[0214] The energy / amplitude de-equalization unit 8034 is used to convert the second channel pair LS e Channel and RS e The energy / amplitude of the channels is balanced, and the specific implementation is the same as that of the first channel pair L e Vocal channel and R e The energy / amplitude of the sound channels are balanced and consistent, which will not be explained here.

[0215] The output of the multi-channel decoding processing unit 803 is the decoded L channel signal, R channel signal, LS channel signal, RS channel signal, C channel signal and LFE channel signal.

[0216] In this embodiment, since the bitstream sent by the encoder does not carry side information regarding energy / amplitude balance of unpaired channels, the number of bits of the energy / amplitude balance side information in the encoded bitstream can be reduced, thereby reducing the number of bits of multi-channel side information. The saved bits can be allocated to other functional modules of the encoder to improve the quality of the audio signal reconstructed by the decoder.

[0217] Based on the same inventive concept as the above method, an embodiment of the present application further provides an audio signal encoding device, which can be applied to an audio encoder.

[0218] Figure 11 This is a structural diagram of an audio signal encoding device according to an embodiment of the present application. Figure 11As shown, the audio signal encoding device 1100 includes: an acquisition module 1101 , an equalization side information generation module 1102 , and an encoding module 1103 .

[0219] An acquisition module 1101 is configured to obtain audio signals of P channels and respective energies / amplitudes of audio signals of P channels of a current frame of a multi-channel audio signal, where P is a positive integer greater than 1, the P channels include K channel pairs, each channel pair includes two channels, K is a positive integer, and P is greater than or equal to K*2.

[0220] The balanced side information generation module 1102 is configured to generate energy / amplitude balanced side information for K channel pairs based on the energy / amplitude of the audio signals of the P channels;

[0221] The encoding module 1103 is configured to encode the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain an encoded bitstream.

[0222] In some embodiments, the K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, where the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance. The energy / amplitude scaling identifier is used to identify whether the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

[0223] In some embodiments, the K channel pairs include a current channel pair, and the equalization side information generation module 1102 is configured to determine the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization. Energy / amplitude equalization side information for the current channel pair is generated based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization and the energy / amplitude of the audio signals of the two channels after energy / amplitude equalization.

[0224] In some embodiments, the current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio of the first channel, a fixed-point energy / amplitude scaling ratio of the second channel, an energy / amplitude scaling identifier of the first channel, and an energy / amplitude scaling identifier of the second channel.

[0225] In some embodiments, the equalization side information generation module 1102 is configured to: determine an energy / amplitude scaling coefficient for the audio signal of the qth channel of the current channel pair based on the energy / amplitude of the qth channel before energy / amplitude equalization and the energy / amplitude of the audio signal of the qth channel after energy / amplitude equalization; determine a fixed-point energy / amplitude scaling ratio for the qth channel based on the energy / amplitude scaling coefficient for the qth channel; and determine an energy / amplitude scaling identifier for the qth channel based on the energy / amplitude of the qth channel before energy / amplitude equalization and the energy / amplitude of the qth channel after energy / amplitude equalization, where q is one or two.

[0226] In some embodiments, the equalization side information generation module 1102 is configured to determine an average energy / amplitude of the audio signals of the current channel pair based on the energy / amplitude of each of the audio signals of the two channels of the current channel pair before energy / amplitude equalization, and determine an average energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the average energy / amplitude of the audio signals of the current channel pair.

[0227] In some embodiments, the encoding module 1103 is configured to encode the energy / amplitude balanced side information of the K channel pairs, the channel pair indexes corresponding to the K and K channel pairs, and the audio signals of the P channels to obtain an encoded bitstream.

[0228] It should be noted that the acquisition module 1101 , the equalization side information generation module 1102 , and the encoding module 1103 can be applied to the audio signal encoding process at the encoding end.

[0229] It should also be noted that the specific implementation process of the acquisition module 1101, the equalization side information generation module 1102, and the encoding module 1103 can refer to the detailed description of the encoding method in the above method embodiment, and for the sake of brevity, it will not be repeated here.

[0230] Based on the same inventive concept as the above method, an embodiment of the present application provides an audio signal encoder, which is used to encode an audio signal, including: an encoder as described in one or more of the above embodiments, wherein the audio signal encoding device is used to encode and generate a corresponding bit stream.

[0231] Based on the same inventive concept as the above method, the embodiment of the present application provides a device for encoding an audio signal, for example, an audio signal encoding device, see Figure 12 As shown, the audio signal encoding device 1200 includes:

[0232] Processor 1201, memory 1202 and communication interface 1203 (wherein the number of processors 1201 in the audio signal encoding device 1200 can be one or more, Figure 12 In some embodiments of the present application, the processor 1201, the memory 1202, and the communication interface 1203 may be connected via a bus or other means, wherein: Figure 12 The bus connection is taken as an example.

[0233] Memory 1202 may include read-only memory and random access memory, and provides instructions and data to processor 1201. A portion of memory 1202 may also include non-volatile random access memory (NVRAM). Memory 1202 stores an operating system and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0234] Processor 1201 controls the operation of the audio encoding device and may also be referred to as a central processing unit (CPU). In specific applications, the various components of the audio encoding device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0235] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1201. Processor 1201 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1201. The above processor 1201 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in storage media such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other well-known storage media in the art. The storage medium is located in the memory 1202 , and the processor 1201 reads the information in the memory 1202 and completes the steps of the above method in combination with its hardware.

[0236] The communication interface 1203 may be used to receive or send digital or character information, and may be, for example, an input / output interface, a pin, or a circuit, etc. For example, the above-mentioned encoded code stream is sent via the communication interface 1203 .

[0237] Based on the same inventive concept as the above method, an embodiment of the present application provides an audio encoding device, comprising: a non-volatile memory and a processor coupled to each other, wherein the processor calls program code stored in the memory to execute some or all steps of the multi-channel audio signal encoding method described in one or more of the above embodiments.

[0238] Based on the same inventive concept as the above method, an embodiment of the present application provides a computer-readable storage medium, which stores program code, wherein the program code includes instructions for executing some or all steps of the multi-channel audio signal encoding method described in one or more embodiments above.

[0239] Based on the same inventive concept as the above method, an embodiment of the present application provides a computer program product. When the computer program product is run on a computer, the computer is caused to perform some or all steps of the multi-channel audio signal encoding method as described in one or more of the above embodiments.

[0240] Based on the same inventive concept as the above method, an embodiment of the present application further provides an audio signal decoding device, which can be applied to an audio decoder.

[0241] Figure 13 This is a structural diagram of an audio signal decoding device according to an embodiment of the present application. Figure 13 As shown, the audio signal decoding device 1300 includes: an acquisition module 1301 , a demultiplexing module 1302 , and a decoding module 1303 .

[0242] The acquisition module 1301 is used to acquire the code stream to be decoded.

[0243] a demultiplexing module 1302 configured to demultiplex a code stream to be decoded to obtain a current frame of a multi-channel audio signal to be decoded, the number K of channel pairs included in the current frame, channel pair indices corresponding to the K channel pairs, and side information of energy / amplitude balance of the K channel pairs;

[0244] The decoding module 1303 is configured to decode a current frame of the multi-channel audio signal to be decoded based on the channel pair indexes corresponding to the K channel pairs and the side information of the energy / amplitude balance of the K channel pairs to obtain a decoded signal of the current frame, where K is a positive integer and each channel pair includes two channels.

[0245] In some embodiments, the K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, where the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance. The energy / amplitude scaling identifier is used to identify whether the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

[0246] In some embodiments, the K channel pairs include a current channel pair, and the decoding module 1303 is configured to: perform stereo decoding processing on a current frame of the multi-channel audio signal to be decoded based on a channel pair index corresponding to the current channel pair to obtain audio signals of two channels of the current channel pair in the current frame; and perform energy / amplitude de-equalization processing on the audio signals of the two channels of the current channel pair based on energy / amplitude balanced side information of the current channel pair to obtain decoded signals of the two channels of the current channel pair.

[0247] In some embodiments, the current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio of the first channel, a fixed-point energy / amplitude scaling ratio of the second channel, an energy / amplitude scaling identifier of the first channel, and an energy / amplitude scaling identifier of the second channel.

[0248] It should be noted that the acquisition module 1301 , the demultiplexing module 1302 , and the decoding module 1303 can be applied to the audio signal decoding process at the decoding end.

[0249] It should also be noted that the specific implementation process of the acquisition module 1301, the demultiplexing module 1302, and the decoding module 1303 can refer to the detailed description of the decoding method in the above method embodiment. For the sake of brevity of the description, it will not be repeated here.

[0250] Based on the same inventive concept as the above method, an embodiment of the present application provides an audio signal decoder, which is used to decode an audio signal, including: executing the decoder as described in one or more of the above embodiments, wherein the audio signal decoding device is used to decode and generate a corresponding bit stream.

[0251] Based on the same inventive concept as the above method, the embodiment of the present application provides a device for decoding an audio signal, for example, an audio signal decoding device, see Figure 14 As shown, the audio signal decoding device 1400 includes:

[0252] Processor 1401, memory 1402 and communication interface 1403 (wherein the number of processors 1401 in the audio signal decoding device 1400 can be one or more, Figure 14 In some embodiments of the present application, the processor 1401, the memory 1402, and the communication interface 1403 may be connected via a bus or other means, wherein: Figure 14 The bus connection is taken as an example.

[0253] Memory 1402 may include read-only memory and random access memory, and provides instructions and data to processor 1401. A portion of memory 1402 may also include non-volatile random access memory (NVRAM). Memory 1402 stores an operating system and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0254] Processor 1401 controls the operation of the audio decoding device and may also be referred to as a central processing unit (CPU). In specific applications, the various components of the audio decoding device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0255] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1401. Processor 1401 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1401. The above processor 1401 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in storage media well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1402 , and the processor 1401 reads the information in the memory 1402 and completes the steps of the above method in combination with its hardware.

[0256] The communication interface 1403 may be used to receive or send digital or character information, and may be, for example, an input / output interface, a pin, or a circuit, etc. For example, the above-mentioned encoded code stream is received via the communication interface 1403 .

[0257] Based on the same inventive concept as the above method, an embodiment of the present application provides an audio decoding device, comprising: a non-volatile memory and a processor coupled to each other, wherein the processor calls program code stored in the memory to execute some or all steps of the multi-channel audio signal decoding method described in one or more of the above embodiments.

[0258] Based on the same inventive concept as the above method, an embodiment of the present application provides a computer-readable storage medium, which stores program code, wherein the program code includes instructions for executing some or all steps of the multi-channel audio signal decoding method described in one or more embodiments above.

[0259] Based on the same inventive concept as the above method, an embodiment of the present application provides a computer program product. When the computer program product is run on a computer, the computer is caused to perform some or all steps of the multi-channel audio signal decoding method as described in one or more embodiments above.

[0260] The processor mentioned in each of the above embodiments can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware coding processor, or can be executed by a combination of hardware and software modules in the coding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0261] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0262] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0263] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0264] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0265] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0266] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0267] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0268] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A multi-channel audio signal encoding method, characterized in that: include: Obtain audio signals of P channels of a current frame of a multi-channel audio signal, where P is a positive integer greater than 1, the P channels include K channel pairs, each channel pair includes two channels, K is a positive integer, and P is greater than or equal to K*2; Obtaining the energy / amplitude of each of the P channels of audio signals; performing energy / amplitude equalization processing on the audio signals of two channels of each of the K channel pairs based on the energy / amplitude of each of the P channels' audio signals, and generating energy / amplitude-balanced side information for the K channel pairs, where the energy / amplitude-balanced side information for the K channel pairs is used to indicate a scaling ratio of the energy / amplitude of the K channel pairs; Encoding the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain a coded bitstream; The K channel pairs include the current channel pair, and the performing energy / amplitude balancing processing on the audio signals of two channels of each of the K channel pairs based on the energy / amplitude of each of the audio signals of the P channels, to generate energy / amplitude balanced side information for the K channel pairs includes: determining, based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization, the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization; Energy / amplitude balanced side information of the current channel pair is generated based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance.

2. The method according to claim 1, characterized in that The K channel pairs include a current channel pair, and the energy / amplitude balanced side information of the current channel pair includes: The fixed-point energy / amplitude scaling ratio and energy / amplitude scaling identifier of the current channel pair, wherein the fixed-point energy / amplitude scaling ratio is a fixed-point value of the energy / amplitude scaling coefficient, the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization and the energy / amplitude of the audio signals of the two channels after energy / amplitude equalization, and the energy / amplitude scaling identifier is used to identify whether the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization is amplified or reduced relative to the energy / amplitude before energy / amplitude equalization.

3. The method according to claim 1, characterized in that The current channel pair includes a first channel and a second channel, and the energy / amplitude balanced side information of the current channel pair includes: A fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the first channel, and a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the second channel.

4. The method according to claim 3, characterized in that Generating energy / amplitude balanced side information of the current channel pair based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance, includes: determining an energy / amplitude scaling coefficient of the qth channel and an energy / amplitude scaling flag of the qth channel according to the energy / amplitude of the audio signal of the qth channel of the current channel pair before energy / amplitude equalization and the energy / amplitude of the audio signal of the qth channel after energy / amplitude equalization; determining a fixed-point energy / amplitude scaling ratio of the qth channel according to an energy / amplitude scaling coefficient of the qth channel; Here, q is one or two.

5. The method according to any one of claims 1 to 4, characterized in that The determining, based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization, the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization comprises: Determine the average energy / amplitude of the audio signals of the current channel pair based on the energy / amplitude of each of the two channels of the current channel pair before energy / amplitude equalization; and determine the energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the average energy / amplitude of the audio signals of the current channel pair.

6. The method according to any one of claims 1 to 4, characterized in that The step of encoding the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain a coded bitstream includes: The energy / amplitude balanced side information of the K channel pairs, K, the channel pair indexes corresponding to the K channel pairs, and the audio signals of the P channels are encoded to obtain the encoded bitstream.

7. A multi-channel audio signal decoding method, characterized in that: include: Get the code stream to be decoded; Demultiplexing the to-be-decoded bitstream to obtain a current frame of a to-be-decoded multi-channel audio signal, the current frame including K number of channel pairs, channel pair indexes corresponding to the K channel pairs, and energy / amplitude balanced side information for the K channel pairs, the energy / amplitude balanced side information for the K channel pairs being used to indicate energy / amplitude scaling ratios for the K channel pairs, where K is a positive integer and each channel pair includes two channels; the energy / amplitude balanced side information for the K channel pairs is generated by performing energy / amplitude balanced processing on audio signals of two channels of each of the K channel pairs; Decoding a current frame of the to-be-decoded multi-channel audio signal according to channel pair indexes corresponding to the K channel pairs and side information of energy / amplitude balance of the K channel pairs to obtain a decoded signal of the current frame; The K channel pairs include the current channel pair, and the energy / amplitude balance side information of the current channel pair is generated based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance; the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is determined based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance.

8. The method according to claim 7, characterized in that The K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, wherein the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of each of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of each of the audio signals of the two channels after energy / amplitude balance; the energy / amplitude scaling identifier is used to identify whether the energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

9. The method according to claim 8, characterized in that The current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the first channel, and a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the second channel.

10. The method according to any one of claims 7 to 9, characterized in that The K channel pairs include a current channel pair, and decoding a current frame of the to-be-decoded multi-channel audio signal according to channel pair indexes corresponding to the K channel pairs and energy / amplitude balanced side information of the K channel pairs to obtain a decoded signal of the current frame, including: performing stereo decoding processing on a current frame of the multi-channel audio signal to be decoded according to a channel pair index corresponding to the current channel pair, so as to obtain audio signals of two channels of the current channel pair of the current frame; Energy / amplitude de-equalization processing is performed on the audio signals of the two channels of the current channel pair according to the energy / amplitude balanced side information of the current channel pair to obtain decoded signals of the two channels of the current channel pair.

11. An audio signal encoding device, characterized in that: include: an acquisition module, configured to acquire audio signals of P channels of a current frame of a multi-channel audio signal and respective energies / amplitudes of the audio signals of the P channels, where P is a positive integer greater than 1, the P channels include K channel pairs, each channel pair includes two channels, K is a positive integer, and P is greater than or equal to K*2; a balanced side information generation module, configured to perform energy / amplitude equalization processing on the audio signals of two channels of each of the K channel pairs based on the energy / amplitude of each of the P channels' audio signals, and generate energy / amplitude balanced side information for the K channel pairs, wherein the energy / amplitude balanced side information for the K channel pairs is used to indicate a scaling ratio of the energy / amplitude of the K channel pairs; an encoding module, configured to encode the energy / amplitude balanced side information of the K channel pairs and the audio signals of the P channels to obtain an encoded bitstream; The K channel pairs include the current channel pair, and the equalization side information generation module is specifically configured to: determine the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization; and generate energy / amplitude equalization side information of the current channel pair based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude equalization and the energy / amplitude of the audio signals of the two channels after energy / amplitude equalization.

12. The device according to claim 11, characterized in that The K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, wherein the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of each of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of each of the audio signals of the two channels after energy / amplitude balance; the energy / amplitude scaling identifier is used to identify whether the energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

13. The device according to claim 11, characterized in that The current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the first channel, and a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the second channel.

14. The device according to claim 13, characterized in that The equalization side information generation module is configured to determine an energy / amplitude scaling coefficient of the qth channel and an energy / amplitude scaling flag of the qth channel based on the energy / amplitude of the audio signal of the qth channel of the current channel pair before energy / amplitude equalization and the energy / amplitude of the audio signal of the qth channel after energy / amplitude equalization; determining a fixed-point energy / amplitude scaling ratio of the qth channel according to an energy / amplitude scaling coefficient of the qth channel; Here, q is one or two.

15. The device according to any one of claims 11 to 14, characterized in that The equalization side information generation module is configured to determine an average energy / amplitude of the audio signals of the current channel pair based on the energy / amplitude of each of the audio signals of the two channels of the current channel pair before energy / amplitude equalization, and determine an energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude equalization based on the average energy / amplitude of the audio signals of the current channel pair.

16. The device according to any one of claims 11 to 14, characterized in that The encoding module is configured to encode the energy / amplitude balanced side information of the K channel pairs, K, the channel pair indexes corresponding to the K channel pairs, and the audio signals of the P channels to obtain the encoded bitstream.

17. An audio signal decoding device, characterized in that: include: An acquisition module is used to obtain the code stream to be decoded; a demultiplexing module, configured to demultiplex the to-be-decoded bitstream to obtain a current frame of the to-be-decoded multi-channel audio signal, the current frame including K number of channel pairs, channel pair indexes corresponding to the K channel pairs, and energy / amplitude balanced side information for the K channel pairs, the energy / amplitude balanced side information for the K channel pairs being used to indicate energy / amplitude scaling ratios for the K channel pairs, where K is a positive integer and each channel pair includes two channels; the energy / amplitude balanced side information for the K channel pairs being generated by performing energy / amplitude balanced processing on audio signals of two channels of each of the K channel pairs; a decoding module, configured to decode a current frame of the to-be-decoded multi-channel audio signal according to the channel pair indexes of the K channel pairs and the side information of the energy / amplitude balance of the K channel pairs, so as to obtain a decoded signal of the current frame; The K channel pairs include the current channel pair, and the energy / amplitude balance side information of the current channel pair is generated based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of the audio signals of the two channels after energy / amplitude balance; the energy / amplitude of the audio signals of the two channels of the current channel pair after energy / amplitude balance is determined based on the energy / amplitude of the audio signals of the two channels of the current channel pair before energy / amplitude balance.

18. The device according to claim 17, characterized in that The K channel pairs include a current channel pair, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the current channel pair, wherein the fixed-point energy / amplitude scaling ratio is a fixed-point value of an energy / amplitude scaling coefficient, and the energy / amplitude scaling coefficient is obtained based on the energy / amplitude of each of the audio signals of the two channels of the current channel pair before energy / amplitude balance and the energy / amplitude of each of the audio signals of the two channels after energy / amplitude balance; the energy / amplitude scaling identifier is used to identify whether the energy / amplitude of each of the audio signals of the two channels of the current channel pair after energy / amplitude balance is amplified or reduced relative to the energy / amplitude before energy / amplitude balance.

19. The device according to claim 18, characterized in that The current channel pair includes a first channel and a second channel, and the side information of the energy / amplitude balance of the current channel pair includes: a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the first channel, and a fixed-point energy / amplitude scaling ratio and an energy / amplitude scaling identifier of the second channel.

20. The device according to any one of claims 17 to 19, characterized in that The K channel pairs include a current channel pair, and the decoding module is configured to: performing stereo decoding processing on a current frame of the multi-channel audio signal to be decoded according to a channel pair index corresponding to the current channel pair, so as to obtain audio signals of two channels of the current channel pair of the current frame; Energy / amplitude de-equalization processing is performed on the audio signals of the two channels of the current channel pair according to the energy / amplitude balanced side information of the current channel pair to obtain decoded signals of the two channels of the current channel pair.

21. An audio signal encoding device, characterized in that: include: A non-volatile memory and a processor coupled to each other, wherein the processor calls a program code stored in the memory to execute the method according to any one of claims 1 to 6.

22. An audio signal decoding device, characterized in that: include: A non-volatile memory and a processor coupled to each other, wherein the processor calls a program code stored in the memory to execute the method according to any one of claims 7 to 10.

23. An audio signal encoding device, characterized in that The method comprises: an encoder configured to execute the method according to any one of claims 1 to 6.

24. An audio signal decoding device, characterized in that: include: A decoder configured to execute the method according to any one of claims 7 to 10.

25. A computer-readable storage medium, characterized in that The method comprises an encoded code stream obtained according to the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi sound channel AF expansion support

    CN1765072A

  • Multisignal audio coding using signal whitening as preprocessing

    WO2020007719A1

Cited By

  • Multi-channel audio signal encoding / decoding method and device

    WO2022012628A1