Encoding methods and apparatus for multi-channel audio signals

By acquiring the correlation values ​​of multi-channel signals and performing energy equalization processing, the problem of insufficient coding efficiency in multi-channel audio encoding and decoding is solved, thereby improving coding efficiency and reducing coding noise.

CN115410584BActive Publication Date: 2026-05-26HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2021-05-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing multi-channel audio codec technologies have shortcomings in coding efficiency, especially in low bitrate environments, where it is difficult to effectively reduce coding noise caused by energy differences between channel signals.

Method used

By acquiring the correlation values ​​of multi-channel signals, selecting channels with high correlation for energy equalization processing, and determining the energy equalization mode based on the fluctuation range value, the first or second energy equalization mode is used to perform energy equalization processing on the channel signals, thereby reducing coding noise and improving coding efficiency.

Benefits of technology

It improves the coding efficiency of multi-channel audio signals, avoids coding noise problems caused by insufficient bits in high-energy channels in low bit rate environments, and reduces redundant bit allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410584B_ABST
    Figure CN115410584B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for encoding multichannel audio signals. The method includes: acquiring a first audio frame to be encoded, the first audio frame including at least five channel signals; acquiring the sum of correlation values ​​for all channel pairs in a target channel pair set, the target channel pair set including at least one channel pair, each channel pair including two of the at least five channel signals, each channel pair having a correlation value representing the correlation between the two channel signals of a channel pair; when the sum of correlation values ​​is greater than a preset threshold, performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals; and encoding the at least two equalized channel signals to obtain an encoded bitstream. This application can improve the encoding efficiency of audio frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to audio processing technology, and more particularly to a method and apparatus for encoding multi-channel audio signals. Background Technology

[0002] Multichannel audio encoding and decoding are techniques used to encode or decode audio that contains two or more channels. Common multichannel audio formats include 5.1 channel audio, 7.1 channel audio, 7.1.4 channel audio, and 22.2 channel audio.

[0003] The MPEG Surround (MPS) standard specifies joint encoding for four channels, but there is still a need for encoding and decoding methods that can handle the various multi-channel audio signals mentioned above. Summary of the Invention

[0004] This application provides a method and apparatus for encoding multi-channel audio signals to improve the encoding efficiency of audio frames.

[0005] In a first aspect, this application provides a method for encoding multichannel audio signals, comprising: acquiring a first audio frame to be encoded, the first audio frame including at least five channel signals; acquiring the sum of correlation values ​​of all channel pairs in a target channel pair set, the target channel pair set including at least one channel pair, a channel pair including two channel signals from the at least five channel signals, the channel pair having a correlation value, the correlation value being used to represent the correlation between the two channel signals of the channel pair; when the sum of the correlation values ​​is greater than a preset threshold, performing energy equalization processing on at least two channel signals from the at least five channel signals to obtain at least two equalized channel signals; and encoding the at least two equalized channel signals to obtain an encoded bitstream.

[0006] In this embodiment, the audio frame contains at least five channel signals to obtain a target channel pair set in order to obtain the maximum sum of correlation values. When the sum of correlation values ​​of the target channel pair set is greater than a preset threshold, energy equalization processing is performed on at least two of the at least five channel signals, and then they are encoded, which can improve the encoding efficiency of the audio frame.

[0007] In one possible implementation, the method further includes: encoding the at least five channel signals to obtain an encoded bitstream when the sum of the correlation values ​​is less than or equal to the preset threshold.

[0008] In this embodiment, if the sum of correlation values ​​is less than or equal to a preset threshold, it indicates that the correlation between the two channel signals in the channel pair in the target channel pair set is low, there is no need for pair coding, and therefore no need to perform energy equalization processing on at least five channel signals. At this time, the object of coding is the at least five channel signals, rather than the equalized channel signals.

[0009] In one possible implementation, the step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: obtaining fluctuation range values ​​of the at least five channel signals; determining an energy equalization mode based on the fluctuation range values ​​of the at least five channel signals; and performing energy equalization processing on the at least two channel signals respectively according to the energy equalization mode to obtain the at least two equalized channel signals.

[0010] The fluctuation range value is used to represent the magnitude of the energy or amplitude difference between at least five channel signals. The energy equalization modes include a first energy equalization mode and a second energy equalization mode. The first energy equalization mode uses two channel signals from a channel pair to acquire the two equalized channel signals corresponding to that channel pair. The second energy equalization mode uses two channel signals from a channel pair and at least one external channel signal from one channel to acquire the two equalized channel signals corresponding to that channel pair.

[0011] In one possible implementation, determining the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals includes: determining the energy equalization mode as a first energy equalization mode when the fluctuation range values ​​meet preset conditions; or determining the energy equalization mode as a second energy equalization mode when the fluctuation range values ​​do not meet preset conditions.

[0012] In one possible implementation, the fluctuation range value includes the energy smoothness of the first audio frame; the fluctuation range value meets a preset condition if the energy smoothness is less than a first threshold; or, the fluctuation range value includes the amplitude smoothness of the first audio frame; the fluctuation range value meets a preset condition if the amplitude smoothness is less than a second threshold; or, the fluctuation range value includes the energy deviation of the first audio frame; the fluctuation range value meets a preset condition if the energy deviation is not within a first preset range; or, the fluctuation range value includes the amplitude deviation of the first audio frame; the fluctuation range value meets a preset condition if the amplitude deviation is not within a second preset range.

[0013] In one possible implementation, when the energy equalization mode is the first energy equalization mode, the step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: performing energy equalization processing on the channel signals corresponding to the target channel pair set to obtain the at least two equalized channel signals.

[0014] In one possible implementation, the step of performing energy equalization processing on the channel signals corresponding to the target channel pair set to obtain the at least two equalized channel signals includes: for the current channel pair in the target channel pair set, calculating the average value of the energy value or amplitude value of the two channel signals contained in the current channel pair, and performing energy equalization processing on the two channel signals contained in the current channel pair according to the average value to obtain the corresponding two equalized channel signals.

[0015] In this way, when the fluctuation range of at least five channel signals is large, energy equalization can be performed only between two relevant channel signals. This makes the bit allocation during stereo processing more consistent with the fluctuation range of the channel signals, avoiding the problem that in low bit rate coding environments, the coding noise of high-energy channel pairs may be much greater than that of low-energy channel pairs due to insufficient bits, while the low-energy channel pairs will have bit redundancy.

[0016] In one possible implementation, when the energy equalization mode is the second energy equalization mode, the step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: calculating the average value of the energy value or amplitude value of the at least five channel signals, and performing energy equalization processing on the at least five channel signals respectively according to the average value to obtain the at least five equalized channel signals.

[0017] In one possible implementation, before determining the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals, the method further includes: determining whether the encoding bitrate corresponding to the first audio frame is greater than a bitrate threshold; when the encoding bitrate is greater than the bitrate threshold, determining the energy equalization mode as a second energy equalization mode; and determining the energy equalization mode based on the fluctuation range values ​​only when the encoding bitrate is less than or equal to the bitrate threshold.

[0018] In one possible implementation, the method further includes encoding the channel signals among the at least five channel signals that have not undergone energy equalization processing.

[0019] Secondly, this application provides an encoding apparatus, comprising: an acquisition module, configured to acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; acquire the sum of correlation values ​​of all channel pairs in a target channel pair set, the target channel pair set including at least one channel pair, a channel pair including two channel signals from the at least five channel signals, the channel pair having a correlation value, the correlation value being used to represent the correlation between the two channel signals of the channel pair; a processing module, configured to perform energy equalization processing on at least two channel signals from the at least five channel signals to obtain at least two equalized channel signals when the sum of the correlation values ​​is greater than a preset threshold; and an encoding module, configured to encode the at least two equalized channel signals to obtain an encoded bitstream.

[0020] In one possible implementation, the encoding module is further configured to encode the at least five channel signals to obtain an encoded bitstream when the sum of the correlation values ​​is less than or equal to the preset threshold.

[0021] In one possible implementation, the processing module is specifically configured to acquire the fluctuation range values ​​of the at least five channel signals; determine an energy equalization mode based on the fluctuation range values ​​of the at least five channel signals; and perform energy equalization processing on the at least two channel signals according to the energy equalization mode to obtain the at least two equalized channel signals.

[0022] In one possible implementation, the processing module is specifically used to determine the energy balance mode as a first energy balance mode when the fluctuation range value meets the preset conditions; or, when the fluctuation range value does not meet the preset conditions, determine the energy balance mode as a second energy balance mode.

[0023] In one possible implementation, the fluctuation range value includes the energy smoothness of the first audio frame; the fluctuation range value meets a preset condition if the energy smoothness is less than a first threshold; or, the fluctuation range value includes the amplitude smoothness of the first audio frame; the fluctuation range value meets a preset condition if the amplitude smoothness is less than a second threshold; or, the fluctuation range value includes the energy deviation of the first audio frame; the fluctuation range value meets a preset condition if the energy deviation is not within a first preset range; or, the fluctuation range value includes the amplitude deviation of the first audio frame; the fluctuation range value meets a preset condition if the amplitude deviation is not within a second preset range.

[0024] In one possible implementation, when the energy equalization mode is the first energy equalization mode, the processing module is specifically used to perform energy equalization processing on the channel signals corresponding to the target channel pair set to obtain the at least two equalized channel signals.

[0025] In one possible implementation, the processing module is specifically used to calculate the average value of the energy value or amplitude value of the two channel signals contained in the current channel pair in the target channel pair set, and perform energy equalization processing on the two channel signals contained in the current channel pair according to the average value to obtain the corresponding two equalized channel signals.

[0026] In one possible implementation, when the energy equalization mode is the second energy equalization mode, the processing module is specifically used to calculate the average value of the energy value or amplitude value of the at least five channel signals, and perform energy equalization processing on the at least five channel signals according to the average value to obtain the at least five equalized channel signals.

[0027] In one possible implementation, the processing module is further configured to determine whether the encoding bitrate corresponding to the first audio frame is greater than a bitrate threshold; when the encoding bitrate is greater than the bitrate threshold, the energy balancing mode is determined to be a second energy balancing mode; and when the encoding bitrate is less than or equal to the bitrate threshold, the energy balancing mode is determined based on the fluctuation range value.

[0028] In one possible implementation, the encoding module is further configured to encode the channel signals among the at least five channel signals that have not undergone energy equalization processing.

[0029] Thirdly, this application provides an apparatus comprising: one or more processors; a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in any one of the first aspects above.

[0030] Fourthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the first aspects above.

[0031] Fifthly, this application provides a computer-readable storage medium including an encoded bitstream obtained according to the encoding method for multi-channel audio signals as described in any one of the first aspects above. Attached Figure Description

[0032] Figure 1 A schematic block diagram of the audio decoding system 10 used in this application is provided as an example;

[0033] Figure 2 A schematic block diagram of the audio decoding device 200 used in this application is provided as an example;

[0034] Figure 3 This is a flowchart of an exemplary embodiment of the multi-channel audio signal encoding method provided in this application;

[0035] Figure 4a This is an exemplary structural diagram of the encoding device used in the multi-channel audio signal encoding method provided in this application;

[0036] Figure 4b This is an exemplary structural diagram of a multi-channel adaptive pairing module;

[0037] Figure 4c This is an exemplary structural diagram of the group pair processing module;

[0038] Figure 5 This is an exemplary structural diagram of the decoding device used in the multi-channel audio decoding method provided in this application;

[0039] Figure 6 This is a schematic diagram of the structure of an embodiment of the encoding device of this application;

[0040] Figure 7 This is a schematic diagram of the structure of an embodiment of the device in this application. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0043] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0044] Explanation of relevant terms used in this application:

[0045] Audio frame: Audio data is streamed. In practical applications, in order to facilitate audio processing and transmission, the amount of audio data within a certain duration is usually taken as an audio frame. This duration is called the "sampling time". Its value can be determined according to the codec and the specific application requirements. For example, this duration is 2.5ms to 60ms, where ms stands for millisecond.

[0046] Audio signals: Audio signals are information carriers of the frequency and amplitude variations of sound waves, containing speech, music, and sound effects. Audio is a continuously changing analog signal, which can be represented by a continuous curve called a sound wave. Audio signals are digital signals generated by analog-to-digital conversion or computers. Sound waves have three important parameters: frequency, amplitude, and phase, which determine the characteristics of audio signals.

[0047] Audio channel signal: refers to the independent audio signals collected or played back from different spatial locations during recording or playback. Therefore, the number of audio channels is the number of sound sources during recording or the number of speakers during playback.

[0048] The following is the system architecture used in this application.

[0049] Figure 1 A schematic block diagram of the audio decoding system 10 used in this application is provided as an example. Figure 1 As shown, the audio decoding system 10 may include a source device 12 and a destination device 14. The source device 12 generates an encoded bitstream, and therefore, the source device 12 may be referred to as an audio encoding device. The destination device 14 can decode the encoded bitstream generated by the source device 12, and therefore, the destination device 14 may be referred to as an audio decoding device.

[0050] The source device 12 includes an encoder 20, and optionally may include an audio source 16, an audio preprocessor 18, and a communication interface 22.

[0051] The audio source 16 may include or may be any type of audio capture device for capturing real-world speech, music, and sound effects, and / or any type of audio generation device, such as an audio processor or device for generating speech, music, and sound effects. The audio source may be any type of memory or storage device for storing the aforementioned audio.

[0052] The audio preprocessor 18 receives (raw) audio data 17 and preprocesses the audio data 17 to obtain preprocessed audio data 19. For example, the preprocessing performed by the audio preprocessor 18 may include trimming or noise reduction. It is understood that the audio preprocessing unit 18 may be an optional component.

[0053] Encoder 20 is used to receive preprocessed audio data 19 and provide encoded audio data 21.

[0054] The communication interface 22 in the source device 12 can be used to receive encoded audio data 21 and send the encoded audio data 21 to the destination device 14 through the communication channel 13 for storage or direct reconstruction.

[0055] The target device 14 includes a decoder 30, and optionally may include a communication interface 28, an audio post-processor 32, and a playback device 34.

[0056] The communication interface 28 in the destination device 14 is used to directly receive encoded audio data 21 from the source device 12 and provide the encoded audio data 21 to the decoder 30.

[0057] Communication interfaces 22 and 28 can be used to send or receive encoded audio data 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.

[0058] For example, the communication interface 22 can be used to encapsulate the encoded audio data 21 into a suitable format such as a message, and / or process the encoded audio data 21 using any type of transmission encoding or processing, so as to transmit it on a communication link or communication network.

[0059] Communication interface 28 corresponds to communication interface 22. For example, it can be used to receive transmitted data and process and / or decapsulate the transmitted data using any type of corresponding transmission decoding or processing to obtain encoded audio data 21.

[0060] Both communication interface 22 and communication interface 28 can be configured as follows: Figure 1The arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14 indicates a one-way or two-way communication interface, which can be used to send and receive messages, establish connections, acknowledge and exchange any other information related to the data transmission of communication links and / or encoded audio data, etc.

[0061] Decoder 30 is used to receive encoded audio data 21 and provide decoded audio data 31.

[0062] The audio post-processor 32 is used to post-process the decoded audio data 31 to obtain post-processed audio data 33. The post-processing performed by the audio post-processor 32 may include, for example, trimming or resampling.

[0063] Playback device 34 is used to receive post-processed audio data 33 to play audio to a user or listener. Playback device 34 can be or include any type of player for playing reconstructed audio, such as an integrated or external speaker. For example, a speaker may include a loudspeaker, a stereo, etc.

[0064] Figure 2 A schematic block diagram of the audio decoding device 200 used in this application is provided as an example. In one embodiment, the audio decoding device 200 may be an audio decoder (e.g., Figure 1 decoder 30) or audio encoder (e.g. Figure 1 The encoder 20).

[0065] The audio decoding device 200 includes: an input port 210 and a receiving unit (Rx) 220 for receiving data; a processor, logic unit, or central processing unit 230 for processing data; a transmitting unit (Tx) 240 and an output port 250 for transmitting data; and a memory 260 for storing data. The audio decoding device 200 may also include photoelectric conversion components and electro-optical (EO) components coupled to the input port 210, the receiving unit 220, the transmitting unit 240, and the output port 250 for the input or output of optical or electrical signals.

[0066] Processor 230 is implemented in both hardware and software. Processor 230 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 230 communicates with ingress port 210, receiver unit 220, transmitter unit 240, egress port 250, and memory 260. Processor 230 includes a decoding module 270 (e.g., an encoding module or a decoding module). Decoding module 270 implements the embodiments disclosed in this application to implement the multi-channel audio signal encoding method provided in this application. For example, decoding module 270 implements, processes, or provides various encoding operations. Therefore, decoding module 270 provides a substantial improvement to the functionality of audio decoding device 200 and affects the transitions of audio decoding device 200 to different states. Alternatively, decoding module 270 can be implemented with instructions stored in memory 260 and executed by processor 230.

[0067] Memory 260 includes one or more disks, tape drives, and solid-state drives, which can be used as overflow data storage devices to store programs while they are selectively executed, and to store instructions and data read during program execution. Memory 260 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).

[0068] Based on the description of the above embodiments, this application provides a method for encoding multi-channel audio signals.

[0069] Figure 3 This is a flowchart of an exemplary embodiment of the multi-channel audio signal encoding method provided in this application. This process 300 can be performed by the source device 12 or the audio decoding device 200 in the audio decoding system 10. Process 300 is described as a series of steps or operations; it should be understood that process 300 can be performed in various orders and / or occur simultaneously, and is not limited to... Figure 3 The execution order is shown. Figure 3 As shown, the method includes:

[0070] Step 301: Obtain the first audio frame to be encoded.

[0071] The first audio frame in this embodiment can be any frame of the multi-channel audio to be encoded, and the first audio frame includes five or more channel signals. For example, a 5.1 channel system includes six channel signals: center channel (C), front left channel (left, L), front right channel (right, R), rear left surround channel (left surround, LS), rear right surround channel (right surround, RS), and 0.1 channel low frequency effects (LFE). A 7.1 channel system includes eight channel signals: C, L, R, LS, RS, LB, RB, and LFE. The LFE channel is an audio channel from 3-120Hz, which is typically sent to speakers specifically designed for bass.

[0072] Step 302: Obtain the sum of the correlation values ​​of all channel pairs in the target channel pair set.

[0073] The target channel pair set is obtained with the aim of obtaining the maximum sum of correlation values. This target channel pair set includes at least one channel pair, where each channel pair comprises two channel signals from at least five channel signals. Each channel pair has a correlation value that represents the correlation between the two channel signals of that channel pair.

[0074] Encoding two channel signals with high correlation together can reduce redundancy and improve coding efficiency. Therefore, in this embodiment, the pairing is determined based on the correlation value between the two channel signals. To find the pairing method with the highest correlation, the correlation values ​​between each pair of at least five channel signals in the first audio frame can be calculated to obtain the correlation value set of the first audio frame. For example, five channel signals can form a total of 10 channel pairs, and the corresponding correlation value set can include 10 correlation values.

[0075] Optionally, the correlation values ​​can be normalized so that the correlation values ​​of all channel pairs are limited to a specific range, which facilitates the setting of a unified judgment standard for the correlation values, such as a pair threshold. This pair threshold can be set to a value greater than or equal to 0.2 and less than or equal to 1, for example, 0.3. In this way, as long as the normalized correlation value of two channel signals is less than the pair threshold, the correlation between the two channel signals is considered to be poor, and pair coding is not required.

[0076] In one possible implementation, the correlation value between two audio channel signals (e.g., ch1 and ch2) can be calculated using the following formula:

[0077]

[0078] Where corr(ch1,ch2) represents the normalized correlation value between channel signals ch1 and ch2, spec_ch1(i) represents the frequency domain coefficient of the i-th frequency point of channel signal ch1, spec_ch2(i) is the frequency domain coefficient of the i-th frequency point of channel signal ch2, and N represents an integer value not exceeding the total number of frequency points of an audio frame.

[0079] It should be noted that other algorithms or formulas can also be used to calculate the correlation value between the two channel signals, and this application does not make any specific limitations on this.

[0080] The method for obtaining the target channel pair set includes: selecting channel pairs from channel pairs corresponding to at least five channel signals and adding them to the target channel pair set with the aim of obtaining the maximum sum of correlation values. The sum of correlation values ​​of the target channel pair set is the sum of correlation values ​​of all channel pairs in the target channel pair set obtained by pairing at least five channel signals according to the above pairing method. The pairing method in this embodiment can include the following two implementation methods:

[0081] (1) Select the M largest correlation values ​​from the correlation value set. These M correlation values ​​must be greater than or equal to the pairing threshold. This is because correlation values ​​less than the pairing threshold indicate that the correlation between the two channel signals in the corresponding channel pair is low, and there is no need for pairing encoding. In order to improve encoding efficiency, it is not necessary to select all correlation values ​​greater than or equal to the pairing threshold. Therefore, an upper limit N of M is set, that is, at most N correlation values ​​can be selected.

[0082] N can be an integer greater than or equal to 2, and the maximum value of N cannot exceed the number of all channel pairs corresponding to all channel signals in the first audio frame. The larger the value of N, the more computational effort is required, while the smaller the value of N, the more likely channel pairs will be lost, thus reducing coding efficiency.

[0083] Optionally, N can be set to the maximum number of channels plus one, i.e. CH represents the number of channels in the first audio frame. For example, a 5.1 channel audio frame contains five channels, so N = 3; a 7.1 channel audio frame contains seven channels, so N = 4.

[0084] Then, based on the M correlation values, M channel pair sets are obtained. Each channel pair set includes at least one of the M channel pairs corresponding to the M correlation values, and when a channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same channel signal. For example, for a 5.1 channel, the three channel pairs corresponding to the largest correlation value selected from the correlation value set are (L,R), (R,C), and (LS,RS). Among them, the correlation value of (LS,RS) is less than the pairing threshold, so it is excluded. Then, the remaining two channel pairs (L,R) and (R,C) can yield two channel pair sets, one of which includes (L,R), and the other includes (R,C).

[0085] Taking any one of the M channel pairs (e.g., the first channel pair) corresponding to M correlation values ​​greater than or equal to the pairing threshold as an example, the method for obtaining the set of M channel pairs in this embodiment may include: adding the first channel pair to the target channel pair set, the set of M channel pairs including the target channel pair set, and when the channel pairs other than the associated channel pair include channel pairs with correlation values ​​greater than the pairing threshold, the channel pair with the largest correlation value is selected from the other channel pairs and added to the target channel pair set, the associated channel pair including any one of the channel signals included in the channel pair already added to the target channel pair set.

[0086] Except for the step of adding the first channel pair to the target channel pair set, the above process consists entirely of iterative steps.

[0087] a. Determine whether, among multiple channel pairs, excluding the associated channel, there are channel pairs with a correlation value greater than the pair threshold.

[0088] b. If the channel pairs include those with a correlation value greater than the pair threshold, then select the channel pair with the largest correlation value from the other channel pairs and add it to the target channel pair set.

[0089] At this point, as long as other channel pairs include channel pairs with correlation values ​​greater than the pair threshold, step b above can be executed iteratively.

[0090] Optionally, to reduce computation, relevant values ​​smaller than the pair threshold can be removed from the relevant value set, which reduces the number of channel pairs and thus the number of iterations.

[0091] (2) Obtain the set of all channel pairs corresponding to at least five channel signals based on multiple channel pairs, obtain the sum of the correlation values ​​of all channel pairs contained in any channel pair set based on the correlation value set, and determine the channel pair set corresponding to the largest sum of correlation values ​​in all channel pair sets as the target channel pair set.

[0092] The correlation value set includes the correlation values ​​of multiple channel pairs of at least five channel signals of the first audio frame. By combining these multiple channel pairs in a regular manner (i.e., multiple channel pairs in the same channel pair set cannot contain the same channel signal), a set of multiple channel pairs corresponding to the at least five channel signals can be obtained.

[0093] In one possible implementation, when the number of channel signals is odd, the number of all channel pairs can be calculated using the following formula:

[0094]

[0095] In one possible implementation, when the number of channel signals is even, the number of all channel pairs can be calculated using the following formula:

[0096]

[0097] Where Pair_num represents the number of all channel pairs, and CH represents the number of channel signals participating in multichannel processing in the first audio frame, which is the result after multichannel masking.

[0098] Optionally, to reduce computational load, after obtaining the set of relevant values, multiple channel pair sets can be obtained based on the other channel pairs among the multiple channel pairs, where the correlation value of the non-relevant channel pair is less than the pairing threshold. This reduces the number of channel pairs involved in the calculation when obtaining the channel pair sets, thereby reducing the number of channel pair sets. In subsequent steps, this also reduces the computational load of the sum of correlation values.

[0099] Step 303: When the sum of the correlation values ​​is greater than a preset threshold, perform energy equalization processing on at least two of the five channel signals to obtain at least two equalized channel signals.

[0100] In one possible implementation, the fluctuation range values ​​of at least five channel signals can be obtained first, then the energy equalization mode can be determined based on the fluctuation range values ​​of at least five channel signals, and then the energy equalization processing of at least five channel signals can be performed on each of the energy equalization modes to obtain at least five equalized channel signals.

[0101] The fluctuation range value is used to represent the magnitude of the difference in energy or amplitude between at least five channel signals.

[0102] The energy equalization modes include a first energy equalization mode and a second energy equalization mode. The first energy equalization mode uses two channel signals from a channel pair to acquire the two equalized channel signals corresponding to that channel pair. The second energy equalization mode uses two channel signals from a channel pair and at least one external channel signal from one channel to acquire the two equalized channel signals corresponding to that channel pair.

[0103] Determining the energy equalization mode based on the fluctuation range values ​​of at least five channel signals may include: when the fluctuation range values ​​meet preset conditions, determining the energy equalization mode as the first energy equalization mode; when the fluctuation range values ​​do not meet preset conditions, determining the energy equalization mode as the second energy equalization mode.

[0104] The aforementioned fluctuation range values ​​include the energy smoothness of the first audio frame, and the fluctuation range value meets the preset condition if the energy smoothness is less than a first threshold; or, the fluctuation range value includes the amplitude smoothness of the first audio frame, and the fluctuation range value meets the preset condition if the amplitude smoothness is less than a second threshold; or, the fluctuation range value includes the energy deviation of the first audio frame, and the fluctuation range value meets the preset condition if the energy deviation is not within a first preset range; or, the fluctuation range value includes the amplitude deviation of the first audio frame, and the fluctuation range value meets the preset condition if the amplitude deviation is not within a second preset range.

[0105] In this embodiment of the invention, energy smoothness represents the fluctuation of the frame energy after the normalization of the frequency domain coefficient energy of multiple channels in the current frame, which has been filtered by the multi-channel filtering unit. It can be measured by the smoothness calculation formula. When the energy of all channels in the current frame is the same, the energy smoothness of the current frame is 1; when the energy of a certain channel in the current frame is 0, the energy smoothness of the current frame is 0. Therefore, the range of energy smoothness between channels is [0,1]. The greater the fluctuation of energy between channels, the smaller the value of its energy smoothness. In one embodiment, a uniform first threshold can be set for all channel formats (e.g., 5.1, 7.1, 9.1, 11.1), for example, it can be 0.483, 0.492, or 0.504, etc. In another embodiment, different first thresholds are set for different channel formats. For example, the first threshold for 5.1 channel format is 0.511, the first threshold for 7.1 channel format is 0.563, the first threshold for 9.1 channel format is 0.608, and the first threshold for 11.1 channel format is 0.654.

[0106] Amplitude smoothness represents the fluctuation of the frame amplitude after the normalization of the frequency domain coefficient amplitude of multiple channels in the current frame, which has been filtered by the multi-channel filtering unit. It can be measured by the smoothness calculation formula. When all channels have the same frame amplitude, its smoothness is 1; when the frame amplitude of one of the channels is 0, its smoothness is 0. Therefore, the range of amplitude smoothness is between [0,1]. The greater the fluctuation of the amplitude between channels, the smaller the smoothness value. In one implementation, a uniform second threshold can be set for all channel formats (such as 5.1, 7.1, 9.1, 11.1), for example, it can be 0.695, 0.701, or 0.710, etc. In another implementation, different second thresholds can be given for different channel formats. For example, the second threshold for 5.1 channel format can be 0.715, the second threshold for 7.1 channel format can be 0.753, the second threshold for 9.1 channel format can be 0.784, and the second threshold for 11.1 channel format can be 0.809.

[0107] Since there is a square relationship between amplitude and energy, there is also a square relationship between amplitude flatness and energy flatness. That is, the fluctuation of the inter-channel frame amplitude corresponding to the square of amplitude flatness is approximately equal to the fluctuation of the inter-channel frame energy corresponding to energy flatness.

[0108] This embodiment can determine the energy equalization mode by using the above-mentioned multiple information representing fluctuation range values ​​of at least five channel signals, including energy smoothness, amplitude smoothness, energy deviation, or amplitude deviation.

[0109] (1) Calculate the energy values ​​of at least five channel signals, obtain the energy flatness of the first audio frame based on the energy values ​​of at least five channel signals, and determine the energy equalization mode as the first energy equalization mode when the energy flatness of the first audio frame is less than the first threshold; and determine the energy equalization mode as the second energy equalization mode when the energy flatness of the first audio frame is greater than or equal to the first threshold.

[0110] (2) Calculate the amplitude values ​​of at least five channel signals, obtain the amplitude flatness of the first audio frame based on the amplitude values ​​of at least five channel signals, and determine the energy equalization mode as the first energy equalization mode when the amplitude flatness of the first audio frame is less than the second threshold; and determine the energy equalization mode as the second energy equalization mode when the amplitude flatness of the first audio frame is greater than or equal to the second threshold.

[0111] (3) Calculate the energy values ​​of at least five channel signals, obtain the energy deviation of the first audio frame based on the energy values ​​of at least five channel signals, and determine the energy equalization mode as the first energy equalization mode when the energy deviation of the first audio frame is not within the first preset range; and determine the energy equalization mode as the second energy equalization mode when the energy deviation of the first audio frame is within the first preset range.

[0112] (4) Calculate the amplitude values ​​of at least five channel signals, obtain the amplitude deviation of the first audio frame based on the amplitude values ​​of at least five channel signals, and determine the energy equalization mode as the first energy equalization mode when the amplitude deviation of the first audio frame is not within the second preset range; and determine the energy equalization mode as the second energy equalization mode when the amplitude deviation of the first audio frame is within the second preset range.

[0113] It should be noted that other energy balance models may also be used in this application, and no specific limitations are imposed on them.

[0114] In one possible implementation, before determining the energy equalization mode based on the fluctuation range values ​​of at least five channel signals, the energy equalization mode can be determined first based on the encoding bitrate corresponding to the first audio frame, that is, whether the encoding bitrate is greater than the bitrate threshold. When the encoding bitrate is greater than the bitrate threshold, the energy equalization mode is determined to be the second energy equalization mode; when the encoding bitrate is less than or equal to the bitrate threshold, the energy equalization mode is determined based on the fluctuation range values ​​of at least five channel signals.

[0115] When the energy equalization mode is the first energy equalization mode, the average value of the energy or amplitude of the two channel signals contained in the current channel pair in the target channel pair set corresponding to the pairing mode can be calculated. Based on the average value, the two channel signals are subjected to energy equalization processing to obtain the corresponding two equalized channel signals.

[0116] In this way, when the fluctuation range of at least five channel signals is large, energy equalization can be performed only between two relevant channel signals. This makes the bit allocation during stereo processing more consistent with the fluctuation range of the channel signals, avoiding the problem that in low bit rate coding environments, the coding noise of high-energy channel pairs may be much greater than that of low-energy channel pairs due to insufficient bits, while the low-energy channel pairs will have bit redundancy.

[0117] When the energy equalization mode is the second energy equalization mode, the average value of the energy or amplitude of at least five channel signals can be calculated, and the energy equalization processing of at least five channel signals can be performed on the at least five channel signals according to the average value to obtain at least five equalized channel signals.

[0118] It should be noted that step 303 mainly involves performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals. These at least two channel signals are the channel signals that have already been paired in the target channel pair set. Apart from the channel signals that have already been paired in the target channel pair set, the remaining unpaired channel signals are directly encoded.

[0119] Step 304: Encode at least two equalized channel signals to obtain an encoded bitstream.

[0120] Step 305: When the sum of the correlation values ​​is less than or equal to a preset threshold, encode at least five channel signals to obtain an encoded bitstream.

[0121] In this embodiment, if the sum of correlation values ​​is less than or equal to a preset threshold, it indicates that the correlation between the two channel signals in the channel pair in the target channel pair set is low, there is no need for pair coding, and therefore no need to perform energy equalization processing on at least five channel signals. At this time, the object of coding is the at least five channel signals, rather than the equalized channel signals.

[0122] In this embodiment, the at least five channel signals contained in the audio frame are paired to obtain a target channel pair set in order to obtain the maximum sum of correlation values. When the sum of correlation values ​​of the target channel pair set is greater than a preset threshold, energy equalization processing is performed on the at least five channel signals, and then they are encoded, which can improve the encoding efficiency of the audio frame.

[0123] The following two specific examples illustrate... Figure 3 The process of determining the pairing method and energy equalization mode in the illustrated embodiment is described. Taking a 5.1 channel as an example, the 5.1 channel includes a center channel (C), a front left channel (left, L), a front right channel (right, R), a rear left surround channel (left surround, LS), a rear right surround channel (right surround, RS), and a 0.1 channel low frequency effects (LFE). The channel indexes of the above six channel signals are set as shown in Table 1, for example.

[0124] Table 1

[0125] Channel Index vocal tract signal 0 L 1 R 2 LS 3 RS 4 C 5 LFE

[0126] Figure 4aThis is an exemplary structural diagram of the encoding device used in the multi-channel audio signal encoding method provided in this application. The encoding device can be the encoder 20 of the source device 12 in the audio decoding system 10, or the decoding module 270 in the audio decoding device 200. The encoding device may include a multi-channel adaptive pairing module, a channel encoding module, and a stream multiplexing interface.

[0127] The input to the multi-channel adaptive pairing module includes six channel signals (L, R, C, LS, RS, LFE) of 5.1 channels, and a multi-channel processing indicator (MultiProcFlag). The output includes six paired channel signals (M1, S1, M2, S2, C, LFE), where M1 and S1 are a pair of paired channels, M2 and S2 are a pair of paired channels, and multi-channel side information (sideInfoMc) which includes the set of channel pairs.

[0128] The channel encoding module uses a mono encoding unit (or mono channel box, mono tool) to encode the channel signals (M1, S1, M2, S2, C, LFE) output from the multi-channel adaptive group module, outputting corresponding encoded channel signals (E1-E6). During the encoding process, the mono encoding unit allocates more bits to channel signals with higher energy (or higher amplitude) and fewer bits to channel signals with lower energy (or lower amplitude). Optionally, the channel encoding module can also use a stereo encoding unit, such as a parametric stereo encoder or a lossy stereo encoder, to encode the processed channel signals output from the multi-channel processing module.

[0129] It should be noted that unpaired channel signals (such as C and LFE) can be directly input into the channel encoding module to obtain encoded channel signals E5 and E6.

[0130] The stream multiplexing interface generates an encoded multichannel signal, which includes the encoded channel signals (E1-E6) output by the channel encoding module and side information (including multichannel side information). Optionally, the stream multiplexing interface can process the encoded multichannel signal into a serial signal or a serial bit stream.

[0131] Figure 4b This is an exemplary structural diagram of a multi-channel adaptive pairing module, such as... Figure 4b As shown, the multi-channel adaptive pairing module includes: a multi-channel filtering unit, a global correlation value statistics unit, a multi-channel energy equalization selection module, and a pairing processing module.

[0132] The multi-channel filtering unit selects five channel signals (L, R, C, LS, RS, LFE) from the six channel signals (L, R, C, LS, RS) to participate in multi-channel processing, namely L, R, C, LS, RS.

[0133] The global correlation value statistics unit first calculates the normalized correlation value between any two channel signals involved in multi-channel processing, i.e., between any two channel signals from L, R, C, LS, and RS. This application can use the following formula to calculate the correlation value between two channel signals (e.g., channel signal ch1 and channel signal ch2):

[0134]

[0135] Here, corr(ch1,ch2) represents the normalized correlation value between channel signals ch1 and ch2, spec_ch1(i) represents the frequency domain coefficient of the i-th frequency point of channel signal ch1, spec_ch2(i) represents the frequency domain coefficient of the i-th frequency point of channel signal ch2, and N represents an integer value not exceeding the total number of frequency points in one audio frame. Then, based on the normalized correlation value between any two channel signals, the system determines the channel pair set with the largest sum of correlation values ​​(i.e., the sum of correlation values ​​of all channel pairs contained in the channel pair set) and the channel pair set corresponding to this largest value (considered as the target channel pair set). Finally, the system outputs global correlation value side information, which includes the maximum sum of correlation values ​​corr_sum_max and the target channel pair set. Assuming the target channel pair set includes (R,C) and (LS,RS), the maximum sum of correlation values ​​corr_sum_max = corr(L,R) + corr(LS,RS).

[0136] It should be noted that after obtaining the normalized correlation value between any two channel signals, the global correlation value statistics unit can filter the correlation value according to the set pairing threshold. That is, correlation values ​​greater than or equal to the pairing threshold are retained, while correlation values ​​less than the pairing threshold are deleted or set to 0. This reduces the amount of computation.

[0137] The multi-channel energy equalization selection module determines whether energy equalization processing is needed for the five channel signals based on the encoding bitrate and the five channel signals. The five channel signals are grouped globally, aiming to obtain the maximum sum of correlation values, as described in step 302. When the sum of the correlation values ​​of the target channel pairs is greater than a preset threshold, it is determined that energy equalization processing is needed for the five channel signals; when the sum of the correlation values ​​of the target channel pairs is less than or equal to the preset threshold, it is determined that energy equalization processing is not needed for the five channel signals. When it is determined that energy equalization processing is needed for the five channel signals, the energy equalization mode is determined.

[0138] Figure 4c This is an exemplary structural diagram of the group pair processing module, such as Figure 4c As shown, the pair processing module includes a pair decision unit, an energy equalization unit, and a stereo processing box.

[0139] The pairing decision unit first calculates the energy or amplitude value of each channel signal. This application can use the following formula to calculate the energy or amplitude value of the channel signal (ch):

[0140]

[0141] Where energy(ch) represents the energy or amplitude value of the vocal channel signal ch, sepc_coeff(ch,i) represents the frequency domain coefficient of the i-th frequency point of the vocal channel signal ch, and N represents an integer value that does not exceed the total number of frequency points of an audio frame.

[0142] Then, the normalized energy or amplitude value of each channel signal is calculated. This application can use the following formula to calculate the normalized energy or amplitude value of the channel signal (ch):

[0143]

[0144] Wherein, energy_uniform(ch) represents the normalized energy or amplitude value of the vocal channel signal ch, and energy_max represents the maximum of the energy or amplitude values ​​of the five vocal channel signals (i.e., energy(L), energy(R), energy(C), energy(LS), and energy(RS)). If energy_max = 0, then energy_uniform(ch) is 0 for all channels.

[0145] Next, the fluctuation range values ​​of the five channel signals are calculated. Optionally, the fluctuation range value can refer to the energy smoothness. This application can use the following formula to calculate the energy smoothness of the five channel signals:

[0146]

[0147] Here, efm represents the energy flatness of the five channel signals. The channel indices of L, R, C, LS, and RS are shown in Table 1.

[0148] Optionally, the fluctuation range value can also refer to the energy deviation. Based on the normalized energy or amplitude value energy_uniform(ch) obtained from the above calculation, this application can use the following formula to calculate the average energy or amplitude value of the five channel signals:

[0149]

[0150] Here, avg_energy_uniform represents the average energy or amplitude value of the five channel signals. The channel indices of L, R, C, LS, and RS are shown in Table 1.

[0151] The energy deviation of the vocal tract signal (ch) is calculated using the following formula:

[0152]

[0153] Wherein, deviation(ch) represents the energy deviation of the vocal tract signal ch. The largest of the energy deviations of L, R, C, LS, and RS is determined as the energy deviation of the five vocal tract signals.

[0154] Optionally, the fluctuation range value can also refer to the amplitude value or amplitude deviation, and its principle is similar to the energy-related values ​​mentioned above, which will not be elaborated here.

[0155] As described above, the energy equalization mode of this application includes two implementation methods. The Pair energy equalization mode uses two channel signals from a channel pair to obtain the two equalized channel signals corresponding to that channel pair in the target channel pair set corresponding to the pairing method determined by the module selection unit. The overall energy equalization mode uses two channel signals from a channel pair and at least one external channel signal from each channel to obtain the two equalized channel signals corresponding to that channel pair. For channel signals without pairing, the corresponding equalized channel signal is the channel signal itself.

[0156] The pair decision maker determines the energy balance mode based on the fluctuation range value, including the following two judgment methods:

[0157] (1) When efm is less than the first threshold, the energy balance mode is the Pair energy balance mode; when efm is greater than or equal to the first threshold, the energy balance mode is the overall energy balance mode.

[0158] (2) When deviation is within the value range [threshold, 1 / threshold], the energy balance mode is the overall energy balance mode; when deviation is not within the value range [threshold, 1 / threshold], the energy balance mode is the Pair energy balance mode. The value range of threshold can be (0, 1).

[0159] Deviation can be represented as the ratio of the frequency domain amplitude of each channel in the current frame to the average frequency domain amplitude of each channel in the current frame, i.e., the amplitude deviation. When the ratio between the frequency domain amplitude of the current channel and the average frequency domain amplitude of each channel in the current frame is less than 5 (corresponding to threshold = 0.2), it can be divided into two cases: First, the frequency domain amplitude of the current channel is less than or equal to the average frequency domain amplitude of each channel in the current frame, and the condition "frequency domain amplitude of the current channel / average frequency domain amplitude of each channel in the current frame" is between (0.2, 1], that is, between (threshold, 1]; Second, the frequency domain amplitude of the current channel is greater than the average frequency domain amplitude of each channel in the current frame, and the condition "frequency domain amplitude of the current channel / average frequency domain amplitude of each channel in the current frame" is between (1, 5). Combining the above two cases, when the ratio between the frequency domain amplitude of the current channel and the average frequency domain amplitude of each channel in the current frame is less than 5, The condition "frequency domain amplitude of the current channel / average frequency domain amplitude of all channels in the current frame" must be within the range of (0.2, 5), which is also within the range of (threshold, 1 / threshold). (threshold, 1 / threshold) is the second preset range mentioned above. The threshold value can be between (0, 1). A smaller threshold value indicates a greater fluctuation in the frequency domain amplitude of the current channel relative to the average frequency domain amplitude of all channels in the current frame, while a larger threshold value indicates a smaller fluctuation. The threshold value can be 0.2, 0.15, 0.125, 0.11, or 0.1, etc.

[0160] Deviation can also represent the ratio of the frequency domain energy of each channel to the average frequency domain energy of each channel, i.e., the energy deviation. When the ratio of the frequency domain energy of the current channel to the average frequency domain energy of each channel in the current frame is less than 25 (threshold = 0.04), it can be divided into two cases: First, the frequency domain energy of the current channel is less than or equal to the average frequency domain energy of each channel in the current frame, satisfying the condition that "the frequency domain energy of the current channel / the average frequency domain energy of each channel in the current frame" is between (0.04, 1], i.e., (threshold, 1]; Second, the frequency domain energy of the current channel is greater than the average frequency domain energy of each channel in the current frame, satisfying the condition that "the frequency domain energy of the current channel / the average frequency domain energy of each channel in the current frame" is between (0.04, 1], i.e., (threshold, 1]. The ratio of the frequency domain energy of the current channel to the average frequency domain energy of each channel in the current frame is between (1, 25). Considering both cases, when the ratio of the frequency domain energy of the current channel to the average frequency domain energy of each channel in the current frame is less than 25, the range of "frequency domain energy of the current channel / average frequency domain energy of each channel in the current frame" that satisfies the condition is between (0.04, 25), which is between (threshold, 1 / threshold). (threshold, 1 / threshold) is the first preset range mentioned above. Where, threshold... The threshold value can be between (0,1). A smaller threshold value indicates a greater fluctuation in the frequency domain energy of the current channel relative to the average frequency domain energy of all channels in the current frame, while a larger threshold value indicates a smaller fluctuation. Threshold values ​​can be 0.04, 0.0225, 0.015625, 0.0121, or 0.01, etc.

[0161] Since there is a square relationship between amplitude and energy, there is also a square relationship between amplitude deviation and energy deviation. That is, the fluctuation of inter-channel frame amplitude corresponding to the square of amplitude deviation is approximately equivalent to the fluctuation of inter-channel frame energy corresponding to energy deviation.

[0162] In another implementation, the first preset range can also be extended to (0, 1 / threshold). In this case, the range of Pair energy equalization is [1 / threshold, +∞). This indicates that Pair energy equalization is performed only when the frequency domain energy of the current channel is greater than the average frequency domain energy of each channel in the current frame, and "the frequency domain energy of the current channel / the average frequency domain energy of each channel in the current frame" is greater than 1 / threshold.

[0163] In another implementation, the second preset range can also be extended to (0, 1 / threshold). In this case, the range of Pair amplitude equalization is [1 / threshold, +∞). This indicates that Pair amplitude equalization is performed only when the frequency domain amplitude of the current channel is greater than the average frequency domain amplitude of each channel in the current frame, and "the frequency domain amplitude of the current channel / the average frequency domain amplitude of each channel in the current frame" is greater than 1 / threshold.

[0164] It should be noted that the pairing decision unit can calculate normalized energy or amplitude values ​​based on the five channel signals to obtain energy smoothness or energy deviation; it can also calculate normalized energy or amplitude values ​​based only on the successfully paired channel signals to obtain energy smoothness or energy deviation; or it can calculate normalized energy or amplitude values ​​based on a portion of the five channel signals to obtain energy smoothness or energy deviation. This application does not impose specific limitations on these methods.

[0165] The stereo processing unit can employ prediction-based or Karhunen-Loeve Transform (KLT)-based processing, in which the two input channel signals are rotated (e.g. via a 2×2 rotation matrix) to maximize energy compression, thereby concentrating the signal energy into one channel.

[0166] After processing the two input channel signals, the stereo processing unit outputs the processed channel signals (P1-P4) corresponding to the two channel signals, as well as multi-channel side information, which includes the sum of correlation values ​​and the target channel pair set.

[0167] Figure 5 This is an exemplary structural diagram of the decoding device used in the multi-channel audio decoding method provided in this application. The decoding device can be the decoder 30 of the destination device 14 in the audio decoding system 10, or the decoding module 270 in the audio decoding device 200. The decoding device may include a stream demultiplexing interface, a channel decoding module, and a multi-channel processing module.

[0168] The stream demultiplexing interface receives encoded multichannel signals (e.g., serial bitstreams) from the encoding device, and demultiplexes them to obtain encoded channel signals (E) and multichannel parameters (SIDE_PAIR). For example, E1, E2, E3, E4, ..., Ei1, Ei, and SIDE_PAIR1, SIDE_PAIR2, ..., SIDE_PAIRm.

[0169] The channel decoding module uses a mono decoding unit (or mono box, mono tool) to decode the encoded channel signals output from the code stream demultiplexing interface and output decoded channel signals (D). For example, E1, E2, E3, E4, ..., Ei1, Ei are each decoded by a mono decoding unit to obtain E1, and then decoded to obtain D1, D2, D3, D4, ..., Di1, Di.

[0170] The multi-channel processing module includes multiple stereo processing units, which can employ prediction-based or KLT-based processing, where the two input channel signals are inversely rotated (e.g., via a 2×2 rotation matrix) to transform the signals back to their original direction.

[0171] The decoded channel signals output by the channel decoding module can identify which two decoded channel signals are paired using multi-channel parameters. The paired decoded channel signals are then input to the stereo processing unit. The stereo processing unit processes the two input decoded channel signals and outputs the corresponding channel signals (CH). For example, stereo processing unit 1 processes D1 and D2 according to SIDE_PAIR1 to obtain CH1 and CH2; stereo processing unit 2 processes D3 and D4 according to SIDE_PAIR2 to obtain CH3 and CH4; and stereo processing unit m processes Di-1 and Di according to SIDE_PAIRm to obtain CHi-1 and CHi.

[0172] It should be noted that unpaired channel signals (such as CHj) do not need to be processed by the stereo processing unit in the multi-channel processing module and can be directly output after decoding.

[0173] Figure 6 This is a schematic diagram of the structure of an embodiment of the encoding device of this application, as shown below. Figure 6 As shown, this device can be applied to the source device 12 or the audio decoding device 200 in the above embodiments. The encoding device in this embodiment may include: an acquisition module 601, a processing module 602, and an encoding module 603.

[0174] The acquisition module 601 is used to acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; acquire the sum of correlation values ​​of a target channel pair set, the target channel pair set being obtained with the aim of obtaining the maximum sum of correlation values, the target channel pair set including at least one channel pair, a channel pair including two channel signals from the at least five channel signals, the channel pair having a correlation value, the correlation value being used to represent the correlation between the two channel signals of the channel pair; the processing module 602 is used to perform energy equalization processing on the at least five channel signals to obtain at least five equalized channel signals when the sum of the correlation values ​​is greater than a preset threshold; the encoding module 603 is used to encode the at least five equalized channel signals.

[0175] In one possible implementation, the encoding module 603 is further configured to encode the at least five channel signals when the sum of the relevant values ​​is less than or equal to the preset threshold.

[0176] In one possible implementation, the processing module 602 is specifically used to obtain the fluctuation range values ​​of the at least five channel signals; determine the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals; and perform energy equalization processing on the at least five channel signals according to the energy equalization mode to obtain the at least five equalized channel signals.

[0177] In one possible implementation, the processing module 602 is specifically used to determine the energy balance mode as a first energy balance mode when the fluctuation range value meets the preset conditions; or, when the fluctuation range value does not meet the preset conditions, determine the energy balance mode as a second energy balance mode.

[0178] In one possible implementation, the fluctuation range value includes the energy smoothness of the first audio frame; the fluctuation range value meets a preset condition if the energy smoothness is less than a first threshold; or, the fluctuation range value includes the amplitude smoothness of the first audio frame; the fluctuation range value meets a preset condition if the amplitude smoothness is less than a second threshold; or, the fluctuation range value includes the energy deviation of the first audio frame; the fluctuation range value meets a preset condition if the energy deviation is not within a first preset range; or, the fluctuation range value includes the amplitude deviation of the first audio frame; the fluctuation range value meets a preset condition if the amplitude deviation is not within a second preset range.

[0179] In one possible implementation, when the energy equalization mode is the first energy equalization mode, the processing module 602 is specifically used to calculate the average value of the energy or amplitude values ​​of the two channel signals contained in the current channel pair in the target channel pair set, and perform energy equalization processing on the two channel signals according to the average value to obtain the corresponding two equalized channel signals.

[0180] In one possible implementation, when the energy equalization mode is the second energy equalization mode, the processing module 602 is specifically used to calculate the average value of the energy or amplitude value of the at least five channel signals, and perform energy equalization processing on the at least five channel signals according to the average value to obtain the at least five equalized channel signals.

[0181] In one possible implementation, the processing module 602 is further configured to determine whether the encoding bitrate corresponding to the first audio frame is greater than a bitrate threshold; when the encoding bitrate is greater than the bitrate threshold, the energy balancing mode is determined to be a second energy balancing mode; and when the encoding bitrate is less than or equal to the bitrate threshold, the energy balancing mode is determined based on the fluctuation range value.

[0182] The apparatus of this embodiment can be used to perform Figure 3 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0183] Figure 7 This is a schematic diagram of the structure of an embodiment of the device in this application, as shown below. Figure 7 As shown, the device can be the encoding device in the above embodiments. The device in this embodiment may include: a processor 701 and a memory 702, the memory 702 being used to store one or more programs; when the one or more programs are executed by the processor 701, the processor 701 performs the following... Figure 3 The technical solution of the method embodiment shown.

[0184] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or implemented by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0185] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0186] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0187] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0188] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0189] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0190] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0191] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0192] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for encoding multi-channel audio signals, characterized in that, include: Acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; Obtain the sum of correlation values ​​for all channel pairs in a target channel pair set, the target channel pair set including at least one channel pair, a channel pair including two channel signals from the at least five channel signals, a channel pair having a correlation value, the correlation value being used to represent the correlation between the two channel signals of the channel pair, the target channel pair set being obtained with the aim of obtaining the maximum sum of correlation values; When the sum of the correlation values ​​is greater than a preset threshold, energy equalization processing is performed on at least two of the at least five channel signals to obtain at least two equalized channel signals. The at least two equalized channel signals are encoded to obtain an encoded bitstream.

2. The method according to claim 1, characterized in that, The method further includes: When the sum of the correlation values ​​is less than or equal to the preset threshold, the at least five channel signals are encoded to obtain an encoded bitstream.

3. The method according to claim 1, characterized in that, The step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: Obtain the fluctuation range values ​​of the at least five channel signals; The energy equalization mode is determined based on the fluctuation range values ​​of the at least five channel signals; The at least two channel signals are subjected to energy equalization processing according to the energy equalization mode to obtain the at least two equalized channel signals.

4. The method according to claim 3, characterized in that, Determining the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals includes: When the fluctuation range value meets the preset conditions, the energy balance mode is determined to be the first energy balance mode; or... When the fluctuation range value does not meet the preset conditions, the energy balance mode is determined to be the second energy balance mode.

5. The method according to claim 4, characterized in that, The fluctuation range value includes the energy smoothness of the first audio frame; the fluctuation range value meets the preset condition if the energy smoothness is less than a first threshold; or... The fluctuation range value includes the amplitude flatness of the first audio frame; the fluctuation range value meets the preset condition if the amplitude flatness is less than a second threshold; or... The fluctuation range value includes the energy deviation of the first audio frame; the fluctuation range value meeting the preset condition means that the energy deviation is not within a first preset range; or... The fluctuation range value includes the amplitude deviation of the first audio frame; the fluctuation range value meets the preset condition if the amplitude deviation is not within a second preset range.

6. The method according to claim 4 or 5, characterized in that, When the energy equalization mode is the first energy equalization mode, the step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: The target channel pair is subjected to energy equalization processing to obtain the at least two equalized channel signals.

7. The method according to claim 6, characterized in that, The step of performing energy equalization processing on the channel signals corresponding to the target channel pair set to obtain the at least two equalized channel signals includes: For the current channel pair in the target channel pair set, calculate the average value of the energy or amplitude value of the two channel signals contained in the current channel pair, and perform energy equalization processing on the two channel signals contained in the current channel pair according to the average value to obtain two equalized channel signals.

8. The method according to claim 4 or 5, characterized in that, When the energy equalization mode is the second energy equalization mode, the step of performing energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals includes: Calculate the average value of the energy or amplitude value of the at least five channel signals, and perform energy equalization processing on the at least five channel signals according to the average value to obtain the at least five equalized channel signals.

9. The method according to any one of claims 3-5 and 7, characterized in that, Before determining the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals, the method further includes: Determine whether the encoding bitrate corresponding to the first audio frame is greater than the bitrate threshold; When the coding bitrate is greater than the bitrate threshold, the energy balancing mode is determined to be the second energy balancing mode. The energy balance mode is determined based on the fluctuation range value only when the encoding bitrate is less than or equal to the bitrate threshold.

10. The method according to any one of claims 1-5 and 7, characterized in that, The method further includes: Encode the channel signals that have not undergone energy equalization processing from the at least five channel signals.

11. An encoding device, characterized in that, include: An acquisition module is used to acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; acquire the sum of correlation values ​​of all channel pairs in a target channel pair set, the target channel pair set including at least one channel pair, a channel pair including two channel signals from the at least five channel signals, the channel pair having a correlation value, the correlation value being used to represent the correlation between the two channel signals of the channel pair, the target channel pair set being obtained with the aim of obtaining the maximum sum of correlation values; The processing module is used to perform energy equalization processing on at least two of the at least five channel signals to obtain at least two equalized channel signals when the sum of the correlation values ​​is greater than a preset threshold. An encoding module is used to encode the at least two equalized channel signals to obtain an encoded bitstream.

12. The apparatus according to claim 11, characterized in that, The encoding module is further configured to encode the at least five channel signals to obtain an encoded bitstream when the sum of the correlation values ​​is less than or equal to the preset threshold.

13. The apparatus according to claim 11, characterized in that, The processing module is specifically used to obtain the fluctuation range values ​​of the at least five channel signals; determine the energy equalization mode based on the fluctuation range values ​​of the at least five channel signals; and perform energy equalization processing on the at least two channel signals according to the energy equalization mode to obtain the at least two equalized channel signals.

14. The apparatus according to claim 13, characterized in that, The processing module is specifically used to determine the energy balance mode as a first energy balance mode when the fluctuation range value meets the preset conditions; or to determine the energy balance mode as a second energy balance mode when the fluctuation range value does not meet the preset conditions.

15. The apparatus according to claim 14, characterized in that, The fluctuation range value includes the energy smoothness of the first audio frame; the fluctuation range value meets the preset condition if the energy smoothness is less than a first threshold; or... The fluctuation range value includes the amplitude flatness of the first audio frame; the fluctuation range value meets the preset condition if the amplitude flatness is less than a second threshold; or... The fluctuation range value includes the energy deviation of the first audio frame; the fluctuation range value meeting the preset condition means that the energy deviation is not within a first preset range; or... The fluctuation range value includes the amplitude deviation of the first audio frame; the fluctuation range value meets the preset condition if the amplitude deviation is not within a second preset range.

16. The apparatus according to claim 14 or 15, characterized in that, When the energy equalization mode is the first energy equalization mode, the processing module is specifically used to perform energy equalization processing on the channel signals corresponding to the target channel pair set to obtain the at least two equalized channel signals.

17. The apparatus according to claim 16, characterized in that, The processing module is specifically used to calculate the average value of the energy value or amplitude value of the two channel signals contained in the current channel pair in the target channel pair set, and perform energy equalization processing on the two channel signals contained in the current channel pair according to the average value to obtain the corresponding two equalized channel signals.

18. The apparatus according to claim 14 or 15, characterized in that, When the energy equalization mode is the second energy equalization mode, the processing module is specifically used to calculate the average value of the energy value or amplitude value of the at least five channel signals, and perform energy equalization processing on the at least five channel signals according to the average value to obtain the at least five equalized channel signals.

19. The apparatus according to any one of claims 13-15, 17, characterized in that, The processing module is further configured to determine whether the encoding bitrate corresponding to the first audio frame is greater than the bitrate threshold; when the encoding bitrate is greater than the bitrate threshold, the energy balancing mode is determined to be the second energy balancing mode; when the encoding bitrate is less than or equal to the bitrate threshold, the energy balancing mode is determined based on the fluctuation range value.

20. The apparatus according to any one of claims 11-15, 17, characterized in that, The encoding module is also used to encode the channel signals that have not undergone energy equalization processing among the at least five channel signals.

21. A device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, It includes a computer program that, when executed on a computer, causes the computer to perform the method of any one of claims 1-10.

23. A computer-readable storage medium, characterized in that, This includes the bitstream obtained according to the encoding method for multichannel audio signals as described in any one of claims 1-10.