Multi-channel audio coding and decoding method, device, equipment, medium and product

By using two downmixing parameters to process multi-channel audio signals, surround channel correlation signals, channel fusion signals, and stereo channel signals are generated, solving the problems of multi-channel audio signal encoding quality and backward compatibility, and achieving high-quality audio encoding and decoding restoration.

CN121747587APending Publication Date: 2026-03-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Multi-channel audio signals have poor encoding quality and are difficult to backward compatible. Traditional audio encoders cannot parse the bitstream with three-channel side information, resulting in reduced encoding quality and compatibility issues.

Method used

Two different downmixing parameters are used to process multi-channel audio signals separately, generating surround channel related signals, channel fusion signals and stereo channel signals. By encapsulating the audio spatial parameter stream, the center channel and low-frequency channel are encoded separately, reducing crosstalk between channels and ensuring backward compatibility on the decoding side.

Benefits of technology

It improves the encoding and decoding quality of multi-channel audio signals, ensuring that traditional decoders can decode stereo and the new decoder can fully recover multi-channel audio signals, achieving backward compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747587A_ABST
    Figure CN121747587A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing in some cases, and discloses a multi-channel audio coding and decoding method, device, equipment, medium and product, the coding method comprises the following steps: obtaining a multi-channel audio signal, a first down-mixing parameter and a second down-mixing parameter; performing down-mixing processing on the multi-channel audio signal by using the first down-mixing parameter to obtain a first audio down-mixing signal; processing the multi-channel audio signal by using the first audio down-mixed signal to obtain an audio space parameter stream; performing down-mixing processing on the multi-channel audio signal by using a second down-mixing parameter to obtain a second audio down-mixing signal; and packaging the second audio down-mixed signal and the audio space parameter stream to obtain an audio code corresponding to the multi-channel audio signal. Based on this, the parameterized coding quality of the multi-channel audio signal is improved, the backward compatibility of the decoding side is ensured, and the decoding restoration quality of audio coding is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In some cases, this relates to the field of audio processing technology, specifically to multi-channel audio encoding and decoding methods, devices, equipment, media, and products. Background Technology

[0002] In the parametric encoding of multi-channel low-bitrate audio signals, the bitstream obtained by adding side information to the transmitted signal may produce unresolved errors for traditional audio encoders, affecting the encoding quality and backward compatibility of multi-channel audio signals. Summary of the Invention

[0003] In view of this, a multi-channel audio encoding and decoding method, apparatus, device, medium and product are provided to solve the problems of poor encoding quality and difficulty in backward compatibility of multi-channel audio signals.

[0004] In a first aspect, a multi-channel audio encoding method is provided, comprising: acquiring a multi-channel audio signal, a first downmixing parameter, and a second downmixing parameter, wherein the multi-channel audio signal is a signal to be encoded, and the first and second downmixing parameters correspond to the multi-channel audio signal; performing downmixing processing on the multi-channel audio signal using the first downmixing parameter to obtain a first audio downmixed signal, the first audio downmixed signal including a surround channel related signal and a channel fusion signal, the channel fusion signal including a center channel and a low-frequency channel, the surround channel related signal and the channel fusion signal being in different transmission channels; processing the multi-channel audio signal using the first audio downmixed signal to obtain an audio spatial parameter stream; performing downmixing processing on the multi-channel audio signal using the second downmixing parameter to obtain a second audio downmixed signal, the second audio downmixed signal including a stereo channel signal and a channel fusion signal, the stereo channel signal and the channel fusion signal being in different transmission channels; and encapsulating the second audio downmixed signal and the audio spatial parameter stream to obtain an audio code, the audio code corresponding to the multi-channel audio signal.

[0005] Secondly, a multi-channel audio decoding method is provided, comprising: acquiring an audio code to be decoded and a decoding end type, wherein the audio code to be decoded is encoded based on a multi-channel audio encoding method according to the first aspect or any corresponding embodiment thereof; parsing the audio code based on the decoding end type to obtain a first channel signal; and decoding the first channel signal to obtain a multi-channel audio signal, wherein the multi-channel audio signal corresponds to the audio code to be decoded.

[0006] Thirdly, a multi-channel audio encoding apparatus is provided, comprising: a first acquisition module, configured to acquire a multi-channel audio signal, a first downmixing parameter, and a second downmixing parameter, wherein the multi-channel audio signal is a signal to be encoded, and the first downmixing parameter and the second downmixing parameter correspond to the multi-channel audio signal; and a first downmixing processing module, configured to perform downmixing processing on the multi-channel audio signal using the first downmixing parameter to obtain a first audio downmixed signal, wherein the first audio downmixed signal includes a surround channel correlation signal and a channel fusion signal, the channel fusion signal includes a center channel and a low-frequency channel, and the surround channel correlation signal is fused with the channel fusion signal. The signals are located in different transmission channels; the spatial parameter determination module is used to process the multi-channel audio signal using the first audio downmixing signal to obtain an audio spatial parameter stream; the second downmixing processing module is used to downmix the multi-channel audio signal using the second downmixing parameters to obtain a second audio downmixing signal, which includes stereo channel signals and channel fusion signals, and the stereo channel signals and channel fusion signals are located in different transmission channels; the bitstream encapsulation module is used to encapsulate the second audio downmixing signal and the audio spatial parameter stream to obtain an audio code, which corresponds to the multi-channel audio signal.

[0007] Fourthly, a multi-channel audio decoding device is provided, comprising: a second acquisition module for acquiring an audio code to be decoded and a decoding end type, wherein the audio code to be decoded is encoded based on the multi-channel audio encoding method of the first aspect or any corresponding embodiment thereof; a signal matching module for parsing the audio code based on the decoding end type to obtain a first channel signal; and an audio decoding module for decoding the first channel signal to obtain a multi-channel audio signal, wherein the multi-channel audio signal corresponds to the audio code to be decoded.

[0008] Fifthly, an electronic device is provided, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the multi-channel audio encoding method of the first aspect or any corresponding embodiment thereof, or to perform the multi-channel audio decoding method of the second aspect or any corresponding embodiment thereof.

[0009] In a sixth aspect, a computer-readable storage medium is provided, on which computer instructions are stored, the computer instructions being configured to cause a computer to perform the multi-channel audio encoding method of the first aspect or any corresponding embodiment thereof, or to perform the multi-channel audio decoding method of the second aspect or any corresponding embodiment thereof.

[0010] In a seventh aspect, a computer program product is provided, including computer instructions for causing a computer to execute the multi-channel audio encoding method of the first aspect or any corresponding embodiment thereof, or to execute the multi-channel audio decoding method of the second aspect or any corresponding embodiment thereof.

[0011] On the encoding side, two different downmixing parameters (i.e., the first downmixing parameter and the second downmixing parameter) are used to determine the first audio downmixing signal and the second audio downmixing signal, respectively. The surround channel related signal in the first audio downmixing signal is located in a different channel from the channel fusion signal composed of the center channel and the low-frequency channel. That is, the stereo channel signal in the first audio downmixing signal is separated from the channel fusion signal composed of the center channel and the low-frequency channel. This achieves separate encoding of the center channel and the low-frequency channel, which helps reduce inter-channel crosstalk. Subsequently, the audio spatial parameter stream corresponding to the multi-channel audio signal is determined using the first audio downmixing signal and the multi-channel audio signal, realizing independent encoding of the audio spatial parameters used to characterize side information. By encapsulating the second audio downmixing signal and the audio spatial parameter stream, the audio code corresponding to the multi-channel audio signal is obtained, so that the audio code simultaneously includes the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream, improving the parameterized encoding quality of the multi-channel audio signal. Subsequently, after receiving the audio encoding, the decoding side can parse it according to the decoding end type, ensuring backward compatibility of the audio encoding with the decoding side, improving the decoding and restoration quality of the audio encoding, and ensuring the playback experience of the audio corresponding to the multi-channel audio signal. Attached Figure Description

[0012] To more clearly illustrate the specific implementation methods or technical solutions in the prior art under certain circumstances, the accompanying drawings used in the description of the specific implementation methods or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are implementation methods under certain circumstances. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0013] Figure 1 These are schematic diagrams illustrating application scenarios in various situations;

[0014] Figure 2 This is a flowchart illustrating the first type of multichannel audio coding method under certain circumstances; Figure 3 This is a schematic diagram of the second process for multi-channel audio coding methods in some situations; Figure 4 These are schematic diagrams illustrating the encoding of multi-channel audio signals under certain conditions; Figure 5 This is a flowchart illustrating the first type of multi-channel audio decoding method under certain circumstances; Figure 6 This is a second flowchart illustrating a multi-channel audio decoding method under certain circumstances; Figure 7 These are decoding diagrams of the first type of decoder in some scenarios; Figure 8 This is a flowchart illustrating the third method for multi-channel audio decoding in some situations; Figure 9 These are decoding diagrams for a second-type decoder in some scenarios; Figure 10 These are block diagrams of multi-channel audio encoding devices in some scenarios; Figure 11 These are block diagrams of multi-channel audio decoding devices under certain conditions; Figure 12 These are schematic diagrams of the hardware structure of electronic devices in some scenarios. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments in some situations clearer, the technical solutions in some situations will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments in some situations, not all embodiments. Based on the embodiments in some situations, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection.

[0016] It is understandable that before using publicly available technical solutions, users should be informed of the types, scope of use, and usage scenarios of personal information involved in certain situations, and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0017] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform certain technical solutions.

[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation method. Other methods that comply with relevant laws and regulations may also be applied to the implementation method in the corresponding situation.

[0020] It is understandable that the data involved (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.

[0021] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature marked "first" or "second" may explicitly or implicitly include one or more of that feature. In some descriptions, "multiple" means two or more, unless otherwise explicitly specified.

[0022] The most direct encoding method for multi-channel audio signals is to quantize and encode the waveforms of the multi-channel audio signals directly without any preprocessing. However, this method requires a large bit overhead. Although it is possible to combine the spatial layout of the multi-channel signals and jointly encode the symmetrical channels in the space to appropriately reduce the required bit rate, the bit rate overhead is still too high.

[0023] Based on this, the current main approach is to use some parametric encoding methods, which use metadata to describe multi-channel audio signals instead of directly encoding the multi-channel audio signals themselves. Examples include MPEG-Surround (MPS), Binaural Cue Coding (BCC), and Directional Audio Coding (DirAC).

[0024] However, dual-channel parametric coding methods often suffer from crosstalk in the center channel, and the bitstream obtained by three-channel coding with added side information cannot guarantee backward compatibility. When a traditional audio decoder receives the encoded bitstream with added side information from three transmission channels, it will encounter unresolved errors.

[0025] In summary, the parametric coding methods in related technologies cannot simultaneously guarantee the coding quality and backward compatibility of multi-channel audio signals.

[0026] Based on this, two different downmixing matrices are used on the encoding side to encode spatial parameter information and the transmitted signal, respectively. Specifically, one downmixing matrix is ​​used so that the first two channels contain only surround sound. The center channel (C) and the low-frequency channel (LFE) are mixed and used as the third transmission channel. The audio spatial parameter information is calculated using this audio downmixed signal and the original signal. The other downmixing matrix is ​​used to mix the center channel (C) and the LFE into the first two channels of the three transmission channels. The center channel (C) and the LFE are then mixed and used as the third transmission channel. The first two channels of the transmitted signal are stereo-coded, and the third transmission channel is encoded as a single channel. This reduces inter-channel crosstalk and improves the quality of parametric coding.

[0027] Traditional audio decoders, upon receiving a bitstream, can only parse the stereo signal from the transmitted signal to achieve a stereo playback experience. However, when a new decoder receives a bitstream, it obtains the three-channel transmission signal and spatial parameters. By upmixing the transmission signal and spatial parameters, it can reconstruct the corresponding multi-channel audio signal from the bitstream. This ensures backward compatibility, guaranteeing both that traditional decoders can decode stereo to maintain a basic playback experience and that new decoders can fully recover multi-channel audio signals.

[0028] In summary, the center channel signal will not be mixed into the stereo transmission channel, avoiding crosstalk between the center channel and the surround channel, improving the encoding quality of multi-channel audio signals and backward compatibility on the decoding side, and improving the decoding quality of audio encoding.

[0029] As an optional application scenario in some situations, such as Figure 1 As shown, this application scenario may include at least one electronic device and at least one server. Figure 1 The example illustrates that the application scenario includes a computer 101, a mobile terminal 102, and a server 103, and that electronic devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 via a network 110.

[0030] Specifically, electronic devices can be smartphones, tablets, laptops, PDAs, desktop computers, game consoles, smart TVs, smart wearable devices, in-vehicle terminals, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranets, local area networks, wide area networks, mobile communication networks, and combinations thereof.

[0031] According to an embodiment of a multi-channel audio encoding method provided in some cases, it should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0032] In some cases, a multi-channel audio encoding method is provided, which can be used in electronic devices with audio encoding capabilities, such as encoders. Figure 2 These are flowcharts of multi-channel audio coding methods in some scenarios, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the multi-channel audio signal, the first downmixing parameter, and the second downmixing parameter. The multi-channel audio signal is the signal to be encoded, and the first and second downmixing parameters correspond to the multi-channel audio signal.

[0033] A multi-channel audio signal can be an audio signal with multiple audio channels, and each audio channel is independent of the others. For example, a multi-channel audio signal corresponding to 5.1 channels or a multi-channel audio signal corresponding to 7.1.4 channels.

[0034] Both the first and second downmixing parameters are used to compress multi-channel audio signals into audio signals with fewer channels. The first downmixing parameter is used to extract spatial audio parameters from the multi-channel audio signal, and the second downmixing parameter is used to extract stereo signals from the multi-channel audio signal. The number of channels corresponding to the first and second downmixing parameters is the same.

[0035] In certain specific situations, different audio systems have corresponding channel layouts to achieve a three-dimensional surround sound field. Examples include audio systems used in home theaters, professional recording studios, or high-end car audio systems. Based on the channel layout of the audio system, the corresponding multi-channel audio signals can be obtained.

[0036] A multi-channel audio signal includes audio parameters corresponding to multiple audio channels. By combining the spatial audio parameters and stereo signal audio parameter information required for audio encoding, the corresponding first and second downmixing parameters are obtained.

[0037] Step S202: Use the first downmixing parameter to downmix the multi-channel audio signal to obtain the first audio downmixed signal.

[0038] The first audio downmix signal includes a surround channel related signal and a channel fusion signal. The channel fusion signal includes the center channel and the low-frequency channel. The surround channel related signal and the channel fusion signal are in different transmission channels.

[0039] The audio parameters in the multi-channel audio signal are down-mixed using the first down-mixing parameter to convert and compress the number of channels in the multi-channel audio signal, resulting in the corresponding first audio down-mixed signal. The number of channels in the first audio down-mixed signal is the same as the number of channels provided by the first down-mixing parameter.

[0040] Specifically, the multi-channel audio signal includes at least the left channel L, the right channel R, the center channel C, the low-frequency channel LFE, the left surround Ls, and the right surround Rs.

[0041] Surround channel related signals are the surround audio signals extracted from the surround channels of a multi-channel audio signal. These signals can include the left channel related surround signal and the right channel related surround signal. Taking a 5.1 channel layout as an example, the surround channel related signals can include the audio signal corresponding to the left surround Ls and the audio signal corresponding to the right surround Rs.

[0042] Channel fusion signal is a signal generated by fusing the central audio signal extracted from the center channel of a multi-channel audio signal with the low-frequency audio signal extracted from the low-frequency channel. Taking a 5.1 channel layout as an example, the central audio signal of the center channel C is superimposed with the low-frequency audio signal of the low-frequency channel LFE to generate a channel fusion signal.

[0043] In the first audio downmix signal, the surround channel related signal and the channel blending signal are located in different transmission channels. Specifically, the left channel signal and the left channel related surround signal are located in transmission channel 0 of the first audio downmix signal, the right channel signal and the right channel related surround signal are located in transmission channel 1 of the first audio downmix signal, and the channel blending signal is located in transmission channel 2 of the first audio downmix signal.

[0044] Taking a 5.1 channel layout as an example, the audio signals corresponding to the left channel L and the left surround Ls are in transmission channel 0 of the first audio downmixing signal, the audio signals corresponding to the right channel R and the right surround Rs are in transmission channel 1 of the first audio downmixing signal, and the channel fusion signal composed of the center channel C and the low frequency channel LFE is in transmission channel 2 of the first audio downmixing signal.

[0045] Step S203: The multi-channel audio signal is processed using the first audio downmixing signal to obtain the audio spatial parameter stream.

[0046] Audio spatial parameter streams are used to characterize the spatial properties of multi-channel audio signals. These streams include inter-channel coherence (ICC) and inter-channel level difference (ICLD). ICC describes the linear correlation between multi-channel audio signals; ICLD describes the level differences between multi-channel audio signals, and the level differences characterize the differences in energy distribution.

[0047] In some specific situations, the level difference between channels can be quantized as the energy ratio of the target channel to the reference channel (or downmixer). By acquiring the first audio downmixer signal and the original multi-channel audio signal, and combining each target channel and its corresponding reference channel in the multi-channel audio signal, the level difference between channels in the multi-channel audio signal can be calculated.

[0048] In some specific cases, inter-channel coherence can be quantified as the degree of linear correlation between two channel signals. By acquiring multiple channel pairs corresponding to the multi-channel audio signal, and combining the first audio downmix signal and the original multi-channel audio signal, the inter-channel coherence between all channel pairs is calculated. The value of this inter-channel coherence ranges from [0,1]. If the inter-channel coherence is 1, it indicates that the two channel signals are completely linearly correlated; if the inter-channel coherence is 0, it indicates that the two channel signals are completely unrelated.

[0049] Step S204: Use the second downmixing parameter to downmix the multi-channel audio signal to obtain the second audio downmixed signal.

[0050] The second audio downmix signal includes a stereo channel signal and a channel fusion signal, with the stereo channel signal and the channel fusion signal located in different transmission channels.

[0051] The audio parameters in the multi-channel audio signal are down-mixed using the second down-mixing parameter, which converts and compresses the number of channels of the multi-channel audio signal to the number of channels provided by the second down-mixing parameter, thus obtaining the corresponding second audio down-mixed signal. That is, the number of channels of the second audio down-mixed signal is the same as the number of channels provided by the second down-mixing parameter.

[0052] As described above, the multi-channel audio signal includes at least the left channel L, the right channel R, the center channel C, the low-frequency channel LFE, the left surround Ls, and the right surround Rs.

[0053] Stereo signal is an audio signal that includes stereo components extracted from a multi-channel audio signal. For example, the stereo signal left channel is the fusion of audio information from the left channel L, left surround Ls, center channel C, and low frequency channel LFE; the stereo signal right channel is the fusion of audio information from the right channel R, right surround Rs, center channel C, and low frequency channel LFE.

[0054] In the second audio downmix signal, the stereo channel signal and the channel blending signal are in different transmission channels. For example, the left channel of the stereo signal is in transmission channel 0 of the second audio downmix signal, the left channel of the stereo signal is in transmission channel 1 of the second audio downmix signal, and the channel blending signal is in transmission channel 2 of the second audio downmix signal.

[0055] Step S205: Encapsulate the second audio downmixing signal and the audio spatial parameter stream to obtain the audio code, which corresponds to the multi-channel audio signal.

[0056] Audio encoding is the digital bitstream obtained by converting multi-channel audio signals. This audio encoding includes a base layer bitstream and an enhancement layer bitstream. The base layer bitstream represents the core audio information corresponding to the multi-channel audio signals and can be decoded independently. The enhancement layer bitstream represents the detailed supplementary information corresponding to the multi-channel audio signals, and may include spatial enhancement layers (such as multi-channel spatial parameters), sound quality enhancement layers (such as high-frequency harmonics, quantization residual optimization, etc.), and channel expansion layers (such as expanding from stereo to 5.1 channels / 7.1 channels, etc.).

[0057] In some specific cases, stereo encoding is performed on the stereo signal in the second audio downmix signal to obtain a stereo bitstream, and the stereo bitstream is encapsulated as audio basic information into the base layer bitstream; at the same time, single-channel encoding is performed on the channel fusion signal in the second audio downmix signal, and the single-channel bitstream obtained by single-channel encoding and the audio spatial parameter stream are encapsulated as audio detail information into the enhancement layer bitstream to form the audio encoding corresponding to the multi-channel audio signal.

[0058] By using two different downmixing parameters (i.e., the first downmixing parameter and the second downmixing parameter) on the encoding side to determine the first audio downmixing signal and the second audio downmixing signal respectively, and with the surround channel related signal in the first audio downmixing signal located in a different channel from the channel fusion signal composed of the center channel and the low-frequency channel, the stereo channel signal in the first audio downmixing signal is separated from the channel fusion signal composed of the center channel and the low-frequency channel. This achieves separate encoding of the center channel and the low-frequency channel, which helps reduce inter-channel crosstalk. Subsequently, the audio spatial parameter stream corresponding to the multi-channel audio signal is determined by using the first audio downmixing signal and the multi-channel audio signal together, realizing independent encoding of the audio spatial parameters used to characterize side information. By encapsulating the second audio downmixing signal and the audio spatial parameter stream, the audio code corresponding to the multi-channel audio signal is obtained, so that the audio code simultaneously contains the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream, improving the parameterized encoding quality of the multi-channel audio signal, while ensuring backward compatibility on the decoding side, which is beneficial to improving the decoding and restoration quality of the audio code.

[0059] In some cases, a multi-channel audio encoding method is provided, which can be used in electronic devices with encoding functions such as encoders. Figure 3 These are flowcharts of multi-channel audio coding methods in some scenarios, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain the multi-channel audio signal, the first downmixing parameter, and the second downmixing parameter. The multi-channel audio signal is the signal to be encoded, and the first and second downmixing parameters correspond to the multi-channel audio signal.

[0060] Specifically, step S301 includes: Step S3011: Obtain the multi-channel audio signal to be encoded. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above; they will not be repeated here.

[0061] Step S3012: Analyze the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters.

[0062] Among them, the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters are all from the multi-channel audio signal; the first surround signal related parameters correspond to the first channel signal parameters, and the second surround signal related parameters correspond to the second channel signal parameters.

[0063] A multi-channel audio signal includes channel signals, corresponding surround channel related signals, center channel signals, and low-frequency channel signals. Each channel of the multi-channel audio signal has corresponding signal parameters. When a multi-channel audio signal is received from an audio system, the signal parameters of each channel are analyzed to obtain the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters corresponding to the first channel signal parameters, the second surround signal related parameters corresponding to the second channel signal parameters, the center channel signal parameters, and the low-frequency channel signal parameters.

[0064] The first channel signal parameter can be the audio signal parameter of the left channel L; the second channel signal parameter can be the audio signal parameter of the right channel R; the first surround signal related parameter can be the left channel related surround signal parameter corresponding to the left channel L; and the second surround signal related parameter can be the right channel related surround signal parameter corresponding to the right channel R.

[0065] In a specific example, a multi-channel audio signal with a 5.1 channel layout is represented as follows: The first surround signal related parameters can be the left surround signal. The corresponding audio signal parameters, the second surround signal related parameters can be right surround. The corresponding audio signal parameters.

[0066] Step S3013: Based on the positions of the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters in the multi-channel audio signal, the first downmixing parameters are obtained.

[0067] The first downmixing parameter is represented by the first downmixing matrix; the first downmixing matrix is ​​used to incorporate the first channel signal parameters and the first surround signal related parameters into the first transmission channel, incorporate the second channel signal parameters and the second surround signal related parameters into the second transmission channel, and incorporate the center channel signal parameters and the low frequency channel signal parameters into the third transmission channel; the first audio downmixing signal includes the first transmission channel, the second transmission channel, and the third transmission channel.

[0068] To integrate the first channel signal parameters and their corresponding first surround signal parameters into the first transmission channel, the second channel signal parameters and their corresponding second surround signal parameters into the second transmission channel, and the center channel signal parameters and low-frequency channel signal parameters into the third transmission channel, the positions of the first channel signal parameters, second channel signal parameters, first surround signal parameters, second surround signal parameters, center channel signal parameters, and low-frequency channel signal parameters in the multi-channel audio signal are identified respectively. This allows for the determination of the corresponding downmixing parameter values ​​for each parameter, and these downmixing parameter values ​​are then used to construct a first downmixing matrix. This first downmixing matrix is ​​used to characterize the first downmixing parameters.

[0069] In a specific example, if a multi-channel audio signal with a 5.1 channel layout is represented as: In order to incorporate the audio signal parameters corresponding to the left channel L and the left surround Ls into the first transmission channel of the first audio downmixing signal, to incorporate the audio signal parameters corresponding to the right channel R and the right surround Rs into the second transmission channel of the first audio downmixing signal, and to incorporate the center channel C signal parameters and the low-frequency channel LFE signal parameters into the third transmission channel of the first audio downmixing signal, the first downmixing matrix used to characterize the first downmixing parameters can be determined as follows:

[0070] in, This is the signal gain.

[0071] Step S3014: Based on the positions of the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters in the multi-channel audio signal, the second downmixing parameters are obtained.

[0072] The second downmixing parameter is represented by the second downmixing matrix. The second downmixing matrix is ​​used to incorporate the first channel signal parameters, the first surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fourth transmission channel, incorporate the second channel signal parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fifth transmission channel, and incorporate the center channel signal parameters and the low-frequency channel signal parameters into the sixth transmission channel. The second audio downmixing signal includes the fourth, fifth, and sixth transmission channels.

[0073] To integrate the first channel signal parameters and their corresponding first surround signal parameters, center channel signal parameters, and low-frequency channel signal parameters into the fourth transmission channel, integrate the second channel signal parameters and their corresponding second surround signal parameters, center channel signal parameters, and low-frequency channel signal parameters into the fifth transmission channel, and integrate the center channel signal parameters and low-frequency channel signal parameters into the sixth transmission channel, the positions of the first channel signal parameters, second channel signal parameters, first surround signal parameters, second surround signal parameters, center channel signal parameters, and low-frequency channel signal parameters in the multi-channel audio signal are identified respectively. This allows for the determination of the corresponding downmixing parameter values, which are then used to construct a second downmixing matrix. This second downmixing matrix is ​​used to characterize the second downmixing parameters.

[0074] In a specific example, if a multi-channel audio signal with a 5.1 channel layout is represented as: In order to incorporate the audio signal parameters corresponding to the left channel L, the left surround Ls, the center channel C, and the low-frequency channel LFE into the fourth transmission channel of the second audio downmixing signal; to incorporate the audio signal parameters corresponding to the right channel R, the right surround Rs, the center channel C, and the low-frequency channel LFE into the fifth transmission channel of the second audio downmixing signal; and to incorporate the audio signal parameters corresponding to the center channel C and the low-frequency channel LFE into the sixth transmission channel of the second audio downmixing signal, the second downmixing matrix M2 used to characterize the second downmixing parameters can be determined as follows:

[0075] in, This is the signal gain.

[0076] Step S302: Use the first downmixing parameter to downmix the multi-channel audio signal to obtain the first audio downmixed signal.

[0077] The first audio downmix signal includes surround channel related signals, center channel and low frequency channel channel fusion signals, with the surround channel related signals and channel fusion signals located in different transmission channels.

[0078] Multiplying the matrix representation of the multi-channel audio signal with the first downmixing matrix representing the first downmixing parameter enables downmixing of the multi-channel audio signal, resulting in a first audio downmixed signal that separates the surround channel related signal from the channel fusion signal.

[0079] Specifically, the matrix representation of the multi-channel audio signal Y is as follows: The first undermixing matrix M1, representing the first undermixing parameter, is: , like Figure 4 As shown, multiplying the first downmixing matrix M1 with the matrix representation of the multi-channel audio signal Y yields the first audio downmixing signal X1, i.e.: .

[0080] in," "This indicates the left channel audio parameters of the first transmission channel of the first audio downmix signal;" "This indicates the right channel audio parameters of the second transmission channel in the first audio downmix signal; This refers to the channel fusion signal of the third transmission channel in the first audio downmix signal.

[0081] Step S303: The multi-channel audio signal is processed using the first audio downmixing signal to obtain the audio spatial parameter stream. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above; they will not be repeated here.

[0082] Step S304: Use the second downmixing parameter to downmix the multi-channel audio signal to obtain the second audio downmixed signal.

[0083] The second audio downmix signal includes a stereo channel signal and a channel fusion signal, with the stereo channel signal and the channel fusion signal located in different transmission channels.

[0084] By multiplying the matrix representation of the multi-channel audio signal with the second downmixing matrix representing the second downmixing parameter, downmixing processing of the multi-channel audio signal can be achieved, resulting in a second audio downmixed signal that separates the stereo channel signal from the channel fusion signal.

[0085] Specifically, the matrix representation of the multi-channel audio signal Y is as follows: The second undermixing matrix M2, representing the second undermixing parameter, is: , like Figure 4 As shown, multiplying the second downmixing matrix M2 with the matrix representation of the multi-channel audio signal Y yields the second audio downmixing signal X2, i.e.: .

[0086] in," "This indicates the left stereo signal in the first transmission channel of the second audio downmix signal;" "Indicates the right stereo signal in the second transmission channel of the second audio sub-mix signal; This indicates the channel fusion signal of the third transmission channel in the second audio downmix signal.

[0087] Step S305: Encapsulate the second audio downmixing signal and the audio spatial parameter stream to obtain the audio code, which corresponds to the multi-channel audio signal.

[0088] Specifically, audio encoding includes a first-layer bitstream and a second-layer bitstream, where the first-layer bitstream represents the base layer bitstream and the second-layer bitstream represents the enhancement layer bitstream. Accordingly, step S305 includes: Step S3051: Pair the stereo signal in the second audio downmix signal to obtain paired stereo.

[0089] The audio encoding information corresponding to each transmission channel in the second audio downmix signal is identified, and the audio encoding information representing stereo is paired to obtain paired stereo. Specifically, if the first transmission channel in the second audio downmix signal is a left stereo signal and the second transmission channel in the second audio downmix signal is a right stereo signal, then the left stereo signal and the right stereo signal are paired stereo.

[0090] Step S3052: Perform stereo encoding on the paired stereo to obtain a stereo bitstream, and encapsulate the stereo bitstream into the first layer bitstream.

[0091] The paired stereo signals are encoded using a stereo coding method (such as joint stereo coding, parametric stereo coding, etc.) to obtain the corresponding stereo bitstream. This stereo bitstream is then encapsulated as the basic audio information into the first layer bitstream of the audio bitstream (i.e., the base layer bitstream).

[0092] Step S3053: The channel fusion signal is encoded in a single channel to obtain the channel fusion bitstream, and the channel fusion bitstream and audio spatial parameter stream are encapsulated into the second layer bitstream.

[0093] The channel fusion signal is encoded using a single-channel encoding method to obtain the corresponding channel fusion bitstream. Then, the channel fusion bitstream and the audio spatial parameter stream are used as supplementary information to the basic audio information and encapsulated into the second layer bitstream (i.e., the enhancement layer bitstream) of the audio bitstream.

[0094] By analyzing the channel signal parameters carried by the multi-channel audio signal, a first downmixing matrix is ​​determined to characterize the first downmixing parameter and a second downmixing matrix is ​​determined to characterize the second downmixing parameter. The first downmixing matrix is ​​used to participate in the determination process of the audio spatial parameters of the multi-channel audio signal, and the second downmixing matrix is ​​used to participate in the determination process of the transmission signal (i.e., stereo signal and channel fusion signal) of the multi-channel audio signal. This achieves the separation of audio spatial parameters and transmission signals, greatly improving the coding quality.

[0095] During the encoding process, the stereo bitstream in the transmitted signal is encapsulated into the base layer bitstream, and the channel fusion bitstream and audio spatial parameter bitstream are encapsulated into the enhancement layer bitstream. This ensures that the generated audio code is backward compatible with both traditional and new decoders. Traditional decoders can decode stereo signals to ensure the basic audio playback experience, while new decoders can completely recover multi-channel audio signals.

[0096] In some cases, a multi-channel audio decoding method is provided, which can be used in electronic devices such as decoders that have audio decoding capabilities. Figure 5 This is a flowchart based on multi-channel audio decoding methods in some situations, such as... Figure 5 As shown, the process includes the following steps: Step S501: Obtain the audio encoding to be decoded and the decoding end type. The audio encoding to be decoded is obtained based on the multi-channel audio encoding method described in the above embodiment.

[0097] The decoder corresponds to the encoder. The decoder type indicates the type of decoder, such as new type decoder and old type decoder. The new type decoder supports parsing stereo bitstreams of the base layer bitstream, as well as single-channel bitstreams and spatial parameter streams of the enhancement layer bitstream; the old type decoder only supports parsing stereo bitstreams of the base layer bitstream.

[0098] When the encoder obtains an audio code based on a multi-channel audio encoding method, it can output the audio code to the decoder. Correspondingly, the decoder can receive the audio code sent by the encoder and perform the decoding process to restore the audio code to the original multi-channel audio signal.

[0099] Step S502: Analyze the audio code to be decoded based on the decoding end type to obtain the first channel signal.

[0100] The first channel signal is the signal to be parsed determined from the audio encoding. For example, the first channel signal can be a stereo signal, a channel-blended signal, or audio spatial parameters. Specifically, different decoding end types use different methods to parse the audio encoding, resulting in different first channel signals. After the decoder obtains the audio encoding, it determines the first channel signal to be parsed from the audio encoding based on the decoding end type of the decoder.

[0101] Step S503: Decode the first channel signal to obtain a multi-channel audio signal, which corresponds to the audio code to be decoded.

[0102] The audio component corresponding to the first channel signal is decoded from the first channel signal, and the audio signal is reconstructed in reverse according to the audio component to restore the corresponding multi-channel audio signal.

[0103] After obtaining the audio code to be decoded, the audio code is analyzed in conjunction with the decoding end type to determine the first channel signal that needs to be analyzed, so as to restore the multi-channel audio signal from the first channel signal, avoid the problem of being unable to be decoded, ensure the backward compatibility of the audio code, improve the decoding quality of the audio code for multi-channel audio signals, and ensure the playback experience of the audio corresponding to the multi-channel audio signal.

[0104] In some cases, a multi-channel audio decoding method is provided, which can be used in electronic devices such as decoders that have audio decoding capabilities. Figure 6 This is a flowchart based on multi-channel audio decoding methods in some situations, such as... Figure 6 As shown, the process includes the following steps: Step S601: Obtain the audio encoding to be decoded and the decoding end type, wherein the audio encoding to be decoded is obtained based on the multi-channel audio encoding method of the above embodiments. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above, which will not be repeated here.

[0105] Step S602: Analyze the audio code to be decoded based on the decoding end type to obtain the first channel signal.

[0106] Specifically, the first channel signal includes stereo channel signals, as described above, which are encapsulated as basic audio information in the base layer bitstream.

[0107] Accordingly, step S602 includes: in response to the decoding end type being the first type, parsing the stereo channel signal and refusing to parse the channel fusion signal and the audio spatial parameter stream. Here, the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream are all part of the audio encoding to be decoded.

[0108] The first type of decoder is an older type of decoder that only supports parsing stereo bitstreams of the base layer bitstream, such as... Figure 7 As shown, when the decoding end type is the first type, the audio encoding output by the encoder is decapsulated to obtain the stereo bitstream, single-channel bitstream, and audio spatial parameter stream corresponding to the audio encoding. The stereo bitstream is input into the first type of decoder for decoding, and the stereo channel signal corresponding to the stereo bitstream is output; at the same time, the recognition and parsing of the single-channel bitstream and the recognition and parsing of the audio spatial parameter stream are skipped.

[0109] Step S603: Decode the first channel signal to obtain a multi-channel audio signal, which corresponds to the audio code to be decoded.

[0110] Stereo signals carry all the basic audio information; that is, they contain all channel components. By decoding the stereo signal corresponding to the stereo bitstream, the audio signals corresponding to each channel are reconstructed. These audio signals are then combined to obtain the multi-channel audio signal corresponding to the target channel signal.

[0111] Since the stereo signal is encapsulated in the base layer bitstream, while the channel fusion signal and audio spatial parameters are encapsulated in the enhancement layer bitstream, when the decoding end type is the first type (i.e., the traditional decoder), only the base layer bitstream can be parsed to obtain the stereo signal, without parsing the enhancement layer bitstream, thus achieving a stereo playback experience and achieving decoding compatibility with traditional decoders.

[0112] In some cases, a multi-channel audio decoding method is provided, which can be used in the aforementioned electronic devices, such as desktop computers and laptop computers. Figure 8 These are flowcharts of multi-channel audio decoding methods in some situations, such as... Figure 8 As shown, the process includes the following steps: Step S801: Obtain the audio encoding to be decoded and the decoding end type. The audio encoding to be decoded is obtained based on the multi-channel audio encoding method described in the above embodiments. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above; they will not be repeated here.

[0113] Step S802: Analyze the audio code to be decoded based on the decoding end type to obtain the first channel signal.

[0114] Specifically, the first channel signal includes a stereo channel signal, a channel fusion signal, and an audio spatial parameter stream. As described above, the stereo channel signal is encapsulated as basic audio information in the base layer bitstream, while the channel fusion signal and the audio spatial parameter stream are encapsulated as audio detail information in the enhancement layer bitstream.

[0115] Accordingly, step S802 includes: in response to the decoding end type being the second type, parsing the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream.

[0116] The second type of decoder is a new type of decoder that supports parsing stereo streams of the base layer bitstream, single-channel streams of the enhancement layer bitstream, and spatial parameter streams. For example... Figure 9 As shown, when the decoder is a type II decoder, the audio encoding output by the encoder is decapsulated to obtain the stereo bitstream, single-channel bitstream, and audio spatial parameter stream corresponding to the audio encoding. Spatial parameter information is recovered from the audio spatial parameter stream; single-channel decoding is performed on the single-channel bitstream to obtain the corresponding channel fusion signal.

[0117] Step S803: Decode the first channel signal to obtain a multi-channel audio signal, which corresponds to the audio code to be decoded.

[0118] Specifically, step S803 includes: Step S8031: Remove the channel fusion signal from the stereo channel signal to obtain the first stereo signal.

[0119] As described above, the stereo channel signal is obtained by encoding paired stereo signals. Therefore, the stereo channel signal that can be parsed from the stereo bitstream includes the left channel signal, the left-side related surround signal, the right channel signal, the right-side related surround signal, the center channel signal, and the low-frequency channel signal.

[0120] After obtaining the stereo channel signal, a separation algorithm (such as spectral subtraction) is used to remove the channel fusion signal composed of the center channel signal and the low-frequency channel signal from the stereo channel signal, resulting in the first stereo signal after removing the channel fusion signal.

[0121] The matrix representation of the multi-channel audio signal Y is as follows: Encoding example, if the decoded stereo signal is ( \ ), that is, paired stereo is ( \ Specifically, it can be expressed as: .

[0122] Decoded channel fusion signal It can be represented as:

[0123] Subsequently, the component of the first stereo signal can be separated from the stereo signal. \ Specifically, it can be expressed as: .

[0124] Step S8032: Fuse the first stereo signal with the channel fusion signal to obtain a multi-channel signal.

[0125] As described above, the channel fusion signal is obtained by performing single-channel encoding. Therefore, the channel fusion signal parsed from the single-channel bitstream can be a single-channel signal. The first stereo signal separated from the stereo channel signal is combined with the channel fusion signal to form a multi-channel signal.

[0126] Following the previous example, the single-channel signal component will be separated from the stereo signal. \ and channel fusion signals By combining these signals, a three-channel signal is formed, as shown below: .

[0127] Step S8033: Use a preset upmixing algorithm to upmix the multi-channel signal and the audio spatial parameter stream to obtain a multi-channel audio signal.

[0128] The preset upmixing algorithm is a pre-defined upmixing synthesis algorithm used to convert low-channel-number audio signals into multi-channel audio signals. Examples include passive upmixing based on signal correlation, active upmixing based on signal characteristics, and AI-based intelligent upmixing. No specific limitation is made to the preset upmixing algorithm here.

[0129] A preset upmixing algorithm is used to upmix the multi-channel signal and the audio spatial parameter stream to reconstruct the corresponding multi-channel audio signal.

[0130] Since the stereo signal is encapsulated in the base layer bitstream, while the channel fusion signal and audio spatial parameters are encapsulated in the enhancement layer bitstream, when the decoding end type is the second type (i.e., the new decoder), the base layer bitstream can be parsed to obtain the stereo signal, and the enhancement layer bitstream can be parsed to obtain the audio spatial parameters and channel fusion parameters. By combining the stereo signal, channel fusion parameters, and audio spatial parameters, the original multi-channel audio signal can be synthesized, achieving decoding compatibility with the new decoder and improving the decoding quality of the multi-channel audio signal.

[0131] As a specific application scenario, taking 7.1.4 channels as an example, the above-mentioned multi-channel audio encoding and decoding methods will be explained. Specifically, the input multi-channel audio signal... The matrix representation is as follows: The first undermixing matrix M1 and the second undermixing matrix M2 are represented as follows:

[0132]

[0133] When the input multi-channel audio signal After being downmixed by M1 and M2 respectively, the following results were obtained:

[0134]

[0135] As can be seen from the above process, the left and right channels in channel 7.1.4 are respectively mixed with... The first two channels of stereo ( and In ) and Includes only the center channel (C) and the low-frequency channel (LFE), utilizing and Extract spatial parameters such as ICC and ICLD, and then calculate. The left channel (L), right channel (R), surround channels, and ICC and ICLD components between channels are not interfered with by the center channel, and encoding the center channel separately can also improve the synthesis quality of the upmixing algorithm on the decoding side.

[0136] Then it is sent as a transmission signal to the core encoder, in The first two channels ( and In addition to retaining In addition to the planar surround channels arranged in the same direction, a center channel (C) and a low-frequency channel (LFE) are also included. It's still a mix of center channel C and low-frequency channel LFE, then... The data is fed into the core encoder for encoding. Within the core encoder, the data is... and The paired stereo encoding is performed and encapsulated into the base layer bitstream, while Perform single-channel encoding and encapsulate it together with the spatial parameter stream into the enhancement layer bitstream.

[0137] On the decoder side, the stereo bitstream of the base layer bitstream and the single-channel bitstream and spatial parameters of the enhancement layer bitstream are decapsulated and sent to the core decoder and the side information recovery module respectively to obtain the stereo transmission signal. , Single-channel transmission signal ( ) and spatial parameters.

[0138] On the encoding side, the 3-channel transmission signal is obtained through the downmixing matrix M2, while the spatial parameters reflect the transmission of the original multi-channel audio signal through the downmixing matrix. Received and the original multi-channel audio signal The relationship between the various channels. Therefore, at this time, a mismatch occurs between the transmitted signal on the decoding side and the spatial parameter side information.

[0139] Due to the undermixing matrix It is known, and In It contains the complete components of the center channel C and the low-frequency channel LFE, at which point it can be based on and decoded Channel, using spectral subtraction to... and The central channel component is separated.

[0140]

[0141]

[0142] in, For the decoded transmission channel, This is the transmitted signal after channel separation. , The components of the center channel (C) and the low-frequency channel (LFE) have been removed as much as possible. and encoding side To maintain consistency, The restored spatial parameters are fed into the mixing algorithm to synthesize a multi-channel audio signal for output.

[0143] Of course, for other multi-channel systems such as 5.1, 7.1, 5.1.2, and 5.1.4, the encoding and decoding methods in some cases can also be supported. The processing methods are similar to those for 7.1.4 channels, so they will not be described in detail here.

[0144] In some cases, a multi-channel audio encoding apparatus is also provided for implementing the embodiments and preferred embodiments described above, which will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the apparatus described below is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0145] In some cases, a multi-channel audio encoding device is provided, such as Figure 10 As shown, it includes: The first acquisition module 901 is used to acquire a multi-channel audio signal, a first downmixing parameter, and a second downmixing parameter. The multi-channel audio signal is a signal to be encoded, and the first downmixing parameter and the second downmixing parameter correspond to the multi-channel audio signal.

[0146] The first downmixing processing module 902 is used to downmix the multi-channel audio signal using the first downmixing parameters to obtain the first audio downmixing signal. The first audio downmixing signal includes a surround channel related signal and a channel fusion signal. The channel fusion signal includes the center channel and the low-frequency channel. The surround channel related signal and the channel fusion signal are in different transmission channels.

[0147] The spatial parameter determination module 903 is used to process the multi-channel audio signal using the first audio downmixing signal to obtain the audio spatial parameter stream.

[0148] The second downmixing processing module 904 is used to downmix the multi-channel audio signal using the second downmixing parameters to obtain the second audio downmixing signal. The second audio downmixing signal includes stereo channel signals and channel fusion signals, and the stereo channel signals and channel fusion signals are in different transmission channels.

[0149] The stream encapsulation module 905 is used to encapsulate the second audio downmix signal and the audio spatial parameter stream to obtain the audio code, which corresponds to the multi-channel audio signal.

[0150] In some alternative implementations, the audio encoding includes a first-layer bitstream and a second-layer bitstream, and accordingly, the bitstream encapsulation module 905 includes: The pairing unit is used to pair the stereo signal in the second audio downmix signal to obtain paired stereo.

[0151] The stereo coding unit is used to perform stereo coding on paired stereo signals to obtain a stereo bitstream, and to encapsulate the stereo bitstream into the first layer bitstream.

[0152] The single-channel encoding unit is used to encode the channel fusion signal in a single channel to obtain the channel fusion bitstream, and encapsulate the channel fusion bitstream and audio spatial parameter stream to the second-layer bitstream.

[0153] In some optional implementations, the first acquisition module 901 includes: The parameter parsing unit is used to parse the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters. The first surround signal related parameters correspond to the first channel signal parameters, and the second surround signal related parameters correspond to the second channel signal parameters. All the parameters—first channel signal parameters, second channel signal parameters, first surround signal related parameters, second surround signal related parameters, center channel signal parameters, and low-frequency channel signal parameters—are from the multi-channel audio signal.

[0154] The first downmixing parameter determination unit is used to obtain the first downmixing parameter based on the position of the first channel signal parameter, the second channel signal parameter, the first surround signal related parameter, the second surround signal related parameter, the center channel signal parameter and the low frequency channel signal parameter in the multi-channel audio signal. The first downmixing parameter is represented by the first downmixing matrix. The first downmixing matrix is ​​used to incorporate the parameters of the first channel signal and the parameters related to the first surround signal into the first transmission channel, incorporate the parameters of the second channel signal and the parameters related to the second surround signal into the second transmission channel, and incorporate the parameters of the center channel signal and the low-frequency channel signal into the third transmission channel; the first audio downmixing signal includes the first transmission channel, the second transmission channel, and the third transmission channel.

[0155] The second downmixing parameter determination unit is used to obtain the second downmixing parameters based on the positions of the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters in the multi-channel audio signal; the second downmixing parameters are represented by the second downmixing matrix. The second downmixing matrix is ​​used to incorporate the first channel signal parameters, the first surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fourth transmission channel; to incorporate the second channel signal parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fifth transmission channel; and to incorporate the center channel signal parameters and the low-frequency channel signal parameters into the sixth transmission channel. The second audio downmixing signal includes the fourth, fifth, and sixth transmission channels.

[0156] The multi-channel audio encoding device can execute the multi-channel audio encoding method provided above, and has the corresponding functional modules and beneficial effects of the method.

[0157] On the encoding side, two different downmixing parameters (i.e., the first downmixing parameter and the second downmixing parameter) are used to determine the first audio downmixing signal and the second audio downmixing signal, respectively. The surround channel related signal in the first audio downmixing signal is located in a different channel from the channel fusion signal composed of the center channel and the low-frequency channel; that is, the stereo channel signal in the first audio downmixing signal is separated from the channel fusion signal composed of the center channel and the low-frequency channel. This achieves separate encoding of the center channel and the low-frequency channel, which helps reduce inter-channel crosstalk. Subsequently, the first audio downmixing signal and the multi-channel audio signal are used together to determine the audio spatial parameter stream corresponding to the multi-channel audio signal, realizing independent encoding of the audio spatial parameters used to characterize side information. By encapsulating the second audio downmixing signal and the audio spatial parameter stream, the audio code corresponding to the multi-channel audio signal is obtained, so that the audio code simultaneously includes the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream. This improves the parameterized encoding quality of the multi-channel audio signal and ensures backward compatibility on the decoding side, which is beneficial to improving the decoding and restoration quality of the audio code.

[0158] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0159] In some cases, a multi-channel audio decoding device is also provided for implementing the embodiments and preferred embodiments described above, which will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0160] In some cases, a multi-channel audio decoding device is provided, such as Figure 11 As shown, it includes: The second acquisition module 1001 is used to acquire the audio encoding to be decoded and the decoding end type. The audio encoding to be decoded is obtained based on a multi-channel audio encoding method.

[0161] The signal matching module 1002 is used to parse the audio code to be decoded based on the decoding end type to obtain the first channel signal.

[0162] The audio decoding module 1003 is used to decode the first channel signal to obtain multi-channel audio signals, and the multi-channel audio signals correspond to the audio codes to be decoded.

[0163] In some optional cases, the first channel signal includes the stereo channel signal, and accordingly, the signal matching module 1002 includes: The first identification unit is configured to, in response to the decoding end being of type 1, parse the stereo channel signal and refuse to parse the channel fusion signal and the audio spatial parameter stream. Here, the stereo channel signal, channel fusion signal, and audio spatial parameter stream are all part of the audio encoding to be decoded.

[0164] In some optional cases, the first channel signal includes a stereo channel signal, a channel fusion signal, and an audio spatial parameter stream; accordingly, the signal matching module 1002 includes: The second identification unit is used to parse the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream in response to the decoding end being of the second type.

[0165] In some alternative implementations, the audio decoding module 1003 includes: The signal removal unit is used to remove the channel fusion signal from the stereo channel signal to obtain the first stereo signal.

[0166] The signal fusion unit is used to fuse the first stereo signal with the channel fusion signal to obtain a multi-channel signal.

[0167] The upmixing unit is used to upmix multi-channel signals and audio spatial parameter streams using a preset upmixing algorithm to obtain multi-channel audio signals.

[0168] The multi-channel audio decoding device can execute the multi-channel audio decoding method described above, possessing the corresponding functional modules and beneficial effects. After acquiring the audio encoding to be decoded, it analyzes the audio encoding in conjunction with the decoder's end type to determine the target channel signal to be decoded, thereby reconstructing the multi-channel audio signal from the target channel signal. This avoids decoding failures, ensures backward compatibility of the audio encoding, improves the decoding quality of audio encoding for multi-channel audio signals, and guarantees the playback experience of the audio corresponding to the multi-channel audio signal.

[0169] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0170] Figure 12 This is a schematic diagram of the structure of an electronic device provided in certain situations.

[0171] The following is a detailed reference. Figure 12 This diagram illustrates a suitable structure for implementing an electronic device in various scenarios. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1101, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 1102 or a program loaded from memory 1108 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of the electronic device. The processor 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0172] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 1108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0173] In particular, under certain circumstances, the processes described in the above-referenced flowchart can be implemented as computer software programs. For example, in some cases, a computer program product may be included, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program may be downloaded and installed from a network via communication device 1109, or installed from memory 1108, or installed from ROM 1102. When the computer program is executed by processor 1101, it performs the functions defined in the multi-channel audio encoding method or multi-channel audio decoding method.

[0174] Figure 12 The electronic device shown is merely an example and should not impose any limitations on the aforementioned functions and scope of use.

[0175] In some cases, a computer-readable storage medium is also provided, in which the above-described methods can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that a computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the multi-channel audio encoding method or multi-channel audio decoding method shown in the above embodiments.

[0176] In some cases, certain components can be applied as computer program products, such as computer program instructions. When executed by a computer, these instructions, through the operation of the computer, can invoke or provide the aforementioned methods and / or technical solutions. Those skilled in the art should understand that the forms in which computer program instructions exist in computer-readable media include, but are not limited to, source files, executable files, and installation package files. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instruction; the computer compiling the instruction and then executing the corresponding compiled program; the computer reading and executing the instruction; or the computer reading and installing the instruction and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0177] While some embodiments have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the embodiments, all of which fall within the scope defined by the appended claims.

Claims

1. A multi-channel audio encoding method, comprising: Acquire a multi-channel audio signal, a first downmixing parameter, and a second downmixing parameter, wherein the multi-channel audio signal is a signal to be encoded, and the first downmixing parameter and the second downmixing parameter correspond to the multi-channel audio signal; The multi-channel audio signal is downmixed using the first downmixing parameter to obtain a first audio downmixing signal. The first audio downmixing signal includes a surround channel related signal and a channel fusion signal. The channel fusion signal includes a center channel and a low-frequency channel. The surround channel related signal and the channel fusion signal are in different transmission channels. The multi-channel audio signal is processed using the first audio downmixing signal to obtain an audio spatial parameter stream; The multi-channel audio signal is downmixed using the second downmixing parameter to obtain a second audio downmixed signal. The second audio downmixed signal includes a stereo channel signal and the channel fusion signal, and the stereo channel signal and the channel fusion signal are in different transmission channels. The second audio downmix signal and the audio spatial parameter stream are encapsulated to obtain an audio code, which corresponds to the multi-channel audio signal.

2. The method according to claim 1, wherein encapsulating the second audio downmixing signal and the audio spatial parameter stream to obtain audio encoding comprises: The stereo channel signals in the second audio downmix signal are paired to obtain paired stereo; The paired stereo signals are stereo encoded to obtain a stereo bitstream, and the stereo bitstream is encapsulated into a first-layer bitstream. The channel fusion signal is encoded in a single channel to obtain a channel fusion bitstream, and the channel fusion bitstream and the audio spatial parameter stream are encapsulated into a second-layer bitstream; The audio encoding includes a first layer bitstream and a second layer bitstream.

3. The method according to claim 1, wherein obtaining the first downmixing parameter comprises: The parameters of the first channel signal, the second channel signal, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters are analyzed. The first surround signal related parameters correspond to the first channel signal parameters, and the second surround signal related parameters correspond to the second channel signal parameters. The first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters are all from multi-channel audio signals. Based on the positions of the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters in the multi-channel audio signal, the first downmixing parameters are obtained, and the first downmixing parameters are represented by a first downmixing matrix. The first downmixing matrix is ​​used to incorporate the first channel signal parameters and the first surround signal related parameters into the first transmission channel, incorporate the second channel signal parameters and the second surround signal related parameters into the second transmission channel, and incorporate the center channel signal parameters and the low-frequency channel signal parameters into the third transmission channel; the first audio downmixing signal includes the first transmission channel, the second transmission channel, and the third transmission channel.

4. The method according to claim 3, wherein obtaining the second downmixing parameter comprises: Based on the positions of the first channel signal parameters, the second channel signal parameters, the first surround signal related parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters in the multi-channel audio signal, the second downmixing parameters are obtained, and the second downmixing parameters are represented by a second downmixing matrix. The second downmixing matrix is ​​used to incorporate the first channel signal parameters, the first surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fourth transmission channel; to incorporate the second channel signal parameters, the second surround signal related parameters, the center channel signal parameters, and the low-frequency channel signal parameters into the fifth transmission channel; and to incorporate the center channel signal parameters and the low-frequency channel signal parameters into the sixth transmission channel. The second audio downmixing signal includes the fourth, fifth, and sixth transmission channels.

5. A multi-channel audio decoding method, comprising: Obtain the audio encoding to be decoded and the type of the decoding end, wherein the audio encoding to be decoded is obtained based on the multi-channel audio encoding method according to any one of claims 1-4; Based on the type of the decoding terminal, the audio encoding is analyzed to obtain the first channel signal; Decode the first channel signal to obtain a multi-channel audio signal, the multi-channel audio signal corresponding to the audio code to be decoded.

6. The method according to claim 5, wherein parsing the audio encoding based on the decoding end type to obtain the first channel signal comprises: In response to the decoding end type being the first type, the stereo channel signal is parsed, and the channel fusion signal and audio spatial parameter stream are rejected for parsing. The stereo channel signal, the channel fusion signal, and the audio spatial parameter stream are all part of the audio encoding to be decoded. The first channel signal includes the stereo channel signal.

7. The method according to claim 6, wherein parsing the audio encoding based on the decoding end type to obtain the first channel signal comprises: In response to the decoding end type being the second type, the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream are parsed. The first channel signal includes the stereo channel signal, the channel fusion signal, and the audio spatial parameter stream.

8. The method according to claim 7, wherein decoding the first channel signal to obtain a multi-channel audio signal comprises: Remove the channel fusion signal from the stereo channel signal to obtain the first stereo signal; The first stereo signal is fused with the channel fusion signal to obtain a multi-channel signal; The multi-channel signal and the audio spatial parameter stream are up-mixed using a preset up-mixing algorithm to obtain the multi-channel audio signal.

9. A multi-channel audio encoding device, comprising: The first acquisition module is used to acquire a multi-channel audio signal, a first downmixing parameter and a second downmixing parameter, wherein the multi-channel audio signal is a signal to be encoded, and the first downmixing parameter and the second downmixing parameter correspond to the multi-channel audio signal. The first downmixing processing module is used to downmix the multi-channel audio signal using the first downmixing parameters to obtain a first audio downmixing signal. The first audio downmixing signal includes a surround channel related signal and a channel fusion signal. The channel fusion signal includes a center channel and a low-frequency channel. The surround channel related signal and the channel fusion signal are in different transmission channels. The spatial parameter determination module is used to process the multi-channel audio signal using the first audio downmixing signal to obtain an audio spatial parameter stream; The second downmixing processing module is used to downmix the multi-channel audio signal using the second downmixing parameters to obtain a second audio downmixing signal. The second audio downmixing signal includes a stereo channel signal and the channel fusion signal, and the stereo channel signal and the channel fusion signal are in different transmission channels. The stream encapsulation module is used to encapsulate the second audio downmixing signal and the audio spatial parameter stream to obtain an audio code, which corresponds to the multi-channel audio signal.

10. A multi-channel audio decoding device, comprising: The second acquisition module is used to acquire the audio encoding to be decoded and the decoding end type, wherein the audio encoding to be decoded is obtained based on the multi-channel audio encoding method according to any one of claims 1-4; The signal matching module is used to parse the audio code based on the decoding end type to obtain the first channel signal; An audio decoding module is used to decode the first channel signal to obtain a multi-channel audio signal, wherein the multi-channel audio signal corresponds to the audio code to be decoded.

11. An electronic device, comprising: The system includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the multi-channel audio encoding method of any one of claims 1 to 4, or the multi-channel audio decoding method of any one of claims 5 to 8.

12. A computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a computer to perform the multi-channel audio encoding method of any one of claims 1 to 4, or the multi-channel audio decoding method of any one of claims 5 to 8.

13. A computer program product comprising computer instructions for causing a computer to perform the multi-channel audio encoding method of any one of claims 1 to 4, or the multi-channel audio decoding method of any one of claims 5 to 8.