A method and apparatus for encoding and decoding a multi-channel signal and a terminal apparatus

By acquiring the silence marker information of multi-channel signals for encoding processing and dynamically adjusting the bit allocation, the problems of low encoding efficiency and resource waste in existing technologies are solved, and more efficient multi-channel signal encoding is achieved.

CN116798438BActive Publication Date: 2026-01-27HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210699863.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-14
Filing Date
2022-06-20
Publication Date
2026-01-27
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

Existing multi-channel signal coding schemes use a fixed bit allocation method, resulting in low coding efficiency and wasted coding bit resources, and are unable to adapt to the differences in input signal characteristics at different times and between different channels.

Method used

By acquiring the mute marker information of the multi-channel signal, including the mute enable flag and the mute flag, multi-channel encoding processing is performed, and a bitstream is generated. The bit allocation is dynamically adjusted to improve encoding efficiency and resource utilization.

Benefits of technology

It improves the coding efficiency and utilization of coding bit resources for multi-channel signals, reduces the waste of coding bit resources, and adapts to the changes in input signal characteristics at different times and between different channels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116798438B_ABST
    Figure CN116798438B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a kind of encoding method and codec equipment and terminal equipment of multi-channel signal, wherein, a kind of encoding method of multi-channel signal, comprising: obtaining the silence mark information of multi-channel signal, the silence mark information includes: silence enable flag, and / or silence flag;The multi-channel signal is carried out multi-channel encoding processing, to obtain the transmission channel signal of each transmission channel;According to the transmission channel signal of each transmission channel and the silence mark information generation bitstream, the bitstream includes: the silence mark information and the multi-channel encoding result of transmission channel signal.The embodiment of the present application is encoded according to the silence mark information to each transmission channel transmission channel signal to generate bitstream, considers the silence condition of multi-channel signal, so as to improve encoding efficiency and coding bit resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202210254868.9, filed on March 14, 2022, entitled "A method for encoding and decoding multi-channel signals, a terminal device, and a network device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of audio encoding and decoding, and in particular to a method, device, and terminal device for encoding and decoding multi-channel signals. Background Technology

[0003] Audio data compression is an indispensable part of media applications such as media communication and media broadcasting. Audio data compression can be achieved through multichannel coding, which can encode a bed signal with multiple channels, or it can encode multiple object audio signals. Multichannel coding can also encode a mixed signal that simultaneously contains bed signals and object audio signals.

[0004] A bed signal, an object signal, or a mixed signal containing both bed and object audio signals can all be input into the audio channel as a multi-channel signal. However, the characteristics of multi-channel signals cannot be exactly the same, and the characteristics of multi-channel signals are constantly changing.

[0005] Currently, the multi-channel signals mentioned above are processed using a fixed encoding scheme, such as a uniform bit allocation scheme, and the multi-channel signals are quantized and encoded based on the bit allocation results. While the uniform bit allocation scheme has the advantages of simplicity and ease of operation, it suffers from low encoding efficiency and wasted encoding bit resources. Summary of the Invention

[0006] This application provides a method, apparatus, and terminal device for encoding and decoding multi-channel signals, which can improve encoding efficiency and utilization of encoding bit resources.

[0007] To address the aforementioned technical problems, this application provides the following technical solutions:

[0008] In a first aspect, embodiments of this application provide a method for encoding multi-channel signals, comprising:

[0009] Obtain mute flag information from multi-channel signals to obtain mute flag information, wherein the mute flag information includes: mute enable flag, and / or mute flag;

[0010] The multi-channel signal is subjected to multi-channel encoding processing to obtain the transmission channel signal of each transmission channel;

[0011] A bitstream is generated based on the transmission channel signals of each transmission channel and the silence marker information. The bitstream includes the multi-channel quantization encoding results of the silence marker information and the transmission channel signals of each transmission channel.

[0012] In the above scheme, the silence marker information of the multi-channel signal includes: a silence enable flag and / or a silence flag; the multi-channel signal is subjected to multi-channel encoding processing to obtain the transmission channel signal of each transmission channel; a bitstream is generated based on the transmission channel signal of each transmission channel and the silence marker information, the bitstream including: the silence marker information and the multi-channel quantization encoding result of the transmission channel signal of each transmission channel. In this embodiment, the transmission channel signal of each transmission channel is encoded based on the silence marker information to generate a bitstream, taking into account the silence situation of the multi-channel signal, thus improving the encoding efficiency and the utilization rate of encoding bit resources.

[0013] In one possible implementation, the multi-channel signal includes: a sound bed signal, and / or an object signal;

[0014] The mute flag information includes: the mute enable flag; the mute enable flag includes: a global mute enable flag, or a partial mute enable flag, wherein...

[0015] The global mute enable flag is a mute enable flag that applies to the multi-channel signal; or,

[0016] The partial mute enable flag is a mute enable flag that operates on a portion of the channels in the multi-channel signal.

[0017] In one possible implementation, when the mute enable flag is the partial mute enable flag...

[0018] The partial mute enable flag is either an object mute enable flag applied to the object signal, or a sound bed mute enable flag applied to the sound bed signal, or a mute enable flag applied to other channel signals in the multi-channel signal that do not contain non-low frequency effect LFE channel signals, or a mute enable flag applied to channel signals participating in pairing in the multi-channel signal.

[0019] In the above scheme, the global mute enable flag or the partial mute enable flag can be used to indicate the mute of the sound bed signal and / or the object signal, thereby improving the encoding efficiency by performing subsequent encoding processing, such as bit allocation, based on the global mute enable flag or the partial mute enable flag.

[0020] In one possible implementation, the multi-channel signal includes: a sound bed signal and an object signal;

[0021] The mute flag information includes: the mute enable flag; the mute enable flag includes: a sound bed mute enable flag and an object mute enable flag.

[0022] The mute enable flag occupies a first bit and a second bit. The first bit is used to carry the value of the sound bed mute enable flag, and the second bit is used to carry the value of the object mute enable flag.

[0023] In the above scheme, the mute enable flag can use different bits to indicate the specific implementation of the mute enable flag. For example, a first bit and a second bit can be predefined. Through the above different bits, the mute enable flag can be indicated as a bed mute enable flag and an object mute enable flag.

[0024] In one possible implementation, the mute flag information includes: the mute enable flag;

[0025] The mute enable flag is used to indicate whether the mute marker detection function is enabled; or...

[0026] The mute enable flag is used to indicate whether each channel of the multi-channel signal needs to be muted; or,

[0027] The mute enable flag is used to indicate whether each channel of the multi-channel signal is a non-mute channel.

[0028] In the above scheme, the mute enable flag is used to indicate whether the mute detection function is enabled. For example, when the mute enable flag is at a first value (e.g., 1), it indicates that the mute detection function is enabled, further detecting the mute status of each channel of the multi-channel signal. When the mute enable flag is at a second value (e.g., 0), it indicates that the mute detection function is disabled.

[0029] In the above scheme, the mute enable flag can also be used to indicate whether each channel of the multi-channel signal is a non-mute channel. For example, when the mute enable flag is a first value (e.g., 1), it indicates that the mute flag of each channel needs to be further checked. When the mute enable flag is a second value (e.g., 0), it indicates that each channel of the multi-channel signal is a non-mute channel.

[0030] In one possible implementation, acquiring the silence marker information of the multi-channel signal includes:

[0031] The silence marker information is obtained according to the control signaling of the input encoding device; or...

[0032] The silence marker information is obtained according to the encoding parameters of the encoding device; or...

[0033] Silence marker detection is performed on each channel of the multi-channel signal to obtain the silence marker information.

[0034] In the above scheme, control signals can be input into the encoding device to determine silence marker information. The silence marker information can be controlled by external input, or the encoding device can include encoding parameters (also called encoder parameters). These encoding parameters can be used to determine the silence marker information and can be preset based on encoder parameters such as encoding rate and encoding bandwidth. Alternatively, the silence marker information can also be determined based on the silence detection results of each channel. This application does not limit the implementation method of the silence marker information in the embodiments.

[0035] In one possible implementation, the mute flag information includes: the mute enable flag and the mute flag;

[0036] The process of detecting silence markers in each channel of the multi-channel signal to obtain silence marker information includes:

[0037] The mute marker detection is performed on each channel of the multi-channel signal to obtain the mute marker of each channel;

[0038] The mute enable flag is determined based on the mute flag of each channel.

[0039] In the above scheme, the encoder can first detect the mute flag of each channel, which indicates whether each channel is muted. After determining the mute flag of each channel, the mute enable flag is determined based on the mute flag of each channel. Based on the above method, the mute enable flag can be generated, thereby generating mute marking information.

[0040] In one possible implementation, the mute flag information includes: the mute flag; or, the mute flag information includes: the mute enable flag and the mute flag;

[0041] The mute flag is used to indicate whether each channel affected by the mute enable flag is a mute channel, wherein the mute channel is a channel that does not require encoding or a channel that requires low-bit encoding.

[0042] In the above scheme, when the mute flag value is a first value (e.g., 1), it indicates that the channel to which the mute enable flag is applied is a mute channel; when the mute flag value is a second value (e.g., 0), it indicates that the channel to which the mute enable flag is applied is a non-mute channel. When the mute flag value is a first value (e.g., 1), the channel is not encoded or is encoded using lower bits.

[0043] In one possible implementation, before acquiring the silence marker information of the multi-channel signal, the method further includes:

[0044] The multi-channel signal is preprocessed to obtain a preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding.

[0045] The acquisition of silence marker information for multi-channel signals includes:

[0046] The preprocessed multi-channel signal is subjected to silence marker detection to obtain the silence marker information.

[0047] In the above scheme, the coding efficiency of multi-channel signals can be improved through the above preprocessing process.

[0048] In one possible implementation, the method further includes:

[0049] The multi-channel signal is preprocessed to obtain a preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding.

[0050] The silence marker information is corrected based on the preprocessed multi-channel signal.

[0051] In the above scheme, after preprocessing, the silence mark information can be corrected based on the preprocessing results. For example, after frequency domain noise shaping, the energy of a certain channel of the multi-channel signal changes, and the silence mark detection result of that channel can be adjusted to correct the silence mark information.

[0052] In one possible implementation, generating the bitstream based on the transmission channel signals of each transmission channel and the silence marker information includes:

[0053] The initial multi-channel processing method is adjusted according to the mute mark information to obtain the adjusted multi-channel processing method;

[0054] The multi-channel signal is encoded according to the adjusted multi-channel processing method to obtain the bitstream.

[0055] In the above scheme, the encoding end can adjust the initial multi-channel processing method based on the silence flag information, and then encode the multi-channel signal according to the adjusted multi-channel processing method, thereby improving encoding efficiency. For example, in the process of screening multi-channel signals, channels with a silence flag of 1 do not participate in pairing screening.

[0056] In one possible implementation, generating the bitstream based on the transmission channel signals of each transmission channel and the silence marker information includes:

[0057] Based on the mute flag information, the number of available bits, and the multi-channel side information, bit allocation is performed for each transmission channel to obtain the bit allocation result for each transmission channel;

[0058] The transmission channel signals of each transmission channel are encoded according to the bit allocation results of each channel to obtain the bit stream.

[0059] In the above scheme, the encoding end allocates bits based on silence marker information, available bits, and multi-channel side information; it then encodes the bitstream based on the bit allocation results for each transmission channel. The specific content of this bit allocation strategy is not limited. For example, the encoding of the transmission channel signal can be multi-channel quantization encoding. In this embodiment, the specific implementation of multi-channel quantization encoding can be that the group-pair downmixed signal is transformed by a neural network to obtain latent features; the latent features are then quantized and interval encoded. Alternatively, multi-channel quantization encoding can be based on vector quantization to quantize and encode the group-pair downmixed signal.

[0060] In one possible implementation, the bit allocation for each transmission channel based on the silence marker information, the number of available bits, and the multi-channel side information includes:

[0061] Based on the available number of bits and multi-channel side information, bits are allocated to each transmission channel according to the bit allocation strategy corresponding to the mute marker information.

[0062] In the above scheme, bit allocation based on silence marker information can be performed by first allocating bits based on the total available bits and the signal characteristics of each transmission channel, combined with a bit allocation strategy. Then, the bit allocation result is adjusted according to the silence marker information. By adjusting the bit allocation, the transmission efficiency of multi-channel signals can be improved.

[0063] In one possible implementation, the multi-channel side information includes: a channel bit allocation ratio field.

[0064] The channel bit allocation ratio field is used to indicate the bit allocation ratio between non-low frequency effect (LFE) channels in a multi-channel signal.

[0065] In the above scheme, the channel bit allocation ratio field can indicate the bit allocation ratio of all channels in the multi-channel signal except for the LFE channel, thereby determining the number of bits for each non-LFE channel.

[0066] In one possible implementation, the silence marker detection for each channel of the multi-channel signal includes:

[0067] The signal energy of each channel in the current frame is determined based on the input signal of each channel in the current frame of the multi-channel signal.

[0068] Based on the signal energy of each channel in the current frame, determine the silence detection parameters for each channel in the current frame;

[0069] Based on the silence detection parameters of each channel in the current frame and the preset silence detection threshold, the silence flag of each channel in the current frame is determined.

[0070] In the above scheme, the silence detection parameters of each channel in the current frame are compared with the silence detection threshold. Taking the silence flag detection of the first channel of the current frame as an example, if the silence detection parameter of the first channel of the current frame is less than the silence detection threshold, then the first channel of the current frame is a silent frame, that is, the first channel is a silent channel at the current moment, and the silence flag muteFlag[1] of the first channel of the current frame is a first value (e.g., 1). If the silence detection parameter of the first channel of the current frame is greater than or equal to the silence detection threshold, then the first channel of the current frame is a non-silent frame, that is, the first channel is a non-silent channel at the current moment, and the silence flag muteFlag[1] of the first channel of the current frame is a second value (e.g., 0).

[0071] In one possible implementation, the multi-channel encoding process performed on the multi-channel signal to obtain the transmission channel signal of each transmission channel includes:

[0072] The multi-channel signal is filtered to obtain the filtered multi-channel signal;

[0073] The filtered multi-channel signals are then processed to obtain multi-channel paired signals and multi-channel side information.

[0074] The multi-channel group signal is downmixed based on the multi-channel side information to obtain the transmission channel signal of each transmission channel.

[0075] In the above scheme, the encoding device filters the multi-channel signals, for example, filtering out multi-channel signals that do not participate in multi-channel pairing, to obtain filtered multi-channel signals. The filtered multi-channel signals can be multi-channel signals that participate in pairing; for example, the filtered channels do not include the LFE channel. After filtering the multi-channel signals, they can be paired, for example, ch1 and ch2 can be paired to obtain multi-channel paired signals. After generating the multi-channel paired signals, downmixing is performed. The specific downmixing process will not be described in detail, but the transmission channel signals of each transmission channel can be obtained. In this embodiment, the transmission channel can be the downmixed channel of the multi-channel pairing.

[0076] In one possible implementation, the multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index;

[0077] The inter-channel amplitude difference parameter quantization codebook index is used to indicate the codebook index for the quantization of the inter-channel amplitude difference (ILD) parameter of each channel in the multi-channel signal.

[0078] The number of channel pairs is used to represent the number of channel pairs in the current frame of the multichannel signal.

[0079] The channel pair index is used to represent the index of a channel pair.

[0080] In the above scheme, the number of bits occupied by the codebook index of the inter-channel amplitude difference parameter quantization is not limited in this embodiment. For example, the inter-channel amplitude difference parameter quantization codebook index occupies 5 bits. The inter-channel amplitude difference parameter quantization codebook index can be represented as mcIld[ch1] and mcIld[ch2], occupying 5 bits. It is the codebook index of the inter-channel amplitude difference ILD parameter quantization for each channel in the current channel pair, used to recover the amplitude of the decoded spectrum. The number of bits occupied by the number of channel pairs is not limited in this embodiment. For example, the number of channel pairs occupies 4 bits, represented as pairCnt, occupying 4 bits, used to represent the number of channel pairs in the current frame. The number of bits occupied by the channel pair index is not limited in this embodiment. For example, the channel pair index is represented as channelPairIndex. The number of bits of channelPairIndex is related to the total number of channels and is used to represent the index of the channel pair. The index values ​​of the two channels in the current channel pair, namely ch1 and ch2, can be obtained by parsing.

[0081] Secondly, embodiments of this application provide a method for decoding multi-channel signals, including:

[0082] The mute flag information is parsed from the bitstream of the encoding device, and the encoding information of each transmission channel is determined based on the mute flag information. The mute flag information includes: a mute enable flag, and / or a mute flag.

[0083] The encoded information of each transmission channel is decoded to obtain the decoded signal of each transmission channel;

[0084] The decoding signals of each transmission channel are subjected to multi-channel decoding processing to obtain multi-channel decoded output signals.

[0085] In the above scheme, the decoding end can obtain the silence mark information from the bitstream of the encoding end in the embodiment of this application, so that the decoding end can perform decoding processing in the same way as the encoding end.

[0086] In one possible implementation, parsing the silence marker information from the bitstream of the encoding device includes:

[0087] Parse the mute flags of each channel from the bitstream; or,

[0088] The mute enable flag is parsed from the bitstream. If the mute enable flag is a first value, the mute flag is parsed from the bitstream; or...

[0089] Parse the bed mute enable flag and / or object mute enable flag, as well as the mute flag for each channel, from the bitstream; or,

[0090] Parse the bed mute enable flag and / or object mute enable flag from the bitstream; based on the bed mute enable flag and / or object mute enable flag, parse the mute flags of some channels of each channel from the bitstream.

[0091] In the above scheme, the code-end parses the mute flag information from the bitstream of the encoding device. Depending on the specific content of the mute flag information generated by the encoding device, the mute flag information obtained by the decoding end corresponds to that of the encoding side. Specifically, in one approach, the mute flag indicates whether each channel is a mute channel. A mute channel is a channel that does not need encoding or a channel that needs to be encoded using low-bit encoding. The decoding end can parse the mute flag of each channel from the bitstream. In another approach, the mute enable flag can also be used to indicate whether each channel is a non-mute channel. For example, when the mute enable flag is a first value (e.g., 1), it indicates that further detection of the mute flag of each channel is needed. When the mute enable flag is a second value (e.g., 0), it indicates that each channel is a non-mute channel. The decoding end parses the mute enable flag from the bitstream; if the mute enable flag is the first value, the mute flag is parsed from the bitstream. In one embodiment, the mute enable flag includes: a bed mute enable flag and / or an object mute enable flag. The decoder parses the bed mute enable flag and / or the object mute enable flag, as well as the mute flags for each channel, from the bitstream. In another embodiment, the decoder parses the bed mute enable flag and / or the object mute enable flag from the bitstream; based on the bed mute enable flag and / or the object mute enable flag, it parses the mute flags for some channels from the bitstream.

[0092] In one possible implementation, decoding the encoded information of each transmission channel includes:

[0093] Multi-channel side information is parsed from the bitstream;

[0094] Bit allocation is performed on each transmission channel based on the multi-channel side information and the mute flag information to obtain the number of encoded bits for each channel;

[0095] The encoded information of each transmission channel is decoded according to the number of encoded bits of each channel.

[0096] In the above scheme, the bitstream can also include multi-channel side information. The decoding end can allocate bits to each transmission channel according to the multi-channel side information and the mute flag information to obtain the number of encoded bits for each transmission channel. The number of encoded bits obtained by the decoding end is the same as the number of encoded bits preset by the encoding end. Then, the encoded information of each transmission channel is decoded according to the number of encoded bits for each transmission channel, thereby realizing the decoding of the transmission channel signal of each transmission channel.

[0097] In one possible implementation, after performing multi-channel decoding processing on the decoded signals of each transmission channel to obtain multi-channel decoded output signals, the method further includes:

[0098] The multi-channel decoded output signal is post-processed, and the post-processing includes at least one of the following: bandwidth extension decoding, inverse time-domain noise shaping, inverse frequency-domain noise shaping, and inverse time-frequency transformation.

[0099] In the above scheme, the post-processing of the multi-channel decoded output signal is the reverse of the pre-processing process at the encoding end, and the specific processing method is no longer limited.

[0100] In one possible implementation, the multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index;

[0101] The inter-channel amplitude difference parameter quantization codebook index is used to indicate the codebook index for the quantization of the inter-channel amplitude difference (ILD) parameter of each channel.

[0102] The number of channel pairs is used to represent the number of channel pairs in the current frame of the multichannel signal.

[0103] The channel pair index is used to represent the index of a channel pair.

[0104] Thirdly, embodiments of this application provide an encoding device, the encoding device comprising:

[0105] A mute marker detection module is used to acquire mute marker information of multi-channel signals, wherein the mute marker information includes: a mute enable flag, and / or a mute flag;

[0106] A multi-channel encoding module is used to perform multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel;

[0107] The bitstream generation module is used to generate a bitstream based on the transmission channel signals of each transmission channel and the silence marker information. The bitstream includes the multi-channel quantization encoding results of the silence marker information and the transmission channel signals.

[0108] Fourthly, embodiments of this application provide a decoding device, the decoding device comprising:

[0109] The parsing module is used to parse the mute marker information from the bitstream of the encoding device and determine the encoding information of each transmission channel based on the mute marker information. The mute marker information includes: a mute enable flag and / or a mute flag.

[0110] The inverse quantization module is used to decode the encoded information of each transmission channel to obtain the decoded signal of each transmission channel;

[0111] The multi-channel decoding module is used to perform multi-channel decoding processing on the decoding signals of each transmission channel to obtain a multi-channel decoded output signal.

[0112] Fifthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect above.

[0113] In a sixth aspect, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the first or second aspect above.

[0114] In a seventh aspect, embodiments of this application provide a communication device, which may include entities such as terminal devices or chips. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the first or second aspects above.

[0115] Eighthly, embodiments of this application provide a computer-readable storage medium storing a bitstream generated by the method of the first aspect.

[0116] Ninthly, this application provides a chip system including a processor for supporting encoding / decoding devices in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the encoding / decoding device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description

[0117] Figure 1 A schematic diagram of the composition structure of a multi-channel signal processing system provided in an embodiment of this application;

[0118] Figure 2a A schematic diagram illustrating the application of the audio encoder and audio decoder provided in this application to a terminal device;

[0119] Figure 2b A schematic diagram illustrating the application of the audio encoder provided in this application to a wireless device or a core network device;

[0120] Figure 2c A schematic diagram illustrating the application of the audio decoder provided in this application embodiment to a wireless device or core network device;

[0121] Figure 3a A schematic diagram illustrating the application of the multi-channel encoder and multi-channel decoder provided in the embodiments of this application to a terminal device;

[0122] Figure 3b A schematic diagram illustrating the application of a multi-channel encoder provided in this application to a wireless device or a core network device;

[0123] Figure 3c A schematic diagram illustrating the application of the multi-channel decoder provided in this application to a wireless device or a core network device;

[0124] Figure 4 A schematic diagram illustrating a multi-channel signal encoding method provided in an embodiment of this application;

[0125] Figure 5 A schematic diagram illustrating a multi-channel signal decoding method provided in an embodiment of this application;

[0126] Figure 6 A schematic diagram illustrating a multi-channel signal encoding process provided in an embodiment of this application;

[0127] Figure 7 A schematic diagram illustrating a multi-channel signal encoding process provided in an embodiment of this application;

[0128] Figure 8 A schematic diagram illustrating a multi-channel signal decoding process provided in an embodiment of this application;

[0129] Figure 9 A schematic diagram illustrating a multi-channel signal decoding process provided in an embodiment of this application;

[0130] Figure 10 This is a schematic diagram of the composition structure of an encoding device provided in an embodiment of this application;

[0131] Figure 11 This is a schematic diagram of the composition structure of a decoding device provided in an embodiment of this application;

[0132] Figure 12 This is a schematic diagram of the composition structure of another encoding device provided in an embodiment of this application;

[0133] Figure 13 This is a schematic diagram of the composition structure of another decoding device provided in an embodiment of this application. Detailed Implementation

[0134] This application provides a method for encoding and decoding multi-channel signals, a terminal device, and a network device to improve encoding efficiency and the utilization rate of encoding bit resources.

[0135] The embodiments of this application will now be described with reference to the accompanying drawings.

[0136] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0137] Sound is a continuous wave produced by the vibration of an object. The object that produces vibrations and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solids, or liquids), the auditory organs of humans or animals can perceive the sound.

[0138] Sound waves are characterized by pitch, intensity, and timbre. Pitch indicates the highness or lowness of a sound. Intensity indicates the loudness or volume of a sound. The unit of intensity is the decibel (dB). Timbre is also known as tone color.

[0139] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is hertz (Hz). The human ear can distinguish sounds with frequencies between 20 Hz and 20,000 Hz.

[0140] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.

[0141] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, and pulse waves, among others.

[0142] Based on the characteristics of sound waves, sound can be divided into regular sound and irregular sound. Irregular sound refers to sound emitted by the irregular vibration of a sound source. Irregular sound is, for example, noise that affects people's work, study, and rest. Regular sound refers to sound emitted by the regular vibration of a sound source. Regular sound includes speech and musical tones. When sound is represented electronically, regular sound is an analog signal that varies continuously in the time and frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.

[0143] Because human hearing has the ability to distinguish the location of sound sources in space, when a listener hears a sound in space, in addition to being able to perceive the pitch, intensity, and timbre of the sound, they can also perceive the location of the sound.

[0144] Sound can also be categorized as mono or stereo. Mono has one sound channel, using a single microphone to pick up sound and a single speaker to reproduce it. Stereo has multiple sound channels, with different channels transmitting different sound waveforms. A sound channel can also be simply referred to as a channel or audio path. For example, a multi-channel signal can include signals from each of its channels, which can also be called a channel. In subsequent embodiments of this application, "channel" and "audio path" have the same meaning. After multi-channel signal undergoes multi-channel encoding, the transmission channel signals for each transmission channel are obtained. This transmission channel refers to the channel after multi-channel encoding. Furthermore, multi-channel encoding can include channel pairing and downmixing; therefore, the transmission channel can also be called the channel pairing and the downmixed channel. See the description of the multi-channel encoding process in subsequent embodiments for details.

[0145] This application applies to the field of audio encoding and decoding, particularly multichannel encoding. Multichannel encoding can encode a bed signal with multiple channels, such as 5.1 channels, 5.1.4 channels, 7.1 channels, 7.1.4 channels, 22.2 channels, etc. Multichannel encoding can also encode multiple object audio signals. Furthermore, multichannel encoding can encode a mixed signal that simultaneously contains a bed signal and / or object audio signals.

[0146] The 5.1 channel includes the center channel (C), the front left channel (L), the front right channel (R), the rear left surround channel (LS), the rear right surround channel (RS), and the 0.1 (LFE) channel.

[0147] The 5.1.4 channel is based on the 5.1 channel and adds the following channels: left high channel, right high channel, left high surround channel, and right high surround channel.

[0148] 7.1 channels include the center channel (C), front left channel (L), front right channel (R), rear left surround channel (LS), rear right surround channel (RS), left rear channel (LB), right rear channel (RB), and 0.1 channel LFE channel.

[0149] 7.1.4 channels are based on 7.1 channels with the addition of 4 height channels.

[0150] 22.2 channel is a multi-channel format that includes three layers with a total of 22 channels and 2 LFE channels.

[0151] A mixed signal of bed-mode and object audio is a signal combination in 3D audio, used to meet the audio recording, transmission, and playback needs of complex scenarios such as filmmaking, sports events, and concerts. For example, the sound content of a sports broadcast is usually represented by a bed-mode signal, while the commentary of different commentators is usually represented by multiple object audio signals. Whether it is a bed-mode signal, an object audio signal, or a mixed signal containing both bed-mode and object audio signals, the characteristics of the input signals from different channels are not entirely the same at any given time, and the characteristics of the input signal from the same channel also change continuously from time to time.

[0152] Current multi-channel signals use a fixed encoding scheme, without considering the differences in input signal characteristics at different times and / or between different channels. For example, a uniform bit allocation scheme is used for processing, and the multi-channel signal is quantized and encoded based on the bit allocation result.

[0153] Using the same bit allocation scheme cannot adapt to changes in the characteristics of input signals between different channels at different times, resulting in low coding efficiency. For example, a multi-channel audio signal to be encoded contains 5.1.4 channel bed signals and 4 object signals. Among the 14 channels to be encoded, channels 0-9 are bed signals, and channels 10-13 are object signals. At one time, channels 6-9 and channels 11, 12, and 13 are silent channels (with little information perceptible to the ear), while the other channels contain the main audio information, i.e., non-silent channels. At another time, the silent channels become channels 10, 12, and 13, and the other channels contain the main audio information.

[0154] If the same bit allocation scheme is used at different times, some channels containing the main audio information may not have enough bits for encoding, while some silent channels may be allocated too many encoding bits, resulting in a waste of encoding bit resources.

[0155] This application provides an audio processing technology, particularly an audio coding technology for multi-channel signals, to improve traditional audio coding systems. A multi-channel signal refers to an audio signal comprising multiple channels, such as a stereo signal. Audio processing includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and includes encoding (e.g., compressing) the raw audio to reduce the amount of data required to represent the audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are collectively referred to as encoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0156] The technical solutions of this application embodiment can be applied to various audio processing systems, such as... Figure 1 The diagram shown is a schematic representation of the composition of an audio processing system provided in an embodiment of this application. The audio processing system 100 may include a multi-channel signal encoding device 101 and a multi-channel signal decoding device 102. The multi-channel signal encoding device 101, also known as an audio encoding device, is used to generate a bitstream. This audio encoded bitstream can then be transmitted to the multi-channel signal decoding device 102 via an audio transmission channel. The multi-channel signal decoding device 102, also known as a multi-audio decoding device, receives the bitstream and performs the audio decoding function of the multi-channel signal decoding device 102 to finally obtain the reconstructed signal.

[0157] In the embodiments of this application, the multi-channel signal encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the multi-channel signal encoding device can be an audio encoder for the aforementioned terminal device, wireless device, or core network device. Similarly, the multi-channel signal decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the multi-channel signal decoding device can be an audio decoder for the aforementioned terminal device, wireless device, or core network device. For example, the audio encoder can include a wireless access network, a core network media gateway, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The audio encoder can also be an audio encoder used in virtual reality (VR) streaming services.

[0158] In the application embodiment, taking the audio encoding module (audio encoding and audio decoding) applicable to virtual reality streaming (VR streaming) services as an example, the end-to-end audio signal encoding and decoding process includes: after the audio signal A passes through the acquisition module, it undergoes a preprocessing operation (audioPReprocessing). The preprocessing operation includes filtering out the low-frequency part in the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the location information in the signal, and then performing encoding processing (audio encoding), packaging (file / segment encapsulation), and then sending (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), and performs binaural rendering processing on the decoded signal. The rendered signal is mapped onto the listener's headphones, which can be independent headphones or headphones on a glasses device.

[0159] like Figure 2a The diagram illustrates the application of the audio encoder and audio decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used for channel encoding of the audio signal, and the channel decoder is used for channel decoding of the audio signal. For example, the first terminal device 20 may include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 may include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and a wireless or wired second network communication device 23 are connected via a digital channel. The second terminal device 21 is connected to the wireless or wired second network communication device 23. The aforementioned wireless or wired network communication device can broadly refer to signal transmission devices, such as communication base stations, data exchange devices, etc.

[0160] In audio communication, the transmitting terminal device first acquires audio, encodes the acquired audio signal, and then performs channel coding before transmitting it over a digital channel via a wireless network or core network. The receiving terminal device, acting as the receiver, decodes the received signal to obtain the bitstream, then recovers the audio signal through audio decoding for playback.

[0161] like Figure 2b The diagram illustrates the application of the audio encoder provided in this embodiment of the application in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, the audio encoder 253 provided in this embodiment of the application, and a channel encoder 254. The other audio decoders 252 refer to audio decoders other than the standard audio decoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded using the channel decoder 251, then audio-decoded using the other audio decoder 252, then audio-encoded using the audio encoder 253 provided in this embodiment of the application, and finally channel-encoded using the channel encoder 254. After channel encoding, the signal is transmitted out. The other audio decoders 252 perform audio decoding on the bitstream decoded by the channel decoder 251.

[0162] like Figure 2c The diagram illustrates the application of the audio decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in this embodiment, other audio encoders 256, and a channel encoder 254. The other audio encoders 256 refer to audio encoders other than the standard audio encoder. Within the wireless device or core network device 25, the channel decoder 251 first performs channel decoding on the incoming signal. Then, the audio decoder 255 decodes the received audio encoded bitstream. Next, the other audio encoders 256 perform audio encoding. Finally, the channel encoder 254 performs channel encoding on the audio signal before transmission. If transcoding is required in the wireless device or core network device, corresponding audio encoding processing is necessary. The wireless device refers to radio frequency (RF) related equipment in communication, and the core network device refers to core network related equipment in communication.

[0163] In some embodiments of this application, the multi-channel signal encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the multi-channel signal encoding device can be a multi-channel encoder of the aforementioned terminal device, wireless device, or core network device. Similarly, the multi-channel signal decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the multi-channel signal decoding device can be a multi-channel decoder of the aforementioned terminal device, wireless device, or core network device.

[0164] like Figure 3aThe diagram illustrates the application of the multi-channel encoder and multi-channel decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in this embodiment of the application, and the multi-channel decoder can execute the audio decoding method provided in this embodiment of the application. Specifically, the channel encoder is used to perform channel encoding on the multi-channel signal, and the channel decoder is used to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. The second terminal device 31 may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and a wireless or wired second network communication device 33 are connected via a digital channel. The second terminal device 31 is connected to the wireless or wired second network communication device 33. The aforementioned wireless or wired network communication equipment can broadly refer to signal transmission equipment, such as communication base stations and data switching equipment. In audio communication, the transmitting terminal device performs multi-channel encoding on the acquired multi-channel signal, followed by channel encoding, and then transmits it through a wireless network or core network in a digital channel. The receiving terminal device performs channel decoding on the received signal to obtain the multi-channel signal encoded bitstream, and then recovers the multi-channel signal through multi-channel decoding for playback by the receiving terminal device.

[0165] like Figure 3b The diagram shown illustrates the application of the multi-channel encoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, as described above. Figure 2b Similarly, this will not be elaborated upon here.

[0166] like Figure 3c The diagram shown illustrates the application of the multi-channel decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, as described above. Figure 2c Similarly, this will not be elaborated upon here.

[0167] The audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the acquired multi-channel signal can involve processing the acquired multi-channel signal to obtain an audio signal, and then encoding the obtained audio signal according to the method provided in this application embodiment. The decoding end decodes the audio signal based on the multi-channel signal encoded bitstream, and recovers the multi-channel signal after upmixing. Therefore, this application embodiment can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding is required, corresponding multi-channel encoding processing is necessary.

[0168] This application first introduces a multi-channel signal encoding method provided in its embodiments. This method can be executed by a terminal device, such as a multi-channel signal encoding device (hereinafter referred to as an encoding end, encoder, or encoding device; for example, the encoding end can be an artificial intelligence (AI) encoder). In this application embodiment, the multi-channel signal can include multiple channels, such as a first channel and a second channel, or multiple channels can include a first channel, a second channel, and a third channel, etc. Figure 4 The following describes the encoding process executed by the encoding device (or encoding end) in the embodiments of this application:

[0169] 401. Obtain the mute flag information of the multi-channel signal. The mute flag information includes: mute enable flag, and / or mute flag.

[0170] In this embodiment, after the encoding end receives a multi-channel signal, it can obtain the silence marker information of the multi-channel signal. This silence marker information indicates the silence status of the channels in the multi-channel signal. For example, silence marker detection can be performed on the multi-channel signal to determine whether it supports silence marking. The encoding end can generate silence marker information based on the multi-channel signal. This silence marker information can be used to guide subsequent encoding processes, such as bit allocation. The silence marker information can also be written into the bitstream by the encoding end and transmitted to the decoding end to ensure consistency between encoding and decoding processes.

[0171] In this embodiment, the mute marker information is used to indicate the mute status of multi-channel signals. The mute marker information can be implemented in various ways; for example, it can include a mute enable flag and / or a mute flag. The mute enable flag indicates whether mute detection is enabled, and the mute flag indicates whether each channel affected by the mute enable flag is a mute channel. The mute channel is either a channel that does not require encoding or a channel that requires low-bit encoding.

[0172] In some embodiments of this application, the multi-channel signal includes a sound bed signal and / or an object signal. Current encoding schemes do not consider the differences in input signal characteristics at different times and / or between different channels, and use a uniform encoding scheme for processing, resulting in low encoding efficiency. The mute enable flag provided in the embodiments of this application can indicate mute for the sound bed signal and / or object signal. Specifically, the mute flag information includes: a mute enable flag; the mute enable flag includes: a global mute enable flag, or a partial mute enable flag, wherein...

[0173] The global mute enable flag is a mute enable flag that applies to multi-channel signals; or,

[0174] The partial mute enable flag is a mute enable flag that applies to only some channels in a multi-channel signal.

[0175] The mute enable flag is denoted as HasSilFlag, and it can be a global mute enable flag or a partial mute enable flag. Using the aforementioned global or partial mute enable flag, mute indication can be provided for the bed signal and / or object signal. Subsequent encoding processing, such as bit allocation, can then be performed based on the global or partial mute enable flag, thereby improving encoding efficiency.

[0176] In some specific implementations, when the mute enable flag is a partial mute enable flag,

[0177] The partial mute enable flag is an object mute enable flag applied to the object signal, or a sound bed mute enable flag applied to the sound bed signal, or a partial mute enable flag applied to other channels in the multi-channel signal that do not contain low frequency effects (LFE) channels, or the partial mute enable flag applied to the mute enable flag of the channel signals participating in the pairing in the multi-channel signal.

[0178] For example, a global mute enable flag applies to all channels, while a partial mute enable flag applies to only certain channels. For example, an object mute enable flag applies to the channel corresponding to the object signal in a multi-channel signal, while a bed mute enable flag applies to the channel corresponding to the bed signal in a multi-channel signal. For example, an object mute enable flag that applies only to the object signal in a multi-channel signal is denoted as objMuteEna. Similarly, a bed mute enable flag that applies only to the bed signal in a multi-channel signal is denoted as bedMuteEna.

[0179] For example, the global mute enable flag is a mute enable flag applied to the multi-channel signal: when the multi-channel signal contains only the sound bed signal, the global mute enable flag is a mute enable flag applied to the sound bed signal; when the multi-channel signal contains only the object signal, the global mute enable flag is a mute enable flag applied to the object signal; when the multi-channel signal contains both the sound bed signal and the object signal, the global mute enable flag is a mute enable flag applied to both the sound bed signal and the object signal.

[0180] The partial mute enable flag applies to a portion of the channels in the multi-channel signal. These channels are pre-defined; for example, the partial mute enable flag may apply to the target signal as an object mute enable flag, or it may apply to the bed-of-sound mute enable flag, or it may apply to other channels in the multi-channel signal that do not contain the LFE channel signal. The partial mute enable flag may also apply to the channels in the multi-channel signal that participate in pairing. The specific method of pairing the multi-channel signals in this embodiment is not limited.

[0181] In some embodiments of this application, the multi-channel signal includes: a sound bed signal and an object signal;

[0182] The mute flag information includes: a mute enable flag; the mute enable flag includes: a sound bed mute enable flag and an object mute enable flag.

[0183] The mute enable flag occupies the first and second bits. The first bit is used to carry the value of the sound bed mute enable flag, and the second bit is used to carry the value of the object mute enable flag.

[0184] The mute enable flag can use different bits to indicate the specific implementation of the mute enable flag. For example, a first bit and a second bit can be predefined. The first bit is used to carry the value of the sound bed mute enable flag, and the second bit is used to carry the value of the object mute enable flag. Through the above different bits, it is possible to indicate that the mute enable flag is the sound bed mute enable flag and the object mute enable flag.

[0185] In some embodiments of this application, step 401, obtaining the silence marker information of the multi-channel signal, includes:

[0186] A1. Obtain the silence marker information according to the control signaling of the input encoding device; or...

[0187] A2. Obtain the silence marker information according to the encoding parameters of the encoding device; or...

[0188] A3. Perform silence mark detection on each channel of the multi-channel signal to obtain the silence mark information.

[0189] The encoding device can input control signals to determine silence marker information. This silence marker information can be controlled by external input, or the encoding device can include encoding parameters (also called encoder parameters). These parameters can be used to determine the silence marker information and can be preset based on encoder parameters such as encoding rate and encoding bandwidth. Alternatively, the silence marker information can be determined based on the silence detection results of each channel. This application does not limit the implementation method of the silence marker information in this embodiment.

[0190] In some embodiments of this application, the mute flag information includes: a mute enable flag;

[0191] The mute enable flag indicates whether the mute marker detection function is enabled.

[0192] The mute enable flag is used to indicate whether each channel sending multi-channel signals needs to be muted; or,

[0193] The mute enable flag is used to indicate whether each channel of a multi-channel signal is a non-mute channel.

[0194] The mute enable flag indicates whether mute detection is enabled. For example, a mute enable flag of 1 indicates that the mute detection function is enabled, further detecting the mute flag of each channel. A mute enable flag of 0 indicates that the mute detection function is disabled. Alternatively, the mute enable flag can be used to indicate whether all channels are non-mute channels. For example, a mute enable flag of 1 indicates that further detection of the mute flag of each channel is required. A mute enable flag of 0 indicates that all channels are non-mute channels.

[0195] In some embodiments of this application, the mute flag information includes: a mute enable flag and a mute flag;

[0196] Step A3 performs silence marker detection on each channel of the multi-channel signal to obtain silence marker information, including:

[0197] A31. Perform mute marker detection on each channel of the multi-channel signal to obtain the mute marker for each channel;

[0198] A32. Determine the mute enable flag based on the mute flag of each channel.

[0199] The encoding end can first detect the mute flag of each channel. The mute flag of each channel is used to indicate whether each channel is a mute channel. The mute flag of each channel is denoted as mutexg[ch], where ch is the channel number, ch = 0…N-1, and N is the total number of channels of the input signal to be encoded. The number of channels for the bed signal is M, the number of channels for the target channel is P, and the total number of channels N = M + P. The channel number of the bed signal is used to indicate whether each channel is a mute channel. For example, the signal to be encoded is a mixed signal containing the bed signal and the target signal. The bed signal is a 5.1.4 channel signal with 10 channels; the number of target signals is 4 with 4 channels; the total number of channels is 14. The channel numbers of the bed signal are from 0 to 9, and the channel numbers of the target signals are from 10 to 13. The mute flag mutexg[ch], ch = 0…13, corresponds to the mute flag of each channel and is used to indicate whether each channel is a mute channel. After determining the mute flag for each channel, the mute enable flag is determined based on the mute flag for each channel.

[0200] In some embodiments of this application, the mute flag information includes: a mute flag; or, the mute flag information includes: a mute enable flag and a mute flag;

[0201] The mute flag indicates whether each channel is a mute channel when the mute enable flag is active. A mute channel is either a channel that does not require encoding or a channel that requires low-bit encoding.

[0202] For example, the channel numbers for the sound bed signal are from 0 to 9, and the channel numbers for the object signal are from 10 to 13. The mute flag `muteflag[ch]`, where `ch` = 0…13, corresponds to the mute flag for each channel and indicates whether the channel to which the mute enable flag is active is a mute channel. A mute channel is a channel whose signal energy, decibels, or loudness is below the hearing threshold; it is a channel that does not require encoding or requires encoding with lower bits. When the mute flag value is the first value (e.g., 1), it indicates that the channel is a mute channel; when the mute flag value is the second value (e.g., 0), it indicates that the channel is not a mute channel. When the mute flag value is the first value (e.g., 1), the channel is not encoded or is encoded with lower bits.

[0203] In some embodiments of this application, step A3 performs silence marker detection on each channel of the multi-channel signal, including:

[0204] B1. Determine the signal energy of each channel in the current frame based on the input signals of each channel in the current frame of the multi-channel signal.

[0205] Based on the input signals of each channel in the current frame, the signal energy of each channel in the current frame is determined. In this embodiment, the value of the frame length is not limited.

[0206] B2. Determine the silence detection parameters for each channel of the current frame based on the signal energy of each channel in the current frame.

[0207] The silence detection parameters for each channel in the current frame are used to characterize the energy value, power value, decibel value, or loudness value of each channel signal in the current frame.

[0208] B3. Determine the mute flag for each channel of the current frame based on the mute detection parameters of each channel and the preset mute detection threshold.

[0209] The silence detection parameters of each channel in the current frame are compared with the silence detection threshold. Taking the silence flag detection of the first channel in the current frame as an example, if the silence detection parameter of the first channel in the current frame is less than the silence detection threshold, then the first channel in the current frame is a silent frame, that is, the first channel is a silent channel at the current moment, and the silence flag muteFlag[1] of the first channel in the current frame is the first value (e.g., 1). If the silence detection parameter of the first channel in the current frame is greater than or equal to the silence detection threshold, then the first channel in the current frame is a non-silent frame, that is, the first channel is a non-silent channel at the current moment, and the silence flag muteFlag[1] of the first channel in the current frame is the second value (e.g., 0).

[0210] 402. Perform multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel.

[0211] In this embodiment of the application, the encoding device can perform multi-channel encoding processing on multi-channel signals. There are various multi-channel encoding processes, as detailed in the examples of the following embodiments. Through the above encoding process, the transmission channel signals of each transmission channel can be obtained.

[0212] The specific implementation of multichannel quantization coding can be as follows: the group-pair downmixed signal is transformed by a neural network to obtain latent features; the latent features are then quantized and interval encoded. Alternatively, multichannel quantization coding can be based on vector quantization to quantize and encode the group-pair downmixed signal. This application does not limit this implementation.

[0213] In some embodiments of this application, step 402 performs multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel, including:

[0214] C1. Perform multi-channel signal filtering on the multi-channel signal to obtain the filtered multi-channel signal.

[0215] For example, the encoding device completes the screening of multi-channel signals. The screened signals are multi-channel signals that participate in the pairing. For example, the screened channels do not include LFE channels. There are no restrictions on the specific screening method.

[0216] C2. Perform pairing processing on the filtered multi-channel signals to obtain multi-channel pairing signals and multi-channel side information.

[0217] For example, the encoding device filters multi-channel signals. The filtered multi-channel signals can be multi-channel signals participating in pairing. After filtering, the multi-channel signals can be paired, for example, channel ch1 and channel ch2 can form a channel pair to obtain a multi-channel paired signal. The specific method of pairing processing is not limited in this invention. Multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index. Among them, the inter-channel amplitude difference parameter quantization codebook index is used to indicate the codebook index of the inter-channel amplitude difference (ILD) parameter quantization of each channel in each channel of the multi-channel signal; the number of channel pairs is used to indicate the number of channel pairs in the current frame of the multi-channel signal; and the channel pair index is used to indicate the index of the channel pair.

[0218] C3. Perform downmixing on the multi-channel group signal based on the multi-channel side information to obtain the transmission channel signal of each transmission channel.

[0219] After generating the multichannel pair signal and multichannel side information, the multichannel side information can be used to downmix the multichannel pair signal. The specific downmixing process will not be described in detail. Through the aforementioned multichannel pair and downmixing, the transmission channel signals of each transmission channel after multichannel pair downmixing can be obtained. Specifically, the transmission channel can refer to the channel after multichannel pair and downmixing.

[0220] In some embodiments of this application, before step 401 obtains the silence marker information of the multi-channel signal, the encoding method of the multi-channel signal executed by the encoding end further includes:

[0221] D1. Preprocess the multi-channel signal to obtain the preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding.

[0222] In the aforementioned implementation scenario of step D1, step 401 obtains the silence marker information of the multi-channel signal, including:

[0223] The preprocessed multi-channel signal is subjected to silence marker detection to obtain silence marker information.

[0224] The input signal for the silence marker detection can be the original multi-channel signal or a pre-processed multi-channel signal. Pre-processing may include, but is not limited to, transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding. The multi-channel signal can be either a time-domain signal or a frequency-domain signal. Through the above pre-processing, the coding efficiency of the multi-channel signal can be improved.

[0225] In some embodiments of this application, the encoding method for multi-channel signals executed at the encoding end further includes:

[0226] E1. Preprocess the multi-channel signal to obtain the preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding.

[0227] E2. Correct the silence marker information based on the preprocessed multi-channel signal.

[0228] The encoding end can preprocess the multi-channel signal. Preprocessing may include, but is not limited to, transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding. The multi-channel signal can be a time-domain signal or a frequency-domain signal. After preprocessing, the silence marker information in step 401 can be corrected based on the preprocessed multi-channel signal. For example, after frequency domain noise shaping, the signal energy of a certain channel of the multi-channel signal changes, and the silence marker detection result of that channel can be adjusted.

[0229] 403. Generate a bitstream based on the transmission channel signals and silence marker information of each transmission channel. The bitstream includes: silence marker information and multi-channel quantization encoding results of the transmission channel signals of each transmission channel.

[0230] The encoding end generates a bitstream that includes silence marker information, which allows the decoding end to obtain the silence marker information and decode the bitstream based on the silence marker information. This makes it easier for the decoding end to perform decoding processing in the same way as the encoding end, such as bit allocation.

[0231] In some embodiments of this application, step 403 generates a bitstream based on the transmission channel signal and silence marker information of each transmission channel, including:

[0232] F1. Adjust the initial multi-channel processing method according to the mute mark information to obtain the adjusted multi-channel processing method;

[0233] F2. Encode the multi-channel signal according to the adjusted multi-channel processing method to obtain the bitstream.

[0234] The encoding end can adjust the initial multi-channel processing method based on the silence flag information, and then encode the multi-channel signal according to the adjusted multi-channel processing method, thereby improving encoding efficiency. For example, in the process of filtering multi-channel signals, channels with a silence flag of 1 do not participate in pair filtering.

[0235] In some embodiments of this application, step 403 generates a bitstream based on the transmission channel signal and silence marker information of each transmission channel, including:

[0236] G1. Based on the mute flag information, the number of available bits, and the multi-channel side information, perform bit allocation for each transmission channel to obtain the bit allocation result for each transmission channel;

[0237] G2. Encode the transmission channel signals of each transmission channel according to the bit allocation results of each channel to obtain the bit stream.

[0238] The encoding end can use the silence marker information for bit allocation of the transmission channels. First, it performs initial bit allocation for each transmission channel based on the number of available bits and multi-channel side information. Then, it performs bit allocation again based on the silence marker information to obtain the bit allocation result of each transmission channel. The transmission channel signal is encoded based on the bit allocation result of each transmission channel to obtain the bit stream. This bit stream can be called the encoded bit stream or the bit stream of the multi-channel signal.

[0239] Furthermore, in some embodiments of this application, step G1 allocates bits to each transmission channel based on the silence marker information, the number of available bits, and the multi-channel side information, including:

[0240] G11. Based on the number of available bits and multi-channel side information, allocate bits to each transmission channel according to the bit allocation strategy corresponding to the silence marker information.

[0241] The encoding end can allocate bits to each transmission channel based on the silence flag information. The silence enable flag can be used to select different bit allocation strategies. The specific content of this bit allocation strategy is not limited, but an example is given below: Assuming the silence enable flags include the bed silence enable flag `bedMuteEna` and the object silence enable flag `objMuteEna`, bit allocation based on the silence flag information can first perform an initial bit allocation based on the total available bits and the signal characteristics of each transmission channel. Then, the bit allocation result is adjusted according to the silence flag information. This adjustment of bit allocation can improve the transmission efficiency of multi-channel signals. For example, if the object silence enable flag `objMuteEna` is 1, the bits initially allocated to the channel in the object signal where `muteflag` is 1 are allocated to the bed signal or other object channels. If both the bedmute enable flag (bedMuteEna) and the object mute enable flag are 1, the bits initially allocated to the channel with mutexlag set to 1 in the object channel can be redistributed to other object channels, and the bits initially allocated to the channel with mutexlag set to 1 in the bedmute signal can be redistributed to other bedmute channels.

[0242] Furthermore, in some embodiments of this application, the multi-channel side information includes: channel bit allocation ratio,

[0243] The channel bit allocation ratio is used to indicate the bit allocation ratio between non-low frequency effect (LFE) channels in a multi-channel signal.

[0244] The Low Frequency Effect (LFE) channel is an audio channel with a bass sound range of 3-120Hz. This channel can be used to send signals to speakers specifically designed for bass tones. The channel bit allocation ratio is used to indicate the bit allocation ratio of non-LFE channels. For example, the channel bit allocation ratio occupies 6 bits. The number of bits occupied by the channel bit allocation ratio is not limited in the embodiments of this application.

[0245] For example, the channel bit allocation ratio can be the channel bit allocation ratio field in the multi-channel side information, represented as chBitRatios, occupying 6 bits, used to indicate the bit allocation ratio of all channels in the multi-channel signal except for the LFE channel. The channel bit allocation ratio field indicates the bit allocation ratio of each transmission channel, thereby determining the number of bits received by each transmission channel. Notably, this number of bits can also be further converted to the number of bytes.

[0246] In some embodiments of this application, the multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index;

[0247] Among them, the codebook index for the quantization of the interaural level difference (ILD) parameter is used to indicate the codebook index for the quantization of the interaural level difference (ILD) parameter of each channel in each channel.

[0248] Channel pair count, used to indicate the number of channel pairs in the current frame of a multichannel signal;

[0249] Channel pair index, used to represent the index of a channel pair.

[0250] In this embodiment, the number of bits occupied by the inter-channel amplitude difference parameter quantization codebook index is not limited. For example, the inter-channel amplitude difference parameter quantization codebook index occupies 5 bits. The inter-channel amplitude difference parameter quantization codebook index can be represented as mcIld[ch1] and mcIld[ch2], occupying 5 bits. It is the codebook index of the inter-channel amplitude difference ILD parameter quantization for each channel in the current channel pair, used to recover the amplitude of the decoded spectrum.

[0251] In this embodiment, the number of bits occupied by the channel pair number is not limited. For example, the channel pair number occupies 4 bits, represented as pairCnt, which occupies 4 bits and is used to represent the number of channel pairs in the current frame.

[0252] In this embodiment, the number of bits occupied by the channel pair index is not limited. For example, the channel pair index is represented as channelPairIndex. The number of bits in channelPairIndex is related to the total number of channels. It is used to represent the index of the channel pair and can be parsed to obtain the index values ​​of the two channels in the current channel pair, namely ch1 and ch2.

[0253] In some embodiments of this application, in addition to performing the aforementioned steps, the encoding method for multi-channel signals performed by the encoding device further includes:

[0254] Send the bitstream to the decoding device.

[0255] In this embodiment of the application, after the encoding end obtains the transmission channel signal and silence mark information of each transmission channel, it can generate a bit stream, which carries the silence mark information. The encoding end can send the bit stream to the decoding end.

[0256] As illustrated by the foregoing embodiments, the process involves detecting silence markers in a multi-channel signal to obtain silence marker information, which includes a silence enable flag and / or a silence flag. Multi-channel encoding is then performed on the multi-channel signal to obtain transmission channel signals for each transmission channel. A bitstream is generated based on the transmission channel signals and the silence marker information. The bitstream includes the multi-channel quantization encoding results of the silence marker information and the transmission channel signals of each transmission channel. Using the silence marker information for subsequent encoding processing can improve encoding efficiency.

[0257] This application also provides a method for decoding multi-channel signals. This method can be executed by a terminal device, such as a multi-channel signal decoding device (hereinafter referred to as a decoding end or decoder, for example, the decoding end can be an AI decoder). Figure 5 As shown, the method executed at the decoding end in this embodiment mainly includes:

[0258] 501. Parse the mute flag information from the bitstream of the encoding device, and determine the encoding information of each transmission channel based on the mute flag information. The mute flag information includes: mute enable flag, and / or mute flag.

[0259] The decoding end employs a processing method reversed by the encoding end. It first receives the bitstream from the encoding device. Since this bitstream carries silence marker information, the encoding information for each transmission channel is determined based on this information. The silence marker information includes a silence enable flag and / or a silence flag. For a detailed explanation of the silence enable flag and the silence flag, please refer to the aforementioned embodiment of the encoding end; it will not be repeated here.

[0260] In some embodiments of this application, step 501, parsing silence marker information from the bitstream of the encoding device, includes:

[0261] H1, parse the mute flags of each channel from the bitstream; or,

[0262] H2. Parse the mute enable flag from the bitstream. If the mute enable flag is the first value, parse the mute flag from the bitstream; or...

[0263] H3. Parse the bed mute enable flag and / or object mute enable flag from the bitstream, as well as the mute flag for each channel; or,

[0264] H4. Extract the bed mute enable flag and / or object mute enable flag from the bitstream; based on the bed mute enable flag and / or object mute enable flag, extract the mute flags of some channels of each channel from the bitstream.

[0265] The decoding end parses the mute flag information from the bitstream of the encoding device. Depending on the specific content of the mute flag information generated by the encoding device, the mute flag information obtained by the decoding end corresponds to that of the encoding side. Specifically, in one method, the mute flag indicates whether each channel is a mute channel. A mute channel is a channel that does not need encoding or a channel that needs to be encoded using low-bit encoding. The decoding end can parse the mute flag of each channel from the bitstream. In another method, the mute enable flag can also be used to indicate whether each channel is a non-mute channel. For example, when the mute enable flag is a first value (e.g., 1), it indicates that further detection of the mute flag of each channel is needed. When the mute enable flag is a second value (e.g., 0), it indicates that each channel is a non-mute channel. The decoding end parses the mute enable flag from the bitstream; if the mute enable flag is a first value, the mute flag is parsed from the bitstream. In one approach, the mute enable flags include: a bed mute enable flag and / or an object mute enable flag. The decoder parses the bed mute enable flag and / or the object mute enable flag, as well as the mute flags for each channel, from the bitstream. In another approach, the decoder parses the bed mute enable flag and / or the object mute enable flag from the bitstream; based on the bed mute enable flag and / or the object mute enable flag, it parses the mute flags for some channels from the bitstream. There is no limitation on which specific channel's mute flag is obtained.

[0266] 502. Decode the encoded information of each transmission channel to obtain the decoded signal of each transmission channel.

[0267] In this process, after obtaining the encoding information of each transmission channel from the bitstream, the decoding end can decode the encoding information of each transmission channel. The decoding and dequantization process is the reverse of the quantization and encoding process of the encoding end, so as to obtain the decoded signal of each transmission channel.

[0268] In some embodiments of this application, step 502 decodes the encoded information of each transmission channel, including:

[0269] I1. Extract multi-channel side information from the bitstream;

[0270] I2. Allocate bits to each transmission channel based on the multi-channel side information and the mute flag information to obtain the number of encoded bits for each channel;

[0271] I3. Decode the encoded information of each transmission channel according to the number of encoded bits of each channel.

[0272] The bitstream may also include multi-channel side information. The decoder can allocate bits to each transmission channel based on the multi-channel side information and the mute flag information to obtain the number of encoded bits for each channel. The number of encoded bits obtained by the decoder is the same as the number of encoded bits preset by the encoder. Then, the decoder decodes the encoded information of each transmission channel based on the number of encoded bits for each transmission channel, thereby realizing the decoding of the transmission channel signals of each transmission channel.

[0273] Furthermore, in some embodiments of this application, the multi-channel side information includes: a channel bit allocation ratio field.

[0274] The channel bit allocation ratio field is used to indicate the bit allocation ratio of the low frequency effects (LFE) channels in each channel.

[0275] The Low Frequency Effect (LFE) channel is an audio channel with a bass sound range of 3-120Hz, which can be used to send signals to speakers specifically designed for bass tones. For example, the channel bit allocation ratio field occupies 6 bits. The number of bits occupied by the channel bit allocation ratio field is not limited in the embodiments of this application.

[0276] For example, the channel bit allocation ratio field, denoted as chBitRatios, occupies 6 bits and is used to indicate the bit allocation ratio of non-LFE channels in each channel. The channel bit allocation ratio field indicates the bit allocation ratio of each channel, thereby determining the number of bits allocated to each channel. Notably, this number of bits can be further converted to a number of bytes.

[0277] In some embodiments of this application, the multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index;

[0278] Among them, the codebook index for the quantization of the inter-channel amplitude difference parameter (ILD) is used to indicate the codebook index for the quantization of the inter-channel amplitude difference (ILD) parameter of each channel.

[0279] Channel pair count, used to indicate the number of channel pairs in the current frame of a multichannel signal;

[0280] Channel pair index, used to represent the index of a channel pair.

[0281] In this embodiment, the number of bits occupied by the inter-channel amplitude difference parameter quantization codebook index is not limited. For example, the inter-channel amplitude difference parameter quantization codebook index occupies 5 bits. The inter-channel amplitude difference parameter quantization codebook index can be represented as mcIld[ch1] and mcIld[ch2], occupying 5 bits. It is the codebook index of the inter-channel amplitude difference ILD parameter quantization for each channel in the current channel pair, used to recover the amplitude of the decoded spectrum.

[0282] In this embodiment, the number of bits occupied by the channel pair number is not limited. For example, the channel pair number occupies 4 bits, represented as pairCnt, which occupies 4 bits and is used to represent the number of channel pairs in the current frame.

[0283] In this embodiment, the number of bits occupied by the channel pair index is not limited. For example, the channel pair index is represented as channelPairIndex. The number of bits in channelPairIndex is related to the total number of channels. It is used to represent the index of the channel pair and can be parsed to obtain the index values ​​of the two channels in the current channel pair, namely ch1 and ch2.

[0284] In some embodiments of this application, step I2 allocates bits to each transmission channel based on multi-channel side information and mute flag information, including:

[0285] I21. Determine the first remaining number of bits based on the number of available bits and the number of safe bits;

[0286] There is no restriction on the value of the number of safe bits. For example, the number of safe bytes is represented as safeBits, which is 8 bits. Subtracting the number of safe bits from the number of available bits will give you the first number of remaining bits.

[0287] I22. Allocate the first remaining number of bits to each channel according to the channel bit allocation ratio field in the multi-channel side information. The channel bit allocation ratio field is used to indicate the bit allocation ratio of each channel.

[0288] I23. When there are still a second number of remaining bits after the first number of remaining bits have been allocated to each channel, the second number of remaining bits shall be allocated to each channel according to the channel bit allocation ratio field.

[0289] The second remaining number of bits can be obtained by subtracting the number of bits allocated to each channel from the first remaining number of bits.

[0290] I24. When there are still a third remaining bit number after the second remaining bit number has been allocated to each channel, the third remaining bit number shall be allocated to the channel that has the most bits allocated when the first remaining bit number was used for bit allocation.

[0291] The third remaining number of bits can be obtained by subtracting the number of bits allocated to each channel from the second remaining number of bits.

[0292] I25. When the number of bits allocated to the first channel in each channel exceeds the upper limit of the number of bits in a single channel, the excess bits are allocated to the other channels in each channel except the first channel.

[0293] There is no limit to the upper limit of the number of bits in a single channel. The first channel can be any one of the channels.

[0294] 503. Perform multi-channel decoding processing on the decoding signals of each transmission channel to obtain multi-channel decoded output signals.

[0295] In this process, the decoding end obtains the decoded signals of each transmission channel through decoding, and then further decodes the decoded signals of each transmission channel to obtain the decoded output signal.

[0296] In some embodiments of this application, after step 503 performs multi-channel decoding processing on the decoded signals of each transmission channel to obtain multi-channel decoded output signals, the multi-channel signal decoding method executed by the decoding end further includes:

[0297] J1. Post-processing of the multi-channel decoded output signal, including at least one of the following: bandwidth extension decoding, inverse time-domain noise shaping, inverse frequency-domain noise shaping, and inverse time-frequency transformation.

[0298] The post-processing of the output signal described above is the reverse of the pre-processing process at the encoding end, and the specific processing method is no longer limited.

[0299] As can be seen from the foregoing examples, in the embodiments of this application, the decoding end can obtain the silence marker information from the bitstream of the encoding end, thereby facilitating the decoding end to perform decoding processing in a manner consistent with that of the encoding end, such as bit allocation.

[0300] To facilitate a better understanding and implementation of the above-described solutions in the embodiments of this application, specific examples of corresponding application scenarios are provided below.

[0301] Multi-channel audio encoders; products include mobile terminals, chips, and wireless networks.

[0302] Example 1 Encoding end as follows Figure 6 As shown, it includes a silence marker detection unit, a multi-channel encoding processing unit, a multi-channel quantization encoding unit, and a stream multiplexing interface.

[0303] The silence marker detection unit is mainly used to detect silence marker information based on the input signal and determine the silence marker information. The silence marker information may include a silence enable flag and / or a silence flag.

[0304] The mute enable flag is denoted as HasSilFlag. It can be a global mute enable flag or a partial mute enable flag. For example, an object mute enable flag that only applies to the target signal in a multi-channel signal is denoted as objMuteEna. Similarly, a bed mute enable flag that only applies to the target signal in a multi-channel signal is denoted as bedMuteEna.

[0305] The global mute enable flag applies to the mute enable flag of the multi-channel signal. When the multi-channel signal contains only the sound bed signal, the global mute enable flag applies to the sound bed signal. When the multi-channel signal contains only the object signal, the global mute enable flag applies to the mute enable flag of the object signal. When the multi-channel signal contains both the sound bed signal and the object signal, the global mute enable flag applies to both the sound bed signal and the object signal.

[0306] The partial mute enable flag applies to a portion of the channels in the multi-channel signal. These channels are pre-defined; for example, the partial mute enable flag may apply to an object mute enable flag of the target signal, or it may apply to a bed mute enable flag of the bed signal, or it may apply to other channels in the multi-channel signal that do not contain the LFE channel signal. The partial mute enable flag may also apply to the mute enable flags of the channels in the multi-channel signal that participate in pairing. The specific method of pairing the multi-channel signals is not limited in this embodiment.

[0307] The mute enable flag indicates whether mute detection is enabled. For example, when the mute enable flag is at its first value (e.g., 1), it means that the mute detection function is enabled, further detecting the mute status of each channel. When the mute enable flag is at its second value (e.g., 0), it means that the mute detection function is disabled.

[0308] The mute enable flag can also be used to indicate whether further transmission of the mute flag for each channel is required. For example, when the mute enable flag is at its first value (e.g., 1), it indicates that further transmission of the mute flag for each channel is required. When the mute enable flag is at its second value (e.g., 0), it indicates that further transmission of the mute flag for each channel is not required.

[0309] The mute enable flag can also be used to indicate whether each channel is a non-mute channel. For example, when the mute enable flag is at its first value (e.g., 1), it indicates that further testing of the mute flag of each channel is required. When the mute enable flag is at its second value (e.g., 0), it indicates that each channel is a non-mute channel.

[0310] The global mute enable flag applies to all channels, while the partial mute enable flag applies to only certain channels. For example, the object mute enable flag applies to the channel corresponding to the object signal in a multi-channel signal, and the bed mute enable flag applies to the channel corresponding to the bed signal in a multi-channel signal.

[0311] The mute enable flag can be controlled by external input. It can be preset according to encoder parameters such as encoding rate and encoding bandwidth, or determined according to the mute detection results of each channel.

[0312] The mute flag for each channel indicates whether it is a mute channel. The mute flag for each channel is denoted as silFlag[i], where ch is the channel number, ch = 0…N-1, and N is the total number of channels in the input signal to be encoded. The number of channels for the bed signal is M, and the number of channels for the target channel is P. The total number of channels N = M + P. For example, the signal to be encoded is a mixed signal containing bed and target signals, where: the bed signal is a 5.1.4 channel signal with M = 10 channels; the number of target signals is 4 with P = 4 channels; the total number of channels is 14. The channel numbers for the bed signals are from 0 to 9, and the channel numbers for the target signals are from 10 to 13. The mute flag silFlag[i], ch = 0…13, corresponds to the mute flag for each channel and is used to indicate whether each channel is a mute channel. A silent channel is a channel whose signal energy / decibels / loudness is below the hearing threshold. It is a channel that does not require encoding or only requires encoding with lower bit values. A mute flag value of 1 indicates a silent channel; a mute flag value of 0 indicates a non-silent channel. When the mute flag value is 1, the channel is either not encoded or encoded with lower bit values.

[0313] The input signal for silence sign detection can be the original input signal or a pre-processed signal. Pre-processing may include, but is not limited to, transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding. The input signal can be a time-domain signal or a frequency-domain signal. Taking the time-domain signals of each channel in a multi-channel signal as an example, one method for detecting the silence sign of each channel could be:

[0314] The energy of each channel signal in the current frame is determined based on the input signals of each channel in the current frame.

[0315] Assuming a frame length of FRAME_LEN, the energy (ch) of the ch-th channel in the current frame is:

[0316]

[0317] Among them, orig ch is the input signal of the ch-th channel of the current frame, and energy(ch) is the energy of the ch-th channel of the current frame.

[0318] Based on the energy of the signals in each channel of the current frame, determine the silence detection parameters for each channel of the current frame.

[0319] The silence detection parameters for each channel in the current frame are used to characterize the energy value, power value, decibel value, or loudness value of the signal in each channel of the current frame.

[0320] For example, the silence detection parameters for each channel in the current frame can be values ​​in the logarithmic domain of the energy of each channel's signal in the current frame, such as log2(energy(ch)) or log10(energy(ch)). Based on the energy of each channel's signal in the current frame, the silence detection parameters for each channel satisfy the following conditions:

[0321] energyDB[ch]=10*log10(energy[ch] / Bit_Depth / Bit_Depth);

[0322] Wherein, energyDB[ch] is the silence detection parameter of the ch-th channel of the current frame, energy(ch) is the energy of the ch-th channel of the current frame, and Bit_Depth is the full-width bias value. For example, if the sampling bit depth is 16 bits, then the full-width bias value is 2^16 = 65536.

[0323] Based on the silence detection parameters and silence detection thresholds of each channel in the current frame, determine the silence flag for each channel in the current frame.

[0324] The silence detection parameters of each channel in the current frame are compared with the silence detection threshold: If the silence detection parameter of the ch-th channel in the current frame is less than the silence detection threshold, then the ch-th channel in the current frame is a silent frame, that is, the ch-th channel at the current moment is a silent channel, and the silence flag silFlag[i] of the ch-th channel in the current frame is the first value (e.g., 1). If the silence detection parameter of the ch-th channel in the current frame is greater than or equal to the silence detection threshold, then the ch-th channel in the current frame is a non-silent frame, that is, the ch-th channel at the current moment is a non-silent channel, and the silence flag silFlag[i] of the ch-th channel in the current frame is the second value (e.g., 0).

[0325] The pseudocode for determining the mute flag of the current frame's ch-th channel based on the mute detection parameters and mute detection threshold is as follows:

[0326] silFlag[i] = 0;

[0327] if(energyDB[ch] <g_MuteThrehold)

[0328] {silFlag[i] = 1;}

[0329] Mute flag information may include a mute enable flag and / or a mute flag. Examples of different mute flag information are shown below:

[0330] Method 1: The mute flag information is the mute flag silFlag[i] for each channel. Determine the mute flag silFlag[i] for each channel, write the mute flag silFlag[i] for each channel into the bitstream, and transmit it to the decoding end.

[0331] Method 2: The mute flag information includes the mute enable flag HasSilFlag and the mute flag silFlag[i].

[0332] The mute enable flag (HasSilFlag) indicates whether the mute detection function is enabled in the current frame, and can also be used to indicate whether the mute detection results of each channel are transmitted in the current frame.

[0333] Determine the mute enable flag HasSilFlag, write it into the bitstream, and transmit it to the decoder; based on the value of the mute enable flag, determine whether to write the mute flag silFlag[i] into the bitstream.

[0334] When the mute enable flag HasSilFlag is 0, the mute flag silFlag[i] is not written into the bitstream and transmitted to the decoder.

[0335] When the mute enable flag HasSilFlag is 1, the mute flag silFlag[i] is written into the bitstream and transmitted to the decoder.

[0336] Method 3: The mute flag information includes the bed mute enable flag bedMuteEna, the object mute enable flag objMuteEna, and the mute flag silFlag[i] for each channel.

[0337] The `bedMuteEna` flag indicates whether the silence detection function for the corresponding channel of the bed sound signal is enabled in the current frame. Similarly, the `objMuteEna` flag indicates whether the silence detection function for the corresponding channel of the object signal is enabled in the current frame. For example:

[0338] When the bed silence enable flag `bedMuteEna` is 0 and the object silence enable flag `objMuteEna` is 1, the silence flag values ​​for the corresponding channels of the bed signal are all set to 0, i.e., non-silent channels. The silence flag value for the corresponding channel of the object signal is the silence detection result.

[0339] When the bed silence enable flag `bedMuteEna` is 1 and the object silence enable flag `objMuteEna` is 0, the silence flag values ​​for the corresponding channels of the object signal are all set to 0, i.e., non-silent channels. The silence flag value for the corresponding channel of the bed signal is the silence detection result.

[0340] When the bed mute enable flag bedMuteEna is 0, the object mute enable flag objMuteEna is 0, and the mute flag values ​​of each channel are all set to 0, that is, it is a non-mute channel.

[0341] When the bed mute enable flag bedMuteEna is 1 and the object mute enable flag objMuteEna is 1, the mute flag for each channel is the mute detection result.

[0342] When the mute flag information includes the bed mute enable flag bedMuteEna, the object mute enable flag objMuteEna, and the mute flag, the mute flags for each channel can be transmitted.

[0343] Method 4: The mute flag information includes the bed mute enable flag bedMuteEna, the object mute enable flag objMuteEna, and the mute flag silFlag[i] for some channels.

[0344] The difference between Method 4 and Method 3 is that only the mute flags of some channels are transmitted. For example, when the bed mute enable flag `bedMuteEna` is 0 and the object mute enable flag `objMuteEna` is 1, only the mute flag of the channel corresponding to the object signal can be transmitted, and the mute flag of the channel corresponding to the bed signal is not transmitted; when the bed mute enable flag `bedMuteEna` is 1 and the object mute enable flag `objMuteEna` is 0, only the mute flag of the channel corresponding to the bed signal can be transmitted; when both the bed mute enable flag `bedMuteEna` and the object mute enable flag `objMuteEna` are 0, it is not necessary to transmit the mute flags of each channel; when both the bed mute enable flag `bedMuteEna` and the object mute enable flag `objMuteEna` are 1, then the mute flags of each channel are transmitted.

[0345] Method 5: The bed mute enable flag `bedMuteEna` and the object mute enable flag `objMuteEna` can be replaced with `HasSilFlag = {HasSilFlag(0), HasSilFlag(1)}`, where `HasSilFlag(0)` and `HasSilFlag(1)` correspond to `bedMuteEna` and `objMuteEna`, respectively. Alternatively, a single 2-bit mute enable flag `HasSilFlag` can represent both the bed mute enable flag `bedMuteEna` and the object mute enable flag `objMuteEna`. This application does not limit the specific implementation of the method.

[0346] Method 6: First determine the mute flag for each channel, and then determine the mute enable flag based on the mute flag for each channel.

[0347] For example, the mute enable flag can be a global mute enable flag. If the mute flags of all channels are 0, the global mute enable flag is set to 0. Only the global mute enable flag needs to be written to the bitstream and transmitted to the decoding side; the mute flags of each channel do not need to be transmitted. If at least one of the mute flags of each channel is 1, the global mute enable flag is set to 1. Only the global mute enable flag needs to be written to the bitstream and transmitted to the decoding side; the mute flags of each channel do not need to be transmitted.

[0348] For example, the mute enable flag can be either the bed mute enable flag `bedMuteEna` or the object mute enable flag `objMuteEna`. Taking the bed mute enable flag `bedMuteEna` as an example, if the mute flags of all channels corresponding to the bed signal are 0, then the bed mute enable flag is set to 0. Only the bed mute enable flag needs to be written into the bitstream and transmitted to the decoding side; there is no need to transmit the mute flags of each channel corresponding to the bed signal. If at least one of the mute flags of each channel corresponding to the bed signal is 1, then the bed mute enable flag is set to 1. Only the bed mute enable flag needs to be written into the bitstream and transmitted to the decoding side; there is no need to transmit the mute flags of each channel corresponding to the bed signal. The object mute enable flag `objMuteEna` can be handled similarly, and will not be elaborated here.

[0349] The embodiments in this application only illustrate some implementation methods. There may be other possible implementation methods, which are not limited.

[0350] The multi-channel encoding processing unit completes the screening, pairing, downmixing, and multi-channel side information generation of multi-channel signals, and obtains the signals of each transmission channel after multi-channel pairing and downmixing.

[0351] Optionally, preprocessing may be included between the silence marker detection processing and the multi-channel encoding processing to preprocess the input signal to obtain a preprocessed signal, which is then used as the input for the multi-channel encoding processing. Preprocessing may include, but is not limited to, transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding, etc., as described in this embodiment. Figure 7 As shown, based on the multi-channel input signal or the pre-processed multi-channel signal, the multi-channel signal is filtered to obtain the filtered multi-channel signal. The filtered multi-channel signal is then paired to obtain the multi-channel paired signal. Finally, the multi-channel paired signal is down-mixed (e.g., center-side information (MIDSIDE, MS) processing) to obtain the down-mixed multi-channel paired signal to be encoded.

[0352] Optionally, the silence marker information can be modified during preprocessing. For example, after frequency domain noise shaping, the energy of the signal in a certain transmission channel changes, and the silence detection result of that channel can be adjusted.

[0353] Multichannel side information includes, but is not limited to: number of pairs, list of pair channel indexes, list of pair channel interaural intensity difference (ILD) coefficients, and list of pair channel ILD big-small end.

[0354] Optionally, the initial multi-channel processing method can be adjusted based on the mute flag information. For example, during the screening process of multi-channel signals, channels with a mute flag of 1 do not participate in the pair screening.

[0355] The multi-channel quantization encoding unit performs quantization encoding on the signals of each transmission channel after the multi-channel group has been downmixed.

[0356] Multichannel quantization coding includes bit allocation processing and encoding.

[0357] Optionally, bit allocation is performed based on the silence marker information, the number of available bits, and the multi-channel side information; encoding is then performed based on the bit allocation results for each channel to obtain the encoded bitstream.

[0358] The specific implementation of multichannel quantization coding can be as follows: the group-pair downmixed signal is transformed by a neural network to obtain latent features; the latent features are then quantized and interval encoded. Alternatively, multichannel quantization coding can be based on vector quantization to quantize and encode the group-pair downmixed signal. This application does not limit this implementation.

[0359] Optionally, bit allocation can be performed based on silence flag information. For example, different bit allocation strategies can be selected based on the silence enable flag.

[0360] Assuming the mute enable flags include the bed mute enable flag `bedMuteEna` and the object mute enable flag `objMuteEna`, bit allocation based on the mute flag information can be performed first, based on the total available bits and the signal characteristics of each channel. Then, the bit allocation results are adjusted according to the mute flag information. For example, if the object mute enable flag `objMuteEna` is 1, the bits initially allocated to the channel with mute flag 1 in the object signal are allocated to the bed signal or other object channels. If both the bed mute enable flag `bedMuteEna` and the object mute enable flag are 1, the bits initially allocated to the channel with mute flag 1 in the object channel can be reallocated to other object channels, and the bits initially allocated to the channel with mute flag 1 in the bed signal can be reallocated to other bed channels.

[0361] The bitstream multiplexing interface multiplexes the encoded audio channels to form a serial bitstream for easy transmission in the channel or storage in digital media.

[0362] In this embodiment, the decoding end is as follows: Figure 8As shown, it includes a stream demultiplexing unit, a channel decoding and dequantization unit, a multi-channel decoding processing unit, and a multi-channel post-processing unit.

[0363] The stream demultiplexing unit parses the silence flag information from the received stream and determines the encoding information for each channel.

[0364] The mute flag information is parsed from the received bitstream. The parsing process is the reverse process of the encoder writing the mute flag information into the bitstream.

[0365] For example, if the encoding end uses method one, then the decoding end: parses the mute flag silFlag[i] of each channel from the bitstream, ch=0…N-1, where N is the number of channels of the multi-channel signal to be decoded.

[0366] Alternatively, if the encoding end adopts method two, then the decoding end: firstly, parse the mute enable flag HasSilFlag from the bitstream; if the mute enable flag HasSilFlag is the first value (e.g., 1), parse the mute flag silFlag[i] from the bitstream, ch=0…N-1, where N is the number of channels of the multi-channel signal to be decoded.

[0367] Alternatively, if the encoding end adopts method three, then the decoding end: firstly, parse the bed mute enable flag bedMuteEna and the object mute enable flag objMuteEna and the mute flag silFlag[i] of each channel from the bitstream, ch=0…N-1, where N is the number of channels of the multi-channel signal to be decoded.

[0368] Alternatively, if the encoding end adopts method four, then the decoding end: firstly, parse the bed mute enable flag bedMuteEna and the object mute enable flag objMuteEna from the bitstream; then, based on the parsed bed mute enable flag bedMuteEna and object mute enable flag objMuteEna, parse the mute flag of the corresponding channel from the bitstream. For example: when the bed mute enable flag `bedMuteEna` is 0 and the object mute enable flag `objMuteEna` is 1, the mute flag of the corresponding channel of the object signal is parsed from the bitstream; when the bed mute enable flag `bedMuteEna` is 1 and the object mute enable flag `objMuteEna` is 0, the mute flag of the corresponding channel of the bed signal is parsed from the bitstream; when the bed mute enable flag `bedMuteEna` is 0 and the object mute enable flag `objMuteEna` is 0, there is no need to parsed the mute flag from the bitstream; when the bed mute enable flag `bedMuteEna` is 1 and the object mute enable flag `objMuteEna` is 1, the mute flag of each channel is parsed from the bitstream, and the number of parsed channels is the sum of the number of channels corresponding to the bed signal and the number of channels corresponding to the object signal.

[0369] Taking the following example, the specific syntax for the decoding end to parse the silence marker information from the bitstream is as follows:

[0370] Parse multi-channel side information from the received bitstream.

[0371] Bit allocation is performed based on multi-channel side information to determine the number of encoded bits for each channel. Optionally, if the encoding end allocates bits based on silence flag information, the decoding end also needs to allocate bits based on silence flag information to determine the number of encoded bits for each channel.

[0372] Based on the number of encoded bits for each channel, the encoding information for each channel is determined from the received bitstream.

[0373] The decoding unit performs inverse encoding and inverse quantization on each encoded channel to obtain the decoded signal of the multi-channel group downmixing.

[0374] Inverse encoding and inverse quantization are the reverse processes of multichannel quantization encoding at the encoding end.

[0375] The multi-channel decoding processing unit performs multi-channel decoding processing on the downmixed decoding signal to obtain a multi-channel output signal.

[0376] Multichannel decoding is the reverse process of multichannel encoding. It reconstructs the multichannel output signal based on the decoded signal of the downmixed multichannel group using multichannel side information.

[0377] like Figure 9 As shown, if preprocessing is included before multi-channel encoding at the encoding end, then post-processing is included after multi-channel decoding at the decoding end, such as: bandwidth extension decoding, inverse time-domain noise shaping, inverse frequency-domain noise shaping, inverse time-frequency transformation, etc., to obtain the final output signal.

[0378] As illustrated by the examples above, detecting and determining the silence marker information of a multi-channel input signal, and then performing subsequent encoding processing such as bit allocation based on the silence marker information, can improve encoding efficiency.

[0379] This application proposes a method for generating a silence flag bitstream based on input signal characteristics. The encoding end detects and determines the silence flag information from the multi-channel input signal; transmits the silence flag information to the decoding end; allocates bits according to the silence flag information, and encodes the multi-channel signal. The decoding end parses the silence flag information from the bitstream; allocates bits according to the silence flag information, and decodes the multi-channel signal.

[0380] In the technical solution included in this application embodiment, a mute flag bit is calculated for each input signal to guide the bit allocation for encoding and decoding. The input signal is checked to determine if it is a mute frame. If it is, the channel is not encoded or only a small number of bits are allocated for encoding. The decibel or loudness value of the signal is calculated at the input end and compared with a set hearing threshold. If it is below the hearing threshold, the mute flag is set to 1; otherwise, the mute flag is set to 0. When the mute flag is 1, the channel is not encoded or encoded using lower bits. The data before quantization for channels with a mute bit of 1 can be cleared to 0. The mute flag is transmitted as side information to the decoding end to guide the bit demultiplexing. The transmission syntax at the encoding end is as follows: HasSilFlag is used to indicate that the mute flag is enabled, and 1 bit can be used to transmit HasSilFlag. When HasSilFlag = 1, the mute flags for each channel are further transmitted; when HasSilFlag = 0, the mute flags for each channel are not transmitted. For example, in a 5.1.4 channel configuration, a 10-bit mute flag is transmitted in the side information of the multi-channel configuration, with 1 bit per channel, and the order is consistent with the order of the input channels. Other modules at the encoding end can modify the mute flag, changing it from 1 to 0 and transmitting it in the bitstream.

[0381] The embodiments of this application have the following advantages: the mute marker information is detected for the multi-channel input signal, the mute marker information is determined, and subsequent encoding processing is performed based on the mute marker information, such as bit allocation. For the mute channel, no encoding or encoding with a lower bit value can be performed, saving the number of encoding bits and improving encoding efficiency.

[0382] The silence marker information is transmitted to the decoding end so that the decoding end can perform decoding processing in the same way as the encoding end, such as bit allocation.

[0383] In other embodiments of this application, the hybrid coding improvement scheme is described as follows:

[0384] A hybrid-mode codec supports encoding and decoding of both bed signals and object signals. The specific implementation consists of three parts:

[0385] Hybrid encoding bit pre-allocation: Based on the multi-channel side information bedBitsRatio, the pre-allocated bit count bedAvailableBytes of the sound bed signal and the pre-allocated bit count objAvailableBytes of the object signal are obtained.

[0386] Hybrid encoding bit allocation consists of four steps, in the following order: silent frame bit allocation, non-silent frame bit allocation adaptation, non-silent frame bit allocation, and non-silent frame bit allocation adaptation restoration.

[0387] Silence Frame Bit Allocation: If a silence frame exists, allocate bits to the silence frame channel according to the silence flag silFlag[i] of the side information and the mixed allocation strategy mixAllocStrategy, and update the pre-allocated bit count bedAvailableBytes of the sound bed signal and the pre-allocated total bit count objAvailableBytes of the object signal.

[0388] Non-silent frame bit allocation adaptation: This involves mapping the channel parameters in sequence to facilitate the processing of non-silent frame bit allocation.

[0389] Non-silent frame bit allocation: Bits are allocated based on the updated pre-allocated bit count bedAvailableBytes of the bed signal, the updated pre-allocated bit count objAvailableBytes of the object signal, and the channel bit allocation scaling factor chBitRatios.

[0390] Non-silent frame bit allocation adaptation and restoration: reverse mapping of channel parameters in order, which facilitates subsequent interval decoding, inverse quantization and neural network inverse transformation steps.

[0391] Mixing and upmixing: Based on the channel pair index channelPairIndex indicating the two paired channels ch1 and ch2, perform M / S upmixing to obtain the upmixed channel signal.

[0392] The syntax for multi-channel stereo side information is shown in Table 1 below, which is the DecodeMcSideBits() syntax.

[0393]

[0394]

[0395] The semantic explanation is as follows: bedBitsRatio occupies 4 bits and represents the ratio factor index of the sound bed signal to the total number of bits, with a value of 0-15. The corresponding floating-point ratios are as follows: 1:0.0625 2:0.125 3:0.1875 4:0.25 5:0.3125 6:0.375 7: 0.4375 8:0.5 9:0.5625 10:0.625 11: 0.6875 12:0.75 13:0.8125 14:0.875

[0410] 15: 0.9375.

[0411] `mixAllocStrategy` occupies 2 bits and represents the allocation strategy for the mixed signal of the sound bed signal and the object signal. This allocation strategy can be predetermined, or it can be predefined according to coding parameters, including: coding rate and signal characteristic parameters. The coding parameters are predetermined. The value range and meaning of the allocation strategy are as follows:

[0412] 0: Excessive bed bits generated by the Mute mechanism (mute flag) are assigned to the bed signal, excess object bits are assigned to the object signal, and silent bed bits are assigned to non-silent bed bits.

[0413] 1: The excess bed bits generated by the Mute mechanism are allocated to the bed signal, and the excess object bits are allocated to the bed signal.

[0414] 2: The extra bed bits generated by the Mute mechanism are given to the object signal, and the extra object bits are given to the object signal.

[0415] 3: Retain.

[0416] HasSilFlag occupies 1 bit, where 0 indicates that silence frame processing is disabled or there are no silence frames; and 1 indicates that silence frame processing is enabled and there are silence frames.

[0417] silFlag[i] occupies 1 bit and represents the mute frame flag for the corresponding channel. 0 indicates a non-mute frame and 1 indicates a mute frame.

[0418] soundBedType occupies 1 bit, type of sound bed, 0f is only the object signal or none (onlyobjs), 1 is the sound bed signal or HOA signal or mc or hoa.

[0419] The codingProfile occupies 3 bits: 0 for mono, or stereo signal or bed signal for mono / stereo / mc; 1 for a mixed signal of bed and object for channel+obj mix; and 2 for hoa.

[0420] pairCnt occupies 4 bits and is used to represent the number of channel pairs in the current frame.

[0421] The number of bits in `channelPairIndex` is related to the total number of channels, as shown in Note 1 of the table above. It represents the index of a channel pair and can be parsed to obtain the index values ​​of the two channels in the current channel pair, namely `ch1` and `ch2`.

[0422] mcIld[ch1] and mcIld[ch2] occupy 4 bits and are the amplitude difference parameters between each channel in the current channel pair, used to recover the amplitude of the decoded spectrum.

[0423] scaleFlag[ch1] and scaleFlag[ch2] occupy 1 bit and represent the scaling flag parameter for each channel in the current channel pair, indicating whether the amplitude of the current channel is reduced or increased.

[0424] chBitRatios occupies 4 bits and represents the bit allocation ratio for each channel.

[0425] The decoding process is as follows: first, the mixed encoding bits are pre-allocated.

[0426] The function of the hybrid encoding bit pre-allocation module is to calculate the number of pre-allocated bytes for the acoustic bed and the number of pre-allocated bytes for the object based on the ratio factor index parameter of the acoustic bed signal obtained by decoding in the bit stream to the total number of bits, and provide it to subsequent modules.

[0427] The number of available bytes remaining after deducting other side information in the current frame is denoted as availableBytes, where the number of pre-allocated bytes for the acoustic bed is bedAvailableBytes, and the number of pre-allocated bytes for the object is objAvailableBytes. The scaling factor index parameter for the acoustic bed signal as a percentage of the total number of bits is bedBitsRatio, and the floating-point scaling factor corresponding to bedBitsRatio is bedBitsRatioFloat. The correspondence between bedBitsRatio and bedBitsRatioFloat is explained in the bedBitsRatio section of the aforementioned semantics.

[0428] The formulas for calculating the pre-allocated bytes bedAvailableBytes and the pre-allocated bytes objAvailableBytes of the object, based on the available bytes and the floating-point scaling factor bedBitsRatioFloat (which represents the proportion of the total bits to the bed signal), are as follows:

[0429] bedAvailbleBytes=floor(availableBytes*bedBitsRatioFloat);

[0430] objAvailbleBytes=availableBytes–bedAvailbleBytes.

[0431] The hybrid coding bit allocation process is as follows: Hybrid coding bit allocation utilizes parameters such as bit allocation parameters in the bitstream and the number of available bytes to allocate the available bits to each lower channel of the hybrid coded multi-channel stereo, thereby completing subsequent interval decoding, inverse quantization, and neural network inverse transform steps. Hybrid coding bit allocation includes the following parts:

[0432] Bit allocation for the silence frame channel. The function of the bit allocation processing module for the silence frame channel is to complete the bit allocation of the mixed signal silence frame based on the allocation strategy parameter mixAllocStrategy of the mixed signal of the bed signal and the object signal obtained by decoding in the bit stream, and the silence enable flag HasSilFlag and the silence flag silFlag obtained by decoding in the bit stream.

[0433] Step 1: Mixed-encoded silence frame bit allocation processing.

[0434] The hybrid-coded silence frame bit allocation processing submodule completes the bit allocation of hybrid-coded silence frames based on the silence frame marker parameters HasSilFlag and silFlag obtained from the decoded bitstream. The following situations and corresponding handling exist:

[0435] Case 1: When HasSilFlag is parsed to be 0, it means that the current frame does not have the silent frame processing mode enabled or there is no silent frame in the current frame. The mixed encoding silent frame bit allocation processing submodule does not perform any other operations.

[0436] Case 2: When HasSilFlag is parsed to be 1, it indicates that the current frame has enabled silent frame processing and a silent frame exists. At this time, silFlag[i] of all channels is traversed. When silFlag[i] is 1, the number of bytes of the channel channelBytes[i] is set to the minimum safe number of bytes, safetyBytes. The value of the minimum safe number of bytes, safetyBytes, is related to the requirements of the quantization and interval encoding modules on the number of input bytes. For example, it can be set to 10 bytes here.

[0437] Update the pre-allocated bytes of the object `objAvailableBytes`. Iterate through the object channels where `silFlag[i]` is 1. For each object channel where `silFlag[i]` is 1, perform the following operations:

[0438] objAvailbleBytes-=safetyBytes;

[0439] Update the pre-allocated bytes of the audio bed, `bedAvailableBytes`. Iterate through the audio bed channels where `silFlag[i]` is 1. For each audio bed channel where `silFlag[i]` is 1, perform the following operations:

[0440] bedAvailbleBytes -= safetyBytes;

[0441] Step 2: Silent frame remaining bit allocation strategy.

[0442] The function of the silent frame bit allocation strategy sub-module is to decide whether to allocate the remaining bits generated by the silent frame to the bed signal or the object signal according to the allocation strategy parameter mixAllocStrategy of the mixed signal of the bed signal and the object signal obtained by decoding in the bitstream when there is a silent frame. The specific allocation strategy is determined by the value of mixAllocStrategy. For the meaning of the values of mixAllocStrategy, please refer to the mixAllocStrategy section.

[0443] The embodiments of this application support 2 different silent frame remaining bit allocation strategies. First, perform pre-calculation:

[0444] Calculate the average number of bytes allocated to each object channel objAvgBytes according to the pre-allocated number of bytes for the object objAvailbleBytes and the number of object channels objNum. The calculation formula is as follows:

[0445] objAvgBytes[i] = floor(objAvailbleBytes / objNum);

[0446] If there are remaining bytes after equal distribution, split the remaining bytes into multiple 1Byte and re-allocate them in ascending order of the object signal numbers. That is, when sum(objAvgBytes[i]) < objAvailbleBytes,

[0447] objAvgBytes[0] += 1, and the same operation is performed on other object channels objAvgBytes[i] until sum(objAvgBytes[i]) == objAvailbleBytes.

[0448] Scheme 1: When mixAllocStrategy is 0, define the remaining bits of the object silent frame objSilLeftBytes with an initial value of 0. Traverse all silFlag[i] corresponding to the object channels. When silFlag[i] = 1, update the value of objSilLeftBytes, that is,

[0449] objSilLeftBytes += objAvailbleBytes[i] – safetyBytes; 0 <= i < objNum;

[0450] Continue until all the OBJ audio channels have been traversed.

[0451] Solution 2: When mixAllocStrategy is 1, define the remaining bits of the object's silent frame, objSilLeftBytes, with an initial value of 0. Iterate through silFlag[i] corresponding to all object channels. When silFlag[i] = 1, update the value of objSilLeftBytes.

[0452] objSilLeftBytes+=objAvailbleBytes[i]–safetyBytes; 0<=i <objNum;

[0453] Continue until all the OBJ audio channels have been traversed.

[0454] Update the pre-allocated bytes for the sound bed (bedAvailableBytes) and the pre-allocated bytes for the object (objAvailableBytes), for example, as follows:

[0455] bedAvailbleBytes+=objSilLeftBytes;

[0456] objAvailbleBytes-=objSilLeftBytes.

[0457] Adaptation before non-silent frame bit allocation. The input parameters for non-silent frame channel bit allocation are mapped to a continuous channel arrangement (the presence of silent frame channels will cause non-silent frame channels to be physically discrete), which facilitates the subsequent module's non-silent frame channel bit allocation processing.

[0458] Bit allocation for non-silent frame channels. Bit allocation for non-silent frame channels of the sound bed is performed using a general bit allocation module. Its function is to allocate the available bits to each downmixer channel in the multi-channel stereo sound of the sound bed object based on parameters such as the updated pre-allocated byte count bedAvailableBytes of the sound bed and the channel bit allocation ratio.

[0459] The number of available bytes input is denoted as availableBytes. Multi-channel stereo mode may have LFE channels. Generally, LFE channels have less effective spectral information and do not need to participate in the bit allocation process of multi-channel stereo mode; a fixed number of bits can be pre-allocated. The number of pre-allocated bits for LFE channels is related to the coding rate. Let the average code rate per channel pair be cpeRate, where cpeRate is the result of the total coding rate converted to a single channel pair. If cpeRate < 64kb / s, the number of bytes allocated to the LFE channel is 10; if cpeRate < 96kb / s, the number of bytes allocated to the LFE channel is 15; if cpeRate >= 96kb / s, the number of bytes allocated to the LFE channel is 20. If an LFE channel exists, the pre-allocated bytes for the LFE channel are deducted from the available bytes, and the remaining bytes are then allocated to the other channels besides the LFE channel.

[0460] The process of allocating available bytes to the remaining channels consists of four steps, as follows:

[0461] The first step is to allocate bits to each channel according to chBitRatios.

[0462] The number of bytes per channel can be expressed as:

[0463] channelBytes[i]=availableBytes*chBitRatios[i] / (1<<4).

[0464] Where (1<<4) represents the maximum range of values ​​for the channel bit allocation ratio chBitRatios.

[0465] The second step is to redistribute the remaining bytes to each channel according to the proportion represented by chBitRatios[i] if not all bytes were allocated in the first step.

[0466] Third step: If there are still bits remaining after the second step, then allocate the remaining bits to the channel that received the most bytes in the first step.

[0467] Step 4: If the number of bytes allocated to some channels exceeds the upper limit of the number of bytes for a single channel, the excess portion will be allocated to the remaining channels.

[0468] The bit allocation process for the non-silent frame channels of the object utilizes a general bit allocation module. Its function is to allocate available bits to each downmixer channel in the multi-channel stereo sound of the sound bed object based on parameters such as the updated available bytes (objAvailableBytes) and the channel bit allocation ratio. The specific bit allocation process for the non-silent frame channels of the object is the same as the bit allocation process for the non-silent frame channels of the sound bed signal.

[0469] Non-silent frame channel adaptation and restoration. The number of bytes output from the non-silent frame channel bit allocation processing is reverse mapped to a physical arrangement according to the aforementioned rules (the presence of silent frame channels will cause non-silent frame channels to be physically discretely arranged), which facilitates subsequent module interval decoding, inverse quantization and neural network inverse transformation steps.

[0470] Mix and upmix. Perform center / side (M / S) upmix on the two paired channels ch1 and ch2 indicated by the channel pair index channelPairIndex. The upmixing method is consistent with the M / S upmixing in two-channel stereo mode.

[0471] After M / S upmixing, the Modified Discrete Cosine Transform (MDCT) spectrum of the upmixed channels needs to be processed by inverse interaural level difference (ILD) to restore the amplitude differences between the channels. The inverse ILD processing procedure is as follows:

[0472] if (scaleFlag[i] == 1) {

[0473] factor = mcIld[i] / (1 << 4)

[0474] }else{

[0475] factor = (1 << 4) / mcIld[i]

[0476] }

[0477] mdctSpectrum[i]=factor*mdctSpectrum[i].

[0478] Where factor is the amplitude adjustment factor corresponding to the ILD parameter of the i-th channel, (1<<4) is the maximum quantization range of mcIld, and mdctSpectrum[i] represents the MDCT coefficient vector of the i-th channel.

[0479] The technical effects of the embodiments of this application are as follows: when the multi-channel signal is a mixed signal containing a sound bed signal and an object signal and the multi-channel signal contains a silent frame, different allocation strategies (mixAllocStrategy) for the mixed signal including the sound bed signal and the object signal are adopted to allocate the bits saved by the silent frame to other non-silent frames, thereby improving coding efficiency.

[0480] The improvements in this application embodiment are as follows: determining the pre-allocated bit count bedAvailableBytes of the sound bed and the pre-allocated total bit count objAvailableBytes of the object; determining whether the sound bed and the object include a silent frame; if a silent frame exists, allocating bits to the silent frame channel according to the side information silFlag[i] and mixAllocStrategy, and updating the pre-allocated bit count bedAvailableBytes of the sound bed and the pre-allocated total bit count objAvailableBytes of the object.

[0481] This application proposes a method for bit allocation mode bitstream in a bed-and-object hybrid mode. The method involves parsing the allocation strategy mixAllocStrategy, which includes a hybrid signal comprising a bed signal and an object signal, from the bitstream; and allocating bits to the silent frame channel according to the allocation strategy for the hybrid signal comprising the bed signal and the object signal.

[0482] Determine the pre-allocated bit count bedAvailableBytes of the sound bed and the pre-allocated total bit count objAvailableBytes of the object; determine whether the sound bed and the object include a silent frame; if a silent frame exists, allocate bits to the silent frame channel according to the side information silFlag[i] and mixAllocStrategy, and update the pre-allocated bit count bedAvailableBytes of the sound bed and the pre-allocated total bit count objAvailableBytes of the object.

[0483] Parse the silence flag information (including HasSilFlag and silFlag[i]) from the bitstream; determine whether a silence frame exists based on the silence flag information.

[0484] The silent frame channel is allocated bits based on the edge information silFlag[i] and mixAllocStrategy, and the pre-allocated bit count bedAvailableBytes of the sound bed and the total pre-allocated bit count objAvailableBytes of the object are updated.

[0485] The allocation strategy parameter mixAllocStrategy, which includes the mixed signal of the sound bed signal and the object signal, determines whether to allocate the remaining bits generated by the silence frame to the sound bed signal or the object signal.

[0486] The `mixAllocStrategy` 2-bit parameter represents the allocation strategy for the mixed signal, which includes the bed signal and the object signal. Its value range and meaning are as follows:

[0487] 0: If the extra bits generated by the Mute mechanism belong to the acoustic bed signal, the extra bits are allocated to other acoustic bed signals; if the extra bits belong to the object signal, the extra bits are allocated to other object signals.

[0488] 1: If the extra bits generated by the Mute mechanism belong to the acoustic bed signal, the extra bits are allocated to other acoustic bed signals; if the extra bits belong to the object signal, the extra bits are allocated to other acoustic bed signals.

[0489] 2: If the extra bits generated by the Mute mechanism belong to the sound bed signal, the extra bits will be allocated to other object signals.

[0490] 3: Retain.

[0491] The specific remaining bit allocation methods correspond to two different silent frame remaining bit allocation strategies. When the multi-channel signal is a mixed signal containing both the sound bed signal and the object signal, the object signal is treated as the sound bed signal and allocated bits together according to a unified bit allocation strategy. However, the sound bed signal and the object signal interfere with each other, resulting in a deterioration in quality for both.

[0492] This application provides a method for bit allocation of bitstreams in a hybrid mode for a sound bed object, specifically:

[0493] When the multi-channel signal is a mixed signal containing both the bed signal and the object signal, the bit allocation scaling factor is obtained from the bit stream decoding. The bit allocation scaling factor is used to characterize the relationship between the number of encoded bits of the bed signal and / or the object channel signal and the total number of available bits.

[0494] Based on the bit allocation ratio factor, determine the pre-allocated bit count bedAvailableBytes for the acoustic bed signal and the pre-allocated bit count objAvailableBytes for the object signal;

[0495] The number of bits allocated to each channel is determined based on the pre-allocated number of bits bedAvailableBytes of the acoustic bed signal and the pre-allocated number of bits objAvailableBytes of the object signal.

[0496] Decoding is performed based on the bit allocation and bitstream of each channel to obtain the decoded multi-channel signal.

[0497] The bit allocation ratio factor is the ratio of the number of encoded bits of the acoustic bed signal to the total number of available bits (bedBitsRatioFloat in the embodiment), or the ratio of the number of encoded bits of the target signal to the total number of available bits, or the ratio of the number of encoded bits of the acoustic bed signal to the number of encoded bits of the target signal, or the ratio of the number of encoded bits of the target signal to the number of encoded bits of the acoustic bed signal.

[0498] The bit allocation ratio factor is the ratio of the number of encoded bits of the acoustic bed signal to the total number of available bits. The specific method for determining the bit allocation ratio factor is as follows: parse the bit allocation ratio factor index (such as bedBitsRatio in the embodiment) from the bit stream, and determine the bit allocation ratio factor (such as bedBitsRatioFloat in the embodiment) based on the bit allocation ratio factor index.

[0499] The bit allocation scaling factor index can be either a coding index obtained by uniformly quantizing the bit allocation scaling factor or a coding index obtained by non-uniformly quantizing the bit allocation scaling factor.

[0500] The bit allocation scaling factor index and the bit allocation scaling factor can be linear or non-linear.

[0501] The formulas for calculating the pre-allocated bytes bedAvailableBytes and the pre-allocated bytes objAvailableBytes of the sound bed, based on the available bytes and the floating-point scaling factor bedBitsRatioFloat (which represents the proportion of the sound bed to the total bits), are as follows:

[0502] bedAvailbleBytes=floor(availableBytes*bedBitsRatioFloat);

[0503] objAvailbleBytes=availableBytes–bedAvailbleBytes.

[0504] The silence flag information (including HasSilFlag and silFlag[i]) is parsed from the bitstream. Based on the pre-allocated number of bits bedAvailableBytes of the sound bed signal, the pre-allocated number of bits objAvailableBytes of the object signal, and the silence flag information, bit allocation is performed to determine the number of bits allocated to each channel.

[0505] The steps of mixed encoding bit allocation are as follows: determine whether a silent frame exists based on the silence flag information; if a silent frame exists, allocate bits to the silent frame channel according to the side information silFlag[i] (and mixAllocStrategy), and update the pre-allocated bit count bedAvailableBytes of the sound bed signal and the pre-allocated total bit count objAvailableBytes of the object signal; allocate bits to the non-silent frame channel according to the non-silent frame bit allocation principle (including three steps: non-silent frame bit allocation adaptation, non-silent frame bit allocation and non-silent frame bit allocation adaptation restoration).

[0506] The encoding end determines the bit allocation scaling factor;

[0507] The factor is quantized and encoded to obtain the index of the bit allocation ratio factor;

[0508] Write the index into the bitstream.

[0509] The bit allocation scaling factor index and the bit allocation scaling factor can be linear or non-linear.

[0510] The scaling factor is predefined according to the coding parameters.

[0511] The encoding parameters include: encoding rate and signal characteristic parameters. These encoding parameters are predetermined.

[0512] The encoding parameters are adaptively determined based on the characteristics of each frame of signal, such as the type of signal.

[0513] The encoder determines the hybrid allocation strategy and carries it in the bitstream. The encoder then sends the strategy to the decoder.

[0514] When the mute enable flag includes both the object mute enable flag and the sound bed mute enable flag, the allocation strategy for the sound bed object mixing signal can also include other modes, such as:

[0515] Mode 1: The object mute enable flag is set to 1, and the extra bits generated by the existence of a mute channel in the object signal are allocated to other non-mute channels in the object channel.

[0516] Mode 2: The object mute enable flag is set to 1, and the extra bits generated by the existence of a mute channel in the object signal are allocated to the channel where the sound bed signal is located.

[0517] Mode 3: The mute enable flag of the sound bed is set to 1, and the extra bits generated by the presence of a mute channel in the sound bed signal are allocated to other non-mute channels in the sound bed channel.

[0518] Mode 4: The sound bed mute enable flag is set to 1, and the extra bits generated by the existence of a mute channel in the sound bed signal are allocated to the channel where the object signal is located.

[0519] Mode 5: Both the sound bed mute enable flag and the object mute enable flag are 1, and the extra bits generated by the existence of a mute channel in the object signal are allocated to other non-mute channels in the object channel.

[0520] Mode 6: Both the bed mute enable flag and the object mute enable flag are 1, and the extra bits generated by the existence of a mute channel in the object signal are allocated to other non-mute channels in the bed mute channel.

[0521] In other embodiments of this application, the mixed-signal coding improvement scheme is as follows:

[0522] The mixed-signal coding mode in the AVS3P3 standard supports encoding and decoding of both bed and object signals. In practical applications, a large number of silent frames exist in both bed and object signals; proper handling of these silent frames can effectively improve the coding efficiency of the mixed signal. Therefore, this proposal presents an efficient mixed-signal coding method that improves the coding quality by rationally allocating bits between silent and non-silent frames in the bed and object signals. Furthermore, the bit allocation strategy for the mixed signal is implemented at the encoding end, while the decoding end does not distinguish between bed and object signals during the bit allocation stage. Specific implementation schemes include:

[0523] The mute enable flag is denoted as HasSilFlag, and the mute flag for the i-th channel in each channel is denoted as silFlag[i]. The mute enable flag operates on the mute enable flags of other channels in the multi-channel signal that do not contain the LFE channel signal. For example, HasSilFlag is used to indicate whether there is a mute frame in each channel other than the LFE channel. In each channel, excluding the LFE channel, the SilFlag corresponding to each channel is used to indicate whether that channel is a mute channel.

[0524] The field chBitRatios[i] was changed from appearing only on non-LFE channels to appearing only on non-LFE and non-mute channels; the number of bits in chBitRatios[i] was changed from 4 to 6.

[0525] The ILD side information has been changed from 4 bits of inter-channel amplitude difference parameter and 1 bit of scaling flag parameter to 5 bits of scaling factor codebook index.

[0526] The syntax for multi-channel stereo decoding is shown in Table 2 below, which is the Avs3McDec() syntax.

[0527]

[0528] The syntax for multi-channel stereo side information is shown in Table 3 below, which is the syntax for DecodeMcSideBits().

[0529]

[0530]

[0531] The semantic McBitsAllocationHasSiL() function allocates multi-channel stereo bits.

[0532] coupleChNum is the number of all other channels in the multichannel signal that do not include the LFE channel.

[0533] HasSilFlag occupies 1 bit and indicates whether there are silent frames in each channel of the current frame of the audio signal. 0 indicates that there are no silent frames and 1 indicates that there are silent frames.

[0534] silFlag[i] occupies 1 bit, where 0 indicates that the i-th channel is a non-silent frame, and 1 indicates that the i-th channel is a silent frame.

[0535] mcIld[ch1] and mcIld[ch2] occupy 5 bits and are codebook indices for the quantization of the ILD parameter of the inter-channel amplitude difference of each channel in the current channel pair, used to recover the amplitude of the decoded spectrum.

[0536] pairCnt occupies 4 bits and is used to represent the number of channel pairs in the current frame.

[0537] The channel pair index is represented as channelPairIndex. The number of bits in channelPairIndex is related to the total number of channels, as shown in Note 1 of the table above. The index used to represent the channel pair can be parsed to obtain the index values ​​of the two channels in the current channel pair, namely ch1 and ch2.

[0538] chBitRatios occupies 6 bits and represents the bit allocation ratio for each channel.

[0539] The decoding process is as follows:

[0540] Mixed signal bit allocation. Based on the mute channel markers obtained from decoding in the bitstream and the bit allocation ratio parameters, the remaining available bits after removing other side information are allocated to each downmixer in the multichannel stereo, thereby completing the subsequent interval decoding, inverse quantization, and neural network inverse transform steps.

[0541] The number of available bytes remaining after deducting other edge information in the current frame is denoted as availableBytes.

[0542] Multi-channel stereo mode may have a mute channel. The mute channel does not need to participate in the bit allocation process of multi-channel stereo mode; it only needs to be pre-allocated a fixed number of bytes, which is 8 bytes. If a mute channel exists, the pre-allocated number of bytes for the mute channel is deducted from the available bytes, and the remaining bytes are then allocated to the other channels except for the mute channel.

[0543] The process of allocating available bytes to the remaining channels consists of five steps, as follows:

[0544] The first step is to pre-allocate 8 safe bytes (safeBits) to each channel. These safe bytes are then deducted from the available bytes (availableBytes), and the remaining bytes (availableBytes) are used for allocation in subsequent steps.

[0545] The second step is to allocate bits to each channel according to chBitRatios. The number of bytes per channel can be expressed as:

[0546] channelBytes[i]=availableBytes*chBitRatios[i] / (1<<6).

[0547] Where (1<<6) represents the maximum range of values ​​for the channel bit allocation ratio chBitRatios.

[0548] Third, if not all bytes were allocated in the second step, the remaining bytes are redistributed to each channel according to the ratio represented by chBitRatios[i].

[0549] Fourth step: If there are still bits remaining after the third step, then allocate the remaining bits to the channel that received the most bytes in step 1.

[0550] Fifth, if the number of bytes allocated to some channels exceeds the upper limit of the number of bytes for a single channel, the excess portion will be allocated to the remaining channels.

[0551] The upmixing process will be explained next. For the two paired channels ch1 and ch2 indicated by the channel pair index `channelPairIndex`, M / S upmixing is performed, using the same upmixing method as in two-channel stereo M / S upmixing. After M / S upmixing, the MDCT spectrum of the upmixed channels needs to be processed by inverse ILD to restore the amplitude differences between the channels. The pseudocode for inverse ILD processing is as follows:

[0552] factor=mcIldCodebook[mcIld[i]],

[0553] mdctSpectrum[i]=factor*mdctSpectrum[i].

[0554] Where factor is the amplitude adjustment factor corresponding to the ILD parameter of the i-th channel, and mcIldCodebook is the quantization codebook of the ILD parameter as shown in Table 4 below, where mcIld[i] represents the codebook index corresponding to the ILD parameter of the i-th channel, and mdctSpectrum[i] represents the MDCT coefficient vector of the i-th channel. Table 4 below shows the mcILD codebook:

[0555]

[0556]

[0557] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0558] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.

[0559] Please see Figure 10 As shown in the embodiment of this application, an encoding device 1000 may include: a silence marker information acquisition module 1001, a multi-channel encoding module 1002, and a bitstream generation module 1003, wherein...

[0560] A mute flag information acquisition module is used to acquire mute flag information of multi-channel signals, wherein the mute flag information includes: a mute enable flag, and / or a mute flag;

[0561] A multi-channel encoding module is used to perform multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel;

[0562] The bitstream generation module is used to generate a bitstream based on the transmission channel signals of each transmission channel and the silence marker information. The bitstream includes the multi-channel encoding results of the silence marker information and the transmission channel signals.

[0563] Please see Figure 11 As shown in the embodiment of this application, a decoding device 1100 may include: a parsing module 1101 and a processing module 1102, wherein,

[0564] The parsing module is used to parse the mute marker information from the bitstream of the encoding device and determine the encoding information of each transmission channel based on the mute marker information. The mute marker information includes: a mute enable flag and / or a mute flag.

[0565] The processing module is used to decode the encoded information of each transmission channel to obtain the decoded signal of each transmission channel;

[0566] The processing module is also used to perform multi-channel decoding processing on the decoding signals of each transmission channel to obtain multi-channel decoding output signals.

[0567] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.

[0568] This application also provides a computer storage medium storing a program that performs some or all of the steps described in the above method embodiments.

[0569] The following describes another encoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 12 As shown, the encoding device 1200 includes:

[0570] Receiver 1201, transmitter 1202, processor 1203, and memory 1204 (wherein the encoding device 1200 may contain one or more processors 1203). Figure 12 (Taking a processor as an example). In some embodiments of this application, the receiver 1201, transmitter 1202, processor 1203, and memory 1204 can be connected via a bus or other means, wherein, Figure 12 Taking the example of a connection between China and Israel via a bus.

[0571] Memory 1204 may include read-only memory and random access memory, and provides instructions and data to processor 1203. A portion of memory 1204 may also include non-volatile random access memory (NVRAM). Memory 1204 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.

[0572] Processor 1203 controls the operation of the encoding device; processor 1203 can also be called a central processing unit (CPU). In specific applications, the various components of the encoding device are coupled together through a bus system, which includes not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.

[0573] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1203 or by instructions in the form of software. The processor 1203 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1204. Processor 1203 reads the information in memory 1204 and completes the steps of the above method in conjunction with its hardware.

[0574] The receiver 1201 can be used to receive input digital or character information and generate signal inputs related to the settings and function control of the encoding device. The transmitter 1202 may include a display device such as a display screen and can be used to output digital or character information through an external interface.

[0575] In this embodiment, processor 1203 is used to execute the aforementioned embodiments. Figure 4 , Figure 6 , Figure 7 The method shown is executed by the encoding device.

[0576] The following describes another decoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 13 As shown, the decoding device 1300 includes:

[0577] Receiver 1301, transmitter 1302, processor 1303, and memory 1304 (wherein the decoding device 1300 may contain one or more processors 1303). Figure 13 (Taking a processor as an example). In some embodiments of this application, the receiver 1301, transmitter 1302, processor 1303, and memory 1304 can be connected via a bus or other means, wherein... Figure 13 Taking the example of a connection between China and Israel via a bus.

[0578] Memory 1304 may include read-only memory and random access memory, and provides instructions and data to processor 1303. A portion of memory 1304 may also include NVRAM. Memory 1304 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.

[0579] Processor 1303 controls the operation of the decoding device; processor 1303 can also be referred to as a CPU. In specific applications, the various components of the decoding device are coupled together through a bus system, which includes not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.

[0580] The methods disclosed in the embodiments of this application can be applied to processor 1303, or implemented by processor 1303. Processor 1303 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1303 or by instructions in the form of software. The processor 1303 can be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1304, and processor 1303 reads the information in memory 1304 and completes the steps of the above method in combination with its hardware.

[0581] In this embodiment, processor 1303 is used to execute the aforementioned embodiments. Figure 5 , Figure 8 , Figure 9 The method shown is performed by the decoding device.

[0582] In another possible design, when the encoding or decoding device is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer-executable instructions stored in the storage unit to cause the chip within the terminal to execute the audio encoding method of any of the first aspects or the audio decoding method of any of the second aspects described above. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the terminal, such as read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0583] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of programs in the first or second aspect of the above methods.

[0584] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0585] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0586] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0587] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method for encoding multi-channel signals, characterized in that, include: Obtain mute marker information for multi-channel signals. The mute marker information includes: a mute enable flag and / or a mute flag. The mute enable flag is used to indicate whether mute detection is enabled. The mute flag is used to indicate whether each channel affected by the mute enable flag is a mute channel. The mute channel is a channel that does not need to be encoded or a channel that needs to be encoded according to low bits. The multi-channel signal is subjected to multi-channel encoding processing to obtain the transmission channel signal of each transmission channel; A bitstream is generated based on the transmission channel signals of each transmission channel and the silence marker information. The bitstream includes the multi-channel encoding results of the silence marker information and the transmission channel signals.

2. The method according to claim 1, characterized in that, The multi-channel signal includes: a sound bed signal, and / or an object signal; The mute flag information includes: the mute enable flag; the mute enable flag includes: a global mute enable flag, or a partial mute enable flag, wherein... The global mute enable flag is a mute enable flag that applies to the multi-channel signal; or, The partial mute enable flag is a mute enable flag that operates on a portion of the channels in the multi-channel signal.

3. The method according to claim 2, characterized in that, When the mute enable flag is the partial mute enable flag. The partial mute enable flag is either an object mute enable flag applied to the object signal, or a sound bed mute enable flag applied to the sound bed signal, or a mute enable flag applied to other channel signals in the multi-channel signal that do not contain non-low frequency effect LFE channel signals, or a mute enable flag applied to channel signals participating in pairing in the multi-channel signal.

4. The method according to any one of claims 1 to 3, characterized in that, The multi-channel signal includes: a sound bed signal and an object signal; The mute flag information includes: the mute enable flag; the mute enable flag includes: a sound bed mute enable flag and an object mute enable flag. The mute enable flag occupies a first bit and a second bit. The first bit is used to carry the value of the sound bed mute enable flag, and the second bit is used to carry the value of the object mute enable flag.

5. The method according to any one of claims 1 to 3, characterized in that, The mute flag information includes: the mute enable flag; The mute enable flag is used to indicate whether the mute marker detection function is enabled; or... The mute enable flag is used to indicate whether each channel of the multi-channel signal needs to be muted; or, The mute enable flag is used to indicate whether each channel of the multi-channel signal is a non-mute channel.

6. The method according to any one of claims 1 to 3, characterized in that, The acquisition of silence marker information for multi-channel signals includes: The silence marker information is obtained according to the control signaling of the input encoding device; or... The silence marker information is obtained according to the encoding parameters of the encoding device; or... Silence marker detection is performed on each channel of the multi-channel signal to obtain the silence marker information.

7. The method according to claim 6, characterized in that, The mute flag information includes: the mute enable flag and the mute flag; The step of detecting silence markers in each channel of the multi-channel signal to obtain the silence marker information includes: The mute marker detection is performed on each channel of the multi-channel signal to obtain the mute marker of each channel; The mute enable flag is determined based on the mute flag of each channel.

8. The method according to any one of claims 1 to 3, characterized in that, Before acquiring the silence marker information of the multi-channel signal, the method further includes: The multi-channel signal is preprocessed to obtain a preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding. The acquisition of silence marker information for multi-channel signals includes: The preprocessed multi-channel signal is subjected to silence marker detection to obtain the silence marker information.

9. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The multi-channel signal is preprocessed to obtain a preprocessed multi-channel signal. The preprocessing includes at least one of the following: transient detection, window type determination, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and bandwidth extension coding. The silence marker information is corrected based on the preprocessed multi-channel signal.

10. The method according to any one of claims 1 to 3, characterized in that, The step of generating a bitstream based on the transmission channel signals of each transmission channel and the silence marker information includes: The initial multi-channel processing method is adjusted according to the mute mark information to obtain the adjusted multi-channel processing method; The transmission channel signals of each transmission channel are encoded according to the adjusted multi-channel processing method to obtain the bitstream.

11. The method according to any one of claims 1 to 3, characterized in that, The step of generating a bitstream based on the transmission channel signals of each transmission channel and the silence marker information includes: Based on the mute flag information, the number of available bits, and the multi-channel side information, bit allocation is performed for each transmission channel to obtain the bit allocation result for each transmission channel; The transmission channel signals of each transmission channel are encoded according to the bit allocation results of each transmission channel to obtain the bit stream.

12. The method according to claim 11, characterized in that, The step of allocating bits to each transmission channel based on the silence marker information, the number of available bits, and the multi-channel side information includes: Based on the available number of bits and multi-channel side information, bits are allocated to each transmission channel according to the bit allocation strategy corresponding to the mute marker information.

13. The method according to claim 11, characterized in that, The multi-channel side information includes: channel bit allocation ratio. The channel bit allocation ratio is used to indicate the bit allocation ratio between the non-low frequency effect (LFE) channels in the multi-channel signal.

14. The method according to claim 6, characterized in that, The step of detecting silence markers for each channel of the multi-channel signal includes: Based on the signals of each channel in the current frame of the multi-channel signal, determine the signal energy of each channel in the current frame; Based on the signal energy of each channel in the current frame, determine the silence detection parameters for each channel in the current frame; Based on the silence detection parameters of each channel in the current frame and the preset silence detection threshold, the silence flag of each channel in the current frame is determined.

15. The method according to any one of claims 1 to 3, characterized in that, The step of performing multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel includes: The multi-channel signal is filtered to obtain the filtered multi-channel signal; The filtered multi-channel signals are then processed to obtain multi-channel paired signals and multi-channel side information. The multi-channel group signal is downmixed based on the multi-channel side information to obtain the transmission channel signal of each transmission channel.

16. The method according to claim 15, characterized in that, The multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index; The codebook index for the inter-channel amplitude difference parameter quantization is used to indicate the codebook index for the quantization of the inter-channel amplitude difference (ILD) parameter of each channel in the multi-channel signal. The number of channel pairs is used to represent the number of channel pairs in the current frame of the multichannel signal; The channel pair index is used to represent the index of a channel pair.

17. A method for decoding multi-channel signals, characterized in that, include: The mute flag information is parsed from the bitstream of the encoding device, and the encoding information of each transmission channel is determined based on the mute flag information. The mute flag information includes: a mute enable flag and / or a mute flag. The mute enable flag is used to indicate whether mute detection is enabled, and the mute flag is used to indicate whether each channel affected by the mute enable flag is a mute channel. The mute channel is a channel that does not need to be encoded or a channel that needs to be encoded according to low bits. The encoded information of each transmission channel is decoded to obtain the decoded signal of each transmission channel; The decoding signals of each transmission channel are subjected to multi-channel decoding processing to obtain multi-channel decoded output signals.

18. The method according to claim 17, characterized in that, The step of parsing the silence marker information from the bitstream of the encoding device includes: Parse the mute flags of each channel from the bitstream; or, The mute enable flag is parsed from the bitstream. If the mute enable flag is a first value, the mute flag is parsed from the bitstream; or... Parse the bed mute enable flag and / or object mute enable flag, as well as the mute flag for each channel, from the bitstream; or, Parse the bed mute enable flag and / or object mute enable flag from the bitstream; based on the bed mute enable flag and / or object mute enable flag, parse the mute flags of some channels of each channel from the bitstream.

19. The method according to claim 17, characterized in that, Decoding the encoded information of each transmission channel includes: Multi-channel side information is parsed from the bitstream; Bit allocation is performed on each transmission channel based on the multi-channel side information and the mute flag information to obtain the number of encoded bits for each transmission channel; The encoded information of each transmission channel is decoded according to the number of encoded bits of each transmission channel.

20. The method according to claim 17, characterized in that, After performing multi-channel decoding processing on the decoded signals of each transmission channel to obtain a multi-channel decoded output signal, the method further includes: The multi-channel decoded output signal is post-processed, and the post-processing includes at least one of the following: bandwidth extension decoding, inverse time-domain noise shaping, inverse frequency-domain noise shaping, and inverse time-frequency transformation.

21. The method according to claim 19, characterized in that, The multi-channel side information includes at least one of the following: inter-channel amplitude difference parameter quantization codebook index, number of channel pairs, and channel pair index; The codebook index for the inter-channel amplitude difference parameter quantization is used to indicate the codebook index for the quantization of the inter-channel amplitude difference (ILD) parameter of each channel. The number of channel pairs is used to represent the number of channel pairs in the current frame of the multichannel signal; The channel pair index is used to represent the index of a channel pair.

22. An encoding device, characterized in that, The encoding device includes: A mute flag information acquisition module is used to acquire mute flag information of multi-channel signals. The mute flag information includes: a mute enable flag and / or a mute flag. The mute enable flag is used to indicate whether mute detection is enabled. The mute flag is used to indicate whether each channel affected by the mute enable flag is a mute channel. The mute channel is a channel that does not need to be encoded or a channel that needs to be encoded according to low bits. A multi-channel encoding module is used to perform multi-channel encoding processing on the multi-channel signal to obtain the transmission channel signal of each transmission channel; The bitstream generation module is used to generate a bitstream based on the transmission channel signals of each transmission channel and the silence marker information. The bitstream includes the multi-channel encoding results of the silence marker information and the transmission channel signals.

23. A decoding device, characterized in that, The decoding device includes: The parsing module is used to parse the mute marker information from the bitstream of the encoding device and determine the encoding information of each transmission channel based on the mute marker information. The mute marker information includes: a mute enable flag and / or a mute flag. The mute enable flag is used to indicate whether mute detection is enabled. The mute flag is used to indicate whether each channel affected by the mute enable flag is a mute channel. The mute channel is a channel that does not need to be encoded or a channel that needs to be encoded according to low bits. The processing module is used to decode the encoded information of each transmission channel to obtain the decoded signal of each transmission channel; The processing module is also used to perform multi-channel decoding processing on the decoding signals of each transmission channel to obtain multi-channel decoding output signals.

24. A terminal device, characterized in that, The terminal device includes: a processor and a memory; the processor and the memory communicate with each other. The memory is used to store instructions; The processor is configured to execute the instructions in the memory and perform the method as described in any one of claims 1 to 16.

25. A terminal device, characterized in that, The terminal device includes: a processor and a memory; the processor and the memory communicate with each other. The memory is used to store instructions; The processor is configured to execute the instructions in the memory, performing the method as described in any one of claims 17 to 21.

26. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 16, or 17 to 21.

27. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 16, or 17 to 21.

28. A computer-readable storage medium, characterized in that, The device stores a bitstream generated by the method as described in any one of claims 1 to 16.

Citation Information

Patent Citations

  • Apparatus and method for encoding / decoding audio signal using mute interval information

    KR1020120069906A

  • Apparatus for processing media signal and method thereof

    TW200803588A