Audio encoding method, apparatus, electronic device, and storage medium
The audio encoding method addresses redundancy issues in multi-channel audio transmission by grouping channels, transforming frequency domains, and applying decorrelation, resulting in reduced transmission and storage costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2024-04-12
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional 2D mid-side audio algorithms increase transmission costs and data waste during multi-channel audio transmission due to redundancy.
An audio encoding method that groups channel sequences into groups with identical channels, performs frame-by-frame frequency domain transformation, determines a target transformation matrix for each frequency band, and applies same-band decorrelation processing to reduce redundancy and obtain encoded streams for decoding.
Reduces redundancy between channels, easing the encoder's load and lowering transmission and storage costs by compressing multi-channel audio signals effectively.
Smart Images

Figure 2026514117000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to the audio processing technology, and more particularly to audio encoding methods, audio decoding methods and apparatus thereof, electronic devices, storage media, computer program products and computer programs. [Background technology]
[0002] With the advancement of multimedia technology, the demands on audio signals are increasing. While conventional 2D mid-side audio algorithms (2D M / S) can effectively reduce data redundancy between multiple channels, in multiple different audio scenarios, conventional algorithms increase transmission costs during data transmission, resulting in data waste. [Overview of the project] [Problems that the invention aims to solve]
[0003] The embodiments of this disclosure provide an audio encoding method, an audio decoding method and apparatus therefor, an electronic device, a computer-readable storage medium, a computer program product, and a computer program, thereby solving problems such as the waste of transmission and storage media during the transmission of multi-channel audio. The technical proposal of this disclosure is as follows. [Means for solving the problem]
[0004] According to a first aspect, an embodiment of the present disclosure provides an audio coding method performed by an encoder, the method comprising: grouping a channel sequence to obtain a plurality of channel groups, each of which includes a plurality of consecutive channels in the channel sequence, wherein one or more identical channels exist between adjacent channel groups; performing a frame-by-frame frequency domain transformation on the audio signals of each channel in the channel sequence to obtain frequency domain coefficients for each frame of each channel; determining a target transformation matrix for each frequency band in a frequency band set corresponding to the channel group from a set of transformation matrices based on the frequency domain coefficients of each channel; performing a same-band decorrelation process on the frequency domain coefficients of the channels in the channel group based on the target transformation matrix for each frequency band to obtain coding information for the channel group; and obtaining a coded stream based on the coding information for the channel group and transmitting the coded stream to a decoder for decoding.
[0005] According to a second aspect, an embodiment of the present disclosure provides an audio decoding method performed by a decoder, the method comprising: receiving an encoded stream transmitted from an encoder, the encoded stream comprising encoded information for a plurality of channel groups, the channel groups being obtained by sequentially grouping a channel sequence, each channel group comprising a plurality of consecutive channels in the channel sequence, with one or more identical channels between adjacent channel groups; decoding the plurality of channel groups sequentially, and for the decoded current channel group, determining a target decoding matrix for each frequency band in a set of frequency bands corresponding to the current channel group based on the encoded information of the current channel group; obtaining decoding frequency domain coefficients for the current channel group based on the target decoding matrix of the current channel group in each frequency band, for the encoded information of the current channel group; and obtaining a decoded audio signal for each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups.
[0006] According to a third aspect, an embodiment of the present disclosure provides an audio encoding device comprising: a channel grouping module configured to group the channel sequence to obtain a plurality of channel groups, each of which includes a plurality of consecutive channels in the channel sequence, and where one or more identical channels exist between adjacent channel groups; a frequency domain processing module configured to perform a frame-by-frame frequency domain transformation on the audio signals of each channel in the channel sequence to obtain frequency domain coefficients for each frame of each channel; a matrix determination module configured to determine a target transformation matrix for each frequency band in a frequency band set corresponding to the channel group from a set of transformation matrices based on the frequency domain coefficients of each channel; an encoding module configured to perform a same-band decorrelation process on the frequency domain coefficients of the channels in the channel group based on the target transformation matrix for each frequency band to obtain encoding information for the channel group; and a transmission module configured to obtain an encoded stream based on the encoding information for the channel group and transmit the encoded stream to a decoder for decoding.
[0007] According to a fourth aspect, an embodiment of the present disclosure provides an audio decoding device, the device comprising: a receiving module configured to receive an encoded stream transmitted from an encoder, the encoded stream comprising encoded information for a plurality of channel groups, the channel groups being obtained by sequentially grouping a channel sequence, each channel group comprising a plurality of consecutive channels in the channel sequence, and one or more identical channels existing between adjacent channel groups; a matrix determination module configured to sequentially decode the plurality of channel groups and, for the decoded current channel group, determine a target decoding matrix for each frequency band in a frequency band set corresponding to the current channel group based on the encoded information of the current channel group; and a decoding module configured to obtain the decoding frequency domain coefficients for the current channel group based on the target decoding matrix of the current channel group in each frequency band, and to obtain the decoded audio signals for each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups.
[0008] According to a fifth aspect, an embodiment of the present disclosure provides an encoder comprising a processor and a memory for storing instructions that can be executed by the processor, wherein the processor is configured to perform steps of the method described in the first aspect of an embodiment of the present disclosure.
[0009] According to a sixth aspect, an embodiment of the present disclosure provides a decoder comprising a processor and a memory for storing instructions executable by the processor, wherein the processor is configured to perform steps of the method described in a second aspect of an embodiment of the present disclosure.
[0010] According to the seventh aspect, an embodiment of the present disclosure provides a computer-readable storage medium in which computer program instructions are stored, and when the program instructions are executed by a processor, steps of the method of the first aspect of the embodiment of the present disclosure are realized.
[0011] According to the eighth aspect, an embodiment of the present disclosure is one in which computer program instructions are stored A computer-readable storage medium is provided, and the program instructions are executed by a processor, thereby realizing the steps of the method according to a second embodiment of the embodiments of the present disclosure.
[0012] According to the ninth aspect, an embodiment of the present disclosure provides an encoder, the device comprising a processor and an interface circuit, the interface circuit being used to receive and transmit code instructions to the processor, the processor being used to cause the device to perform the method described in the first aspect by executing the code instructions.
[0013] According to the tenth aspect, an embodiment of the present disclosure provides a decoder, the device comprising a processor and an interface circuit, the interface circuit being used to receive and transmit code instructions to the processor, the processor being used to cause the device to perform the method described in the second aspect by executing the code instructions.
[0014] According to the eleventh aspect, an embodiment of the present disclosure provides an encoding and decoding system, the system comprising an encoding device described in the third aspect and a decoding device described in the fourth aspect, or the system comprising an encoder described in the fifth aspect and a decoder described in the sixth aspect, or the system comprising an encoder described in the seventh aspect and an encoding device described in the eighth aspect, or the system comprising an encoder described in the ninth aspect and a decoder described in the tenth aspect.
[0015] According to the 12th aspect, an embodiment of the present invention provides a computer-readable storage medium storing instructions used by the above encoder, and when the instructions are executed, the encoder is caused to execute the method described in the 1st aspect.
[0016] According to the 13th aspect, an embodiment of the present invention provides a computer-readable storage medium storing instructions used by the above decoder, and when the instructions are executed, the decoder is caused to execute the method described in the 2nd aspect.
[0017] According to the 14th aspect, an embodiment of the present disclosure further provides a computer program product including a computer program, and when it is executed on a computer, the computer is caused to execute the method described in the 1st aspect.
[0018] According to the 15th aspect, an embodiment of the present disclosure further provides a computer program product including a computer program, and when it is executed on a computer, the computer is caused to execute the method described in the 2nd aspect.
[0019] According to the 16th aspect, an embodiment of the present disclosure provides a chip system, and the chip system includes at least one processor and an interface to support a network device to realize the function according to the 1st aspect, for example, to determine or process at least one of the data and information related to the above method. In one possible design, the chip system further includes a memory, and the memory is used to store a computer program and data required by the network device. The chip system may be composed of a chip or may include a chip and other individual components.
[0020] According to the 17th aspect, an embodiment of the present disclosure provides a chip system, which includes at least one processor and an interface to support a terminal device in realizing the functions according to the 2nd aspect, for example, determining or processing at least one of the data and information according to the above method. In one possible design, the chip system further includes a memory, which is used to store computer programs and data required by the terminal device. The chip system may be composed of a chip or may include a chip and other individual components.
[0021] According to the 18th aspect, an embodiment of the present disclosure provides a computer program, which, when executed on a computer, causes the computer to execute the method described in the 1st aspect above.
[0022] According to the 19th aspect, an embodiment of the present disclosure provides a computer program, which, when executed on a computer, causes the computer to execute the method described in the 2nd aspect above.
Advantages of the Invention
[0023] The technical solutions provided by the embodiments of the present disclosure have at least the following beneficial effects. The encoder obtains frequency domain coefficients by performing frequency band division and grouping on the channel signal, and can determine the target transformation matrix for each frequency band corresponding to the channel group based on the frequency domain coefficients of each channel. Further, based on the target transformation matrix, decorrelation processing is performed on the frequency domain coefficients of the channel to obtain the encoded information of the channel group, and an encoded stream is obtained based on the encoded information and transmitted to the decoder for decoding. In the embodiments of the present disclosure, through the target transformation matrix, the frequency domain coefficients in each frequency band within the channel group are encoded, thereby realizing the compression of the audio signals of multiple channels, reducing the redundancy between multiple channels, reducing the burden on the encoder, and reducing the costs of transmission and storage.
[0024] The above general explanation and the following detailed explanation are illustrative and explanatory, and do not limit this disclosure. [Brief explanation of the drawing]
[0025] The drawings herein are incorporated into the specification and constitute part of this specification, illustrating embodiments consistent with this disclosure and illustrating the principles of this application together with the specification. [Figure 1] This is a flowchart of an audio encoding method according to one exemplary embodiment. [Figure 2] This is a flowchart of an audio encoding method according to another exemplary embodiment. [Figure 3] This is a schematic diagram of a cross-correlation matrix according to one exemplary embodiment. [Figure 4] This is a flowchart of an audio encoding method according to another exemplary embodiment. [Figure 5] This is a flowchart of an audio encoding method according to another exemplary embodiment. [Figure 6] This is a schematic diagram of an audio encoding method according to one exemplary embodiment. [Figure 7] This is a flowchart of an audio decoding method according to one exemplary embodiment. [Figure 8] This is a flowchart of an audio decoding method according to another exemplary embodiment. [Figure 9] This is a schematic diagram of an audio decoding method according to one exemplary embodiment. [Figure 10] This is a flowchart of an audio encoding method according to another exemplary embodiment. [Figure 11] This is a block diagram of an audio encoding device according to one exemplary embodiment. [Figure 12] This is a block diagram of an audio decoding device according to one exemplary embodiment. [Figure 13] This is a block diagram of an audio processing device according to one exemplary embodiment. [Figure 14] This is a block diagram of another audio processing chip according to an exemplary embodiment. [Modes for carrying out the invention]
[0026] Here, an exemplary embodiment is described, and this example is shown in the drawings. The following description is related to the drawings. Where applicable, unless otherwise specified, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary examples do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure, as detailed in the appended claims.
[0027] The terms used in the embodiments of this disclosure are intended to describe specific embodiments and are not intended to limit the embodiments of this disclosure. Unless otherwise clearly indicated in the context, the singular forms “one kind” and “the” used in the embodiments of this disclosure and the appended claims are intended to include the plural forms. The terms “and / or” as used herein refer to and include any or all possible combinations of one or more related items listed.
[0028] The terms "First," "Second," etc., in the specification, claims, and drawings of this disclosure are for the purpose of distinguishing similar subjects and are not necessarily used to describe a specific order or priority. The data used in this manner may, where appropriate, be used interchangeably so that the embodiments of this disclosure described herein are carried out in an order other than that illustrated or described herein.
[0029] Depending on the context, the term “in the case” as used herein may be understood as “when…” or “on the occasion of…” or “in response to a decision,” and for the sake of brevity and ease of understanding, the terms used herein to characterize greater than or less than relationships are “greater” or “less,” “higher” or “lower.” However, as will be understood by those skilled in the art, the term “greater” means “greater than or equal to,” the term “less than or equal to,” the term “higher” means “greater than or equal to,” and the term “lower” means “less than or equal to.”
[0030] The audio encoding / decoding methods disclosed in the embodiments of this disclosure can be applied to a variety of communication systems, such as 3rd Generation (3G) Universal Mobile Telecommunications Systems (UMTS), Long Term Evolution (LTE) systems, 5th Generation (5G) mobile communication systems, 5G New Radio (NR) systems, 6th Generation (6G) mobile communication systems, or other future advanced mobile communication systems. The audio encoding / decoding methods disclosed in the embodiments of this disclosure can also be applied to streaming media transmission systems and OTT (Over The Top) media transmission systems.
[0031] Figure 1 is a schematic flowchart of an audio encoding method provided by an embodiment of the present disclosure. The audio encoding method may be performed by an encoder. As shown in Figure 1, the method may include, but is not limited to, steps S101 to S105.
[0032] In S101, the channel sequence is grouped to obtain multiple channel groups, each channel group containing multiple consecutive channels within the channel sequence, with one or more identical channels existing between adjacent channel groups.
[0033] In one embodiment, the encoder can group M channels in a channel sequence to obtain a plurality of channel groups. In one embodiment, each channel group contains a plurality of consecutive channels in the channel sequence, which may include, for example, three consecutive channels, and in the embodiments of the present disclosure, adjacent channels. There are one or more identical channels between adjacent channel groups. The adjacent channel groups are the first channel group and the second channel group, respectively. Here, the first channel group and the second channel group each contain three consecutive channels in the channel sequence, and the first channel group and the second channel group each contain two identical channels. For example, five channels in a channel sequence are divided into channel group 1, channel group 2, and channel group 3. Here, channel group 1 contains channels 1, 2, and 3; channel group 2 contains channels 2, 3, and 4; and channel group 3 contains channels 3, 4, and 5.
[0034] In S102, frequency domain conversion is performed on the audio signal of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients for each frame of each channel.
[0035] In the embodiments of this disclosure, the encoder divides the audio signal of each channel in the channel sequence into a plurality of fixed-length frames, and performs a Modified Discrete Cosine Transform (MDCT) on each frame to obtain a frequency-domain representation of each frame. Based on the frequency-domain representation of each frame, the MDCT coefficients for each frame can be extracted from the frequency-domain representation and used as the frequency-domain coefficients for each frame.
[0036] In one embodiment, the channel sequence of the embodiment of the present disclosure includes M channels.
[0037] In one embodiment, the audio data for each frame of each channel may include 2N sampling points, and the sampling rate is f s Each frame after MDCT transformation may contain N frequency points, and correspondingly, the frequency spectral distribution range of the MDCT coefficients is (0, f s ( / 2), and the frequency resolution is f s It is / 2N.
[0038] In S103, based on the frequency domain coefficients of each channel, the target transformation matrix for each frequency band within the frequency band set corresponding to the channel group is determined from the transformation matrix set.
[0039] In one embodiment, the frequency band may first be pre-divided according to a psychoacoustic frequency band division method to obtain a set of frequency bands, where the set of frequency bands may contain multiple divided frequency bands, for example, the set of frequency bands may contain b divided frequency bands, where b is an integer of 1 or more. Each frequency band in the set of frequency bands has a different frequency range, and the frequency ranges of adjacent frequency bands are continuous.
[0040] In one embodiment, after performing an MDCT transformation on each frame for the channel, the frequency value of each frequency domain coefficient is determined by multiplying the sampling point order of each frequency domain coefficient by the frequency resolution.
[0041] In one embodiment, the formula for calculating the frequency value of any one of the frequency domain coefficients is as follows: f=n*f s / 2N (1)
[0042] Here, n represents the nth sampling point corresponding to any one frequency domain coefficient, the value of n is from 1 to N, and N is the number of sampling points, and according to equation (1), The frequency value f corresponding to the frequency domain coefficient can be obtained.
[0043] In one embodiment, the frequency values of the frequency domain coefficients are compared with the frequency ranges of each frequency band to obtain the frequency range in which the frequency domain coefficients are located, thereby determining the frequency domain coefficients within each frequency band.
[0044] In one embodiment, the cross-correlation coefficient between different channels in the same frequency band is calculated based on the frequency-domain coefficients between channels in the same frequency band. That is, for each frequency band in a set of frequency bands, the cross-correlation coefficient between any two channels corresponding to that frequency band can be obtained based on the frequency-domain coefficients between channels in the same frequency band.
[0045] In one embodiment, for any one frequency band b within a frequency band set, the cross-correlation coefficient between any two channels in a channel group in frequency band b can be determined from the cross-correlation coefficient between any two channels corresponding to frequency band b, based on the channels included in the channel group. In one embodiment, the target transformation matrix in frequency band b corresponding to the channel group is determined from a set of transformation matrices based on the cross-correlation coefficient between any two channels in a channel group in frequency band b. The frequency band set includes B frequency bands, and the target transformation matrix for the channel group in each frequency band can be obtained.
[0046] In S104, based on the target transformation matrix for each frequency band, a same-band decorrelation process is performed on the frequency domain coefficients of the channels within the channel group to obtain the coding information for the channel group.
[0047] In one embodiment, the transformation matrix set includes multiple transformation matrices, each of which can correspond to one decorrelated mode. In one embodiment, the transformation matrix set may include M0, M1, M2, M3, and M4. The specific values of the transformation matrices are as follows:
[0048]
number
[0049] In one embodiment, the decorrelation mode of a channel group in each frequency band can be determined based on the target transformation matrix of the channel group in each frequency band. If the target transformation matrices corresponding to different frequency bands are different, the decorrelation modes corresponding to those different frequency bands are also different. If the target transformation matrices corresponding to different frequency bands are the same, the decorrelation modes corresponding to those different frequency bands are the same. For example, if the target transformation matrix for frequency band 1 is M1, the target transformation matrix for frequency band 2 is M2, and the target transformation matrix for frequency band 3 is M1, then the decorrelation modes for frequency band 1 and frequency band 3 are the same, but the decorrelation modes for frequency band 1 and frequency band 3 are different from those for frequency band 2.
[0050] In one embodiment, frequency domain coefficients of channels within a channel group are distinguished by frequency band, frequency domain coefficients in each frequency band are obtained, and decorrelation is performed on the frequency domain coefficients of channels in the same frequency band within the channel group based on the target transformation matrix corresponding to the same frequency band to obtain the encoding information of the channel group.
[0051] For example, a channel group is defined as containing three consecutive channels. The three consecutive channels can be described as channel L, channel C, and channel R. For any one frequency band b within the frequency band set, it can be determined that the target transformation matrix for the channel group in frequency band b is M4. A frequency domain coefficient matrix can be constructed using the frequency domain coefficients of channels L, channel C, and channel R in frequency band b. Matrix operations can be performed using the frequency domain coefficient matrix for the channel group in frequency band b and the target transformation matrix M4 corresponding to frequency band b to obtain the first encoding information for the channel group in frequency band b. In other words, the target transformation matrix M4 corresponding to frequency band b can be used to decorrelate the frequency domain coefficients of channels L, channel C, and channel R within the channel group in frequency band b to obtain the encoding information for the channel group in frequency band b. By performing the same frequency decorrelation on channels L, channel C, and channel R within the channel group, interference at the same frequency can be reduced, thereby reducing redundancy and transmission costs.
[0052] Furthermore, by using the target transformation matrix corresponding to each frequency band, decorrelation can be performed on the frequency-domain coefficients of the channel group in each frequency band to obtain the encoded information of the channel group in each frequency band. After obtaining the encoded information corresponding to each frequency band, the encoded information of the channel group across the entire frequency band can be obtained based on the encoded information in each frequency band. Note that the encoded information of a channel group includes the encoded information of that channel group in all frequency bands.
[0053] In S105, the encoded stream is obtained based on the encoding information of the channel group, and the encoded stream is sent to the decoder for decoding.
[0054] In one embodiment, the encoding information of a channel group can be encoded in binary format to obtain a binary encoded stream. That is, the encoding information of each channel is converted into a binary code, and the binary codes of all channel groups are concatenated to form an encoded stream. In one embodiment, the original channel signals are restored by sending the encoded stream to a decoder for decoding.
[0055] Furthermore, in order for the decoder to perform decoding, it is necessary to transmit the target transformation matrix for the channel group in each frequency band. In one embodiment, the target transformation matrix for each frequency band may be written to the encoding stream and transmitted together with the channel encoding information, or it may be transmitted to the decoder independently in synchronization with the encoding stream.
[0056] In the audio encoding method according to the embodiments of this disclosure, the encoder obtains frequency domain coefficients by performing frequency band division and grouping on the channel signals, and can determine a target transformation matrix for each frequency band corresponding to the channel group based on the frequency domain coefficients of each channel. Furthermore, based on the target transformation matrix, it performs decorrelation on the frequency domain coefficients of the channels to obtain encoding information for the channel group, obtains an encoded stream based on the encoding information, and transmits it to the decoder for decoding. In the embodiments of this disclosure, the frequency domain coefficients in each frequency band within the channel group are encoded via the target transformation matrix, thereby enabling compression of audio signals of multiple channels, reducing redundancy between multiple channels, easing the load on the encoder, and lowering transmission and storage costs.
[0057] Figure 2 is a schematic flowchart of an audio encoding method provided by an embodiment of the present disclosure. The audio encoding method may be performed by an encoder. As shown in Figure 2... The method may include, but is not limited to, steps S201 to S207.
[0058] In S201, channel sequences are grouped to obtain multiple channel groups, each channel group containing multiple consecutive channels within the channel sequence, with one or more identical channels existing between adjacent channel groups.
[0059] In the embodiments of this disclosure, the implementation of S201 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0060] In S202, frequency domain transformation is performed on the audio signal of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients for each frame of each channel.
[0061] In the embodiments of this disclosure, the implementation of S202 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0062] In S203, the first cross-correlation matrix between channels corresponding to each frequency band is determined based on the frequency domain coefficients of each channel.
[0063] In one embodiment, the energy value of one channel in each frequency band is determined based on the frequency domain coefficient of any one channel, and a first cross-correlation matrix corresponding to each frequency band is determined based on the energy value of each channel in each frequency band.
[0064] In one embodiment, the root mean square (RMS) energy of any one channel can be calculated based on the frequency domain coefficient of any one channel to determine the energy value of any one channel in each frequency band. The formula for calculating RMS is as follows:
[0065]
number
[0066] Here, c represents the channel index value, which ranges from 1 to M, where M is the number of input channels, b is the frequency band index value, and X i N is the frequency domain coefficient of the corresponding channel. c,b This is the number of frequency points within the channel frequency band.
[0067] The number of frequency points can be calculated from the bandwidth and the frequency resolution corresponding to that bandwidth.
[0068] Typically, the data length of each channel is the same, and therefore, using the same frequency band division method, the data length within the same frequency band between different channels is also the same, and consequently, the number of frequency points within the same frequency band is the same, i.e., N c,b =N b That is the case.
[0069] In one embodiment, after obtaining the energy values of each channel in each frequency band, the energy ratio of any two channels in the channel sequence in any one frequency band b is obtained, and in one embodiment, the cross-correlation coefficient between any two channels in any one frequency band b is determined based on the energy ratio between channels in any one frequency band b.
[0070] In one embodiment, the energy ratio Q of any two channels in a channel sequence in any one frequency band b is obtained, and the formula for calculating the energy ratio Q is as follows:
[0071]
number
[0072] Here, c represents the channel index value, which ranges from 1 to M, where M is the number of input channels, b is the number of frequency band divisions, c1 may be the same as or different from c2, and b1 may be the same as or different from b2.
[0073] In one embodiment, energy discrimination can be performed using the energy ratio Q, and based on the energy discrimination result, the cross-correlation coefficient between channels in any one frequency band b is determined. If the energy ratio in any one frequency band b is less than or equal to a first set threshold, it is determined that the cross-correlation coefficient between any two channels in any one frequency band b is zero. If the energy ratio in any one frequency band b is greater than or equal to a second set threshold, it is determined that the cross-correlation coefficient between any two channels in any one frequency band b is zero. Here, the first set threshold is smaller than the second set threshold.
[0074] If the energy ratio in any one of the frequency bands b is between the first setting threshold and the second setting threshold, that is, if the energy ratio is greater than the first setting threshold and less than the second setting threshold, the cross-correlation coefficient of any two channels in any one of the frequency bands b is determined based on the frequency domain coefficients of any two channels in any one of the frequency bands b.
[0075] In other words, if the energy ratio Q of any two channels in any one frequency band b is too large or too small, the cross-correlation coefficient of those two channels is 0, and if the energy ratio Q of any two channels in any one frequency band b is (0.5, 2), the cross-correlation coefficient of those two channels is further calculated.
[0076] In one embodiment, the cross-correlation coefficient can be determined based on the magnitude of Q. The formula for calculating the cross-correlation coefficient is as follows:
[0077]
number
[0078] Here, [x, y] represents the index value of the channel, b represents the index value of the frequency band, and N b is the number of frequency points within the channel frequency band.
[0079] In one embodiment, based on the cross-correlation coefficients of any two channels in any one frequency band b, a first cross-correlation matrix corresponding to any one frequency band b can be obtained. For example, the channel sequence includes M channels, and the following first cross-correlation matrix can be obtained.
[0080]
Equation
[0081] Both the first row and the first column of the first cross-correlation matrix correspond to channel 1, both the second row and the second column correspond to channel 2, and thus, both the Mth row and the Mth column correspond to channel M.
[0082] Note that each frequency band in the frequency band set can determine one corresponding first cross-correlation matrix according to the above method. If the frequency band set includes b frequency bands, there are b first cross-correlation matrices.
[0083] In S204, from the first cross-correlation matrix of each frequency band, a second cross-correlation matrix of the channel group is determined respectively, and the second cross-correlation matrix includes the cross-correlation coefficients between the channels within the channel group.
[0084] In one embodiment, the M*M first cross-correlation matrix described above allows for the extraction of a second cross-correlation matrix corresponding to a channel group from the first cross-correlation matrix based on the channels within the channel group. In one embodiment, the channel index of the channels within the channel group is determined, and the second cross-correlation matrix for the channel group is extracted from the first cross-correlation matrix based on the channel index. In one embodiment, the channel group contains three channels, and the second cross-correlation matrix corresponding to the channel group can be determined from the first cross-correlation matrix based on the cross-correlation coefficient between any two channels included in the channel group, and this second cross-correlation matrix is a 3*3 matrix.
[0085] For example, a channel sequence contains five channels, arranged in the order of channel 1, channel 2, channel 3, channel 4, and channel 5. A channel group contains three channels, where channel group 1 may include channels 1, 2, and 3; channel group 2 may include channels 2, 3, and 4; and channel group 3 may include channels 3, 4, and 5. Here, the first cross-correlation matrix is a 5*5 matrix, as shown in Figure 3. For channel group 2, the matrix elements at the intersection of rows 2-4 and columns 2-4 can be extracted from the first cross-correlation matrix to obtain the second cross-correlation matrix for channel group 2. For example, the second cross-correlation matrix for channel group 2 may be the portion within the dotted frame in the first cross-correlation matrix as shown in Figure 3, i.e., the second cross-correlation matrix for channel group 2 is as shown below.
[0086]
number
[0087] Furthermore, each channel group has one second cross-correlation matrix in each frequency band.
[0088] In S205, the target transformation matrix for each channel group in each frequency band is determined from the transformation matrix set based on the second cross-correlation matrix of the channel group in each frequency band.
[0089] In one embodiment, for any one frequency band b, if the cross-correlation coefficient between any two channels included in the second cross-correlation matrix in any one frequency band b satisfies the condition for selecting a specified transformation matrix in the set of transformation matrices, then the specified transformation matrix is selected as the target transformation matrix for any one frequency band b.
[0090] If the cross-correlation coefficient between any two channels in the second cross-correlation matrix for any one frequency band b does not satisfy the condition, a transformation matrix other than the specified one is selected from the set of transformation matrices as the target transformation matrix for any one frequency band b, based on the maximum cross-correlation coefficient between any two channels.
[0091] In one embodiment, a cross-correlation coefficient threshold may be set, and based on the cross-correlation coefficient threshold The system then determines the target transformation matrix for a channel group from the set of transformation matrices. The second cross-correlation matrix for the channel group in each frequency band contains the cross-correlation coefficient between any two channels in the channel group. If the cross-correlation coefficient between any two channels is greater than a set threshold, the specified transformation matrix is selected as the target transformation matrix for any one of the frequency bands b. If the cross-correlation coefficients between any two channels are both less than or equal to the threshold, a transformation matrix other than the specified one is selected from the set of transformation matrices as the target transformation matrix for any one of the frequency bands b, based on the maximum cross-correlation coefficient between any two channels.
[0092] For example, the set of transformation matrices may include M0, M1, M2, M3, and M4. Here, the specified transformation matrix may be M4, and if the cross-correlation coefficients of all three channels are greater than the threshold, the target transformation matrix is determined to be M4, and if the maximum value of the cross-correlation coefficients in the three channels is greater than the threshold, the target transformation matrix is determined based on the maximum value.
[0093] In S206, based on the target transformation matrix for each frequency band, same-band decorrelation processing is performed on the frequency domain coefficients of the channels within the channel group to obtain the coding information of the channel group.
[0094] In the embodiments of this disclosure, the implementation of S206 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0095] In S207, the encoded stream is obtained based on the encoding information of the channel group, and the encoded stream is sent to the decoder for decoding.
[0096] In the embodiments of this disclosure, the implementation of S207 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0097] In embodiments of this disclosure, frequency domain coefficients in each frequency band within a channel group are encoded via a target transformation matrix, thereby enabling compression of multi-channel audio signals, reducing redundancy between channels, alleviating the encoder load, and lowering transmission and storage costs.
[0098] Figure 4 is a schematic flowchart of an audio encoding method provided by an embodiment of the present disclosure. The audio encoding method may be performed by an encoder. As shown in Figure 4, the method may include, but is not limited to, steps S401 to S408.
[0099] In S401, channel sequences are grouped to obtain multiple channel groups, each channel group containing multiple consecutive channels within the channel sequence, with one or more identical channels existing between adjacent channel groups.
[0100] In the embodiments of this disclosure, the implementation of S401 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0101] In S402, frequency domain conversion is performed on the audio signal of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients for each frame of each channel.
[0102] In the embodiments of this disclosure, the implementation of S402 is any one of the methods of each embodiment of this disclosure This can be achieved by [method / method], and is not limited to [this method], so a detailed explanation will be omitted.
[0103] In S403, the first cross-correlation matrix between channels corresponding to each frequency band is determined based on the frequency domain coefficients of each channel.
[0104] In the embodiments of this disclosure, the implementation of S403 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0105] In S404, the first cross-correlation matrix for each frequency band is normalized, and based on the channels included in the channel group, the second cross-correlation matrix for each frequency band corresponding to the channel group is extracted from the normalized first cross-correlation matrix for each frequency band.
[0106] In one embodiment, the first cross-correlation matrix for each frequency band between channels is normalized to obtain a normalized first cross-correlation matrix for each frequency band. From this normalized first cross-correlation matrix, a second cross-correlation matrix for each frequency band corresponding to a channel group is extracted. The second cross-correlation matrix is a normalized matrix.
[0107] In one embodiment, a channel identifier associated with any one matrix element in a first cross-correlation matrix is determined. In one embodiment, based on the associated channel identifier, a normalized matrix element corresponding to any one matrix element can be determined, where any one matrix element is the cross-correlation coefficient between two channels, and the channel identifier corresponding to the row in which any one matrix element is located may be one channel identifier associated with any one matrix element, or the channel identifier corresponding to the row in which any one matrix element is located may be another channel identifier associated with any one matrix element.
[0108] For example, any one matrix element Corr in the first cross-correlation matrix corresponding to frequency band b [2,3]、b Let's explain using this as an example, where any one matrix element Corr [2,3]、b The channel identifiers associated with it are 2 and 3, meaning that either one of the matrix elements Corr [2,3]、b The channels associated with it are channel 2 and channel 3.
[0109] Any one of the matrix elements Corr [2,3]、b If the associated channels are channel 2 and channel 3, then one of the matrix elements Corr [2,3]、b The normalized matrix elements are Corr [2,2]、b and Corr [3,3]、b It can be determined that this is the case.
[0110] In one embodiment, the normalization result of one of the matrix elements is obtained based on either one matrix element and the normalized matrix element. In one embodiment, the normalization formula corresponding to either one matrix element is as follows:
[0111]
number
[0112] Here, b represents the index value of the frequency band.
[0113]
number
[0114] Note that the cross-correlation coefficient matrix has equal values at symmetrical positions along the diagonal, so the lower left of the diagonal is... There is no need to calculate the cross-correlation coefficients at the corners; normalization can be performed using the element at the upper right corner of the diagonal of the cross-correlation coefficient matrix. When performing the above calculation, if the denominator is greater than 0, the normalization calculation is performed again; if the denominator is less than 0, the cross-correlation coefficient at the corresponding position is 0. Here, the diagonal is the diagonal from the upper left corner to the lower right corner.
[0115] In S405, the cross-correlation coefficient between any two channels in a channel group in any one frequency band b is determined based on the second cross-correlation matrix in any one frequency band b.
[0116] In one embodiment, the elements in the second cross-correlation matrix are the cross-correlation coefficients between any two channels in a channel group in any one frequency band b, so that the cross-correlation coefficients between any two channels in a channel group in any one frequency band b can be determined based on the second cross-correlation matrix.
[0117] In S406, the target transformation matrix for a channel group in any one frequency band b is determined from the set of transformation matrices based on the cross-correlation coefficient between any two channels within the channel group in any one frequency band b.
[0118] In one embodiment, a threshold value of Thr can be set for the cross-correlation coefficient between any two channels in a channel group within any one frequency band b, and the target transformation matrix is determined based on a preset condition by comparing the cross-correlation coefficient between any two channels in the channel group with the threshold value.
[0119] In one embodiment, if the cross-correlation coefficient between any two channels satisfies the condition for selecting a specified transformation matrix in the set of transformation matrices, that is, if the cross-correlation coefficients between any two channels in a channel group are both greater than Thr, then the specified transformation matrix M4 is selected as the target transformation matrix for any one of the frequency bands b.
[0120] For example, if [L,C,R] are three channels in a channel group, and the cross-correlation coefficients between channel L and channel C, between channel L and channel R, and between channel C and channel R are all greater than Thr, then the specified transformation matrix M4 is selected as the target transformation matrix for any one of the frequency bands b.
[0121] In one embodiment, if the cross-correlation coefficient between any two channels in any one frequency band b does not satisfy the condition, a transformation matrix other than the specified one is selected as the target transformation matrix for any one frequency band b based on the maximum cross-correlation coefficient between any two channels.
[0122] In other words, the maximum cross-correlation coefficient is selected from three cross-correlation coefficients: the cross-correlation coefficient between channel L and channel C, the cross-correlation coefficient between channel L and channel R, and the cross-correlation coefficient between channel C and channel R. If the maximum cross-correlation coefficient is greater than Thr, one target transformation matrix for frequency band b is selected from any of the transformation matrices other than the specified transformation matrix. For example, if the specified transformation matrix is M4, one target transformation matrix is selected from M0 to M3 based on the maximum cross-correlation coefficient.
[0123] In one embodiment, if the maximum cross-correlation coefficient is greater than Thr, the target transformation matrix is selected based on equation (6).
[0124]
number
[0125] For example, if the maximum cross-correlation coefficient between any two channels in a channel group within any one frequency band b is the cross-correlation coefficient between channel L and channel C, then M1 is selected as the target transformation matrix. If the maximum cross-correlation coefficient between any two channels in a channel group within any one frequency band b is the cross-correlation coefficient between channel L and channel R, then M2 is selected as the target transformation matrix. If the maximum cross-correlation coefficient between any two channels in a channel group within any one frequency band b is the cross-correlation coefficient between channel C and channel R, then M3 is selected as the target transformation matrix.
[0126] In S407, based on the target transformation matrix for each frequency band, same-band decorrelation processing is performed on the frequency domain coefficients of the channels within the channel group to obtain the coding information of the channel group.
[0127] In the embodiments of this disclosure, the implementation of S407 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0128] In S408, the encoded stream is obtained based on the encoding information of the channel group, and the encoded stream is sent to the decoder for decoding.
[0129] In the embodiments of this disclosure, the implementation of S408 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0130] In embodiments of this disclosure, frequency domain coefficients in each frequency band within a channel group are encoded via a target transformation matrix, thereby enabling compression of multi-channel audio signals, reducing redundancy between channels, alleviating the encoder load, and lowering transmission and storage costs.
[0131] Figure 5 is a schematic flowchart of an audio encoding method provided by an embodiment of the present disclosure. The audio encoding method may be performed by an encoder. As shown in Figure 5, the method may include, but is not limited to, steps S501 to S507.
[0132] In S501, channel sequences are grouped to obtain multiple channel groups, each channel group containing multiple consecutive channels within the channel sequence, and there are one or more overlapping channels between adjacent channel groups.
[0133] In the embodiments of this disclosure, the implementation of S501 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0134] In S502, frequency domain conversion is performed on the audio signal of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients for each frame of each channel.
[0135] In the embodiments of this disclosure, the implementation of S502 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0136] In S503, based on the frequency domain coefficients of each channel, the target transformation matrix for each frequency band in the frequency band set corresponding to the channel group is determined from the transformation matrix set. To determine.
[0137] In the embodiments of this disclosure, the implementation of S503 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0138] In S504, the frequency domain coefficients of the channels within the channel group in any one frequency band b are obtained, and the first encoding information of the channel group in any one frequency band b is obtained based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b.
[0139] In one embodiment, for a first channel group, the first channel group includes channel 1, channel 2, and channel 3, and first coded information for the first channel group in any one frequency band b is obtained based on the frequency domain coefficients of the channels in the first channel group in any one frequency band b, and based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b, and the first coded information includes center information M1, side information S1, and first information T1 of the first channel group.
[0140] Furthermore, for each of the remaining channel groups excluding the first channel group, based on the frequency domain coefficients of the channels in the remaining channel group in any one frequency band b, and based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b, the first encoded information of the remaining channel group in any one frequency band b is obtained, and the first encoded information is the first information T of the remaining channel group. i Includes.
[0141] Furthermore, the first encoded information for the first channel group may include all of the central information, side information, and first information, while the first encoded information for the remaining channel groups may include only the first information.
[0142] The formula for calculating channel coding information is as follows: [LCR]*M=[MST] (7)
[0143] Here, [L,C,R] are the three channels in the channel group, M is the target transformation matrix determined based on frequency domain coefficients, and [MST] is the encoded information of the channel group.
[0144] During the decorrelation calculation, the first decorrelation coding unit outputs all of the coded information [MST], while the remaining coding units output only the first information [T].
[0145] In S505, the second encoding information for a channel group is obtained based on the first encoding information for the channel group in each frequency band.
[0146] In one embodiment, the target transformation matrix of a channel group can be determined based on the first coding information of the channel group in each frequency band, and the decorrelation mode of the channel group in each frequency band can be determined based on the target transformation matrix, that is, the second coding information of the channel group can be determined.
[0147] Furthermore, if the target transformation matrices corresponding to different frequency bands are different, the decorrelations corresponding to those frequency bands will also be different. If the target transformation matrices corresponding to different frequency bands are the same, then the decorrelations corresponding to those frequency bands will also be the same.
[0148] In S506, the encoding information for the channel group is obtained based on the second encoding information and the target transformation matrix corresponding to each frequency band.
[0149] In one embodiment, a target transformation matrix corresponding to each frequency band is used to decorrelate the frequency domain coefficients of the channel group in each frequency band based on the second encoding information, thereby obtaining the encoding information for the channel group in each frequency band. After obtaining the encoding information corresponding to each frequency band, the encoding information for the channel group across the entire frequency band can be obtained based on the encoding information for each frequency band. Note that the encoding information for the channel group includes the encoding information for the channel group in all frequency bands.
[0150] In S507, the encoded stream is obtained based on the encoding information of the channel group, and the encoded stream is sent to the decoder for decoding.
[0151] In the embodiments of this disclosure, the implementation of S507 can be achieved by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0152] In embodiments of this disclosure, frequency domain coefficients in each frequency band within a channel group are encoded via a target transformation matrix, thereby enabling compression of multi-channel audio signals, reducing redundancy between channels, alleviating the encoder load, and lowering transmission and storage costs.
[0153] Figure 6 shows one possible encoding flowchart of an embodiment of the present disclosure. An MDCT transform is performed on the audio signal within each channel to obtain the MDCT coefficients (frequency domain coefficients) for each frame of each channel. The channel signals are input to a frequency band division processing unit to perform frequency band division and obtain the frequency domain coefficients in each frequency band. Furthermore, an energy calculation unit calculates the energy value of the channels in each frequency band, and the energy values are input to a cross-correlation calculation unit to obtain the cross-correlation coefficients between channels in each frequency band. This obtains the first cross-correlation matrix between channels in each frequency band. Each channel group corresponds to one decorrelation unit, and the frequency domain coefficients of the three channels within the channel group and the first cross-correlation matrix for each frequency band are input to the decorrelation unit. The decorrelation unit then performs the same frequency decorrelation process on the channel group to obtain the encoding information of the channel. As shown in Figure 6, the frequency band division results of the frequency domain coefficients of channels 1, 2, and 3 in the first channel group, and the first cross-correlation matrix for each frequency band are input to decorrelation unit 1. Decorrelation unit 1 performs identical frequency decorrelation processing and outputs the encoded information for channel group 1, which includes central information M1, side information S1, and first information T1. The frequency band division results of the frequency domain coefficients of channels 2, 3, and 4 in channel group 2, and the first cross-correlation matrix for each frequency band are input to decorrelation unit 2. Decorrelation unit 2 performs identical frequency decorrelation processing and outputs the encoded information for channel group 2, which includes first information T2. The frequency band division results of the frequency domain coefficients of channels 3, 4, and 5 in channel group 3, and the first cross-correlation matrix for each frequency band are input to decorrelation unit 3. Decorrelation unit 3 performs identical frequency decorrelation processing and outputs the encoded information for channel group 3, which includes T3.In this way, for the last channel group M-2, the frequency band division results of the frequency domain coefficients of channel M-2, channel M-1 and channel M within channel group M-2, and the first cross-correlation matrix of each frequency band are input to the decorrelation unit M-2, and the decorrelation unit M-2 performs the same frequency decorrelation process to output the encoded information of channel group M-2, which is T. M-2 Includes.
[0154] In embodiments of this disclosure, frequency domain coefficients in each frequency band within a channel group are encoded via a target transformation matrix, thereby enabling compression of multi-channel audio signals, reducing redundancy between channels, alleviating the encoder load, and lowering transmission and storage costs.
[0155] Figure 7 is a schematic flowchart of an audio decoding method provided by an embodiment of the present disclosure. The audio decoding method may be performed by a decoder. As shown in Figure 7, the method may include, but is not limited to, steps S701 to S704.
[0156] S701 receives the encoded stream transmitted from the encoder, and the encoded stream contains encoded information for multiple channel groups.
[0157] In the embodiments of this disclosure, a channel group is obtained by sequentially grouping channel sequences, and each channel group includes a plurality of consecutive channels in a channel sequence, with one or more identical channels between adjacent channel groups.
[0158] In embodiments of this disclosure, the decoder receives an encoded stream transmitted from the encoder, reads the encoded information of multiple channel groups from the encoded stream, and performs an inverse transform on the input channel signal to obtain the original channel signal.
[0159] According to the above embodiment, the encoder can group M channels in a channel sequence to obtain multiple channel groups. Each channel group contains three consecutive channels in the channel sequence, and one or more identical channels exist between adjacent channel groups. For details on the process by which the encoder encodes channel groups, please refer to the above embodiment, and a detailed explanation is omitted here.
[0160] In S702, multiple channel groups are decoded sequentially, and for the currently decoded channel group, the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group is determined based on the encoding information of the current channel group.
[0161] In one embodiment, second encoding information for channel groups across the entire frequency band can be obtained from the encoding information, and this second encoding information includes first encoding information for each frequency band and a target transformation matrix corresponding to each frequency band. For any one frequency band b in the frequency band set, the target decoding matrix for the current channel group in any one frequency band b is obtained by querying the correspondence between the transformation matrix and the decoding matrix based on the target transformation matrix of any one frequency band b.
[0162] Note that the transformation matrix and the decoding matrix correspond one-to-one, and in one embodiment, the decoding matrix may include the following matrix.
[0163]
number
[0164] For example, transformation matrix M0 corresponds to decoding matrix J0, transformation matrix M1 corresponds to decoding matrix J1, transformation matrix M2 corresponds to decoding matrix J2, and transformation matrix M3 corresponds to decoding matrix J3. The transformation matrix M4 corresponds to the decoding matrix J4. If the target transformation matrix for the current channel group is M4, then the target decoding matrix for the current channel group in any one frequency band b is J4.
[0165] In S703, the decoding frequency domain coefficients for the current channel group are obtained for the encoding information of the current channel group, based on the target decoding matrix of the current channel group in each frequency band.
[0166] In one embodiment, the current channel group encoding information includes a second encoding information for the channel group across the entire frequency band, the second encoding information includes a first encoding information for the channel group in each frequency band.
[0167] For any one frequency band b within the frequency band set, the first encoded information in that frequency band b is decoded based on the target decoding matrix corresponding to that frequency band b to obtain the first decoded frequency domain coefficients for the current channel group in that frequency band b. In one embodiment, the decoded frequency domain coefficients for the channel group across the entire frequency band can be obtained based on the first decoded frequency domain coefficients in each frequency band.
[0168] In S704, the decoded audio signal for each channel in the channel sequence is obtained based on the decoding frequency domain coefficients of multiple channel groups.
[0169] In one embodiment, a frequency-domain to time-domain conversion can be performed on the signals of a group of channels based on the decoding frequency-domain coefficients of the group. In one embodiment, using inverse MDCT conversion, the frequency-domain signals of the channels can be converted to time-domain signals based on the decoding frequency-domain coefficients, thereby obtaining the decoded audio signals of each channel.
[0170] In the audio decoding method provided by the embodiments of this disclosure, the decoder receives an encoded stream transmitted from the encoder and obtains encoded information for each channel group therefrom. The decoder obtains the decoded frequency domain coefficients for each channel group by decoding the encoded information via a decoding unit in the order of the multiple channel groups, and performs a frequency-domain-to-time-domain conversion on the decoded frequency domain coefficients for each channel group to obtain the decoded audio signal for each channel. During decoding, the decoder can use frequency division decoding similar to that on the encoder side to achieve the reconstruction of the multi-channel audio signal, and because the encoder has performed compression, the transmission of the multi-channel signal becomes easier and transmission space is saved.
[0171] Figure 8 is a schematic flowchart of an audio decoding method provided by an embodiment of the present disclosure. The audio decoding method may be performed by a decoder. As shown in Figure 8, the method may include, but is not limited to, steps S801 to S809.
[0172] In S801, the first encoding information for the first channel group in any one frequency band b is obtained from the encoding information of the first channel group.
[0173] In one embodiment, each channel group includes three consecutive channels, and the first channel group includes channel 1, channel 2, and channel 3. In one embodiment, the encoding information of the first channel group can be obtained from the encoded stream, and the encoding information of the first channel group includes the first encoding information of the first channel group in each frequency band, and the first code The numbering information includes at least the central information M1, side information S1, and first information T1 of the first channel group.
[0174] In S802, based on the target decoding matrix of the first channel group in each frequency band, the first encoded information of the first channel group in each frequency band is decoded to obtain the first decoded frequency domain coefficients of the first channel group in any one of the frequency bands b.
[0175] In S803, the second decoding frequency domain coefficient of the first channel group is obtained based on the first decoding frequency domain coefficient of the first channel group in each frequency band, and the second decoding frequency domain coefficient includes three outputs.
[0176] In one embodiment, the decoder can obtain the target transformation matrix for the first channel group in any one frequency band b from the encoded information corresponding to the first channel group. In one embodiment, the correspondence between the transformation matrix and the decoding matrix can be established in advance, and in one embodiment, the target decoding matrix for the first channel group in any one frequency band b can be determined based on the target transformation matrix in any one frequency band b.
[0177] Based on the target decoding matrix in any one frequency band b, an inverse transform is performed on the first encoded information in any one frequency band b to obtain the first decoding frequency domain coefficients for the first channel group in any one frequency band b.
[0178] In one embodiment, for any one frequency band b in the frequency band set, the first encoded information for any one frequency band b is decoded based on the target decoding matrix corresponding to any one frequency band b to obtain the first decoded frequency domain coefficients of the first channel group in any one frequency band b. The first encoded information of the first channel group includes M1, S1, and T1, and after decoding, the first decoded frequency domain coefficients for any one frequency band b include three outputs, and the first decoded frequency domain coefficients are,
[0179]
number
[0180] In one embodiment, based on the first decoding frequency domain coefficient in each frequency band, the decoding frequency domain coefficient for the first channel group across the entire frequency band can be obtained, and the decoding frequency domain coefficient for the first channel group across the entire frequency band is,
[0181]
number
[0182] In one embodiment, the decoding formula for the first channel group is as follows:
[0183]
number
[0184] Here, [M1S1T1] represents the second coded information of the first channel group, and J is the target decoding matrix.
[0185]
number
[0186] In S804, multiple decoded channels adjacent to and consecutive to the current channel group The channel group is determined to be the upmix channel group corresponding to the current channel group.
[0187] In one embodiment, if the current channel group is a channel group other than the first channel group among multiple channel groups, a decoding operation can be performed on multiple decoded channel groups adjacent to and contiguous with the current channel group as upmix channel groups corresponding to the current channel group. In one embodiment, the multiple decoded channel groups contain two channel groups.
[0188] For example, if the current channel group is channel group 2, the corresponding upmix channel group includes the decoded frequency domain coefficients of the last two outputs of channel group 1; if the current channel group is channel group 3, the corresponding upmix channel group includes the decoded frequency domain coefficient of the last output of channel group 1 and the decoded frequency domain coefficient output by channel group 2; and if the current channel group is channel group 4, the corresponding upmix channel group includes the decoded frequency domain coefficient output by channel group 2 and the decoded frequency domain coefficient output by channel group 3.
[0189] In S805, the first encoding information for the current channel group in any one frequency band b is obtained from the encoding information.
[0190] As can be seen from the encoding process on the encoder side, each channel group output from channel group 2 is a single output. The decoder can obtain the first encoding information of the current channel group in any one frequency band b from the encoding information of the current channel group, and this first encoding information is the first information T of the channel group in any one frequency band b. i That is the case.
[0191] S806 obtains the decoding frequency domain coefficients of the upmix channel group in each frequency band.
[0192] In one embodiment, the target transformation matrix for the current upmix channel group in each frequency band can be determined from the encoded information received from the decoder, and the target decoding matrix for the upmix channel group can be determined based on the target transformation matrix. The decoder can then perform an inverse transform on the target decoding matrix to obtain the decoding frequency-domain coefficients for the upmix channel group in each frequency band.
[0193] In S807, based on the target decoding matrix corresponding to one of the frequency bands b and the decoding frequency domain coefficients in one of the frequency bands b, the first encoded information in one of the frequency bands b is decoded, and the first decoding frequency domain coefficients in one of the frequency bands b are obtained.
[0194] In one embodiment, the decoder can obtain the target transformation matrix for the current channel group in each frequency band from the encoded information, and further determine the target decoding matrix for the current channel group in each frequency band based on the target transformation matrix.
[0195] For any one frequency band b within the frequency band set, any one frequency band Based on the target decoding matrix of region b, the first coded information T in any one of the frequency bands b i Then, the first decoded frequency domain coefficient of the upmix channel group in any one frequency band b is decoded to obtain the first decoded frequency domain coefficient of the current channel group in any one frequency band b.
[0196] Note that the first encoded information of the current channel group is the first information T. i The first decoded frequency domain coefficient in any one frequency band b after decoding includes one output, and the first decoded frequency domain coefficient is
[0197]
number
[0198] In S808, the decoding frequency domain coefficients of the current channel group are obtained based on the first decoding frequency domain coefficients of the current channel group in each frequency band, and the decoding frequency domain coefficients of the current channel group include one output.
[0199] In one embodiment, based on a first decoding frequency domain coefficient in each frequency band, the decoding frequency domain coefficient for the first channel group across the entire frequency band can be obtained, and the decoding frequency domain coefficient for the first channel group across the entire frequency band is,
[0200]
number
[0201] In one embodiment, the decoding formula for the current channel group i is as follows:
[0202]
number
[0203] Furthermore, since the current channel group uses a different encoding mode and the decoding unit outputs only one decoding frequency domain coefficient, T i Only the value of is required, and therefore the value of the decoding matrix J is also different. The specific values are as follows:
[0204]
number
[0205] In S809, the decoded audio signal for each channel in the channel sequence is obtained based on the decoding frequency domain coefficients of multiple channel groups.
[0206] In the embodiments of this disclosure, the implementation of S809 can be realized by any one of the embodiments of this disclosure, and is not limited thereto; therefore, a detailed explanation is omitted.
[0207] As shown in the decoding flowchart in Figure 9, decoding by the decoding side must be performed sequentially according to the decoding units, and upmix decoding unit 1 is the first upmix decoding unit, and the first coded information [M1S1T1] is input to the first upmix decoding unit, and the three output decoding frequency domain coefficients are,
[0208]
number
[0209] In the embodiments of this disclosure, the decoder can use frequency division decoding similar to that of the encoder during decoding to achieve the reconstruction of a multi-channel audio signal, and because the encoder has performed compression, the transmission of the multi-channel signal becomes easier and transmission space is saved.
[0210] Figure 10 is a schematic flowchart of an audio encoding method provided by an embodiment of the present disclosure. As shown in Figure 10, the method may include, but is not limited to, steps S1001 to S1015.
[0211] In S1001, the channel sequence is grouped to obtain multiple channel groups, each channel group containing multiple consecutive channels within the channel sequence, and there are one or more overlapping channels between adjacent channel groups.
[0212] In S1002, a frequency domain transformation is performed on the audio signal of each channel in the channel sequence, frame by frame, to obtain candidate frequency domain coefficients for each frame of each channel.
[0213] In S1003, a first cross-correlation matrix between channels corresponding to each frequency band is determined based on the frequency domain coefficients of each channel.
[0214] In S1004, the second cross-correlation matrix for each channel group is determined from the first cross-correlation matrix for each frequency band, and the second cross-correlation matrix includes the cross-correlation coefficients between channels within the channel group.
[0215] In S1005, the target transformation matrix for each channel group in each frequency band is determined from the transformation matrix set based on the second cross-correlation matrix of the channel group in each frequency band.
[0216] In S1006, the frequency domain coefficients of the channels within the channel group in any one frequency band b are obtained, and the first encoding information of the channel group in any one frequency band b is obtained based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b.
[0217] In S1007, the second encoding information for a channel group is obtained based on the first encoding information for the channel group in each frequency band.
[0218] In S1008, the encoding information for the channel group is obtained based on the second encoding information and the target transformation matrix corresponding to each frequency band.
[0219] In S1009, the encoded stream is obtained based on the encoding information of the channel group, and the encoded stream is sent to the decoder for decoding.
[0220] S1010 receives the encoded stream transmitted from the encoder.
[0221] In S1011, multiple channel groups are decoded sequentially, and then applied to the currently decoded channel group.
[0222] In S1012, the target transformation matrix for each frequency band is obtained from the encoded information.
[0223] In S1013, the mapping relationship between the transformation matrix and the decoding matrix is queried based on the target transformation matrix of any one of the frequency bands b, and the target decoding matrix for the current channel group in any one of the frequency bands b is obtained.
[0224] In S1014, target decoding of the current channel group in each frequency band. Based on the matrix, the decoding frequency domain coefficients for the current channel group are obtained for the encoding information of the current channel group.
[0225] In S1015, the decoded audio signal for each channel in the channel sequence is obtained based on the decoding frequency domain coefficients of multiple channel groups.
[0226] In the embodiments of this disclosure, frequency domain coefficients in each frequency band within a channel group are encoded via a target transformation matrix, thereby achieving compression of multi-channel audio signals, reducing redundancy between channels, lessening the encoder load, and lowering transmission and storage costs. During decoding, the decoder can use frequency division decoding similar to that used by the encoder to restore the multi-channel audio signal.
[0227] Figure 11 is a block diagram of an audio encoding device according to an exemplary embodiment. Referring to Figure 11, the audio encoding device 1100 of the embodiment of this disclosure includes a channel grouping module 1101, a frequency domain conversion module 1102, a matrix determination module 1103, an encoding module 1104, and a transmission module 1105.
[0228] The channel grouping module 1101 is configured to group channel sequences to obtain multiple channel groups, each channel group containing multiple consecutive channels in the channel sequence, with one or more identical channels between adjacent channel groups.
[0229] The frequency domain processing module 1102 is configured to perform a frequency domain transformation on each audio signal of each channel in the channel sequence, frame by frame, to obtain the frequency domain coefficients for each frame of each channel.
[0230] The matrix determination module 1103 is configured to determine the target transformation matrix for each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set, based on the frequency domain coefficients of each channel.
[0231] The encoding module 1104 is configured to obtain encoding information for a channel group by performing same-band decorrelation on the frequency domain coefficients of the channels within the channel group, based on the target transformation matrix for each frequency band.
[0232] The transmitting module 1105 is configured to acquire an encoded stream based on the encoding information of the channel group and to send the encoded stream to the decoder for decoding.
[0233] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to determine a first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel, to determine a second cross-correlation matrix for each channel group from the first cross-correlation matrix for each frequency band, the second cross-correlation matrix including the cross-correlation coefficients between channels within the channel group, and to determine a target transformation matrix for the channel group in each frequency band from a set of transformation matrices based on the second cross-correlation matrix for the channel group in each frequency band.
[0234] In one embodiment of the present disclosure, the encoding module 1104 further obtains the frequency domain coefficients of the channels in the channel group in any one frequency band b, and the frequency domain coefficients in any one frequency band b corresponding to any one frequency band b The system is configured to obtain first coding information for a channel group in any one frequency band b based on a target transformation matrix, obtain second coding information for a channel group based on the first coding information for the channel group in each frequency band, and obtain coding information for a channel group based on the second coding information and the target transformation matrix corresponding to each frequency band.
[0235] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to determine the cross-correlation coefficient between any two channels in a channel group in any one frequency band b based on a second cross-correlation matrix in any one frequency band b, and to determine the target transformation matrix for a channel group in any one frequency band b from a set of transformation matrices based on the cross-correlation coefficient between any two channels in a channel group in any one frequency band b.
[0236] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to select a specified transformation matrix as the target transformation matrix for any one frequency band b if the cross-correlation coefficient between any two channels satisfies the condition for selecting a specified transformation matrix in the transformation matrix set, and, if the cross-correlation coefficient between any two channels does not satisfy the condition, to select a transformation matrix other than the specified transformation matrix from the transformation matrix set as the target transformation matrix for any one frequency band b based on the maximum cross-correlation coefficient between any two channels.
[0237] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to determine the energy value of any one channel in each frequency band based on the frequency domain coefficient of any one channel, and to determine a first cross-correlation matrix corresponding to each frequency band based on the energy value of each channel in each frequency band.
[0238] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to obtain the energy ratio of any two channels in a channel sequence in any one frequency band b, and to determine that the cross-correlation coefficient between any two channels in any one frequency band b is zero if the energy ratio in any one frequency band b is less than or equal to a first setting threshold, or if the energy ratio in any one frequency band b is greater than or equal to a second setting threshold, where the first setting threshold is less than the second setting threshold and the energy ratio is between the first and second setting thresholds, and to determine the cross-correlation coefficient between any two channels in any one frequency band b based on the frequency domain coefficients of any two channels in any one frequency band b, and to obtain a first cross-correlation matrix corresponding to any one frequency band b based on the cross-correlation coefficient between any two channels in any one frequency band b.
[0239] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to perform a normalization process on the first cross-correlation matrix for each frequency band and to extract a second cross-correlation matrix for each frequency band corresponding to a channel group from the normalized first cross-correlation matrix for each frequency band based on the channels contained within the channel group.
[0240] In one embodiment of the present disclosure, the matrix determination module 1103 is further configured to determine a channel identifier associated with any one matrix element in a first cross-correlation matrix, to determine a normalized matrix element corresponding to any one matrix element based on the associated channel identifier, and to obtain a normalized result for any one matrix element based on any one matrix element and the normalized matrix element.
[0241] In one embodiment of the present disclosure, the encoding module 1104 further performs the first channeling For each of the remaining channel groups other than the first channel group, the system is configured to obtain first coded information for the remaining channel groups in any one frequency band b based on the frequency domain coefficients of the channels in the remaining channel group in any one frequency band b, and based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b, wherein the first coded information includes the first information for the remaining channel groups.
[0242] In one embodiment of the present disclosure, the channel grouping module 1101 is further configured to determine that adjacent channel groups include a first channel group and a second channel group, the first channel group and the second channel group each include three consecutive channels in the channel sequence, and the first channel group and the second channel group include two identical channels.
[0243] In an embodiment of the present disclosure, through the target conversion matrix, the frequency domain coefficients in each frequency band within the channel group are encoded, thereby realizing compression of the audio signals of multiple channels, reducing redundancy between multiple channels, reducing the burden on the encoder, and reducing the costs of transmission and storage.
[0244] FIG. 12 is a block diagram of an audio decoding apparatus according to an exemplary embodiment. Referring to FIG. 12, the audio decoding apparatus 1200 of an embodiment of the present disclosure includes a receiving module 1201, a matrix determination module 1202, and a decoding module 1203.
[0245] The receiving module 1201 is configured to receive an encoded stream transmitted from an encoder. The encoded stream includes encoded information of multiple channel groups. The channel groups are obtained by sequentially grouping a channel sequence. Each channel group includes a plurality of consecutive channels in the channel sequence, and there is one or more identical channels between adjacent channel groups.
[0246] The matrix determination module 1202 is configured to sequentially decode multiple channel groups, and for the currently decoded channel group, determine the target decoding matrix of each frequency band within the frequency band set corresponding to the current channel group based on the encoded information of the current channel group.
[0247] The decoding module 1203 obtains the decoding frequency domain coefficients of the current channel group for the encoded information of the current channel group based on the target decoding matrix of the current channel group in each frequency band, and is configured to obtain the decoded audio signals of each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups.
[0248] In one embodiment of the present disclosure, the matrix determination module 1202 further obtains the target transformation matrix of each frequency band from the encoded information, queries the mapping relationship between the transformation matrix and the decoding matrix based on the target transformation matrix of any one frequency band b, and is configured to obtain the target decoding matrix of the current channel group in any one frequency band b.
[0249] In one embodiment of the present disclosure, the decoding module 1203 further obtains the first encoded information of the first channel group in any one frequency band b from the encoded information of the first channel group, decodes the first encoded information of the first channel group in any one frequency band b based on the target decoding matrix of the first channel group in any one frequency band b to obtain the first decoding frequency domain coefficients of the first channel group in any one frequency band b, and is configured to obtain the first decoding frequency domain coefficients of the first channel group in each frequency band and the decoding frequency domain coefficients of the first channel group, where the first decoding frequency domain coefficients of the first channel group and the decoding frequency domain coefficients include three outputs.
[0250] In one embodiment of the present disclosure, the decoding module 1203 is further configured such that the first encoded information of the first channel group in any one frequency band b includes at least the center information, side information, and first information in any one frequency band b.
[0251] In one embodiment of the present disclosure, the decoding module 1203 further determines that a plurality of decoded channel groups adjacent to and contiguous with the current channel group are upmix channel groups corresponding to the current channel group, obtains first coded information for the current channel group in any one frequency band b from the coded information, obtains the decoded frequency domain coefficients for the upmix channel groups in each frequency band, decodes the first coded information in any one frequency band b based on the target decoding matrix corresponding to any one frequency band b and the decoded frequency domain coefficients in any one frequency band b, obtains first decoded frequency domain coefficients in any one frequency band b, obtains the decoded frequency domain coefficients for the current channel group based on the first decoded frequency domain coefficients for the current channel group in each frequency band, and is configured such that the decoded frequency domain coefficients for the current channel group include one output.
[0252] In one embodiment of the present disclosure, the decoding module 1203 is further configured to include first encoded information of the current channel group in any one frequency band b, the first information of the current channel group in any one frequency band b.
[0253] In the embodiments of this disclosure, the decoder can use frequency division decoding similar to that of the encoder during decoding to achieve the reconstruction of a multi-channel audio signal, and because the encoder has performed compression, the transmission of the multi-channel signal becomes easier and transmission space is saved.
[0254] Figure 13 is a schematic diagram of another audio processing device 1300 provided by an embodiment of the present disclosure. The audio processing device 1300 may be an encoder, a decoder, the encoder may be a chip, chip system, or processor that supports implementing the above method, or the decoder may be a chip, chip system, or processor that supports implementing the above method. The device may be used to implement the method described in the above embodiment of the method, for which refer to the description of the above embodiment of the method.
[0255] The audio processing unit 1300 may include one or more processors 1301. The processors 1301 may be general-purpose processors or dedicated processors, for example. They may be baseband processors or central processing units. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control audio processing units (e.g., base stations, baseband chips, decoders, decoder chips, DUs or CUs, etc.), execute computer programs, and process data from computer programs.
[0256] In some embodiments, the audio processing unit 1300 may include one or more memories 1302 in which a computer program 1304 is stored, and the processor 1301 executes the computer program 1304 to cause the audio processing unit 1300 to perform the method described in the above embodiment. In some embodiments, data may be stored in the memories 1302. The audio processing unit 1300 and the memories 1302 may be installed separately or integrated as a single unit.
[0257] In some embodiments, the audio processing unit 1300 may further include a transceiver 1305 and an antenna 1306. The transceiver 1305 may also be called a transceiver unit, transceiver, or transceiver circuit, and is used to implement a transceiver function. The transceiver 1305 may include a receiver and a transmitter, the receiver may also be called a receiving device or receiving circuit, and is used to implement a receiving function, and the transmitter may also be called a transmitting device or transmitting circuit, and is used to implement a transmitting function.
[0258] In some embodiments, the audio processing unit 1300 may further include one or more interface circuits 1307. The interface circuits 1307 receive code instructions and transmit them to the processor 1301. The processor 1301 executes the code instructions, causing the audio processing unit 1300 to perform the method described in the above embodiment of the method.
[0259] In one implementation, the processor 1301 may include a transceiver for implementing receiving and transmitting functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing receiving and transmitting functions may be separate or integrated as a single unit. The transceiver circuit, interface, or interface circuit may be used for reading and writing code / data, or the transceiver circuit, interface, or interface circuit may be used for transmitting or transmitting signals.
[0260] In one implementation, the processor 1301 may store a computer program 1303, which is executed in the processor 1301, thereby causing the audio processing unit 1300 to execute the method described in the above embodiment. The computer program 1303 may be fixed in the processor 1301, in which case the processor 1301 may be implemented by hardware.
[0261] In one embodiment, the audio processing device 1300 may include a circuit that can implement the transmission, reception, or communication functions described in the method embodiment described above. The processor and transceiver described herein can be integrated into an integrated circuit (IC), analog IC, high-frequency integrated circuit (RFIC), mixed-signal IC, application-specific integrated circuit (ASIC), printed circuit board (PCB), electronic device, etc. The processor and transceiver can be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), n-metal oxide semiconductor (NMOS), positive channel metal oxide semiconductor (PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (Gas), etc.
[0262] The audio processing device in the above description of embodiments may be an encoder or a decoder, but the scope of the audio processing device described in this disclosure is not limited to these, and the structure of the audio processing device is not limited to that shown in Figure 13. The audio processing device may be an independent device or part of a larger device. For example, the audio processing device may be as follows: (1) Independent integrated circuit IC, or chip, or chip system or subsystem, (2) A set having one or more ICs, in some embodiments the set of ICs may include a storage component for storing data, computer programs, (3) ASIC, for example, modem, (4) A module that can be incorporated into other devices, (5) Receivers, decoders, intelligent decoders, cellular phones, wireless devices, handhelds, mobile units, in-vehicle devices, encoders, cloud devices, artificial intelligence devices, etc., (6) Others.
[0263] Regarding the case where the audio processing device may be a chip or a chip system, refer to the schematic configuration diagram of the chip shown in FIG. 14. The chip shown in FIG. 14 includes a processor 1401 and an interface 1402. The number of processors 1401 may be one or more, and the number of interfaces 1402 may be more than one.
[0264] In some embodiments, the chip further includes a memory 1403 for storing the necessary computer programs and data.
[0265] In some implementations, the chip is for realizing the function of the decoder in the embodiments of the present disclosure.
[0266] In some implementations, the chip is for realizing the function of the encoder in the embodiments of the present disclosure.
[0267] As can be understood by those skilled in the art, the various illustrative logical blocks and steps listed in the embodiments of the present disclosure can be realized by electronic hardware, computer software, or a combination of both. Whether such functions are realized by hardware or by software depends on the specific application and the overall design requirements of the system. Those skilled in the art can use various methods to realize the above functions for each specific application, but such realizations should not be understood as exceeding the protection scope of the embodiments of the present disclosure.
[0268] Embodiments of the present disclosure further provide an audio processing system which includes an audio processing device as an encoder and an audio processing device as a decoder as described in the embodiment of Figure 13 described above, or an audio processing device as an encoder and an audio processing device as an encoder as described in the embodiment of Figure 14 described above.
[0269] This disclosure further provides a readable storage medium on which instructions are stored, which, when the instructions are executed by a computer, enables the functionality of an embodiment of any one of the methods described above.
[0270] Embodiments of the present disclosure further provide a computer program product, which, when executed by a computer, implements the functionality of any one of the above-described embodiment of the method.
[0271] Embodiments of the present disclosure further provide computer programs that, when executed on a computer, cause the computer to perform the functions of any one of the above-described embodiment of the method.
[0272] Furthermore, the above description of the methods and apparatus embodiments is also applicable to the electronic devices, computer-readable storage media, computer program products, and computer programs of the above embodiments, and a detailed explanation is omitted here.
[0273] In the embodiments described above, all or part of them can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of them can be implemented in the form of a computer program product. The computer program product includes one or more computer programs. When the computer programs are loaded and executed on a computer, all or part of the flows or functions described in the embodiments of this disclosure are generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer programs can be stored on a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer programs can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, radio, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device such as a server or data center containing one or more available media. The usable media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0274] As those skilled in the art will understand, the various numerical designations such as "First," "Second," etc., in this disclosure are classifications made for the sake of clarity and do not limit the scope of the embodiments of this disclosure, but rather represent priority.
[0275] In this disclosure, “at least one” may also be described as “one or more,” where “more” may be two, three, four or more, and is not limited to this disclosure. In embodiments of this disclosure, for a single technical feature, technical features in that type of technical feature are distinguished by “first,” “second,” “third,” “A,” “B,” “C,” and “D,” and there is no priority or size order among the technical features described by “first,” “second,” “third,” “A,” “B,” “C,” and “D.”
[0276] The correspondences shown in each table in this disclosure may be pre-configured or pre-defined. The possible values of the information in each table are merely examples and may be set to other values, and are not limited in this disclosure. When setting the correspondence between information and each parameter, it is not necessary to set all the correspondences shown in each table. For example, in the tables of this application, the correspondences shown in some rows may not be set. Also, appropriate modifications and adjustments such as splitting and joining may be made to the above tables. The names of the parameters shown in the themes of each table above may also be called by other names that are understandable to the communication device, and the possible values or representations of those parameters may also be other possible values or representations that are understandable to the communication device. When the above tables are implemented, other data structures such as arrays, queues, containers, stacks, linked lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables may be used.
[0277] In this disclosure, pre-definitions may be understood as definition, pre-definition, memory, pre-storage, pre-agreement, pre-setting, curing, or pre-firing.
[0278] As those skilled in the art will understand, the units and algorithmic steps described in each example disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or in software depends on the specific application of the proposed technology and the design constraints. Those skilled in the art may implement the described functions in different ways for each specific application, but such implementations should not be considered beyond the scope of this disclosure.
[0279] For the convenience and simplification of the explanation, and so that those skilled in the art can clearly understand, the specific working processes of the systems, apparatus, and units described above are omitted here, and should be referred to by the corresponding processes in the embodiments of the methods described above.
[0280] The above description is merely a specific embodiment of the present application, and the scope of protection of this application is not limited thereto. Any modification or substitution that a person skilled in the art could easily conceive of, as long as it does not deviate from the technical scope disclosed herein, should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be the same as the scope of the claims described above.
[0281] All embodiments of this disclosure may be performed independently or in combination with other embodiments, and both shall be considered to fall within the scope of protection claimed by this disclosure.
[0282] This disclosure claims priority to the Chinese Patent Application No. 202310403661.8, filed in China on April 14, 2023, and its entire contents are incorporated herein by reference.
Claims
1. An audio encoding method performed by an encoder, A step of grouping channel sequences to obtain multiple channel groups, wherein each channel group includes multiple consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups. The steps include: performing a frequency domain transformation on each audio signal of each channel in the channel sequence for each frame to obtain the frequency domain coefficients for each channel; A step of determining, from the set of transformation matrices, the target transformation matrix for each frequency band in the set of frequency bands corresponding to the channel group, based on the frequency domain coefficients of each channel, A step of obtaining coding information for the channel group by performing same-band decorrelation on the frequency domain coefficients of the channels in the channel group based on the target transformation matrix for each frequency band, The process includes the steps of obtaining an encoded stream based on the encoded information of the channel group and sending the encoded stream to a decoder for decoding, Audio encoding method.
2. The step of determining the target transformation matrix for each frequency band corresponding to the channel group from the set of transformation matrices based on the frequency domain coefficients of each channel is: The steps include determining a first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel, A step of determining a second cross-correlation matrix for each channel group from the first cross-correlation matrix for each frequency band, wherein the second cross-correlation matrix includes the cross-correlation coefficients between channels within the channel group. The step of determining a target transformation matrix for the channel group in each frequency band from the set of transformation matrices, based on the second cross-correlation matrix of the channel group in each frequency band, is included. The audio encoding method according to feature 1.
3. The step of obtaining coding information for the channel group by performing same-band decorrelation on the frequency domain coefficients of the channels in the channel group based on the target transformation matrix of each frequency band is as follows: The steps include: obtaining frequency domain coefficients of channels within the channel group in any one frequency band b, and obtaining first encoding information of the channel group in any one frequency band b based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b; A step of obtaining second encoding information for the channel group based on the first encoding information for the channel group in each frequency band, The step of obtaining the encoding information of the channel group based on the second encoding information and the target transformation matrix corresponding to each frequency band, The audio encoding method according to claim 1 or 2, characterized by the above.
4. The step of determining the target transformation matrix for the channel group in each frequency band from the set of transformation matrices based on the second cross-correlation matrix of the channel group in each frequency band is: Based on the second cross-correlation matrix in any one of the frequency bands b, any one The steps include determining the cross-correlation coefficient between any two channels in the channel group in the frequency band, The process includes the step of determining the target transformation matrix for the channel group in any one of the frequency bands b from the set of transformation matrices, based on the cross-correlation coefficient between any two channels in the channel group in any one of the frequency bands b, The audio encoding method according to feature 2.
5. The step of determining the target transformation matrix for the channel group in any one frequency band from the set of transformation matrices based on the cross-correlation coefficient between any two channels in the channel group in any one frequency band b is: If the cross-correlation coefficient between any two channels satisfies the condition for selecting a specified transformation matrix in the set of transformation matrices, the step of selecting the specified transformation matrix as the target transformation matrix for any one of the frequency bands b, If the cross-correlation coefficient between any two channels does not satisfy the condition, the step of selecting a transformation matrix other than the specified transformation matrix from the set of transformation matrices as the target transformation matrix for any one of the frequency bands b, based on the maximum cross-correlation coefficient between the two channels, is included. The audio encoding method according to feature 4.
6. The step of determining a first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel is as follows: A step of determining the energy value of any one channel in each frequency band based on the frequency domain coefficient of any one channel, The process includes the step of determining the first cross-correlation matrix corresponding to each frequency band based on the energy value of each channel in each frequency band. The audio encoding method according to feature 2.
7. The step of determining the first cross-correlation matrix corresponding to each frequency band based on the energy values of each channel in each frequency band is: A step of obtaining the energy ratio of any two channels in the channel sequence in any one frequency band b, If the energy ratio in any one of the aforementioned frequency bands b is less than or equal to a first setting threshold, or if the energy ratio in any one of the aforementioned frequency bands b is greater than or equal to a second setting threshold, a step of determining that the cross-correlation coefficient between any two channels in any one of the aforementioned frequency bands b is zero, wherein the first setting threshold is less than the second setting threshold, If the energy ratio in any one of the frequency bands b is between the first setting threshold and the second setting threshold, the step of determining the cross-correlation coefficient of any two channels in any one of the frequency bands b based on the frequency domain coefficients of any two channels in any one of the frequency bands b, The steps include obtaining the first cross-correlation matrix corresponding to the one frequency band b based on the cross-correlation coefficients of any two channels in any one frequency band b, The audio encoding method according to feature 6.
8. The step of determining the second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band is as follows: The process further includes performing a normalization operation on the first cross-correlation matrix for each frequency band, and extracting the second cross-correlation matrix for each frequency band corresponding to the channel group from the normalized first cross-correlation matrix for each frequency band based on the channels included in the channel group. The audio encoding method according to feature 2.
9. The step of performing normalization on the first cross-correlation matrix is: The steps include determining a channel identifier associated with any one matrix element in the first cross-correlation matrix, The steps include determining a normalized matrix element corresponding to any one of the matrix elements based on the associated channel identifier, The method includes the step of obtaining the normalization result of any one of the matrix elements based on any one of the matrix elements and the normalized matrix element, The audio encoding method according to feature 7.
10. A step of obtaining first encoded information for the first channel group in any one frequency band b, based on the frequency domain coefficients of the channels in the first channel group in any one frequency band b, and based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b, wherein the first encoded information includes center information, side information and first information of the first channel group. A step of obtaining first encoded information for each of the remaining channel groups excluding the first channel group, based on the frequency domain coefficients of the channels in the remaining channel group in any one frequency band b, and based on the frequency domain coefficients in any one frequency band b and the target transformation matrix corresponding to any one frequency band b, further comprising the step of the first encoded information including first information for the remaining channel groups, The audio encoding method according to any one of claims 3 to 9, characterized by the features described above.
11. Between the adjacent channel groups, there are one or more overlapping channels, and the method is as follows: A step of determining that the adjacent channel group includes a first channel group and a second channel group, wherein the first channel group and the second channel group each include three consecutive channels in the channel sequence, and the first channel group and the second channel group each include two identical channels. The audio encoding method according to any one of claims 1 to 9.
12. An audio decoding method performed by a decoder, A step of receiving an encoded stream transmitted from an encoder, wherein the encoded stream includes encoded information for a plurality of channel groups, the channel groups are obtained by sequentially grouping channel sequences, each channel group includes a plurality of consecutive channels in the channel sequence, and there is one or more identical channels between adjacent channel groups. The steps include sequentially decoding the plurality of channel groups, and determining the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group based on the encoding information of the current channel group, Based on the target decoding matrix of the current channel group in each frequency band, the current channel group The steps include obtaining the decoding frequency domain coefficients of the P, The step of obtaining the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups, An audio decoding method characterized by the following features.
13. The step of determining the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group, based on the encoding information of the current channel group, is: The steps include obtaining a target transformation matrix for each frequency band from the aforementioned encoding information, The process includes the steps of querying the mapping relationship between the transformation matrix and the decoding matrix based on the target transformation matrix of any one frequency band b, and obtaining the target decoding matrix of the current channel group in any one frequency band b, The audio decoding method according to claim 12.
14. If the current channel is the first channel group among the plurality of channel groups, the step of obtaining the decoding frequency domain coefficients of the current channel group for the encoding information of the current channel group based on the target decoding matrix of the current channel group in each frequency band is: The steps include obtaining first coding information for the first channel group in any one of the frequency bands b from the coding information of the first channel group, A step of decoding the first encoded information of the first channel group in any one of the frequency bands b based on the target decoding matrix of the first channel group in any one of the frequency bands b, and obtaining the first decoded frequency domain coefficients of the first channel group in any one of the frequency bands b, A step of obtaining the decoded frequency domain coefficient of the first channel group based on the first decoded frequency domain coefficient of the first channel group in each frequency band, the first decoded frequency domain coefficient of the first channel group and the decoded frequency domain coefficient having three outputs, The audio decoding method according to claim 12 or 13, characterized by the features described herein.
15. The first encoded information of the first channel group in any one of the aforementioned frequency bands b includes at least central information, side information, and first information in any one of the aforementioned frequency bands. The audio decoding method according to feature 14.
16. If the current channel is a channel group excluding the first channel group among the plurality of channel groups, the step of obtaining the decoding frequency domain coefficients of the current channel group for the encoding information of the current channel group based on the target decoding matrix of the current channel group in each frequency band is: The steps include determining that a plurality of decoded channel groups adjacent to and contiguous with the current channel group are upmix channel groups corresponding to the current channel group, The steps include obtaining first encoding information for the current channel group in any one frequency band b from the encoding information, The steps include obtaining the decoding frequency domain coefficients of the upmix channel group in each frequency band, Based on the target decoding matrix corresponding to any one of the frequency bands b and the decoding frequency domain coefficients in any one of the frequency bands b, the one of the frequency The steps include decoding the first encoded information in band b to obtain a first decoded frequency domain coefficient in any one of the frequency bands b, A step of obtaining the decoding frequency domain coefficients of the current channel group based on the first decoding frequency domain coefficients in each frequency band of the current channel group, the step of the decoding frequency domain coefficients of the current channel group having one output, The audio decoding method according to claim 12 or 13, characterized by the features described herein.
17. The first encoded information of the current channel group in any one frequency band b includes the first information of the current channel group in any one frequency band b. The audio decoding method according to feature 16.
18. An audio encoding device, A channel grouping module configured to group channel sequences and obtain multiple channel groups, wherein each channel group includes multiple consecutive channels in the channel sequence, and one or more identical channels exist between adjacent channel groups. A frequency domain processing module is configured to perform a frequency domain transformation on each audio signal of each channel in the channel sequence, frame by frame, to obtain the frequency domain coefficients of each frame of each channel. A matrix determination module configured to determine, from a set of transformation matrices, the target transformation matrix for each frequency band in a set of frequency bands corresponding to the channel group, based on the frequency domain coefficients of each channel, An encoding module configured to obtain encoding information for a channel group by performing same-band decorrelation on the frequency domain coefficients of channels within a channel group based on the target transformation matrix for each frequency band, A transmitting module is configured to acquire an encoded stream based on the encoded information of the channel group and send it to a decoder to decode the encoded stream, An audio encoding device characterized by the following features.
19. An audio decoding device, A receiving module configured to receive an encoded stream transmitted from an encoder, wherein the encoded stream includes encoded information for a plurality of channel groups, the channel groups are obtained by sequentially grouping channel sequences, each channel group includes a plurality of consecutive channels in the channel sequence, and one or more identical channels exist between adjacent channel groups. A matrix determination module configured to sequentially decode the plurality of channel groups and, for the decoded current channel group, determine the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group based on the encoding information of the current channel group, A decoding module is configured to obtain decoding frequency domain coefficients for the current channel group based on the target decoding matrix of the current channel group in each frequency band, and to obtain the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups. An audio decoding device characterized by the following features.
20. It is an encoder, Processor and Includes memory for storing instructions that can be executed by the processor, The processor is configured to perform the steps of the method according to any one of claims 1 to 11. An encoder characterized by the following features.
21. It is a decoder, Processor and Includes memory for storing instructions that can be executed by the processor, The processor is configured to implement the steps of the method according to any one of claims 12 to 17. A decoder characterized by the following features.
22. A computer-readable storage medium storing computer program instructions, wherein when the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 11 are realized. A computer-readable storage medium characterized by the following features.
23. A computer-readable storage medium storing computer program instructions, wherein when the program instructions are executed by a processor, the steps of the method according to any one of claims 12 to 17 are realized. A computer-readable storage medium characterized by the following features.
24. A computer program product including a computer program, which, when executed on a computer, causes the computer to execute the audio encoding method described in any one of claims 1 to 11. A computer program product characterized by the following features.
25. A computer program product including a computer program, which, when executed on a computer, causes the computer to execute the audio decoding method described in any one of claims 12 to 17. A computer program product characterized by the following features.
26. A computer program that, when executed on a computer, causes the computer to perform the audio encoding method described in any one of claims 1 to 11. A computer program characterized by the following features.
27. A computer program that, when executed on a computer, causes the computer to execute the audio decoding method described in any one of claims 12 to 17. A computer program characterized by the following features.