An audio encoding method, apparatus, electronic device and storage medium

CN116434760BActive Publication Date: 2026-08-14BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]本申请提供一种音频编码方法、装置、电子设备、计算机可读存储介质及计算机程序产品,以解决在多声道音频传输过程中,对传输和存储介质造成浪费等问题

Benefits of technology

[0023] The technical solution provided by the embodiments of this application brings at least the following beneficial effects: The encoder obtains frequency domain coefficients by dividing and grouping the channel signals into frequency bands. Based on the frequency domain coefficients of each channel, the target transform matrix of each frequency band corresponding to the channel group can be determined. Further, decorrelation processing is performed on the frequency domain coefficients of the channels according to the target transform matrix to obtain the encoding information of the channel group. Based on the encoding information, the encoded bitstream is obtained and then sent to the decoder for decoding. In the embodiments of this application, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transform matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, reduces the burden on the encoder, and lowers transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434760B_ABST
    Figure CN116434760B_ABST
Patent Text Reader

Abstract

This application relates to an audio encoding method, apparatus, electronic device, and storage medium, belonging to the field of audio processing technology. The method includes: grouping a channel sequence to obtain multiple channel groups, each channel group comprising several consecutive channels in the channel sequence, with one or more identical channels existing between adjacent channel groups; performing frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel per frame; determining the target transformation matrix for each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set based on the frequency domain coefficients of each channel; performing same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group based on the target transformation matrix of each frequency band to obtain the encoding information of the channel group; obtaining the encoded bitstream based on the encoding information of the channel group, and sending the encoded bitstream to the decoder for decoding. Therefore, this solution can achieve compressed transmission of audio signals from multiple channels, reducing transmission and storage costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio encoding method, apparatus, electronic device and storage medium. Background Technology

[0002] With the development of multimedia technology, the requirements for audio signals are getting higher and higher. Although the existing two-dimensional audio algorithm (2D mid-side, 2D M / S) can effectively reduce data redundancy between multiple channels, it increases the transmission cost and wastes data in the process of transmitting data in various different audio scenarios. Summary of the Invention

[0003] This application provides an audio encoding method, apparatus, electronic device, computer-readable storage medium, and computer program product to solve the problem of waste of transmission and storage media during multi-channel audio transmission. The technical solution of this application is as follows:

[0004] In a first aspect, embodiments of this application provide an audio encoding method executed by an encoder. The method includes: grouping a channel sequence to obtain multiple channel groups, each channel group including several consecutive channels in the channel sequence, with one or more identical channels existing between adjacent channel groups; performing frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel per frame; determining the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set based on the frequency domain coefficients of each channel; performing same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group based on the target transformation matrix of each frequency band to obtain the encoding information of the channel group; obtaining an encoded bitstream based on the encoding information of the channel group, and sending the encoded bitstream to a decoder for decoding.

[0005] Secondly, embodiments of this application provide an audio decoding method executed by a decoder. The method includes: receiving an encoded bitstream sent by an encoder, the encoded bitstream including encoding information of multiple channel groups, the channel groups being obtained by sequentially grouping a channel sequence, each channel group including several consecutive channels in the channel sequence, and one or more identical channels existing between adjacent channel groups; sequentially decoding the multiple channel groups; for the decoded current channel group, determining the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group based on the encoding information of the current channel group; obtaining the decoding frequency domain coefficients of the current channel group based on the target decoding matrix of the current channel group in each frequency band; and obtaining the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of the multiple channel groups.

[0006] Thirdly, embodiments of this application provide an audio encoding apparatus, comprising: a channel grouping module configured to group the channel sequence to obtain multiple channel groups, each channel group including a plurality of consecutive channels in the channel sequence, and one or more identical channels existing between adjacent channel groups; a frequency domain processing module configured to perform frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain frequency domain coefficients of each channel per frame; a matrix determination module configured to determine the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set based on the frequency domain coefficients of each channel; an encoding module configured to perform same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group based on the target transformation matrix of each frequency band to obtain the encoding information of the channel group; and a transmission module configured to obtain an encoded bitstream based on the encoding information of the channel group and send the encoded bitstream to a decoder for decoding.

[0007] Fourthly, embodiments of this application provide an audio decoding apparatus, including a receiving module configured to receive an encoded bitstream sent by an encoder, the encoded bitstream including encoding information of multiple channel groups, the channel groups being obtained by sequentially grouping channel sequences, each channel group including a plurality of consecutive channels in the channel sequence, and adjacent channel groups having one or more identical channels; a matrix determination module configured to sequentially decode the plurality of channel groups, and for the decoded current channel group, determine the target decoding matrix for each frequency band in the frequency band set corresponding to the current channel group based on the encoding information of the current channel group; and a decoding module configured to, based on the target decoding matrix of the current channel group in each frequency band, obtain the decoding frequency domain coefficients of the current channel group from the encoding information of the current channel group, and obtain the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of the plurality of channel groups.

[0008] Fifthly, embodiments of this application provide an encoder, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the method described in the first aspect of embodiments of this application.

[0009] In a sixth aspect, embodiments of this application provide a decoder, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the method described in the second aspect of embodiments of this application.

[0010] In a seventh aspect, embodiments of this application provide a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of the method described in the first aspect of embodiments of this application.

[0011] Eighthly, embodiments of this application provide a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of the method described in the second aspect of embodiments of this application.

[0012] Ninthly, embodiments of this application provide an encoder, the device including a processor and an interface circuit, the interface circuit being used to receive code instructions and transmit them to the processor, the processor being used to execute the code instructions to cause the device to perform the method described in the first aspect above.

[0013] In a tenth aspect, embodiments of this application provide a decoder. The device includes a processor and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processor. The processor is used to execute the code instructions to cause the device to perform the method described in the second aspect above.

[0014] Eleventhly, embodiments of this application provide an encoding / decoding system, which includes the encoding device described in the third aspect and the decoding device described in the fourth aspect; or, the system includes the encoder described in the fifth aspect and the decoder described in the sixth aspect; or, the system includes the encoder described in the seventh aspect and the encoding device described in the eighth aspect; or, the system includes the encoder described in the ninth aspect and the decoder described in the tenth aspect.

[0015] In a twelfth aspect, embodiments of the present invention provide a computer-readable storage medium for storing instructions for use by the encoder described above, which, when executed, cause the encoder to perform the method described in the first aspect.

[0016] In a thirteenth aspect, embodiments of the present invention provide a readable storage medium for storing instructions for use by the decoder described above, which, when executed, cause the decoder to perform the method described in the second aspect above.

[0017] In a fourteenth aspect, this application also provides a computer program product including a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect above.

[0018] In a fifteenth aspect, this application also provides a computer program product including a computer program, which, when run on a computer, causes the computer to perform the method described in the second aspect above.

[0019] In a sixteenth aspect, this application provides a chip system including at least one processor and an interface for supporting a network device in implementing the functions involved in the first aspect, such as determining or processing at least one of the data and information involved in the above methods. In one possible design, the chip system further includes a memory for storing computer programs and data necessary for the network device. The chip system may be composed of chips or may include chips and other discrete devices.

[0020] In a seventeenth aspect, this application provides a chip system including at least one processor and an interface for supporting a terminal device in implementing the functions involved in the second aspect, such as determining or processing at least one of the data and information involved in the above methods. In one possible design, the chip system further includes a memory for storing computer programs and data necessary for the terminal device. The chip system may be composed of chips or may include chips and other discrete devices.

[0021] In an eighteenth aspect, this application provides a computer program that, when run on a computer, causes the computer to perform the method described in the first aspect above.

[0022] In a nineteenth aspect, this application provides a computer program that, when run on a computer, causes the computer to perform the method described in the second aspect above.

[0023] The technical solution provided by the embodiments of this application brings at least the following beneficial effects: The encoder obtains frequency domain coefficients by dividing and grouping the channel signals into frequency bands. Based on the frequency domain coefficients of each channel, the target transform matrix of each frequency band corresponding to the channel group can be determined. Further, decorrelation processing is performed on the frequency domain coefficients of the channels according to the target transform matrix to obtain the encoding information of the channel group. Based on the encoding information, the encoded bitstream is obtained and then sent to the decoder for decoding. In the embodiments of this application, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transform matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, reduces the burden on the encoder, and lowers transmission and storage costs.

[0024] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of the application.

[0026] Figure 1 This is a flowchart illustrating an audio encoding method according to an exemplary embodiment.

[0027] Figure 2 This is a flowchart illustrating an audio encoding method according to another exemplary embodiment.

[0028] Figure 3 This is a schematic diagram of a cross-correlation matrix according to an exemplary embodiment.

[0029] Figure 4 This is a flowchart illustrating an audio encoding method according to another exemplary embodiment.

[0030] Figure 5 This is a flowchart illustrating an audio encoding method according to another exemplary embodiment.

[0031] Figure 6 This is a schematic diagram illustrating an audio encoding method according to an exemplary embodiment.

[0032] Figure 7 This is a flowchart illustrating an audio decoding method according to an exemplary embodiment.

[0033] Figure 8 This is a flowchart illustrating an audio decoding method according to another exemplary embodiment.

[0034] Figure 9 This is a schematic diagram illustrating an audio decoding method according to an exemplary embodiment.

[0035] Figure 10 This is a flowchart illustrating an audio encoding method according to another exemplary embodiment.

[0036] Figure 11 This is a block diagram illustrating an audio encoding apparatus according to an exemplary embodiment.

[0037] Figure 12 This is a block diagram illustrating an audio decoding device according to an exemplary embodiment.

[0038] Figure 13 This is a block diagram illustrating an audio processing apparatus according to an exemplary embodiment.

[0039] Figure 14 This is a block diagram illustrating another audio processing chip according to an exemplary embodiment. Detailed Implementation

[0040] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0041] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. The singular forms “a” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0043] Depending on the context, the word "if" as used herein can be interpreted as "when," "when," or "in response to determination." For the purposes of brevity and ease of understanding, this document uses the terms "greater than" or "less than," "higher than," or "lower than" to characterize size relationships. However, it will be understood by those skilled in the art that the term "greater than" also encompasses the meaning of "greater than or equal to," and "less than" also encompasses the meaning of "less than or equal to"; the term "higher than" also encompasses the meaning of "higher than or equal to," and "lower than" also encompasses the meaning of "lower than or equal to."

[0044] The audio encoding / decoding method disclosed in this application is applicable to various communication systems, such as: 3rd Generation (3G) Universal Mobile Telecommunications System (UMTS) Long Term Evolution (LTE) system, 5th Generation (5G) mobile communication system, 5G New Radio (NR) system, 6th Generation (6G) mobile communication system, or other future new mobile communication systems. The audio encoding / decoding method disclosed in this application is also applicable to streaming media transmission systems and OTT (Over-The-Top) media transmission systems.

[0045] Figure 1 This is a flowchart illustrating an audio encoding method provided in an embodiment of this application. This audio encoding method can be executed by an encoder. Figure 1 As shown, the method may include, but is not limited to, the following steps:

[0046] S101, group the vocal tract sequence to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence. There are one or more identical vocal tracts between adjacent vocal tract groups.

[0047] Optionally, the encoder can group the M channels in the channel sequence to obtain multiple channel groups. In one embodiment, each channel group contains several consecutive channels in the channel sequence, such as three consecutive channels. In this embodiment, there are one or more identical channels between adjacent channel groups. It can be understood that adjacent channel groups are the first channel group and the second channel group. The first channel group and the second channel group each include three consecutive channels in the channel sequence, and the first channel group and the second channel group include two identical channels. For example, the five channels in the channel sequence are divided into channel group 1, channel group 2, and channel group 3. Channel group 1 includes channel 1, channel 2, and channel 3; channel group 2 includes channel 2, channel 3, and channel 4; and channel group 3 includes channel 3, channel 4, and channel 5.

[0048] S102, perform frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame.

[0049] In this embodiment, the encoder divides the audio signal of each channel in the channel sequence into multiple fixed-length frames, and performs a modified discrete cosine transform (MDCT) on each frame to obtain the frequency domain representation of each frame. Based on the frequency domain representation of each frame, the MDCT coefficients of each frame can be extracted from the frequency domain representation as the frequency domain coefficients of each frame.

[0050] In one embodiment, the channel sequence of this application includes M channels.

[0051] In one implementation, each frame of audio data for each channel may include 2N sampling points, with a sampling rate of f. s Each frame after MDCT transformation can include N frequency points, and correspondingly, the spectral distribution range of the MDCT coefficients is (0, f). s / 2), frequency resolution is f s / 2N.

[0052] S103, based on the frequency domain coefficients of each channel, determine the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set.

[0053] In one implementation, the frequency bands can be pre-divided according to a psychoacoustic frequency banding method to obtain a frequency band set. The frequency band set may include multiple divided frequency bands; for example, the frequency band set may include b divided frequency bands, where b is an integer greater than or equal to 1. Each frequency band in the frequency band set has a different frequency range, and the frequency ranges of adjacent frequency bands are continuous.

[0054] In one implementation, after performing MDCT transformation on each channel for each frame, the sampling points of each frequency domain coefficient are sequentially multiplied by the frequency resolution to determine the frequency value of each frequency domain coefficient.

[0055] In one implementation, the formula for calculating the frequency value of any frequency domain coefficient is as follows:

[0056] f = n * f s / 2N (1)

[0057] Where n represents the nth sampling point corresponding to any frequency domain coefficient, and the value of n ranges from 1 to N, where N is the number of sampling points. The frequency value f corresponding to the frequency domain coefficient can be obtained by formula (1).

[0058] Furthermore, the frequency values ​​of the frequency domain coefficients are compared with the frequency range of each frequency band to obtain the frequency range in which the frequency domain coefficients are located, thereby determining the frequency domain coefficients in each frequency band.

[0059] In one implementation, the cross-correlation coefficients between different channels in the same frequency band are calculated based on the frequency domain coefficients between the channels in the same frequency band. That is, for each frequency band in the frequency band set, the cross-correlation coefficients between each pair of channels in that frequency band can be obtained based on the frequency domain coefficients between the channels in the same frequency band.

[0060] Furthermore, for any frequency band b in the frequency band set, the cross-correlation coefficients between any two channels within the channel group in frequency band b can be determined from the cross-correlation coefficients between any two channels corresponding to frequency band b. Further, based on the cross-correlation coefficients between any two channels within the channel group in frequency band b, the target transformation matrix corresponding to the channel group in frequency band b can be determined from the transformation matrix set. It can be understood that if the frequency band set includes B frequency bands, the target transformation matrix of the channel group in each frequency band can be obtained.

[0061] S104: Based on the target transformation matrix of each frequency band, perform same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group to obtain the coding information of the channel group.

[0062] In one implementation, the transformation matrix set includes multiple transformation matrices, each of which can correspond to a decorrelation mode. Optionally, the transformation matrix set may include M0, M1, M2, M3, and M4. The specific values ​​of the transformation matrices are as follows:

[0063]

[0064]

[0065]

[0066] In one implementation, the decorrelation mode of a channel group in each frequency band can be determined using the target transform matrix of the channel group in each frequency band. When the target transform matrices for different frequency bands are different, the decorrelation modes for those frequency bands are also different. When the target transform matrices for different frequency bands are the same, the decorrelation modes for those frequency bands are also the same. For example, if the target transform matrix for frequency band 1 is M1, the target transform matrix for frequency band 2 is M2, and the target transform matrix for frequency band 3 is M1, then the decorrelation modes for frequency band 1 and frequency band 3 are the same, but the decorrelation modes for frequency bands 1 and 3 are different from those for frequency band 2.

[0067] In one implementation, the frequency domain coefficients of the channels within the channel group are divided into frequency bands to obtain the frequency domain coefficients in each frequency band. Based on the target transformation matrix corresponding to the same frequency band, the frequency domain coefficients in the same frequency band of the channels within the channel group are decorrelated to obtain the channel group coding information.

[0068] For example, a channel group can be defined as three consecutive channels, labeled as channel L, channel C, and channel R. For any frequency band b in the frequency band set, the target transformation matrix of the channel group in frequency band b can be determined as M4. A frequency domain coefficient matrix can be formed from the frequency domain coefficients of channels L, C, and R in frequency band b. A matrix operation is then performed between this frequency domain coefficient matrix and the target transformation matrix M4 corresponding to frequency band b to obtain the first coding information of the channel group in frequency band b. In other words, decorrelation processing is performed on the frequency domain coefficients of channels L, C, and R in frequency band b using the target transformation matrix M4 corresponding to frequency band b to obtain the coding information of the channel group in frequency band b. Co-channel decorrelation processing can be performed on channels L, C, and R within the channel group to reduce co-channel interference, redundancy, and transmission costs.

[0069] It is understandable that the frequency domain coefficients of the channel group in each frequency band can be decorrelated using the target transformation matrix corresponding to each frequency band to obtain the coding information of the channel group in each frequency band. After obtaining the coding information corresponding to each frequency band, the coding information of the channel group in the entire frequency band can be obtained based on the coding information in each frequency band. It is understandable that the coding information of the channel group includes the coding information of that channel group in all frequency bands.

[0070] S105 obtains the encoded bitstream based on the encoding information of the channel group and sends the encoded bitstream to the decoder for decoding.

[0071] In one implementation, the encoding information of the channel groups can be encoded in binary to obtain a binary encoded bitstream. That is, the encoding information of each channel is converted into binary code, and the binary codes of all channel groups are concatenated to form an encoded bitstream. Further, the encoded bitstream is sent to a decoder for decoding to recover the original channel signals.

[0072] It should be noted that in order for the decoder to perform decoding, it is also necessary to send the target transformation matrix of each channel group in each frequency band. Optionally, the target transformation matrix of each frequency band can be written into the encoded bitstream along with the channel's encoding information and sent, or it can be sent to the decoder separately and synchronously with the encoded bitstream.

[0073] In the audio encoding method provided in this application embodiment, the encoder divides and groups the channel signals into frequency bands to obtain frequency domain coefficients. Based on the frequency domain coefficients of each channel, the target transform matrix for each frequency band corresponding to the channel group can be determined. Further, decorrelation processing is performed on the frequency domain coefficients of the channels according to the target transform matrix to obtain the encoding information of the channel group. Based on the encoding information, the encoded bitstream is obtained and then sent to the decoder for decoding. In this application embodiment, the frequency domain coefficients in each frequency band within the channel group are encoded using the target transform matrix, thereby enabling compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, reduces the burden on the encoder, and lowers transmission and storage costs.

[0074] Figure 2 This is a flowchart illustrating an audio encoding method provided in an embodiment of this application. This audio encoding method can be executed by an encoder. Figure 2 As shown, the method may include, but is not limited to, the following steps:

[0075] S201, group the vocal tract sequence to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence. There are one or more identical vocal tracts between adjacent vocal tract groups.

[0076] In the embodiments of this application, S201 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0077] S202, perform frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame.

[0078] In the embodiments of this application, S202 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0079] S203, based on the frequency domain coefficients of each channel, determine the first cross-correlation matrix between the channels corresponding to each frequency band.

[0080] In one implementation, the energy value of any channel in each frequency band is determined based on the frequency domain coefficient of any channel, and the first cross-correlation matrix corresponding to each frequency band is determined based on the energy value of each channel in each frequency band.

[0081] Optionally, the root mean square energy (RMS) of a channel can be calculated based on its frequency domain coefficients to determine the energy value of that channel in each frequency band. The formula for calculating RMS is shown below:

[0082]

[0083] Where c represents the channel index value, which ranges from 1 to M, where M is the number of input channels; b is the frequency band index value; X i N represents the frequency domain coefficients of the corresponding audio channel; c,b This represents the number of frequency points within the frequency band of this channel.

[0084] It should be noted that the number of frequency points can be calculated based on the bandwidth and the frequency resolution corresponding to the bandwidth.

[0085] Typically, since each channel has the same data length, the data length within the same frequency band across different channels is also the same under the same frequency band division method. Therefore, the number of frequency points within the same frequency band is the same, i.e., N. c,b =N b .

[0086] In one implementation, after obtaining the energy value of each channel in each frequency band, the energy ratio of each pair of channels in the channel sequence in any frequency band b is obtained. Further, based on the energy ratio between channels in any frequency band b, the cross-correlation coefficient between each pair of channels in any frequency band b is determined.

[0087] Optionally, the energy ratio Q of each pair of channels in the channel sequence at any frequency band b can be obtained, wherein the formula for calculating the energy ratio Q is as follows:

[0088]

[0089] Where c represents the channel index value, which ranges from 1 to M, where M is the number of input channels; b is the number of frequency band divisions; c1 can be the same as or different from c2, and similarly b1 can be the same as or different from b2.

[0090] In one implementation, energy discrimination can be performed using the energy ratio Q, and based on the energy discrimination result, the cross-correlation coefficient between channels in any frequency band b can be determined. If the energy ratio in any frequency band b is less than or equal to a first preset threshold, the cross-correlation coefficient between any two channels in any frequency band b is determined to be zero; if the energy ratio in any frequency band b is greater than or equal to a second preset threshold, the cross-correlation coefficient between any two channels in any frequency band b is determined to be zero. The first preset threshold is less than the second preset threshold.

[0091] If the energy ratio in any frequency band b is between the first set threshold and the second set threshold, that is, the energy ratio is greater than the first set threshold and less than the second set threshold, the cross-correlation coefficient of the two channels in any frequency band b is determined according to the frequency domain coefficients of the two channels in any frequency band b.

[0092] In other words, if the energy ratio Q between any two channels in any frequency band b is too large or too small, the cross-correlation coefficient between the two channels will be 0; if the energy ratio Q between any two channels in any frequency band b is between (0.5, 2), the cross-correlation coefficient between the two channels will be further calculated.

[0093] Optionally, the cross-correlation coefficient can be determined based on the magnitude of Q. The formula for calculating the cross-correlation coefficient is shown below:

[0094]

[0095] Where [x,y] represents the channel index; b represents the frequency band index; N b The number of frequency points within the frequency band of this channel.

[0096] Furthermore, based on the cross-correlation coefficients between any two channels in any frequency band b, the first cross-correlation matrix corresponding to any frequency band b can be obtained. For example, if the channel sequence contains M channels, the first cross-correlation matrix can be obtained as shown below:

[0097]

[0098] In the first cross-correlation matrix, the first row and first column correspond to channel 1, the second row and second column correspond to channel 2, and so on, with the Mth row and Mth column corresponding to channel M.

[0099] It should be noted that each frequency band in the frequency band set can be assigned a corresponding first cross-correlation matrix in the manner described above. If the frequency band set includes b frequency bands, then it has b first cross-correlation matrices.

[0100] S204, determine the second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band. The second cross-correlation matrix includes the cross-correlation coefficients between channels within the channel group.

[0101] In one implementation, a second cross-correlation matrix corresponding to the channel group can be extracted from the first cross-correlation matrix based on the channels within the channel group, according to the M*M first cross-correlation matrix shown above. Optionally, the channel index of the channel within the channel group is determined, and the second cross-correlation matrix of the channel group is extracted from the first cross-correlation matrix based on the channel index. Optionally, if the channel group includes three channels, the second cross-correlation matrix corresponding to the channel group can be determined from the first cross-correlation matrix based on the cross-correlation coefficients between any two channels included in the channel group, and the second cross-correlation matrix is ​​a 3*3 matrix.

[0102] For example, the channel sequence includes 5 channels, arranged sequentially as channel 1, channel 2, channel 3, channel 4, and channel 5. Each channel group contains three channels: channel group 1 includes channel 1, channel 2, and channel 3; channel group 2 includes channel 2, channel 3, and channel 4; and channel group 3 includes channel 3, channel 4, and channel 5. The first cross-correlation matrix is ​​a 5x5 matrix, which can be represented as follows... Figure 3 As shown. For channel group 2, the matrix elements intersecting rows 2 to 4 and columns 2 to 4 of the first cross-correlation matrix can be used as the second cross-correlation matrix for channel group 2. For example, the second cross-correlation matrix for channel group 2 can be as follows: Figure 3 The portion within the dashed box in the first cross-correlation matrix is ​​shown below; the second cross-correlation matrix for channel group 2 is shown below:

[0103]

[0104] It should be noted that each channel group has a second cross-correlation matrix in each frequency band.

[0105] S205, based on the second cross-correlation matrix of the channel group in each frequency band, determines the target transformation matrix of the channel group in each frequency band from the transformation matrix set.

[0106] In one implementation, for any frequency band b, if the cross-correlation coefficients between any two channels included in the second cross-correlation matrix of any frequency band b satisfy the condition for selecting the specified transformation matrix from the set of transformation matrices, then the specified transformation matrix is ​​selected as the target transformation matrix for any frequency band b.

[0107] If the cross-correlation coefficients between any two channels included in the second cross-correlation matrix of any frequency band b do not meet the conditions, then based on the maximum cross-correlation coefficients between any two channels, a transformation matrix other than the specified transformation matrix is ​​selected from the set of transformation matrices as the target transformation matrix for any frequency band b.

[0108] Optionally, a cross-correlation threshold can be set, and the target transformation matrix for the channel group can be determined from the transformation matrix set based on the cross-correlation threshold. The second cross-correlation matrix of the channel group in each frequency band contains the cross-correlation coefficients between each pair of channels within the channel group. If the cross-correlation coefficients between each pair of channels are greater than the set threshold, then the specified transformation matrix is ​​selected as the target transformation matrix for any frequency band b; if the cross-correlation coefficients between each pair of channels are not both greater than the threshold, then based on the maximum cross-correlation coefficient between each pair of channels, a transformation matrix other than the specified transformation matrix is ​​selected from the transformation matrix set as the target transformation matrix for any frequency band b.

[0109] For example, the transformation matrix set can include M0, M1, M2, M3, and M4. The specified transformation matrix can be M4. If the cross-correlation coefficients of the three channels are all greater than a threshold, then the target transformation matrix is ​​determined to be M4; if the maximum value of the cross-correlation coefficients of the three channels is greater than the threshold, then the target transformation matrix is ​​determined based on the maximum value.

[0110] S206, based on the target transformation matrix of each frequency band, performs same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group to obtain the coding information of the channel group.

[0111] In the embodiments of this application, S206 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0112] S207: Obtain the encoded bitstream based on the encoding information of the channel group, and send the encoded bitstream to the decoder for decoding.

[0113] In the embodiments of this application, S207 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0114] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transformation matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, decreases the burden on the encoder, and lowers transmission and storage costs.

[0115] Figure 4 This is a flowchart illustrating an audio encoding method provided in an embodiment of this application. This audio encoding method can be executed by an encoder. Figure 4 As shown, the method may include, but is not limited to, the following steps:

[0116] S401, group the vocal tract sequence to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence. There are one or more identical vocal tracts between adjacent vocal tract groups.

[0117] In the embodiments of this application, S401 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0118] S402 performs frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame.

[0119] In the embodiments of this application, S402 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0120] S403, based on the frequency domain coefficients of each channel, determines the first cross-correlation matrix between the channels corresponding to each frequency band.

[0121] In the embodiments of this application, S403 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0122] S404: Normalize the first cross-correlation matrix of each frequency band, and extract the second cross-correlation matrix of each frequency band corresponding to the channel group from the normalized first cross-correlation matrix of each frequency band according to the channels included in the channel group.

[0123] In one implementation, the first cross-correlation matrix of each frequency band between channels is normalized to obtain a normalized first cross-correlation matrix for each frequency band. Then, the second cross-correlation matrix corresponding to each frequency band of the channel group is extracted from the normalized first cross-correlation matrix of each frequency band. It is understood that the second cross-correlation matrix is ​​a normalized matrix.

[0124] In one implementation, the channel identifier associated with any matrix element in the first cross-correlation matrix is ​​determined. Further, based on the associated channel identifier, a normalized matrix element corresponding to any matrix element can be determined. Here, any matrix element is the cross-correlation coefficient between two channels, the channel identifier corresponding to the row containing any matrix element can be one channel identifier associated with that matrix element, and the channel identifier corresponding to the row containing any matrix element can be another channel identifier associated with that matrix element.

[0125] For example, any matrix element Corr in the first cross-correlation matrix corresponding to frequency band b. [2,3],b Let's take an example to explain, where any matrix element Corr [2,3],b The associated channel identifiers are 2 and 3, meaning that any matrix element Corr [2,3],b The associated audio channels are channel 2 and channel 3.

[0126] Any matrix element Corr [2,3],b The associated channels are channel 2 and channel 3, which can determine any matrix element Corr. [2,3],b The normalized matrix elements are Corr [2,2],b and Corr [3,3],b .

[0127] Furthermore, based on any matrix element and the normalized matrix element, the normalization result of any matrix element is obtained. Optionally, the normalization formula corresponding to any matrix element is as follows:

[0128]

[0129] Where b represents the frequency band index value; Represents any matrix element; Represents the normalized matrix elements.

[0130] It should be noted that since the cross-correlation coefficients are equal at symmetrical positions along the diagonal, it is unnecessary to calculate the cross-correlation coefficients at the lower left of the diagonal. Instead, the elements at the upper right of the diagonal can be used for normalization. During this calculation, if the denominator is greater than 0, normalization continues; if the denominator is less than 0, the corresponding cross-correlation coefficient is set to 0. The diagonal runs from the upper left to the lower right corner.

[0131] S405, based on the second cross-correlation matrix on any frequency band b, determine the cross-correlation coefficient between any two channels in the channel group on any frequency band b.

[0132] In one implementation, since the elements in the second cross-correlation matrix are the cross-correlation coefficients between any two channels within a channel group on any frequency band b, the cross-correlation coefficients between any two channels within a channel group on any frequency band b can be determined based on the second cross-correlation matrix.

[0133] S406, based on the cross-correlation coefficients between any two channels within the channel group on any frequency band b, determine the target transformation matrix of the channel group on any frequency band b from the transformation matrix set.

[0134] In one implementation, a threshold value of the cross-correlation coefficient between any two channels within a channel group on any frequency band b can be set as Thr. By comparing the cross-correlation coefficient between any two channels within a channel group with the threshold value, the target transformation matrix can be determined according to preset conditions.

[0135] Optionally, if the cross-correlation coefficients between any two channels satisfy the condition for selecting the specified transformation matrix from the set of transformation matrices, that is, the cross-correlation coefficients between any two channels within the channel group are all greater than Thr, then the specified transformation matrix M4 is selected as the target transformation matrix for any frequency band b.

[0136] For example, [L,C,R] represents three channels within a channel group. If the cross-correlation coefficients between channel L and channel C, between channel L and channel R, and between channel C and channel R are all greater than Thr, then a specified transformation matrix M4 is selected as the target transformation matrix for any frequency band b.

[0137] Optionally, if the cross-correlation coefficient between any two channels in any frequency band b does not meet the condition, then a transformation matrix other than the specified transformation matrix is ​​selected as the target transformation matrix for any frequency band b based on the maximum cross-correlation coefficient between any two channels.

[0138] In other words, the largest cross-correlation coefficient is selected from the three cross-correlation coefficients: the cross-correlation coefficient between channel L and channel C, the cross-correlation coefficient between channel L and channel R, and the cross-correlation coefficient between channel C and channel R. If the largest cross-correlation coefficient is greater than Thr, then a target transformation matrix for any frequency band b is selected from the transformation matrices other than the specified transformation matrix. For example, if the specified transformation matrix is ​​M4, then a target transformation matrix is ​​selected from M0 to M3 based on the largest cross-correlation coefficient.

[0139] Optionally, if the maximum cross-correlation coefficient is greater than Thr, then the target transformation matrix is ​​selected according to formula (6):

[0140]

[0141] For example, if the maximum cross-correlation coefficient between any two channels in a channel group on any frequency band b is the same as the cross-correlation coefficient between channel L and channel C, then M1 is selected as the target transformation matrix; if the maximum cross-correlation coefficient between any two channels in a channel group on any frequency band b is the same as the cross-correlation coefficient between channel L and channel R, then M2 is selected as the target transformation matrix; if the maximum cross-correlation coefficient between any two channels in a channel group on any frequency band b is the same as the cross-correlation coefficient between channel C and channel R, then M3 is selected as the target transformation matrix.

[0142] S407, based on the target transformation matrix of each frequency band, performs same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group to obtain the coding information of the channel group.

[0143] In the embodiments of this application, S407 can be implemented in any of the ways described in the various embodiments of this application. This is not limited here and will not be described in detail.

[0144] S408 obtains the encoded bitstream based on the encoding information of the channel group and sends the encoded bitstream to the decoder for decoding.

[0145] In the embodiments of this application, S408 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.

[0146] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transformation matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, decreases the burden on the encoder, and lowers transmission and storage costs.

[0147] Figure 5 This is a flowchart illustrating an audio encoding method provided in an embodiment of this application. This audio encoding method can be executed by an encoder. Figure 5As shown, the method may include, but is not limited to, the following steps:

[0148] S501, group the vocal tract sequence to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence. There are one or more overlapping vocal tracts between adjacent vocal tract groups.

[0149] In the embodiments of this application, S501 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0150] S502 performs frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame.

[0151] In the embodiments of this application, S502 can be implemented in any of the ways described in the various embodiments of this application. This is not limited here and will not be described in detail.

[0152] S503, based on the frequency domain coefficients of each channel, determine the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set.

[0153] In the embodiments of this application, S503 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0154] S504, obtain the frequency domain coefficients of the channels within the channel group on any frequency band b, and obtain the first coding information of the channel group on any frequency band b based on the frequency domain coefficients on any frequency band b and the target transformation matrix corresponding to any frequency band b.

[0155] In one implementation, for the first channel group, which includes channel 1, channel 2 and channel 3, the first coding information of the first channel group in any frequency band b is obtained based on the frequency domain coefficients of the channels in the first channel group in any frequency band b, and according to the frequency domain coefficients in any frequency band b and the target transformation matrix corresponding to any frequency band b. The first coding information includes the center information M1, the side information S1 and the first information T1 of the first channel group.

[0156] It should be noted that, for each remaining channel group other than the first channel group, based on the frequency domain coefficients of the channels within the remaining channel group in any frequency band b, and according to the frequency domain coefficients in any frequency band b and the target transformation matrix corresponding to any frequency band b, the first coding information of the remaining channel group in any frequency band b is obtained. The first coding information includes the first information T of the remaining channel group. i .

[0157] It is understandable that the first coding information of the first channel group fully includes the center information, side information and first information, while the first coding information of the remaining channel groups may only include the first information.

[0158] The formula for calculating vocal tract coding information is shown below:

[0159] [LCR]*M=[MST] (7)

[0160] Where [L,C,R] represents the three channels within the channel group; M is the target transformation matrix determined based on the frequency domain coefficients; and [MST] represents the encoding information of the channel group.

[0161] It should be noted that during decorrelation calculation, the first decorrelation coding unit outputs the complete coding information [MST], while the remaining coding units only output the first information [T].

[0162] S505: Based on the first coding information of each frequency band of the channel group, the second coding information of the channel group is obtained.

[0163] In one implementation, the target transformation matrix of the channel group can be determined based on the first coding information of the channel group in each frequency band, and the decorrelation mode of the channel group in each frequency band, which is the second coding information of the channel group, can be determined based on the target transformation matrix.

[0164] It is understandable that when the target transformation matrix is ​​different for different frequency bands, the decorrelation modes corresponding to those frequency bands will also be different. Conversely, when the target transformation matrix is ​​the same for different frequency bands, the decorrelation modes corresponding to those frequency bands will also be the same.

[0165] S506, based on the second coding information and the target transformation matrix corresponding to each frequency band, the coding information of the channel group is obtained.

[0166] In one implementation, the frequency domain coefficients of the channel group in each frequency band can be decorrelated using the target transformation matrix corresponding to each frequency band, based on the second coding information, to obtain the coding information of the channel group in each frequency band. After obtaining the coding information corresponding to each frequency band, the coding information of the channel group in the entire frequency band can be obtained based on the coding information in each frequency band. It is understood that the coding information of the channel group includes the coding information of the channel group in all frequency bands.

[0167] S507 obtains the encoded bitstream based on the encoding information of the channel group and sends the encoded bitstream to the decoder for decoding.

[0168] In the embodiments of this application, S507 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0169] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transformation matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, decreases the burden on the encoder, and lowers transmission and storage costs.

[0170] like Figure 6 The diagram illustrates a possible encoding flowchart of an embodiment of this application. The audio signal in each channel undergoes MDCT transformation to obtain the MDCT coefficients (frequency domain coefficients) for each frame of each channel. Then, the channel signals are input to a band-segmentation processing unit for frequency band division, obtaining the frequency domain coefficients for each band. Subsequently, an energy calculation unit is used to calculate the energy value of each channel in each frequency band. This energy value is then input to a cross-correlation calculation unit to obtain the cross-correlation coefficients between channels in each frequency band, thus obtaining the first cross-correlation matrix between channels in each frequency band. It can be understood that each channel group corresponds to a decorrelation unit. The frequency domain coefficients of the three channels within the channel group and the first cross-correlation matrix of each frequency band are input to the decorrelation unit, which performs same-frequency decorrelation processing on the channel group to obtain the encoded information for that channel. For example... Figure 6 As shown, the frequency domain coefficients of channels 1, 2, and 3 in the first channel group, along with the first cross-correlation matrix of each frequency band, are input into decorrelation unit 1. Decorrelation unit 1 performs same-frequency decorrelation processing and outputs the encoding information of channel group 1. The encoding information of channel group 1 includes center information M1, side information S1, and first information T1. The frequency domain coefficients of channels 2, 3, and 4 in channel group 2, along with the first cross-correlation matrix of each frequency band, are input into decorrelation unit 2. Decorrelation unit 2 performs same-frequency decorrelation processing and outputs the encoding information of channel group 2. The encoding information of channel group 2 includes first information T2. The frequency domain coefficients of channels 3, 4, and 5 in channel group 3, along with the first cross-correlation matrix of each frequency band, are input into decorrelation unit 3. Decorrelation unit 3 performs same-frequency decorrelation processing and outputs the encoded information of channel group 3, which includes T3. Similarly, for the last channel group M-2, the frequency domain coefficients of channels M-2, M-1, and M, along with the first cross-correlation matrix of each frequency band, are input into decorrelation unit M-2. Decorrelation unit M-2 performs same-frequency decorrelation processing and outputs the encoded information of channel group M-2, which includes T3. M-2 .

[0171] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transformation matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, decreases the burden on the encoder, and lowers transmission and storage costs.

[0172] Figure 7 This is a schematic flowchart illustrating an audio decoding method provided in an embodiment of this application. This audio decoding method can be executed by a decoder. Figure 7 As shown, the method may include, but is not limited to, the following steps:

[0173] S701 receives the encoded bitstream sent by the encoder, which includes encoding information for multiple channel groups.

[0174] In this embodiment of the application, the channel group is obtained by sequentially grouping the channel sequence. Each channel group includes several consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups.

[0175] In this embodiment, the decoder receives the encoded bitstream sent by the encoder, reads the encoding information of multiple channel groups from the encoded bitstream, and performs an inverse transformation on the input channel signal to obtain the original channel signal.

[0176] Referring to the description in the above embodiments, the encoder can group the M channels in the channel sequence to obtain multiple channel groups. Each channel group contains three consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups. The specific process of the encoder encoding the channel groups can be found in the above embodiments and will not be repeated here.

[0177] S702 decodes multiple channel groups sequentially. For the current channel group that has been decoded, the target decoding matrix of each frequency band in the frequency band set corresponding to the current channel group is determined based on the encoding information of the current channel group.

[0178] In one implementation, second encoding information for the channel group across the entire frequency band can be obtained from the encoding information. This second encoding information includes first encoding information for each frequency band and the target transform matrix corresponding to each frequency band. For any frequency band b in the frequency band set, based on the target transform matrix of any frequency band b, the target decoding matrix of the current channel group in any frequency band b is obtained by querying the correspondence between the transform matrix and the decoding matrix.

[0179] It should be noted that there is a one-to-one correspondence between the transformation matrix and the decoding matrix. In one implementation, the decoding matrix may include the following matrices:

[0180]

[0181]

[0182]

[0183] For example, transform matrix M0 corresponds to decoding matrix J0, transform matrix M1 corresponds to decoding matrix J1, transform matrix M2 corresponds to decoding matrix J2, transform matrix M3 corresponds to decoding matrix J3, and transform matrix M4 corresponds to decoding matrix J4. If the target transform matrix of the current channel group is M4, then the target decoding matrix of the current channel group in any frequency band b is J4.

[0184] S703, based on the target decoding matrix of the current channel group in each frequency band, obtains the decoding frequency domain coefficients of the current channel group from the encoding information of the current channel group.

[0185] In one embodiment, the encoding information of the current channel group includes second encoding information of the channel group across the entire frequency band, wherein the second encoding information includes first encoding information of the channel group in each frequency band.

[0186] For any frequency band b in the frequency band set, based on the target decoding matrix corresponding to any frequency band b, the first encoded information of any frequency band b is decoded to obtain the first decoded frequency domain coefficients of the current channel group on any frequency band b. Furthermore, based on the first decoded frequency domain coefficients on each frequency band, the decoded frequency domain coefficients of the channel group on the entire frequency band can be obtained.

[0187] S704 obtains the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of multiple channel groups.

[0188] In one implementation, the signal of a channel group can be converted from the frequency domain to the time domain based on the decoding frequency domain coefficients of multiple channel groups. Optionally, using inverse MDCT transform, the frequency domain signal of a channel can be converted into a time domain signal based on the decoding frequency domain coefficients, thereby obtaining the decoded audio signal of each channel.

[0189] In the audio decoding method provided in this application embodiment, the decoder receives the encoded bitstream sent by the encoder and obtains the encoded information of each channel group from it. The encoded information is decoded by the decoding unit in the order of multiple channel groups to obtain the decoded frequency domain coefficients of each channel group. The decoded frequency domain coefficients of each channel group are then converted from frequency domain to time domain to obtain the decoded audio signal for each channel. During the decoding process, the decoder employs frequency division decoding similar to that on the encoder side, which enables the recovery of multi-channel audio signals. Because the encoder performs compression, the multi-channel signal is easier to transmit, saving transmission space.

[0190] Figure 8 This is a schematic flowchart illustrating an audio decoding method provided in an embodiment of this application. This audio decoding method can be executed by a decoder. Figure 8 As shown, the method may include, but is not limited to, the following steps:

[0191] S801: Obtain the first coding information of the first channel group on any frequency band b from the coding information of the first channel group.

[0192] In one implementation, each channel group contains three consecutive channels, wherein the first channel group contains channel 1, channel 2, and channel 3. Optionally, the encoding information of the first channel group can be obtained from the encoded bitstream. The encoding information of the first channel group includes the first encoding information of the first channel group in each frequency band. The first encoding information includes at least the center information M1, the side information S1, and the first information T1 of the first channel group.

[0193] S802, based on the target decoding matrix of the first channel group in each frequency band, decodes the first encoded information of the first channel group in each frequency band to obtain the first decoded frequency domain coefficient of the first channel group in any frequency band b.

[0194] S803 obtains the second decoding frequency domain coefficients of the first channel group based on the first decoding frequency domain coefficients of the first channel group in each frequency band, wherein the second decoding frequency domain coefficients are three-way outputs.

[0195] In one implementation, the decoder can obtain the target transform matrix of the first channel group in any frequency band b from the encoded information corresponding to the first channel group. Optionally, a correspondence between the transform matrix and the decoding matrix is ​​established in advance. Further, the target decoding matrix of the first channel group in any frequency band b can be determined based on the target transform matrix in any frequency band b.

[0196] Based on the target decoding matrix in any frequency band b, the first encoded information in any frequency band b is inversely transformed to obtain the first decoding frequency domain coefficients of the first channel group in any frequency band b.

[0197] In one implementation, for any frequency band b in the frequency band set, based on the target decoding matrix corresponding to any frequency band b, the first encoded information of any frequency band b is decoded to obtain the first decoded frequency domain coefficients of the first channel group on any frequency band b. It is understood that the first encoded information of the first channel group includes M1S1T1, and the first decoded frequency domain coefficients on any frequency band b after decoding include three outputs. The first decoded frequency domain coefficients may include...

[0198] Furthermore, based on the first decoding frequency domain coefficients in each frequency band, the decoding frequency domain coefficients of the first channel group in the entire frequency band can be obtained. The decoding frequency domain coefficients of the first channel group in the entire frequency band include the three outputs.

[0199] In one implementation, the decoding formula for the first channel group is as follows:

[0200]

[0201] Where [M1 S1 T1] represents the second encoded information of the first channel group; J is the target decoding matrix. This represents the three-channel decoding frequency domain coefficients of the first channel group.

[0202] S804 identifies several decoded channel groups that are adjacent to and consecutive to the current channel group, which are the upmix channel groups corresponding to the current channel group.

[0203] In one implementation, when the current channel group is one of multiple channel groups excluding the first channel group, several adjacent and consecutive decoded channel groups can be identified as the upmix channel groups corresponding to the current channel group for decoding operations. Optionally, the several decoded channel groups may include two channel groups.

[0204] For example, if the current channel group is channel group 2, the corresponding upmix channel group includes the decoding frequency domain coefficients of the last two outputs of channel group 1; if the current channel group is channel group 3, the corresponding upmix channel group includes the decoding frequency domain coefficient of the last output of channel group 1 and one decoding frequency domain coefficient of channel group 2; if the current channel group is channel group 4, the corresponding upmix channel group includes one decoding frequency domain coefficient of channel group 2 and one decoding frequency domain coefficient of channel group 3.

[0205] S805: Obtain the first coding information of the current channel group in any frequency band b from the coding information.

[0206] Based on the encoding process on the encoder side, it can be seen that the output of each channel group starting from channel group 2 is a single output. The decoder can obtain the first encoded information of the current channel group in any frequency band b from the encoded information of the current channel group. This first encoded information is the first information T of the previous channel group in any frequency band b. i .

[0207] S806, obtains the decoding frequency domain coefficients of the upper mixing channel group in each frequency band.

[0208] In one implementation, the target transform matrix of the current upmix channel group in each frequency band can be determined from the encoded information received from the decoder, and the target decoding matrix of the upmix channel group is determined based on the target transform matrix. The decoder performs an inverse transform on the target decoding matrix to obtain the decoding frequency domain coefficients of the upmix channel group in each frequency band.

[0209] S807, decode the first coded information on any frequency band b according to the target decoding matrix corresponding to any frequency band b and the decoding frequency domain coefficients on any frequency band b, and obtain the first decoding frequency domain coefficients on any frequency band b.

[0210] In one implementation, the decoder obtains the target transformation matrix of the current channel group in each frequency band from the encoded information, and then determines the target decoding matrix of the current channel group in each frequency band based on the target transformation matrix.

[0211] For any frequency band b in the frequency band set, based on the target decoding matrix of any frequency band b, the first encoded information T on any frequency band b is... i Decode the first decoding frequency domain coefficients of the current channel group in any frequency band b by decoding the first decoding frequency domain coefficients of the current channel group in any frequency band b.

[0212] It is understandable that the first encoded information of the current channel group includes the first information T. i After decoding, the first decoded frequency domain coefficients in any frequency band b include one output. The first decoded frequency domain coefficients may include...

[0213] S808 obtains the decoding frequency domain coefficients of the current channel group based on the first decoding frequency domain coefficients of each frequency band of the current channel group, wherein the decoding frequency domain coefficients of the current channel group are one output.

[0214] Furthermore, based on the first decoding frequency domain coefficients in each frequency band, the decoding frequency domain coefficients of the first channel group in the entire frequency band can be obtained. The decoding frequency domain coefficients of the first channel group in the entire frequency band include the three outputs.

[0215] In one implementation, the decoding formula for the current channel group i is as follows:

[0216]

[0217] in, This represents the frequency domain coefficient of one output channel corresponding to the (i-2)th channel group (upper mix channel group). T is the decoding frequency domain coefficient of one output channel corresponding to the (i-2)th channel group (upper mix channel group), i It is the first encoded information of the current channel group i, and J is the target decoding matrix. This represents the decoding frequency domain coefficient of the current channel group's corresponding output.

[0218] It should be noted that since the current channel group uses different encoding modes, and the decoding unit only outputs one channel of decoding frequency domain coefficients, only T is needed. i The value of the decoding matrix J varies depending on the value of the matrix. The specific values ​​are shown below:

[0219]

[0220]

[0221]

[0222] In this context, * represents meaningless.

[0223] S809 obtains the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of multiple channel groups.

[0224] In the embodiments of this application, S809 can be implemented in any of the embodiments of this application, and no limitation is made here, nor will it be described in detail.

[0225] like Figure 9 As shown in the decoding flowchart, the decoding end needs to perform decoding sequentially according to the decoding units. Upmixing decoding unit 1 is the first upmixing decoding unit. The first encoded information [M1 S1 T1] is input into the first upmixing decoding unit, and the output of the three-channel decoding frequency domain coefficients is... The first encoded information T2 and the decoding frequency domain coefficients output by the first upmixing decoding unit are used to... The input is fed into the second upmixing decoding unit, and the output decoding frequency domain coefficients are... The first encoded information T3 and the decoding frequency domain coefficients output by the first upmixing decoding unit are used to... The decoding frequency domain coefficients output by the second upmixing decoding unit The input is fed into the third upmixing decoding unit, and the output decoding frequency domain coefficients are... The first encoded information T i The decoding frequency domain coefficients output by the (i-2)th upmixing decoding unit The decoding frequency domain coefficients output by the (i-1)th upmixing decoding unit The input is given to the i-th upmixing decoding unit, and the output decoding frequency domain coefficients are: And so on, the first encoded information T M-2 The decoding frequency domain coefficients output by the (M-4)th upmixing decoding unit The decoding frequency domain coefficients output by the (M-3)th upmixing decoding unit The input is fed into the (M-2)th upmixing decoding unit, and the output decoding frequency domain coefficients are...

[0226] In this embodiment, the decoder uses frequency division decoding similar to that on the encoder side during the decoding process, which can realize the recovery of multi-channel audio signals. Since the encoder has compressed the signal, the multi-channel signal is easier to transmit, saving transmission space.

[0227] Figure 10 This is a flowchart illustrating an audio encoding method provided in an embodiment of this application. Figure 10As shown, the method may include, but is not limited to, the following steps:

[0228] S1001, group the vocal tract sequence to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence. There are one or more overlapping vocal tracts between adjacent vocal tract groups.

[0229] S1002 performs frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the candidate frequency domain coefficients for each channel in each frame.

[0230] S1003, based on the frequency domain coefficients of each channel, determine the first cross-correlation matrix between the channels corresponding to each frequency band.

[0231] S1004, determine the second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band. The second cross-correlation matrix includes the cross-correlation coefficients between channels within the channel group.

[0232] S1005, based on the second cross-correlation matrix of the channel group in each frequency band, determine the target transformation matrix of the channel group in each frequency band from the transformation matrix set.

[0233] S1006, obtain the frequency domain coefficients of the channels within the channel group on any frequency band b, and obtain the first coding information of the channel group on any frequency band b based on the frequency domain coefficients on any frequency band b and the target transformation matrix corresponding to any frequency band b.

[0234] S1007: Based on the first coding information of each frequency band of the channel group, the second coding information of the channel group is obtained.

[0235] S1008, based on the second coding information and the target transformation matrix corresponding to each frequency band, the coding information of the vocal tract group is obtained.

[0236] S1009 obtains the encoded bitstream based on the encoding information of the channel group and sends the encoded bitstream to the decoder for decoding.

[0237] S1010 receives the encoded bitstream sent by the encoder.

[0238] S1011 decodes multiple channel groups sequentially, targeting the currently decoded channel group.

[0239] S1012, obtain the target transformation matrix of each frequency band from the encoded information.

[0240] S1013, based on the target transformation matrix of any frequency band b, query the mapping relationship between the transformation matrix and the decoding matrix, and obtain the target decoding matrix of the current channel group in any frequency band b.

[0241] S1014: Based on the target decoding matrix of the current channel group in each frequency band, obtain the decoding frequency domain coefficients of the current channel group from the encoding information of the current channel group.

[0242] S1015: Obtain the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of multiple channel groups.

[0243] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using a target transformation matrix, thereby enabling compression of audio signals from multiple channels. This also reduces redundancy between channels, decreases the encoder's burden, and lowers transmission and storage costs. During the decoding process, the decoder employs frequency division decoding similar to that used on the encoder side, enabling the recovery of the multi-channel audio signals.

[0244] Figure 11 This is a block diagram illustrating an audio encoding apparatus according to an exemplary embodiment. (Refer to...) Figure 11 The audio encoding device 1100 of this application embodiment includes: a channel grouping module 1101, a frequency domain conversion module 1102, a matrix determination module 1103, an encoding module 1104, and a transmission module 1105.

[0245] The channel grouping module 1101 is configured to perform grouping of the channel sequence to obtain multiple channel groups, each channel group including several consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups.

[0246] The frequency domain processing module 1102 is configured to perform frequency domain conversion on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame.

[0247] The matrix determination module 1103 is configured to perform the determination of the target transformation matrix for each frequency band in the frequency band set corresponding to the channel group from the transformation matrix set based on the frequency domain coefficients of each channel.

[0248] The encoding module 1104 is configured to perform co-band decorrelation processing on the frequency domain coefficients of the channels within the channel group based on the target transformation matrix of each frequency band, so as to obtain the encoding information of the channel group.

[0249] The sending module 1105 is configured to execute the encoding information based on the channel group to obtain the encoded bitstream, and send the encoded bitstream to the decoder for decoding.

[0250] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: determining a first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel; determining a second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band, the second cross-correlation matrix including the cross-correlation coefficients between channels within the channel group; and determining a target transformation matrix of the channel group in each frequency band from the transformation matrix set based on the second cross-correlation matrix of the channel group in each frequency band.

[0251] In one embodiment of this application, the encoding module 1104 is further configured to perform: obtaining the frequency domain coefficients of the channels within the channel group in any frequency band b, and obtaining the first encoding information of the channel group in any frequency band b based on the frequency domain coefficients in any frequency band b and the target transformation matrix corresponding to any frequency band b; obtaining the second encoding information of the channel group based on the first encoding information in each frequency band of the channel group; and obtaining the encoding information of the channel group based on the second encoding information and the target transformation matrix corresponding to each frequency band.

[0252] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: determining the cross-correlation coefficient between any two channels in the channel group in any frequency band b based on the second cross-correlation matrix in any frequency band b; and determining the target transformation matrix of the channel group in any frequency band b from the transformation matrix set according to the cross-correlation coefficient between any two channels in the channel group in any frequency band b.

[0253] In one embodiment of this application, the matrix determination module 1103 is further configured to perform the following: if the cross-correlation coefficient between any two channels satisfies the condition for selecting a specified transformation matrix from the set of transformation matrices, then the specified transformation matrix is ​​selected as the target transformation matrix for any frequency band b; if the cross-correlation coefficient between any two channels does not satisfy the condition, then based on the maximum cross-correlation coefficient between any two channels, a transformation matrix other than the specified transformation matrix is ​​selected from the set of transformation matrices as the target transformation matrix for any frequency band b.

[0254] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: determining the energy value of any channel in each frequency band based on the frequency domain coefficients of any channel; and determining the first cross-correlation matrix corresponding to each frequency band based on the energy value of each channel in each frequency band.

[0255] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: obtaining the energy ratio of each pair of channels in the channel sequence in any frequency band b; if the energy ratio in any frequency band b is less than or equal to a first preset threshold, or if the energy ratio in any frequency band b is greater than or equal to a second preset threshold, determining that the cross-correlation coefficient between each pair of channels in any frequency band b is zero; wherein the first preset threshold is less than the second preset threshold, and if the energy ratio is between the first preset threshold and the second preset threshold, determining the cross-correlation coefficient between each pair of channels in any frequency band b based on the frequency domain coefficients of each pair of channels in any frequency band b; and obtaining the first cross-correlation matrix corresponding to any frequency band b based on the cross-correlation coefficient between each pair of channels in any frequency band b.

[0256] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: normalizing the first cross-correlation matrix of each frequency band, and extracting the second cross-correlation matrix of each frequency band corresponding to the channel group from the normalized first cross-correlation matrix of each frequency band according to the channels included in the channel group.

[0257] In one embodiment of this application, the matrix determination module 1103 is further configured to perform: determining the channel identifier associated with any matrix element in the first cross-correlation matrix; determining the normalized matrix element corresponding to any matrix element based on the associated channel identifier; and obtaining the normalization result of any matrix element based on the any matrix element and the normalized matrix element.

[0258] In one embodiment of this application, the encoding module 1104 is further configured to perform: for the first channel group, based on the frequency domain coefficients of the channels within the first channel group in any frequency band b, and according to the frequency domain coefficients in any frequency band b and the target transformation matrix corresponding to any frequency band b, to obtain first encoding information of the first channel group in any frequency band b, the first encoding information including center information, side information and first information of the first channel group; for each remaining channel group other than the first channel group, based on the frequency domain coefficients of the channels within the remaining channel group in any frequency band b, and according to the frequency domain coefficients in any frequency band b and the target transformation matrix corresponding to any frequency band b, to obtain first encoding information of the remaining channel group in any frequency band b, the first encoding information including first information of the remaining channel group.

[0259] In one embodiment of this application, the channel grouping module 1101 is further configured to perform: determining that adjacent channel groups include a first channel group and a second channel group, wherein the first channel group and the second channel group each include three consecutive channels in the channel sequence, and the first channel group and the second channel group include two identical channels.

[0260] In this embodiment, the frequency domain coefficients of each frequency band within the channel group are encoded using the target transformation matrix, thereby enabling the compression of audio signals from multiple channels. Furthermore, it reduces redundancy between multiple channels, decreases the burden on the encoder, and lowers transmission and storage costs.

[0261] Figure 12 This is a block diagram illustrating an audio decoding apparatus according to an exemplary embodiment. (Refer to...) Figure 11 The audio decoding device 1200 of this application embodiment includes: a receiving module 1201, a matrix determination module 1202, and a decoding module 1203.

[0262] The receiving module 1201 is configured to receive the encoded bitstream sent by the encoder. The encoded bitstream includes the encoding information of multiple channel groups. The channel groups are obtained by sequentially grouping the channel sequences. Each channel group includes several consecutive channels in the channel sequence. There are one or more identical channels between adjacent channel groups.

[0263] The matrix determination module 1202 is configured to perform sequential decoding of multiple channel groups, and for the current channel group that has been decoded, determine the target decoding matrix of each frequency band in the frequency band set corresponding to the current channel group based on the encoding information of the current channel group.

[0264] The decoding module 1203 is configured to execute the target decoding matrix of the current channel group in each frequency band, obtain the decoding frequency domain coefficients of the current channel group based on the encoding information of the current channel group, and obtain the decoded audio signal of each channel in the channel sequence based on the decoding frequency domain coefficients of multiple channel groups.

[0265] In one embodiment of this application, the matrix determination module 1202 is further configured to perform: obtaining the target transformation matrix of each frequency band from the encoding information; querying the mapping relationship between the transformation matrix and the decoding matrix according to the target transformation matrix of any frequency band b, and obtaining the target decoding matrix of the current channel group in any frequency band b.

[0266] In one embodiment of this application, the decoding module 1203 is further configured to perform: obtaining first encoded information of the first channel group in any frequency band b from the encoded information of the first channel group; decoding the first encoded information of the first channel group in any frequency band b based on the target decoding matrix of the first channel group in any frequency band b to obtain the first decoding frequency domain coefficients of the first channel group in any frequency band b; and obtaining the decoding frequency domain coefficients of the first channel group according to the first decoding frequency domain coefficients of the first channel group in each frequency band, wherein the first decoding frequency domain coefficients and the decoding frequency domain coefficients of the first channel group include three outputs.

[0267] In one embodiment of this application, the decoding module 1203 is further configured to execute: first encoded information of the first channel group on any frequency band b, including at least center information, side information and first information on any frequency band b.

[0268] In one embodiment of this application, the decoding module 1203 is further configured to perform: determining a plurality of decoded channel groups that are adjacent to and consecutive to the current channel group as the upmixing channel group corresponding to the current channel group; obtaining the first encoding information of the current channel group in any frequency band b from the encoding information; obtaining the decoding frequency domain coefficients of the upmixing channel group in each frequency band; decoding the first encoding information in any frequency band b according to the target decoding matrix corresponding to any frequency band b and the decoding frequency domain coefficients in any frequency band b to obtain the first decoding frequency domain coefficients in any frequency band b; obtaining the decoding frequency domain coefficients of the current channel group according to the first decoding frequency domain coefficients in each frequency band of the current channel group, wherein the decoding frequency domain coefficients of the current channel group include one output.

[0269] In one embodiment of this application, the decoding module 1203 is further configured to perform: the first encoded information of the current channel group in any frequency band b includes the first information of the current channel group in the any frequency band b.

[0270] In the embodiments of this application, the decoder uses frequency division decoding similar to that on the encoder side during the decoding process, which can realize the recovery of multi-channel audio signals. Since the encoder has compressed the signal, the multi-channel signal is easier to transmit, saving transmission space.

[0271] Figure 13 This is a schematic diagram of another audio processing device 1300 provided in an embodiment of this application. The audio processing device 1300 can be an encoder, a decoder, a chip, chip system, or processor that supports the encoder in implementing the above methods, or a chip, chip system, or processor that supports the decoder in implementing the above methods. This device can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.

[0272] The audio processing device 1300 may include one or more processors 1301. The processor 1301 may be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control the audio processing device (e.g., base station, baseband chip, decoder, decoder chip, DU or CU, etc.), execute computer programs, and process data from the computer programs.

[0273] Optionally, the audio processing device 1300 may further include one or more memories 1302, which may store a computer program 1304. The processor 1301 executes the computer program 1304 to cause the audio processing device 1300 to perform the methods described in the above method embodiments. Optionally, the memory 1302 may also store data. The audio processing device 1300 and the memory 1302 may be provided separately or integrated together.

[0274] Optionally, the audio processing device 1300 may also include a transceiver 1305 and an antenna 1306. The transceiver 1305 may be referred to as a transceiver unit, transceiver, or transceiver circuit, etc., and is used to implement the transmission and reception functions. The transceiver 1305 may include a receiver and a transmitter. The receiver may be referred to as a receiver or receiving circuit, etc., and is used to implement the receiving function; the transmitter may be referred to as a transmitter or transmitting circuit, etc., and is used to implement the transmitting function.

[0275] Optionally, the audio processing apparatus 1300 may further include one or more interface circuits 1307. The interface circuits 1307 are used to receive code instructions and transmit them to the processor 1301. The processor 1301 executes the code instructions to cause the audio processing apparatus 1300 to perform the methods described in the above method embodiments.

[0276] In one implementation, the processor 1301 may include a transceiver for implementing receive and transmit functions. For example, the transceiver may be a transceiver circuit, an interface, or an interface circuit. The transceiver circuit, interface, or interface circuit for implementing receive and transmit functions may be separate or integrated. The aforementioned transceiver circuit, interface, or interface circuit can be used for reading and writing code / data, or it can be used for transmitting or relaying signals.

[0277] In one implementation, processor 1301 may store computer program 1303, which runs on processor 1301 and causes audio processing device 1300 to perform the methods described in the above method embodiments. Computer program 1303 may be embedded in processor 1301; in this case, processor 1301 may be implemented in hardware.

[0278] In one implementation, the audio processing device 1300 may include circuitry capable of performing the transmitting, receiving, or communicating functions described in the foregoing method embodiments. The processor and transceiver described in this application can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits (RFICs), mixed-signal ICs, application-specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductors (CMOS), n-metal-oxide-semiconductor (NMOS), positive channel metal oxide semiconductors (PMOS), bipolar junction transistors (BJTs), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0279] The audio processing device described in the above embodiments may be an encoder or a decoder, but the scope of the audio processing device described in this application is not limited thereto, and the structure of the audio processing device may vary. Figure 13 The audio processing device can be a standalone device or part of a larger device. For example, the audio processing device could be:

[0280] (1) Independent integrated circuit IC, or chip, or chip system or subsystem;

[0281] (2) A collection of one or more ICs, optionally including storage components for storing data and computer programs;

[0282] (3) ASIC, such as modem;

[0283] (4) Modules that can be embedded in other devices;

[0284] (5) Receivers, decoders, smart decoders, cellular phones, wireless devices, handheld devices, mobile units, vehicle-mounted devices, encoders, cloud devices, artificial intelligence devices, etc.;

[0285] (6) Others, etc.

[0286] For cases where the audio processing device can be a chip or a chip system, please refer to [link / reference]. Figure 14 The diagram shows the structure of the chip. Figure 14 The chip shown includes a processor 1401 and an interface 1402. There can be one or more processors 1401, and multiple interfaces 1402.

[0287] Optionally, the chip also includes a memory 1403 for storing necessary computer programs and data.

[0288] In some implementations, the chip can be used to implement the functions of the decoder described in the embodiments of this application.

[0289] In some implementations, the chip can be used to implement the encoder functions described in the embodiments of this application.

[0290] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0291] This application also provides an audio processing system, which includes the aforementioned... Figure 13 The embodiments include an audio processing device as an encoder and an audio processing device as a decoder, or the system includes the aforementioned... Figure 14 The embodiments include an audio processing device as an encoder and an audio processing device as an encoder.

[0292] This application also provides a readable storage medium having instructions stored thereon that, when executed by a computer, implement the functions of any of the above method embodiments.

[0293] This application also provides a computer program product that, when executed by a computer, implements the functions of any of the above method embodiments.

[0294] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0295] Those skilled in the art will understand that the various numerical designations such as "first," "second," etc., involved in this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application, nor do they indicate the order of sequence.

[0296] At least one in this application can also be described as one or more, and multiple can be two, three, four or more, and this application does not impose any limitation. In the embodiments of this application, for a technical feature, the technical features in that technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", and there is no order or size among the technical features described by "first", "second", "third", "A", "B", "C" and "D".

[0297] The correspondences shown in the tables of this application can be configured or predefined. The values ​​of the information in each table are merely examples and can be configured to other values; this application is not limited to these values. When configuring the correspondences between information and parameters, it is not necessarily required to configure all the correspondences shown in each table. For example, the correspondences shown in some rows of the tables in this application may not be configured. Furthermore, appropriate modifications and adjustments can be made based on the above tables, such as splitting, merging, etc. The names of the parameters shown in the headings of the above tables can also use other names that the audio processing device can understand, and the values ​​or representations of the parameters can also be other values ​​or representations that the audio processing device can understand. In the implementation of the above tables, other data structures can also be used, such as arrays, queues, containers, stacks, linear lists, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables, or hash tables, etc.

[0298] The term "predefined" in this application can be understood as definition, pre-defined, stored, pre-stored, pre-negotiated, pre-configured, solidified, or pre-burned.

[0299] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0300] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0301] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An audio encoding method, characterized in that, Performed by the encoder, the method includes: The vocal tract sequence is grouped to obtain multiple vocal tract groups. Each vocal tract group includes several consecutive vocal tracts in the vocal tract sequence, and there are one or more identical vocal tracts between adjacent vocal tract groups. The audio signals of each channel in the channel sequence are converted in the frequency domain frame by frame to obtain the frequency domain coefficients of each channel for each frame. Based on the frequency domain coefficients of each channel, determine the first cross-correlation matrix between the channels corresponding to each frequency band; The second cross-correlation matrix of the channel group is determined from the first cross-correlation matrix of each frequency band, and the second cross-correlation matrix includes the cross-correlation coefficient between the channels within the channel group; Based on the second cross-correlation matrix of the channel group in each frequency band, the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group is determined from the transformation matrix set. Based on the target transformation matrix of each frequency band, the frequency domain coefficients of the channels within the channel group are subjected to same-frequency band decorrelation processing to obtain the coding information of the channel group; The encoded bitstream is obtained based on the encoding information of the channel group, and the encoded bitstream is sent to the decoder for decoding.

2. The method according to claim 1, characterized in that, The target transform matrix based on each frequency band is used to perform in-band decorrelation processing on the frequency domain coefficients of the channels in the channel group to obtain the coding information of the channel group, including: Obtain the channels within the channel group in any frequency band b The frequency domain coefficients on, and according to the any frequency band b The frequency domain coefficients and any frequency band b The corresponding target transformation matrix is ​​used to obtain the channel group in any frequency band. b The first encoded information on; The second coding information of the channel group is obtained based on the first coding information in each frequency band of the channel group; Based on the second encoding information and the target transformation matrix corresponding to each frequency band, the encoding information of the vocal tract group is obtained.

3. The method according to claim 1, characterized in that, The step of determining the target transformation matrix of the channel group in each frequency band from the transformation matrix set based on the second cross-correlation matrix of the channel group in each frequency band includes: Based on any frequency band b The second cross-correlation matrix is ​​used to determine the cross-correlation coefficients between any two channels within the channel group in any frequency band. According to the frequency band between any two channels within the aforementioned channel group b The cross-correlation coefficients on the transform matrix set are used to determine the channel group in any frequency band. b The target transformation matrix.

4. The method according to claim 3, characterized in that, The basis for the frequency band between any two channels within the channel group b The cross-correlation coefficients on the transform matrix set are used to determine the target transform matrix of the channel group in any frequency band, including: If the cross-correlation coefficients between any two channels satisfy the condition for selecting a specified transformation matrix from the set of transformation matrices, then the specified transformation matrix is ​​selected as any frequency band. b The target transformation matrix; If the cross-correlation coefficient between any two channels does not meet the condition, then based on the maximum cross-correlation coefficient between any two channels, a transformation matrix other than the specified transformation matrix is ​​selected from the set of transformation matrices as any frequency band. b The target transformation matrix.

5. The method according to claim 1, characterized in that, The step of determining the first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel includes: Based on the frequency domain coefficients of any channel, determine the energy value of the channel in each frequency band; The first cross-correlation matrix corresponding to each frequency band is determined based on the energy value of each channel in each frequency band.

6. The method according to claim 5, characterized in that, The step of determining the first cross-correlation matrix corresponding to each frequency band based on the energy values ​​of each channel in each frequency band includes: Obtain the pairwise channels in the channel sequence in any frequency band b The energy ratio on; If any of the frequency bands b The energy ratio on the frequency band is less than or equal to a first set threshold, or, in any frequency band b The energy ratio on the two channels is greater than or equal to the second set threshold to determine the energy ratio between the two channels in any frequency band. b The cross-correlation coefficient is zero, wherein the first set threshold is less than the second set threshold; If any of the frequency bands b The energy ratio is between the first set threshold and the second set threshold, according to the two channels in any frequency band. b The frequency domain coefficients determine the frequency domain coefficients of the two channels in any frequency band. b The cross-correlation coefficient; Based on the two channels in any frequency band b The cross-correlation coefficient is used to obtain the cross-correlation coefficient of any frequency band. b The corresponding first cross-correlation matrix.

7. The method according to claim 1, characterized in that, Determining the second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band further includes: The first cross-correlation matrix of each frequency band is normalized, and the second cross-correlation matrix of each frequency band corresponding to the channel group is extracted from the normalized first cross-correlation matrix of each frequency band according to the channels included in the channel group.

8. The method according to claim 7, characterized in that, The normalization process for the first cross-correlation matrix of each frequency band includes: Determine the vocal tract identifier associated with any element in the first cross-correlation matrix; Based on the associated channel identifier, determine the normalized matrix element corresponding to any matrix element; Based on any matrix element and the normalized matrix element, the normalization result of any matrix element is obtained.

9. The method according to any one of claims 2-8, characterized in that, The method further includes: For the first channel group, based on the channels within the first channel group in any frequency band b The frequency domain coefficients on, and according to the any frequency band b The frequency domain coefficients and any frequency band b The corresponding target transformation matrix is ​​used to obtain the first channel group in any frequency band. b The first encoded information includes the center information, side information and first information of the first channel group; For each remaining channel group other than the first channel group, based on the channels within the remaining channel group in any frequency band b The frequency domain coefficients on, and according to the any frequency band b The frequency domain coefficients and any frequency band b The corresponding target transformation matrix is ​​used to obtain the remaining channel group in any frequency band. b The first encoded information includes the first information of the remaining channel group.

10. The method according to any one of claims 1-8, characterized in that, There are one or more overlapping channels between the adjacent groups of channels, including: The adjacent channel groups are defined as a first channel group and a second channel group, wherein the first channel group and the second channel group each include three consecutive channels in the channel sequence, and the first channel group and the second channel group include two identical channels.

11. An audio decoding method, characterized in that, The method, executed by the decoder, includes: The encoder sends an encoded bitstream, which includes encoding information for multiple channel groups. Each channel group is obtained by sequentially grouping a channel sequence. Each channel group includes several consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups. The multiple channel groups are decoded sequentially. For the current channel group, the target transform matrix of each frequency band is obtained from the encoding information of the current channel group, and then based on any frequency band in the frequency band set... b The target transform matrix is ​​used to query the mapping relationship between the transform matrix and the decoding matrix, and to determine the current channel group in any frequency band. b The target decoding matrix on; Based on the target decoding matrix of the current channel group in each frequency band, the decoding frequency domain coefficients of the current channel group are obtained by encoding the current channel group. Based on the decoding frequency domain coefficients of the multiple channel groups, the decoded audio signal of each channel in the channel sequence is obtained.

12. The method according to claim 11, characterized in that, When the current channel is the first channel group among the plurality of channel groups, the step of obtaining the decoding frequency domain coefficients of the current channel group based on the target decoding matrix of the current channel group in each frequency band, using the encoding information of the current channel group, includes: From the encoding information of the first channel group, obtain the first channel group in any frequency band. b The first encoded information on; Based on the first channel group in any frequency band b The target decoding matrix on the first channel group in any frequency band b Decode the first encoded information to obtain the first channel group in any frequency band. b The first decoding frequency domain coefficients; The decoding frequency domain coefficients of the first channel group are obtained based on the first decoding frequency domain coefficients of the first channel group in each frequency band. The first decoding frequency domain coefficients and the decoding frequency domain coefficients of the first channel group include three outputs.

13. The method according to claim 12, characterized in that, The first channel group is in any frequency band b The first encoding information includes at least the center information, side information, and first information on any of the frequency bands.

14. The method according to claim 11, characterized in that, When the current channel is a channel group other than the first channel group among the plurality of channel groups, the step of obtaining the decoding frequency domain coefficients of the current channel group based on the target decoding matrix of the current channel group in each frequency band and the encoding information of the current channel group includes: Identify a number of decoded channel groups that are adjacent to and consecutive to the current channel group, and define them as the upper mixing channel group corresponding to the current channel group; Obtain the current channel group in any frequency band from the encoded information. b The first encoded information on; Obtain the decoding frequency domain coefficients of the upper mixing channel group in each frequency band; According to any of the frequency bands b The corresponding target decoding matrix and any frequency band b The decoding frequency domain coefficients on the above, for any frequency band b Decode the first encoded information to obtain the arbitrary frequency band. b The first decoding frequency domain coefficients; Based on the first decoding frequency domain coefficients in each frequency band of the current channel group, the decoding frequency domain coefficients of the current channel group are obtained, and the decoding frequency domain coefficients of the current channel group are output in one channel.

15. The method according to claim 14, characterized in that, The current channel group in any frequency band b The first encoding information includes the current channel group in any frequency band. b The first piece of information.

16. An audio encoding device, characterized in that, include: The channel grouping module is configured to group the channel sequence to obtain multiple channel groups, each of which includes a number of consecutive channels in the channel sequence, and there are one or more identical channels between adjacent channel groups. The frequency domain processing module is configured to perform frequency domain transformation on the audio signals of each channel in the channel sequence frame by frame to obtain the frequency domain coefficients of each channel for each frame. The matrix determination module is configured to perform the following: determine a first cross-correlation matrix between channels corresponding to each frequency band based on the frequency domain coefficients of each channel; and determine a second cross-correlation matrix of the channel group from the first cross-correlation matrix of each frequency band, wherein the second cross-correlation matrix includes the cross-correlation coefficients between channels within the channel group. Based on the second cross-correlation matrix of the channel group in each frequency band, the target transformation matrix of each frequency band in the frequency band set corresponding to the channel group is determined from the transformation matrix set. The encoding module is configured to execute the target transformation matrix based on each frequency band, perform same-band decorrelation processing on the frequency domain coefficients of the channels within the channel group, and obtain the encoding information of the channel group; The sending module is configured to execute the encoding information based on the channel group to obtain the encoded bitstream, and send the encoded bitstream to the decoder for decoding.

17. An audio decoding device, characterized in that, include: The receiving module is configured to receive the encoded bitstream sent by the encoder. The encoded bitstream includes the encoding information of multiple channel groups. The channel groups are obtained by sequentially grouping the channel sequences. Each channel group includes several consecutive channels in the channel sequence. There are one or more identical channels between adjacent channel groups. The matrix determination module is configured to perform sequential decoding of the plurality of channel groups, and for the current channel group decoded, obtain the target transform matrix of each frequency band from the encoding information of the current channel group, and determine the target transform matrix of each frequency band according to any frequency band in the frequency band set. b The target transform matrix is ​​used to query the mapping relationship between the transform matrix and the decoding matrix, and to obtain the current channel group in any frequency band. b The target decoding matrix is ​​used to determine the target decoding matrix of the current channel group in any frequency band b. The decoding module is configured to execute the target decoding matrix of the current channel group in each frequency band, obtain the decoding frequency domain coefficients of the current channel group based on the encoding information of the current channel group, and obtain the decoded audio signal of each channel in the channel sequence according to the decoding frequency domain coefficients of the multiple channel groups.

18. An encoder, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the steps of the method according to any one of claims 1-10.

19. A decoder, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the steps of the method according to any one of claims 11-15.

20. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1-10.

21. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 11-15.

Citation Information

Patent Citations

  • Encoding and decoding method and system for multi-channel three-dimensional voice frequency

    CN103400582A

  • Stereo audio signal processing method and device, encoding equipment, decoding equipment and storage medium

    CN114258568A