A method and apparatus for processing three-dimensional audio signals

By spatially encoding three-dimensional audio signals to generate transmission channel signals and attribute information, the problem of bit allocation in three-dimensional audio signal encoding is solved, encoding efficiency is improved, and storage and transmission requirements are reduced.

CN115472170BActive Publication Date: 2026-04-03HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the bit allocation of signals cannot be effectively determined during the encoding process of three-dimensional audio signals, resulting in high storage and transmission bandwidth requirements.

Method used

By spatially encoding the three-dimensional audio signal, transmission channel signals and attribute information are generated. The bit allocation ratio is determined by using attribute information such as the energy ratio, coding efficiency, and identification of the virtual speaker signal group and the residual signal group.

Benefits of technology

It achieves efficient bit allocation for three-dimensional audio signals, improves coding efficiency, and reduces storage and transmission requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115472170B_ABST
    Figure CN115472170B_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for processing three-dimensional audio signals, used to implement bit allocation of the signal. The method includes: spatially encoding the three-dimensional audio signal to be encoded to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group; and determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202110657283.7, filed on June 11, 2021, entitled "A Method and Apparatus for Processing Three-Dimensional Audio Signals", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of audio processing technology, and in particular to a method and apparatus for processing three-dimensional audio signals. Background Technology

[0003] 3D audio technology has been widely applied in wireless communication voice, virtual reality / augmented reality, and media audio. It is an audio technology that acquires, processes, transmits, renders, and plays back sound events and 3D sound field information from the real world. 3D audio technology gives sound a strong sense of space, immersion, and surround sound, providing an extraordinary "sound-immersive" auditory experience. Higher-order ambisonics (HOA) technology, with its speaker layout independence during recording, encoding, and playback, and the rotatable playback characteristics of HOA format data, offers greater flexibility in 3D audio playback, thus attracting wider attention and research.

[0004] Acquisition devices (such as microphones) collect large amounts of data to record 3D sound field information and transmit the 3D audio signals to playback devices (such as speakers and headphones) for playback. Because the 3D sound field information is large in volume, it requires significant storage space and high bandwidth for transmission. To address these issues, the 3D audio signals can be compressed, and the compressed data can be stored or transmitted.

[0005] Currently, encoders can use multiple pre-configured virtual speakers to encode 3D audio signals. However, how to allocate the bits of the signal after the encoder has encoded the 3D audio signal remains an unsolved problem. Summary of the Invention

[0006] This application provides a method and apparatus for processing three-dimensional audio signals, used to implement bit allocation of the signal.

[0007] To address the aforementioned technical problems, this application provides the following technical solutions:

[0008] In a first aspect, embodiments of this application provide a method for processing three-dimensional audio signals, comprising: spatially encoding a three-dimensional audio signal to be encoded to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group; and determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information. In the above solution, embodiments of this application obtain transmission channel signals and transmission channel attribute information through three-dimensional audio signal encoding. The transmission channel signal may include at least one virtual speaker signal group and at least one residual signal group. The transmission channel attribute information can be used to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, thereby solving the problem of not being able to determine the bit allocation of the signal.

[0009] In one possible implementation, the transmission channel attribute information includes: virtual loudspeaker coding efficiency; the spatial coding of the three-dimensional audio signal to be encoded to obtain the transmission channel attribute information includes: using a virtual loudspeaker to reconstruct the three-dimensional audio signal to be encoded to obtain a reconstructed three-dimensional audio signal; obtaining the energy characterization value of the reconstructed three-dimensional audio signal and the energy characterization value of the three-dimensional audio signal to be encoded; and obtaining the virtual loudspeaker coding efficiency based on the energy characterization value of the reconstructed three-dimensional audio signal and the energy characterization value of the three-dimensional audio signal to be encoded. In the above scheme, the encoding end first performs signal reconstruction using a virtual loudspeaker to obtain the reconstructed three-dimensional audio signal. The encoding end can calculate the energy characterization value of the signal for each transmission channel, for example, it can obtain the energy characterization value of the reconstructed three-dimensional audio signal and the energy characterization value of the three-dimensional audio signal to be encoded. The energy characterization value of the three-dimensional audio signal is different before and after signal reconstruction, so the virtual loudspeaker coding efficiency can be calculated by the change in the energy characterization value before and after signal reconstruction.

[0010] In one possible implementation, the transmission channel attribute information includes: the energy percentage of the virtual speaker signal group; the method further includes: obtaining the energy characterization value of the virtual speaker signal group based on the energy characterization value of each virtual speaker signal in the virtual speaker signal group; obtaining the energy characterization value of the residual signal group based on the energy characterization value of each residual signal in the residual signal group; and obtaining the energy percentage of the virtual speaker signal group based on the energy characterization value of the virtual speaker signal group and the energy characterization value of the residual signal group. In the above scheme, the encoding end first obtains the energy characterization value of each virtual speaker signal in the virtual speaker signal group, and then adds the energy characterization values ​​of all virtual speaker signals in the same group to obtain the energy characterization value of the virtual speaker signal group. If there are multiple virtual speaker signal groups, the energy characterization value of each group can be calculated in the above manner. Similarly, the encoding end can obtain the energy characterization value of the residual signal group based on the energy characterization value of each residual signal in the residual signal group. Finally, the encoding end can obtain the energy percentage of the virtual speaker signal group based on the energy characterization value of the virtual speaker signal group and the energy characterization value of the residual signal group. The energy percentage of the virtual speaker signal group indicates its proportion in the total signal energy of the transmission channel. If the energy percentage of the virtual speaker signal group is high, it means that the virtual speaker signal group is dominant in the total signal energy of the transmission channel. If the energy percentage of the virtual speaker signal group is low, it means that the virtual speaker signal group is not dominant (i.e., weak) in the total signal energy of the transmission channel.

[0011] In one possible implementation, the transmission channel attribute information includes: a virtual speaker encoding identifier, which indicates whether the bit allocation of the virtual speaker signal group is dominant; the spatial encoding of the three-dimensional audio signal to be encoded to obtain the transmission channel attribute information includes: spatially encoding the three-dimensional audio signal to be encoded to obtain the number of dissimilar sound sources and the virtual speaker encoding efficiency of the transmission channel signal; and obtaining the virtual speaker encoding identifier based on the number of dissimilar sound sources and the virtual speaker encoding efficiency of the transmission channel signal. In the above scheme, after obtaining the number of dissimilar sound sources and the virtual speaker encoding efficiency of the transmission channel signal, the encoding end obtains the specific value of the virtual speaker encoding identifier based on the decision conditions satisfied by the number of dissimilar sound sources and the virtual speaker encoding efficiency of the transmission channel signal.

[0012] In one possible implementation, obtaining the virtual speaker encoding identifier based on the number of dissimilar sound sources in the transmission channel signal and the virtual speaker encoding efficiency includes: determining the virtual speaker encoding identifier as dominant when the number of dissimilar sound sources in the transmission channel signal is less than or equal to a preset dissimilar sound source number threshold, and the virtual speaker encoding efficiency is greater than or equal to a preset first virtual speaker encoding efficiency threshold; or determining the virtual speaker encoding identifier as non-dominant when the number of dissimilar sound sources in the transmission channel signal is greater than the preset dissimilar sound source number threshold, or the virtual speaker encoding efficiency is less than the preset first virtual speaker encoding efficiency threshold. In the above scheme, the encoding end can determine the virtual speaker encoding identifier by comparing the number of dissimilar sound sources, the virtual speaker encoding efficiency, and the above decision conditions, thereby using the virtual speaker encoding identifier to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group.

[0013] In one possible implementation, dominance includes sub-dominance or strong dominance; determining that the virtual speaker encoding identifier is dominant includes: determining that the virtual speaker encoding identifier is sub-dominant when the virtual speaker encoding efficiency is greater than or equal to the first virtual speaker encoding efficiency threshold and the virtual speaker encoding efficiency is less than or equal to a preset second virtual speaker encoding efficiency threshold; or determining that the virtual speaker encoding identifier is strong dominance when the virtual speaker encoding efficiency is greater than or equal to the first virtual speaker encoding efficiency threshold and the virtual speaker encoding efficiency is greater than the preset second virtual speaker encoding efficiency threshold; wherein, the second virtual speaker encoding efficiency threshold is greater than the first virtual speaker encoding efficiency threshold. In the above scheme, the encoding end can further divide the cases where the virtual speaker encoding identifier is dominant, that is, it can obtain two cases: sub-dominant and strong dominance of the virtual speaker encoding identifier. It is understood that if the virtual speaker encoding identifier is strong dominance, then the virtual speaker signal group needs to be allocated more bits, for example, after determining the initial bit ratio of the virtual speaker signal group, the bit ratio can be increased. If the virtual speaker code is subdominant, then the virtual speaker signal group needs to be allocated fewer bits than when the virtual speaker code is strongly dominant. However, the number of bits allocated to the virtual speaker signal group still needs to be greater than the number of bits allocated when the virtual speaker code is nondominant. For example, after determining the initial bit percentage of the virtual speaker signal group, this bit percentage can be increased. In comparison, the increased bit percentage in the strongly dominant case is greater than the increased bit percentage in the subdominant case.

[0014] In one possible implementation, the transmission channel attribute information includes: the energy percentage of the virtual speaker signal group, and / or the virtual speaker encoding identifier; determining the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group based on the transmission channel attribute information includes: when the energy percentage of the virtual speaker signal group is greater than or equal to a preset first energy percentage threshold, and / or the virtual speaker encoding identifier is strongly dominant, determining the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group according to a preset first signal group bit allocation algorithm; ... determining the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group according to a preset first signal group bit allocation algorithm. When the energy percentage of the virtual speaker signal group is equal to or less than a preset second energy percentage threshold and less than a preset first energy percentage threshold, and / or the virtual speaker encoding identifier is subdominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset second signal group bit allocation algorithm; wherein the second energy percentage threshold is less than the first energy percentage threshold; or, when the energy percentage of the virtual speaker signal group is less than the preset first energy percentage threshold, or the virtual speaker encoding identifier is non-dominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset third signal group bit allocation algorithm. In the above scheme, the encoding end can preset multiple signal group bit allocation algorithms. Under different conditions of transmission channel attribute information, different signal group bit allocation algorithms can be used, thereby allocating bit allocation percentages to the virtual speaker signal group and the residual signal group that are adapted to these conditions when the transmission channel attribute information meets certain conditions, thus improving the encoding efficiency of the encoding end for three-dimensional audio signals.

[0015] In one possible implementation, when the energy proportion of the virtual speaker signal group is greater than or equal to a preset first energy proportion threshold, and / or the virtual speaker encoding identifier is strongly dominant, determining the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group according to a preset first signal group bit allocation algorithm includes: when directionalNrgRatio≥TH1, and / or S≤TH0 and η>TH2, calculating the bit allocation proportion of the virtual speaker signal group as follows: Ratio1_1=FAC1*directionalNrgRatio+(1–FAC1)*maxdirectionalNrgRatio; where, directionalNrgRatio represents the virtual speaker signal... The energy proportion of the group, where S is the number of dissimilar sound sources, η represents the virtual loudspeaker coding efficiency, maxdirectionalNrgRatio is a preset maximum virtual loudspeaker signal group bit allocation proportion, FAC1 is a preset first adjustment factor, Ratio1_1 is the bit allocation proportion of the virtual loudspeaker signal group, * represents multiplication, TH1 is the first energy proportion threshold, TH0 is the threshold for the number of dissimilar sound sources, and TH2 is the second virtual loudspeaker coding efficiency threshold; the bit allocation proportion of the residual signal group is calculated as follows: Ratio2 = 1 - Ratio1_1; where Ratio1_1 is the bit allocation proportion of the virtual loudspeaker signal group, and Ratio2 is the bit allocation proportion of the residual signal group. In the above scheme, as can be seen from the calculation process of Ratio1_1, the bit allocation proportion of the virtual loudspeaker signal group increases, therefore the encoder can allocate more bits to the virtual loudspeaker signal group. The transmission channel signal includes a virtual speaker signal group and a residual signal group. After obtaining the bit allocation ratio Ratio1_1 of the virtual speaker signal group, the bit allocation ratio of the residual signal group can be obtained through the above formula for Ratio2.

[0016] In one possible implementation, after obtaining the bit allocation ratio of the virtual speaker signal group, the method further includes updating the bit allocation ratio of the virtual speaker signal group as follows: Ratio1_2 = min(Ratio1_1, maxdirectionalNrgRatio + FAC2 * Ratio1_1); where Ratio1_2 represents the updated bit allocation ratio of the virtual speaker signal group, FAC2 is a preset second adjustment factor, maxdirectionalNrgRatio is a preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * represents multiplication, and min is the minimum value operation. In the above scheme, as can be seen from the calculation process of Ratio1_2, the bit allocation ratio of the virtual speaker signal group can be safely limited, restricting Ratio1_2 within a safe bit range, thereby enabling the encoding end to safely and usably allocate bits for the virtual speaker signal group.

[0017] In one possible implementation, when the energy proportion of the virtual speaker signal group is greater than or equal to a preset second energy proportion threshold and less than a preset first energy proportion threshold, and / or the virtual speaker encoding identifier is second dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to a preset second signal group bit allocation algorithm; wherein, the second energy proportion threshold being less than the first energy proportion threshold includes: when TH3≤directionalNrgRatio<TH1, and / or, when S≤TH0 and TH4≤η≤TH2, Ratio1_1 is calculated as follows: Ratio1_1=FAC3*directionalNrgRatio+(1–FAC3)*maxdirectionalNrgRatio; wherein, the maxdirectionalNrgRatio is a preset virtual speaker signal group bit allocation proportion. The bit allocation ratio of the virtual speaker signal group is defined as follows: FAC3 is a preset third adjustment factor; directionalNrgRatio represents the energy ratio of the virtual speaker signal group; S is the number of dissimilar sound sources; η represents the virtual speaker coding efficiency; Ratio1_1 is the bit allocation ratio of the virtual speaker signal group; * represents multiplication; TH0 is the threshold for the number of dissimilar sound sources; TH1 is the first energy ratio threshold; TH2 is the second virtual speaker coding efficiency threshold; TH3 is the second energy ratio threshold; and TH4 is the first virtual speaker coding efficiency threshold. The bit allocation ratio of the residual signal group is calculated as follows: Ratio2 = 1 - Ratio1_1; where Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group. In the above scheme, as can be seen from the calculation process of Ratio1_1, the bit allocation ratio of the virtual speaker signal group increases, therefore the encoding end can allocate more bits to the virtual speaker signal group. The transmission channel signal includes a virtual speaker signal group and a residual signal group. After obtaining the bit allocation ratio Ratio1_1 of the virtual speaker signal group, the bit allocation ratio of the residual signal group can be obtained through the above formula for Ratio2.

[0018] In one possible implementation, after obtaining the bit allocation ratio of the virtual speaker signal group, the method further includes updating the bit allocation ratio of the virtual speaker signal group as follows: Ratio1_2 = min(Ratio1_1, maxdirectionalNrgRatio + FAC4 * Ratio1_1); where Ratio1_2 represents the updated bit allocation ratio of the virtual speaker signal group, FAC4 is a preset fourth adjustment factor, maxdirectionalNrgRatio is a preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * represents multiplication, and min is the minimum value operation. In the above scheme, as can be seen from the calculation process of Ratio1_2, the bit allocation ratio of the virtual speaker signal group can be safely limited, restricting Ratio1_2 within a safe bit range, thereby enabling the encoding end to safely and usably allocate bits for the virtual speaker signal group.

[0019] In one possible implementation, the method further includes: if there are multiple residual signal groups, the bit allocation ratio of the i-th residual signal group is calculated as follows: Ratio2_i = Ratio2 * (R_i / C); where R_i represents the number of transmission channels included in the i-th residual signal group, C is the total number of transmission channels in all residual signal groups, Ratio2_i is the bit allocation ratio of the i-th residual signal group, * represents multiplication, and Ratio2 is the bit allocation ratio of all residual signal groups. In the above scheme, when there are multiple residual signal groups, the bit allocation ratio of each residual signal group in all residual signal groups can be determined based on the number of transmission channels in each residual signal group. For example, R_i / C represents the transmission channel ratio of the i-th residual signal group to all residual signal groups, and the bit allocation ratio of the i-th residual signal group can be obtained through (R_i / C) and Ratio2.

[0020] In one possible implementation, when the energy percentage of the virtual speaker signal group is less than a preset first energy percentage threshold, or the virtual speaker encoding identifier is not dominant, determining the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group according to a preset third signal group bit allocation algorithm includes: when directionalNrgRatio < TH3, or when S > TH0, or η < TH4, calculating the bit allocation percentage of the virtual speaker signal group as follows: Ratio1_1 = directionalNrgRatio; where directionalNrgRatio represents The energy proportion of the virtual speaker signal group, Ratio1_1 is the bit allocation proportion of the virtual speaker signal group, TH3 is the second energy proportion threshold, TH4 is the first virtual speaker coding efficiency threshold, S is the number of dissimilar sound sources, η represents the virtual speaker coding efficiency, and TH0 is the threshold for the number of dissimilar sound sources; the bit allocation proportion of the residual signal group is calculated as follows: Ratio2_1 = D / (F+D); where Ratio2_1 is the bit allocation proportion of the residual signal group, F represents the energy characterization value of the virtual speaker signal group, and D is the energy characterization value of the residual signal group. In the above scheme, as can be seen from the calculation process of Ratio1_1, the bit allocation proportion of the virtual speaker signal group is equal to the energy proportion of the virtual speaker signal group. Therefore, when the bit allocation of the virtual speaker signal group is not dominant, the encoder will not allocate more bits to the virtual speaker signal group, thereby ensuring the rationality of the parallel allocation at the encoder.

[0021] In a possible implementation, the method further includes: after obtaining the bit allocation ratio of the virtual speaker signal group, updating the bit allocation ratio of the virtual speaker signal group in the following manner: when Ratio1_1 < groupBitsRatio1, Ratio1_2 = groupBitsRatio1; when Ratio1_1 ≥ groupBitsRatio1, Ratio1_2 = FAC5 * groupBitsRatio1 + (1 – FAC5) * Ratio1_1; where, the Ratio1_2 represents the updated bit allocation ratio of the virtual speaker signal group, the FAC5 is a preset fifth adjustment factor, the Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before update, the * represents multiplication operation, and the groupBitsRatio1 is the preset bit allocation ratio of the virtual speaker signal group; after obtaining the bit allocation ratio of the residual signal group, updating the bit allocation ratio of the residual signal group in the following manner: when Ratio2_1 < groupBitsRatio2, Ratio2_2 = groupBitsRatio2; when Ratio2_1 ≥ groupBitsRatio2, Ratio2_2 = FAC6 * groupBitsRatio2 + (1 – FAC6) * Ratio2_1; where, the Ratio2_2 represents the updated bit allocation ratio of the residual signal group, the FAC6 is a preset sixth adjustment factor, the Ratio2_1 is the bit allocation ratio of the residual signal group before update, the * represents multiplication operation, and the groupBitsRatio2 is the preset bit allocation ratio of the residual signal group. In the above solution, from the calculation process of the above Ratio1_2, it can be seen that the bit allocation ratio of the virtual speaker signal group can be safely restricted, and Ratio1_2 is restricted within the safe bit range, so that the encoding end can safely and usably perform the bit allocation of the virtual speaker signal group. From the calculation process of the above Ratio2_2, it can be seen that the bit allocation ratio of the residual signal group can be safely restricted, and Ratio2_2 is restricted within the safe bit range, so that the encoding end can safely and usably perform the bit allocation of the residual signal group.

[0022] In one possible implementation, the method further includes: determining the number of bits in the virtual speaker signal group and the number of bits in the residual signal group based on the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and the total number of transmission channel bits, respectively; allocating bits to the virtual speaker signal group based on the number of bits in the virtual speaker signal group, and allocating bits to the residual signal group based on the number of bits in the residual signal group. In the above scheme, the encoding end allocates bits to the virtual speaker signal group based on the number of bits in the virtual speaker signal group, and allocates bits to the residual signal group based on the number of bits in the residual signal group, thus solving the problem that the encoding end cannot allocate bits for the virtual speaker signal and the residual signal.

[0023] In one possible implementation, determining the number of bits in the virtual speaker signal group and the number of bits in the residual signal group based on the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and the total number of transmission channel bits includes: calculating the number of bits in the virtual speaker signal group as follows: F_bitnum = Ratio1 * C_bitnum; where F_bitnum is the number of bits in the virtual speaker signal group, Ratio1 is the bit allocation ratio of the virtual speaker signal group, and C_bitnum is the total number of transmission channel bits; and calculating the number of bits in the residual signal group as follows: D_bitnum = Ratio2 * C_bitnum; where D_bitnum is the number of bits in the residual signal group, Ratio2 is the bit allocation ratio of the residual signal group, and C_bitnum is the total number of transmission channel bits. In the above scheme, the encoding end can predetermine the total number of transmission channel bits, and there is no limitation on the value of the total number of transmission channel bits. The encoding end can calculate the number of bits of the virtual speaker signal group and the number of bits of the residual signal group through the above calculation formula, thus realizing the bit allocation problem of the virtual speaker signal and the residual signal at the encoding end.

[0024] In one possible implementation, the method further includes: encoding the transmission channel signal, the bit allocation ratio of the virtual speaker signal group, and the bit allocation ratio of the residual signal group, and writing them into the bitstream. In the above scheme, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group can be encoded into the bitstream. After the encoding end sends the bitstream to the decoding end, the decoding end can obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group by parsing the bitstream. The decoding end can obtain the number of bits allocated to the virtual speaker signal group and the number of bits allocated to the residual signal group by the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, thereby decoding the bitstream to obtain a three-dimensional audio signal.

[0025] Secondly, embodiments of this application also provide a method for processing three-dimensional audio signals, including: receiving a bitstream; decoding the bitstream to obtain the bit allocation ratio of a virtual speaker signal group and the bit allocation ratio of a residual signal group; and decoding the virtual speaker signal and the residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain a decoded three-dimensional audio signal. In the above scheme, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group can be encoded into the bitstream. After the encoding end sends the bitstream to the decoding end, the decoding end can obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group by parsing the bitstream. The decoding end can obtain the number of bits allocated to the virtual speaker signal group and the number of bits allocated to the residual signal group through the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, thereby decoding the bitstream to obtain a three-dimensional audio signal.

[0026] In one possible implementation, decoding the virtual speaker signal and residual signal in the bitstream based on the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group includes: determining the number of available bits based on the bitstream; determining the number of bits in the virtual speaker signal group based on the number of available bits and the bit allocation ratio of the virtual speaker signal group; decoding the virtual speaker signal in the bitstream based on the number of bits in the virtual speaker signal group; determining the number of bits in the residual signal group based on the number of available bits and the bit allocation ratio of the residual signal group; and decoding the residual signal in the bitstream based on the number of bits in the residual signal group.

[0027] Thirdly, embodiments of this application also provide a three-dimensional audio signal processing apparatus, comprising: an encoding module, configured to spatially encode the three-dimensional audio signal to be encoded to obtain a transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group; and a bit allocation ratio determination module, configured to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information.

[0028] In a third aspect of this application, the constituent modules of the three-dimensional audio signal processing apparatus may also perform the steps described in the first aspect and various possible implementations, as detailed in the foregoing description of the first aspect and various possible implementations.

[0029] Fourthly, embodiments of this application also provide a three-dimensional audio signal processing apparatus, comprising: a receiving module for receiving a bitstream; a decoding module for decoding the bitstream to obtain a bit allocation ratio of a virtual speaker signal group and a bit allocation ratio of a residual signal group; and a signal generation module for decoding the virtual speaker signal and the residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain a decoded three-dimensional audio signal.

[0030] In the fourth aspect of this application, the constituent modules of the three-dimensional audio signal processing apparatus may also perform the steps described in the second aspect and various possible implementations, as detailed in the foregoing description of the second aspect and various possible implementations.

[0031] Fifthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect above.

[0032] Sixthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the first or second aspect above.

[0033] In a seventh aspect, embodiments of this application provide a computer-readable storage medium including a bitstream generated by the method described in the first aspect above.

[0034] Eighthly, embodiments of this application provide a communication device, which may include a terminal device or a chip, etc. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the first or second aspects above.

[0035] Ninthly, this application provides a chip system including a processor for supporting an audio encoder or audio decoder in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing necessary program instructions and data for the audio encoder or audio decoder. This chip system may be composed of chips or may include chips and other discrete devices.

[0036] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0037] In this embodiment, the three-dimensional audio signal to be encoded is first spatially encoded to obtain a transmission channel signal and transmission channel attribute information. The transmission channel signal includes at least one virtual speaker signal group and at least one residual signal group. Then, the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group are determined based on the transmission channel attribute information. This embodiment obtains the transmission channel signal and transmission channel attribute information through three-dimensional audio signal encoding. The transmission channel signal may include at least one virtual speaker signal group and at least one residual signal group. The transmission channel attribute information can be used to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, thereby solving the problem of not being able to determine the bit allocation of the signal. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the composition structure of the audio processing system provided in the embodiments of this application;

[0039] Figure 2a A schematic diagram illustrating the application of the audio encoder and audio decoder provided in this application to a terminal device;

[0040] Figure 2b A schematic diagram illustrating the application of the audio encoder provided in this application to a wireless device or a core network device;

[0041] Figure 2c A schematic diagram illustrating the application of the audio decoder provided in this application embodiment to a wireless device or core network device;

[0042] Figure 3a A schematic diagram illustrating the application of the multi-channel encoder and multi-channel decoder provided in the embodiments of this application to a terminal device;

[0043] Figure 3b A schematic diagram illustrating the application of a multi-channel encoder provided in this application to a wireless device or a core network device;

[0044] Figure 3c A schematic diagram illustrating the application of the multi-channel decoder provided in this application to a wireless device or a core network device;

[0045] Figure 4 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;

[0046] Figure 5 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;

[0047] Figure 6 This is a schematic diagram illustrating an application scenario of a three-dimensional audio signal provided in an embodiment of this application;

[0048] Figure 7 This is a schematic diagram of the composition structure of an audio encoding device provided in an embodiment of this application;

[0049] Figure 8 This is a schematic diagram of the composition structure of an audio decoding device provided in an embodiment of this application;

[0050] Figure 9 This is a schematic diagram of the composition structure of another audio encoding device provided in the embodiments of this application;

[0051] Figure 10 This is a schematic diagram of the composition structure of another audio decoding device provided in an embodiment of this application. Detailed Implementation

[0052] The embodiments of this application will now be described with reference to the accompanying drawings.

[0053] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0054] Sound is a continuous wave produced by the vibration of an object. The object that produces vibrations and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solids, or liquids), the auditory organs of humans or animals can perceive the sound.

[0055] Sound waves are characterized by pitch, intensity, and timbre. Pitch indicates the highness or lowness of a sound. Intensity indicates the loudness or volume of a sound. The unit of intensity is the decibel (dB). Timbre is also known as tone color.

[0056] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is hertz (Hz). The human ear can distinguish sounds with frequencies between 20 Hz and 20,000 Hz.

[0057] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.

[0058] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, and pulse waves, among others.

[0059] Based on the characteristics of sound waves, sound can be divided into regular sound and irregular sound. Irregular sound refers to sound emitted by the irregular vibration of a sound source. Irregular sound is, for example, noise that affects people's work, study, and rest. Regular sound refers to sound emitted by the regular vibration of a sound source. Regular sound includes speech and musical tones. When sound is represented electronically, regular sound is an analog signal that varies continuously in the time and frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.

[0060] Because human hearing has the ability to distinguish the location of sound sources in space, when a listener hears a sound in space, in addition to being able to perceive the pitch, intensity, and timbre of the sound, they can also perceive the location of the sound.

[0061] As people pay increasing attention to and demand higher quality in their auditory experience, three-dimensional audio technology has emerged to enhance the depth, presence, and spatial feel of sound. This allows listeners to not only perceive sounds from front, back, left, and right sources, but also to feel surrounded by the spatial sound field created by these sources, and to experience the sound spreading outwards, creating an immersive audio experience as if the listener were in a cinema or concert hall.

[0062] Three-dimensional audio technology refers to the concept of the space outside the human ear as a system, where the signal received at the eardrum is a three-dimensional audio signal output after the sound emitted from the sound source has been filtered by this external system. For example, the system outside the human ear can be defined as the system impulse response h(n), any sound source can be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal described in this application can refer to a higher-order ambisonics (HOA) signal or a first-order ambisonics (FOA) signal. Three-dimensional audio can also be called three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio, or binaural audio, etc.

[0063] Sound waves propagate in an ideal medium with a wave number of k = w / c and an angular frequency of w = 2πf, where f is the sound wave frequency and c is the speed of sound. The sound pressure p satisfies formula (1). For the Laplace operator.

[0064]

[0065] Assuming the spatial system outside the human ear is a sphere, with the listener at the center, the sound from outside the sphere has a projection on the sphere's surface. Filtering out sounds from outside the sphere, and assuming the sound sources are distributed on this sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound source. That is, three-dimensional audio technology is a method of fitting a sound field. Specifically, in spherical coordinates, equation (1) is solved. In the passive spherical region, the solution to equation (1) is as follows: equation (2).

[0066]

[0067] Where r represents the radius of the sphere, and θ represents the horizontal angle. The elevation angle is represented by k, the wave number by s, the amplitude of the ideal plane wave by m, and the order number of the three-dimensional audio signal (or the order number of the HOA signal). Let represent the spherical Bessel function, also known as the radial basis function, where the first 'j' represents the imaginary unit. It does not change with the angle. Represents θ, spherical harmonic function of direction, The spherical harmonic function represents the direction of the sound source. The coefficients of the three-dimensional audio signal satisfy formula (3).

[0068]

[0069] Substituting formula (3) into formula (2), formula (2) can be transformed into formula (4).

[0070]

[0071] in, The coefficients of the three-dimensional audio signal of order N are used to approximate the sound field. A sound field refers to the region in a medium where sound waves exist. N is an integer greater than or equal to 1. For example, the value of N ranges from 2 to 6. The coefficients of the three-dimensional audio signal described in the embodiments of this application may refer to HOA coefficients or ambisonic coefficients.

[0072] A three-dimensional audio signal is an information carrier that carries the spatial location information of the sound source in the sound field, describing the sound field of the listener in space. Equation (4) shows that the sound field can be expanded on a sphere according to the spherical harmonic function, that is, the sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional audio signal can be expressed by the superposition of multiple plane waves, and the sound field can be reconstructed through the coefficients of the three-dimensional audio signal.

[0073] Compared to a 5.1 channel audio signal or a 7.1 channel audio signal, an Nth-order HOA signal has (N+1) 2 With multiple channels, the HOA signal contains a large amount of data describing the spatial information of the sound field. If the acquisition device (e.g., a microphone) transmits this 3D audio signal to the playback device (e.g., a speaker), it consumes a significant amount of bandwidth. Currently, encoders can use spatial squeezed surround audio coding (S3AC), directional audio coding (DirAC), or virtual speaker selection-based coding methods to compress and encode the 3D audio signal into a bitstream, which is then transmitted to the playback device. The virtual speaker selection-based coding method can also be called match projection (MP) coding; we will use virtual speaker selection-based coding as an example later. The playback device decodes the bitstream, reconstructs the 3D audio signal, and plays the reconstructed 3D audio signal. This reduces the amount of data transmitted to the playback device and the bandwidth usage.

[0074] Currently, it is impossible to classify the sound field of the aforementioned three-dimensional audio signal. How to classify the sound field of a three-dimensional audio signal is a technical problem that this application aims to solve. In this application, the linear decomposition of the three-dimensional audio signal enables sound field classification, thereby accurately classifying the sound field and achieving the goal of obtaining the sound field classification result of the current frame.

[0075] Furthermore, current encoders suffer from the inability to achieve high compression ratios when compressing 3D audio signals. Therefore, improving the compression ratio of 3D audio signals with different sound fields is another problem addressed by the embodiments of this application.

[0076] This application provides an audio encoding technique, particularly a three-dimensional audio encoding technique for three-dimensional audio signals. Specifically, it provides an encoding technique that uses fewer channels to represent three-dimensional audio signals, thereby improving traditional audio encoding systems. Audio encoding (or commonly referred to as encoding) includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and includes processing (e.g., compressing) the raw audio to reduce the amount of data required to represent the audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are also collectively referred to as encoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0077] The technical solutions of this application embodiment can be applied to various audio processing systems, such as... Figure 1 The diagram shown is a schematic representation of the composition of an audio processing system provided in this embodiment. The audio processing system 100 may include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 generates a bitstream, which is then transmitted to the audio decoding device 102 via an audio transmission channel. The audio decoding device 102 receives the bitstream, performs its audio decoding function, and finally obtains the reconstructed signal.

[0078] In the embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be an audio encoder for the aforementioned terminal devices, wireless devices, or core network devices. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be an audio decoder for the aforementioned terminal devices, wireless devices, or core network devices. For example, the audio encoder can include a wireless access network, a core network media gateway, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The audio encoder can also be an audio encoder used in virtual reality (VR) streaming services.

[0079] In the application embodiment, taking the audio encoding module (audio encoding and audio decoding) applicable to virtual reality streaming (VR streaming) services as an example, the end-to-end audio signal processing flow includes: after the audio signal A passes through the acquisition module, it undergoes a preprocessing operation (audioPReprocessing). The preprocessing operation includes filtering out the low-frequency part in the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the directional information in the signal, and then performing encoding processing (audio encoding), packaging (file / segment encapsulation), and then sending (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), and performs binaural rendering processing on the decoded signal. The rendered signal is mapped onto the listener's headphones, which can be independent headphones or headphones on the glasses device.

[0080] like Figure 2a The diagram illustrates the application of the audio encoder and audio decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used for channel encoding of the audio signal, and the channel decoder is used for channel decoding of the audio signal. For example, the first terminal device 20 may include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 may include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and a wireless or wired second network communication device 23 are connected via a digital channel. The second terminal device 21 is connected to the wireless or wired second network communication device 23. The aforementioned wireless or wired network communication device can broadly refer to signal transmission devices, such as communication base stations, data exchange devices, etc.

[0081] In audio communication, the transmitting terminal device first acquires audio, encodes the acquired audio signal, and then performs channel coding before transmitting it over a digital channel via a wireless network or core network. The receiving terminal device, acting as the receiver, decodes the received signal to obtain the bitstream, then recovers the audio signal through audio decoding for playback.

[0082] like Figure 2b The diagram illustrates the application of the audio encoder provided in this embodiment of the application in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, the audio encoder 253 provided in this embodiment of the application, and a channel encoder 254. The other audio decoders 252 refer to audio decoders other than the standard audio decoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded using the channel decoder 251, then audio-decoded using the other audio decoder 252, then audio-encoded using the audio encoder 253 provided in this embodiment of the application, and finally channel-encoded using the channel encoder 254. After channel encoding, the signal is transmitted out. The other audio decoders 252 perform audio decoding on the bitstream decoded by the channel decoder 251.

[0083] like Figure 2c The diagram illustrates the application of the audio decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in this embodiment, other audio encoders 256, and a channel encoder 254. The other audio encoders 256 refer to audio encoders other than the standard audio encoder. Within the wireless device or core network device 25, the channel decoder 251 first performs channel decoding on the incoming signal. Then, the audio decoder 255 decodes the received audio encoded bitstream. Next, the other audio encoders 256 perform audio encoding. Finally, the channel encoder 254 performs channel encoding on the audio signal before transmission. If transcoding is required in the wireless device or core network device, corresponding audio encoding processing is necessary. The wireless device refers to radio frequency (RF) related equipment in communication, and the core network device refers to core network related equipment in communication.

[0084] In some embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be a multi-channel encoder of the aforementioned terminal device, wireless device, or core network device. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be a multi-channel decoder of the aforementioned terminal device, wireless device, or core network device.

[0085] like Figure 3aThe diagram illustrates the application of the multi-channel encoder and multi-channel decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in this embodiment of the application, and the multi-channel decoder can execute the audio decoding method provided in this embodiment of the application. Specifically, the channel encoder is used to perform channel encoding on the multi-channel signal, and the channel decoder is used to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. The second terminal device 31 may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and a wireless or wired second network communication device 33 are connected via a digital channel. The second terminal device 31 is connected to the wireless or wired second network communication device 33. The aforementioned wireless or wired network communication equipment can broadly refer to signal transmission equipment, such as communication base stations and data switching equipment. In audio communication, the transmitting terminal device performs multi-channel encoding on the acquired multi-channel signal, followed by channel encoding, and then transmits it through a wireless network or core network in a digital channel. The receiving terminal device performs channel decoding on the received signal to obtain the multi-channel signal encoded bitstream, and then recovers the multi-channel signal through multi-channel decoding for playback by the receiving terminal device.

[0086] like Figure 3b The diagram shown illustrates the application of the multi-channel encoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, as described above. Figure 2b Similarly, this will not be elaborated upon here.

[0087] like Figure 3c The diagram shown illustrates the application of the multi-channel decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, as described above. Figure 2c Similarly, this will not be elaborated upon here.

[0088] The audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the acquired multi-channel signal can involve processing the acquired multi-channel signal to obtain an audio signal, and then encoding the obtained audio signal according to the method provided in this application embodiment. The decoding end decodes the audio signal based on the multi-channel signal encoded bitstream, and recovers the multi-channel signal after upmixing. Therefore, this application embodiment can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding is required, corresponding multi-channel encoding processing is necessary.

[0089] This application first introduces a method for processing three-dimensional audio signals, provided in an embodiment. This method can be executed by a terminal device, such as an audio encoding device (hereinafter referred to as an encoder). Not limited to this, the terminal device can also be a three-dimensional audio signal processing device. Figure 4 As shown, the main methods for processing three-dimensional audio signals include the following:

[0090] 401. Spatial encoding is performed on the three-dimensional audio signal to be encoded to obtain the transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group.

[0091] The encoding end can acquire a three-dimensional audio signal, such as a scene audio signal. Specifically, the three-dimensional audio signal can be a time-domain signal or a frequency-domain signal. Additionally, the three-dimensional audio signal can also be a downsampled signal.

[0092] In this embodiment of the invention, the virtual speaker signal and the virtual speaker are in one-to-one correspondence. After determining the virtual speakers encoding the three-dimensional audio signal from the candidate virtual speaker set, the virtual speaker signals corresponding to these virtual speakers can be obtained, and then these virtual speaker signals can be grouped to obtain at least one virtual speaker signal group; or, after determining the virtual speakers encoding the three-dimensional audio signal from the candidate virtual speaker set, these virtual speakers can be grouped to obtain at least one virtual speaker group, and then the virtual speaker signals corresponding to each virtual speaker in the at least one virtual speaker group can be obtained respectively to obtain the at least one virtual speaker signal group.

[0093] In some embodiments of this application, the three-dimensional audio signal includes: a high-order stereo reverberation (HOA) signal or a first-order stereo reverberation (FOA) signal. It is not limited to this; the three-dimensional audio signal can also be other types of signals. This is merely an example and is not intended to limit the embodiments of this application.

[0094] For example, a 3D audio signal can be a time-domain HOA signal or a frequency-domain HOA signal. Furthermore, a 3D audio signal can contain all channels of a HOA signal or only some HOA channels (e.g., FOA channels). Additionally, a 3D audio signal can be all samples of a HOA signal or 1 / Q downsampled points of the HOA signal to be analyzed. Here, Q is the downsampling interval, and 1 / Q is the downsampling rate.

[0095] In this embodiment, the three-dimensional audio signal includes multiple frames. The following example focuses on processing one frame of the three-dimensional audio signal. For instance, if this frame is the current frame, then the three-dimensional audio signal contains a previous frame before the current frame and a subsequent frame after the current frame. Furthermore, the processing methods for other frames of the three-dimensional audio signal besides the current frame in this embodiment are similar to the processing method for the current frame; the following example will also focus on the processing of the current frame.

[0096] In this embodiment, after acquiring the three-dimensional audio signal, spatial encoding is first performed on the three-dimensional audio signal to obtain the transmission channel signal and transmission channel attribute information. The specific process of spatial encoding will not be described here. The process of outputting the virtual speaker signal and residual signal after spatial encoding will also not be described.

[0097] In this embodiment, after acquiring the three-dimensional audio signal to be encoded, the encoding end can perform spatial encoding on the three-dimensional audio signal and output a transmission channel signal and transmission channel attribute information. The transmission channel signal includes virtual speaker signals and residual signals. For example, the virtual speaker signals can be grouped to obtain at least one virtual speaker signal group. Similarly, the residual signals can be grouped to obtain at least one residual signal group. In this embodiment, the number of virtual speaker signal groups and the number of residual signal groups in the transmission channel signal are not limited.

[0098] In this embodiment of the application, spatial coding can also be used to output transmission channel attribute information corresponding to the transmission channel signal. This transmission channel attribute information is used to indicate the attributes of the transmission channel signal. There are various ways to implement the transmission channel attribute information, as detailed in the examples of the following embodiments.

[0099] In some embodiments of this application, the transmission channel attribute information includes: virtual loudspeaker coding efficiency; the virtual loudspeaker coding efficiency represents the efficiency of reconstructing the three-dimensional audio signal using a virtual loudspeaker. The transmission channel attribute information output by the encoder (which can also be an encoding end) through spatial coding includes the virtual loudspeaker coding efficiency, and the calculation method of this virtual loudspeaker coding efficiency will be described below.

[0100] Step 401 spatially encodes the three-dimensional audio signal to be encoded to obtain transmission channel attribute information, including:

[0101] A virtual loudspeaker is used to reconstruct the three-dimensional audio signal to be encoded, so as to obtain the reconstructed three-dimensional audio signal; wherein, the virtual loudspeaker used to reconstruct the three-dimensional audio signal to be encoded can be the virtual loudspeaker determined from the candidate virtual loudspeaker set for encoding the three-dimensional audio signal.

[0102] Obtain the energy characterization value of the reconstructed 3D audio signal and the energy characterization value of the 3D audio signal to be encoded;

[0103] The virtual loudspeaker coding efficiency is obtained based on the energy characterization values ​​of the reconstructed 3D audio signal and the energy characterization values ​​of the 3D audio signal to be encoded.

[0104] The encoding end first performs signal reconstruction using a virtual loudspeaker to obtain the reconstructed 3D audio signal. The encoding end can calculate the energy characterization value of the signal for each transmission channel. For example, it can obtain the energy characterization value of the reconstructed 3D audio signal and the energy characterization value of the 3D audio signal to be encoded. The energy characterization value of the 3D audio signal is different before and after signal reconstruction. Therefore, by observing the change in the energy characterization value before and after signal reconstruction, the encoding efficiency of the virtual loudspeaker can be calculated.

[0105] The following example illustrates the method for calculating the virtual loudspeaker coding efficiency. Taking the 3D audio signal as a HOA signal, the energy representation value of each transmission channel of the reconstructed HOA signal calculated by the encoder can be represented as R1, R2, ..., Rt. The energy representation value of each transmission channel of the original HOA signal calculated by the encoder can be represented as N1, N2, ..., Nt. Finally, the virtual loudspeaker coding efficiency η is: η = sum(R) / sum(N), where sum(R) represents the summation of R1 to Rt, and sum(N) represents the summation of N1 to Nt. The virtual loudspeaker coding efficiency can be calculated using the above formula.

[0106] In some embodiments of this application, the transmission channel attribute information includes: the energy percentage of the virtual speaker signal group; the energy percentage of the virtual speaker signal group refers to the proportion of the energy of all virtual speaker signals in the virtual speaker signal group to the total energy of all transmission channel signals. The calculation method for the energy percentage of the virtual speaker signal group will be described below.

[0107] The methods executed at the encoding end also include:

[0108] The energy characterization value of the virtual loudspeaker signal group is obtained based on the energy characterization value of each virtual loudspeaker signal in the virtual loudspeaker signal group;

[0109] The energy characterization value of the residual signal group is obtained based on the energy characterization value of each residual signal in the residual signal group;

[0110] The energy percentage of the virtual loudspeaker signal group is obtained based on the energy characterization values ​​of the virtual loudspeaker signal group and the residual signal group.

[0111] The encoding end first obtains the energy representation value of each virtual speaker signal in the virtual speaker signal group, and then adds the energy representation values ​​of all virtual speaker signals in the same group to obtain the energy representation value of the virtual speaker signal group. If there are multiple virtual speaker signal groups, the energy representation value of each group can be calculated in the above manner.

[0112] Similarly, the encoder can obtain the energy characterization value of the residual signal group based on the energy characterization value of each residual signal in the residual signal group. Finally, the encoder can obtain the energy proportion of the virtual loudspeaker signal group based on the energy characterization values ​​of the virtual loudspeaker signal group and the residual signal group. The energy proportion of the virtual loudspeaker signal group indicates its share in the total transmission channel signal energy. If the energy proportion of the virtual loudspeaker signal group is high, it means that the virtual loudspeaker signal group is dominant in the total transmission channel signal energy; if the energy proportion of the virtual loudspeaker signal group is low, it means that the virtual loudspeaker signal group is not dominant (i.e., weak) in the total transmission channel signal energy.

[0113] In some embodiments of this application, the transmission channel attribute information includes a virtual speaker code identifier, which indicates whether the bit allocation of a virtual speaker signal group is dominant. Specifically, the virtual speaker code identifier indicates whether the bit allocation of at least one virtual speaker signal group is dominant. For example, the virtual speaker code identifier can be represented as a flag. The virtual speaker code identifier can indicate whether the bit allocation of the virtual speaker signal group is dominant or not dominant. Different values ​​of the virtual speaker code identifier can indicate whether the bit allocation of the virtual speaker signal group is dominant or not dominant. Furthermore, the dominant situation can be further divided into strong dominance and secondary dominance (i.e., slight dominance).

[0114] Spatial encoding is performed on the three-dimensional audio signal to be encoded to obtain transmission channel attribute information, including:

[0115] Spatial encoding is performed on the three-dimensional audio signal to be encoded in order to obtain the number of dissimilar sound sources in the transmission channel signal and the encoding efficiency of the virtual loudspeaker.

[0116] The virtual speaker encoding identifier is obtained based on the number of dissimilar sound sources in the transmission channel signal and the virtual speaker encoding efficiency.

[0117] The encoding end, through spatial encoding, can classify the sound field of the transmission channel signal and generate a sound field classification result. This sound field classification result may include the number of dissimilar sound sources. The specific calculation process for the number of dissimilar sound sources is not limited here. The method for determining the virtual loudspeaker encoding efficiency is detailed in the foregoing embodiments and will not be repeated here. After obtaining the number of dissimilar sound sources and the virtual loudspeaker encoding efficiency of the transmission channel signal, the encoding end obtains the specific value of the virtual loudspeaker encoding identifier based on the decision conditions met by the number of dissimilar sound sources and the virtual loudspeaker encoding efficiency. In this application embodiment, there are multiple implementation methods for obtaining the virtual loudspeaker encoding identifier, as detailed in the examples of subsequent embodiments.

[0118] In some embodiments of this application, the virtual loudspeaker encoding identifier is further obtained based on the number of dissimilar sound sources in the transmission channel signal and the virtual loudspeaker encoding efficiency, including:

[0119] When the number of dissimilar sound sources in the transmission channel signal is less than or equal to a preset threshold for the number of dissimilar sound sources, and the virtual loudspeaker coding efficiency is greater than or equal to a preset first virtual loudspeaker coding efficiency threshold, the virtual loudspeaker coding identifier is determined to be dominant; or,

[0120] When the number of dissimilar sound sources in the transmission channel signal is greater than the preset threshold for the number of dissimilar sound sources, or when the virtual speaker coding efficiency is less than the preset first virtual speaker coding efficiency threshold, the virtual speaker coding is identified as not dominant.

[0121] In this embodiment, the specific implementation of the dissimilar sound source quantity threshold and the first virtual loudspeaker coding efficiency threshold can be tailored to the application scenario and is not limited here. For example, the dissimilar sound source quantity threshold can be represented as TH0, and the first virtual loudspeaker coding efficiency threshold can be represented as TH4.

[0122] Specifically, a dominant virtual speaker code indicates that the virtual speaker signal group is dominant in the total transmission channel signal. Therefore, this virtual speaker signal group needs to be allocated more bits, for example, by increasing the initial bit percentage of the virtual speaker signal group after determining it. Conversely, a non-dominant virtual speaker code indicates that the virtual speaker signal group is not dominant in the total transmission channel signal. In this case, fewer bits can be allocated to the virtual speaker signal group. For example, by decreasing the initial bit percentage of the virtual speaker signal group after determining it. In this embodiment, the encoding end can determine the virtual speaker code identifier by comparing the number of dissimilar sound sources, the virtual speaker coding efficiency, and the above-mentioned decision conditions. Therefore, the virtual speaker code identifier can be used to determine the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group.

[0123] Furthermore, in some embodiments of this application, the dominance includes secondary dominance or strong dominance; determining the virtual speaker encoding identifier as dominant includes:

[0124] When the virtual speaker encoding efficiency is greater than or equal to the first virtual speaker encoding efficiency threshold, and the virtual speaker encoding efficiency is less than or equal to the preset second virtual speaker encoding efficiency threshold, the virtual speaker encoding identifier is determined to be second dominant; or,

[0125] When the virtual speaker coding efficiency is greater than or equal to the first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is greater than the preset second virtual speaker coding efficiency threshold, the virtual speaker coding identifier is determined to be strongly dominant.

[0126] The second virtual speaker coding efficiency threshold is greater than the first virtual speaker coding efficiency threshold.

[0127] Specifically, when the number of dissimilar sound sources in the transmission channel signal is less than or equal to a preset threshold for the number of dissimilar sound sources, and the virtual speaker coding efficiency is greater than or equal to a preset first virtual speaker coding efficiency threshold, the virtual speaker coding identifier is determined to be dominant. The encoding end can further divide the case where the virtual speaker coding identifier is dominant, resulting in two cases: secondary dominance and strong dominance. It is understandable that if the virtual speaker coding identifier is strong dominance, the virtual speaker signal group needs to be allocated more bits; for example, after determining the initial bit percentage of the virtual speaker signal group, this bit percentage can be increased. If the virtual speaker coding identifier is secondary dominance, the virtual speaker signal group needs to be allocated fewer bits than when the virtual speaker coding identifier is strong dominance, but the number of bits allocated to the virtual speaker signal group still needs to be greater than when the virtual speaker coding identifier is not dominant; for example, after determining the initial bit percentage of the virtual speaker signal group, this bit percentage can be increased. Comparatively, the increased bit percentage in the strong dominance case is greater than the increased bit percentage in the secondary dominance case.

[0128] For example, the second virtual speaker coding efficiency threshold can be represented as TH2.

[0129] 402. Determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information.

[0130] In this embodiment, after acquiring the transmission channel signal and transmission channel attribute information, the encoding end can use the transmission channel attribute information, which carries attribute parameters of the transmission channel signal, to allocate bits to the virtual speaker signal group and the residual signal group. For example, the encoding end determines the bit allocation ratio of the virtual speaker signal group and the residual signal group based on the transmission channel attribute information. The bit allocation ratio refers to the proportion of bits allocated to a signal group to the total number of bits in the transmission channel signal; it can also be called the "bit allocation proportion." In this embodiment, the transmission channel signal includes at least one virtual speaker signal group and at least one residual signal group, thus the bit allocation ratio of the virtual speaker signal group and the residual signal group can be obtained. Subsequent embodiments will illustrate the process of determining the bit allocation ratio of one virtual speaker signal group and two residual signal groups as an example.

[0131] For example, in this embodiment of the application, spatial coding can output transmission channel signals and transmission channel attribute information. The core encoder obtains the transmission channel signals and transmission channel attribute information, and then the core encoder can obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group through the transmission channel signals and transmission channel attribute information.

[0132] In some embodiments of this application, the transmission channel attribute information includes: the energy percentage of the virtual speaker signal group, and / or the virtual speaker encoding identifier;

[0133] The bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group are determined based on the transmission channel attribute information, including:

[0134] When the energy proportion of the virtual speaker signal group is greater than or equal to the preset first energy proportion threshold, and / or the virtual speaker encoding identifier is strongly dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to the preset first signal group bit allocation algorithm.

[0135] When the energy proportion of the virtual speaker signal group is greater than or equal to the preset second energy proportion threshold and less than the preset first energy proportion threshold, and / or the virtual speaker encoding identifier is second dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to the preset second signal group bit allocation algorithm; wherein, the second energy proportion threshold is less than the first energy proportion threshold.

[0136] When the energy proportion of the virtual speaker signal group is less than the preset first energy proportion threshold, or when the virtual speaker encoding is not dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to the preset third signal group bit allocation algorithm.

[0137] In this embodiment, the encoding end can preset multiple signal group bit allocation algorithms. Under different conditions of the transmission channel attribute information, different signal group bit allocation algorithms can be used. Thus, when the transmission channel attribute information meets certain conditions, the bit allocation ratio of the virtual speaker signal group and the residual signal group can be allocated to the virtual speaker signal group and the residual signal group in accordance with these conditions. Therefore, the encoding efficiency of the encoding end for three-dimensional audio signals can be improved.

[0138] For example, the first energy percentage threshold can be represented as TH1, and the second energy percentage threshold can be represented as TH3.

[0139] In some embodiments of this application, when the energy proportion of the virtual speaker signal group is greater than or equal to a preset first energy proportion threshold, and / or the virtual speaker encoding identifier is strongly dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to a preset first signal group bit allocation algorithm, including:

[0140] When directionalNrgRatio≥TH1 and / or S≤TH0 and η>TH2, the bit allocation ratio of the virtual loudspeaker signal group is calculated as follows:

[0141] Ratio1_1=FAC1*directionalNrgRatio+(1–FAC1)*maxdirectionalNrgRatio;

[0142] Wherein, directionalNrgRatio represents the energy proportion of the virtual loudspeaker signal group, S is the number of dissimilar sound sources, η represents the virtual loudspeaker coding efficiency, maxdirectionalNrgRatio is the preset maximum virtual loudspeaker signal group bit allocation proportion, FAC1 is the preset first adjustment factor, Ratio1_1 is the bit allocation proportion of the virtual loudspeaker signal group, * represents multiplication operation, TH1 is the first energy proportion threshold, TH0 is the dissimilar sound source number threshold, and TH2 is the second virtual loudspeaker coding efficiency threshold.

[0143] The bit allocation percentage of the residual signal group is calculated as follows:

[0144] Ratio2 = 1 - Ratio1_1;

[0145] Where Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

[0146] As can be seen from the calculation process of Ratio1_1 above, the bit allocation ratio of the virtual speaker signal group is increased, so the encoding end can allocate more bits to the virtual speaker signal group.

[0147] The transmission channel signal includes a virtual speaker signal group and a residual signal group. After obtaining the bit allocation ratio Ratio1_1 of the virtual speaker signal group, the bit allocation ratio of the residual signal group can be obtained through the above formula for Ratio2.

[0148] It should be noted that in the embodiments of this application, FAC1 can be flexibly determined according to the specific application scenario, and no limitation is made here.

[0149] In some embodiments of this application, after obtaining the bit allocation ratio of the virtual speaker signal group, the method executed by the encoding end further includes:

[0150] The bit allocation percentage of the virtual speaker signal group is updated as follows:

[0151] Ratio1_2=min(Ratio1_1,maxdirectionalNrgRatio+FAC2*Ratio1_1)

[0152] Where Ratio1_2 represents the bit allocation ratio of the updated virtual speaker signal group, FAC2 is the preset second adjustment factor, maxdirectionalNrgRatio is the preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * indicates multiplication operation, and min is the minimum value operation.

[0153] It should be noted that in this embodiment, FAC2 can be flexibly determined according to the specific application scenario, and no limitation is made here.

[0154] As can be seen from the above calculation process of Ratio1_2, the bit allocation ratio of the virtual speaker signal group can be safely restricted, limiting Ratio1_2 to a safe bit range, thereby enabling the encoding end to safely and usably allocate bits of the virtual speaker signal group.

[0155] In some embodiments of this application, when the energy proportion of the virtual speaker signal group is greater than or equal to a preset second energy proportion threshold and less than a preset first energy proportion threshold, and / or the virtual speaker encoding identifier is subdominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to a preset second signal group bit allocation algorithm; wherein, the second energy proportion threshold being less than the first energy proportion threshold includes:

[0156] When TH3 ≤ directionalNrgRatio < TH1, and / or S ≤ TH0 and TH4 ≤ η ≤ TH2, Ratio1_1 is calculated as follows:

[0157] Ratio1_1=FAC3*directionalNrgRatio+(1–FAC3)*maxdirectionalNrgRatio;

[0158] Wherein, maxdirectionalNrgRatio is the preset bit allocation ratio of the virtual loudspeaker signal group, FAC3 is the preset third adjustment factor, directionalNrgRatio represents the energy ratio of the virtual loudspeaker signal group, S is the number of dissimilar sound sources, η represents the virtual loudspeaker coding efficiency, Ratio1_1 is the bit allocation ratio of the virtual loudspeaker signal group, * represents the multiplication operation, TH0 is the threshold for the number of dissimilar sound sources, TH1 is the first energy ratio threshold, TH2 is the second virtual loudspeaker coding efficiency threshold, TH3 is the second energy ratio threshold, and TH4 is the first virtual loudspeaker coding efficiency threshold.

[0159] The bit allocation percentage of the residual signal group is calculated as follows:

[0160] Ratio2 = 1 - Ratio1_1;

[0161] Where Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

[0162] It should be noted that in this embodiment, FAC3 can be flexibly determined according to the specific application scenario, and is not limited here. For example, 0≤FAC3≤0.5, FAC3>FAC1.

[0163] As can be seen from the calculation process of Ratio1_1 above, the bit allocation ratio of the virtual speaker signal group is increased, so the encoding end can allocate more bits to the virtual speaker signal group.

[0164] The transmission channel signal includes a virtual speaker signal group and a residual signal group. After obtaining the bit allocation ratio Ratio1_1 of the virtual speaker signal group, the bit allocation ratio of the residual signal group can be obtained through the above formula for Ratio2.

[0165] In some embodiments of this application, after obtaining the bit allocation ratio of the virtual speaker signal group, the method provided in the embodiments of this application further includes:

[0166] The bit allocation percentage of the virtual speaker signal group is updated as follows:

[0167] Ratio1_2=min(Ratio1_1,maxdirectionalNrgRatio+FAC4*Ratio1_1).

[0168] Where Ratio1_2 represents the bit allocation ratio of the updated virtual speaker signal group, FAC4 is the preset fourth adjustment factor, maxdirectionalNrgRatio is the preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * indicates multiplication operation, and min is the minimum value operation.

[0169] It should be noted that in this embodiment, FAC4 can be flexibly determined according to the specific application scenario, and no limitation is made here.

[0170] As can be seen from the above calculation process of Ratio1_2, the bit allocation ratio of the virtual speaker signal group can be safely restricted, limiting Ratio1_2 to a safe bit range, thereby enabling the encoding end to safely and usably allocate bits of the virtual speaker signal group.

[0171] In some embodiments of this application, the method provided in this application further includes:

[0172] There are multiple residual signal groups. The bit allocation ratio of the i-th residual signal group is calculated as follows:

[0173] Ratio2_i = Ratio2 * (R_i / C);

[0174] Where R_i represents the number of transmission channels included in the i-th residual signal group, C is the total number of transmission channels in all residual signal groups, Ratio2_i is the bit allocation ratio of the i-th residual signal group, * represents the multiplication operation, and Ratio2 is the bit allocation ratio of all residual signal groups.

[0175] When there are multiple residual signal groups, the proportion of bit allocation for each residual signal group in all residual signal groups can be determined based on the number of transmission channels for each residual signal group. For example, R_i / C represents the ratio of transmission channels for the i-th residual signal group to all residual signal groups, and the bit allocation proportion of the i-th residual signal group can be obtained through (R_i / C) and Ratio2.

[0176] In some embodiments of this application, when the energy proportion of the virtual speaker signal group is less than a preset first energy proportion threshold, or when the virtual speaker encoding identifier is not dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to a preset third signal group bit allocation algorithm, including:

[0177] When directionalNrgRatio < TH3, or S > TH0, or η < TH4, the bit allocation ratio of the virtual loudspeaker signal group is calculated as follows:

[0178] Ratio1_1 = directionalNrgRatio;

[0179] Where, directionalNrgRatio represents the energy ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, TH3 is the second energy ratio threshold, TH4 is the first virtual speaker coding efficiency threshold, S is the number of dissimilar sound sources, η represents the virtual speaker coding efficiency, and TH0 is the threshold of the number of dissimilar sound sources;

[0180] Calculate the bit allocation ratio of the residual signal group in the following way:

[0181] Ratio2_1 = D / (F + D);

[0182] Where, Ratio2_1 is the bit allocation ratio of the residual signal group, F represents the energy characterization value of the virtual speaker signal group, and D is the energy characterization value of the residual signal group.

[0183] From the above calculation process of Ratio1_1, it can be seen that the bit allocation ratio of the virtual speaker signal group is equal to the energy ratio of the virtual speaker signal group. Therefore, when the bit allocation of the virtual speaker signal group at the encoding end is not dominant, more bits will not be allocated to the virtual speaker signal group, thus ensuring the rationality of the bit allocation at the encoding end.

[0184] In some embodiments of the present application, the method provided by the embodiments of the present application further includes:

[0185] After obtaining the bit allocation ratio of the virtual speaker signal group, update the bit allocation ratio of the virtual speaker signal group in the following way:

[0186] When Ratio1_1 < groupBitsRatio1, Ratio1_2 = groupBitsRatio1;

[0187] When Ratio1_1 ≥ groupBitsRatio1, Ratio1_2 = FAC5 * groupBitsRatio1 + (1 - FAC5) * Ratio1_1;

[0188] Where, Ratio1_2 represents the updated bit allocation ratio of the virtual speaker signal group, FAC5 is a preset fifth adjustment factor, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before update, * represents multiplication operation, and groupBitsRatio1 is the preset bit allocation ratio of the virtual speaker signal group;

[0189] After obtaining the bit allocation ratio of the residual signal group, the bit allocation ratio of the residual signal group is updated as follows:

[0190] When Ratio2_1 < groupBitsRatio2, Ratio2_2 = groupBitsRatio2;

[0191] When Ratio2_1 ≥ groupBitsRatio2, Ratio2_2 = FAC6 * groupBitsRatio2 + (1 – FAC6) * Ratio2_1;

[0192] Among them, Ratio2_2 represents the bit allocation ratio of the updated residual signal group, FAC6 is a preset sixth adjustment factor, Ratio2_1 is the bit allocation ratio of the residual signal group before update, * represents multiplication operation, and groupBitsRatio2 is the preset bit allocation ratio of the residual signal group.

[0193] It should be noted that in the embodiments of the present application, FAC5 can be flexibly determined according to specific application scenarios and is not limited here.

[0194] From the above calculation process of Ratio1_2, it can be seen that the bit allocation ratio of the virtual speaker signal group can be safely restricted, and Ratio1_2 is restricted within the safe bit range, so that the encoding end can safely and usably perform the bit allocation of the virtual speaker signal group.

[0195] From the above calculation process of Ratio2_2, it can be seen that the bit allocation ratio of the residual signal group can be safely restricted, and Ratio2_2 is restricted within the safe bit range, so that the encoding end can safely and usably perform the bit allocation of the residual signal group.

[0196] In some embodiments of the present application, in addition to performing the foregoing method, the method provided in the embodiments of the present application further includes the following steps:

[0197] According to the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and the total number of transmission channel bits, respectively determine the number of bits of the virtual speaker signal group and the number of bits of the residual signal group;

[0198] Perform bit allocation for the virtual speaker signal group according to the number of bits of the virtual speaker signal group, and perform bit allocation for the residual signal group according to the number of bits of the residual signal group.

[0199] In this process, after obtaining the bit allocation ratios of the virtual speaker signal group and the residual signal group, the encoding end can perform bit allocation for each group separately to determine the bit allocation results. For example, the encoding end obtains the bit allocation ratios of the virtual speaker signal group and the residual signal group, and then, combined with the total number of bits in the transmission channel, determines the number of bits for each group. The number of bits for the virtual speaker signal group represents the actual number of bits that the encoding end can allocate to it, and the number of bits for the residual signal group represents the actual number of bits that the encoding end can allocate to it. Finally, the encoding end performs bit allocation for both the virtual speaker signal group and the residual signal group based on their respective bit counts, thus solving the problem of the encoding end being unable to allocate bits for the virtual speaker signal and the residual signal.

[0200] Furthermore, in some embodiments of this application, the number of bits in the virtual speaker signal group and the number of bits in the residual signal group are determined based on the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and the total number of transmission channel bits, including:

[0201] The number of bits in the virtual speaker signal group is calculated as follows:

[0202] F_bitnum = Ratio1 * C_bitnum;

[0203] Where F_bitnum is the number of bits in the virtual speaker signal group, Ratio1 is the bit allocation ratio of the virtual speaker signal group, and C_bitnum is the total number of bits in the transmission channel;

[0204] The number of bits in the residual signal group is calculated as follows:

[0205] D_bitnum = Ratio2 * C_bitnum;

[0206] Where D_bitnum is the number of bits in the residual signal group, Ratio2 is the bit allocation ratio of the residual signal group, and C_bitnum is the total number of bits in the transmission channel.

[0207] Specifically, the encoding end can predetermine the total number of transmission channel bits, without limiting the value of the total number of transmission channel bits. The encoding end can calculate the number of bits of the virtual speaker signal group and the number of bits of the residual signal group using the above calculation formula, thus realizing the bit allocation problem of the virtual speaker signal and the residual signal at the encoding end.

[0208] It is not limited to the above calculation formula, which is only one possible way to implement it and is not intended to limit the embodiments of this application. For example, the number of bits of the virtual speaker signal group and the number of bits of the residual signal group can be calculated by the above formula. The values ​​of the number of bits of the virtual speaker signal group and the number of bits of the residual signal group can also be adjusted by a preset adjustment factor to obtain the final values. The above calculation process is not limited.

[0209] In some embodiments of this application, in addition to performing the aforementioned steps, the method performed by the encoding end may also include the following steps:

[0210] The bit allocation ratios of the transmission channel signal, the virtual speaker signal group, and the residual signal group are encoded and written into the bitstream.

[0211] The bit allocation ratios of the virtual speaker signal group and the residual signal group can be encoded into the bitstream. After the encoding end sends the bitstream to the decoding end, the decoding end can obtain the bit allocation ratios of the virtual speaker signal group and the residual signal group by parsing the bitstream. The decoding end can then obtain the number of bits allocated to the virtual speaker signal group and the number of bits allocated to the residual signal group, thereby decoding the bitstream to obtain the three-dimensional audio signal.

[0212] In some embodiments of this application, the bit allocation ratios of the transmission channel signal, the virtual speaker signal group, and the residual signal group are encoded. Specifically, this can include directly encoding the transmission channel signal, or processing the transmission channel signal first, and then encoding the virtual speaker signal and residual signal after obtaining them. For example, the encoding end can be a core encoder, which encodes the virtual speaker signal, the residual signal, and the bit allocation ratios of the virtual speaker signal group and the residual signal group to obtain the bitstream. This bitstream can also be called an audio signal encoded bitstream.

[0213] The three-dimensional audio signal processing method provided in this application embodiment may include: an audio encoding method and an audio decoding method, wherein the audio encoding method is executed by an audio encoding device, the audio decoding method is executed by an audio decoding device, and the audio encoding device and the audio decoding device can communicate with each other. (The foregoing...) Figure 4 The processing method for three-dimensional audio signals, executed by the audio decoding device (hereinafter referred to as the decoding end) in the embodiments of this application, is described below. Figure 5 As shown, the main steps include the following:

[0214] 501. Receive bitstream.

[0215] The decoding end receives the bitstream from the encoding end. This bitstream carries the bit allocation percentages of the virtual speaker signal group and the residual signal group.

[0216] 502. Decode the bitstream to obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group.

[0217] The decoding end parses the bitstream and obtains the bit allocation ratios of the virtual speaker signal group and the residual signal group from it. These bit allocation ratios are determined by the encoding end according to the aforementioned... Figure 4 The example shown is obtained.

[0218] 503. Decode the virtual speaker signal and residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain the decoded three-dimensional audio signal.

[0219] After obtaining the bit allocation ratios of the virtual speaker signal group and the residual signal group, the decoding end uses these ratios to parse the bitstream and obtain the decoded 3D audio signal. In this embodiment, the decoding process for the virtual speaker signal and residual signal in the bitstream is not limited. In this embodiment, the decoding end can determine the number of bits allocated to the virtual speaker signal and the number of bits allocated to the residual signal using the bit allocation ratios of the virtual speaker signal group and the residual signal group. The decoding end uses a decoding method corresponding to the encoding method of the encoding end to perform decoding, thereby obtaining the 3D audio signal sent by the encoding end and realizing the transmission of the 3D audio signal from the encoding end to the decoding end.

[0220] For example, the decoding end can determine the number of bits allocated to the virtual speaker signal and the number of bits allocated to the residual signal based on the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group transmitted in the bit stream, thus solving the problem that the decoding end cannot determine the allocated bits of the signal.

[0221] In some embodiments of this application, step 503 decodes the virtual speaker signal and residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group, including:

[0222] The number of available bits is determined based on the bitstream;

[0223] The number of bits in the virtual speaker signal group is determined based on the available number of bits and the bit allocation ratio of the virtual speaker signal group; the virtual speaker signal in the bitstream is decoded based on the number of bits in the virtual speaker signal group;

[0224] The number of bits in the residual signal group is determined based on the number of available bits and the bit allocation ratio of the residual signal group; the residual signal in the bit stream is decoded based on the number of bits in the residual signal group.

[0225] The decoding end first determines the number of available bits, which is the total number of bits that can be allocated to the transmission channel. By parsing the bitstream, the decoding end can obtain the bit allocation ratio of the virtual speaker signal group. Based on the available bits and the bit allocation ratio of the virtual speaker signal group, the decoding end can determine the number of bits for the virtual speaker signal group. This number of bits for the virtual speaker signal group is the number of bits used by the encoding end when encoding the virtual speaker signal group. The decoding end can also decode the virtual speaker signal in the bitstream based on the number of bits for the virtual speaker signal group, thus decoding the virtual speaker signal from the bitstream.

[0226] Similarly, the decoding end can obtain the bit allocation ratio of the residual signal group by parsing the bit stream. Then, based on the number of available bits and the bit allocation ratio of the residual signal group, the number of bits of the residual signal group can be determined. The number of bits of the residual signal group is the number of bits used by the encoding end when encoding the residual signal group. The decoding end can also decode the residual signal in the bit stream based on the number of bits of the residual signal group, so that the decoding end can decode the residual signal from the bit stream.

[0227] For example, during the decoding process at the decoding end, the following two parameters can be parsed from the bitstream: `groupBitsRatio` and `bitsRatio`. `groupBitsRatio` occupies 4 bits and represents the inter-group bit allocation ratio parameter, which includes the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group. `bitsRatio` occupies 4 bits and represents the intra-group bit allocation ratio parameter, which includes the bit allocation ratio of each virtual speaker signal group within all virtual speaker signal groups and the bit allocation ratio of each residual signal group within all residual signal groups.

[0228] For example, the decoding end may include a bit allocation module. The main function of this bit allocation module is to allocate the remaining available bits after removing other side information to each transmission channel according to the bit allocation ratio parameter obtained from decoding in the bit stream. The encoding of other side information will also occupy bits.

[0229] First, we need to calculate the number of available bits remaining in the current frame after deducting other side information, denoted as availableBits.

[0230] The general algorithm for calculating availableBits is expressed as follows:

[0231] availableBits=bitsPerFrame-bitsUsed;

[0232] Where bitsPerFrame is the initial number of bits per frame, and bitsUsed is the number of bits already used before bit allocation.

[0233] The calculation process for HOA bit allocation HoaSplitBytesGroup() is as follows.

[0234] First, calculate the number of bits (groupBytes) for each channel based on the total available bits (availableBits) and groupBitsRatio, as follows:

[0235]

[0236] in, It can represent the bit allocation percentage of the virtual speaker signal group in all transmission channel signals, or it can represent the bit allocation percentage of the residual signal group in all transmission channel signals.

[0237] Then, the number of bits per channel (bytesChannels) is calculated based on bitsRatio, as follows:

[0238]

[0239] For example, groupBytes represents the total number of allocated bits for a virtual speaker signal group.

[0240] This indicates the percentage of bits allocated to each virtual speaker signal group within all virtual speaker signal groups, while bytesChannels indicates the number of bits in each virtual speaker signal group.

[0241] For example, groupBytes represents the total number of allocated bits for the residual signal group.

[0242] This indicates the percentage of bits allocated to each residual signal group within all residual signal groups, while bytesChannels indicates the number of bits in each residual signal group.

[0243] The number of bits in each channel can be calculated through the above process.

[0244] It should be noted that the decoding end can also calculate the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group in a similar way to the encoding end, for example, by using the aforementioned calculation process of Ratio1 and Ratio2, which will not be repeated here.

[0245] To facilitate a better understanding and implementation of the above-described solutions in the embodiments of this application, specific examples of corresponding application scenarios are provided below.

[0246] In this embodiment of the application, taking the three-dimensional audio signal as the HOA signal as an example, this embodiment of the application provides a bit allocation method for virtual speaker signals and residual signals. First, the virtual speaker signals and residual signals are grouped. Then, the bit allocation ratio between groups is obtained according to the signal characteristics and sound field characteristics. Finally, channel bit allocation is realized.

[0247] The purpose of this application embodiment is to obtain the bit allocation result of the transmission channel signal, which consists of a virtual speaker signal and a residual signal. This application embodiment first groups the transmission channel signal into a virtual speaker signal group and a residual signal group.

[0248] The inter-group bit allocation ratio is obtained based on signal characteristics and sound field characteristics, and then the number of bits in the virtual speaker signal group and the number of bits in the residual signal group are obtained from the total number of bits. When the encoder encodes at a certain rate, the total number of bits allocated to each frame is fixed. In this embodiment, the bits are allocated based on the available bits in that frame. For example, in constant bitrate (CBR) mode, the bitrate is 384kbps, and the number of bits in each frame is approximately 7680 bits. The actual number of available bits is less than 7680 bits, and in this embodiment, these bits less than 7680 bits can be allocated.

[0249] When the virtual loudspeaker encoding efficiency is high, for example, when the number of dissimilar sound sources is less than or equal to the number of transmission channels of the virtual loudspeaker signal, the number of encoded bits of the virtual loudspeaker signal can be increased by increasing the inter-group bit allocation ratio of the virtual loudspeaker signal group.

[0250] In the above calculation method, the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal can conform to the actual situation of the sound field classification of the current frame, thus solving the problem of needing to determine the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal when encoding the current frame.

[0251] In this embodiment of the application, the execution flow of the core codec will be described below.

[0252] Please see Figure 6 As shown, the specific implementation steps are given below:

[0253] S1. The HOA signal to be encoded is spatially encoded using HOA to obtain the transmission channel signal and attribute information.

[0254] The transmission channel signals include: virtual speaker signals and residual signals;

[0255] The attribute information is the aforementioned single-channel transmission attribute information, including sound field classification results and virtual loudspeaker coding efficiency η.

[0256] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; the virtual loudspeaker coding efficiency η represents the efficiency of reconstructing the HOA signal using a virtual loudspeaker in the current frame.

[0257] The following is a method for calculating the encoding efficiency of virtual loudspeakers:

[0258] Calculate the energy representation values ​​R1, R2, ..., Rt for each channel of the reconstructed HOA signal, where Rt = norm(SRt), norm() is the norm operation, SRt is the improved discrete cosine transform (MDCT) coefficient of the t-th channel of the reconstructed HOA signal, and t is (HOA order + 1). 2 .

[0259] Calculate the energy representation values ​​N1, N2, ..., Nt of the original HOA signal, where Nt = norm(SNt), norm() is the norm operation, SNt is the MDCT coefficient of the t-th channel of the original HOA signal, and t is (HOA order + 1). 2 .

[0260] The virtual loudspeaker encoding efficiency η = sum(R) / sum(N); sum(R) represents the summation of R1 to Rt, and sum(N) represents the summation of N1 to Nt.

[0261] S2. Obtain the percentage of packet bit allocation in the transmission channel.

[0262] First, the transmission channel signals are grouped, assuming they consist of M virtual speaker signals and N residual signals. Further, the N residual signals can be divided into K groups. If the M virtual speaker signals are divided into one group, then the transmission channel is divided into K+1 groups. The number of channels in each group can be the same or different, and the grouping in each frame can be the same or different; neither affects the subsequent processes in this embodiment.

[0263] The following example uses K equal to 2. However, K can also be 3 or other values.

[0264] Taking a transmission channel of 11 as an example, the virtual speaker signal group contains 2 virtual speakers, the residual signal group 1 contains 4 residual signals, and the residual signal group 2 contains 5 residual signals.

[0265] Step S2 includes the following steps S21 to S23.

[0266] S21. Calculate the energy characterization value for each group.

[0267] The energy characterization value of each channel can be calculated using the method in S1. Then, the energy characterization values ​​of the channels in each group are added together to obtain the energy characterization value of each group. For example, the energy characterization value of the virtual speaker signal group is F, the energy characterization value of residual signal group 1 is D1, and the energy characterization value of residual signal group 2 is D2.

[0268] S22. Calculate the energy ratio of the virtual loudspeaker signal group: directionalNrgRatio.

[0269] directionalNrgRatio=F / (F+D1+D2).

[0270] S23. Determine the bit allocation ratio between transmission channel groups.

[0271] The bit allocation ratio between transmission channel groups is determined based on at least one of the following: the virtual speaker signal group energy ratio (directionalNrgRatio) and / or the virtual speaker coding flag (Flag). Assuming the virtual speaker signal group bit allocation ratio is Ratio1, the residual signal group 1 bit allocation ratio is Ratio2, and the residual signal group 2 bit allocation ratio is Ratio3, when the current frame's virtual speaker signal group bit allocation is determined to be dominant based on the virtual speaker signal group energy ratio (directionalNrgRatio) and / or the virtual speaker coding efficiency (η), the virtual speaker signal group bit allocation ratio needs to be increased, and the residual signal group bit allocation ratio needs to be decreased. Different adjustment methods can be selected to increase the virtual speaker signal group bit allocation ratio under different preset conditions.

[0272] The judgment criteria include the speaker signal group energy ratio (directionalNrgRatio) and / or the virtual speaker encoding flag.

[0273] The virtual speaker encoding identifier Flag is obtained through the following method:

[0274] When the number of dissimilar sound sources is ≤ TH0 and the virtual loudspeaker coding efficiency η > TH2, Flag = strong dominance (High).

[0275] When the number of dissimilar sound sources is ≤ TH0 and the virtual loudspeaker coding efficiency is TH4 ≤ η ≤ TH2, Flag = Middle Dominant. Otherwise, Flag = Low Dominant.

[0276] The following examples illustrate the above judgment conditions. For instance, the judgment conditions may include conditions 1 to 6.

[0277] Condition 1: When directionalNrgRatio≥TH1 is satisfied, 0.9≤TH1≤1, for example, TH1=0.9375.

[0278] First, calculate the virtual speaker signal group bit allocation ratio Ratio1:

[0279] Ratio1=FAC1*directionalNrgRatio+(1–FAC1)*maxdirectionalNrgRatio.

[0280] Where maxdirectionalNrgRatio is the preset maximum virtual speaker signal group bit allocation ratio, and FAC1 is the preset first adjustment factor, 0≤FAC1≤0.5.

[0281] Optionally, restrict the security bits for Ratio1, for example:

[0282] Ratio1=min(Ratio1,maxdirectionalNrgRatio+FAC2*Ratio1).

[0283] FAC2 is a preset second adjustment factor, where 0 ≤ FAC2 ≤ 0.5.

[0284] Then, calculate the allocation ratio of 1 bit of the residual signal group Ratio2 and the allocation ratio of 2 bits of the residual signal group Ratio3:

[0285] Ratio2 = (1 - Ratio1) * number of channels in residual signal group 1 / (number of channels in residual signal group 1 + number of channels in residual signal group 2);

[0286] Ratio3 = (1 - Ratio1) * number of channels in residual signal group 2 / (number of channels in residual signal group 1 + number of channels in residual signal group 2).

[0287] Condition 2: When the number of dissimilar sound sources is ≤ TH0 and the virtual loudspeaker coding efficiency η > TH2, i.e., Flag = High, TH0 is the number of virtual loudspeakers matched by the codec or the number of virtual loudspeaker signals in the codec. For example, TH0 = 2. 0.8 ≤ TH1 ≤ 1, for example, TH2 = 0.875. It can be considered that the bit allocation of the virtual loudspeaker signal group is strongly dominant. At this time, the bit allocation ratio between transmission channel groups is adjusted as follows:

[0288] The steps for calculating Ratio1, Ratio2, and Ratio3 are the same as those for condition 1.

[0289] Condition 3: When TH3≤directionalNrgRatio<TH1 is satisfied, 0.5≤TH3<0.9, for example, TH3=0.75.

[0290] First, calculate the virtual speaker signal group bit allocation ratio Ratio1:

[0291] Ratio1=FAC3*directionalNrgRatio+(1–FAC3)*maxdirectionalNrgRatio.

[0292] Where maxdirectionalNrgRatio is the preset bit allocation ratio of the virtual speaker signal group, FAC3 is the preset third adjustment factor, 0≤FAC3≤0.5; FAC3>FAC1.

[0293] Optionally, restrict the security bits for Ratio1, for example:

[0294] Ratio1=min(Ratio1,maxdirectionalNrgRatio+TH8FAC4*Ratio1).

[0295] Wherein, FAC4 is the preset fourth adjustment factor, 0≤FAC4≤0.5, FAC4<FAC2;

[0296] Then, calculate the allocation ratio of 1 bit of the residual signal group Ratio2 and the allocation ratio of 2 bits of the residual signal group Ratio3:

[0297] Ratio2 = (1 - Ratio1) * number of channels in residual signal group 1 / (number of channels in residual signal group 1 + number of channels in residual signal group 2);

[0298] Ratio3 = (1 - Ratio1) * number of channels in residual signal group 2 / (number of channels in residual signal group 1 + number of channels in residual signal group 2).

[0299] Condition 4: When the number of dissimilar sound sources ≤ TH0 and the virtual loudspeaker coding efficiency TH4 ≤ η ≤ TH2, i.e., when Flag = Middle, 0.5 ≤ TH4 < 0.8, for example, TH4 = 0.6875. It can be considered that the virtual loudspeaker signal group bit allocation is slightly advantageous. In this case, the bit allocation ratio between transmission channel groups is adjusted as follows:

[0300] The steps for calculating Ratio1, Ratio2, and Ratio3 are the same as those for condition 3.

[0301] Condition 5: When directionalNrgRatio < TH3, it can be considered that the residual group bit allocation is dominant. At this time, the following adjustments are made to the bit allocation ratio between transmission channel groups:

[0302] Ratio1 = directionalNrgRatio.

[0303] Ratio2 = D1 / (F + D1 + D2).

[0304] Ratio3 = D2 / (F + D1 + D2).

[0305] Optionally, safety bits are restricted for Ratio1, Ratio2, and Ratio3. For example:

[0306] When Ratio1 < groupBitsRatio1, Ratio1 = groupBitsRatio1;

[0307] When Ratio1 ≥ groupBitsRatio1, Ratio1 = FAC5 * groupBitsRatio1 + (1 – FAC5) * Ratio1;

[0308] When Ratio2 < groupBitsRatio2, Ratio2 = groupBitsRatio2;

[0309] When Ratio2 ≥ groupBitsRatio2, Ratio2 = FAC6 * groupBitsRatio2 + (1 – FAC6) * Ratio2;

[0310] When Ratio3 < groupBitsRatio3, Ratio3 = groupBitsRatio3;

[0311] When Ratio3 ≥ groupBitsRatio3, Ratio3 = FAC7 * groupBitsRatio3 + (1 – FAC7) * Ratio3;

[0312] Wherein, groupBitsRatio1, groupBitsRatio2, and groupBitsRatio3 are the preset bit allocation ratios for virtual speaker signal groups, the preset bit allocation ratios for residual signal group 1, and the preset bit allocation ratios for residual signal group 2, respectively. FAC5 is the preset fifth adjustment factor, 0.5 < FAC5 ≤ 1; FAC6 is the preset sixth adjustment factor, 0.5 < FAC6 ≤ 1; and FAC7 is the preset seventh adjustment factor, 0.5 < FAC7 ≤ 1. FAC5, FAC6, and FAC7 can be equal or unequal.

[0313] Condition 6: When the number of dissimilar sound sources > TH0, or the virtual loudspeaker coding efficiency η < TH4, i.e., Flag = Low, the residual group bit allocation can be considered dominant. In this case, the bit allocation ratio between transmission channel groups is adjusted as follows:

[0314] The steps for calculating Ratio1, Ratio2, and Ratio3 are the same as those for condition 5.

[0315] After obtaining Ratio1, Ratio2, and Ratio3, Ratio1, Ratio2, and Ratio3 can be quantized and written into the bitstream.

[0316] S3. Downmix the transmission channel signal.

[0317] The specific process of downmixing the transmission channel signal will not be described further. The original channel signal is used to calculate the downmixed channel using a downmixing algorithm, and then bit allocation is performed. Step S3 is optional, and the execution order of step S3 can be before or after step S2.

[0318] S4. Perform bit allocation on the transmission channel signal.

[0319] First, the number of bits in each group is determined by the inter-group bit allocation ratio in step S2 and the total number of available bits, for example:

[0320] Virtual speaker signal group bit count = Ratio1 * total available bits.

[0321] The number of bits in the residual signal group 1 = Ratio2 * total available bits.

[0322] The number of bits in the residual signal group is equal to Ratio3 * the total number of available bits.

[0323] Then, the number of bits for each channel is determined, which can be done in several ways, such as allocating bits according to the energy proportion of each channel.

[0324] The signal decoding process executed at the decoding end will be explained next.

[0325] The decoding end receives the bitstream sent by the encoding end, and then parses Ratio1, Ratio2, and Ratio3 from the bitstream. Then, it can perform bit allocation on the transmission channel signal. For example, bit allocation on the transmission channel signal can be done using the method of obtaining the number of bits for each channel in the aforementioned step S4.

[0326] As illustrated by the foregoing examples, the encoding end of this application can group the transmission channels and determine the bit allocation ratio of each group based on the virtual loudspeaker signal group energy, the number of dissimilar sound sources, and the reconstructed HOA signal. This application embodiment can adjust the inter-group allocation ratio through the aforementioned conditions. Therefore, this application embodiment can effectively improve the bit allocation efficiency of the transmission channels.

[0327] The decoding process executed by the decoding end will not be described in detail in this embodiment.

[0328] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0329] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.

[0330] Please see Figure 7 As shown in the embodiment of this application, a three-dimensional audio signal processing device is provided. For example, the three-dimensional audio signal processing device is specifically an audio encoding device 700, which may include: an encoding module 701 and a bit allocation ratio determination module 702, wherein...

[0331] The encoding module is used to spatially encode the three-dimensional audio signal to be encoded to obtain the transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group;

[0332] The bit allocation ratio determination module is used to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information.

[0333] Please see Figure 8As shown in the embodiment of this application, a three-dimensional audio signal processing device is provided. For example, the three-dimensional audio signal processing device is specifically an audio decoding device 800, which may include: a receiving module 801, a decoding module 802, and a signal generation module 803, wherein...

[0334] The receiving module is used to receive the bitstream;

[0335] The decoding module is used to decode the bitstream to obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group;

[0336] The signal generation module is used to decode the virtual speaker signal and the residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain the decoded three-dimensional audio signal.

[0337] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.

[0338] This application also provides a computer storage medium storing a program that performs some or all of the steps described in the above method embodiments.

[0339] The following describes another audio encoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 9 As shown, the audio encoding device 900 includes:

[0340] Receiver 901, transmitter 902, processor 903, and memory 904 (wherein the audio encoding device 900 may contain one or more processors 903). Figure 9 (Taking a processor as an example). In some embodiments of this application, the receiver 901, transmitter 902, processor 903, and memory 904 can be connected via a bus or other means, wherein... Figure 9 Taking the example of a connection between China and Israel via a bus.

[0341] Memory 904 may include read-only memory and random access memory, and provides instructions and data to processor 903. A portion of memory 904 may also include non-volatile random access memory (NVRAM). Memory 904 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.

[0342] Processor 903 controls the operation of the audio encoding device; processor 903 can also be called a central processing unit (CPU). In specific applications, the various components of the audio encoding device are coupled together through a bus system. This bus system includes not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.

[0343] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 903. The processor 903 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the processor 903. The processor 903 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 904. Processor 903 reads the information in memory 904 and, in conjunction with its hardware, completes the steps of the above method.

[0344] The receiver 901 can be used to receive input digital or character information and generate signal inputs related to the settings and function control of the audio encoding device. The transmitter 902 may include a display device such as a display screen and can be used to output digital or character information through an external interface.

[0345] In this embodiment, processor 903 is used to execute the aforementioned embodiments. Figure 4 The method shown is performed by the audio encoding device.

[0346] The following describes another audio decoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 10 As shown, the audio decoding device 1000 includes:

[0347] Receiver 1001, transmitter 1002, processor 1003, and memory 1004 (wherein the audio decoding device 1000 may contain one or more processors 1003). Figure 10 (Taking a processor as an example). In some embodiments of this application, the receiver 1001, transmitter 1002, processor 1003, and memory 1004 can be connected via a bus or other means, wherein... Figure 10 Taking the example of a connection between China and Israel via a bus.

[0348] Memory 1004 may include read-only memory and random access memory, and provides instructions and data to processor 1003. A portion of memory 1004 may also include NVRAM. Memory 1004 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.

[0349] Processor 1003 controls the operation of the audio decoding device; processor 1003 can also be referred to as a CPU. In specific applications, the various components of the audio decoding device are coupled together through a bus system. This bus system includes not only a data bus but also a power bus, control bus, and status signal bus, etc. However, for clarity, all buses in the diagram are referred to as the bus system.

[0350] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1003. The processor 1003 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1003 or by instructions in the form of software. The processor 1003 can be a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers and other mature storage media in the art. The storage medium is located in memory 1004, and the processor 1003 reads the information in memory 1004 and completes the steps of the above method in combination with its hardware.

[0351] In this embodiment, processor 1003 is used to execute the aforementioned embodiments. Figure 5 The method shown is performed by the audio decoding device.

[0352] In another possible design, when the audio encoding or decoding device is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer-executable instructions stored in the storage unit to cause the chip within the terminal to execute the audio encoding method of any of the first aspects or the audio decoding method of any of the second aspects described above. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the terminal, such as read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0353] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of programs in the first or second aspect of the above methods.

[0354] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0355] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0356] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0357] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method for processing three-dimensional audio signals, characterized in that, include: Spatial encoding is performed on the three-dimensional audio signal to be encoded to obtain transmission channel signals and transmission channel attribute information, wherein the transmission channel signals include: at least one virtual speaker signal group and at least one residual signal group; The bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group are determined based on the transmission channel attribute information.

2. The method according to claim 1, characterized in that, The transmission channel attribute information includes: virtual speaker encoding efficiency; The spatial encoding of the three-dimensional audio signal to be encoded to obtain transmission channel attribute information includes: A virtual loudspeaker is used to reconstruct the three-dimensional audio signal to be encoded, so as to obtain the reconstructed three-dimensional audio signal; Obtain the energy characterization value of the reconstructed three-dimensional audio signal and the energy characterization value of the three-dimensional audio signal to be encoded; The encoding efficiency of the virtual speaker is obtained based on the energy characterization value of the reconstructed three-dimensional audio signal and the energy characterization value of the three-dimensional audio signal to be encoded.

3. The method according to claim 1 or 2, characterized in that, The transmission channel attribute information includes: the energy percentage of the virtual speaker signal group; The method further includes: The energy characterization value of the virtual speaker signal group is obtained based on the energy characterization value of each virtual speaker signal in the virtual speaker signal group; The energy characterization value of the residual signal group is obtained based on the energy characterization value of each residual signal in the residual signal group; The energy percentage of the virtual loudspeaker signal group is obtained based on the energy characterization value of the virtual loudspeaker signal group and the energy characterization value of the residual signal group.

4. The method according to claim 1, characterized in that, The transmission channel attribute information includes: a virtual speaker encoding identifier, which is used to indicate whether the bit allocation of the virtual speaker signal group is dominant; The spatial encoding of the three-dimensional audio signal to be encoded to obtain transmission channel attribute information includes: The three-dimensional audio signal to be encoded is spatially encoded to obtain the number of dissimilar sound sources and the encoding efficiency of the virtual speaker in the transmission channel signal. The virtual speaker encoding identifier is obtained based on the number of dissimilar sound sources in the transmission channel signal and the virtual speaker encoding efficiency.

5. The method according to claim 4, characterized in that, The step of obtaining the virtual speaker encoding identifier based on the number of dissimilar sound sources in the transmission channel signal and the virtual speaker encoding efficiency includes: When the number of dissimilar sound sources in the transmission channel signal is less than or equal to a preset threshold for the number of dissimilar sound sources, and the virtual speaker coding efficiency is greater than or equal to a preset first virtual speaker coding efficiency threshold, the virtual speaker coding identifier is determined to be dominant; or When the number of dissimilar sound sources in the transmission channel signal is greater than a preset threshold for the number of dissimilar sound sources, or when the virtual speaker coding efficiency is less than a preset first virtual speaker coding efficiency threshold, the virtual speaker coding is determined to be non-dominant.

6. The method according to claim 5, characterized in that, The dominance includes secondary dominance or strong dominance; The step of determining that the virtual speaker encoding identifier is dominant includes: When the virtual speaker encoding efficiency is greater than or equal to the first virtual speaker encoding efficiency threshold, and the virtual speaker encoding efficiency is less than or equal to a preset second virtual speaker encoding efficiency threshold, the virtual speaker encoding identifier is determined to be second dominant; or When the virtual speaker coding efficiency is greater than or equal to the first virtual speaker coding efficiency threshold and the virtual speaker coding efficiency is greater than the preset second virtual speaker coding efficiency threshold, the virtual speaker coding identifier is determined to be strongly dominant. Wherein, the second virtual speaker coding efficiency threshold is greater than the first virtual speaker coding efficiency threshold.

7. The method according to any one of claims 1 to 2, characterized in that, The transmission channel attribute information includes: the energy percentage of the virtual speaker signal group, and / or the virtual speaker encoding identifier; Determining the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information includes: When the energy proportion of the virtual speaker signal group is greater than or equal to the preset first energy proportion threshold, and / or the virtual speaker encoding identifier is strongly dominant, the bit allocation proportion of the virtual speaker signal group and the bit allocation proportion of the residual signal group are determined according to the preset first signal group bit allocation algorithm. When the energy percentage of the virtual speaker signal group is greater than or equal to a preset second energy percentage threshold and less than a preset first energy percentage threshold, and / or the virtual speaker encoding identifier is subdominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset second signal group bit allocation algorithm; wherein, the second energy percentage threshold is less than the first energy percentage threshold; or When the energy percentage of the virtual speaker signal group is less than the preset first energy percentage threshold, or when the virtual speaker encoding identifier is not dominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to the preset third signal group bit allocation algorithm.

8. The method according to claim 7, characterized in that, When the energy percentage of the virtual speaker signal group is greater than or equal to a preset first energy percentage threshold, and / or the virtual speaker encoding identifier is strongly dominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset first signal group bit allocation algorithm, including: When directionalNrgRatio≥TH1 and / or S≤TH0 and η>TH2, the bit allocation ratio of the virtual speaker signal group is calculated as follows: Ratio1_1=FAC1*directionalNrgRatio+(1–FAC1)*maxdirectionalNrgRatio; Wherein, directionalNrgRatio represents the energy proportion of the virtual speaker signal group, S is the number of dissimilar sound sources, η represents the virtual speaker coding efficiency, maxdirectionalNrgRatio is the preset maximum virtual speaker signal group bit allocation proportion, FAC1 is the preset first adjustment factor, Ratio1_1 is the bit allocation proportion of the virtual speaker signal group, * represents multiplication operation, TH1 is the first energy proportion threshold, TH0 is the dissimilar sound source number threshold, and TH2 is the second virtual speaker coding efficiency threshold; The bit allocation percentage of the residual signal group is calculated as follows: Ratio2 = 1 - Ratio1_1; Wherein, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

9. The method according to claim 8, characterized in that, After obtaining the bit allocation ratio of the virtual speaker signal group, the method further includes: The bit allocation ratio of the virtual speaker signal group is updated in the following manner: Ratio1_2=min(Ratio1_1,maxdirectionalNrgRatio+FAC2*Ratio1_1) Wherein, Ratio1_2 represents the bit allocation ratio of the updated virtual speaker signal group, FAC2 is a preset second adjustment factor, maxdirectionalNrgRatio is a preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * represents multiplication operation, and min represents minimum value operation.

10. The method according to claim 7, characterized in that, When the energy percentage of the virtual speaker signal group is greater than or equal to a preset second energy percentage threshold and less than a preset first energy percentage threshold, and / or the virtual speaker encoding identifier is second dominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset second signal group bit allocation algorithm; wherein, the second energy percentage threshold being less than the first energy percentage threshold includes: When TH3 ≤ directionalNrgRatio < TH1, and / or S ≤ TH0 and TH4 ≤ η ≤ TH2, Ratio1_1 is calculated as follows: Ratio1_1=FAC3*directionalNrgRatio+(1–FAC3)*maxdirectionalNrgRatio; Wherein, maxdirectionalNrgRatio is a preset bit allocation ratio for virtual speaker signal groups, FAC3 is a preset third adjustment factor, directionalNrgRatio represents the energy ratio of the virtual speaker signal group, S is the number of dissimilar sound sources, η represents the virtual speaker coding efficiency, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, * represents multiplication operation, TH0 is the threshold for the number of dissimilar sound sources, TH1 is the first energy ratio threshold, TH2 is the second virtual speaker coding efficiency threshold, TH3 is the second energy ratio threshold, and TH4 is the first virtual speaker coding efficiency threshold. The bit allocation percentage of the residual signal group is calculated as follows: Ratio2 = 1 - Ratio1_1; Wherein, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group, and Ratio2 is the bit allocation ratio of the residual signal group.

11. The method according to claim 10, characterized in that, After obtaining the bit allocation ratio of the virtual speaker signal group, the method further includes: The bit allocation ratio of the virtual speaker signal group is updated in the following manner: Ratio1_2=min(Ratio1_1,maxdirectionalNrgRatio+FAC4*Ratio1_1) Wherein, Ratio1_2 represents the bit allocation ratio of the updated virtual speaker signal group, FAC4 is the preset fourth adjustment factor, maxdirectionalNrgRatio is the preset maximum bit allocation ratio of the virtual speaker signal group, Ratio1_1 is the bit allocation ratio of the virtual speaker signal group before the update, * represents multiplication operation, and min is the minimum value operation.

12. The method according to any one of claims 8 to 11, characterized in that, The method further includes: There are multiple residual signal groups, and the bit allocation ratio of the i-th residual signal group is calculated as follows: Ratio2_i = Ratio2 * (R_i / C); Wherein, R_i represents the number of transmission channels included in the i-th residual signal group, C is the total number of transmission channels in all residual signal groups, Ratio2_i is the bit allocation ratio of the i-th residual signal group, * represents multiplication operation, and Ratio2 is the bit allocation ratio of all residual signal groups.

13. The method according to claim 7, characterized in that, When the energy percentage of the virtual speaker signal group is less than a preset first energy percentage threshold, or the virtual speaker encoding identifier is not dominant, the bit allocation percentage of the virtual speaker signal group and the bit allocation percentage of the residual signal group are determined according to a preset third signal group bit allocation algorithm, including: When directionalNrgRatio < TH3, or S > TH0, or η < TH4, the bit allocation ratio of the virtual speaker signal group is calculated as follows: Ratio1_1 = directionalNrgRatio; Where, the directionalNrgRatio represents the energy proportion of the virtual speaker signal group, Ratio1_1 is the bit allocation proportion of the virtual speaker signal group, TH3 is the second energy proportion threshold, TH4 is the first virtual speaker coding efficiency threshold, S is the number of different sound sources, η represents the virtual speaker coding efficiency, and TH0 is the threshold of the number of different sound sources; Calculate the bit allocation proportion of the residual signal group in the following way: Ratio2_1 = D / (F + D); Where, Ratio2_1 is the bit allocation proportion of the residual signal group, F represents the energy characterization value of the virtual speaker signal group, and D is the energy characterization value of the residual signal group.

14. The method according to claim 13, characterized in that, The method further includes: After obtaining the bit allocation proportion of the virtual speaker signal group, update the bit allocation proportion of the virtual speaker signal group in the following way: When Ratio1_1 < groupBitsRatio1, Ratio1_2 = groupBitsRatio1; When Ratio1_1 ≥ groupBitsRatio1, Ratio1_2 = FAC5 * groupBitsRatio1 + (1 – FAC5) * Ratio1_1; Where, Ratio1_2 represents the updated bit allocation proportion of the virtual speaker signal group, FAC5 is a preset fifth adjustment factor, Ratio1_1 is the bit allocation proportion of the virtual speaker signal group before update, * represents multiplication operation, and groupBitsRatio1 is the preset bit allocation proportion of the virtual speaker signal group; After obtaining the bit allocation proportion of the residual signal group, update the bit allocation proportion of the residual signal group in the following way: When Ratio2_1 < groupBitsRatio2, Ratio2_2 = groupBitsRatio2; When Ratio2_1 ≥ groupBitsRatio2, Ratio2_2 = FAC6 * groupBitsRatio2 + (1 – FAC6) * Ratio2_1; Where, Ratio2_2 represents the updated bit allocation proportion of the residual signal group, FAC6 is a preset sixth adjustment factor, Ratio2_1 is the bit allocation proportion of the residual signal group before update, * represents multiplication operation, and groupBitsRatio2 is the preset bit allocation proportion of the residual signal group.

15. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Determine the number of bits of the virtual speaker signal group and the number of bits of the residual signal group respectively according to the bit allocation proportion of the virtual speaker signal group, the bit allocation proportion of the residual signal group, and the total number of bits of the transmission channel; The virtual speaker signal group is bit-allocated according to the number of bits in the virtual speaker signal group, and the residual signal group is bit-allocated according to the number of bits in the residual signal group.

16. The method according to claim 15, characterized in that, The step of determining the number of bits in the virtual speaker signal group and the number of bits in the residual signal group based on the bit allocation ratio of the virtual speaker signal group, the bit allocation ratio of the residual signal group, and the total number of transmission channel bits includes: The number of bits in the virtual speaker signal group is calculated as follows: F_bitnum = Ratio1 * C_bitnum; Wherein, F_bitnum is the number of bits in the virtual speaker signal group, Ratio1 is the bit allocation ratio of the virtual speaker signal group, and C_bitnum is the total number of transmission channel bits; The number of bits in the residual signal group is calculated as follows: D_bitnum = Ratio2 * C_bitnum; Wherein, D_bitnum is the number of bits in the residual signal group, Ratio2 is the bit allocation ratio of the residual signal group, and C_bitnum is the total number of bits in the transmission channel.

17. The method according to any one of claims 1 to 2, characterized in that, The method further includes: The bit allocation ratios of the transmission channel signal, the virtual speaker signal group, and the residual signal group are encoded and written into the bitstream.

18. A method for processing three-dimensional audio signals, characterized in that, include: Receive bitstream; Decode the bitstream to obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group; The virtual speaker signal and the residual signal in the bitstream are decoded according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain the decoded three-dimensional audio signal.

19. The method according to claim 18, characterized in that, Decoding the virtual speaker signal and residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group includes: Determine the number of available bits; The number of bits in the virtual speaker signal group is determined based on the available number of bits and the bit allocation ratio of the virtual speaker signal group; the virtual speaker signal in the bitstream is decoded based on the number of bits in the virtual speaker signal group; The number of bits in the residual signal group is determined based on the number of available bits and the bit allocation ratio of the residual signal group; the residual signal in the bit stream is decoded based on the number of bits in the residual signal group.

20. A three-dimensional audio signal processing device, characterized in that, include: The encoding module is used to spatially encode the three-dimensional audio signal to be encoded to obtain the transmission channel signal and transmission channel attribute information, wherein the transmission channel signal includes: at least one virtual speaker signal group and at least one residual signal group; The bit allocation ratio determination module is used to determine the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group based on the transmission channel attribute information.

21. A three-dimensional audio signal processing device, characterized in that, include: The receiving module is used to receive the bit stream; The decoding module is used to decode the bitstream to obtain the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group; The signal generation module is used to decode the virtual speaker signal and the residual signal in the bitstream according to the bit allocation ratio of the virtual speaker signal group and the bit allocation ratio of the residual signal group to obtain the decoded three-dimensional audio signal.

22. A three-dimensional audio signal processing device, characterized in that, The three-dimensional audio signal processing device includes at least one processor, the at least one processor being coupled to a memory to read and execute instructions in the memory to implement the method as described in any one of claims 1 to 17.

23. The three-dimensional audio signal processing apparatus according to claim 22, characterized in that, The three-dimensional audio signal processing device further includes the memory.

24. A three-dimensional audio signal processing device, characterized in that, The three-dimensional audio signal processing device includes at least one processor, the at least one processor being coupled to a memory to read and execute instructions in the memory to implement the method as described in any one of claims 18 to 19.

25. The three-dimensional audio signal processing apparatus according to claim 24, characterized in that, The three-dimensional audio signal processing device further includes the memory.

26. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 17 or 18 to 19.

27. A computer-readable storage medium comprising a bitstream generated by the method as described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Apparatus and method for processing audio signal

    KR1020140017344A

  • Apparatus and method for audio signal processing

    KR1020140128565A

Cited By

  • Three-dimensional audio signal processing method and apparatus

    WO2022257824A1