A method and apparatus for processing three-dimensional audio signals
By performing linear decomposition and sound field classification of three-dimensional audio signals, the problem of the inability to identify three-dimensional audio signals before encoding is solved, thereby improving encoding efficiency and auditory quality, and reducing storage and transmission requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-05-31
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, three-dimensional audio signals cannot be effectively identified before encoding, resulting in problems such as large data volume, high storage space and transmission bandwidth requirements.
By performing linear decomposition on the current frame of the 3D audio signal, sound field classification parameters are obtained, and the sound field classification result is determined, thereby achieving accurate identification and classification of the 3D audio signal.
It improves the encoding efficiency and auditory quality of three-dimensional audio signals, while reducing storage space and transmission bandwidth requirements.
Smart Images

Figure CN115938388B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a method and apparatus for processing three-dimensional audio signals. Background Technology
[0002] 3D audio technology has been widely applied in wireless communication voice, virtual reality / augmented reality, and media audio. It is an audio technology that acquires, processes, transmits, renders, and plays back sound events and 3D sound field information from the real world. 3D audio technology gives sound a strong sense of space, immersion, and surround sound, providing an extraordinary "sound-immersive" auditory experience. Higher-order ambisonics (HOA) technology, with its speaker layout independence during recording, encoding, and playback, and the rotatable playback characteristics of HOA format data, offers greater flexibility in 3D audio playback, thus attracting wider attention and research.
[0003] Acquisition devices (such as microphones) collect large amounts of data to record 3D sound field information and transmit the 3D audio signals to playback devices (such as speakers and headphones) for playback. Because the 3D sound field information is large in volume, it requires significant storage space and high bandwidth for transmission. To address these issues, the 3D audio signals can be compressed, and the compressed data can be stored or transmitted.
[0004] Currently, encoders can use multiple pre-configured virtual speakers to encode 3D audio signals. However, before the encoder encodes the 3D audio signals, it is impossible to classify the 3D audio signals, resulting in the inability to effectively identify them. Summary of the Invention
[0005] This application provides a method and apparatus for processing three-dimensional audio signals, which is used to classify the sound field of three-dimensional audio signals, thereby enabling accurate identification of three-dimensional audio signals.
[0006] To address the aforementioned technical problems, the embodiments of this application provide the following technical solutions:
[0007] Firstly, embodiments of this application provide a method for processing three-dimensional audio signals, including: performing linear decomposition on the current frame of the three-dimensional audio signal to obtain a linear decomposition result; obtaining sound field classification parameters corresponding to the current frame based on the linear decomposition result; and determining the sound field classification result of the current frame based on the sound field classification parameters. In the above scheme, firstly, the current frame of the three-dimensional audio signal is linearly decomposed to obtain a linear decomposition result; then, the sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition result; and finally, the sound field classification result of the current frame is determined based on the sound field classification parameters. Since embodiments of this application obtain a linear decomposition result of the current frame by performing linear decomposition on the current frame of the three-dimensional audio signal, and then obtain the sound field classification parameters corresponding to the current frame through the linear decomposition result, the sound field classification result of the current frame is determined through the sound field classification parameters, and the sound field classification of the current frame can be achieved through the sound field classification result. Embodiments of this application perform sound field classification on three-dimensional audio signals, thereby accurately identifying three-dimensional audio signals.
[0008] In one possible implementation, the three-dimensional audio signal includes: a high-order stereo reverberation (HOA) signal, or a first-order stereo reverberation (FOA) signal.
[0009] In one possible implementation, the linear decomposition of the current frame of the three-dimensional audio signal to obtain a linear decomposition result includes: performing singular value decomposition on the current frame to obtain singular values corresponding to the current frame, wherein the linear decomposition result includes the singular values; or performing principal component analysis on the current frame to obtain a first eigenvalue corresponding to the current frame, wherein the linear decomposition result includes the first eigenvalue; or performing independent component analysis on the current frame to obtain a second eigenvalue corresponding to the current frame, wherein the linear decomposition result includes the second eigenvalue. In the above schemes, linear decomposition can be singular value decomposition. Linear decomposition can also be principal component analysis to obtain eigenvalues, and linear decomposition can also be independent component analysis to obtain a second eigenvalue. Linear decomposition of the current frame can be achieved through any of the above three methods, providing linear analysis results for subsequent channel determination.
[0010] In one possible implementation, there are multiple linear decomposition results and multiple sound field classification parameters; obtaining the sound field classification parameters corresponding to the current frame based on the linear decomposition results includes: obtaining the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer; and obtaining the i-th sound field classification parameter corresponding to the current frame based on the ratio.
[0011] Furthermore, the i-th linear analysis result and the (i+1)-th linear analysis result are two consecutive linear analysis results of the current frame.
[0012] In the above scheme, the encoding end can calculate the sound field classification parameters corresponding to the current frame based on the linear decomposition results. For example, if there are multiple linear decomposition results for the current frame, and two consecutive linear analysis results are represented as the i-th linear analysis result and the (i+1)-th linear analysis result of the current frame, then the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result of the current frame can be calculated, without any limitation on the specific value of i. After obtaining the above ratio, the i-th sound field classification parameter corresponding to the current frame can be obtained using this ratio.
[0013] In one possible implementation, the sound field classification parameters are multiple; the sound field classification result includes: sound field type; determining the sound field classification result of the current frame based on the sound field classification parameters includes: determining the sound field type as a diffuse sound field when the values of the multiple sound field classification parameters all satisfy a preset diffuse sound source decision condition; or, determining the sound field type as a dissimilar sound field when at least one value of the multiple sound field classification parameters satisfies a preset dissimilar sound source decision condition. In the above scheme, the sound field type may include a dissimilar sound field and a diffuse sound field. In this embodiment, a diffuse sound source decision condition and a dissimilar sound source decision condition are preset. The diffuse sound source decision condition is used to determine whether the sound field type is a diffuse sound field, and the dissimilar sound source decision condition is used to determine whether the sound field type is a dissimilar sound field. After obtaining the multiple sound field classification parameters of the current frame, a judgment is made based on the values of the multiple sound field classification parameters and the preset conditions mentioned above.
[0014] In one possible implementation, the diffusion sound source determination condition includes: the value of the sound field classification parameter is less than a preset dissimilar sound source determination threshold; or, the dissimilar sound source determination condition includes: the value of the sound field classification parameter is greater than or equal to the preset dissimilar sound source determination threshold. In the above scheme, the dissimilar sound source determination threshold can be a pre-set threshold, and its specific value is not limited. The diffusion sound source determination condition includes: the value of the sound field classification parameter is less than the preset dissimilar sound source determination threshold; therefore, when the values of multiple sound field classification parameters are all less than the preset dissimilar sound source determination threshold, the sound field type is determined to be a diffusion sound field. The dissimilar sound source determination condition includes: the value of the sound field classification parameter is greater than or equal to the preset dissimilar sound source determination threshold; therefore, when at least one of the values of multiple sound field classification parameters is greater than or equal to the preset dissimilar sound source determination threshold, the sound field type is determined to be a dissimilar sound field.
[0015] In one possible implementation, the sound field classification parameters are multiple; the sound field classification result includes: sound field type; or, the sound field classification result includes: number of dissimilar sound sources and sound field type; determining the sound field classification result of the current frame based on the sound field classification parameters includes: obtaining the number of dissimilar sound sources corresponding to the current frame based on the values of the multiple sound field classification parameters; and determining the sound field type based on the number of dissimilar sound sources corresponding to the current frame. In the above scheme, after the encoding end obtains multiple generation classification parameters corresponding to the current frame, the encoding end can obtain the number of dissimilar sound sources corresponding to the current frame through the values of the multiple sound field classification parameters. Dissimilar sound sources are point sound sources with different positions and / or directions. The number of dissimilar sound sources included in the current frame is called the number of dissimilar sound sources. The sound field of the current frame can be classified by the number of dissimilar sound sources. After obtaining the number of dissimilar sound sources corresponding to the current frame and determining the sound field type, the sound field type corresponding to the current frame can be determined by analyzing the number of dissimilar sound sources corresponding to the current frame.
[0016] In one possible implementation, the sound field classification parameters are multiple; the sound field classification result includes the number of dissimilar sound sources; determining the sound field classification result of the current frame based on the sound field classification parameters includes obtaining the number of dissimilar sound sources corresponding to the current frame based on the values of the multiple sound field classification parameters. In the above scheme, after the encoding end obtains multiple generation classification parameters corresponding to the current frame, the encoding end can obtain the number of dissimilar sound sources corresponding to the current frame through the values of the multiple sound field classification parameters. Dissimilar sound sources are point sound sources with different positions and / or directions. The number of dissimilar sound sources included in the current frame is called the number of dissimilar sound sources.
[0017] In one possible implementation, the plurality of sound field classification parameters are temp[i], where i = 0, 1, ..., min(L, K)-2, where L represents the number of channels in the current frame, K is the number of signal points corresponding to each channel in the current frame, and min represents the minimum value operation. Obtaining the number of dissimilar sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters includes: starting from i = 0, sequentially executing the following judgment process: determining whether temp[i] is greater than a preset dissimilar sound source determination threshold; when temp[i] is less than the dissimilar sound source determination threshold in this judgment process, updating the value of i to i+1, and continuing to execute the next judgment process; or, when temp[i] is greater than or equal to the dissimilar sound source determination threshold in this judgment process, terminating the judgment process, and determining that i plus 1 in this judgment process equals the number of dissimilar sound sources. In the above scheme, the number of dissimilar sound sources is obtained by repeatedly executing the above judgment process and determining whether to terminate the judgment process each time.
[0018] In one possible implementation, determining the sound field type based on the number of dissimilar sound sources corresponding to the current frame includes: determining the sound field type as a first sound field type when the number of dissimilar sound sources meets a first preset condition; and determining the sound field type as a second sound field type when the number of dissimilar sound sources does not meet the first preset condition. The number of dissimilar sound sources corresponding to the first sound field type is different from the number of dissimilar sound sources corresponding to the second sound field type. In the above scheme, the sound field type can be divided into two types according to the different numbers of dissimilar sound sources: a first sound field type and a second sound field type. The encoding end obtains the preset condition, determines whether the number of dissimilar sound sources meets the preset condition, and determines the sound field type as a first sound field type when the number of dissimilar sound sources meets the first preset condition; and determines the sound field type as a second sound field type when the number of dissimilar sound sources does not meet the first preset condition. In this embodiment of the application, the sound field type of the current frame can be divided by determining whether the number of dissimilar sound sources meets the first preset condition, thereby accurately identifying whether the sound field type of the current frame belongs to the first sound field type or the second sound field type.
[0019] In one possible implementation, the first preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or, the first preset condition includes the number of dissimilar sound sources being neither greater than the first threshold nor less than the second threshold, wherein the second threshold is greater than the first threshold. In the above scheme, the specific values of the first and second thresholds are not limited and can be determined based on the application scenario. Since the second threshold is greater than the first threshold, the first and second thresholds can constitute a preset range. Therefore, the first preset condition can be that the number of dissimilar sound sources is within this preset range, or the first preset condition can be that the number of dissimilar sound sources is outside this preset range. By using the first and second thresholds in the above first preset condition, the number of dissimilar sound sources can be judged to determine whether the number of dissimilar sound sources meets the first preset condition, thereby accurately identifying whether the sound field type of the current frame belongs to the first sound field type or the second sound field type.
[0020] In one possible implementation, the method further includes: determining the encoding mode corresponding to the current frame based on the sound field classification result. In the above scheme, the encoding end can determine the encoding mode corresponding to the current frame based on the sound field classification result. This encoding mode refers to the mode used when encoding the current frame of the three-dimensional audio signal. There are multiple encoding modes, and different encoding modes can be used depending on the sound field classification result of the current frame. In this embodiment, a suitable encoding mode is selected for different sound field classification results of the current frame, and this encoding mode is used to encode the current frame, thereby improving the compression efficiency and auditory quality of the audio signal.
[0021] In one possible implementation, determining the encoding mode corresponding to the current frame based on the sound field classification result includes: when the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the encoding mode corresponding to the current frame based on the number of dissimilar sound sources; or, when the sound field classification result includes the sound field type, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the encoding mode corresponding to the current frame based on the sound field type; or, when the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the encoding mode corresponding to the current frame based on the number of dissimilar sound sources and the sound field type. In the above scheme, the encoding end can determine the encoding mode corresponding to the current frame through the number of dissimilar sound sources and / or the sound field type, so that the encoding end can determine the corresponding encoding mode based on the sound field classification result of the current frame, making the determined encoding mode compatible with the current frame of the three-dimensional audio signal, thereby improving encoding efficiency.
[0022] In one possible implementation, determining the encoding mode corresponding to the current frame based on the number of dissimilar sound sources includes: when the number of dissimilar sound sources meets a second preset condition, determining the encoding mode as a first encoding mode; when the number of dissimilar sound sources does not meet the second preset condition, determining the encoding mode as a second encoding mode; wherein, the first encoding mode is a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional audio coding, and the second encoding mode is a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional audio coding, and the first encoding mode and the second encoding mode are different encoding modes. In the above scheme, the encoding mode can be divided into two types according to the different number of dissimilar sound sources: a first encoding mode and a second encoding mode. The encoding end obtains the second preset condition, determines whether the number of dissimilar sound sources meets the second preset condition, and when the number of dissimilar sound sources meets the second preset condition, determines the encoding mode as the first encoding mode; when the number of dissimilar sound sources does not meet the second preset condition, determines the encoding mode as the second encoding mode. In this embodiment of the application, the encoding mode of the current frame can be divided by determining whether the number of dissimilar sound sources meets the second preset condition, thereby accurately identifying whether the encoding mode of the current frame belongs to the first encoding mode or the second encoding mode.
[0023] In one possible implementation, the second preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or, the second preset condition includes the number of dissimilar sound sources being neither greater than the first threshold nor less than the second threshold, wherein the second threshold is greater than the first threshold.
[0024] In one possible implementation, determining the encoding mode corresponding to the current frame based on the sound field type includes: when the sound field type is an anisotropic sound field, determining the encoding mode as a HOA encoding mode selected based on a virtual loudspeaker; when the sound field type is a diffuse sound field, determining the encoding mode as a HOA encoding mode based on directional audio coding.
[0025] In one possible implementation, determining the encoding mode corresponding to the current frame based on the sound field classification result includes: determining the initial encoding mode corresponding to the current frame based on the sound field classification result of the current frame; obtaining a sliding window containing the current frame, the sliding window including: the initial encoding mode of the current frame and the encoding modes of the N-1 frames preceding the current frame, where N is the length of the sliding window; and determining the encoding mode of the current frame based on the initial encoding mode of the current frame and the encoding modes of the N-1 frames. In the above scheme, in this embodiment of the application, the initial encoding mode of the current frame is corrected by a sliding window to obtain the encoding mode of the current frame, so as to ensure that the encoding mode between consecutive frames does not switch frequently, thereby improving encoding efficiency.
[0026] In one possible implementation, the method further includes: determining the encoding parameters corresponding to the current frame based on the sound field classification result. In the above scheme, the encoding end can determine the encoding parameters corresponding to the current frame based on the sound field classification result. These encoding parameters refer to the parameters used when encoding the current frame of the three-dimensional audio signal. There are various encoding parameters, and different encoding parameters can be used depending on the sound field classification result of the current frame. In this embodiment, appropriate encoding parameters are selected for different sound field classification results of the current frame, and these encoding parameters are used to encode the current frame, thereby improving the compression efficiency and auditory quality of the audio signal.
[0027] In one possible implementation, the encoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of encoded bits of the virtual speaker signal, the number of encoded bits of the residual signal, or the number of voting rounds for the best matching speaker search; wherein the virtual speaker signal and the residual signal are signals generated based on the three-dimensional audio signal.
[0028] In one possible implementation, the number of voting rounds satisfies the following relationship: 1 ≤ I ≤ d, where I is the number of voting rounds and d is the number of dissimilar sound sources included in the sound field classification result. In the above scheme, the encoding end determines the number of voting rounds for the best matching speaker search based on the number of dissimilar sound sources in the current frame. This number of voting rounds is less than or equal to the number of dissimilar sound sources in the current frame, thus ensuring that the number of voting rounds conforms to the actual situation of the sound field classification in the current frame, solving the problem of needing to determine the number of voting rounds for the best matching speaker search when encoding the current frame.
[0029] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources and the sound field type. When the sound field type is a dissimilar sound field, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S, PF), where F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the encoder. Alternatively, when the sound field type is a diffuse sound field, the number of channels of the virtual speaker signal satisfies the following relationship: F = 1, where F is the number of channels of the virtual speaker signal. In the above scheme, the number of channels of the virtual speaker signal refers to the number of channels used to transmit the virtual speaker signal. The number of channels of the virtual speaker signal can be determined by the number of dissimilar sound sources and the sound field type. In the above calculation method, when the sound field type is a diffuse sound field, the number of channels of the virtual speaker signal is determined to be 1, thereby improving the coding efficiency of the current frame. When the sound field type is a dissimilar sound field, min represents the minimum value operation, that is, taking the minimum value from S and PF as the number of channels of the virtual speaker signal, so that the number of channels of the virtual speaker signal can conform to the actual situation of the sound field classification of the current frame, and solving the problem of determining the number of channels of the virtual speaker signal when encoding the current frame.
[0030] In one possible implementation, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1, PR), where PR is the number of residual signal channels preset by the encoder, and C is the sum of the number of residual signal channels preset by the encoder and the number of virtual speaker signal channels preset by the encoder; or, when the sound field type is a dissimilar sound field, the number of channels of the residual signal satisfies the following relationship: R = C – F, where R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the encoder and the number of virtual speaker signal channels preset by the encoder, and F is the number of channels of the virtual speaker signal. In the above scheme, after obtaining the number of channels of the virtual speaker signal, the number of channels of the residual signal can be calculated based on the preset number of channels of the residual signal, the sum of the preset number of channels of the virtual speaker signal, and the preset number of channels of the residual signal. The value of PR can be preset at the encoding end. The value of R can be obtained through the above max(C-1, PR) calculation formula. The preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal are preset at the encoding end. In addition, C can also be simply referred to as the total number of transmission channels.
[0031] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources; the number of channels of the virtual loudspeaker signal satisfies the following relationship: F = min(S, PF), where F is the number of channels of the virtual loudspeaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual loudspeaker signal channels preset by the encoder.
[0032] In one possible implementation, the number of channels of the residual signal satisfies the following relationship: R = C – F, where R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal. In the above scheme, after obtaining the number of channels of the virtual speaker signal, the number of channels of the residual signal can be calculated based on the preset number of channels of the residual signal, the preset sum of the number of channels of the virtual speaker signal, and the number of channels of the virtual speaker signal. The preset number of channels of the residual signal and the preset sum of the number of channels of the virtual speaker signal are preset by the encoding end. Alternatively, C can also be simply referred to as the total number of transmission channels.
[0033] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; the number of encoded bits of the virtual loudspeaker signal is obtained by the ratio of the number of encoded bits of the virtual loudspeaker signal to the number of encoded bits of the transmission channel; the number of encoded bits of the residual signal is obtained by the ratio of the number of encoded bits of the virtual loudspeaker signal to the number of encoded bits of the transmission channel; wherein, the number of encoded bits of the transmission channel includes the number of encoded bits of the virtual loudspeaker signal and the number of encoded bits of the residual signal, and when the number of dissimilar sound sources is less than or equal to the number of channels of the virtual loudspeaker signal, the ratio of the number of encoded bits of the virtual loudspeaker signal to the number of encoded bits of the transmission channel is obtained by increasing the initial ratio of the number of encoded bits of the virtual loudspeaker signal to the number of encoded bits of the transmission channel.
[0034] In one possible implementation, the method further includes: encoding the current frame and the sound field classification result, and writing them into a bitstream.
[0035] Secondly, embodiments of this application also provide a method for processing three-dimensional audio signals, including: receiving a bitstream; decoding the bitstream to obtain a sound field classification result of the current frame; and obtaining a decoded three-dimensional audio signal of the current frame based on the sound field classification result. In the above scheme, the sound field classification result can be used for decoding the current frame in the bitstream. Therefore, the decoding end uses a decoding method that matches the sound field of the current frame to perform decoding, thereby obtaining the three-dimensional audio signal sent by the encoding end, realizing the transmission of the audio signal from the encoding end to the decoding end.
[0036] In one possible implementation, obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result includes: determining the decoding mode of the current frame based on the sound field classification result; and obtaining the decoded three-dimensional audio signal of the current frame based on the decoding mode.
[0037] In one possible implementation, determining the decoding mode of the current frame based on the sound field classification result includes: when the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the decoding mode of the current frame based on the number of dissimilar sound sources; or, when the sound field classification result includes the sound field type, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the decoding mode of the current frame based on the sound field type; or, when the sound field classification result includes the number of dissimilar sound sources and the sound field type, determining the decoding mode of the current frame based on the number of dissimilar sound sources and the sound field type.
[0038] In one possible implementation, determining the decoding mode corresponding to the current frame based on the number of dissimilar sound sources includes: when the number of dissimilar sound sources meets a preset condition, determining the decoding mode as a first decoding mode; when the number of dissimilar sound sources does not meet the preset condition, determining the decoding mode as a second decoding mode; wherein, the first decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional audio coding, the second decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional audio coding, and the first decoding mode and the second decoding mode are different decoding modes.
[0039] In one possible implementation, the preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or, the preset condition includes the number of dissimilar sound sources being neither greater than the first threshold nor less than the second threshold, wherein the second threshold is greater than the first threshold.
[0040] In one possible implementation, obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result includes: determining the decoding parameters of the current frame based on the sound field classification result; and obtaining the decoded three-dimensional audio signal of the current frame based on the decoding parameters.
[0041] In one possible implementation, the decoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signal, or the number of decoded bits of the residual signal; wherein the virtual speaker signal and the residual signal are obtained by decoding the bitstream.
[0042] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources and the sound field type; when the sound field type is a dissimilar sound field, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S, PF), where F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder; or, when the sound field type is a diffuse sound field, the number of channels of the virtual speaker signal satisfies the following relationship: F = 1, where F is the number of channels of the virtual speaker signal.
[0043] In one possible implementation, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1, PR), where PR is the number of residual signal channels preset by the decoder, and C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder; or, when the sound field type is a dissimilar sound field, the number of channels of the residual signal satisfies the following relationship: R = C – F, where R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0044] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources; the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S, PF), where F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder.
[0045] In one possible implementation, the number of channels of the residual signal satisfies the following relationship: R = C – F, where R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0046] In one possible implementation, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; the number of decoded bits of the virtual speaker signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel; the number of decoded bits of the residual signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel; wherein, the number of decoded bits of the transmission channel includes the number of decoded bits of the virtual speaker signal and the number of decoded bits of the residual signal, and when the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel is obtained by increasing the initial ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel.
[0047] Thirdly, embodiments of this application also provide a three-dimensional audio signal processing apparatus, comprising: a linear analysis module for performing linear decomposition on the three-dimensional audio signal to obtain a linear decomposition result; a parameter generation module for obtaining sound field classification parameters corresponding to the current frame based on the linear decomposition result; and a sound field classification module for determining the sound field classification result of the current frame based on the sound field classification parameters.
[0048] In a third aspect of this application, the constituent modules of the three-dimensional audio signal processing apparatus may also perform the steps described in the first aspect and various possible implementations, as detailed in the foregoing description of the first aspect and various possible implementations.
[0049] Fourthly, embodiments of this application also provide a three-dimensional audio signal processing apparatus, comprising: a receiving module for receiving a bitstream; a decoding module for decoding the bitstream to obtain a sound field classification result of the current frame; and a signal generation module for obtaining a decoded three-dimensional audio signal of the current frame based on the sound field classification result.
[0050] In the fourth aspect of this application, the constituent modules of the three-dimensional audio signal processing apparatus may also perform the steps described in the second aspect and various possible implementations, as detailed in the foregoing description of the second aspect and various possible implementations.
[0051] In one possible implementation, the number of encoded bits of the virtual speaker signal satisfies the following relationship:
[0052]
[0053] Wherein, core_numbit is the number of encoded bits of the virtual speaker signal, fac1 is the weighting factor for the encoded bits of the virtual speaker signal, fac2 is the weighting factor for the encoded bits of the residual signal, round represents rounding down, F is the number of channels of the virtual speaker signal, R represents the number of channels of the residual signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal; the number of encoded bits of the residual signal satisfies the following relationship:
[0054] res_numbit=numbit-core_numbit.
[0055] Wherein, res_numbit is the number of encoded bits of the residual signal, core-numbit is the number of encoded bits of the virtual speaker signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal.
[0056] In one possible implementation, fac1 > fac2.
[0057] In one possible implementation, the number of encoded bits of the residual signal satisfies the following relationship:
[0058]
[0059] Wherein, res_numbit is the number of encoded bits of the residual signal, fac1 is the weighting factor for the encoded bits of the virtual speaker signal, fac2 is the weighting factor for the encoded bits of the residual signal, round represents rounding down, F is the number of channels of the virtual speaker signal, R represents the number of channels of the residual signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal;
[0060] The number of encoded bits of the virtual speaker signal satisfies the following relationship:
[0061] core_numbit=numbit-res_numbit;
[0062] Wherein, core_numbit is the number of encoded bits of the virtual speaker signal, res_numbit is the number of encoded bits of the residual signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal.
[0063] In one possible implementation, the number of encoded bits for each virtual speaker signal satisfies the following relationship:
[0064]
[0065] Wherein, core-ch_numbit is the number of encoded bits for each virtual speaker signal, fac1 is the weighting factor for the encoded bits of the virtual speaker signal, fac2 is the weighting factor for the encoded bits of the residual signal, round represents rounding down, F is the number of channels of the virtual speaker signal, R represents the number of channels of the residual signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal;
[0066] The number of coded bits for each residual signal satisfies the following relationship:
[0067]
[0068] Wherein, res_numbit is the number of encoded bits for each residual signal, fac1 is the weighting factor for the encoded bits of the virtual speaker signal, fac2 is the weighting factor for the encoded bits of the residual signal, round represents rounding down, F is the number of channels of the virtual speaker signal, R represents the number of channels of the residual signal, and numbit is the sum of the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal.
[0069] Fifthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect above.
[0070] In a sixth aspect, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the first or second aspect above.
[0071] In a seventh aspect, embodiments of this application provide a computer-readable storage medium including a bitstream generated by the method described in the first aspect above.
[0072] Eighthly, embodiments of this application provide a communication device, which may include entities such as terminal devices or chips. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the first or second aspects above.
[0073] Ninthly, this application provides a chip system including a processor for supporting an audio encoder or audio decoder in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing necessary program instructions and data for the audio encoder or audio decoder. This chip system may be composed of chips or may include chips and other discrete devices.
[0074] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0075] In this embodiment, the current frame of the 3D audio signal is first linearly decomposed to obtain the linear decomposition result; then, the sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition result; finally, the sound field classification result of the current frame is determined based on the sound field classification parameters. Since this embodiment obtains the linear decomposition result of the current frame by performing linear decomposition on the current frame of the 3D audio signal, and then obtains the sound field classification parameters corresponding to the current frame through these linear decomposition results, the sound field classification result of the current frame is determined through these sound field classification parameters. This sound field classification result allows for sound field classification of the current frame. This embodiment of the application performs sound field classification on the 3D audio signal, thereby enabling accurate identification of the 3D audio signal. Attached Figure Description
[0076] Figure 1 This is a schematic diagram of the composition structure of the audio processing system provided in the embodiments of this application;
[0077] Figure 2a A schematic diagram illustrating the application of the audio encoder and audio decoder provided in this application to a terminal device;
[0078] Figure 2b A schematic diagram illustrating the application of the audio encoder provided in this application to a wireless device or a core network device;
[0079] Figure 2c A schematic diagram illustrating the application of the audio decoder provided in this application embodiment to a wireless device or core network device;
[0080] Figure 3a A schematic diagram illustrating the application of the multi-channel encoder and multi-channel decoder provided in the embodiments of this application to a terminal device;
[0081] Figure 3b A schematic diagram illustrating the application of a multi-channel encoder provided in this application to a wireless device or a core network device;
[0082] Figure 3c A schematic diagram illustrating the application of the multi-channel decoder provided in this application to a wireless device or a core network device;
[0083] Figure 4 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;
[0084] Figure 5 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;
[0085] Figure 6 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;
[0086] Figure 7 A schematic diagram illustrating a method for processing three-dimensional audio signals provided in an embodiment of this application;
[0087] Figure 8 A schematic diagram of the encoding process of a hybrid HOA encoder provided in an embodiment of this application;
[0088] Figure 9 A flowchart illustrating the process of determining the encoding mode of an HOA signal, provided for an embodiment of this application;
[0089] Figure 10 This application provides a schematic diagram of the decoding process of a hybrid HOA decoder.
[0090] Figure 11 A schematic diagram of the encoding process of an MP-based HOA encoder provided for an embodiment of this application;
[0091] Figure 12 This is a schematic diagram of the composition structure of an audio encoding device provided in an embodiment of this application;
[0092] Figure 13 This is a schematic diagram of the composition structure of an audio decoding device provided in an embodiment of this application;
[0093] Figure 14 This is a schematic diagram of the composition structure of another audio encoding device provided in the embodiments of this application;
[0094] Figure 15 This is a schematic diagram of the composition structure of another audio decoding device provided in an embodiment of this application. Detailed Implementation
[0095] The embodiments of this application will now be described with reference to the accompanying drawings.
[0096] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0097] Sound is a continuous wave produced by the vibration of an object. The object that produces vibrations and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solids, or liquids), the auditory organs of humans or animals can perceive the sound.
[0098] Sound waves are characterized by pitch, intensity, and timbre. Pitch indicates the highness or lowness of a sound. Intensity indicates the loudness or volume of a sound. The unit of intensity is the decibel (dB). Timbre is also known as tone color.
[0099] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is hertz (Hz). The human ear can distinguish sounds with frequencies between 20 Hz and 20,000 Hz.
[0100] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.
[0101] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, and pulse waves, among others.
[0102] Based on the characteristics of sound waves, sound can be divided into regular sound and irregular sound. Irregular sound refers to sound emitted by the irregular vibration of a sound source. Irregular sound is, for example, noise that affects people's work, study, and rest. Regular sound refers to sound emitted by the regular vibration of a sound source. Regular sound includes speech and musical tones. When sound is represented electronically, regular sound is an analog signal that varies continuously in the time and frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.
[0103] Because human hearing has the ability to distinguish the location of sound sources in space, when a listener hears a sound in space, in addition to being able to perceive the pitch, intensity, and timbre of the sound, they can also perceive the location of the sound.
[0104] As people pay increasing attention to and demand higher quality in their auditory experience, three-dimensional audio technology has emerged to enhance the depth, presence, and spatial feel of sound. This allows listeners to not only perceive sounds from front, back, left, and right sources, but also to feel surrounded by the spatial sound field created by these sources, and to experience the sound spreading outwards, creating an immersive audio experience as if the listener were in a cinema or concert hall.
[0105] Three-dimensional audio technology refers to the concept of the space outside the human ear as a system, where the signal received at the eardrum is a three-dimensional audio signal output after the sound emitted from the sound source has been filtered by this external system. For example, the system outside the human ear can be defined as the system impulse response h(n), any sound source can be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal described in this application can refer to a higher-order ambisonics (HOA) signal or a first-order ambisonics (FOA) signal. Three-dimensional audio can also be called three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio, or binaural audio, etc.
[0106] Sound waves propagate in an ideal medium with a wave number of k = w / c and an angular frequency of w = 2πf, where f is the sound wave frequency and c is the speed of sound. The sound pressure p satisfies formula (1). For the Laplace operator.
[0107]
[0108] Assuming the spatial system outside the human ear is a sphere, with the listener at the center, the sound from outside the sphere has a projection on the sphere's surface. Filtering out sounds from outside the sphere, and assuming the sound sources are distributed on this sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound source. That is, three-dimensional audio technology is a method of fitting a sound field. Specifically, in spherical coordinates, equation (1) is solved. In the passive spherical region, the solution to equation (1) is as follows: equation (2).
[0109]
[0110] Where r represents the radius of the sphere, and θ represents the horizontal angle. The elevation angle is represented by k, the wave number by s, the amplitude of the ideal plane wave by m, and the order number of the three-dimensional audio signal (or the order number of the HOA signal). Let represent the spherical Bessel function, also known as the radial basis function, where the first 'j' represents the imaginary unit. It does not change with the angle. Represents θ, spherical harmonic function of direction, The spherical harmonic function represents the direction of the sound source. The coefficients of the three-dimensional audio signal satisfy formula (3).
[0111]
[0112] Substituting formula (3) into formula (2), formula (2) can be transformed into formula (4).
[0113]
[0114] in, The coefficients of the three-dimensional audio signal of order N are used to approximate the sound field. A sound field refers to the region in a medium where sound waves exist. N is an integer greater than or equal to 1. For example, the value of N ranges from 2 to 6. The coefficients of the three-dimensional audio signal described in the embodiments of this application may refer to HOA coefficients or ambisonic coefficients.
[0115] A three-dimensional audio signal is an information carrier that carries the spatial location information of the sound source in the sound field, describing the sound field of the listener in space. Equation (4) shows that the sound field can be expanded on a sphere according to the spherical harmonic function, that is, the sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional audio signal can be expressed by the superposition of multiple plane waves, and the sound field can be reconstructed through the coefficients of the three-dimensional audio signal.
[0116] Compared to a 5.1 channel audio signal or a 7.1 channel audio signal, an Nth-order HOA signal has (N+1) 2With multiple channels, the HOA signal contains a large amount of data describing the spatial information of the sound field. If the acquisition device (e.g., a microphone) transmits this 3D audio signal to the playback device (e.g., a speaker), it consumes a significant amount of bandwidth. Currently, encoders can use spatial squeezed surround audio coding (S3AC), directional audio coding (DirAC), or virtual speaker selection-based coding methods to compress and encode the 3D audio signal into a bitstream, which is then transmitted to the playback device. The virtual speaker selection-based coding method can also be called match projection (MP) coding; we will use virtual speaker selection-based coding as an example later. The playback device decodes the bitstream, reconstructs the 3D audio signal, and plays the reconstructed 3D audio signal. This reduces the amount of data transmitted to the playback device and the bandwidth usage.
[0117] Currently, it is impossible to classify the sound field of the aforementioned three-dimensional audio signal. How to classify the sound field of a three-dimensional audio signal is a technical problem that this application aims to solve. In this application, the linear decomposition of the three-dimensional audio signal enables sound field classification, thereby accurately classifying the sound field and achieving the goal of obtaining the sound field classification result of the current frame.
[0118] Furthermore, current encoders suffer from the inability to achieve high compression ratios when compressing 3D audio signals. Therefore, improving the compression ratio of 3D audio signals with different sound fields is another problem addressed by the embodiments of this application.
[0119] This application provides an audio encoding technique, particularly a three-dimensional audio encoding technique for three-dimensional audio signals. Specifically, it provides an encoding technique that uses fewer channels to represent three-dimensional audio signals, thereby improving traditional audio encoding systems. Audio encoding (or commonly referred to as encoding) includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and includes processing (e.g., compressing) the raw audio to reduce the amount of data required to represent the audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are also collectively referred to as encoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0120] The technical solutions of this application embodiment can be applied to various audio processing systems, such as... Figure 1The diagram shown is a schematic representation of the composition of an audio processing system provided in this embodiment. The audio processing system 100 may include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 generates a bitstream, which is then transmitted to the audio decoding device 102 via an audio transmission channel. The audio decoding device 102 receives the bitstream, performs its audio decoding function, and finally obtains the reconstructed signal.
[0121] In the embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be an audio encoder for the aforementioned terminal devices, wireless devices, or core network devices. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be an audio decoder for the aforementioned terminal devices, wireless devices, or core network devices. For example, the audio encoder can include a wireless access network, a core network media gateway, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The audio encoder can also be an audio encoder used in virtual reality (VR) streaming services.
[0122] In the application embodiment, taking the audio encoding module (audio encoding and audio decoding) applicable to virtual reality streaming (VR streaming) services as an example, the end-to-end audio signal processing flow includes: after the audio signal A passes through the acquisition module, it undergoes a preprocessing operation (audioPReprocessing). The preprocessing operation includes filtering out the low-frequency part in the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the directional information in the signal, and then performing encoding processing (audio encoding), packaging (file / segment encapsulation), and then sending (delivery) to the decoding end. The decoding end first performs unpacking (file / segment decapsulation), then decoding (audio decoding), and performs binaural rendering processing on the decoded signal. The rendered signal is mapped onto the listener's headphones, which can be independent headphones or headphones on the glasses device.
[0123] like Figure 2aThe diagram illustrates the application of the audio encoder and audio decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used for channel encoding of the audio signal, and the channel decoder is used for channel decoding of the audio signal. For example, the first terminal device 20 may include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. The second terminal device 21 may include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and a wireless or wired second network communication device 23 are connected via a digital channel. The second terminal device 21 is connected to the wireless or wired second network communication device 23. The aforementioned wireless or wired network communication device can broadly refer to signal transmission devices, such as communication base stations, data exchange devices, etc.
[0124] In audio communication, the transmitting terminal device first acquires audio, encodes the acquired audio signal, and then performs channel coding before transmitting it over a digital channel via a wireless network or core network. The receiving terminal device, acting as the receiver, decodes the received signal to obtain the bitstream, then recovers the audio signal through audio decoding for playback.
[0125] like Figure 2b The diagram illustrates the application of the audio encoder provided in this embodiment of the application in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, the audio encoder 253 provided in this embodiment of the application, and a channel encoder 254. The other audio decoders 252 refer to audio decoders other than the standard audio decoder. Within the wireless device or core network device 25, the signal entering the device is first channel-decoded using the channel decoder 251, then audio-decoded using the other audio decoder 252, then audio-encoded using the audio encoder 253 provided in this embodiment of the application, and finally channel-encoded using the channel encoder 254. After channel encoding, the signal is transmitted out. The other audio decoders 252 perform audio decoding on the bitstream decoded by the channel decoder 251.
[0126] like Figure 2cThe diagram illustrates the application of the audio decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in this embodiment, other audio encoders 256, and a channel encoder 254. The other audio encoders 256 refer to audio encoders other than the standard audio encoder. Within the wireless device or core network device 25, the channel decoder 251 first performs channel decoding on the incoming signal. Then, the audio decoder 255 decodes the received audio encoded bitstream. Next, the other audio encoders 256 perform audio encoding. Finally, the channel encoder 254 performs channel encoding on the audio signal before transmission. If transcoding is required in the wireless device or core network device, corresponding audio encoding processing is necessary. The wireless device refers to radio frequency (RF) related equipment in communication, and the core network device refers to core network related equipment in communication.
[0127] In some embodiments of this application, the audio encoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio encoding device can be a multi-channel encoder of the aforementioned terminal device, wireless device, or core network device. Similarly, the audio decoding device can be applied to various terminal devices requiring audio communication, wireless devices requiring transcoding, and core network devices. For example, the audio decoding device can be a multi-channel decoder of the aforementioned terminal device, wireless device, or core network device.
[0128] like Figure 3aThe diagram illustrates the application of the multi-channel encoder and multi-channel decoder provided in this embodiment of the application to a terminal device. Each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in this embodiment of the application, and the multi-channel decoder can execute the audio decoding method provided in this embodiment of the application. Specifically, the channel encoder is used to perform channel encoding on the multi-channel signal, and the channel decoder is used to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. The second terminal device 31 may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and a wireless or wired second network communication device 33 are connected via a digital channel. The second terminal device 31 is connected to the wireless or wired second network communication device 33. The aforementioned wireless or wired network communication equipment can broadly refer to signal transmission equipment, such as communication base stations and data switching equipment. In audio communication, the transmitting terminal device performs multi-channel encoding on the acquired multi-channel signal, followed by channel encoding, and then transmits it through a wireless network or core network in a digital channel. The receiving terminal device performs channel decoding on the received signal to obtain the multi-channel signal encoded bitstream, and then recovers the multi-channel signal through multi-channel decoding for playback by the receiving terminal device.
[0129] like Figure 3b The diagram shown illustrates the application of the multi-channel encoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, as described above. Figure 2b Similarly, this will not be elaborated upon here.
[0130] like Figure 3c The diagram shown illustrates the application of the multi-channel decoder provided in this embodiment of the invention in a wireless device or core network device. The wireless device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, as described above. Figure 2c Similarly, this will not be elaborated upon here.
[0131] The audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the acquired multi-channel signal can involve processing the acquired multi-channel signal to obtain an audio signal, and then encoding the obtained audio signal according to the method provided in this application embodiment. The decoding end decodes the audio signal based on the multi-channel signal encoded bitstream, and recovers the multi-channel signal after upmixing. Therefore, this application embodiment can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding is required, corresponding multi-channel encoding processing is necessary.
[0132] This application first introduces a method for processing three-dimensional audio signals, provided in an embodiment. This method can be executed by a terminal device, such as an audio encoding device (hereinafter referred to as an encoder). Not limited to this, the terminal device can also be a three-dimensional audio signal processing device. Figure 4 As shown, the main methods for processing three-dimensional audio signals include the following:
[0133] 401. Perform linear decomposition on the current frame of the three-dimensional audio signal to obtain the linear decomposition result.
[0134] The encoding end can acquire a three-dimensional audio signal, such as a scene audio signal. Specifically, the three-dimensional audio signal can be a time-domain signal or a frequency-domain signal. Additionally, the three-dimensional audio signal can also be a downsampled signal.
[0135] In some embodiments of this application, the three-dimensional audio signal includes: a high-order stereo reverberation (HOA) signal or a first-order stereo reverberation (FOA) signal. It is not limited to this; the three-dimensional audio signal can also be other types of signals. This is merely an example and is not intended to limit the embodiments of this application.
[0136] For example, a 3D audio signal can be a time-domain HOA signal or a frequency-domain HOA signal. Furthermore, a 3D audio signal can contain all channels of a HOA signal or only some HOA channels (e.g., FOA channels). Additionally, a 3D audio signal can be all samples of a HOA signal or 1 / Q downsampled points of the HOA signal to be analyzed. Here, Q is the downsampling interval, and 1 / Q is the downsampling rate.
[0137] In this embodiment, the three-dimensional audio signal includes multiple frames. The following example focuses on processing one frame of the three-dimensional audio signal. For instance, if this frame is the current frame, then the three-dimensional audio signal contains a previous frame before the current frame and a subsequent frame after the current frame. Furthermore, the processing methods for other frames of the three-dimensional audio signal besides the current frame in this embodiment are similar to the processing method for the current frame; the following example will also focus on the processing of the current frame.
[0138] In this embodiment, after obtaining the current frame of the three-dimensional audio signal, the current frame is first linearly decomposed to obtain the linear decomposition result of the current frame. There are various methods of linear decomposition, which will be described in detail below.
[0139] In some embodiments of this application, step 401 performs linear decomposition on the current frame of the three-dimensional audio signal to obtain the linear decomposition result, including:
[0140] A1. Perform singular value decomposition on the current frame to obtain the singular values corresponding to the current frame. The linear decomposition result includes singular values.
[0141] or,
[0142] A2. Perform principal component analysis on the current frame to obtain the first eigenvalue corresponding to the current frame. The linear decomposition result includes the first eigenvalue.
[0143] or,
[0144] A3. Perform independent component analysis on the current frame to obtain the second eigenvalue corresponding to the current frame. The linear decomposition result includes the second eigenvalue.
[0145] There are various methods for linear decomposition, including at least one of the following: singular value decomposition (SVD), principal component analysis (PCA), and independent component analysis (ICA). The results obtained from different linear decomposition methods are expressed in different ways, which will be explained in detail below.
[0146] In step A1, linear decomposition can be singular value decomposition. For example, suppose the three-dimensional audio signal is a HOA signal, and matrix A is composed of the HOA signals. Matrix A is an L*K matrix, where L equals the number of channels of the HOA signal, and K is the number of signal points in each channel of the HOA signal in the current frame. For example, the number of signal points can include: the number of frequency points, or the number of time-domain samples, or the number of frequency points or samples after downsampling. Singular value decomposition of matrix A satisfies the following relationship:
[0147] A=UΣV T .
[0148] Where U is an L*L matrix, V is a K*K matrix, the subscript T is the transpose of matrix V, and * denotes multiplication. Σ is an L*K diagonal matrix, where each element on the main diagonal is a singular value of matrix A obtained from singular value decomposition, and all elements outside the main diagonal are 0. The elements on the main diagonal of the diagonal matrix Σ, i.e., the singular values of matrix A, are denoted as v[i], i = 0, 1, ..., min(L, K)-1.
[0149] It should be noted that if the 3D audio signal is a downsampled HOA signal, then K is the number of signal points after downsampling of each channel of the HOA signal in the current frame. For example, the number of signal points can be the number of sample points or the number of frequency points.
[0150] In step A2, linear decomposition can also be principal component analysis (PCA) to obtain eigenvalues. To distinguish them from other eigenvalues in subsequent embodiments, the eigenvalues obtained through PCA are defined as the first eigenvalues. The specific implementation of PCA will not be elaborated here.
[0151] In step A3, linear decomposition can also be performed as independent component analysis to obtain the second eigenvalue. The specific implementation of independent component analysis will not be elaborated here.
[0152] In this embodiment of the application, any of the above-mentioned implementation methods A1 to A3 can be used to achieve linear decomposition of the current frame, thereby obtaining various types of linear decomposition results.
[0153] 402. Obtain the sound field classification parameters corresponding to the current frame based on the linear decomposition results.
[0154] After obtaining the linear analysis result of the current frame, the encoder analyzes this linear decomposition result to obtain the sound field classification parameters corresponding to the current frame. These sound field classification parameters are obtained by analyzing the linear decomposition result of the current frame and are used to determine the sound field classification result of the current frame. Depending on the specific implementation method of the linear decomposition result, these sound field classification parameters can have multiple implementation methods.
[0155] In the embodiments of this application, the linear decomposition result can be one or more. For example, the linear decomposition result includes singular values, singular values v[i], i = 0, 1, ..., min(L, K)-1. When there is only one singular value in the current frame, i has only one value, i.e., v[0]. When there are multiple singular values in the current frame, i has multiple values, i.e., v[i], i = 1, ..., min(L, K)-1.
[0156] In this embodiment, when there are two linear decomposition results, one sound field classification parameter is obtained. When there are N linear decomposition results, N-1 sound field classification parameters are obtained, and the value of N is not limited.
[0157] In some embodiments of this application, step 402 obtains the sound field classification parameters corresponding to the current frame based on the linear decomposition result, including:
[0158] B1. Obtain the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer;
[0159] B2. Obtain the i-th sound field classification parameter corresponding to the current frame based on the ratio.
[0160] The encoding end can calculate the sound field classification parameters corresponding to the current frame based on the linear decomposition results. For example, if there are multiple linear decomposition results for the current frame, and two consecutive linear analysis results are represented as the i-th linear analysis result and the (i+1)-th linear analysis result of the current frame, then the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result of the current frame can be calculated, and the specific value of i is not limited.
[0161] Optionally, the i-th linear analysis result and the (i+1)-th linear analysis result are two consecutive linear analysis results of the current frame.
[0162] After obtaining the aforementioned ratio, the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result of the current frame can be used to obtain the i-th sound field classification parameter corresponding to the current frame. This demonstrates that the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result can be used to calculate the i-th sound field classification parameter, and the ratio of the (i+1)-th linear analysis result to the (i+2)-th linear analysis result can be used to calculate the (i+1)-th sound field classification parameter, and so on. There is a corresponding relationship between the linear analysis results and the sound field classification parameters.
[0163] One possible approach is to use the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result as the i-th sound field classification parameter. However, it is not limited to this; after obtaining the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result, various calculations can be performed on this ratio to calculate the i-th sound field classification parameter. For example, the ratio can be multiplied by a preset adjustment factor to obtain the i-th sound field classification parameter.
[0164] For example, if the linear decomposition uses singular value decomposition, the sound field classification parameters can be obtained by obtaining singular values from the singular value decomposition, and the ratio parameter between two adjacent singular values can be calculated as the sound field classification parameters.
[0165] For example, the ratio temp[i] between singular values is calculated as a sound field classification parameter. For i = 0, 1, ..., min(L, K)-2, temp[i] satisfies:
[0166] temp[i] = v[i] / v[i+1].
[0167] If PCA or ICA is used for linear decomposition, the sound field classification parameters can be determined based on the eigenvalues. The calculation method for the sound field classification parameters is similar to the calculation method for the ratio temp between singular values mentioned above; it can also be based on the eigenvalues obtained from the linear decomposition, calculating the ratio between two consecutive eigenvalues as the sound field classification parameters.
[0168] It should be noted that if the number of eigenvalues or singular values obtained by linear decomposition is greater than 2, then the sound field classification parameter is a vector; otherwise, the sound field classification parameter is a scalar. For example, for v[i], if the value of i is equal to 2, then the calculated temp[i] is a scalar, that is, there is only one temp value; for v[i], if the value of i is greater than 2, then the calculated temp[i] is a vector, and temp has at least two elements.
[0169] 403. Determine the sound field classification result of the current frame based on the sound field classification parameters.
[0170] In this embodiment of the application, after the encoding end obtains the sound field classification parameters corresponding to the current frame, the encoding end can perform sound field classification on the current frame according to the sound field classification parameters. Since the sound field classification parameters corresponding to the current frame can indicate the parameters required for classifying the sound field corresponding to the current frame, the sound field classification result of the current frame can be obtained based on the sound field classification parameters.
[0171] In some embodiments of this application, the sound field classification result may include at least one of the following: sound field type, number of dissimilar sound sources.
[0172] The sound field type refers to the type of sound field in the current frame determined after sound field classification. There are various ways to classify sound field types; for example, it can be divided into first sound field type, second sound field type, or third sound field type, etc. The specific number of types can be determined based on the application scenario. For instance, sound field types can include dissimilar sound fields and diffuse sound fields. A dissimilar sound field refers to a sound field containing point sound sources with different positions and / or directions, while a diffuse sound field refers to a sound field that does not contain dissimilar sound sources. For example, point sound sources with different positions and / or directions are dissimilar sound sources; a sound field containing dissimilar sound sources is a dissimilar sound field, and a sound field without dissimilar sound sources is a diffuse sound field.
[0173] Dissimilar sound sources are point sound sources with different locations and / or orientations. The number of dissimilar sound sources included in the current frame is called the number of dissimilar sound sources. The sound field of the current frame can also be classified based on the number of dissimilar sound sources.
[0174] In some embodiments of this application, there are multiple sound field classification parameters; the sound field classification result includes: sound field type;
[0175] Step 403 determines the sound field classification result of the current frame based on the sound field classification parameters, including:
[0176] When the values of multiple sound field classification parameters all meet the preset criteria for determining a diffuse sound source, the sound field type is determined to be a diffuse sound field.
[0177] or,
[0178] When at least one of the values of multiple sound field classification parameters satisfies the preset dissimilar sound source decision condition, the sound field type is determined to be a dissimilar sound field.
[0179] The sound field type can include dissimilar sound fields and diffuse sound fields. In this embodiment, diffuse sound source decision conditions and dissimilar sound source decision conditions are preset. The diffuse sound source decision conditions are used to determine whether the sound field type is a diffuse sound field, and the dissimilar sound source decision conditions are used to determine whether the sound field type is a dissimilar sound field. After obtaining multiple sound field classification parameters of the current frame, a judgment is made based on the values of the multiple sound field classification parameters and the preset conditions. The specific implementation of the diffuse sound source decision conditions and the dissimilar sound source decision conditions is not limited here.
[0180] After the encoding end obtains multiple sound field classification parameters, the sound field type is determined to be a diffuse sound field when the values of all multiple sound field classification parameters meet the preset diffuse sound source determination conditions. For example, if the current frame corresponds to N sound field classification parameters, the sound field type of the current frame is determined to be a diffuse sound field only when the values of these N sound field classification parameters all meet the preset diffuse sound source determination conditions.
[0181] After the encoding end obtains multiple sound field classification parameters, the sound field type is determined to be an incongruous sound field when at least one of the values of these multiple sound field classification parameters meets the preset incongruous sound source determination condition. For example, if the current frame corresponds to N sound field classification parameters, then the sound field type is determined to be an incongruous sound field as long as at least one of these N sound field classification parameters meets the preset incongruous sound source determination condition.
[0182] Furthermore, in some embodiments of this application, the criteria for determining diffuse sound sources include: the value of the sound field classification parameter is less than a preset threshold for determining dissimilar sound sources;
[0183] or,
[0184] The criteria for determining dissimilar sound sources include: the value of the sound field classification parameter is greater than or equal to the preset dissimilar sound source determination threshold.
[0185] The dissimilar sound source determination threshold can be a pre-set threshold, and its specific value is not limited. The criteria for determining a diffuse sound source include: the value of the sound field classification parameter is less than the pre-set dissimilar sound source determination threshold; therefore, when the values of multiple sound field classification parameters are all less than the pre-set dissimilar sound source determination threshold, the sound field type is determined to be a diffuse sound field. The criteria for determining a dissimilar sound source include: the value of the sound field classification parameter is greater than or equal to the pre-set dissimilar sound source determination threshold; therefore, when at least one of the multiple sound field classification parameters is greater than or equal to the pre-set dissimilar sound source determination threshold, the sound field type is determined to be a dissimilar sound field.
[0186] In some embodiments of this application, there are multiple sound field classification parameters;
[0187] The sound field classification results include: sound field type; or, the sound field classification results include: number of dissimilar sound sources and sound field type.
[0188] Step 403 determines the sound field classification result of the current frame based on the sound field classification parameters, including:
[0189] C1. Obtain the number of dissimilar sound sources corresponding to the current frame based on the values of multiple sound field classification parameters;
[0190] C2. Determine the sound field type based on the number of dissimilar sound sources corresponding to the current frame.
[0191] After the encoding end obtains multiple generation classification parameters corresponding to the current frame, it can determine the number of dissimilar sound sources in the current frame based on the values of these parameters. Dissimilar sound sources are point sound sources with different positions and / or directions. The number of dissimilar sound sources included in the current frame is called the dissimilar sound source count. The sound field of the current frame can be classified using the dissimilar sound source count. After determining the sound field type by obtaining the dissimilar sound source count for the current frame, the sound field type corresponding to the current frame can be determined by analyzing the dissimilar sound source count.
[0192] In some embodiments of this application, there are multiple sound field classification parameters;
[0193] The sound field classification results include: the number of dissimilar sound sources;
[0194] Step 403 determines the sound field classification result of the current frame based on the sound field classification parameters, including:
[0195] D1. Obtain the number of dissimilar sound sources corresponding to the current frame based on the values of multiple sound field classification parameters.
[0196] After the encoding end obtains multiple generation classification parameters corresponding to the current frame, it can obtain the number of dissimilar sound sources corresponding to the current frame through the values of multiple sound field classification parameters. Dissimilar sound sources are point sound sources with different positions and / or directions. The number of dissimilar sound sources included in the current frame is called the number of dissimilar sound sources.
[0197] Furthermore, in some embodiments of this application, multiple sound field classification parameters are temp[i], i = 0, 1, ..., min(L, K)-2, where L represents the number of channels in the current frame, K is the number of signal points corresponding to each channel in the current frame, and min represents the minimum value operation; for example, the number of signal points can be the number of frequency points, the number of time-domain samples, or the number of frequency points or time-domain samples after downsampling.
[0198] The aforementioned step C1 or D1 obtains the number of dissimilar sound sources corresponding to the current frame based on the values of multiple sound field classification parameters, including:
[0199] Starting from i=0, the following judgment process is executed sequentially:
[0200] Determine whether temp[i] is greater than the preset dissimilar sound source determination threshold;
[0201] If temp[i] is less than the dissimilar sound source determination threshold in the current judgment process, update the value of i to i+1 and continue to execute the next judgment process; or,
[0202] When temp[i] is greater than or equal to the dissimilar sound source determination threshold in this judgment process, the judgment process is terminated, and i plus 1 in this judgment process is determined to be equal to the number of dissimilar sound sources.
[0203] Specifically, the encoder can estimate the number of dissimilar sound sources and determine the sound field type based on the sound field classification parameters.
[0204] Sound field types can include dissimilar sound fields and diffuse sound fields. A dissimilar sound field refers to a sound field containing point sound sources with different locations and / or orientations. A diffuse sound field refers to a sound field that does not contain dissimilar sound sources.
[0205] If the values of the sound field classification parameters all satisfy the criteria for a diffuse sound field, then the sound field type is a diffuse sound field.
[0206] If any of the sound field classification parameters contains a value that satisfies the dissimilarity sound field decision condition, then the sound field type is a dissimilarity sound field. The number of dissimilarity sound sources can be estimated based on the index of the value that satisfies the dissimilarity sound field decision condition among the sound field classification parameters.
[0207] For example, the ratio between singular values, temp[i], is used as a sound field classification parameter. Based on the sound field classification parameter, the sound field type and the number of dissimilar sound sources are estimated. Starting from i=0, the value of temp[i] is judged sequentially. When the value of i is m, the value of the m-th sound field classification parameter is represented as temp[m]. When the m-th sound field classification parameter satisfies temp[m]≥TH1, the sound field type is a dissimilar sound field and there are (m+1) dissimilar sound sources in the sound field of the current frame. If temp[m]≥TH1 does not exist, the sound field type is a diffuse sound field. The value of m ranges from [0, 1, ..., min(L, K)-2], and TH1 is a pre-set dissimilar sound source determination threshold. The value of TH1 can be a constant, such as 30 or 100. In this embodiment, the value of TH1 is not limited.
[0208] In some embodiments of this application, step C2, which involves determining the sound field type based on the number of dissimilar sound sources corresponding to the current frame, includes:
[0209] When the number of dissimilar sound sources meets the first preset condition, the sound field type is determined to be the first sound field type;
[0210] When the number of dissimilar sound sources does not meet the first preset condition, the sound field type is determined to be the second sound field type.
[0211] The number of dissimilar sound sources corresponding to the first sound field type is different from the number of dissimilar sound sources corresponding to the second sound field type.
[0212] Specifically, the sound field types can be classified into two types according to the number of different sound sources: the first sound field type and the second sound field type. The encoding end obtains a first preset condition, determines whether the number of different sound sources meets the first preset condition. When the number of different sound sources meets the first preset condition, it is determined that the sound field type is the first sound field type; when the number of different sound sources does not meet the first preset condition, it is determined that the sound field type is the second sound field type. In the embodiments of the present application, by determining whether the number of different sound sources meets the first preset condition, the classification of the sound field type of the current frame can be realized, so that the sound field type of the current frame can be accurately identified as belonging to the first sound field type or the second sound field type.
[0213] In some embodiments of the present application, the first preset condition includes that the number of different sound sources is greater than a first threshold and less than a second threshold, where the second threshold is greater than the first threshold;
[0214] Or,
[0215] The first preset condition includes that the number of different sound sources is not greater than the first threshold or not less than the second threshold, where the second threshold is greater than the first threshold.
[0216] Among them, the specific values of the first threshold and the second threshold are not limited, and can be specifically combined with the application scenario. Since the second threshold is greater than the first threshold, the first threshold and the second threshold can constitute a preset range. Then the first preset condition can be that the number of different sound sources is within this preset range, or the first preset condition can be that the number of different sound sources is outside this preset range. Through the first threshold and the second threshold in the above first preset condition, the number of different sound sources can be judged to determine whether the number of different sound sources meets the first preset condition, so that the sound field type of the current frame can be accurately identified as belonging to the first sound field type or the second sound field type.
[0217] Illustrated as follows, the first threshold is 0, the second threshold is 3, and the number of different sound sources is represented as n. Then the first preset condition can be 0 < n < 3, or the first preset condition can be n >= 3 or n = 0.
[0218] In some embodiments of the present application, determining the sound field classification result of the current frame according to the sound field classification parameter may further include: determining the sound field classification result of the current frame according to the sound field classification parameter and other parameters characterizing the three-dimensional audio signal characteristics.
[0219] Among them, there are multiple implementation manners for other parameters characterizing the three-dimensional audio signal characteristics. For example, other parameters characterizing the three-dimensional audio signal characteristics may include at least one of the following: the energy ratio parameter of the three-dimensional audio signal, the high-frequency and low-frequency feature analysis parameter of the three-dimensional audio signal, etc.
[0220] Such as Figure 5As shown in the embodiments of this application, a method for processing three-dimensional audio signals mainly includes the following:
[0221] 501. Perform linear decomposition on the current frame of the three-dimensional audio signal to obtain the linear decomposition result.
[0222] 502. Obtain the sound field classification parameters corresponding to the current frame based on the linear decomposition results.
[0223] 503. Determine the sound field classification result of the current frame based on the sound field classification parameters.
[0224] The implementation of steps 501 to 503 is similar to that of steps 401 to 403 in the previous embodiment, and will not be described in detail here.
[0225] 504. Determine the encoding mode corresponding to the current frame based on the sound field classification results.
[0226] In this embodiment, the encoding end can execute steps 501 to 503 as described above. After obtaining the sound field classification result of the current frame, the encoding end can determine the encoding mode corresponding to the current frame based on the sound field classification result. This encoding mode refers to the mode used when encoding the current frame of the three-dimensional audio signal. There are multiple encoding modes, and different encoding modes can be used depending on the sound field classification result of the current frame. In this embodiment, a suitable encoding mode is selected for different sound field classification results of the current frame, and this encoding mode is used to encode the current frame, thereby improving the compression efficiency and auditory quality of the audio signal.
[0227] Furthermore, in some embodiments of this application, step 503, which determines the encoding mode corresponding to the current frame based on the sound field classification result, includes:
[0228] E1. When the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, the coding mode corresponding to the current frame is determined based on the number of dissimilar sound sources.
[0229] or,
[0230] E2. When the sound field classification result includes the sound field type, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the coding mode corresponding to the current frame based on the sound field type.
[0231] or,
[0232] E3. When the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the coding mode corresponding to the current frame based on the number of dissimilar sound sources and the sound field type.
[0233] In step E1 above, after the encoding end obtains the number of dissimilar sound sources in the current frame, this number can be used to determine the encoding mode corresponding to the current frame. In step E2 above, after the encoding end obtains the sound field type of the current frame, this sound field type can be used to determine the encoding mode corresponding to the current frame. In step E3 above, after the encoding end obtains the number of dissimilar sound sources and the sound field type of the current frame, these two factors can be used to determine the encoding mode corresponding to the current frame. Therefore, the encoding end can determine the encoding mode corresponding to the current frame through the number of dissimilar sound sources and / or the sound field type. This allows the encoding end to determine the corresponding encoding mode based on the sound field classification results of the current frame, ensuring that the determined encoding mode is compatible with the current frame of the 3D audio signal, thereby improving encoding efficiency.
[0234] Furthermore, in some embodiments of this application, step E1, determining the coding mode corresponding to the current frame based on the number of dissimilar sound sources, includes:
[0235] When the number of dissimilar sound sources meets the second preset condition, the coding mode is determined to be the first coding mode.
[0236] When the number of dissimilar sound sources does not meet the second preset condition, the coding mode is determined to be the second coding mode;
[0237] The first encoding mode is either a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional audio coding. The second encoding mode is either a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional audio coding, and the first and second encoding modes are different encoding modes. The HOA encoding mode based on virtual speaker selection can also be called a HOA encoding mode based on match projection (MP).
[0238] Specifically, the encoding mode can be divided into two types based on the number of dissimilar sound sources: a first encoding mode and a second encoding mode. The encoding end obtains a second preset condition and determines whether the number of dissimilar sound sources meets this condition. When the number of dissimilar sound sources meets the second preset condition, the encoding mode is determined to be the first encoding mode; when the number of dissimilar sound sources does not meet the second preset condition, the encoding mode is determined to be the second encoding mode. In this embodiment, the encoding mode of the current frame can be divided by determining whether the number of dissimilar sound sources meets the second preset condition, thereby accurately identifying whether the encoding mode of the current frame belongs to the first or second encoding mode.
[0239] For example, if the first encoding mode is a HOA encoding mode based on virtual speaker selection, the second encoding mode is a HOA encoding mode based on directional audio coding. Alternatively, if the first encoding mode is a HOA encoding mode based on directional audio coding, the second encoding mode is a HOA encoding mode based on virtual speaker selection. The specific implementation of the first and second encoding modes can be determined according to the application scenario.
[0240] For example, the sound field classification result in this application embodiment can determine the encoding mode selected by the encoder. For instance, the sound field classification result can be used to determine the encoding mode of the HOA signal. For example, the encoding mode can be determined based on the sound field type: HOA signals belonging to dissimilar sound fields are suitable for encoding with the encoder corresponding to encoding mode A, and HOA signals belonging to diffuse sound fields are suitable for encoding with the encoder corresponding to encoding mode B. Another example is the encoding mode determined based on the number of dissimilar sound sources: when the number of dissimilar sound sources meets the decision condition for using encoding mode X, the encoder corresponding to encoding mode X is used for encoding. Yet another example is the encoding mode also determined based on the sound field type and the number of dissimilar sound sources: when the sound field type is a diffuse sound field, the encoder corresponding to encoding mode C is used for encoding; when the sound field type is a dissimilar sound field and the number of dissimilar sound sources meets the decision condition for using encoding mode X, the encoder corresponding to encoding mode X is used for encoding. Encoding modes A, B, C, and X can include multiple different encoding modes. Different sound field classification results in this application embodiment correspond to different encoding modes, and this application embodiment does not impose any limitations on these modes. For example, encoding mode X can be encoding mode 1 when the number of dissimilar sound sources is less than a preset threshold, and encoding mode 2 when the number of dissimilar sound sources is greater than or equal to the preset threshold.
[0241] In some embodiments of this application, the second preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold;
[0242] or,
[0243] The second preset condition includes that the number of dissimilar sound sources is not greater than the first threshold or not less than the second threshold, wherein the second threshold is greater than the first threshold.
[0244] Among them, the specific values of the first threshold and the second threshold are not limited and can be specifically combined with the application scenario. The second threshold is greater than the first threshold. Therefore, the first threshold and the second threshold can form a preset range. Then the second preset condition can be that the number of different sound sources is within this preset range, or the second preset condition can be that the number of different sound sources is outside this preset range. Through the first threshold and the second threshold in the above second preset condition, the number of different sound sources can be judged to determine whether the number of different sound sources meets the second preset condition, so that the sound field type of the current frame can be accurately identified as the first sound field type or the second sound field type.
[0245] For example, the first threshold is 0, the second threshold is 3, and the number of different sound sources is represented as n. Then the second preset condition can be 0 < n < 3, or the second preset condition can be n >= 3 or n = 0.
[0246] It should be noted that in the embodiments of the present application, the first preset condition is a condition set for identifying different sound field types, and the second preset condition is a condition set for identifying different coding modes. The first preset condition and the second preset condition can include the same condition content or different condition content. That is, the first preset condition and the second preset condition can be different preset conditions, or the first preset condition and the second preset condition can be the same preset condition. However, considering the possible differences in actual use, the first preset condition and the second preset condition are distinguished by the first and the second.
[0247] In some embodiments of the present application, step E2 determines the coding mode corresponding to the current frame according to the sound field type, including:
[0248] When the sound field type is a different sound field, determine that the coding mode is the HOA coding mode based on the selection of virtual speakers;
[0249] When the sound field type is a diffused sound field, determine that the coding mode is the HOA coding mode based on the direction audio coding.
[0250] Among these, the HOA encoding mode based on directional audio has lower compression efficiency than the HOA encoding mode based on virtual speakers when there are few dissimilar sound sources in the sound field or when the sound field is diffuse. Conversely, when there are many dissimilar sound sources in the sound field, the HOA encoding mode based on virtual speakers is less efficient than the HOA encoding mode based on directional audio. In this embodiment, when the sound field type is dissimilar, the encoding mode is determined to be the HOA encoding mode based on virtual speakers; when the sound field type is diffuse, the encoding mode is determined to be the HOA encoding mode based on directional audio encoding. This embodiment can select the appropriate encoding mode based on the sound field classification result of the current frame to meet the need for maximum compression efficiency for different types of audio signals.
[0251] In some embodiments of this application, step 503, which determines the encoding mode corresponding to the current frame based on the sound field classification result, includes:
[0252] F1. Determine the initial coding mode corresponding to the current frame based on the sound field classification result of the current frame;
[0253] F2. Get the current frame's hangover. The hangover includes the initial encoding mode of the current frame and the encoding modes of the N-1 frames preceding the current frame, where N is the length of the hangover.
[0254] F3. Determine the encoding mode of the current frame based on the initial encoding mode of the current frame and the encoding mode of frame N-1.
[0255] In step F1, the initial encoding mode can be a mode determined based on the sound field classification result. For example, the encoding mode of the current frame can be determined according to any of the implementation methods in steps E1 to E3, and this encoding mode can be used as the initial encoding mode in F1. After obtaining the initial encoding mode, a sliding window is obtained based on the current frame and the window size of the sliding window. This sliding window includes the initial encoding mode of the current frame and the encoding modes of the N-1 frames preceding the current frame, where N represents the number of frames included in the sliding window. Finally, the encoding mode of the current frame is determined based on the encoding modes corresponding to the N frames within the sliding window. The encoding mode of the current frame obtained in step F3 can be the encoding mode used when encoding the current frame. In this embodiment, the initial encoding mode of the current frame is corrected by the sliding window to obtain the encoding mode of the current frame, so as to ensure that the encoding mode between consecutive frames does not switch frequently and improve encoding efficiency.
[0256] For example, after obtaining the initial encoding mode of the current frame, a sliding window process can be applied to the current frame to ensure that the encoding mode does not switch frequently between consecutive frames. There are many sliding window processing methods, and this application does not limit them. For example, one processing method is to store encoder selection identifiers of length N frames within the sliding window, where N frames include the encoder selection identifiers of the current frame and the previous N-1 frames; when the encoder selection identifiers accumulate to a specified threshold, the encoding type indicator identifier of the current frame is updated. Optionally, in addition to sliding window processing, other post-processing can also be used to correct the current frame. For example, the initial encoding mode can be used as the initial classification, and the initial classification can be corrected based on the speech classification result of the audio signal, signal-to-noise ratio, and other characteristics, and the corrected result can be used as the final encoding mode result.
[0257] like Figure 6 As shown in the embodiments of this application, a method for processing three-dimensional audio signals mainly includes the following:
[0258] 601. Perform linear decomposition on the current frame of the three-dimensional audio signal to obtain the linear decomposition result.
[0259] 602. Obtain the sound field classification parameters corresponding to the current frame based on the linear decomposition results.
[0260] 603. Determine the sound field classification result of the current frame based on the sound field classification parameters.
[0261] The implementation methods of steps 601 to 603 are similar to those of steps 401 to 403 in the previous embodiments, and will not be described in detail here.
[0262] 604. Determine the encoding parameters corresponding to the current frame based on the sound field classification results.
[0263] In this embodiment, the encoding end can execute steps 601 to 603 as described above. After obtaining the sound field classification result of the current frame, the encoding end can determine the encoding parameters corresponding to the current frame based on the sound field classification result. These encoding parameters refer to the parameters used when encoding the current frame of the three-dimensional audio signal. There are various encoding parameters, and different encoding parameters can be used depending on the sound field classification result of the current frame. In this embodiment, appropriate encoding parameters are selected for different sound field classification results of the current frame, and these encoding parameters are used to encode the current frame, thereby improving the compression efficiency and auditory quality of the audio signal.
[0264] Furthermore, in some embodiments of this application, the encoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of encoded bits of the virtual speaker signal, the number of encoded bits of the residual signal, or the number of voting rounds for the best matching speaker search;
[0265] Among them, the virtual speaker signal and the residual signal are signals generated based on the three-dimensional audio signal.
[0266] Specifically, the encoding end can determine the encoding parameters of the current frame based on the sound field classification results, and then use these encoding parameters to encode the current frame. Encoding parameters can be implemented in various ways, including at least one of the following: the number of channels for the virtual speaker signal, the number of channels for the residual signal, the number of encoded bits for the virtual speaker signal, the number of encoded bits for the residual signal, or the number of voting rounds in the best-matching speaker search. Here, the number of channels can also be called the number of transmission channels; the number of channels is the number of transmission channels allocated during signal encoding, and the number of encoded bits is the number of encoded bits allocated during signal encoding.
[0267] This application provides a method for selecting a virtual loudspeaker. The encoder uses the virtual loudspeaker coefficients of the current frame to vote on each virtual loudspeaker in the candidate virtual loudspeaker set, and selects the virtual loudspeaker of the current frame based on the voting values, thereby reducing the computational complexity of virtual loudspeaker search and alleviating the computational burden on the encoder. The number of voting rounds for the best-matching loudspeaker search refers to the number of voting rounds required to search for the best-matching loudspeaker. In one possible implementation, the number of voting rounds can be pre-configured or determined based on the sound field classification results of the current frame. For example, the number of voting rounds for the best-matching loudspeaker search is the number of voting rounds performed during the virtual loudspeaker search process when determining the virtual loudspeaker signal based on the three-dimensional audio signal.
[0268] Furthermore, the virtual speaker signal and residual signal in this embodiment are signals generated based on the three-dimensional audio signal. For example, a first target virtual speaker is selected from a preset set of virtual speakers based on the first scene audio signal; a virtual speaker signal is generated based on the attribute information of the first scene audio signal and the first target virtual speaker; a second scene audio signal is obtained using the attribute information of the first target virtual speaker and the first virtual speaker signal; and a residual signal is generated based on the first scene audio signal and the second scene audio signal.
[0269] In some embodiments of this application, the number of voting rounds satisfies the following relationship:
[0270] 1≤I≤d,
[0271] Where I represents the number of voting rounds, and d represents the number of dissimilar sound sources included in the sound field classification results.
[0272] In this process, the encoding end determines the number of voting rounds for the best matching loudspeaker search based on the number of dissimilar sound sources in the current frame. This number of voting rounds is less than or equal to the number of dissimilar sound sources in the current frame, thus ensuring that the number of voting rounds conforms to the actual situation of the sound field classification in the current frame. This solves the problem of determining the number of voting rounds for the best matching loudspeaker search when encoding the current frame.
[0273] The voting round number I should follow these principles: the minimum number of voting rounds is one, and the maximum number of voting rounds cannot exceed the total number of speakers or the number of virtual speaker signal channels. For example, the total number of speakers could be 1024 speakers generated by the virtual speaker set generation unit in the encoder. The number of virtual speaker signal channels is the virtual speaker signal to be transmitted by the encoder, which is the N transmission channels generated corresponding to the N best-matched speakers. Typically, the number of virtual speaker signal channels is less than the total number of speakers. The voting round number is estimated as follows: the number of voting rounds I for the best-matched speaker search is determined based on the number of dissimilar sound sources in the sound field of the current frame obtained from the sound field classification results. The voting round number I satisfies the following relationship: 1 ≤ I ≤ d, where d is the number of sound sources in different directions in the sound field, i.e., the estimated number of dissimilar sound sources in the sound field classification results. For example, I = d. Or, the voting round number I = min(d, total number of speakers, number of virtual speaker signal channels, preset number of voting rounds). The number of voting rounds I can be obtained by using min(d, total number of speakers, number of virtual speaker signal channels, preset number of voting rounds), so that the encoding end can determine the number of voting rounds for the best matching speaker search according to the value of I.
[0274] In some embodiments of this application, the sound field classification results include the number of dissimilar sound sources and the sound field type;
[0275] When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0276] F = min(S, PF),
[0277] Where F is the number of virtual speaker signal channels, S is the number of dissimilar sound sources, and PF is the preset number of virtual speaker signal channels by the encoder; or,
[0278] When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0279] F = 1,
[0280] Where F is the number of channels for the virtual speaker signal.
[0281] The number of channels for the virtual speaker signal refers to the number of channels used to transmit the virtual speaker signal. This number can be determined by the number of dissimilar sound sources and the sound field type. In the above calculation method, when the sound field type is a diffuse sound field, the number of channels for the virtual speaker signal is set to 1, thus improving the coding efficiency for the current frame. When the sound field type is a dissimilar sound field, "min" represents the minimum value operation, i.e., taking the minimum value from S and PF as the number of channels for the virtual speaker signal. This ensures that the number of channels for the virtual speaker signal conforms to the actual sound field classification of the current frame, solving the problem of determining the number of channels for the virtual speaker signal when encoding the current frame.
[0282] In some embodiments of this application, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship:
[0283] R = max(C-1, PR),
[0284] Wherein, PR is the preset number of residual signal channels of the encoder, and C is the sum of the preset number of residual signal channels and the preset number of virtual speaker signal channels of the encoder; or,
[0285] When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship:
[0286] R = C – F,
[0287] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the encoder and the number of virtual speaker signal channels preset by the encoder, and F is the number of channels of the virtual speaker signal.
[0288] After obtaining the number of channels for the virtual speaker signal, the number of channels for the residual signal can be calculated based on the sum of the preset number of channels for the residual signal and the preset number of channels for the virtual speaker signal, as well as the preset number of channels for the residual signal. The value of PR can be preset at the encoding end. The value of R can be obtained through the above calculation formula max(C-1, PR). The preset number of channels for the residual signal and the preset number of channels for the virtual speaker signal are preset at the encoding end. In addition, C can also be simply referred to as the total number of transmission channels.
[0289] In some embodiments of this application, after obtaining the number of channels of the virtual speaker signal, the number of channels of the residual signal can be calculated based on the sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal, and the number of channels of the virtual speaker signal. The preset number of channels of the residual signal and the preset sum of channels of the virtual speaker signal are preset by the encoding end. In addition, C mentioned above can also be simply referred to as the total number of transmission channels.
[0290] In some embodiments of this application, the sound field classification results include the number of dissimilar sound sources;
[0291] The number of channels for the virtual speaker signal satisfies the following relationship:
[0292] F = min(S, PF),
[0293] Where F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the encoder.
[0294] The number of channels for the virtual speaker signal refers to the number of channels used to transmit the virtual speaker signal. The number of channels for the virtual speaker signal can be determined by the number of dissimilar sound sources. In the above calculation method, min means taking the minimum value operation, that is, taking the minimum value from S and PF as the number of channels for the virtual speaker signal, so that the number of channels for the virtual speaker signal can conform to the actual situation of the sound field classification of the current frame, and solving the problem of needing to determine the number of channels for the virtual speaker signal when encoding the current frame.
[0295] In some embodiments of this application, the number of channels of the residual signal satisfies the following relationship:
[0296] R = C – F,
[0297] Where R represents the number of channels of the residual signal, C is the sum of the number of channels of the encoder's preset residual signal and the number of channels of the encoder's preset virtual speaker signal, and F is the number of channels of the virtual speaker signal. For example, C is the sum of PF and PR mentioned above.
[0298] After obtaining the number of channels for the virtual speaker signal, the number of channels for the residual signal can be calculated based on the sum of the preset number of channels for the residual signal and the preset number of channels for the virtual speaker signal, as well as the number of channels for the virtual speaker signal. This preset sum of the preset number of channels for the residual signal and the preset number of channels for the virtual speaker signal is preset at the encoding end. Furthermore, C can also be simply referred to as the total number of transmission channels.
[0299] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type;
[0300] The number of encoded bits of the virtual speaker signal is obtained by the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel;
[0301] The number of coded bits of the residual signal is obtained by the ratio of the number of coded bits of the virtual speaker signal to the number of coded bits of the transmission channel;
[0302] The number of encoded bits in the transmission channel includes the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel is obtained by increasing the initial ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel.
[0303] In this method, the encoding end presets an initial ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel. The encoding end acquires the number of dissimilar sound sources and determines whether this number is less than or equal to the number of channels for the virtual speaker signal. If the number of dissimilar sound sources is less than or equal to the number of channels for the virtual speaker signal, the initial ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel can be increased. This increased initial ratio is defined as the ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel. This ratio can be used to calculate the number of encoded bits for the virtual speaker signal and the number of encoded bits for the transmission channel. It can also be used to calculate the number of encoded bits for the residual signal. This calculation method ensures that the number of encoded bits for the virtual speaker signal and the residual signal conforms to the actual sound field classification of the current frame, solving the problem of needing to determine the number of encoded bits for the virtual speaker signal and the residual signal when encoding the current frame.
[0304] For example, the encoding end determines the bit allocation method for the virtual speaker signal and residual signal based on the sound field classification results. The transmission channel signal is divided into a virtual speaker signal group and a residual signal group. The pre-set allocation ratio of the virtual speaker signal group is used as the initial ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel. When the number of dissimilar sound sources is less than or equal to the number of channels for the virtual speaker signal, the initial ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel is increased according to a preset adjustment value. This increased ratio is then used as the ratio of the number of encoded bits for the virtual speaker signal to the number of encoded bits for the transmission channel. For example, the increased ratio is equal to the sum of the preset adjustment value and the initial ratio.
[0305] In some embodiments of this application, the ratio of the number of encoded bits of the residual signal to the number of encoded bits of the transmission channel is 1.0 - the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel.
[0306] In some embodiments of this application, in addition to performing the aforementioned steps, the method performed by the encoding end may also include the following steps:
[0307] The current frame and sound field classification results are encoded and written into the bitstream.
[0308] The sound field classification result can be encoded into a bitstream. After the encoding end sends the bitstream to the decoding end, the decoding end can obtain the sound field classification result through the bitstream. By parsing the bitstream, the decoding end can obtain the sound field classification result carried in the bitstream. The decoding end can obtain the sound field distribution of the current frame through the sound field classification result, and thus can decode the current frame to obtain a three-dimensional audio signal.
[0309] In some embodiments of this application, the current frame and the sound field classification result are encoded. Specifically, this may include directly encoding the current frame, or processing the current frame first, and then encoding the virtual speaker signal and the residual signal after obtaining them. For example, the encoding end may be a core encoder, which encodes the virtual speaker signal, the residual signal, and the sound field classification result to obtain a bitstream. This bitstream can also be called an audio signal encoded bitstream.
[0310] The three-dimensional audio signal processing method provided in this application embodiment may include: an audio encoding method and an audio decoding method, wherein the audio encoding method is executed by an audio encoding device, the audio decoding method is executed by an audio decoding device, and the audio encoding device and the audio decoding device can communicate with each other. (The foregoing...) Figures 4 to 6 The processing method for three-dimensional audio signals, executed by the audio decoding device (hereinafter referred to as the decoding end) in the embodiments of this application, is described below. Figure 7 As shown, the main steps include the following:
[0311] 701. Receive bitstream.
[0312] The decoding end receives the bitstream from the encoding end. This bitstream carries the sound field classification results.
[0313] 702. Decode the bitstream to obtain the sound field classification result of the current frame.
[0314] The decoding end parses the bitstream and obtains the sound field classification result of the current frame from it. This sound field classification result is determined by the encoding end according to the aforementioned... Figures 4 to 6 The example shown is obtained.
[0315] 703. Obtain the three-dimensional audio signal after decoding the current frame based on the sound field classification results.
[0316] After obtaining the sound field classification result, the decoding end uses the sound field classification result to parse the bitstream and obtain the decoded three-dimensional audio signal of the current frame. In this embodiment, the decoding process for the current frame is not limited. In this embodiment, the decoding end can decode the current frame using the sound field classification result, which can be used to decode the current frame in the bitstream. Therefore, the decoding end uses a decoding method that matches the sound field of the current frame to decode, thereby obtaining the three-dimensional audio signal sent by the encoding end, realizing the transmission of the audio signal from the encoding end to the decoding end.
[0317] For example, the decoding end can determine the same decoding mode and / or decoding parameters as the encoding end based on the sound field classification results transmitted in the bitstream, which reduces the number of encoding bits compared to the way the encoding end transmits the encoding mode and / or encoding parameters to the decoding end.
[0318] In some embodiments of this application, step 703, obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result, includes:
[0319] G1. Determine the decoding mode of the current frame based on the sound field classification results;
[0320] G2. Obtain the three-dimensional audio signal after decoding the current frame according to the decoding mode.
[0321] The decoding mode corresponds to the encoding mode in the aforementioned embodiments. The implementation of step G1 is similar to step 504 in the aforementioned embodiments, and will not be repeated here. After obtaining the decoding mode, the decoding end can decode the bitstream according to the decoding mode to obtain the three-dimensional audio signal after decoding the current frame.
[0322] Furthermore, in some embodiments of this application, step G1, which determines the decoding mode of the current frame based on the sound field classification result, includes:
[0323] When the sound field classification result includes the number of dissimilar sound sources, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined according to the number of dissimilar sound sources.
[0324] or,
[0325] When the sound field classification result includes a sound field type, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined according to the sound field type;
[0326] or,
[0327] When the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined based on the number of dissimilar sound sources and the sound field type.
[0328] The above implementation method is similar to the implementation method of steps E1 to E3 in the previous embodiment, and will not be repeated here.
[0329] In some embodiments of this application, determining the decoding mode corresponding to the current frame based on the number of dissimilar sound sources includes:
[0330] When the number of dissimilar sound sources meets a preset condition, the decoding mode is determined to be the first decoding mode;
[0331] When the number of dissimilar sound sources does not meet the preset condition, the decoding mode is determined to be the second decoding mode;
[0332] The first decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio encoding, and the second decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio encoding, and the first decoding mode and the second decoding mode are different decoding modes.
[0333] It should be noted that this preset condition is set by the decoding end to identify different decoding modes, and there are no restrictions on how the preset condition is implemented.
[0334] In some embodiments of this application, the preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold;
[0335] or
[0336] The preset conditions include that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
[0337] In some embodiments of this application, step 703, obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result, includes:
[0338] H1. Determine the decoding parameters of the current frame based on the sound field classification results;
[0339] H2. Obtain the three-dimensional audio signal after decoding the current frame according to the decoding parameters.
[0340] The decoding parameters correspond to the encoding parameters in the aforementioned embodiments. The implementation of step H1 is similar to step 604 in the aforementioned embodiments, and will not be repeated here. After obtaining the decoding parameters, the decoding end can decode the bitstream according to the decoding parameters to obtain the three-dimensional audio signal after decoding the current frame.
[0341] In some embodiments of this application, the decoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signal, or the number of decoded bits of the residual signal;
[0342] The virtual speaker signal and the residual signal are obtained by decoding the bitstream.
[0343] In some embodiments of this application, the sound field classification results include the number of dissimilar sound sources and the sound field type;
[0344] When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0345] F = min(S, PF),
[0346] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder; or
[0347] When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0348] F = 1,
[0349] Wherein, F is the number of channels of the virtual speaker signal.
[0350] In some embodiments of this application, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship:
[0351] R = max(C-1, PR),
[0352] Wherein, PR is the number of residual signal channels preset by the decoder, and C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder; or,
[0353] When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship:
[0354] R = C – F,
[0355] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0356] It should be noted that the number of virtual speaker signal channels preset by the decoder is equal to the number of virtual speaker signal channels preset by the encoder. Similarly, the number of residual signal channels preset by the decoder is equal to the number of residual signal channels preset by the encoder.
[0357] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources;
[0358] The number of channels of the virtual speaker signal satisfies the following relationship:
[0359] F = min(S, PF),
[0360] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder.
[0361] In some embodiments of this application, the number of channels of the residual signal satisfies the following relationship:
[0362] R = C – F,
[0363] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0364] It should be noted that the implementation of the above decoding parameters is similar to the implementation of the encoding parameters in the aforementioned embodiments, and will not be described in detail here.
[0365] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type;
[0366] The number of decoded bits of the virtual speaker signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel;
[0367] The number of decoded bits of the residual signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel;
[0368] The number of decoded bits in the transmission channel includes the number of decoded bits of the virtual speaker signal and the number of decoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel is obtained by increasing the initial ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel.
[0369] To facilitate a better understanding and implementation of the above-described solutions in the embodiments of this application, specific examples of corresponding application scenarios are provided below.
[0370] In this embodiment, the three-dimensional audio signal is taken as a HOA signal. The sound field classification method of the HOA signal in this embodiment is applied to a hybrid HOA encoder. The basic encoding process is as follows: Figure 8 As shown, the encoding end classifies the HOA signal to be encoded to determine whether the HOA signal to be encoded in the current frame is suitable for a virtual speaker-based HOA encoding scheme or a DirAC-based HOA encoding scheme, and determines the HOA encoding mode of the current frame based on the sound field classification result. Specifically, the HOA encoder includes an encoder selection unit, which performs sound field classification on the HOA signal to be encoded and determines the encoding mode of the current frame; based on the encoding mode, encoder A or encoder B is selected for encoding to obtain the final encoded bitstream. Here, encoder A and encoder B represent different types of encoders, each of which is adapted to a sound field type of the current frame. When an encoder adapted to the sound field type is used for encoding, the signal compression ratio can be improved.
[0371] The specific process of classifying the sound field of the HOA signal to be encoded and determining the encoding mode includes:
[0372] The sound field classification of the HOA signal to be encoded is performed to obtain the sound field classification results.
[0373] Based on the sound field classification results, determine the encoding mode of the current frame.
[0374] The encoding mode of the current frame indicates the encoder selection method for the current frame. The criteria for determining the encoder selection identifier can be based on the sound field type of the HOA signal to which encoder A and encoder B are applicable. For example, encoder A processes HOA signals with dissimilar sound fields and fewer than 3 dissimilar sound sources, while encoder B processes HOA signals with dissimilar sound fields and 3 or more dissimilar sound sources. Alternatively, encoder B processes signals with diffuse sound fields or HOA signals with 3 or more dissimilar sound sources.
[0375] It is important to note that a hanging over process can also be applied to the sound field classification results to ensure that the encoding mode does not switch frequently between consecutive frames. There are many hanging over processing methods, and this application does not limit them. For example, one approach is to store encoder selection identifiers of length N frames within the hanging over window, where N frames include the encoder selection identifiers of the current frame and the previous N-1 frames; when the encoder selection identifiers accumulate to a specified threshold, the encoding type indicator of the current frame is updated. Optionally, in addition to the hanging over process, other processing methods can be used to correct the sound field classification results.
[0376] like Figure 9 As shown, the process for determining the encoding mode of the HOA signal mainly includes:
[0377] S01. Obtain the HOA signal to be analyzed.
[0378] S02, downsample the HOA signal.
[0379] Not limited to, downsampling the HOA signal to be analyzed is an optional step.
[0380] By downsampling the HOA signal to be analyzed, the computational complexity can be reduced. The HOA signal to be analyzed can be a time-domain HOA signal or a frequency-domain HOA signal. The HOA signal to be analyzed can contain all channels or only some HOA channels (e.g., FOA channels). For example, the HOA signal to be analyzed can be all samples or 1 / Q downsampled points; in this embodiment, 1 / 120 downsampled points are used.
[0381] For example, the HOA signal of the current frame is of order 3, has 16 channels, and the frame length is 20 milliseconds (ms). This means the current frame signal contains 960 samples. After the HOA signal to be encoded in the current frame is downsampled by 1 / 120, each channel contains 8 samples. Therefore, the HOA signal has 16 channels, each with 8 samples, forming the input signal for sound field type analysis, i.e., the HOA signal to be analyzed.
[0382] S03. Analyze the sound field type based on the downsampled signal.
[0383] After downsampling the HOA signal, the sound field type is obtained by analyzing the number of dissimilar sound sources in the HOA signal.
[0384] For example, in the embodiments of this application, the sound field type analysis can be to perform linear decomposition on the HOA signal, obtain the linear decomposition result through the linear decomposition, and then obtain the sound field classification result through the linear decomposition result.
[0385] For example, the number of dissimilar sound sources can be obtained based on the linear decomposition result. For example, the linear decomposition result can include eigenvalues, and the number of dissimilar sound sources is estimated through the ratio between the eigenvalues, specifically including:
[0386] Perform singular value decomposition on the HOA signal to be analyzed to obtain singular values v[i], where i = 0, 1, … min(L, K) - 1.
[0387] Among them, L is equal to the number of channels of the HOA signal, K is the number of signal points of each channel in the current frame. For example, the number of signal points can be the number of frequency points. In this embodiment, L = 16, K = 8, and min(L, K) = 8.
[0388] Calculate the ratio temp[i] between the singular values v as the sound field classification parameter, where i = 0, 1, … min(L, K) - 2:
[0389] temp[i] = v[i] / v[i + 1].
[0390] The determination threshold for dissimilar sound sources is 100, and the estimated number of dissimilar sound sources n can be obtained through the following method:
[0391] Starting from i = 0, determine whether temp[i] is greater than or equal to 100. If temp[i] is greater than or equal to 100, satisfying temp[i] ≥ 100, then stop the determination; otherwise, i = i + 1 and continue the determination. When the determination stops, the sequence number i plus 1 when the determination stops is equal to the number of dissimilar sound sources n. For example, when i = 0, if temp[0] ≥ 100, then stop the determination, and the number of dissimilar sound sources n is equal to 1; otherwise, let i = 1 and continue the determination for i = 1; when i = 1, if temp[1] ≥ 100, then stop the determination, and the number of dissimilar sound sources n is equal to i + 1 = 2.
[0392] S04. Determine the predicted coding mode according to the analysis result of the sound field type.
[0393] Determine the predicted coding mode according to the number of dissimilar sound sources n:
[0394] When 0 < n < 3, the predicted coding mode is coding mode 1;
[0395] When n ≥ 3 or n = 0, the predicted coding mode is coding mode 2.
[0396] For example, coding mode 1 can be an HOA coding scheme based on virtual speaker selection. Coding mode 2 can be an HOA coding scheme based on Directional Audio DirAC.
[0397] S05. Determine the actual coding mode according to the predicted coding mode.
[0398] After determining the expected coding mode of the current frame, the next step is to determine the actual coding mode. For example, a sliding window can be used to determine the actual coding mode. Within the sliding window, when the expected coding mode 2 of multiple frames within the sliding window accumulates to a specified threshold, the actual coding mode of the current frame is coding mode 2; otherwise, the actual coding mode of the current frame is coding mode 1.
[0399] For example, the sliding window contains 10 frames of expected coding mode results, including the coding mode decision result of the current frame in step S03 and the coding mode results of the previous 9 frames. If the expected coding mode of the 10 frames is coding mode 2, the actual coding mode of the current frame is determined to be coding mode 2.
[0400] S06. Obtain the final encoding mode.
[0401] The basic decoding process of a hybrid HOA decoder corresponding to the encoding end is as follows: Figure 10 As shown: The decoding end obtains the bitstream from the encoding end, and then parses the HOA decoding mode of the current frame based on the bitstream. According to the HOA decoding mode of the current frame, the corresponding decoding scheme is selected for decoding to obtain the reconstructed HOA signal. Specifically, the decoding end includes a decoder selection unit, which parses the bitstream to determine the decoding mode; based on the decoding mode, decoder A or decoder B is selected for decoding to obtain the reconstructed HOA signal. Here, decoder A and decoder B represent different types of decoders, each adapted to a sound field type of the current frame. When a decoder adapted to the sound field type is used for decoding, the HOA signal can be correctly reconstructed.
[0402] As explained above, the sound field classification results of the HOA signal to be encoded, and the encoding mode determined based on the sound field classification results, can be matched with the signal types suitable for different encoding modes, so that different types of signals can obtain the maximum compression efficiency.
[0403] The following describes the HOA encoder based on virtual speaker selection provided in the embodiments of this application. The basic encoding process is as follows: Figure 11 As shown.
[0404] The encoding end may include: a virtual loudspeaker configuration unit, an encoding analysis unit, a virtual loudspeaker set generation unit, a virtual loudspeaker selection unit, a virtual loudspeaker signal generation unit, a core encoder processing unit, a signal reconstruction unit, a residual signal generation unit, a selection unit, and a signal compensation unit. The functions of each component unit of the encoding end will be described below. In this embodiment, Figure 11The encoding end shown can generate one virtual speaker signal or multiple virtual speaker signals. The generation process for multiple virtual speaker signals can be based on... Figure 11 The encoder structure shown is generated multiple times. The following example uses the generation process of a virtual speaker signal.
[0405] The virtual speaker configuration unit is used to configure the virtual speakers in the virtual speaker set to obtain multiple virtual speakers.
[0406] The virtual speaker configuration unit outputs virtual speaker configuration parameters based on the encoder configuration information. The encoder configuration information includes, but is not limited to, HOA order, encoding bit rate, and user-defined information. The virtual speaker configuration parameters include, but are not limited to, the number of virtual speakers, the HOA order of the virtual speakers, and the position coordinates of the virtual speakers.
[0407] The virtual speaker configuration parameters output by the virtual speaker configuration unit are used as input to the virtual speaker set generation unit.
[0408] The encoding analysis unit is used to perform encoding analysis on the HOA signal to be encoded, such as analyzing the sound field distribution of the HOA signal to be encoded, including the number of sound sources, directionality, dispersion and other characteristics of the HOA signal to be encoded, as one of the judgment conditions for deciding how to select the target virtual loudspeaker.
[0409] Not limited to, in the embodiments of this application, the encoding end may not include an encoding analysis unit, that is, the encoding end may not analyze the input signal, and a default configuration is used to determine how to select the target virtual speaker.
[0410] The encoder acquires the HOA signal to be encoded. For example, the HOA signal recorded from the actual acquisition device or the HOA signal synthesized using artificial audio objects can be used as the input of the encoder. The HOA signal to be encoded input to the encoder can be a time-domain HOA signal or a frequency-domain HOA signal.
[0411] The virtual speaker set generation unit is used to generate a virtual speaker set, which may include multiple virtual speakers. The virtual speakers in the virtual speaker set may also be referred to as "candidate virtual speakers".
[0412] The virtual speaker set generation unit generates specified candidate virtual speaker HOA coefficients based on the virtual speaker configuration parameters. Generating candidate virtual speaker HOA coefficients requires the coordinates (i.e., position coordinates or position information) of the candidate virtual speakers and their HOA order. Methods for determining the coordinates of candidate virtual speakers include, but are not limited to, generating K virtual speakers according to an equidistant rule, or generating K candidate virtual speakers with a non-uniform distribution based on auditory perception principles. The following example illustrates a method for generating a uniformly distributed fixed number of virtual speakers.
[0413] The coordinates of the candidate virtual speakers are generated based on the number of candidate virtual speakers, and an approximately uniform speaker arrangement can be given by using numerical iterative calculation methods.
[0414] The HOA coefficients of the candidate virtual loudspeakers output by the virtual loudspeaker set generation unit are used as inputs to the virtual loudspeaker selection unit.
[0415] The virtual speaker selection unit is used to select a target virtual speaker from multiple candidate virtual speakers in the set of virtual speakers based on the HOA signal to be encoded. The target virtual speaker can be referred to as the "virtual speaker that matches the HOA signal to be encoded", or simply as the matched virtual speaker.
[0416] The virtual loudspeaker selection unit matches the HOA signal to be encoded with the candidate virtual loudspeaker HOA coefficients output by the virtual loudspeaker set generation unit, and selects the specified matching virtual loudspeaker.
[0417] In this embodiment of the application, the sound field classification of the HOA signal to be encoded is performed, and the encoding parameters are determined based on the sound field classification results.
[0418] The encoding analysis unit performs encoding analysis on the HOA signal to be encoded. This analysis may include: classifying the sound field based on the HOA signal to be encoded. The sound field classification method is detailed in the aforementioned embodiments and will not be repeated here.
[0419] Based on the sound field classification results, the encoding parameters are determined. The encoding parameters may include at least one of the following: the number of channels of the virtual speaker signal in the HOA encoding scheme based on virtual speaker selection, the number of channels of the residual signal, and the number of voting rounds in the best-matching speaker search.
[0420] Specifically, the virtual loudspeaker selection unit, based on the number of voting rounds in the determined best-matching loudspeaker search and the number of channels in the virtual loudspeaker signal, matches the HOA coefficients to be encoded with the candidate virtual loudspeaker HOA coefficients output by the virtual loudspeaker set generation unit, selects the best-matching virtual loudspeaker, and obtains the matching virtual loudspeaker HOA coefficients. The number of best-matching virtual loudspeakers is equal to the number of channels in the virtual loudspeaker signal.
[0421] The virtual loudspeaker selection unit uses a voting-based best-match loudspeaker search method to match the HOA coefficients to be encoded with the candidate virtual loudspeaker HOA coefficients output by the virtual loudspeaker set generation unit, and selects the best-matching virtual loudspeaker. The number of voting rounds I for the best-matching loudspeaker search can be determined based on the sound field classification results.
[0422] The number of voting rounds I should follow the following principles: the minimum number of voting rounds is one, and the maximum number cannot exceed the total number of speakers (e.g., 1024 speakers obtained by the virtual speaker set generation unit) and the number of virtual speaker signal channels (the virtual speaker signals to be transmitted by the encoder, that is, the N transmission channels generated corresponding to the N best-matched speakers). In general, the number of virtual speaker signal channels is less than the total number of speakers.
[0423] The method for estimating the number of voting rounds is as follows:
[0424] Based on the number of dissimilar sound sources in the sound field obtained from the sound field classification results, determine the number of voting rounds I for loudspeaker selection.
[0425] The number of voting rounds I satisfies 1 ≤ I ≤ d, where d is the number of sound sources in different directions in the sound field, i.e., the estimated number of dissimilar sound sources in the sound field classification results. For example, I = d.
[0426] The number of channels for the virtual loudspeaker signal and the number of channels for the residual signal are determined based on the sound field type.
[0427] Next, this application provides a method for selecting the number of channels F of an adaptive virtual speaker signal:
[0428] When the sound field type is a dissimilar sound field, F = min(S, PF), where S is the number of dissimilar sound sources in the sound field, and PF is the number of virtual speaker signal channels preset by the encoder.
[0429] When the sound field type is a diffuse sound field, F = 1.
[0430] Next, this application provides a method for selecting the number of channels R of the adaptive residual signal:
[0431] When the sound field type is a diffuse sound source field, R = max(C-1, PR), where C is the preset total number of transmission channels and PR is the preset number of residual signals of the encoder. For example, C is the sum of PF and PR.
[0432] When the sound field type is anisotropic sound source field, R = CF.
[0433] The bit allocation method for the virtual loudspeaker signal and the residual signal is determined based on the sound field classification results:
[0434] When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the residual signal energy is low, so more bits can be allocated to the virtual speaker signal channels.
[0435] In some embodiments, the virtual speaker signal and the residual signal are divided into two groups, namely the virtual speaker signal group and the residual signal group. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the allocation ratio of the virtual speaker signal group is increased according to a preset ratio adjustment value, and the increased allocation ratio of the virtual speaker signal group is used as the allocation ratio of the virtual speaker signal group.
[0436] The allocation ratio of the residual signal group = 1.0 - the allocation ratio of the virtual speaker signal group.
[0437] Virtual speaker signal generation unit: Calculates the virtual speaker signal using the HOA coefficients to be encoded and the HOA coefficients of the matched virtual speaker.
[0438] Signal reconstruction unit: Reconstructs the HOA signal using the virtual speaker signal and the HOA coefficients of the matched virtual speaker.
[0439] Residual signal generation unit: Based on the number of channels of the residual signal determined in step 1, the residual signal is calculated by using the HOA coefficients to be encoded and the reconstructed HOA signal output by the HOA signal reconstruction unit.
[0440] Signal compensation unit: Since the number of channels with fewer than the Nth order Ambisonic coefficients is selected as the residual signal to be transmitted, there will be information loss compared with the residual signal with Nth order Ambisonic coefficients. Therefore, information compensation is required for the residual signal that is not transmitted.
[0441] Selection Unit: The virtual loudspeaker signal has a high amplitude or energy, while the residual signal to be transmitted has a relatively low amplitude or energy. Therefore, the selection unit pre-allocates all available bits to the virtual loudspeaker signal and the residual signal to be transmitted, and the resulting bit pre-allocation information is used to guide the core encoder processing.
[0442] Core encoder processing unit: Performs core encoder processing on the transmission channel and outputs the transmission bitstream. The transmission channel includes a virtual speaker signal channel and a residual signal channel.
[0443] Based on the sound field classification results, the encoding parameters are determined. The encoding parameters may also include at least one of the bit allocations for the virtual speaker signal and the residual signal in the HOA encoding scheme based on the virtual speaker selection. If the bit allocations for the virtual speaker signal and the residual signal are determined using the sound field classification results, then the bit allocations for the virtual speaker signal and the residual signal need to be determined based on the sound field classification results.
[0444] In some embodiments, the bit allocation method for virtual speaker signals and residual signals is determined based on the sound field classification results as follows: Assuming the number of channels for the virtual speaker signal is F, the number of channels for the residual signal is R, and the total number of bits available for encoding the virtual speaker signal and residual signal is numbit.
[0445] One approach is to first determine the total number of bits for encoding the virtual speaker signal and the total number of bits for encoding the residual signal, and then determine the number of bits for encoding each channel. For example:
[0446] The total number of bits for encoding the virtual speaker signal is:
[0447]
[0448] Where fac1 is the weighting factor for the virtual loudspeaker signal encoding bits, and fac2 is the weighting factor for the residual signal encoding bits. round() represents rounding down. For example, fac1 > fac2. For example, fac1 = 2, fac2 = 1.
[0449] The total number of bits for the residual signal encoding is res_numbit = numbit - core_numbit.
[0450] Then, the encoding bits of each channel of the virtual loudspeaker signal are allocated according to the bit allocation criteria of the virtual loudspeaker signal, and the encoding bits of each channel of the residual signal are allocated according to the bit allocation criteria of the residual signal.
[0451] Alternatively, the total number of bits for encoding the residual signal is:
[0452]
[0453] Where fac1 is the weighting factor for the virtual loudspeaker signal encoding bits, and fac2 is the weighting factor for the residual signal encoding bits. round() represents rounding down. For example, fac1 > fac2. For example, fac1 = 2, fac2 = 1.
[0454] The total number of bits encoded in the virtual speaker signal is core_numbit = numbit - res_numbit.
[0455] Then, the encoding bits of each channel of the virtual loudspeaker signal are allocated according to the bit allocation criteria of the virtual loudspeaker signal, and the encoding bits of each channel of the residual signal are allocated according to the bit allocation criteria of the residual signal.
[0456] Alternatively, the number of encoded bits for each channel can be determined directly. For example, the number of bits encoded for each virtual speaker signal is:
[0457]
[0458] The number of bits encoded for each residual signal is:
[0459]
[0460] It should be noted that the final bit allocation result used for encoding the virtual speaker signal and the residual signal can be determined by adjusting the bit allocation result obtained according to the above method. After obtaining the bit allocation result for encoding the virtual speaker signal and the residual signal, the core encoder processing unit will encode the virtual speaker signal and the residual signal according to the bit allocation result.
[0461] The sound field classification results of the HOA signal to be encoded are used to determine the encoding parameters, and then the signal to be encoded is encoded according to these parameters. The encoding parameters include at least one of the following in the HOA encoding scheme based on virtual loudspeaker selection: the number of channels for the virtual loudspeaker signal, the number of channels for the residual signal, the bit allocation of the virtual loudspeaker signal, the bit allocation of the residual signal, and the number of voting rounds for the best-matching loudspeaker search. For a description of the encoding parameters, please refer to the foregoing content; they will not be repeated here.
[0462] As can be seen from the foregoing examples, the embodiments of this application classify the sound field of the HOA signal to be encoded, thereby selecting appropriate encoding modes and / or encoding parameters for different characteristics of the HOA signal to be encoded, encoding the HOA signal, and improving compression efficiency and auditory quality.
[0463] The decoding process executed by the decoding end will not be described in detail in this embodiment.
[0464] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0465] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.
[0466] Please see Figure 12As shown in the embodiment of this application, a three-dimensional audio signal processing device is provided. For example, the three-dimensional audio signal processing device is specifically an audio encoding device 1200, which may include: a linear analysis module 1201, a parameter generation module 1202, and a sound field classification module 1203, wherein...
[0467] The linear analysis module is used to perform linear decomposition on the three-dimensional audio signal to obtain the linear decomposition result;
[0468] The parameter generation module is used to obtain the sound field classification parameters corresponding to the current frame based on the linear decomposition result;
[0469] The sound field classification module is used to determine the sound field classification result of the current frame based on the sound field classification parameters.
[0470] In some embodiments of this application, the three-dimensional audio signal includes: a high-order stereo reverberation (HOA) signal, or a first-order stereo reverberation (FOA) signal.
[0471] In some embodiments of this application, the linear analysis module is used to perform singular value decomposition on the current frame to obtain singular values corresponding to the current frame, wherein the linear decomposition result includes the singular values; or, to perform principal component analysis on the current frame to obtain a first feature value corresponding to the current frame, wherein the linear decomposition result includes the first feature value; or, to perform independent component analysis on the current frame to obtain a second feature value corresponding to the current frame, wherein the linear decomposition result includes the second feature value.
[0472] In some embodiments of this application, the linear decomposition results are multiple, and the sound field classification parameters are multiple;
[0473] The parameter generation module is used to obtain the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer; and to obtain the i-th sound field classification parameter corresponding to the current frame based on the ratio.
[0474] Optionally, the i-th linear analysis result and the (i+1)-th linear analysis result are two consecutive linear analysis results of the current frame.
[0475] In some embodiments of this application, the sound field classification parameters are multiple; the sound field classification result includes: sound field type; the sound field classification module is used to determine the sound field type as a diffuse sound field when the values of the multiple sound field classification parameters all satisfy a preset diffuse sound source decision condition; or, when at least one of the values of the multiple sound field classification parameters satisfies a preset dissimilar sound source decision condition, determine the sound field type as a dissimilar sound field.
[0476] In some embodiments of this application, the diffusion source determination condition includes: the value of the sound field classification parameter is less than a preset dissimilar sound source determination threshold; or, the dissimilar sound source determination condition includes: the value of the sound field classification parameter is greater than or equal to the preset dissimilar sound source determination threshold.
[0477] In some embodiments of this application, the sound field classification parameters are multiple;
[0478] The sound field classification result includes: sound field type; or, the sound field classification result includes: number of dissimilar sound sources and sound field type;
[0479] The sound field classification module is used to obtain the number of dissimilar sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters; and to determine the sound field type based on the number of dissimilar sound sources corresponding to the current frame.
[0480] In some embodiments of this application, the sound field classification parameters are multiple;
[0481] The sound field classification results include: the number of dissimilar sound sources;
[0482] The sound field classification module is used to obtain the number of dissimilar sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters.
[0483] In some embodiments of this application, the plurality of sound field classification parameters are temp[i], where i = 0, 1, ..., min(L, K)-2, where L represents the number of channels in the current frame, K is the number of signal points corresponding to each channel in the current frame, and min represents the minimum value operation;
[0484] The sound field classification module is used to execute the following judgment process sequentially starting from i=0:
[0485] Determine whether temp[i] is greater than a preset dissimilar sound source determination threshold;
[0486] When temp[i] is less than the dissimilar sound source determination threshold in the current judgment process, the value of i is updated to i+1, and the next judgment process is executed; or,
[0487] When temp[i] in this judgment process is greater than or equal to the dissimilar sound source determination threshold, the judgment process is terminated, and i plus 1 in this judgment process is determined to be equal to the number of dissimilar sound sources.
[0488] In some embodiments of this application, determining the sound field type based on the number of dissimilar sound sources corresponding to the current frame includes:
[0489] When the number of dissimilar sound sources meets the first preset condition, the sound field type is determined to be the first sound field type;
[0490] When the number of dissimilar sound sources does not meet the first preset condition, the sound field type is determined to be the second sound field type;
[0491] The number of dissimilar sound sources corresponding to the first sound field type is different from the number of dissimilar sound sources corresponding to the second sound field type.
[0492] In some embodiments of this application, the first preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold;
[0493] or,
[0494] The first preset condition includes that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
[0495] In some embodiments of this application, the audio encoding apparatus further includes: an encoding mode determination module ( Figure 12 (Not illustrated in the image), the encoding mode determination module is used to determine the encoding mode corresponding to the current frame based on the sound field classification result.
[0496] In some embodiments of this application, the encoding mode determination module is used to determine the encoding mode corresponding to the current frame based on the number of dissimilar sound sources when the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; or, when the sound field classification result includes the sound field type, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the encoding mode corresponding to the current frame based on the sound field type; or, when the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the encoding mode corresponding to the current frame based on the number of dissimilar sound sources and the sound field type.
[0497] In some embodiments of this application, the encoding mode determination module is used to determine the encoding mode as a first encoding mode when the number of dissimilar sound sources meets a second preset condition; and to determine the encoding mode as a second encoding mode when the number of dissimilar sound sources does not meet the second preset condition.
[0498] Wherein, the first encoding mode is either a HOA encoding mode selected based on a virtual speaker or a HOA encoding mode based on directional audio encoding, the second encoding mode is either a HOA encoding mode selected based on a virtual speaker or a HOA encoding mode based on directional audio encoding, and the first encoding mode and the second encoding mode are different encoding modes.
[0499] In some embodiments of this application, the second preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or,
[0500] The second preset condition includes that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
[0501] In some embodiments of this application, the encoding mode determination module is used to determine the encoding mode as a HOA encoding mode based on virtual loudspeaker selection when the sound field type is an anisotropic sound field; and to determine the encoding mode as a HOA encoding mode based on directional audio coding when the sound field type is a diffuse sound field.
[0502] In some embodiments of this application, the encoding mode determination module is used to determine the initial encoding mode corresponding to the current frame based on the sound field classification result of the current frame; obtain the sliding window in which the current frame is located, the sliding window including: the initial encoding mode of the current frame and the encoding modes of N-1 frames before the current frame, where N is the length of the sliding window; and determine the encoding mode of the current frame based on the initial encoding mode of the current frame and the encoding modes of the N-1 frames.
[0503] In some embodiments of this application, the audio encoding apparatus further includes: an encoding parameter determination module ( Figure 12 (Not illustrated in the image), the encoding parameter determination module is used to determine the encoding parameters corresponding to the current frame based on the sound field classification results.
[0504] In some embodiments of this application, the encoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of encoded bits of the virtual speaker signal, the number of encoded bits of the residual signal, or the number of voting rounds for the best matching speaker search;
[0505] The virtual speaker signal and the residual signal are signals generated based on the three-dimensional audio signal.
[0506] In some embodiments of this application, the number of voting rounds satisfies the following relationship:
[0507] 1≤I≤d,
[0508] Wherein, I represents the number of voting rounds, and d represents the number of dissimilar sound sources included in the sound field classification result.
[0509] In some embodiments of this application, the sound field classification results include the number of dissimilar sound sources and the sound field type;
[0510] When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0511] F = min(S, PF),
[0512] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the preset number of virtual speaker signal channels of the encoder; or,
[0513] When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0514] F = 1,
[0515] Wherein, F is the number of channels of the virtual speaker signal.
[0516] In some embodiments of this application, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship:
[0517] R = max(C-1, PR),
[0518] Wherein, PR is the preset number of residual signal channels of the encoder, and C is the sum of the preset number of residual signal channels and the preset number of virtual speaker signal channels of the encoder; or,
[0519] When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship:
[0520] R = C – F,
[0521] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the encoder and the number of virtual speaker signal channels preset by the encoder, and F is the number of channels of the virtual speaker signal.
[0522] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources;
[0523] The number of channels of the virtual speaker signal satisfies the following relationship:
[0524] F = min(S, PF),
[0525] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the encoder.
[0526] In some embodiments of this application, the number of channels of the residual signal satisfies the following relationship:
[0527] R = C – F,
[0528] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal.
[0529] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type;
[0530] The number of encoded bits of the virtual speaker signal is obtained by the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel;
[0531] The number of encoded bits of the residual signal is obtained by the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel;
[0532] The number of encoded bits in the transmission channel includes the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel is obtained by increasing the initial ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel.
[0533] In some embodiments of this application, the audio encoding device further includes: an encoding module ( Figure 12 (Not illustrated in the image), the encoding module is used to encode the current frame and the sound field classification result, and write them into the bitstream.
[0534] As illustrated by the foregoing embodiments, the current frame of the 3D audio signal is first linearly decomposed to obtain the linear decomposition result; then, the sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition result; finally, the sound field classification result of the current frame is determined based on the sound field classification parameters. Since this embodiment obtains the linear decomposition result of the current frame by performing linear decomposition on the current frame of the 3D audio signal, and then obtains the corresponding sound field classification parameters based on these results, the sound field classification result of the current frame is determined using these parameters. This sound field classification result allows for sound field classification of the current frame. This embodiment of the application performs sound field classification on the 3D audio signal, thereby accurately identifying the 3D audio signal.
[0535] Please see Figure 13As shown in the embodiment of this application, a three-dimensional audio signal processing device is provided. For example, the three-dimensional audio signal processing device is specifically an audio decoding device 1300, which may include: a receiving module 1301, a decoding module 1302, and a signal generation module 1303, wherein...
[0536] The receiving module is used to receive the bitstream;
[0537] A decoding module is used to decode the bitstream to obtain the sound field classification result of the current frame;
[0538] The signal generation module is used to obtain the three-dimensional audio signal decoded from the current frame based on the sound field classification result.
[0539] In some embodiments of this application, the signal generation module is used to determine the decoding mode of the current frame based on the sound field classification result; and to obtain the decoded three-dimensional audio signal of the current frame based on the decoding mode.
[0540] In some embodiments of this application, the signal generation module is configured to determine the decoding mode of the current frame based on the number of dissimilar sound sources when the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; or, when the sound field classification result includes the sound field type, or the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the decoding mode of the current frame based on the sound field type; or, when the sound field classification result includes the number of dissimilar sound sources and the sound field type, determine the decoding mode of the current frame based on the number of dissimilar sound sources and the sound field type.
[0541] In some embodiments of this application, the signal generation module is used to determine the decoding mode as a first decoding mode when the number of dissimilar sound sources meets a preset condition; and to determine the decoding mode as a second decoding mode when the number of dissimilar sound sources does not meet the preset condition.
[0542] The first decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio encoding, and the second decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio encoding, and the first decoding mode and the second decoding mode are different decoding modes.
[0543] In some embodiments of this application, the preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold;
[0544] or
[0545] The preset conditions include that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
[0546] In some embodiments of this application, the signal generation module is used to determine the decoding parameters of the current frame based on the sound field classification result; and to obtain the decoded three-dimensional audio signal of the current frame based on the decoding parameters.
[0547] In some embodiments of this application, the decoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signal, or the number of decoded bits of the residual signal;
[0548] The virtual speaker signal and the residual signal are obtained by decoding the bitstream.
[0549] In some embodiments of this application, the sound field classification results include the number of dissimilar sound sources and the sound field type;
[0550] When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0551] F = min(S, PF),
[0552] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder; or,
[0553] When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship:
[0554] F = 1,
[0555] Wherein, F is the number of channels of the virtual speaker signal.
[0556] In some embodiments of this application, when the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship:
[0557] R = max(C-1, PR),
[0558] Wherein, PR is the number of residual signal channels preset by the decoder, and C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder; or,
[0559] When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship:
[0560] R = C – F,
[0561] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0562] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources;
[0563] The number of channels of the virtual speaker signal satisfies the following relationship:
[0564] F = min(S, PF),
[0565] Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder.
[0566] In some embodiments of this application, the number of channels of the residual signal satisfies the following relationship:
[0567] R = C – F,
[0568] Wherein, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
[0569] In some embodiments of this application, the sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type;
[0570] The number of decoded bits of the virtual speaker signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel;
[0571] The number of decoded bits of the residual signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel;
[0572] The number of decoded bits in the transmission channel includes the number of decoded bits of the virtual speaker signal and the number of decoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel is obtained by increasing the initial ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel.
[0573] As can be seen from the examples in the foregoing embodiments, the sound field classification result can be used to decode the current frame in the bitstream. Therefore, the decoding end uses a decoding method that matches the sound field of the current frame to decode, thereby obtaining the three-dimensional audio signal sent by the encoding end, and realizing the transmission of the audio signal from the encoding end to the decoding end.
[0574] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0575] This application also provides a computer storage medium storing a program that performs some or all of the steps described in the above method embodiments.
[0576] The following describes another audio encoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 14 As shown, the audio encoding device 1400 includes:
[0577] Receiver 1401, transmitter 1402, processor 1403, and memory 1404 (wherein the audio encoding device 1400 may contain one or more processors 1403). Figure 14 (Taking a processor as an example). In some embodiments of this application, the receiver 1401, transmitter 1402, processor 1403, and memory 1404 can be connected via a bus or other means, wherein... Figure 14 Taking the example of a connection between China and Israel via a bus.
[0578] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0579] Processor 1403 controls the operation of the audio encoding device; processor 1403 can also be called a central processing unit (CPU). In specific applications, the various components of the audio encoding device are coupled together through a bus system, which includes not only a data bus but also a power bus, control bus, and status signal bus, etc. However, for clarity, all buses are referred to as the bus system in the diagram.
[0580] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1403 or by instructions in the form of software. The processor 1403 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1404. Processor 1403 reads the information in memory 1404 and completes the steps of the above method in conjunction with its hardware.
[0581] The receiver 1401 can be used to receive input digital or character information and generate signal inputs related to the settings and function control of the audio encoding device. The transmitter 1402 may include a display device such as a display screen and can be used to output digital or character information through an external interface.
[0582] In this embodiment, processor 1403 is used to execute the aforementioned embodiments. Figures 4 to 6 The method shown is performed by the audio encoding device.
[0583] The following describes another audio decoding device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 15As shown, the audio decoding device 1500 includes:
[0584] Receiver 1501, transmitter 1502, processor 1503, and memory 1504 (wherein the audio decoding device 1500 may contain one or more processors 1503). Figure 15 (Taking a processor as an example). In some embodiments of this application, the receiver 1501, transmitter 1502, processor 1503, and memory 1504 can be connected via a bus or other means, wherein... Figure 15 Taking the example of a connection between China and Israel via a bus.
[0585] Memory 1504 may include read-only memory and random access memory, and provides instructions and data to processor 1503. A portion of memory 1504 may also include NVRAM. Memory 1504 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0586] Processor 1503 controls the operation of the audio decoding device; processor 1503 can also be referred to as a CPU. In specific applications, the various components of the audio decoding device are coupled together through a bus system. This bus system includes not only a data bus but also a power bus, control bus, and status signal bus, etc. However, for clarity, all buses in the diagram are referred to as the bus system.
[0587] The methods disclosed in the embodiments of this application can be applied to processor 1503, or implemented by processor 1503. Processor 1503 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1503 or by instructions in the form of software. The processor 1503 can be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1504, and processor 1503 reads the information in memory 1504 and completes the steps of the above method in combination with its hardware.
[0588] In this embodiment, processor 1503 is used to execute the aforementioned embodiments. Figure 7 The method shown is performed by the audio decoding device.
[0589] In another possible design, when the audio encoding or decoding device is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer-executable instructions stored in the storage unit to cause the chip within the terminal to execute the audio encoding method of any of the first aspects or the audio decoding method of any of the second aspects described above. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the terminal, such as read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0590] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of programs in the first or second aspect of the above methods.
[0591] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0592] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0593] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0594] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. A method for processing three-dimensional audio signals, characterized in that, include: Perform linear decomposition on the current frame of the 3D audio signal to obtain the linear decomposition result; The sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition results. The sound field classification result of the current frame is determined based on the sound field classification parameters.
2. The method according to claim 1, characterized in that, The three-dimensional audio signal includes: a high-order stereo reverb (HOA) signal or a first-order stereo reverb (FOA) signal.
3. The method according to claim 1 or 2, characterized in that, The linear decomposition of the current frame of the three-dimensional audio signal to obtain the linear decomposition result includes: Singular value decomposition is performed on the current frame to obtain the singular values corresponding to the current frame, wherein the linear decomposition result includes the singular values; or, Principal component analysis is performed on the current frame to obtain the first feature value corresponding to the current frame, wherein the linear decomposition result includes: the first feature value; or, Independent component analysis is performed on the current frame to obtain the second feature value corresponding to the current frame, wherein the linear decomposition result includes the second feature value.
4. The method according to any one of claims 1 to 3, characterized in that, The linear decomposition results are multiple, and the sound field classification parameters are multiple; The step of obtaining the sound field classification parameters corresponding to the current frame based on the linear decomposition result includes: Obtain the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer; The i-th sound field classification parameter corresponding to the current frame is obtained based on the ratio.
5. The method according to any one of claims 1 to 4, characterized in that, The sound field classification parameters are multiple; The sound field classification results include: sound field type; Determining the sound field classification result of the current frame based on the sound field classification parameters includes: When the values of the plurality of sound field classification parameters all satisfy the preset criteria for determining a diffuse sound source, the sound field type is determined to be a diffuse sound field. or, When at least one of the values of the plurality of sound field classification parameters satisfies the preset dissimilar sound source determination condition, the sound field type is determined to be a dissimilar sound field.
6. The method according to claim 5, characterized in that, The criteria for determining diffuse sound sources include: the value of the sound field classification parameter is less than the preset threshold for determining dissimilar sound sources; or, The dissimilar sound source determination conditions include: the value of the sound field classification parameter is greater than or equal to the preset dissimilar sound source determination threshold.
7. The method according to any one of claims 1 to 4, characterized in that, The sound field classification parameters are multiple; The sound field classification result includes: sound field type; or, the sound field classification result includes: number of dissimilar sound sources and sound field type; Determining the sound field classification result of the current frame based on the sound field classification parameters includes: The number of dissimilar sound sources corresponding to the current frame is obtained based on the values of the multiple sound field classification parameters; The sound field type is determined based on the number of dissimilar sound sources corresponding to the current frame.
8. The method according to any one of claims 1 to 4, characterized in that, The sound field classification parameters are multiple; The sound field classification results include: the number of dissimilar sound sources; Determining the sound field classification result of the current frame based on the sound field classification parameters includes: The number of dissimilar sound sources corresponding to the current frame is obtained based on the values of the multiple sound field classification parameters.
9. The method according to claim 7 or 8, characterized in that, The multiple sound field classification parameters are temp[i], where i = 0, 1, ..., min(L, K)-2, where L represents the number of channels in the current frame, K is the number of signal points corresponding to each channel in the current frame, and min represents the minimum value operation. The step of obtaining the number of dissimilar sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters includes: Starting from i=0, the following judgment process is executed sequentially: Determine whether temp[i] is greater than a preset dissimilar sound source determination threshold; When temp[i] is less than the dissimilar sound source determination threshold in the current judgment process, the value of i is updated to i+1, and the next judgment process is executed; or, When temp[i] in this judgment process is greater than or equal to the dissimilar sound source determination threshold, the judgment process is terminated, and i plus 1 in this judgment process is determined to be equal to the number of dissimilar sound sources.
10. The method according to claim 7, characterized in that, Determining the sound field type based on the number of dissimilar sound sources corresponding to the current frame includes: When the number of dissimilar sound sources meets the first preset condition, the sound field type is determined to be the first sound field type; When the number of dissimilar sound sources does not meet the first preset condition, the sound field type is determined to be the second sound field type; The number of dissimilar sound sources corresponding to the first sound field type is different from the number of dissimilar sound sources corresponding to the second sound field type.
11. The method according to claim 10, characterized in that, The first preset condition includes the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or, The first preset condition includes that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: The encoding mode corresponding to the current frame is determined based on the sound field classification results.
13. The method according to claim 12, characterized in that, Determining the encoding mode corresponding to the current frame based on the sound field classification result includes: When the sound field classification result includes the number of dissimilar sound sources, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the coding mode corresponding to the current frame is determined according to the number of dissimilar sound sources. or, When the sound field classification result includes a sound field type, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the coding mode corresponding to the current frame is determined according to the sound field type; or, When the sound field classification result includes the number of dissimilar sound sources and the sound field type, the encoding mode corresponding to the current frame is determined based on the number of dissimilar sound sources and the sound field type.
14. The method according to claim 13, characterized in that, Determining the encoding mode corresponding to the current frame based on the number of dissimilar sound sources includes: When the number of dissimilar sound sources meets the second preset condition, the encoding mode is determined to be the first encoding mode; When the number of dissimilar sound sources does not meet the second preset condition, the encoding mode is determined to be the second encoding mode; Wherein, the first encoding mode is either a HOA encoding mode selected based on a virtual speaker or a HOA encoding mode based on directional audio encoding, the second encoding mode is either a HOA encoding mode selected based on a virtual speaker or a HOA encoding mode based on directional audio encoding, and the first encoding mode and the second encoding mode are different encoding modes.
15. The method according to claim 14, characterized in that, The second preset condition includes that the number of dissimilar sound sources is greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or, The second preset condition includes that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
16. The method according to claim 13, characterized in that, Determining the encoding mode corresponding to the current frame based on the sound field type includes: When the sound field type is anisotropic sound field, the encoding mode is determined to be the HOA encoding mode based on virtual loudspeaker selection; When the sound field type is a diffuse sound field, the encoding mode is determined to be the HOA encoding mode based on directional audio coding.
17. The method according to claim 12, characterized in that, Determining the encoding mode corresponding to the current frame based on the sound field classification result includes: The initial encoding mode corresponding to the current frame is determined based on the sound field classification result of the current frame; Obtain the sliding window in which the current frame is located. The sliding window includes: the initial encoding mode of the current frame and the encoding modes of the N-1 frames preceding the current frame, where N is the length of the sliding window. The encoding mode of the current frame is determined based on the initial encoding mode of the current frame within the sliding window and the encoding mode of the N-1 frames.
18. The method according to any one of claims 1 to 17, characterized in that, The method further includes: The encoding parameters corresponding to the current frame are determined based on the sound field classification results.
19. The method according to claim 18, characterized in that, The encoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of encoded bits of the virtual speaker signal, the number of encoded bits of the residual signal, or the number of voting rounds for the best matching speaker search; The virtual speaker signal and the residual signal are generated based on the three-dimensional audio signal.
20. The method according to claim 19, characterized in that, The number of voting rounds satisfies the following relationship: 1≤I≤d, Wherein, I represents the number of voting rounds, and d represents the number of dissimilar sound sources included in the sound field classification result.
21. The method according to claim 19 or 20, characterized in that, The sound field classification results include the number of dissimilar sound sources and the sound field type; When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship: F = min(S, PF), Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the preset number of virtual speaker signal channels of the encoder; or, When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship: F=1, Wherein, F is the number of channels of the virtual speaker signal.
22. The method according to any one of claims 19 to 21, characterized in that, When the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1, PR), Wherein, PR is the preset number of residual signal channels of the encoder, and C is the sum of the preset number of residual signal channels and the preset number of virtual speaker signal channels of the encoder; or, When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship: R = C – F, Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the encoder and the number of virtual speaker signal channels preset by the encoder, and F is the number of channels of the virtual speaker signal.
23. The method according to claim 19 or 20, characterized in that, The sound field classification results include the number of dissimilar sound sources; The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S, PF), Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the encoder.
24. The method according to claim 19, 20, 21 or 23, characterized in that, The number of channels in the residual signal satisfies the following relationship: R = C – F, Wherein, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal.
25. The method according to any one of claims 19 to 24, characterized in that, The sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; The number of encoded bits of the virtual speaker signal is obtained by the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel; The number of encoded bits of the residual signal is obtained by the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel; The number of encoded bits in the transmission channel includes the number of encoded bits of the virtual speaker signal and the number of encoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel is obtained by increasing the initial ratio of the number of encoded bits of the virtual speaker signal to the number of encoded bits of the transmission channel.
26. The method according to any one of claims 1 to 25, characterized in that, The method further includes: The current frame and the sound field classification result are encoded and written into the bitstream.
27. A method for processing three-dimensional audio signals, characterized in that, include: Receive bitstream; Decode the bitstream to obtain the sound field classification result of the current frame; The three-dimensional audio signal after decoding the current frame is obtained based on the sound field classification results.
28. The method according to claim 27, characterized in that, The step of obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result includes: The decoding mode of the current frame is determined based on the sound field classification results; The three-dimensional audio signal after decoding the current frame is obtained according to the decoding mode.
29. The method according to claim 28, characterized in that, Determining the decoding mode of the current frame based on the sound field classification result includes: When the sound field classification result includes the number of dissimilar sound sources, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined according to the number of dissimilar sound sources. or, When the sound field classification result includes a sound field type, or when the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined according to the sound field type; or, When the sound field classification result includes the number of dissimilar sound sources and the sound field type, the decoding mode of the current frame is determined based on the number of dissimilar sound sources and the sound field type.
30. The method according to claim 29, characterized in that, Determining the decoding mode corresponding to the current frame based on the number of dissimilar sound sources includes: When the number of dissimilar sound sources meets a preset condition, the decoding mode is determined to be the first decoding mode; When the number of dissimilar sound sources does not meet the preset condition, the decoding mode is determined to be the second decoding mode; The first decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio coding, and the second decoding mode is either a HOA decoding mode selected based on a virtual speaker or a HOA decoding mode based on directional audio coding, and the first decoding mode and the second decoding mode are different decoding modes.
31. The method according to claim 30, characterized in that, The preset conditions include the number of dissimilar sound sources being greater than a first threshold and less than a second threshold, wherein the second threshold is greater than the first threshold; or The preset conditions include that the number of dissimilar sound sources is not greater than a first threshold or not less than a second threshold, wherein the second threshold is greater than the first threshold.
32. The method according to claim 27, characterized in that, The step of obtaining the decoded three-dimensional audio signal of the current frame based on the sound field classification result includes: The decoding parameters of the current frame are determined based on the sound field classification results; The three-dimensional audio signal after decoding the current frame is obtained based on the decoding parameters.
33. The method according to claim 32, characterized in that, The decoding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signal, or the number of decoded bits of the residual signal; The virtual speaker signal and the residual signal are obtained by decoding the bitstream.
34. The method according to claim 33, characterized in that, The sound field classification results include the number of dissimilar sound sources and the sound field type; When the sound field type is an anisotropic sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship: F = min(S, PF), Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder; or, When the sound field type is a diffuse sound field, the number of channels of the virtual loudspeaker signal satisfies the following relationship: F=1, Wherein, F is the number of channels of the virtual speaker signal.
35. The method according to claim 33 or 34, characterized in that, When the sound field type is a diffuse sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1, PR), Wherein, PR is the number of residual signal channels preset by the decoder, and C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder; or, When the sound field type is an anisotropic sound field, the number of channels of the residual signal satisfies the following relationship: R = C – F, Wherein, R represents the number of channels of the residual signal, C is the sum of the number of residual signal channels preset by the decoder and the number of virtual speaker signal channels preset by the decoder, and F is the number of channels of the virtual speaker signal.
36. The method according to claim 33 or 35, characterized in that, The sound field classification results include the number of dissimilar sound sources; The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S, PF), Wherein, F is the number of channels of the virtual speaker signal, S is the number of dissimilar sound sources, and PF is the number of virtual speaker signal channels preset by the decoder.
37. The method according to any one of claims 33 to 36, characterized in that, The number of channels in the residual signal satisfies the following relationship: R = C – F, Wherein, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
38. The method according to any one of claims 33 to 37, characterized in that, The sound field classification result includes the number of dissimilar sound sources, or the sound field classification result includes the number of dissimilar sound sources and the sound field type; The number of decoded bits of the virtual speaker signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel; The number of decoded bits of the residual signal is obtained by the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel; The number of decoded bits in the transmission channel includes the number of decoded bits of the virtual speaker signal and the number of decoded bits of the residual signal. When the number of dissimilar sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel is obtained by increasing the initial ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel.
39. A three-dimensional audio signal processing device, characterized in that, include: The linear analysis module is used to perform linear decomposition on the current frame of the 3D audio signal to obtain the linear decomposition result; The parameter generation module is used to obtain the sound field classification parameters corresponding to the current frame based on the linear decomposition result; The sound field classification module is used to determine the sound field classification result of the current frame based on the sound field classification parameters.
40. A three-dimensional audio signal processing device, characterized in that, include: The receiving module is used to receive the bitstream; A decoding module is used to decode the bitstream to obtain the sound field classification result of the current frame; The signal generation module is used to obtain the three-dimensional audio signal decoded from the current frame based on the sound field classification result.
41. A three-dimensional audio signal processing device, characterized in that, The three-dimensional audio signal processing apparatus includes at least one processor, the at least one processor being coupled to a memory to read and execute instructions in the memory to implement the method as described in any one of claims 1 to 26.
42. The three-dimensional audio signal processing apparatus according to claim 41, characterized in that, The three-dimensional audio signal processing device further includes the memory.
43. A three-dimensional audio signal processing device, characterized in that, The three-dimensional audio signal processing device includes at least one processor, the at least one processor being coupled to a memory to read and execute instructions in the memory to implement the method as described in any one of claims 27 to 38.
44. The three-dimensional audio signal processing apparatus according to claim 43, characterized in that, The three-dimensional audio signal processing device further includes the memory.
45. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 26 or 27 to 38.
46. A computer-readable storage medium comprising a bitstream generated by the method as described in any one of claims 1 to 26.