Three-dimensional audio signal processing method and device
By performing linear decomposition and sound field classification on three-dimensional audio signals, the method addresses the inefficiencies in current systems, enabling accurate classification and optimized encoding for efficient compression and enhanced audio quality.
Patent Information
- Application Number
- JP2023573612
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-31
- Filing Date
- 2022-05-30
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2042-05-30
AI Technical Summary
Current three-dimensional audio signal processing systems struggle with efficient classification and compression of three-dimensional audio signals due to the large data amount and high bandwidth requirements, as existing encoders cannot effectively identify and classify these signals before encoding.
Perform linear decomposition on a three-dimensional audio signal to obtain a linear decomposition result, then determine sound field classification parameters based on these results, allowing for accurate classification of the signal into non-uniform and distributed sound fields, and adapt encoding modes and parameters accordingly.
This method enables efficient compression and improved hearing quality by accurately identifying sound field types, optimizing encoding modes and parameters, thereby reducing data requirements and enhancing audio reproduction.
Smart Images

Figure 0007680571000022 
Figure 0007680571000023 
Figure 0007680571000024
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of audio processing technology, and in particular to a three-dimensional audio signal processing method and device. [Background technology]
[0002] This application claims priority to Chinese Patent Application No. 202110602507.4, entitled "Three-dimensional audio signal processing method and device," filed with the China National Intellectual Property Office on May 31, 2021, the contents of which are incorporated herein by reference in their entirety.
[0003] Three-dimensional audio technology is widely used in wireless communication conversation, virtual reality / augmented reality, and media audio. Three-dimensional audio technology is an audio technology that acquires, processes, transmits, renders, and reproduces acoustic events and three-dimensional sound field information in the real world. Three-dimensional audio technology gives sound a strong sense of space, envelopment, and immersion, providing an extraordinary "immersive" auditory experience. Higher-order Ambisonics (HOA) technology has the ability to rotate and reproduce data in HOA format, independent of the placement of speakers during recording, encoding, and playback. Higher-order Ambisonics technology has higher flexibility for three-dimensional audio reproduction, and therefore has attracted more attention and research.
[0004] An imaging device (e.g., a microphone) captures a large amount of data to record the three-dimensional sound field information, and transmits a three-dimensional audio signal to a reproduction device (e.g., a speaker or a microphone), which reproduces the three-dimensional audio signal. Because the data amount of the three-dimensional sound field information is large, a large memory capacity is required to store the data, and a high bandwidth is required to carry the three-dimensional audio signal. To solve the above problem, the three-dimensional audio signal may be compressed, and the compressed data may be stored or transmitted.
[0005] Currently, an encoder can encode a three-dimensional audio signal by using multiple pre-configured virtual speakers. However, before encoding the three-dimensional audio signal, the encoder cannot classify the three-dimensional audio signal, and as a result, cannot effectively identify the three-dimensional audio signal. Summary of the Invention
[0006] The embodiments of the present invention provide a three-dimensional audio signal processing method and apparatus for implementing sound field classification of a three-dimensional audio signal and accurately identifying the three-dimensional audio signal.
[0007] In order to solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0008] According to a first aspect, an embodiment of the present application provides a three-dimensional audio signal processing method, including: performing linear decomposition on a current frame of a three-dimensional audio signal to obtain a linear decomposition result; obtaining sound field classification parameters corresponding to the current frame based on the linear decomposition result; and determining a sound field classification result of the current frame based on the sound field classification parameter. In the above solution, first, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result; then, obtaining sound field classification parameters corresponding to the current frame based on the linear decomposition result; and finally, determining a sound field classification result of the current frame based on the sound field classification parameter. In an embodiment of the present application, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result of the current frame; then, obtaining sound field classification parameters corresponding to the current frame based on the linear decomposition result. Therefore, a sound field classification result of the current frame can be determined based on the sound field classification parameter, and a sound field classification of the current frame can be implemented based on the sound field classification result. In an embodiment of the present application, a sound field classification is performed on a three-dimensional audio signal to accurately identify the three-dimensional audio signal.
[0009] In a possible implementation, the three-dimensional audio signal comprises a higher order Ambisonics HOA signal or a first order Ambisonics FOA signal.
[0010] In a possible implementation, performing a linear decomposition on a current frame of the three-dimensional audio signal to obtain a linear decomposition result includes: performing a singular value decomposition on the current frame to obtain singular values corresponding to the current frame, the linear decomposition result including the singular values; performing a principal component analysis on the current frame to obtain first feature values corresponding to the current frame, the linear decomposition result including the first feature values; or performing an independent component analysis on the current frame to obtain second feature values corresponding to the current frame, the linear decomposition result including the second feature values. In the above solutions, the linear decomposition may be a singular value decomposition; the linear decomposition may alternatively be a principal component analysis to obtain feature values, or the linear decomposition may alternatively be an independent component analysis to obtain second feature values. In any one of these three schemes, a linear decomposition of the current frame may be performed to provide a linear analysis result for subsequent audio channel determination.
[0011] In a possible implementation, there are multiple linear decomposition results and there are multiple sound field classification parameters. The step of obtaining a sound field classification parameter corresponding to a current frame based on the linear decomposition results includes: obtaining a ratio of an i-th linear analysis result of the current frame to an (i+1)-th linear analysis result of the current frame, where i is a positive integer; and obtaining an i-th sound field classification parameter corresponding to the current frame based on the ratio.
[0012] Furthermore, the i-th linear analysis result and the (i+1)-th linear analysis result are two consecutive linear analysis results in the current frame.
[0013] In the above solution, the encoder side may obtain a sound field classification parameter corresponding to the current frame based on the linear decomposition result. For example, there are multiple linear decomposition results in the current frame, and two consecutive linear analysis results in the multiple linear analysis results are expressed as the i-th linear analysis result and the (i+1)-th linear analysis result of the current frame. In this case, the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame may be calculated, and the specific value of i is not particularly limited. After this ratio is obtained, the i-th sound field classification parameter corresponding to the current frame may be obtained based on the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result of the current frame.
[0014] In a possible implementation, there are multiple sound field classification parameters, and the sound field classification result includes a sound field type. The step of determining the sound field classification result of the current frame based on the sound field classification parameters includes: determining that the sound field type is a distributed sound if the values of the multiple sound field classification parameters all satisfy a preset distributed sound source judgment condition; or determining that the sound field type is a non-uniform sound field if at least one of the values of the multiple sound field classification parameters satisfies a preset non-uniform sound source judgment condition. In the above solution, the sound field type may include a non-uniform sound field and a distributed sound field. In this embodiment of the present invention, a distributed sound source judgment condition and a non-uniform sound source judgment condition are preset. The distributed sound source judgment condition is used to judge whether the sound field type is a distributed sound field, and the non-uniform sound source judgment condition is used to judge whether the sound field type is a non-uniform sound field. After the multiple sound field classification parameters in the current frame are obtained, a judgment is performed based on the values of the multiple sound field classification parameters and the preset conditions.
[0015] In a possible implementation, the distributed sound source determination condition includes that the value of the sound field classification parameter is less than a predetermined non-uniform sound source determination threshold, or the non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a predetermined non-uniform sound source determination threshold. In the above-mentioned solution, the non-uniform sound source determination threshold may be a threshold value that is set in advance, and a specific value is not limited. The distributed sound source determination condition includes that the value of the sound field classification parameter is less than a predetermined non-uniform sound source determination threshold. Therefore, when the values of the multiple sound field classification parameters are all less than the predetermined non-uniform sound source determination threshold, the sound field type is determined to be a distributed sound field. The non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a predetermined non-uniform sound source determination threshold. Therefore, when at least one of the values of the multiple sound field classification parameters is equal to or greater than a predetermined non-uniform sound source determination threshold, the sound field type is determined to be a non-uniform sound field.
[0016] In a possible implementation, there are multiple sound field classification parameters, and the sound field classification result includes a sound field type, or the sound field classification result includes a non-uniform sound source number and a sound field type. The step of determining the sound field classification result of the current frame based on the sound field classification parameters includes: obtaining a non-uniform sound source number corresponding to the current frame based on the values of the multiple sound field classification parameters; and determining a sound field type based on the non-uniform sound source number corresponding to the current frame. In the above solution, after obtaining the multiple sound field classification parameters corresponding to the current frame, the encoder side may obtain a non-uniform sound source number corresponding to the current frame based on the values of the multiple sound field classification parameters. A non-uniform sound source is a point sound source having different positions and / or directions, and the number of non-uniform sound sources included in the current frame is called the non-uniform sound source number. The sound field of the current frame can be classified based on the non-uniform sound source number. After the non-uniform sound source number corresponding to the current frame is obtained to determine the sound field type, the sound field type corresponding to the current frame can be determined by analyzing the non-uniform sound source number corresponding to the current frame.
[0017] In a possible implementation, there are multiple sound field classification parameters, and the sound field classification result includes a non-uniform sound source number. The step of determining the sound field classification result of the current frame based on the sound field classification parameters includes: obtaining a non-uniform sound source number corresponding to the current frame based on the values of the multiple sound field classification parameters. In the above solution, after obtaining the multiple sound field classification parameters corresponding to the current frame, the encoder side may obtain a non-uniform sound source number corresponding to the current frame based on the values of the multiple sound field classification parameters. A non-uniform sound source is a point sound source having different positions and / or directions, and the number of non-uniform sound sources included in the current frame is called the non-uniform sound source number.
[0018] In a possible implementation, the multiple sound field classification parameters are temp[i], i=0,1,...,min(L,K)-2, where L represents the number of channels in the current frame, K represents the number of signal points corresponding to each channel in the current frame, and min represents the operation of selecting the minimum value. The step of acquiring the number of non-uniform sound sources corresponding to the current frame based on the values of the multiple sound field classification parameters includes the following: That is, the step of sequentially executing the following judgment procedures from i=0. The step of determining whether temp[i] exceeds a preset non-uniform sound source judgment threshold. If temp[i] is less than the non-uniform sound source judgment threshold in this judgment procedure, the step of updating the value of i to i+1 and continuing to execute the next judgment procedure. Or, if temp[i] is equal to or greater than the non-uniform sound source judgment threshold in this judgment procedure, the step of terminating the execution of the judgment procedure and determining that i incremented by 1 in this judgment procedure is equal to the number of non-uniform sound sources. In the above solution, the determination procedure is executed multiple times, and each time it is determined whether to terminate the execution of the determination procedure so as to obtain the number of non-uniform sound sources.
[0019] In a possible implementation, the step of determining the sound field type based on the number of non-uniform sound sources corresponding to the current frame includes the following: if the number of non-uniform sound sources satisfies a first preset condition, the sound field type is determined to be a first sound field type; or if the number of non-uniform sound sources does not satisfy the first preset condition, the sound field type is determined to be a second sound field type. The number of non-uniform sound sources corresponding to the first sound field type is different from the number of non-uniform sound sources corresponding to the second sound field type. In the above-mentioned solution, the sound field type can be classified into two types, the first sound field type and the second sound field type, based on the difference in the number of non-uniform sound sources. The encoder side acquires the preset condition. That is, it is determined whether the number of non-uniform sound sources satisfies the preset condition, and if the number of non-uniform sound sources satisfies the first preset condition, the sound field type is determined to be a first sound field type, or if the number of non-uniform sound sources does not satisfy the first preset condition, the sound field type is determined to be a second sound field type. In this embodiment of the present application, in order to implement the division of the sound field type of the current frame, and to accurately identify whether the sound field type of the current frame belongs to a first sound field type or a second sound field type, it can be determined whether the number of non-uniform sound sources meets a first preset condition.
[0020] In a possible implementation, the first preset condition includes that the number of non-uniform sound sources exceeds a first threshold and is less than a second threshold, and the second threshold exceeds the first threshold. Or, the first preset condition includes that the number of non-uniform sound sources is equal to or less than the first threshold or equal to or greater than the second threshold, and the second threshold exceeds the first threshold. In the above solution, the specific values of the first threshold and the second threshold are not limited and can be specifically determined based on the application scenario. The second threshold exceeds the first threshold. Therefore, the first threshold and the second threshold may constitute a preset range, and the first preset condition may be that the number of non-uniform sound sources falls within the preset range, or the first preset condition may be that the number of non-uniform sound sources exceeds the preset range. The number of non-uniform sound sources may be determined based on the first threshold and the second threshold in the first preset condition to determine whether the number of non-uniform sound sources satisfies the first preset condition and accurately identify that the sound field type of the current frame belongs to the first sound field type or the second sound field type.
[0021] In a possible implementation, the method further includes: determining an encoding mode corresponding to a current frame based on the sound field classification result. In the above solution, the encoder side may determine an encoding mode corresponding to a current frame based on the sound field classification result. The encoding mode is the mode used in encoding the current frame of the three-dimensional audio signal. There are multiple encoding modes, and different encoding modes may be used based on different sound field classification results of the current frame. In the embodiment of the present application, an appropriate encoding mode is selected according to different sound field classification results of the current frame, so that the current frame is encoded by using the encoding mode. This improves the compression efficiency and hearing quality of the audio signal.
[0022] In a possible implementation, the step of determining an encoding mode corresponding to the current frame based on the sound field classification result includes the following: if the sound field classification result includes a non-uniform sound source number, or if the sound field classification result includes a non-uniform sound source number and a sound field type, determining an encoding mode corresponding to the current frame based on the non-uniform sound source number; if the sound field classification result includes a sound field type, or if the sound field classification result includes a non-uniform sound source number and a sound field type, determining an encoding mode corresponding to the current frame based on the sound field type; or if the sound field classification result includes a non-uniform sound source number and a sound field type, determining an encoding mode corresponding to the current frame based on the non-uniform sound source number and the sound field type. In the above solution, the encoder side may determine an encoding mode corresponding to the current frame based on the non-uniform sound source number and / or the sound field type, and determine the corresponding encoding mode based on the sound field classification result of the current frame, so that the determined encoding mode can be applied to the current frame of the three-dimensional audio signal. This improves the efficiency of encoding.
[0023] In a possible implementation, the step of determining the coding mode corresponding to the current frame based on the number of non-uniform sound sources includes: determining that the coding mode is the first coding mode if the number of non-uniform sound sources meets a second preset condition; or determining that the coding mode is the second coding mode if the number of non-uniform sound sources does not meet the second preset condition; the first coding mode is a HOA coding mode based on virtual speaker selection or a HOA coding mode based on directional sound coding, and the second coding mode is a HOA coding mode based on virtual speaker selection or a HOA coding mode based on directional sound coding, and the first coding mode and the second coding mode are different coding modes. In the above solution, the coding mode can be classified into two types, a first coding mode and a second coding mode, based on the different numbers of non-uniform sound sources. The encoder side obtains a second preset condition; that is, determines whether the number of non-uniform sound sources meets the second preset condition; and determines that the coding mode is the first coding mode if the number of non-uniform sound sources meets the second preset condition. Or, if the number of non-uniform sound sources does not satisfy the second preset condition, determine that the coding mode is the second coding mode. In this embodiment of the present application, the division of the coding mode of the current frame is implemented, and it can be determined whether the number of non-uniform sound sources satisfies the second preset condition, so as to accurately identify whether the coding mode of the current frame belongs to the first preset condition or the second preset condition.
[0024] In a possible implementation, the second preset condition includes that the number of heterogeneous sound sources is greater than a first threshold and less than a second threshold, and the second threshold is greater than the first threshold, or that the number of heterogeneous sound sources is less than or equal to the first threshold or greater than or equal to the second threshold, and the second threshold is greater than the first threshold.
[0025] In a possible implementation, determining the coding mode corresponding to the current frame based on the sound field type includes: determining that the coding mode is a virtual speaker-based HOA coding mode if the sound field type is a non-uniform sound field; or determining that the coding mode is a directional voice coding-based HOA coding mode if the sound field type is a distributed sound field.
[0026] In a possible implementation, the step of determining the coding mode corresponding to the current frame based on the sound field classification result includes: determining the initial coding mode corresponding to the current frame based on the sound field classification result of the current frame; obtaining a hangover time window in which the current frame is located, the hangover time window includes the initial coding mode of the current frame and the coding modes of N-1 frames before the current frame, where N is the length of the hangover time window; and determining the coding mode of the current frame based on the initial coding mode of the current frame and the coding modes of the N-1 frames. In the above solution, in this embodiment of the present application, the initial coding mode of the current frame is modified based on the hangover time window to obtain the coding mode of the current frame. This ensures that the coding modes of successive frames are not frequently switched, improving the efficiency of coding.
[0027] In a possible implementation, the method further includes: determining an encoding parameter corresponding to the current frame based on the sound field classification result. In the above solution, the encoder side may determine an encoding parameter corresponding to the current frame based on the sound field classification result. The encoding parameter is a parameter used in encoding the current frame of the three-dimensional audio signal. There are multiple encoding parameters, and different encoding parameters may be used based on different sound field classification results of the current frame. In this embodiment of the application, for different sound field classification results of the current frame, an appropriate encoding parameter is selected, and the current frame is encoded based on the encoding parameter. This improves the compression efficiency and hearing quality of the audio signal.
[0028] In a possible implementation, the coding parameters include at least one of a number of channels of the virtual speaker signals, a number of channels of the residual signal, a number of coding bits of the virtual speaker signals, a number of coding bits of the residual signal, or a number of votes for searching for the best matching speaker. The virtual speaker signals and the residual signal are generated based on the three-dimensional audio signal.
[0029] The number of votes satisfies the relationship 1≦I≦d, where I is the number of votes and d is the number of non-uniform sound sources included in the sound field classification result. In the above solution, the encoder side determines the number of votes for searching for the best matching speaker based on the number of non-uniform sound sources of the current frame. The number of votes is less than or equal to the number of non-uniform sound sources of the current frame, so that the number of votes can be adapted to the actual situation in the sound field classification of the current frame. This solves the problem that the number of votes for searching for the best matching speaker needs to be determined when the current frame is encoded.
[0030] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources and the sound field type. When the sound field type is a non-uniform sound field, the number of channels of the virtual speaker signal satisfies the relationship of F=min(S,PF), where F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the encoder. Or, when the sound field type is a distributed sound field, the number of channels of the virtual speaker signal satisfies the relationship of F=1, where F is the number of channels of the virtual speaker signal. In the above solution, the number of channels of the virtual speaker signal is the number of channels for transmitting the virtual speaker signal, and the number of channels of the virtual speaker signal can be determined based on the non-uniform sound source and the sound field type. In the above calculation method, when the sound field type is a distributed sound field, the number of channels of the virtual speaker signal is determined to be 1 to improve the coding efficiency of the current frame. When the sound field type is a non-uniform sound source, min represents the operation of selecting the minimum value, that is, the operation of selecting the minimum value of S and PF as the channel number of the virtual speaker signal, so that the channel of the virtual speaker signal can be adapted to the actual situation in the sound field classification of the current frame. This solves the problem that the channel number of the virtual speaker signal needs to be determined when encoding the current frame.
[0031] In a possible implementation, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the relationship R=max(C-1,PR), where PR is the number of channels of the residual signal preset by the encoder, and C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder. Or, when the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the relationship R=CF, where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal. In the above solution, after the number of channels of the virtual speaker signal is obtained, the number of channels of the residual signal can be calculated based on the preset number of channels of the residual signal and the sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal. The value of PR can be preset on the encoder side, and the value of R can be obtained according to the calculation formula of max(C-1,PR). The sum of the number of preset channels of the residual signal and the number of preset channels of the virtual speaker signal is preset on the encoder side. Note that C is sometimes referred to as the total number of transmission channels.
[0032] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources. The number of channels of the virtual speaker signals satisfies the relationship F=min(S, PF), where F is the number of channels of the virtual speaker signals, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signals preset by the encoder.
[0033] In a possible implementation, the number of channels of the residual signal satisfies the relationship R=CF, where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker preset by the encoder, and F is the number of channels of the virtual speaker signal. In the above solution, after the number of channels of the virtual speaker signal is obtained, the number of channels of the residual signal can be calculated based on the number of channels of the virtual speaker signal and the sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal. The sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal is preset on the encoder side. Note that C is sometimes referred to as the total number of transmission channels.
[0034] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type. The number of coding bits of the virtual speaker signal is obtained based on a ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel. The number of coding bits of the residual signal is obtained based on a ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel. The number of coding bits of the transmission channel includes the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal, and when the number of non-uniform sound sources is less than or equal to the number of channels of the virtualized speaker signal, the ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel is obtained by increasing an initial ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel.
[0035] In a possible implementation, the method further includes: encoding the current frame and the sound field classification result; and writing the encoded current frame and the sound field classification result into a bitstream.
[0036] According to a second aspect, an embodiment of the present application further provides a three-dimensional audio signal processing method, including: receiving a bitstream; decoding the bitstream to obtain a sound field classification result of a current frame; and obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result. In the above solution, the sound field classification result can be used to decode the current frame in the bitstream. Therefore, the decoder side performs decoding in a decoding manner that matches the sound field of the current frame to obtain the three-dimensional audio signal sent from the encoder side. This implements the transmission of the audio signal from the encoder side to the decoder side.
[0037] In a possible implementation, obtaining a three-dimensional audio signal of a decoded current frame based on the sound field classification result includes: determining a decoding mode of the current frame based on the sound field classification result; and obtaining a three-dimensional audio signal of the decoded current frame based on the decoding mode.
[0038] In a possible implementation, determining a decoding mode of the current frame based on the sound field classification result includes: determining a decoding mode of the current frame based on the non-uniform sound source number if the sound field classification result includes a non-uniform sound source number or a non-uniform sound source number and a sound field type; determining a decoding mode of the current frame based on the sound field type if the sound field classification result includes a sound field type or a non-uniform sound source number and a sound field type; or determining a decoding mode of the current frame based on the non-uniform sound source number and a sound field type if the sound field classification result includes a non-uniform sound source number and a sound field type.
[0039] In a possible implementation, determining a decoding mode corresponding to a current frame based on the number of non-uniform sound sources includes: determining that the decoding mode is a first decoding mode if the number of non-uniform sound sources meets a preset condition; or determining that the decoding mode is a second decoding mode if the number of non-uniform sound sources does not meet the preset condition, where the first decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional voice coding, and the second decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional voice coding, and the first decoding mode and the second decoding mode are different decoding modes.
[0040] In a possible implementation, the preset condition includes that the number of heterogeneous sound sources is greater than a first threshold and less than a second threshold, and the second threshold is greater than the first threshold, or that the preset condition includes that the number of heterogeneous sound sources is less than or equal to the first threshold or greater than or equal to the second threshold, and the second threshold is greater than the first threshold.
[0041] In a possible implementation, obtaining a three-dimensional audio signal of a decoded current frame based on the sound field classification result includes: determining a decoding parameter of the current frame based on the sound field classification result; and obtaining a three-dimensional audio signal of the decoded current frame based on the decoding parameter.
[0042] In a possible implementation, the decoding parameters include at least one of the following: a number of channels of the virtual speaker signals, a number of channels of the residual signal, a number of decoded bits of the virtual speaker signals, or a number of decoded bits of the virtual speaker signals, wherein the virtual speaker signals and the residual signal are obtained by decoding a bitstream.
[0043] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources and the sound field type. When the sound field type is a non-uniform sound field, the number of channels of the virtual speaker signal satisfies the relationship F=min(S, PF), where F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the decoder. Or, when the sound field type is a distributed sound field, the number of channels of the virtual speaker signal satisfies the relationship F=1, where F is the number of channels of the virtual speaker signal.
[0044] In a possible implementation, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the relationship R=max(C-1,PR), where PR is the number of channels of the residual signal preset by the decoder, and C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder. Or, when the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the relationship R=CF, where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0045] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources. The number of channels of the virtual speaker signals satisfies the relationship F=min(S, PF), where F is the number of channels of the virtual speaker signals, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signals preset by the decoder.
[0046] In a possible implementation, the number of channels of the residual signal satisfies the relationship R=CF, where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0047] In a possible implementation, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type. The number of decoded bits of the virtual speaker signal is obtained based on a ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel. The number of decoded bits of the residual signal is obtained based on a ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel. The number of decoded bits of the transmission channel includes the number of decoded bits of the virtual speaker signal and the number of decoded bits of the residual signal, and when the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signal, the ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel is obtained by increasing an initial ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of the transmission channel.
[0048] According to a third aspect, an embodiment of the present application further provides a three-dimensional audio signal processing apparatus, including: a linear analysis module configured to perform a linear decomposition on the three-dimensional audio signal to obtain a linear decomposition result, a parameter generation module configured to obtain sound field classification parameters corresponding to a current frame based on the linear decomposition result, and a sound field classification module configured to determine a sound field classification result for the current frame based on the sound field classification parameters.
[0049] In a third aspect of the present application, the modules included in the three-dimensional audio signal processing device may further execute the steps described in the first aspect and possible implementations. For details, please refer to the description of the first aspect and possible implementations.
[0050] According to a fourth aspect, an embodiment of the present application further provides a three-dimensional audio signal processing apparatus, including: a receiving module configured to receive a bitstream, a decoding module configured to decode the bitstream to obtain a sound field classification result for a current frame, and a signal generating module configured to obtain a three-dimensional audio signal for the decoded current frame based on the sound field classification result.
[0051] In a fourth aspect of the present application, the modules included in the three-dimensional audio signal processing device may further execute the steps described in the second aspect and possible implementations. For details, please refer to the description of the second aspect and possible implementations.
[0052] In a possible implementation, the number of coding bits of the virtual speaker signals satisfies the following relationship:
[0053]
number
[0054] where core_numbit is the number of coding bits of the virtual speaker signal, fac1 is a weighting factor assigned to the coding bits of the virtual speaker signal, fac2 is a weighting factor assigned to the coding bits of the residual signal, round represents rounding down, F is the number of channels of the virtual speaker signal, R represents the number of channels of the residual signal, and numbit is the sum of the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal. The number of coding bits of the residual signal satisfies the following relationship, i.e.
[0055]
number
[0056] res_numbit is the number of coding bits of the residual signal, core_numbit is the number of coding bits of the virtual speaker signals, and numbit is the sum of the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal.
[0057] A possible implementation would be:
[0058]
number
[0059] It is.
[0060] In a possible implementation, the number of coding bits of the residual signal satisfies the following relationship:
[0061]
number
[0062] res_numbit is the number of coding bits of the residual signal, fac1 is a weighting factor assigned to the coding bits of the virtual speaker signals, fac2 is a weighting factor assigned to the coding bits of the residual signal, round represents rounding, F is the number of channels of the virtual speaker signals, R represents the number of channels of the residual signal, and numbit is the sum of the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal.
[0063] The number of coding bits for the virtual speaker signal satisfies the following relationship:
[0064]
number
[0065] core_numbit is the number of coding bits of the virtual speaker signal, res_numbit is the number of coding bits of the residual signal, and numbit is the sum of the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal.
[0066] In a possible implementation, the number of coding bits for each virtual speaker signal satisfies the following relationship:
[0067]
number
[0068] where core_ch_numbit is the number of coding bits of each virtual speaker signal, fac1 is a weighting factor assigned to the coding bits of the virtual speaker signals, fac2 is a weighting factor assigned to the coding bits of the residual signal, round represents truncation, F is the number of channels of the virtual speaker signals, R represents the number of channels of the residual signal, and numbit is the sum of the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal.
[0069] The number of coding bits for each residual signal satisfies the following relationship:
[0070]
number
[0071] where res_numbit is the number of coding bits of each residual signal, fac1 is a weighting factor assigned to the coding bits of the virtual speaker signals, fac2 is a weighting factor assigned to the coding bits of the residual signals, round represents rounding, F is the number of channels of the virtual speaker signals, R represents the number of channels of the residual signals, and numbit is the sum of the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signals.
[0072] According to a fifth aspect, an embodiment of the present application provides a computer readable storage medium storing instructions which, when executed on a computer, enable the computer to perform the method of the first or second aspect.
[0073] According to a sixth aspect, an embodiment of the present application provides a computer program product comprising instructions, which when executed on a computer, enable the computer to perform the method of the first or second aspect.
[0074] According to a seventh aspect, an embodiment of the present application provides a computer-readable storage medium comprising a bitstream generated in the method of the first aspect.
[0075] According to an eighth aspect, an embodiment of the present application provides a communication device. The communication device may include an entity, such as a terminal device or a chip. The communication device includes a processor and a memory. The memory is configured to store instructions, and the processor is configured to execute the instructions in the memory, enabling the communication device to perform the method in any one of the implementations of the first or second aspect.
[0076] According to a ninth aspect, the present application provides a chip system. The chip system includes a processor configured to support a voice encoder or a voice decoder in implementing the functions of the aforementioned aspects, such as performing data and / or information transmission or processing in the aforementioned methods. In a possible design, the chip system further includes a memory. The memory is configured to store program instructions and data required for the voice encoder or the voice decoder. The chip system may include a chip, or may include a chip and another discrete component.
[0077] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0078] In this embodiment of the present application, first, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result. Then, based on the linear decomposition result, a sound field classification parameter corresponding to the current frame is obtained. Finally, based on the sound field classification parameter, a sound field classification result of the current frame is determined. In this embodiment of the present application, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result of the current frame. Then, based on the linear decomposition result, a sound field classification parameter corresponding to the current frame is obtained. Thus, based on the sound field classification parameter, a sound field classification result of the current frame is determined, and based on the sound field classification result, a sound field classification of the current frame can be implemented. In this embodiment of the present application, a sound field classification is performed on a three-dimensional audio signal to accurately identify the three-dimensional audio signal. [Brief description of the drawings]
[0079] [Figure 1] FIG. 1 is a schematic diagram illustrating a configuration structure of a voice processing system according to an embodiment of the present application. [Figure 2a] 1 is a schematic diagram of a speech encoder and a speech decoder used in a terminal device according to an embodiment of the present application; [Figure 2b] FIG. 2 is a schematic diagram of a voice encoder used in a radio device or core network device according to an embodiment of the present application; [Figure 2c] FIG. 2 is a schematic diagram of a voice decoder used in a radio device or core network device according to an embodiment of the present application; [Figure 3a] 1 is a schematic diagram of a multi-channel encoder and a multi-channel decoder used in a terminal device according to an embodiment of the present application; [Figure 3b] FIG. 2 is a schematic diagram of a multi-channel encoder used in a wireless device or core network device according to an embodiment of the present application; [Figure 3c]FIG. 2 is a schematic diagram of a multi-channel decoder used in a radio device or core network device according to an embodiment of the present application; [Figure 4] FIG. 2 is a schematic diagram illustrating a three-dimensional audio signal processing method according to an embodiment of the present application. [Diagram 5] FIG. 2 is a schematic diagram illustrating a three-dimensional audio signal processing method according to an embodiment of the present application. [Figure 6] FIG. 2 is a schematic diagram illustrating a three-dimensional audio signal processing method according to an embodiment of the present application. [Figure 7] FIG. 2 is a schematic diagram illustrating a three-dimensional audio signal processing method according to an embodiment of the present application. [Figure 8] 4 is a schematic flow chart illustrating encoding of a hybrid HOA encoder according to an embodiment of the present application; [Figure 9] 4 is a schematic flow chart illustrating the determination of an encoding mode of an HOA signal according to an embodiment of the present application; [Figure 10] 4 is a schematic flow chart illustrating decoding of a hybrid HOA decoder according to an embodiment of the present application; [Figure 11] 4 is a schematic flow chart illustrating encoding of an MP-based HOA encoder according to an embodiment of the present application; [Figure 12] 1 is a schematic diagram illustrating a configuration structure of a speech encoding device according to an embodiment of the present application; [Figure 13] 1 is a schematic diagram illustrating a configuration structure of an audio decoding device according to an embodiment of the present application; [Figure 14] FIG. 2 is a schematic diagram illustrating a configuration structure of another audio encoding device according to an embodiment of the present application; [Figure 15] FIG. 2 is a schematic diagram illustrating a configuration structure of another audio decoding device according to an embodiment of the present application; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0080] Hereinafter, embodiments of the present application will be described with reference to the drawings.
[0081] In the specification, claims, and accompanying drawings of this application, terms such as "first" and "second" are intended to distinguish between similar objects, but do not necessarily indicate a particular order or arrangement. Terms used in such aspects are interchangeable under appropriate circumstances, and should be understood as merely distinguishing aspects used in describing objects having the same attributes in the embodiments of this application. Furthermore, the terms "include", "contain", and any other variations are meant to cover a non-exclusive inclusion, and a process, method, system, product, or apparatus that includes a set of units may include, but is not necessarily limited to, other units that are not expressly recited or that are inherent to such a process, method, system, product, or apparatus.
[0082] Sound is a continuous wave produced by the vibration of an object. An object that emits sound waves by vibration is called a sound source. When sound waves propagate through a medium (such as air, a solid, or a liquid), the hearing organs of humans or animals can detect the sound.
[0083] Characteristics of sound waves include tone, sound intensity, and timbre. Tone describes the pitch of a sound. Sound intensity describes the strength of a sound. Sound intensity is also called sound pressure or sound number. The unit of sound intensity is the decibel (dB). Timbre is also called sound quality.
[0084] The frequency of a sound wave determines the pitch of a tone. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is Hertz (Hz). Sound frequencies perceived by the human ear range from 20 Hz to 20,000 Hz.
[0085] The amplitude of a sound wave determines the strength of the sound. The greater the amplitude, the greater the sound intensity. The closer the distance from the sound source, the greater the sound intensity.
[0086] The waveform of a sound wave determines the timbre. The waveforms of sound waves include square waves, sawtooth waves, sine waves, and pulse waves.
[0087] Sounds are classified into regular sounds and irregular sounds based on the characteristics of sound waves. Irregular sounds are those generated by irregular vibrations of the sound source. Irregular sounds are, for example, noises that affect human work, study, and rest. Regular sounds are those generated by regular vibrations of the sound source. Regular sounds include conversations and music. When sounds are expressed electrically, regular sounds are analog signals that change continuously in the time-frequency domain. This analog signal is sometimes called an audio signal (acoustic signal). The audio signal is an information carrier that conveys conversations, music, and sound effects.
[0088] Since the human auditory system can identify the position distribution of sound sources in space, when listening to sounds in space, the listener can perceive not only the pitch, acoustic intensity, and timbre of the sound but also the position of the sound.
[0089] With the increasing attention to and quality requirements for the auditory system experience, three-dimensional audio technologies have emerged to enhance the depth, immersion, and sense of space of sounds. Therefore, the listener can listen to sounds emitted from sound sources in all directions, and feel as if the space where the listener is located is surrounded by the spatial sound field (referred to as the sound field) generated by the sound source, and feel that the sound spreads around. Three-dimensional audio technologies create an "immersive" stereo effect that makes the listener feel as if they are in a place such as a movie theater or concert hall.
[0090] The three-dimensional sound technology is a technology in which the space outside the human ear is assumed as a system, and the signal received by the eardrum is obtained by filtering and extracting the sound emitted from the sound source by the system outside the ear, and outputting the signal. For example, the system outside the human ear can be defined as a system impulse response h(n), an arbitrary sound source can be defined as x(n), and the signal received by the eardrum is the convolution result of x(n) and h(n). In the embodiment of the present application, the three-dimensional sound signal can be a higher order Ambisonics (HOA) signal or a first order Ambisonics (FOA) signal. The three-dimensional sound can also be called a three-dimensional sound effect, spatial sound, three-dimensional sound field reconstruction, virtual 3D sound, or binaural sound.
[0091] Sound waves propagate in an ideal medium with wave number k=w / c and angular frequency w=2πf, where f is the frequency of the sound wave and c is the speed of sound. The sound pressure satisfies equation (1):
[0092]
number
[0093] is the Laplace operator.
[0094]
number
[0095] The spatial system outside the human ear is assumed to be a sphere, and the listener is assumed to be at the center of the sphere. Sound from outside the sphere is projected onto the surface of the sphere, and the sound outside the sphere is filtered out. The sound sources are assumed to be distributed on the sphere. The sound field generated by the sound sources on the surface of the sphere is used to fit the sound field generated by the original sound source, that is, the three-dimensional sound technology is a sound field fitting method. Specifically, the equation (1) is solved in the spherical coordinate system, and the equation (1) is solved in the passive spherical domain as the following equation (2). That is,
[0096]
number
[0097] r represents the spherical radius, θ represents the horizontal angle, φ represents the elevation angle, k represents the wave number, s represents the amplitude of an ideal plane wave, and m represents the order number (also called the order of the HOA signal).
[0098]
number
[0099] represents the spherical Bessel functions, also known as radial basis functions, where the first j represents the imaginary unit,
[0100]
number
[0101] does not change with angle.
[0102]
number
[0103] represents the spherical harmonic function in the θ, φ directions,
[0104]
number
[0105] represents the spherical harmonic function in the direction of the sound source. The coefficients of the three-dimensional sound signal satisfy equation (3). That is,
[0106]
number
[0107] By substituting equation (3) into equation (2), equation (2) can be transformed into equation (4).
[0108]
number
[0109]
number
[0110] represents the coefficients of an N-th order three-dimensional sound signal and is used to approximately describe a sound field. A sound field is a region in which sound waves exist in a medium. N is an integer equal to or greater than 1. For example, the value of N is an integer ranging from 2 to 6. The coefficients of the three-dimensional sound signal in the embodiment of the present application can be HOA coefficients or Ambisonic coefficients.
[0111] The three-dimensional sound signal is an information carrier that carries the spatial location information of the sound source in the sound field, and describes the sound field of the listener in space. Equation (4) shows that the sound field can be expanded on the sphere as a spherical harmonic function, that is, the sound field can be decomposed into a superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional sound signal can be expressed by using a superposition of multiple plane waves, and the sound field can be reconstructed based on the coefficients of the three-dimensional sound signal.
[0112] Compared with a 5.1 channel audio signal or a 7.1 channel audio signal, the Nth order HOA signal is (N+1) 2The HOA signal has channels. Therefore, the HOA signal contains a large amount of data used to describe the spatial information of the sound field. When the collecting device (e.g., microphone, etc.) transmits the three-dimensional sound signal to the reproducing device (e.g., speaker, etc.), a large bandwidth needs to be consumed. Currently, an encoder may compress and encode the three-dimensional sound signal by using a spatial squeezed surround sound coding (S3AC) method, a directional sound coding (DirAC) method, or an encoding method based on virtual speaker selection to obtain a bitstream, and transmit the bitstream to the reproducing device. The encoding method based on virtual speaker selection is sometimes called a match projection (MP) encoding method. In the following, the encoding method based on virtual speaker selection is used as an example for explanation. The reproducing device decodes the bitstream, reconstructs the three-dimensional sound signal, and reproduces the reconstructed three-dimensional sound signal. This reduces the amount of data and bandwidth occupation for transmitting the three-dimensional sound signal to the reproducing device.
[0113] For three-dimensional audio signals, currently, the sound field of the three-dimensional audio signal cannot be classified. How to classify the sound field of the three-dimensional audio signal is a technical problem to be solved in the embodiment of the present application. In the embodiment of the present application, a linear decomposition is performed on the three-dimensional audio signal to implement the sound field classification of the three-dimensional audio signal, which can accurately implement the sound field classification of the three-dimensional audio signal and obtain the sound field classification result of the current frame.
[0114] In addition, current encoders cannot obtain a high compression ratio when compressing and encoding 3D audio signals, so how to increase the compression ratio for performing compression encoding on 3D audio signals of different sound fields is another problem to be solved in the embodiments of this application.
[0115] An embodiment of the present application provides an audio coding technique, and in particular, a three-dimensional audio coding technique targeted at three-dimensional audio signals. Specifically, to improve conventional audio coding systems, an encoding technique is provided that represents three-dimensional audio signals by using a smaller number of channels. Audio coding (or commonly referred to as coding) includes two parts: audio encoding and audio decoding. Audio coding is performed at the transmission source side and involves processing (e.g., compressing) the original audio to reduce the amount of data required to represent the audio. This improves storage and / or transmission efficiency. Audio decoding is performed at the transmission destination side and involves the reverse processing to the encoder to reconstruct the original audio. The encoding part and the decoding part are also referred to as coding. Hereinafter, the implementation of the embodiment of the present application will be described in detail with reference to the accompanying drawings.
[0116] The technical solutions in the embodiments of the present application may be applied to various audio processing systems. Figure 1 is a schematic diagram showing a configuration structure of an audio processing system according to an embodiment of the present application. The audio processing system 100 may include an audio encoding device 101 and an audio decoding device 102. The audio encoding device 101 may be configured to generate a bitstream. The audio coding bitstream may then be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 may receive the bitstream and then perform the audio decoding function of the audio decoding device 102 to obtain a reconstructed signal.
[0117] In this embodiment of the present application, the voice encoding device can be used in various terminal devices that require voice communication, as well as wireless devices and core network devices that require transcoding. For example, the voice encoding device can be a voice encoder of a terminal device, a wireless device, or a core network device. Similarly, the voice decoding device can be used in various terminal devices that require voice communication, as well as wireless devices and core network devices that require transcoding. For example, the voice decoding device can be a voice decoder of a terminal device, a wireless device, or a core network device. For example, the voice encoder can include a radio access network, a media gateway in a core network, a transcoding device, a media resource server, a mobile terminal, and a fixed network terminal, etc. Alternatively, the voice encoder can be a voice encoder used for a virtual reality (VR) streaming media service.
[0118] In this embodiment of the application, an audio coding (audio encoding and audio decoding) module applicable to virtual reality streaming (VR streaming) media services is used as an example. The end-to-end audio signal processing procedure includes: After the audio signal A passes through the collection module, a pre-processing (audio pre-processing) operation is performed. The pre-processing operation includes: Filtering out the low frequency part of the signal, where the filter extraction can be performed by using 20Hz or 50Hz as the boundary point; and Extracting the directional information of the signal. After that, encoding (audio encoding) and encapsulation (file / segment encapsulation) are performed, and the signal is delivered to the decoder side (delivery). The decoder side first performs decapsulation (file / segment decapsulation), then performs decoding (audio decoding), and performs binaural rendering (audio rendering) on the decoded signal. The signal obtained through rendering is mapped to the listener's headset (headphones), which can be an independent headset or a headset on a glasses device.
[0119] FIG. 2a is a schematic diagram in which a voice encoder and a voice decoder are used in a terminal device according to an embodiment of the present application. Each terminal device may include a voice encoder, a channel encoder, a voice decoder, and a channel decoder. Specifically, the channel encoder is configured to perform channel encoding on the voice signal, and the channel decoder is configured to perform channel decoding on the voice signal. For example, a first terminal device 20 may include a first voice encoder 201, a first channel encoder 202, a first voice decoder 203, and a first channel decoder 204. A second terminal device 21 may include a second voice decoder 211, a second channel encoder 212, a second voice decoder 213, and a second channel decoder 214. The first terminal device 20 is connected to a first network communication device 22, which is wireless or wired, the first network communication device 22 is connected to a second network communication device 23, which is wireless or wired, and the second terminal device 21 is connected to the second network communication device 23, which is wireless or wired. A wireless or wired network communication device may generally be a signal transmission device, for example a communication base station or a data switching device.
[0120] In voice communication, the terminal device acting as the transmitting end first performs voice collection, performs voice coding on the collected voice signal, then performs channel coding, and transmits the coded signal in a digital channel through a wireless network or a core network. The terminal device acting as the receiving end performs channel decoding based on the received signal to obtain a bit stream, and then restores the voice signal through voice decoding. The terminal device at the receiving end performs voice playback.
[0121] Fig. 2b is a schematic diagram of a voice encoder used in a radio device or core network device according to an embodiment of the present application. The radio device or core network device 25 includes: a channel decoder 251, another voice decoder 252, a voice encoder 253 provided in this embodiment of the present application, and a channel encoder 254. The another voice decoder 252 is another voice decoder other than the voice decoder. In the radio device or core network device 25, first, the channel decoder 251 performs channel decoding on the signal input to the device, and then the another voice decoder 252 performs voice decoding. After that, the voice encoder 253 provided in this embodiment of the present application performs voice encoding, and finally, the channel encoder 254 performs channel encoding on the voice signal, and then transmits the encoded voice signal after the channel encoding is completed. The another voice decoder 252 performs voice decoding on the bit stream decoded by the channel decoder 251.
[0122] Fig. 2c is a schematic diagram of a voice decoder used in a radio device or core network device according to an embodiment of the present application. The radio device or core network device 25 includes: a channel decoder 251, a voice decoder 255 provided in this embodiment of the present application, another voice encoder 256, and a channel encoder 254. The another voice encoder 256 is another voice encoder other than the voice encoder. In the radio device or core network device 25, first, the channel decoder 251 performs channel decoding on the signal input to the device, and then the voice decoder 255 decodes the received voice coding bit stream. After that, the another voice encoder 256 performs voice coding, and finally, the channel encoder 254 performs channel coding on the voice signal, and then transmits the coded voice signal after the channel coding is completed. In the radio device or core network device, if transcoding needs to be implemented, it is necessary to perform the corresponding voice coding process. The radio device is a radio frequency-related device in communication, and the core network device is a core network-related device in communication.
[0123] In some embodiments of the present application, the voice coding device may be used in various terminal devices that require voice communication, as well as wireless devices and core network devices that require transcoding. For example, the voice coding device may be a multi-channel encoder of a terminal device, a wireless device, or a core network device. Similarly, the voice decoding device may be used in various terminal devices that require voice communication, as well as wireless devices and core network devices that require transcoding. For example, the voice decoding device may be a multi-channel decoder of a terminal device, a wireless device, or a core network device.
[0124] FIG. 3a is a schematic diagram illustrating the application of a multi-channel encoder and a multi-channel decoder to a terminal device according to an embodiment of the present application. Each terminal device may include a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder may perform the audio encoding method provided in the embodiment of the present application, and the multi-channel decoder may perform the audio decoding method provided in the embodiment of the present application. Specifically, the channel encoder is configured to perform channel encoding on the multi-channel signal, and the channel decoder is configured to perform channel decoding on the multi-channel signal. For example, the first terminal device 30 may include a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. The second terminal device 31 may include a second multi-channel encoder 311, a second channel encoder 312, a second multi-channel decoder 313, and a second channel decoder 314. The first terminal device 30 is connected to a first network communication device 32, which is wireless or wired, and the first network communication device 32 is connected to a second network communication device 33, which is wireless or wired, via a digital channel, and the second terminal device 31 is connected to the second network communication device 33, which is wireless or wired. The wireless or wired network communication device can generally be a signal transmission device, for example, a communication base station or a data switching device. In voice communication, the terminal device acting as a transmitting end performs multi-channel coding on the collected multi-channel signal, then performs channel coding, and transmits the coded signal in a digital channel through a wireless network or a core network. The terminal device acting as a receiving end performs channel decoding based on the received signal to obtain a bit stream of multi-channel signal coding, and then restores the multi-channel signal through multi-channel decoding. The terminal device at the receiving end performs reproduction.
[0125] Fig. 3b is a schematic diagram illustrating the application of a multi-channel encoder to a radio device or core network device according to one embodiment of the present application. The radio device or core network device 35 includes: a channel decoder 351, another speech decoder 352, a multi-channel encoder 353, and a channel encoder 354. Fig. 3b is similar to Fig. 2b, and the details will not be described again here.
[0126] Fig. 3c is a schematic diagram illustrating the application of a multi-channel decoder to a radio device or core network device according to one embodiment of the present application. The radio device or core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, a further speech encoder 356, and a channel encoder 354. Fig. 3c is similar to Fig. 2c, and the details will not be described again here.
[0127] The voice encoding may be part of a multi-channel encoder, and the voice decoding may be part of a multi-channel decoder. For example, performing multi-channel encoding on the collected multi-channel signal may be processing the collected multi-channel signal to obtain a voice signal. The obtained voice signal is then encoded according to the method provided in the embodiment of the present application. The decoder side encodes a bitstream based on the multi-channel signal, performs decoding to obtain a voice signal, and restores the multi-channel signal after the up-mix process. Therefore, the embodiment of the present application may also be applied to the multi-channel encoder and multi-channel decoder in the terminal device, the wireless device, or the core network device. In the wireless device or the core network device, if transcoding needs to be implemented, it is necessary to perform the corresponding multi-channel encoding process.
[0128] First, a three-dimensional audio signal processing method provided in an embodiment of the present application is described. The method can be executed by a terminal device. For example, the terminal device can be a voice encoding device (hereinafter referred to as an encoder side or an encoder). Alternatively, it is not limited that the terminal device can be a three-dimensional audio signal processing device. As shown in Figure 4, the three-dimensional audio signal processing method mainly includes the following steps:
[0129] 401: Perform linear decomposition on a current frame of a 3D audio signal to obtain a linear decomposition result.
[0130] The encoder side may obtain a three-dimensional audio signal. For example, the three-dimensional audio signal may be a scene audio signal. Specifically, the three-dimensional audio signal may be a time domain signal or a frequency domain signal. Alternatively, the three-dimensional audio signal may be a signal obtained through downsampling.
[0131] In some embodiments of the present application, the three-dimensional audio signal includes a higher-order Ambisonics HOA signal or a first-order Ambisonics FOA signal. Alternatively, the three-dimensional audio signal can be another type of signal, without limitation. This is merely an example of the present application and is not intended to be a limitation to this embodiment of the present application.
[0132] For example, the three-dimensional audio signal may be a time-domain HOA signal or a frequency-domain HOA signal. As another example, the three-dimensional audio signal may include all channels of the HOA signal, or may include some HOA channels (e.g., FOA channels, etc.). Also, the three-dimensional audio signal may be all sampling points of the HOA signal, or 1 / Q downsampling points in the analyzed HOA signal obtained through downsampling, where Q is the downsampling interval and 1 / Q is the downsampling rate.
[0133] In this embodiment of the present application, the three-dimensional audio signal includes multiple frames. The following uses the processing of one frame of the three-dimensional audio signal as an example. For example, if the frame is a current frame, there is a previous frame before the current frame of the three-dimensional audio signal, and there is a next frame after the current frame. In addition, in this embodiment of the present application, the processing method of another frame other than the current frame in the three-dimensional audio signal is similar to the method for processing the current frame. The following uses the processing of the current frame as an example.
[0134] In this embodiment of the present application, after obtaining a current frame of a three-dimensional audio signal, first perform linear decomposition on the current frame to obtain a linear decomposition result of the current frame. There are several linear decomposition methods, which will be described in detail below.
[0135] In some embodiments of the present application, performing linear decomposition on a current frame of a three-dimensional audio signal to obtain a linear decomposition result in step 401 includes: A1: performing singular value decomposition on a current frame to obtain singular values corresponding to the current frame, where the linear decomposition result includes the singular values. A2: performing a principal component analysis on the current frame to obtain a first feature value corresponding to the current frame, where the linear decomposition result includes the first feature value; or A3: A step of performing independent component analysis on the current frame to obtain a second feature value corresponding to the current frame, where the linear decomposition result includes the second feature value.
[0136] There are several linear decomposition methods. For example, the linear decomposition may include at least one of the following: Singular Value Decomposition (SVD), Principal Component Analysis (PCA), and Independent Component Analysis (ICA). In different linear decomposition methods, the obtained linear decomposition results have different representations, which will be described in detail later.
[0137] In step A1, the linear decomposition may be a singular value decomposition. For example, it is assumed that the three-dimensional audio signal is an HOA signal. The HOA signal forms a matrix A, which is an L*K matrix, where L is equal to the number of channels of the HOA signal, and K is the number of signal points of each channel of the HOA signal in the current frame. For example, the number of signal points may include the number of frequencies, the number of sampling points in the time domain, or the number of frequencies or sampling points after downsampling. A singular value decomposition is performed on the matrix A, and the following relationship is satisfied: A=UΣV T
[0138] U is an L*L matrix, V is a K*K matrix, the superscript T is the transpose of matrix V, and * denotes multiplication. Σ is an L*K diagonal matrix, where each element on the diagonal of the matrix is a singular value of matrix A obtained by singular value decomposition, and all elements outside the diagonal are 0. The elements on the diagonal of diagonal matrix Σ, i.e., the singular values of matrix A, are represented as v[i], where i=0,1,...,min(L,K)-1.
[0139] It should be noted that if the three-dimensional audio signal is an HOA signal obtained through downsampling, K is the signal point number of each channel of the HOA signal in the current frame after downsampling. For example, the signal point number may be the number of sampling points or the number of frequencies.
[0140] In step A2, the linear decomposition may alternatively be a principal component analysis to obtain a feature value. In the following embodiment, the feature value obtained through the principal component analysis is defined as a first feature value to distinguish from another feature value. The specific implementation of the principal component analysis is not described again in this specification.
[0141] In step A3, the linear decomposition may alternatively be an independent component analysis to obtain a second feature value, and the specific implementation of the independent component analysis will not be described again in this specification.
[0142] In this embodiment of the present application, the linear decomposition of the current frame can be implemented by any one of the above implementations A1 to A3 to obtain multiple kinds of linear decomposition results.
[0143] 402: Obtain sound field classification parameters corresponding to the current frame based on the linear decomposition result.
[0144] After obtaining the linear analysis result of the current frame, the encoder side analyzes the linear decomposition result to obtain a sound field classification parameter corresponding to the current frame. The sound field classification parameter is obtained by analyzing the linear decomposition result of the current frame, and the sound field classification parameter is used to determine the sound field classification result of the current frame. Based on different specific implementations of the linear decomposition result, the sound field classification parameter may have multiple implementations.
[0145] In this embodiment of the present application, there may be one or more linear decomposition results. For example, the linear decomposition result includes a singular value, and the singular value is v[i], where i=0,1,...,min(L,K)-1. If the current frame has only one singular value, there is only one value of i, i.e., v[0]. If the current frame has multiple singular values, there are multiple values of i, i.e., v[i], where i=1,...,min(L,K)-1.
[0146] In this embodiment of the present application, when there are two linear decomposition results, one sound field classification parameter is obtained. If the number of linear decomposition results is N, the number of sound field classification parameters obtained is N-1, and the value of N is not limited.
[0147] In some embodiments of the present application, the step of obtaining the sound field classification parameters corresponding to the current frame based on the linear decomposition result in step 402 includes: B1: Obtaining a ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer; and B2: Obtaining the i-th sound field classification parameter corresponding to the current frame based on this ratio.
[0148] The encoder side may obtain sound field classification parameters corresponding to the current frame based on the linear decomposition results. For example, there are multiple linear decomposition results of the current frame, and two consecutive linear analysis results among the multiple linear analysis results are represented as the i-th linear analysis result and the (i+1)-th linear analysis result of the current frame. In this case, the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame may be calculated, and the specific value of i is not limited.
[0149] Optionally, the i-th linear analysis result and the (i+1)-th linear analysis result are two successive linear analysis results in the current frame.
[0150] After this ratio is obtained, the i-th sound field classification parameter corresponding to the current frame can be obtained based on the ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame. It can be seen that the i-th sound field classification parameter can be calculated based on the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result. The (i+1)-th sound field classification parameter can be calculated based on the ratio of the (i+1)-th linear analysis result to the (i+2)-th linear analysis result, and the rest can be inferred by analogy. There is a correspondence between the linear analysis results and the sound field classification parameters.
[0151] In one implementation, a ratio of the i-th linear analysis result to the (i+1)-th linear analysis result may be used as the i-th sound field classification parameter. After the ratio of the i-th linear analysis result to the (i+1)-th linear analysis result is obtained, multiple calculation methods may be further performed on the ratio to obtain the i-th sound field classification parameter. For example, a multiplication operation is performed on the ratio based on a preset adjustment coefficient to obtain the i-th sound field classification parameter.
[0152] For example, when singular value decomposition is used for linear decomposition, based on the sound field classification parameters, singular values can be obtained through singular value decomposition, and a ratio parameter between two adjacent singular values is calculated and used as the sound field classification parameter.
[0153] For example, the ratio between the singular values temp[i] is calculated and used as the sound field classification parameter. For i=0,1,...,min(L,K)-2, temp[i] satisfies: temp[i]=v[i] / v[i+1]
[0154] If PCA or ICA is used for the linear decomposition, the sound field classification parameters can be determined based on the feature values. The method for calculating the sound field classification parameters is similar to the method for calculating the ratio temp between the singular values. Alternatively, the ratio of two consecutive feature values is calculated based on the feature values obtained via the linear decomposition, and the ratio is used as the sound field classification parameter.
[0155] It should be noted that if the number of feature values or singular values obtained via linear decomposition exceeds two, the sound field classification parameter is a vector. Otherwise, the sound field classification parameter is a scalar. For example, for v[i], if the value of i is equal to 2, the calculated temp[i] is a scalar, i.e. there is only one value of temp. For v[i], if the value of i is greater than 2, the calculated temp[i] is a vector, i.e. temp contains at least two elements.
[0156] 403: Determine a sound field classification result for the current frame based on the sound field classification parameters.
[0157] In this embodiment of the present invention, after obtaining the sound field classification parameters corresponding to the current frame, the encoder side may perform sound field classification on the current frame based on the sound field classification parameters. The sound field classification parameters corresponding to the current frame may indicate parameters required for classifying the sound field corresponding to the current frame, so that the sound field classification result of the current frame may be obtained based on the sound field classification parameters.
[0158] In some embodiments of the present application, the sound field classification result may include at least one of a sound field type and a number of non-uniform sound sources.
[0159] The sound field type is the sound field type of the current frame, and is determined after the sound field classification is performed on the current frame. There are multiple ways to classify the sound field type. For example, the sound field type can be classified into a first sound field type and a second sound field type. Alternatively, the sound field type can be classified into a first sound field type, a second sound field type, a third sound field type, etc. In particular, the number of sound field types that can be classified can be determined based on the application scenario. As another example, the sound field type can include a non-uniform sound field and a distributed sound field. A non-uniform sound field means that there are point sound sources with different positions and / or directions in the sound field, and a distributed sound field is a sound field that does not include a non-uniform sound source. For example, a point sound source with different positions and / or directions is a non-uniform sound source, a sound field that includes a non-uniform sound source is a non-uniform sound field, and a sound field that does not include a non-uniform sound source is a distributed sound field.
[0160] A non-uniform sound source is a point sound source having different positions and / or directions, and the number of non-uniform sound sources included in the current frame is called the number of non-uniform sound sources. The sound field of the current frame can also be classified based on the number of non-uniform sound sources.
[0161] In some embodiments of the present application, there are multiple sound field classification parameters. The sound field classification result includes a sound field type.
[0162] In step 403, determining a sound field classification result of the current frame based on the sound field classification parameters includes: determining that the sound field type is a distributed sound field when all of the values of the plurality of sound field classification parameters satisfy a predetermined distributed sound source determination condition; or A step of determining that the sound field type is a non-uniform sound field when at least one of the values of the plurality of sound field classification parameters satisfies a preset non-uniform sound source determination condition.
[0163] The sound field types include a non-uniform sound field and a distributed sound field. In this embodiment of the present invention, a distributed sound source judgment condition and a non-uniform sound source judgment condition are set in advance. The distributed sound source judgment condition is used to judge whether the sound field type is a distributed sound field, and the non-uniform sound source judgment condition is used to judge whether the sound field type is a non-uniform sound field. After a plurality of sound field classification parameters in a current frame are obtained, a judgment is performed based on the values of the plurality of sound field classification parameters and a preset condition. The specific implementation of the distributed sound source judgment condition and the non-uniform sound source judgment condition is not limited in this specification.
[0164] After the multiple sound field classification parameters are obtained, the encoder side determines that the sound field type is a distributed sound field if all the values of the multiple sound field classification parameters satisfy a preset distributed sound source determination condition. For example, the current frame corresponds to N sound field classification parameters. Only when all the values of the N sound field classification parameters satisfy a preset distributed sound source determination condition, the sound field type of the current frame is determined to be a distributed sound field.
[0165] After the multiple sound field classification parameters are obtained, the encoder side determines that the sound field type is a non-uniform sound field if at least one of the values of the multiple sound field classification parameters satisfies a preset non-uniform sound source determination condition. For example, the current frame corresponds to N sound field classification parameters. Only when at least one of the values of the N sound field classification parameters satisfies a preset non-uniform sound source determination condition, the sound field type is determined to be a non-uniform sound field.
[0166] Furthermore, in some embodiments of the present application, the distributed sound source determination condition includes the following: the value of the sound field classification parameter is less than a preset non-uniform sound source determination threshold value; or The non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a preset non-uniform sound source determination threshold value.
[0167] The non-uniform sound source determination threshold may be a preset threshold, and a specific value is not limited. The distributed sound source determination condition includes that the value of the sound field classification parameter is less than a predetermined non-uniform sound source determination threshold. Therefore, when the values of the multiple sound field classification parameters are all less than a predetermined non-uniform sound source determination threshold, the sound field type is determined to be a distributed sound field. The non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a predetermined non-uniform sound source determination threshold. Therefore, when at least one of the values of the multiple sound field classification parameters is equal to or greater than a predetermined non-uniform sound source determination threshold, the sound field type is determined to be a non-uniform sound field.
[0168] In some embodiments of the present application, there are multiple sound field classification parameters.
[0169] The sound field classification result includes a sound field type, or the sound field classification result includes a number of non-uniform sound sources and a sound field type.
[0170] In step 403, determining a sound field classification result of the current frame based on the sound field classification parameters includes: C1: Obtaining a number of non-uniform sound sources corresponding to a current frame according to the values of a plurality of sound field classification parameters; and C2: A step of determining a sound field type based on the number of non-uniform sound sources corresponding to the current frame.
[0171] After obtaining the plurality of sound field classification parameters corresponding to the current frame, the encoder side may obtain a non-uniform sound source number corresponding to the current frame based on the values of the plurality of sound field classification parameters. A non-uniform sound source is a point sound source having different positions and / or directions, and the number of non-uniform sound sources included in the current frame is called the non-uniform sound source number. The sound field of the current frame can be classified based on the non-uniform sound source number. After the non-uniform sound source number corresponding to the current frame is obtained to determine the sound field type, the sound field type corresponding to the current frame can be determined by analyzing the non-uniform sound source number corresponding to the current frame.
[0172] In some embodiments of the present application, there are multiple sound field classification parameters.
[0173] The sound field classification result includes the number of non-uniform sound sources.
[0174] In step 403, determining a sound field classification result of the current frame based on the sound field classification parameters includes: D1: A step of obtaining a number of non-uniform sound sources corresponding to a current frame based on values of a plurality of sound field classification parameters.
[0175] After obtaining the plurality of sound field classification parameters corresponding to the current frame, the encoder side may obtain a non-uniform sound source number corresponding to the current frame according to the values of the plurality of sound field classification parameters. The non-uniform sound sources are point sound sources having different positions and / or directions, and the number of non-uniform sound sources included in the current frame is called the non-uniform sound source number.
[0176] Furthermore, in some embodiments of the present application, the multiple sound field classification parameters are temp[i], i=0,1,...,min(L,K)-2, where L represents the number of channels in the current frame, K represents the number of signal points corresponding to each channel in the current frame, and min represents the operation of selecting the minimum value. For example, the number of signal points can be the number of frequencies, the number of sampling points in the time domain, or the number of frequencies or sampling points in the time domain after downsampling.
[0177] In step C1 or step D1, the step of obtaining the number of non-uniform sound sources corresponding to the current frame according to the values of the plurality of sound field classification parameters includes: A step of sequentially executing the next decision procedure starting from i=0. A step of determining whether temp[i] exceeds a preset non-uniform sound source determination threshold value; and In the present determination procedure, if temp[i] is less than the non-uniform sound source determination threshold, the value of i is updated to i+1, and the next determination procedure is continued. If temp[i] is equal to or greater than the non-uniform sound source determination threshold in this determination procedure, the execution of this determination procedure is terminated, and it is determined that i+1 in this determination procedure, which is incremented by 1, is equal to the number of non-uniform sound sources.
[0178] Specifically, the encoder side can estimate the number of non-uniform sound sources and determine the sound field type based on the sound field classification parameters.
[0179] The types of sound fields include non-uniform sound fields and distributed sound fields. A non-uniform sound field is one in which point sound sources with different positions and directions exist within the sound field. A distributed sound field is one that does not contain a non-uniform sound source.
[0180] All sound field classification parameter values are distributed sound source If the determination condition is satisfied, the sound field type is a distributed sound field.
[0181] Sound field classification parameter values are non-uniform sound sourceIf the judgment condition is satisfied, the sound field type is judged to be a non-uniform sound field. sound source The estimation may be based on the order number of values among the values of the sound field classification parameter that satisfy the judgment condition.
[0182] For example, when the ratio temp[i] between singular values is used as a sound field classification parameter, the sound field type and the number of non-uniform sound sources are estimated based on the sound field classification parameter, and the value of temp[i] is determined sequentially from i=0. When the value of i is m, the value of the m-th sound field classification parameter is expressed as temp[m]. When the m-th sound field classification parameter satisfies temp[m]≧TH1, the sound field type is a non-uniform sound field, and there are (m+1) non-uniform sound sources in the sound field of the current frame. When temp[m]≧TH1, the sound field type is a distributed sound field. The range of the value of m is [0,1,...,min(L,K)-2], TH1 is a preset non-uniform sound source determination threshold, and the value of TH1 is a constant, for example, the value of TH1 may be 30 or 100. In this embodiment of the present application, the value of TH1 is not limited.
[0183] In some embodiments of the present application, the step of determining the sound field type based on the number of non-uniform sound sources corresponding to the current frame in step C2 includes: determining that the sound field type is a first sound field type if the number of non-uniform sound sources satisfies a first preset condition; or If the number of non-uniform sound sources does not satisfy the first preset condition, determining that the sound field type is the second sound field type.
[0184] The number of non-uniform sound sources corresponding to the first sound field type is different from the number of non-uniform sound sources corresponding to the second sound field type.
[0185] Specifically, the sound field type can be classified into two types, a first sound field type and a second sound field type, based on the difference in the number of non-uniform sound sources. The encoder side acquires a first preset condition. That is, determining whether the number of non-uniform sound sources satisfies the first preset condition. And, if the number of non-uniform sound sources satisfies the first preset condition, determining that the sound field type is the first sound field type. Or, if the number of non-uniform sound sources does not satisfy the first preset condition, determining that the sound field type is the second sound field type. In this embodiment of the present application, the division of the sound field type of the current frame is implemented, and it can be determined whether the number of non-uniform sound sources satisfies the first preset condition to accurately identify that the sound field type of the current frame belongs to the first sound field type or the second sound field type.
[0186] In some embodiments of the present application, the first preset condition includes that the number of non-uniform sound sources is greater than a first threshold or less than a second threshold, and the second threshold is greater than the first threshold; or The first preset condition includes that the number of heterogeneous sound sources is equal to or less than a first threshold or equal to or more than a second threshold, and that the second threshold exceeds the first threshold.
[0187] The specific values of the first threshold and the second threshold are not limited and can be specifically determined based on the application scenario. The second threshold exceeds the first threshold. Therefore, the first threshold and the second threshold can constitute a preset range, and the first preset condition can be that the number of non-uniform sound sources falls within the preset range, or the first preset condition can be that the number of non-uniform sound sources exceeds the preset range. The number of non-uniform sound sources is determined based on the first threshold and the second threshold in the first preset condition, and can accurately identify that the sound field type of the current frame belongs to the first sound field type or the second sound field type by determining whether the number of non-uniform sound sources meets the first preset condition.
[0188] For example, the first threshold is 0, the second threshold is 3, and the number of non-uniform sound sources is represented as n. In this case, the first preset condition may be 0 < n < 3, or the first preset condition may be n ≥ 3 or n = 0.
[0189] In some embodiments of the present application, the step of determining the sound field classification result of the current frame based on the sound field classification parameter further includes the following. That is, the step of determining the sound field classification result of the current frame based on the sound field classification parameter and another parameter including the characteristics of the tridimensional voice signal.
[0190] There are multiple implementations for another parameter indicating the characteristics of the tridimensional voice signal. For example, another parameter indicating the characteristics of the tridimensional voice signal may include at least one of the following. That is, the energy ratio parameter of the tridimensional voice signal, the high-frequency analysis parameter of the tridimensional voice signal, and the low-frequency characteristic analysis parameter of the tridimensional voice signal, etc.
[0191] As shown in FIG. 5, the tridimensional voice signal processing method according to an embodiment of the present application mainly includes the following steps.
[0192] 501: Executing linear decomposition on the current frame of the tridimensional voice signal to obtain a linear decomposition result.
[0193] 502: Obtaining the sound field classification parameter corresponding to the current frame based on the linear decomposition result.
[0194] 503: Determining the sound field classification result of the current frame based on the sound field classification parameter.
[0195] The implementation of steps 501 to 503 is the same as the implementation of steps 401 to 403 in the foregoing embodiments. For steps 501 to 503, they will not be described in detail again in this specification.
[0196] 504: Determining an encoding mode corresponding to the current frame based on the sound field classification result.
[0197] The encoder side may perform steps 501 to 503. After obtaining the sound field classification result of the current frame, the encoder side may determine an encoding mode corresponding to the current frame based on the sound field classification result. The encoding mode is a mode used in encoding the current frame of the three-dimensional audio signal. There are multiple encoding modes, and different encoding modes may be used based on different sound field classification results of the current frame. In this embodiment of the present invention, an appropriate encoding mode is selected for different sound field classification results of the current frame, so that the current frame is encoded by using the encoding mode. This improves the compression efficiency and the hearing quality of the audio signal.
[0198] Furthermore, in some embodiments of the present application, the step of determining the coding mode corresponding to the current frame based on the sound field classification result in step 503 includes: E1: When the sound field classification result includes a non-uniform sound source number, or when the sound field classification result includes a non-uniform sound source number and a sound field type, a step of determining an encoding mode corresponding to a current frame based on the non-uniform sound source number. E2: When the sound field classification result includes a sound field type, or when the sound field classification result includes a non-uniform sound source number and a sound field type, determining an encoding mode corresponding to the current frame based on the sound field type; or E3: If the sound field classification result includes the number of non-uniform sound sources and the type of sound field, determining an encoding mode corresponding to the current frame based on the number of non-uniform sound sources and the type of sound field.
[0199] In step E1, after the encoder side obtains the number of non-uniform sound sources of the current frame, the number of non-uniform sound sources can be used to determine the coding mode corresponding to the current frame. In step E2, after the encoder side obtains the sound field type of the current frame, the sound field type can be used to determine the coding mode corresponding to the current frame. In step E3, after the encoder side obtains the number of non-uniform sound sources and the sound field type, the number of non-uniform sound sources and the sound field type can be used to determine the coding mode corresponding to the current frame. Therefore, the encoder side can determine the coding mode corresponding to the current frame based on the number of non-uniform sound sources and / or the sound field type to determine the corresponding coding mode based on the sound field classification result of the current frame, so that the determined coding mode can be adapted to the current frame of the three-dimensional audio signal. This improves the coding efficiency.
[0200] Further, in some embodiments of the present application, the step of determining the coding mode corresponding to the current frame based on the number of non-uniform sound sources in step E1 includes: if the number of non-uniform sound sources satisfies a second preset condition, determining that the coding mode is the first coding mode; or If the number of non-uniform sound sources does not satisfy the second preset condition, the coding mode is determined to be the second coding mode.
[0201] The first encoding mode is a virtual speaker selection-based HOA encoding mode or a directional voice coding-based HOA encoding mode, and the second encoding mode is a virtual speaker selection-based HOA encoding mode or a directional voice coding-based HOA encoding mode, and the first encoding mode and the second encoding mode are different encoding modes. The virtual speaker selection-based HOA encoding mode is sometimes called a match projection (MP)-based HOA encoding mode.
[0202] Specifically, the encoding mode may be classified into two types, a first encoding mode and a second encoding mode, based on the difference in the number of non-uniform excitation sources. The encoder side acquires a second preset condition. That is, determining whether the number of non-uniform excitation sources satisfies the second preset condition. If the number of non-uniform excitation sources satisfies the second preset condition, determining that the encoding mode is the first encoding mode. If the number of non-uniform excitation sources does not satisfy the second preset condition, determining that the encoding mode is the second encoding mode. In this embodiment of the present application, the division of the encoding mode of the current frame is implemented, and it may be determined whether the number of non-uniform excitation sources satisfies the second preset condition, so as to accurately identify that the encoding mode of the current frame belongs to the first encoding mode or the second encoding mode.
[0203] For example, when the first encoding mode is a HOA encoding mode based on virtual speaker selection, the second encoding mode is a HOA encoding mode based on directional voice coding. Alternatively, when the first encoding mode is a HOA encoding mode based on directional voice coding, the second encoding mode is a HOA encoding mode based on virtual speaker selection, and the specific implementation of the first encoding mode and the second encoding mode may be determined based on an application scenario.
[0204] For example, in this embodiment of the present application, the sound field classification result is used to determine the encoding mode selected by the encoder side. For example, the sound field classification result may be used to determine the encoding mode of the HOA signal. For example, the encoding mode is determined based on the sound field type. The HOA signal belonging to a non-uniform sound source is suitable for encoding by using an encoder corresponding to encoding mode A, and the HOA signal belonging to a distributed sound field is suitable for encoding by using an encoder corresponding to encoding mode B. As another example, the encoding mode is determined based on the number of non-uniform sound sources. If the number of non-uniform sound sources meets the judgment condition for using encoding mode X, the encoding is performed by using an encoder corresponding to encoding mode X. As another example, the encoding mode is alternatively selectively determined based on the sound field type and the number of non-uniform sound sources. If the sound field type is a distributed sound field, the encoding is performed by using an encoder corresponding to encoding mode C. If the sound field type is a non-uniform sound field and the number of non-uniform sound sources meets the judgment condition for using encoding mode X, the encoding is performed by using an encoder corresponding to encoding mode X. The coding mode A, the coding mode B, the coding mode C, and the coding mode X may include a plurality of different coding modes. In this embodiment of the present application, different sound field classification results correspond to different coding modes. This is not limited in this embodiment of the present application. For example, the coding mode X may be coding mode 1 when the number of non-uniform sound sources is less than a preset threshold, or may be coding mode 2 when the number of non-uniform sound sources is equal to or greater than a preset threshold.
[0205] In some embodiments of the present application, the second preset condition includes that the number of non-uniform sound sources is greater than a first threshold or less than a second threshold, and the second threshold is greater than the first threshold; or The second preset condition includes that the number of non-uniform sound sources is equal to or less than a first threshold or equal to or greater than a second threshold, and that the second threshold exceeds the first threshold.
[0206] The specific values of the first threshold and the second threshold are not limited and can be specifically determined based on the application scenario. The second threshold exceeds the first threshold. Therefore, the first threshold and the second threshold constitute a preset range, and the second preset condition may be that the number of non-uniform sound sources falls within the preset range, or the second preset condition may be that the number of non-uniform sound sources exceeds the preset range. In order to determine whether the number of non-uniform sound sources meets the second preset condition and accurately identify that the sound field type of the current frame belongs to the first sound field type or the second sound field type, the number of non-uniform sound sources can be determined based on the second threshold and the second threshold in the first preset condition.
[0207] For example, the first threshold is 0, the second threshold is 3, and the number of non-uniform sound sources is expressed as n. In this case, the second preset condition may be 0 < n < 3, or the second preset condition may be n ≥ 3 or n = 0.
[0208] It should be noted that in the present embodiment of the present application, the first preset condition is a set of conditions for identifying different sound field types, and the second preset condition is a set of conditions for identifying different coding modes. The first preset condition and the second preset condition may include the same condition content, or may include different condition content. In other words, the first preset condition and the second preset condition may be different preset conditions, or may be the same preset condition. However, differences may occur during actual use. The first preset condition and the second preset condition are distinguished by using the numerals first and second.
[0209] In some embodiments of the present application, the step of determining the coding mode corresponding to the current frame based on the sound field type in step E2 includes the following. That is, If the sound field type is a non-uniform sound field, determining that the coding mode is a HOA coding mode based on the virtual speaker selection; or If the sound field type is a distributed sound field, determining that the coding mode is a HOA coding mode based on directional audio coding.
[0210] For sound fields with few non-uniform sound sources in the sound field and distributed sound fields, the HOA coding mode based on directional sound has lower compression efficiency than the HOA coding mode based on virtual speaker selection. However, for sound fields with multiple non-uniform sound sources in the sound field, the HOA coding mode based on virtual speaker selection has lower compression efficiency than the HOA coding mode based on directional sound. In this embodiment of the present application, if the sound field type is a non-uniform sound source, the coding mode is determined to be the HOA coding mode based on virtual speaker selection. If the sound field type is a distributed sound field, the coding mode is determined to be the HOA coding mode based on directional sound coding. In this embodiment of the present application, in order to meet the requirement of obtaining the maximum compression efficiency for different types of sound signals, the corresponding coding mode can be selected based on the sound field classification result of the current frame.
[0211] In some embodiments of the present application, determining an encoding mode corresponding to a current frame based on the sound field classification result in step 503 includes: F1: determining an initial encoding mode corresponding to the current frame based on the sound field classification result of the current frame. F2: Obtaining a hangover time frame in which a current frame is located, the hangover time frame including an initial coding mode of the current frame and coding modes of N-1 frames before the current frame, where N is a length of the hangover; and F3: Determining an encoding mode for the current frame based on the initial encoding mode for the current frame and the encoding modes for the N-1 frames.
[0212] In step F1, the initial encoding mode may be an encoding mode determined based on the sound field classification result. For example, the encoding mode of the current frame may be determined based on any one of the above implementations in steps E1 to E3, and the encoding mode may be used as the initial encoding mode in F1. After the initial encoding mod is obtained, a hangover time window is obtained based on the current frame and a window size of the hangover time window. The hangover time window includes the initial encoding mode of the current frame and the encoding modes of N-1 frames before the current frame, where N represents the number of frames included in the hangover time window. Finally, the encoding mode of the current frame is determined based on the encoding modes corresponding to the N frames in the hangover time window individually. The encoding mode of the current frame obtained in step F3 may be the encoding mode used when encoding the current frame. In this embodiment of the present application, the initial encoding mode of the current frame is modified based on the hangover time window to obtain the encoding mode of the current frame. This makes the encoding modes of successive frames not frequently switched, improving the encoding efficiency.
[0213] For example, after the initial encoding mode of the current frame is obtained, a hangover window processing may be performed on the current frame to ensure that the encoding modes of successive frames are not frequently switched. There are multiple ways to process the hangover window, which is not limited in this embodiment of the application. For example, the processing method may be: saving an encoder selection identifier with a length of N frames in the hangover window, where the N frames include the encoder selection identifiers of the current frame and N-1 frames before the current frame, and updating the encoding type indication identifier of the current frame when the encoder selection identifiers are accumulated to a specified threshold. Optionally, in addition to the hangover window processing, other post-processing may be used to perform corrections on the current frame. For example, the initial encoding mode is used as an initial classification, and the initial classification is modified based on features, such as the speech classification result and the signal-to-noise ratio of the voice signal, and the modified result is used as the final result of the encoding mode.
[0214] As shown in FIG. 6 , the three-dimensional audio signal processing method according to an embodiment of the present application mainly includes the following steps:
[0215] 601: Perform linear decomposition on a current frame of a 3D audio signal to obtain a linear decomposition result.
[0216] 602: Obtain sound field classification parameters corresponding to the current frame based on the linear decomposition result.
[0217] 603: Determine a sound field classification result for the current frame based on the sound field classification parameters.
[0218] The implementation of steps 601 to 603 is similar to that of steps 401 to 403 in the previous embodiment, and steps 601 to 603 will not be described again in detail in this specification.
[0219] 604: Determine coding parameters corresponding to the current frame based on the sound field classification result.
[0220] The encoder side may perform steps 601 to 603. After obtaining the sound field classification result of the current frame, the encoder side may determine an encoding parameter corresponding to the current frame based on the sound field classification result. The encoding parameter is a parameter used when encoding the current frame of the three-dimensional audio signal. There are multiple encoding parameters, and different encoding parameters may be used based on different sound field classification results of the current frame. In this embodiment of the present application, for different sound field classification results of the current frame, an appropriate encoding parameter is selected, and the current frame is then encoded based on the encoding parameter. This improves the compression efficiency and hearing quality of the audio signal.
[0221] Furthermore, in some embodiments of the present application, the coding parameters include at least one of the following: the number of channels of the virtual speaker signals, the number of channels of the residual signal, the number of coding bits of the virtual speaker signals, the number of coding bits of the residual signal, or the number of votes to search for the best matching speaker.
[0222] The virtual speaker signals and the residual signals are signals that are generated based on the three-dimensional audio signal.
[0223] Specifically, the encoder side may determine coding parameters for the current frame based on the sound field classification result of the current frame, and the coding parameters may be used to code the current frame. There are multiple implementations of the coding parameters. For example, the coding parameters include at least one of the following: the number of channels of the virtual speaker signal, the number of channels of the residual signal, the number of coding bits of the virtual speaker signal, the number of coding bits of the residual signal, or the number of votes for searching for the best-matching speaker. The number of channels is also called the number of transmission channels. The number of channels is the number of transmission channels allocated when encoding a signal, and the number of coding bits is the number of coding bits allocated when encoding a signal.
[0224] In the method for selecting a virtual speaker provided in this embodiment of the present application, the encoder votes for each virtual speaker in the candidate virtual speaker set based on the virtual speaker coefficient of the current frame, and selects a virtual speaker for the current frame based on the voting value, so as to reduce the computation load for searching for a virtual speaker and reduce the computation load of the encoder. The number of votes for searching for the best-matching speaker is the number of votes required in searching for the best-matching speaker. In a possible implementation, the number of votes may be pre-configured or determined based on the sound field classification result of the current frame. For example, the number of votes for searching for the best-matching speaker is the number of votes for searching for a virtual speaker in the process of determining a virtual speaker signal based on a three-dimensional sound signal.
[0225] In addition, the virtual speaker signal and the residual signal in this embodiment of the present application are signals generated based on a three-dimensional sound signal. For example, a first target virtual speaker is selected from a preset virtual speaker set based on a first scene sound signal, and the virtual speaker signal is generated based on the first scene sound signal and attribute information of the first target virtual speaker. A second scene sound signal is obtained based on the attribute information of the first target virtual speaker and the first virtual speaker signal, and a residual signal is generated based on the first scene sound signal and the second scene sound signal.
[0226] In some embodiments of the present application, the number of votes satisfies the following relationship: 1≦I≦d
[0227] I is the number of votes, and d is the number of inhomogeneous sound sources included in the sound field classification result.
[0228] The encoder determines the number of votes to search for the best matching speaker based on the number of non-uniform sound sources of the current frame. The number of votes is equal to or less than the number of non-uniform sound sources of the current frame, so that the number of votes can be adapted to the actual situation of the sound field classification of the current frame. This solves the problem that the number of votes to search for the best matching speaker needs to be determined when the current frame is encoded.
[0229] For example, the number of votes I must follow the following rules: the minimum number of votes is 1, the maximum number of votes does not exceed the total number of speakers, and the maximum number of votes does not exceed the number of channels of the virtual speaker signal. For example, the total number of speakers is 1024 speakers obtained by the virtual speaker set generation unit in the encoder, and the number of channels of the virtual speaker signal is the number of virtual speaker signals transmitted by the encoder, i.e., the number of N transmission channels correspondingly generated by the N best-matching speakers. Usually, the number of channels of the virtual speaker signal is less than the total number of speakers. The method for estimating the number of votes is as follows: In the sound field of the current frame, a step of determining a number of votes I for searching for a best-matching speaker based on the number of non-uniform sound sources obtained in the sound field classification result. The number of votes I satisfies the relationship of 1≦I≦d. d is the number of sound sources in different directions included in the sound field, i.e., the number of non-uniform sound sources estimated in the sound field classification result. For example, I=d. Or, the number of votes I=min(d, total number of speakers, number of channels of the virtual speaker signal, preset number of votes). The number of votes I can be obtained based on min(d, total number of speakers, number of channels of virtual speaker signals, preset number of voting rounds), so that the encoder side can determine the number of votes to search for the best-matching speaker based on the value of I.
[0230] In some embodiments of the present application, the sound field classification results include the number of non-uniform sound sources and the sound field type.
[0231] When the sound field type is a non-uniform sound source, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the encoder. Or, When the sound field type is a distributed sound field, the number of channels of the virtual speaker signals satisfies the following relationship: F=1 where F is the number of channels of the virtual speaker signal.
[0232] The number of channels of the virtual speaker signal is the number of channels for transmitting the virtual speaker signal, and the number of channels of the virtual speaker signal can be determined based on the number of non-uniform sound sources and the type of sound field. In the above calculation method, when the sound field type is a distributed sound field, the number of channels of the virtual speaker signal is determined to be 1 in order to improve the coding efficiency of the current frame. When the sound field type is a non-uniform sound source, min represents an operation of selecting a minimum value, that is, an operation of selecting a minimum value from S and PF as the number of channels of the virtual speaker signal, so that the channels of the virtual speaker signal can be adapted to the actual situation of the sound field classification of the current frame. This solves the problem that the number of channels of the virtual speaker signal needs to be determined when coding the current frame.
[0233] In some embodiments of the present application, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1,PR) Here, PR is the number of channels of the residual signal preset by the encoder, and C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder. Or, When the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signals preset by the encoder, and F is the number of channels of the virtual speaker signals.
[0234] After the number of channels of the virtual speaker signal is obtained, the number of channels of the residual signal can be calculated based on the preset number of channels of the residual signal and the sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal. The value of PR can be preset on the encoder side, and the value of R can be obtained according to the calculation formula max(C-1,PR). The sum of the preset number of channels of the residual signal and the preset number of channels of the virtual speaker signal is preset on the encoder side. Note that C is also called the total number of transmission channels.
[0235] In some embodiments of the present application, after the number of channels of the virtual speaker signal is obtained, the number of channels of the residual signal can be calculated based on the number of channels of the virtual speaker signal and the sum of the number of preset channels of the residual signal and the number of preset channels of the virtual speaker signal. The sum of the number of preset channels of the residual signal and the number of preset channels of the virtual speaker signal is preset on the encoder side. Note that C is also referred to as the total number of transmission channels.
[0236] In some embodiments of the present application, the sound field classification result includes a number of non-uniform sound sources.
[0237] The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal that is preset by the encoder.
[0238] The number of channels of the virtual speaker signal is the number of channels for transmitting the virtual speaker signal, and the number of channels of the virtual speaker signal can be determined based on the number of non-uniform sound sources. In the above calculation method, min represents the operation of selecting the minimum value, that is, the operation of selecting the minimum value from S and PF as the number of channels of the virtual speaker signal, so that the number of channels of the virtual speaker signal can be adapted to the actual situation of the sound field classification of the current frame. This solves the problem that the number of channels of the virtual speaker signal needs to be determined when encoding the current frame.
[0239] In some embodiments of the present application, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal. For example, C is the sum of PF and PR.
[0240] After the number of channels of the virtual speaker signal is obtained, the number of channels of the residual signal can be calculated based on the number of channels of the virtual speaker signal and the sum of the number of preset channels of the residual signal and the number of preset channels of the virtual speaker signal. The sum of the number of preset channels of the residual signal and the number of preset channels of the virtual speaker is preset on the encoder side. Note that C is also called the total number of transmission channels.
[0241] In some embodiments of the present application, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type.
[0242] The number of coding bits of the virtual speaker signal is obtained based on a ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel.
[0243] The number of coding bits of the residual signal is obtained based on a ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels.
[0244] The number of coding bits of the transmission channels includes the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal, and when the number of non-uniform sound sources is equal to or less than the number of channels of the virtual speaker signals, the ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels is obtained by increasing an initial ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels.
[0245] The encoder side pre-sets an initial ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel, obtains the number of non-uniform sound sources, and determines whether the number of non-uniform sound sources is equal to or less than the number of channels of the virtual speaker signal. If the number of non-uniform sound sources is equal to or less than the number of channels of the virtual speaker signal, the initial ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel may be increased, and the increased initial ratio is defined as the ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel. The ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel may be used to calculate the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal. In the above calculation method, the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal can be adapted to the actual situation of the sound field classification of the current frame. This solves the problem that the number of coding bits of the virtual speaker signal and the number of coding bits of the residual signal need to be determined when encoding the current frame.
[0246] For example, the encoder side determines a bit allocation method for the virtual speaker signals and the residual signals based on the sound field classification result, divides the transmission channel signals into virtual speaker signal groups and residual signal groups, and uses a preset allocation ratio of the virtual speaker signal groups as an initial ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels. If the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signals, the initial ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels is increased based on a preset adjustment value, and the increased ratio is used as the ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels. For example, the increased ratio is equal to the sum of the preset adjustment value and the initial ratio.
[0247] In some embodiments of the present application, the ratio of the number of coding bits of the residual signal to the number of coding bits of the transmission channel=1.0-the ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel.
[0248] In some embodiments of the present application, in addition to performing the steps mentioned above, the method performed by the encoder side may further include: encoding the current frame and the sound field classification result, and writing the encoded current frame and the sound field classification result into a bitstream.
[0249] The sound field classification result may be encoded into a bitstream. After the encoder side sends the bitstream to the decoder side, the decoder side may obtain the sound field classification result based on the bitstream. The decoder side obtains the sound field classification result carried in the bitstream by analyzing the bitstream, and obtains the sound field distribution state of the current frame based on the sound field classification result, so that the current frame can be decoded to obtain a three-dimensional audio signal.
[0250] In some embodiments of the present application, the step of encoding the current frame and the sound field classification result specifically includes: directly encoding the current frame or first processing the current frame; and after obtaining the virtual speaker signals and the residual signals, encoding the virtual speaker signals and the residual signals. For example, the encoder side may specifically be a core encoder. The core encoder encodes the virtual speaker signals, the residual signals, and the sound field classification result to obtain a bitstream. The bitstream may also be called a speech signal encoding bitstream.
[0251] The three-dimensional audio signal processing method provided in this embodiment of the present application may include an audio encoding method and an audio decoding method. The audio encoding method is performed by an audio encoding device, and the audio decoding method is performed by an audio decoding device, and the audio encoding device may communicate with the audio decoding device. Figures 4 to 6 are performed by an audio encoding device. The following describes a three-dimensional audio signal processing method performed by an audio decoding device (referred to as a decoder side) according to one embodiment of the present technology. As shown in Figure 7, the method may mainly include the following steps:
[0252] 701: Receive a bitstream.
[0253] The decoder side receives a bitstream from the encoder side, which includes the sound field classification result.
[0254] 702: Decode the bitstream to obtain the sound field classification result of the current frame.
[0255] The decoder side parses the bitstream and obtains the sound field classification result of the current frame from the bitstream, which is obtained by the encoder side according to the embodiments shown in Figures 4 to 6.
[0256] 703: Obtain a 3D audio signal of the decoded current frame based on the sound field classification result.
[0257] After obtaining the sound field classification result, the decoder side analyzes the bitstream based on the sound field classification result to obtain the three-dimensional audio signal of the decoded current frame. In the embodiment of the present application, the decoding process of the current frame is not limited. In the present embodiment of the present application, the decoder side may decode the current frame based on the sound field classification result. The result of the sound field classification can be used to decode the current frame in the bitstream. Therefore, the decoder side performs decoding in a decoding manner adapted to the sound field of the current frame to obtain the three-dimensional audio signal sent by the encoder side. This implements the transmission of the audio signal from the encoder side to the decoder side.
[0258] For example, the decoder side can determine a decoding mode and / or a decoding parameter that matches the encoding mode and / or the encoding parameter of the encoder side based on the sound field classification result transmitted in the bitstream, which reduces the number of coding bits compared to the method in which the encoder side transmits the encoding mode and / or the encoding parameter to the decoder side.
[0259] In some embodiments of the present application, the step of obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result in step 703 includes: G1: determining a decoding mode for the current frame based on the sound field classification result; and G2: Obtaining a three-dimensional audio signal of the decoded current frame based on a decoding mode.
[0260] The decoding mode corresponds to the encoding mode in the above embodiment. The implementation of step G1 is similar to step 504 in the above embodiment. The details are not described again in this specification. After obtaining the decoding mode, the decoder side can decode the bitstream according to the decoding mode to obtain the decoded 3D audio signal of the current frame.
[0261] Furthermore, in some embodiments of the present application, the step of determining the decoding mode of the current frame based on the sound field classification result in step G1 includes: determining a decoding mode for the current frame based on the number of non-uniform sound sources when the sound field classification result includes the number of non-uniform sound sources, or when the sound field classification result includes the number of non-uniform sound sources and the sound field type; determining a decoding mode for the current frame based on the sound field type if the sound field classification result includes a sound field type, or if the sound field classification result includes a non-uniform sound source number and a sound field type; or determining a decoding mode for the current frame based on the number of non-uniform sound sources and the type of sound field when the sound field classification result includes the number of non-uniform sound sources and the type of sound field;
[0262] The implementation of the above steps is similar to that of steps E1 to E3 in the above embodiments, and the details will not be described again in this specification.
[0263] In some embodiments of the present application, the step of determining a decoding mode for the current frame based on the number of non-uniform sound sources includes: determining that the decoding mode is the first decoding mode if the number of non-uniform sound sources satisfies the preset condition; or If the number of non-uniform sound sources does not satisfy the preset condition, determining that the decoding mode is the second decoding mode.
[0264] The first decoding mode is an HOA decoding mode based on virtual speaker selection or an HOA decoding mode based on directional voice coding, and the second decoding mode is an HOA decoding mode based on virtual speaker selection or an HOA decoding mode based on directional voice coding, and the first decoding mode and the second decoding mode are different decoding modes.
[0265] It should be noted that the preset conditions are conditions set by the decoder side to distinguish different decoding modes, and the implementation of the preset conditions is not limited.
[0266] In some embodiments of the present application, the preset conditions include that the number of non-uniform sound sources is greater than a first threshold or less than a second threshold, and the second threshold is greater than the first threshold; or The preset conditions include that the number of non-uniform sound sources is equal to or less than a first threshold or equal to or greater than a second threshold, and that the second threshold exceeds the first threshold.
[0267] In some embodiments of the present application, obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result in step 703 includes: H1: determining a decoding parameter of a current frame based on the sound field classification result; and H2: Obtaining a three-dimensional audio signal of the decoded current frame based on the decoding parameters.
[0268] The decoding parameters correspond to the encoding parameters in the above embodiment. The implementation of step H1 is similar to step 604 in the above embodiment. The details are not described again in this specification. After obtaining the decoding parameters, the decoder side can decode the bitstream according to the decoding parameters to obtain the decoded 3D audio signal of the current frame.
[0269] In some embodiments of the present application, the decoding parameters include at least one of the following: the number of channels of the virtual speaker signals, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signals, the number of encoding bits of the virtual speaker signals, or the number of decoded bits of the residual signal.
[0270] The virtual speaker signals and the residual signal are obtained by decoding the bitstream.
[0271] In some embodiments of the present application, the sound field classification results include the number of non-uniform sound sources and the sound field type.
[0272] When the sound field type is a non-uniform sound source, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) where F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the decoder. Or, When the sound field type is a distributed sound field, the number of channels of the virtual speaker signals satisfies the following relationship: F=1 where F is the number of channels in the virtual speaker signal.
[0273] In some embodiments of the present application, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1,PR) Here, PR is the number of channels of the residual signal preset by the decoder, and C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder. Or When the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the following relationship: F=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0274] It should be noted that the number of channels of the virtual speaker signals preset by the decoder is equal to the number of channels of the virtual speaker signals preset by the encoder, and similarly, the number of channels of the residual signal preset by the decoder is equal to the number of channels of the residual signal preset by the encoder.
[0275] In some embodiments of the present application, the sound field classification result includes a number of non-uniform sound sources.
[0276] The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal that is preset by the decoder.
[0277] In some embodiments of the present application, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0278] It should be noted that the implementation of the decoding parameters is similar to that of the encoding parameters in the previous embodiment, and the details will not be described again here.
[0279] In some embodiments of the present application, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type.
[0280] The number of decoding bits of the virtual speaker signal is obtained based on a ratio of the number of decoding bits of the virtual speaker signal to the number of decoding bits of the transmission channel.
[0281] The number of decoding bits of the residual signal is obtained based on a ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels.
[0282] The number of decoding bits of the transmission channels includes the number of decoding bits of the virtual speaker signals and the number of decoding bits of the residual signal, and when the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signals, the ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels is obtained by increasing an initial ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels.
[0283] In order to better understand and implement the aforementioned solutions in the embodiments of the present application, a specific description is provided below by using corresponding application scenarios as examples.
[0284] In this embodiment of the application, an example is used in which the three-dimensional audio signal is an HOA signal. The sound field classification method for HOA signals in this embodiment of the application is applied to a hybrid HOA encoder. FIG. 8 shows a basic encoding procedure. The encoder side performs classification on the HOA signal to be encoded to determine whether the HOA signal to be encoded of the current frame is suitable for the HOA encoding scheme based on virtual speaker selection or the HOA encoding scheme based on directional audio coding DirAC, and determines the HOA encoding mode of the current frame based on the sound field classification result. Specifically, the HOA encoder includes an encoder selection unit. The encoder selection unit performs sound field classification on the HOA signal to be encoded to determine the encoding mode of the current frame. Then, selects encoder A or encoder B for encoding based on the encoding mode to obtain a final encoded bitstream. Encoder A and encoder B refer to different types of encoders, and each type of encoder is adapted to the sound field type of the current frame. When an encoder adapted to the sound field type is used for encoding, the compression rate of the signal can be improved.
[0285] The specific process of performing sound field classification on the HOA signal to be encoded and determining the encoding mode includes the following: performing sound field classification on the HOA signal to be encoded to obtain a sound field classification result; and Determining an encoding mode corresponding to the current frame based on the sound field classification result.
[0286] The coding mode of the current frame indicates the selection method of the encoder of the current frame. The criterion for determining the encoder selection identifier may be determined based on the sound field type of the HOA signal to which the encoder A and the encoder B are applicable. For example, the signal type processed by the encoder A is an HOA signal having a non-uniform sound field and the number of non-uniform sound sources is less than three, and the signal type processed by the encoder B is an HOA signal having a non-uniform sound field and the number of non-uniform sound sources is three or more. Alternatively, the signal type processed by the encoder B is an HOA signal having a distributed sound field, or an HOA signal having the number of non-uniform sound sources is three or more.
[0287] It should be noted that hangover time window processing may be performed on the sound field classification result to ensure that the encoding mode between successive frames is not frequently switched. There are multiple hangover time window processing methods, which are not limited in this embodiment of the present application. For example, the processing method may be: save an encoder selection identifier that is N frames long in a hangover time window, where the N frames include the encoder selection identifiers of the current frame and N-1 frames before the current frame, and update the encoding type indication identifier of the current frame when the encoder selection identifiers are accumulated to a specified threshold. Optionally, in addition to the hangover time window processing, other processing may be used to perform correction of the sound field classification result.
[0288] As shown in FIG. 9, the procedure of determining the coding mode of the HOA signal mainly includes:
[0289] S01: Acquire the HOA signal to be branched.
[0290] S02: Downsampling is performed on the HOA signal.
[0291] It is not limited that the step of performing downsampling on the HOA signal to be analyzed is an optional step.
[0292] In order to reduce the computational complexity, downsampling is performed on the analyzed HOA signal. The analyzed HOA signal may be a time domain HOA signal or a frequency domain HOA signal. The analyzed HOA signal may include all channels or may include some HOA channels (such as FOA channels). For example, the analyzed HOA signal may be full sampling points or 1 / Q downsampling points. For example, in this embodiment, a 1 / 120 downsampling point is used.
[0293] For example, the order of the HOA signal in the current frame is 3, the number of channels of the HOA signal is 16, and the frame length of the current frame is 20 milliseconds (ms), that is, the signal of the current frame contains 960 sampling points. After the HOA signal to be coded in the current frame is processed by 1 / 120 downsampling, each channel of the signal contains 8 sampling points. In other words, the HOA signal has 16 channels, each channel has 8 sampling points, which constitute the input signal of the sound field type analysis, i.e., the HOA signal to be analyzed.
[0294] S03: Perform sound field type analysis based on the signal obtained through downsampling.
[0295] After downsampling is performed on the HOA signal, the sound field type is obtained by analyzing the number of non-uniform sound sources in the HOA signal.
[0296] For example, the sound field type analysis in this embodiment of the present application may be a step of performing linear decomposition on the HOA signal, a step of obtaining a linear decomposition result via the linear decomposition, and then a step of obtaining a sound field classification result based on the linear decomposition result.
[0297] For example, the number of non-uniform sound sources can be obtained based on a linear decomposition result. For example, the linear decomposition result may include a feature value. The number of non-uniform sound sources is estimated based on the ratio between the feature values, specifically including: Performing a singular value decomposition on the HOA signal to be analyzed to obtain singular values v[i], where i=0,1,...,min(L,K)-1.
[0298] L is equal to the number of channels of the HOA signal, and K is the number of signal points of each channel in the current frame. For example, the number of signal points can be the number of frequencies. In this embodiment, L=16, K=8, and min(L,K)=8.
[0299] The ratio temp[i] between the singular values v is calculated and used as the sound field classification parameter, i.e. temp[i]=v[i] / v[i+1] Here, let i = 0, 1, ..., min(L,K)-2.
[0300] The non-uniform sound source decision threshold is 100, and the number n of non-uniform sound sources can be estimated in the following manner. The step of determining whether temp[i] exceeds 100 from i = 0. And, when temp[i] is 100 or more and temp[i] ≥ 100 is satisfied, the step of stopping the determination step. Otherwise, as i = i + 1, the step of continuing the execution of the determination step. When stopping the determination step, the non-uniform sound source number n becomes equal to the sequence number i at which the determination step is stopped plus 1. For example, when i = 0 and temp[0] ≥ 100, the determination step is stopped and the non-uniform sound source number n becomes equal to 1. Otherwise, i is set to 1, and when 1 = 1, the determination step is continuously executed. When i = 1 and temp[1] ≥ 100, the determination step is stopped and the non-uniform sound source number n becomes equal to i + 1 = 2.
[0301] S04: Based on the analysis result of the sound field type, determine the predictive coding mode.
[0302] The predictive coding mode is determined based on the non-uniform sound source number n.
[0303] When 0 < n < 3, the predictive coding mode becomes coding mode 1.
[0304] When n ≥ 3 or n = 0, the predictive coding mode becomes coding mode 2.
[0305] For example, coding mode 1 can be an HOA coding mode based on virtual speaker selection. Coding mode 2 can be an HOA coding method based on directional audio coding DirAC.
[0306] S05: Based on the predictive coding mode, determine the actual coding mode.
[0307] After the predictive coding mode of the current frame is determined, then the actual coding mode is determined. For example, the hangover time frame is used to determine the actual coding mode. In the hangover time frame, when the predictive coding mode 2 of multiple frames in the hangover time frame is accumulated to a specified threshold, the actual coding mode of the current frame is coding mode 2. Otherwise, the actual coding mode of the current frame is coding mode 1.
[0308] For example, there are predictive coding mode results for 10 frames in the hangover time frame, including the coding mode determination result of the current frame in step S03 and the coding mode result of the frame 9 frames before the current frame. When up to 7 frames of the predictive coding mode results of the 10 frames whose coding mode is coding mode 2 are accumulated, the actual coding mode of the current frame is determined to be coding mode 2.
[0309] S06: The final encoding mode is obtained.
[0310] The basic decoding procedure of the hybrid HOA decoder corresponding to the encoder side is shown in Figure 10. The decoder side obtains a bitstream from the encoder side, and then analyzes the bitstream to obtain the HOA decoding mode of the current frame. A corresponding decoding scheme is selected for decoding based on the HOA decoding mode of the current frame to obtain a reconstructed HOA signal. Specifically, the decoder side includes a decoder selection unit. The decoder selection unit analyzes the bitstream, determines the decoding mode, and selects decoder A or decoder B for decoding based on the decoding mode to obtain a reconstructed HOA signal. The decoder A and decoder B refer to different types of decoders, and each type of decoder is adapted to the sound field type of the current frame. When the decoder adapted to the sound field type is used for decoding, the HOA signal can be correctly reconstructed.
[0311] From the above description, it can be seen that sound field classification is performed on the HOA signal to be encoded, and the encoding mode is determined based on the result of the sound field classification, so that different encoding modes are used for appropriate signal types to obtain maximum compression efficiency for various types of signals.
[0312] In the following, a HOA encoder based on virtual speaker selection according to one embodiment of the present application is described. Figure 11 shows the basic encoding procedure.
[0313] The encoder side may include the following: a virtual speaker configuration unit, an encoding analysis unit, a virtual speaker set generation unit, a virtual speaker selection unit, a virtual speaker signal generation unit, a core encoder processing unit, a signal reconstruction unit, a residual signal generation unit, a selection unit, and a signal compensation unit. The following describes the functions of the units included in the encoder side individually. In this embodiment of the present application, the encoder side shown in FIG. 11 may generate one virtual speaker signal or multiple virtual speaker signals. The procedure of generating multiple virtual speaker signals may perform the generation based on the configuration of the encoder shown in FIG. 11 multiple times. The following uses the procedure of generating one virtual speaker signal as an example.
[0314] The virtual speaker configuration unit is configured to configure the virtual speakers in the virtual speaker set to obtain a plurality of virtual speakers.
[0315] The virtual speaker configuration unit outputs virtual speaker configuration parameters based on the encoder configuration information. The encoder configuration information includes, but is not limited to, HOA order, coding bit rate, and user-defined information, etc. The virtual speaker configuration parameters include, but are not limited to, the number of virtual speakers, the HOA order of the virtual speakers, and the position coordinates of the virtual speakers, etc.
[0316] The virtual speaker configuration parameters output by the virtual speaker configuration unit are used as input of a virtual speaker set generation unit.
[0317] The coding analysis unit is configured to perform coding analysis on the HOA signal to be coded, for example, to analyze the sound field distribution including features in the HOA to be coded, such as the number of sound sources, directivity, and dispersion degree of the signal to be coded, which are used as one of the decision conditions for determining how to select the target virtual speaker.
[0318] In this embodiment of the present application, it is not limited that the encoder side may alternatively not include a coding analysis unit. In other words, the encoder side may not analyze the input signal, but may use a default configuration to determine how to select a target virtual speaker.
[0319] The encoder side obtains the HOA signal to be encoded. For example, the encoder side may use the HOA signal recorded from the actual collecting device or the HOA signal synthesized by using an artificial voice object as the input of the encoder. In addition, the HOA signal to be encoded input by the encoder may be the HOA signal in the time domain or the HOA signal in the frequency domain.
[0320] The virtual speaker set generation unit is configured to generate a virtual speaker set. The virtual speaker set may include a plurality of virtual speakers, and the virtual speakers in the virtual speaker set may be referred to as "candidate virtual speakers".
[0321] The virtual speaker set generation unit generates HOA coefficients of the designated candidate virtual speakers based on the virtual speaker configuration parameters. To generate the HOA coefficients of the candidate virtual speakers, the coordinates (i.e., position coordinates or position information) of the candidate virtual speakers and the HOA orders of the candidate virtual speakers are required. The method of determining the coordinates of the candidate virtual speakers includes, but is not limited to, generating K virtual speakers according to the equidistance principle and generating K candidate virtual speakers that are unevenly distributed according to the hearing principle. An example of generating a certain number of evenly distributed virtual speakers is described below.
[0322] The coordinates of the evenly distributed candidate virtual speakers are generated based on the number of candidate virtual speakers, for example, an approximately even arrangement of the virtual speakers is obtained by using a numerical iteration method.
[0323] The HOA coefficients of the candidate virtual speakers output by the virtual speaker set generation unit are used as inputs of the virtual speaker selection unit.
[0324] The virtual speaker selection unit is configured to select a target virtual speaker from a plurality of candidate virtual speakers in the virtual speaker set based on the HOA signal to be encoded, where the target virtual speaker may be referred to as a "virtual speaker that matches the HOA signal to be encoded" or a matching virtual speaker.
[0325] The virtual speaker selection unit matches the HOA signal to be encoded with the HOA coefficients of the candidate virtual speakers output by the virtual speaker set generation unit, and selects a designated matching virtual speaker.
[0326] In this embodiment of the present invention, to obtain a sound field classification result, a sound field classification is performed on the HOA signal to be encoded, and encoding parameters are determined based on the sound field classification result.
[0327] The coding analysis unit is configured to perform coding analysis according to the HOA signal to be coded, and the analysis includes: performing sound field classification according to the HOA signal to be coded. For the sound field classification method, please refer to the above embodiment. The details will not be described again in this specification.
[0328] The coding parameters are determined based on the sound field classification result, and may include at least one of the number of channels of the virtual speaker signals, the number of channels of the residual signal, or the number of votes for searching for the best matching speaker in the HOA coding scheme based on the virtual speaker selection.
[0329] Specifically, the virtual speaker selection unit matches the HOA coefficients of the candidate virtual speakers output by the virtual speaker set generation unit to the HOA coefficients of the candidate virtual speakers based on the number of votes determined to search for the best-matching speaker and the channels of the virtual speaker signal, selects the best-matching virtual speaker, and obtains the HOA coefficients of the best-matching virtual speaker. The number of the best-matching virtual speakers is equal to the number of channels of the virtual speaker signal.
[0330] The virtual speaker selection unit may use a voting-based best-matching speaker search method to match the HOA coefficients to be encoded to the HOA coefficients of candidate virtual speakers output by the virtual speaker set generation unit, select the best-matching virtual speaker, and determine the number of votes I for searching for the best-matching speaker based on the sound field classification result.
[0331] The number of votes I must follow the following rule: the minimum number of votes is 1, and the maximum number of votes does not exceed the total number of speakers (e.g., 1024 speakers obtained by the virtual speaker set generation unit) and the number of channels of the virtual speaker signals (the number of virtual speaker signals transmitted by the encoder, i.e., the N transmission channels correspondingly generated by the N best-matching speakers). Usually, the number of channels of the virtual speaker signals is less than the total number of speakers.
[0332] The method for estimating the number of votes is as follows: A step of determining the number of votes I for selecting a speaker based on the number of non-uniform sound sources in the sound field obtained from the sound field classification result.
[0333] The number of votes I satisfies 1≦I≦d, where d is the number of sound sources in different directions included in the sound field, that is, the number of non-uniform sound sources estimated in the sound field classification result. For example, I=d.
[0334] The number of channels of the virtual speaker signals and the number of channels of the residual signal are determined based on the type of sound field.
[0335] An embodiment of the present application then provides a method for selecting the number of channels F of the adaptive virtual speaker signal.
[0336] When the sound field type is a non-uniform sound field, F=min(S, PF), where S is the number of non-uniform sound sources in the sound field, and PF is the number of channels of the virtual speaker signal that is preset by the encoder.
[0337] If the sound field type is a distributed sound field, F=1.
[0338] Then, an embodiment of the present application provides a method for selecting the number of channels R of the adaptive residual signal.
[0339] When the sound field type is a distributed sound source field, R=max(C-1,PR), where C is the total number of transmission channels preset, and PR is the number of residual signals preset by the encoder. For example, C is the sum of PF and PR.
[0340] When the sound field type is a non-uniform sound source, R=CF.
[0341] The method for determining the bit allocation of the virtual speaker signals and the residual signal based on the sound field classification result is as follows.
[0342] If the number of non-uniform sound sources≦the number of channels of the virtual speaker signals, more bits can be allocated to the channels of the virtual speaker signals since the energy of the residual signal is lower.
[0343] In some embodiments, the virtual speaker signals and the residual signals are divided into two groups, namely a virtual speaker signal group and a residual signal group. If the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signals, the pre-set allocation percentage of the virtual speaker signal group is increased based on the preset adjustment value, and the increased allocation percentage of the virtual speaker signal group is used as the allocation percentage of the virtual speaker signal group.
[0344] The residual signal group allocation percentage = 1.0 - the virtual speaker signal group allocation percentage.
[0345] The virtual speaker signal generation unit calculates a virtual speaker signal based on the HOA coefficients to be coded and the HOA coefficients of the best-matching virtual speaker.
[0346] The signal reconstruction unit reconstructs the HOA signal based on the virtual speaker signal and the HOA coefficients of the best-matching virtual speaker.
[0347] The residual signal generation unit calculates a residual signal based on the number of channels of the residual signal determined in step 1, the HOA coefficients to be coded, and the reconstructed HOA signal output by the HOA signal reconstruction unit.
[0348] Compared with a residual signal having an Nth-order Ambisonic coefficient, if a channel number that is less than the Nth-order Ambisonic coefficient is selected as the residual signal to be transmitted, information loss will occur, so the signal compensation unit needs to perform information compensation on the residual signal that is not transmitted.
[0349] The virtual speaker signals have high amplitude or energy, and the residual signal to be transmitted has low amplitude or energy. Therefore, the selection unit pre-allocates all available bits to the virtual speaker signals and the residual signal to be transmitted. The obtained bit pre-allocation information is used to guide the core encoder for processing.
[0350] The core encoder processing unit performs core encoder processing on the transmission channels, and outputs a transmission bitstream, where the transmission channels include channels of the virtual speaker signals and channels of the residual signal.
[0351] The coding parameters are determined based on the sound field classification result. The coding parameters may further include at least one of a bit allocation of the virtual speaker signals and a bit allocation of the residual signal in a HOA coding scheme based on virtual speaker selection. When the bit allocation of the virtual speaker signals and the bit allocation of the residual signal are determined based on the sound field classification result, it is necessary to determine the bit allocation of the virtual speaker signals and the residual signal based on the sound field classification result.
[0352] In some embodiments, the method for determining bit allocation of the virtual speaker signals and the residual signals based on the sound field classification result is as follows: the number of channels of the virtual speaker signals is F, the number of channels of the virtual speakers is R, the number of channels of the residual signal is R, and the total number of bits that can be used to encode the virtual speaker signals and the residual signals is numbit.
[0353] In one method, the total number of coding bits of the virtual speaker signals and the total number of coding bits of the residual signal are first determined, and then the number of coding bits of each channel is determined. For example, the total number of coding bits of the virtual speaker signals is
[0354]
number
[0355] It is.
[0356] fac1 is a weighting factor assigned to the coding bits of the virtual speaker signal, fac2 is a weighting factor assigned to the coding bits of the residual signal, and round() represents rounding down. For example, fac1>fac2. For example, fac1=2 and fac2=1.
[0357] The total number of coding bits of the residual signal is res_numbit=numbit-core_numbit become.
[0358] Then, the coding bits for each channel of the virtual speaker signal are allocated according to a bit allocation criterion of the virtual speaker signal, and the coding bits for each channel of the residual signal are allocated according to a bit allocation criterion of the residual signal.
[0359] Alternatively, the total number of coding bits of the residual signal is
[0360]
number
[0361] become.
[0362] fac1 is a weighting factor assigned to the coding bits of the virtual speaker signal, fac2 is a weighting factor assigned to the coding bits of the residual signal, and round() represents rounding down. For example, fac1>fac2. For example, fac1=2 and fac2=1.
[0363] Then, the total number of coding bits of the virtual speaker signal is core_numbit=numbit-res_numbit It becomes.
[0364] Then, the coding bits for each channel of the virtual speaker signal are allocated according to a bit allocation criterion of the virtual speaker signal, and the coding bits for each channel of the residual signal are allocated according to a bit allocation criterion of the residual signal.
[0365] Alternatively, the number of coding bits for each channel may be determined directly. For example, the number of coding bits for each virtual speaker signal may be
[0366]
number
[0367] It becomes.
[0368] The number of coding bits for each residual signal is
[0369]
number
[0370] It becomes.
[0371] The bit allocation result finally used for encoding the virtual speaker signals and the residual signal may be determined based on the adjusted bit allocation result obtained by using the above-mentioned method. After obtaining the bit allocation result for encoding the virtual speaker signals and the residual signal, the core encoder processing unit encodes the virtual speaker signals and the residual signal based on the bit allocation result.
[0372] A sound field classification is performed on the HOA signal to be coded, coding parameters are determined based on the sound field classification result, and the signal to be coded is coded based on the determined coding parameters. The coding parameters include at least one of the number of channels of the virtual speaker signals, the number of channels of the residual signal, the bit allocation of the virtual speaker signals, the bit allocation of the residual signal, or the number of votes for searching for the best-matching speaker in the HOA coding scheme based on the virtual speaker selection. For a description of the coding parameters, please refer to the above. Details will not be described again in this specification.
[0373] From the above example, it can be seen that in this embodiment of the present application, sound field classification is performed on the HOA signal to be encoded, so that a suitable encoding mode and / or encoding parameters are selected for encoding the HOA signal based on different features in the HOA signal to be encoded, which improves compression efficiency and hearing quality.
[0374] The decoding procedure performed by the decoder side is not described in detail in the embodiments of this application.
[0375] It should be noted that for ease of description, the above method embodiments are expressed as a series of operations. However, those skilled in the art should understand that the present application is not limited to the order of operations described, since some steps may be performed in other orders or simultaneously according to the present application. In addition, those skilled in the art should further understand that the embodiments described herein are all examples of embodiments, and the operations and modules involved are not necessarily required by the present application.
[0376] In order to better implement the solutions of the embodiments of the present application, related apparatuses for implementing the solutions are further provided below.
[0377] 12 shows a three-dimensional audio signal processing apparatus according to an embodiment of the present application. For example, the three-dimensional audio signal processing apparatus is specifically an audio encoding apparatus 1200, and may include a linear analysis module 1201, a parameter generation module 1202, and a sound field classification module 1203.
[0378] The linear analysis module is configured to perform a linear decomposition on the three-dimensional audio signal to obtain a linear decomposition result.
[0379] The parameter generation module is configured to obtain sound field classification parameters corresponding to the current frame based on the linear decomposition result.
[0380] The sound field classification module is configured to determine a sound field classification result for the current frame based on the sound field classification parameters.
[0381] In some embodiments of the present application, the three-dimensional audio signal comprises a higher order Ambisonics HOA signal or a first order Ambisonics FOA signal.
[0382] In some embodiments of the present application, the linear analysis module is configured to: perform singular value decomposition on the current frame to obtain singular values corresponding to the current frame, where the linear decomposition result includes the singular values; perform principal component analysis on the current frame to obtain first feature values corresponding to the current frame, where the linear decomposition result includes the first feature values; or perform independent component analysis on the current frame to obtain second feature values corresponding to the current frame, where the linear decomposition result includes the second feature values.
[0383] In some embodiments of the present application, there are multiple linear decomposition results and there are multiple sound field classification parameters.
[0384] The parameter generation module is configured to obtain a ratio of the i-th linear analysis result of the current frame to the (i+1)-th linear analysis result of the current frame, where i is a positive integer, and obtain the i-th sound field classification parameter corresponding to the current frame based on the ratio.
[0385] Optionally, the i-th linear analysis result and the (i+1)-th linear analysis result are two consecutive linear analysis results in the current frame.
[0386] In some embodiments of the present application, there are multiple sound field classification parameters, and the sound field classification result includes a sound field type. The sound field classification module is configured to: determine the sound field type as a distributed sound field if all values of the multiple sound field classification parameters satisfy a preset distributed sound source determination condition; or determine the sound field type as a non-uniform sound field if at least one of the values of the multiple sound field classification parameters satisfies a preset non-uniform sound source determination condition.
[0387] In some embodiments of the present application, the distributed sound source determination condition includes that the value of the sound field classification parameter is less than a predetermined non-uniform sound source determination threshold. Alternatively, the non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a predetermined non-uniform sound source determination threshold.
[0388] In some embodiments of the present application, there are multiple sound field classification parameters.
[0389] The sound field classification result includes a sound field type, or the sound field classification result includes a number of non-uniform sound sources and a sound field type.
[0390] The sound field classification module is configured to: obtain a non-uniform sound source number corresponding to a current frame based on values of a plurality of sound field classification parameters; and determine a sound field type based on the non-uniform sound source number corresponding to the current frame.
[0391] In some embodiments of the present application, there are multiple sound field classification parameters.
[0392] The sound field classification result includes the number of non-uniform sound sources.
[0393] The sound field classification module is configured to obtain a number of non-uniform sound sources corresponding to a current frame based on values of the plurality of sound field classification parameters.
[0394] In some embodiments of the present application, the multiple sound field classification parameters are temp[i], i=0,1,...,min(L,K)-2, where L represents the number of channels in the current frame, K represents the number of signal points corresponding to each channel in the current frame, and min represents the operation of selecting the minimum value.
[0395] The sound field classification module is configured to execute the following determination processes in sequence starting from i=0: A step of determining whether temp[i] exceeds a preset non-uniform sound source determination threshold value; and If temp[i] is less than the non-uniform sound source determination threshold in this determination procedure, update the value of i to i+1 and continue with the execution of the next determination procedure. If temp[i] is equal to or greater than the non-uniform sound source determination threshold in this determination procedure, the execution of this determination procedure is terminated, and it is determined that i in this determination procedure plus 1 is equal to the number of non-uniform sound sources.
[0396] In some embodiments of the present application, the step of determining the sound field type based on the number of non-uniform sound sources corresponding to the current frame includes: determining that the sound field type is a first sound field type if the number of non-uniform sound sources satisfies a first preset condition; or If the number of non-uniform sound sources does not satisfy the first preset condition, determining that the sound field type is the second sound field type.
[0397] The number of non-uniform sound sources corresponding to the first sound field type is different from the number of non-uniform sound sources corresponding to the second sound field type.
[0398] In some embodiments of the present application, the first preset condition includes that the number of non-uniform sound sources exceeds a first threshold or is less than a second threshold, and the second threshold exceeds the first threshold; or The first preset condition includes that the number of non-uniform sound sources is equal to or less than a first threshold or equal to or greater than a second threshold, and that the second threshold exceeds the first threshold.
[0399] In some embodiments of the present application, the audio encoding apparatus further includes an encoding mode decision module (not shown in FIG. 12 ). The encoding mode decision module is configured to determine an encoding mode corresponding to a current frame based on the sound field classification result.
[0400] In a possible implementation, the coding mode determination module is configured to: determine an coding mode corresponding to a current frame based on a non-uniform sound source number when the sound field classification result includes a non-uniform sound source number or a non-uniform sound source number and a sound field type; determine an coding mode corresponding to a current frame based on a sound field type when the sound field classification result includes a sound field type or a non-uniform sound source number and a sound field type; or determine an coding mode corresponding to a current frame based on a non-uniform sound source number and a sound field type when the sound field classification result includes a non-uniform sound source number and a sound field type.
[0401] In some embodiments of the present application, the encoding mode determination module is configured to: determine the encoding mode to be the first encoding mode if the number of non-uniform sound sources satisfies a second preset condition; or determine the encoding mode to be the second encoding mode if the number of non-uniform sound sources does not satisfy the second preset condition.
[0402] The first encoding mode is an HOA encoding mode based on virtual speaker selection or an HOA encoding mode based on directional voice coding, and the second encoding mode is an HOA encoding mode based on virtual speaker selection or an HOA encoding mode based on directional voice coding, and the first encoding mode and the second encoding mode are different encoding modes.
[0403] In some embodiments of the present application, the second preset condition includes that the number of non-uniform sound sources is greater than a first threshold or less than a second threshold, and the second threshold is greater than the first threshold; or The second preset condition includes that the number of non-uniform sound sources is equal to or less than a first threshold or equal to or greater than a second threshold, and that the second threshold exceeds the first threshold.
[0404] In some embodiments of the present application, the coding mode decision module is configured to: determine the coding mode to be a coding mode HOA based on virtual speaker selection if the sound field type is a non-uniform sound field; or determine the coding mode to be a HOA coding mode based on directional audio coding if the sound field type is a distributed sound field.
[0405] In some embodiments of the present application, the coding mode determination module is configured to: determine an initial coding mode corresponding to a current frame based on the sound field classification result of the current frame; obtain a hangover time window in which the current frame is located, where the hangover time window includes an initial coding mode of the current frame and coding modes of N-1 frames before the current frame, where N is a length of the hangover time window; and determine an coding mode of the current frame based on the initial coding mode of the current frame and the coding modes of the N-1 frames.
[0406] In some embodiments of the present application, the audio encoding apparatus further includes an encoding parameter determination module (not shown in FIG. 12). The encoding parameter determination module is configured to determine an encoding parameter corresponding to a current frame based on the sound field classification result.
[0407] In some embodiments of the present application, the encoding parameters include at least one of the number of channels of the virtual speaker signals, the number of channels of the residual signal, the number of coding bits of the virtual speaker signals, the number of coding bits of the residual signal, or the number of votes to search for the best matching speaker.
[0408] The virtual speaker signals and the residual signals are signals that are generated based on the three-dimensional audio signal.
[0409] In some embodiments of the present application, the number of votes satisfies the following relationship: 1≦I≦d
[0410] I is the number of votes, and d is the number of inhomogeneous sound sources included in the sound field classification result.
[0411] In some embodiments of the present application, the sound field classification results include the number of non-uniform sound sources and the sound field type.
[0412] When the sound field type is a non-uniform sound source, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the encoder. Or When the sound field type is a distributed sound field, the number of channels of the virtual speaker signals satisfies the following relationship: F=1 where F is the number of channels in the virtual speaker signal.
[0413] In some embodiments of the present application, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1,PR) Here, PR is the number of channels of the residual signal preset by the encoder, and C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder. Or When the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signals preset by the encoder, and F is the number of channels of the virtual speaker signals.
[0414] In some embodiments of the present application, the sound field classification result includes a number of non-uniform sound sources.
[0415] The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal that is preset by the encoder.
[0416] In some embodiments of the present application, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signals preset by the encoder, and F is the number of channels of the virtual speaker signals.
[0417] In some embodiments of the present application, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type.
[0418] The number of coding bits of the virtual speaker signal is obtained based on a ratio of the number of coding bits of the virtual speaker signal to the number of coding bits of the transmission channel.
[0419] The number of coding bits of the residual signal is obtained based on a ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels.
[0420] The number of coding bits of the transmission channels includes the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal, and when the number of non-uniform sound sources is equal to or less than the number of channels of the virtual speaker signals, the ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels is obtained by increasing an initial ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels.
[0421] In some embodiments of the present application, the audio encoding apparatus further includes an encoding module (not shown in FIG. 12 ). The encoding module is configured to encode the current frame and the sound field classification result, and write the encoded current frame and the sound field classification result into a bitstream.
[0422] From the examples in the above embodiment, it can be seen that first, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result. Then, sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition result. Finally, a sound field classification result of the current frame is determined based on the sound field classification parameters. In this embodiment of the present application, a linear decomposition is performed on a current frame of a three-dimensional audio signal to obtain a linear decomposition result of the current frame. Then, sound field classification parameters corresponding to the current frame are obtained based on the linear decomposition result. Thus, a sound field classification result of the current frame is determined based on the sound field classification parameters, and a sound field classification of the current frame can be implemented based on the sound field classification result. In this embodiment of the present application, a sound field classification is performed on a three-dimensional audio signal to accurately identify the three-dimensional audio signal.
[0423] 13 shows a three-dimensional audio signal processing device according to an embodiment of the present application. For example, the three-dimensional audio signal processing device is specifically an audio decoding device 1300, and may include: a receiving module 1301, a decoding module 1302, and a signal generating module 1303.
[0424] The receiving module is configured to receive the bitstream.
[0425] The decoding module is configured to decode the bitstream to obtain a sound field classification result for the current frame.
[0426] The signal generation module is configured to obtain a three-dimensional audio signal of the decoded current frame based on the sound field classification result.
[0427] In some embodiments of the present application, the signal generation module is configured to determine a decoding mode for the current frame based on the sound field classification result, and obtain a three-dimensional audio signal for the decoded current frame based on the decoding mode.
[0428] In some embodiments of the present application, the signal generation module is configured to: determine a decoding mode of the current frame based on the non-uniform sound source number if the sound field classification result includes a non-uniform sound source number or the sound field classification result includes a non-uniform sound source number and a sound field type; determine a decoding mode of the current frame based on the sound field type if the sound field classification result includes a sound field type or the sound field classification result includes a non-uniform sound source number and a sound field type; or determine a decoding mode of the current frame based on the non-uniform sound source number and the sound field type if the sound field classification result includes a non-uniform sound source number and a sound field type.
[0429] In some embodiments of the present application, the signal generating module is configured to: determine the decoding mode to be a first decoding mode if the number of non-uniform sound sources meets a preset condition; or determine the decoding mode to be a second decoding mode if the number of non-uniform sound sources does not meet the preset condition.
[0430] The first decoding mode is an HOA decoding mode based on virtual speaker selection or an HOA decoding mode based on directional voice coding, and the second decoding mode is an HOA decoding mode based on virtual speaker selection or an HOA decoding mode based on directional voice coding, and the first decoding mode and the second decoding mode are different decoding modes.
[0431] In some embodiments of the present application, the preset conditions include that the number of non-uniform sound sources is greater than a first threshold or less than a second threshold, and the second threshold is greater than the first threshold; or The preset conditions include that the number of non-uniform sound sources is equal to or less than a first threshold or equal to or more than a second threshold, and that the second threshold exceeds the first threshold.
[0432] In some embodiments of the present application, the signal generation module is configured to determine decoding parameters for the current frame based on the sound field classification result, and obtain a three-dimensional audio signal of the decoded current frame based on the decoding parameters.
[0433] In some embodiments of the present application, the decoding parameters include at least one of the following: the number of channels of the virtual speaker signals, the number of channels of the residual signal, the number of decoded bits of the virtual speaker signals, or the number of decoded bits of the residual signal.
[0434] The virtual speaker signals and the residual signal are obtained by decoding the bitstream.
[0435] In some embodiments of the present application, the sound field classification results include the number of non-uniform sound sources and the sound field type.
[0436] When the sound field type is a non-uniform sound source, the number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the decoder. Or, When the sound field type is a distributed sound field, the number of channels of the virtual speaker signals satisfies the following relationship: F=1 where F is the number of channels in the virtual speaker signal.
[0437] In some embodiments of the present application, when the sound field type is a distributed sound field, the number of channels of the residual signal satisfies the following relationship: R = max(C-1,PR) Here, PR is the number of channels of the residual signal preset by the decoder, and C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder. Or, When the sound field type is a non-uniform sound field, the number of channels of the residual signal satisfies the following relationship: R=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0438] In some embodiments of the present application, the sound field classification result includes a number of non-uniform sound sources.
[0439] The number of channels of the virtual speaker signal satisfies the following relationship: F = min(S,PF) Here, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal that is preset by the decoder.
[0440] In some embodiments of the present application, the number of channels of the residual signal satisfies the following relationship: F=CF Here, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signals preset by the decoder, and F is the number of channels of the virtual speaker signals.
[0441] In some embodiments of the present application, the sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type.
[0442] The number of decoding bits of the virtual speaker signal is obtained based on a ratio of the number of decoding bits of the virtual speaker signal to the number of decoding bits of the transmission channel.
[0443] The number of decoding bits of the residual signal is obtained based on a ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels.
[0444] The number of decoding bits of the transmission channels includes the number of decoding bits of the virtual speaker signals and the number of decoding bits of the residual signal, and when the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signals, the ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels is obtained by increasing the initial ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels.
[0445] From the examples in the above embodiments, it can be seen that the sound field classification result can be used to decode the current frame in the bitstream, so that the decoder side performs decoding in a decoding manner adapted to the sound field of the current frame to obtain the three-dimensional audio signal sent by the encoder side, which implements the transmission of the audio signal from the encoder side to the decoder side.
[0446] It should be noted that the contents of the information exchange between modules / units of the present device and the execution process thereof are based on the same idea as the method embodiment of the present application, and produce the same technical effect as the method embodiment of the present application. For the specific contents, please refer to the above description of the method embodiment of the present application. The details will not be described again in this specification.
[0447] An embodiment of the present application further provides a computer storage medium, which stores a program, which performs some or all of the steps described in the above method embodiments.
[0448] The following describes another speech encoding device according to an embodiment of the present application. Please refer to Figure 14. The speech encoding device 1400 includes: Receiver 1401, transmitter 1402, processor 1403, and memory 1404 (there may be one or more processors 1403 in the speech encoding device 1400, and one processor is used as an example in FIG. 14). In some embodiments of the present application, receiver 1401, transmitter 1402, processor 1403, and memory 1404 may be connected via a bus or in another manner. In FIG. 14, connection via a bus is used as an example.
[0449] The memory 1404 may include read-only memory and random access memory to provide instructions and data to the processor 1403. A portion of the memory 1404 further includes non-volatile random access memory (NVRAM). The memory 1404 stores an operating system and operating instructions, executable modules or data structures, or a subset or an extended set thereof. The operating instructions may include various operating instructions used to realize various operations. The operating system may include various system programs to implement various basic services and handle hardware-based tasks.
[0450] The processor 1403 controls the operation of the audio coding device, and may also be referred to as a central processing unit (CPU). During a particular application, the components of the audio coding device are coupled via a bus system. In addition to a data bus, the bus system may further include a power bus, a control bus, a status signal bus, and so on. However, for clarity of explanation, the various types of buses in the figures are labeled as a bus system.
[0451] The methods disclosed in the embodiments of the present application may be applied to the processor 1403 or may be implemented by using the processor 1403. The processor 1403 may be an integrated circuit chip and has signal processing capabilities. In the implementation process, the steps in the aforementioned methods may be implemented by using hardware integrated logic circuits in the processor 1403 or by using instructions in the form of software. The processor 1403 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, and so on. The steps of the methods disclosed with reference to the embodiments of the present application may be directly performed and achieved by using a hardware decoding processor, or may be performed and achieved by using a combination of hardware and software modules in the decoding processor. The software module may be arranged in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is arranged in the memory 1404, and the processor 1403 reads information in the memory 1404 and completes the steps of the method in combination with the hardware in the processor 1403.
[0452] The receiver 1401 is configured to receive input digital or textual information and generate signal inputs associated with configuration and function control of the audio coding device. The transmitter 1402 may include a display device, such as a display screen, and may be configured to output the digital or textual information via an external interface.
[0453] In this embodiment of the present application, the processor 1403 is configured to execute the methods performed by the audio encoding device in the embodiments shown in Figures 4 to 6 .
[0454] The following describes another audio decoding apparatus according to an embodiment of the present application. Please refer to Figure 15. The audio decoding apparatus 1500 includes: Receiver 1501, transmitter 1502, processor 1503, and memory 1504 (there may be one or more processors 1503 in the audio decoding device 1500, and one processor is used as an example in FIG. 15). In some embodiments of the present application, receiver 1501, transmitter 1502, processor 1503, and memory 1504 may be connected via a bus or in another manner. In FIG. 15, connection via a bus is used as an example.
[0455] The memory 1504 may include read-only memory and random access memory and may provide instructions and data to the processor 1503. A portion of the memory 1504 may further include NVRAM. The memory 1504 stores an operating system and operating instructions, executable modules or data structures, or a subset or an extension thereof. The operating instructions may include various operating instructions used to implement various operations. The operating system may include various system programs to implement various basic services and handle hardware-based tasks.
[0456] The processor 1503 controls the operation of the audio decoding device, and the processor 1503 may be referred to as a CPU. In a particular application, the components of the audio decoding device are coupled via a bus system. In addition to the data bus, the bus system may further include a power bus, a control bus, and a status signal bus, etc. However, for clarity of explanation, the various types of buses in the figures are labeled as a bus system.
[0457] The methods disclosed in the embodiments of the present application may be applied to the processor 1503 or may be implemented by using the processor 1503. The processor 1503 may be an integrated circuit chip and has signal processing capabilities. In the implementation process, the steps in the aforementioned methods may be implemented by using hardware integrated logic circuits in the processor 1503 or by using instructions in the form of software. The aforementioned processor 1503 may be a general-purpose processor, a DSP, an ASIC, an FPGA or another programmable logic component, a discrete gate or transistor logic device, or a discrete hardware component to implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, and so on. The steps of the methods disclosed with reference to the embodiments of the present application may be directly performed and achieved by using a hardware decoding processor, or may be performed and achieved by using a combination of hardware and software modules in the decoding processor. The software module may be arranged in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is arranged in the memory 1504, and the processor 1503 reads information in the memory 1504 and completes the steps of the method in combination with the hardware in the processor 1503.
[0458] In this embodiment of the present application, the processor 1503 is configured to execute the method performed by the audio decoding device in the embodiment shown in FIG.
[0459] In another possible design, when the voice encoding device or the voice decoding device is a chip in a terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit, so that the chip in the terminal executes the voice encoding method in any one of the implementations of the first aspect or the voice decoding method in any one of the implementations of the second aspect. Optionally, the storage unit is a storage unit in the chip, for example, a register or a buffer. Alternatively, the storage unit may be a storage unit in the terminal but outside the chip, for example, a read-only memory (ROM), another type of static storage device that can store static information and instructions, or a random access memory (RAM).
[0460] The processor referred to above may be a general purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits configured to control program execution of the method of the first or second aspect.
[0461] Furthermore, it should be noted that the above-described embodiment of the device is merely an example. The units described as separate parts may or may not be physically separate, and the parts shown as units may or may not be physical units, located in one location, or distributed across multiple network units. Some or all of the modules may be selected based on the actual requirements to achieve the objectives of the solutions of the embodiments. Furthermore, in the accompanying drawings of the embodiments of the device provided by the present application, the connection relationships between the modules indicate that the modules have communication connections with each other, which may be specifically implemented as one or more communication buses or signal cables.
[0462] Based on the above implementation description, those skilled in the art can clearly understand that the present application can be implemented by software in addition to the necessary general-purpose hardware, or by dedicated hardware, including dedicated integrated circuits, dedicated CPUs, dedicated memories, and dedicated components, etc. Memory, dedicated components, etc. In general, any function that can be executed by a computer program can be easily implemented by using corresponding hardware. In addition, the specific hardware configuration used to realize the same function can be in various forms, for example, in the form of an analog circuit, a digital circuit, or a dedicated circuit. However, for the present application, the implementation of a software program is a better implementation in most cases. Based on such understanding, the essential technical solution of the present application, or the part that contributes to the prior art, can be implemented in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a removable hard disk, a ROM, a RAM, a magnetic disk, or an optical disk in a computer, and includes some instructions to instruct a computer device (which may be a personal computer, a server, or a network device) to execute the method described in the embodiments of the present application.
[0463] All or some of the above embodiments may be implemented by using software, hardware, firmware, or any combination thereof. If software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product.
[0464] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the procedures or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from a computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (e.g., coaxial cable, fiber optic, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio, or microwave) manner. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device, such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
Claims
1. A three-dimensional audio signal processing method, comprising: performing a linear decomposition on a current frame of the three-dimensional audio signal to obtain a linear decomposition result; obtaining a sound field classification parameter corresponding to the current frame based on the linear decomposition result; determining a sound field classification result for the current frame based on the sound field classification parameters; determining an encoding mode corresponding to the current frame based on the sound field classification result; Equipped with The step of determining an encoding mode corresponding to the current frame based on the sound field classification result includes: determining the coding mode corresponding to the current frame based on the number of non-uniform sound sources when the sound field classification result includes a number of non-uniform sound sources or when the sound field classification result includes the number of non-uniform sound sources and a sound field type; determining the coding mode corresponding to the current frame based on the sound field type when the sound field classification result includes the sound field type, or when the sound field classification result includes the number of non-uniform sound sources and the sound field type; or determining the encoding mode corresponding to the current frame based on the number of non-uniform sound sources and the type of sound field when the sound field classification result includes the number of non-uniform sound sources and the type of sound field; Including, method.
2. The method of claim 1 , wherein the three-dimensional audio signal comprises a higher order Ambisonics HOA signal or a first order Ambisonics FOA signal.
3. said step of performing a linear decomposition on a current frame of the three-dimensional audio signal to obtain a linear decomposition result comprises: performing a singular value decomposition on the current frame to obtain singular values corresponding to the current frame, the linear decomposition result including the singular values; performing a principal component analysis on the current frame to obtain first feature values corresponding to the current frame, the linear decomposition result including the first feature values; or performing an independent component analysis on the current frame to obtain second feature values corresponding to the current frame, the linear decomposition result including the second feature values; The method of claim 1 , comprising:
4. There are a plurality of linear decomposition results, and there are a plurality of sound field classification parameters, The step of obtaining a sound field classification parameter corresponding to the current frame based on the linear decomposition result includes: obtaining a ratio of an i-th linear analysis result of the current frame to an (i+1)-th linear analysis result of the current frame, where i is a positive integer; obtaining an i-th sound field classification parameter corresponding to the current frame based on the ratio; Including, The method of claim 1.
5. There are a plurality of sound field classification parameters, and the sound field classification result includes a sound field type; The step of determining a sound field classification result for the current frame based on the sound field classification parameters comprises: determining that the sound field type is a distributed sound field when all values of the plurality of sound field classification parameters satisfy a predetermined distributed sound source determination condition; or determining that the sound field type is a non-uniform sound field when at least one value of the plurality of sound field classification parameters satisfies a predetermined non-uniform sound source determination condition; Including, The method of claim 1.
6. The distributed sound source determination condition includes that the value of the sound field classification parameter is less than a predetermined distributed sound source determination threshold value; or The non-uniform sound source determination condition includes that the value of the sound field classification parameter is equal to or greater than a predetermined non-uniform sound source determination threshold value. The method according to claim 5.
7. There are multiple sound field classification parameters, The sound field classification result includes a sound field type, or the sound field classification result includes a number of non-uniform sound sources and a sound field type, The step of determining a sound field classification result for the current frame based on the sound field classification parameters comprises: obtaining a number of non-uniform sound sources corresponding to the current frame according to values of the plurality of sound field classification parameters; determining the sound field type based on the number of non-uniform sound sources corresponding to the current frame; Including, The method of claim 1.
8. There are multiple sound field classification parameters, The sound field classification parameters include a number of non-uniform sound sources; The step of determining a sound field classification result for the current frame based on the sound field classification parameters comprises: obtaining the number of non-uniform sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters; Including, The method of claim 1.
9. the plurality of sound field classification parameters are temp[i], i=0, 1, ..., min(L,K)-2, where L represents the number of channels in the current frame, K represents the number of signal points corresponding to each channel in the current frame, and min represents an operation of selecting the minimum value; The step of obtaining a number of non-uniform sound sources corresponding to the current frame based on the values of the plurality of sound field classification parameters includes: From i=0, the following judgment procedure: A step of determining whether or not temp[i] exceeds a preset non-uniform sound source determination threshold; If temp[i] is less than the non-uniform sound source determination threshold in this determination procedure, updating the value of i to i+1 and performing a next determination procedure; or if temp[i] is equal to or greater than the non-uniform sound source determination threshold in this determination procedure, terminating execution of the determination procedure, and determining that a value obtained by adding 1 to i in this determination procedure is equal to the number of non-uniform sound sources; The method includes the steps of: The method according to claim 7.
10. The step of determining a sound field type based on the number of non-uniform sound sources corresponding to a current frame includes: determining that the sound field type is a first sound field type if the number of the non-uniform sound sources satisfies a first preset condition; or determining that the sound field type is a second sound field type if the number of the non-uniform sound sources does not satisfy a first preset condition; Including, The number of non-uniform sound sources corresponding to the first sound field type is different from the number of non-uniform sound sources corresponding to the second sound field type. The method according to claim 7.
11. The first preset condition includes that the number of the non-uniform sound sources is greater than a first threshold and less than a second threshold, and the second threshold is greater than the first threshold; or the first preset condition includes that the number of the non-uniform sound sources is equal to or less than the first threshold or equal to or more than a second threshold, and the second threshold exceeds the first threshold; The method of claim 10.
12. The step of determining the coding mode corresponding to the current frame based on the number of non-uniform sound sources comprises: determining that the encoding mode is a first encoding mode if the number of the non-uniform sound sources satisfies a second preset condition; or determining that the encoding mode is a second encoding mode if the number of the non-uniform sound sources does not satisfy a second preset condition; Including, The first encoding mode is a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional voice coding, and the second encoding mode is a HOA encoding mode based on virtual speaker selection or a HOA encoding mode based on directional voice coding, and the first encoding mode and the second encoding mode are different encoding modes. The method of claim 1.
13. The second preset condition includes that the number of the non-uniform sound sources is greater than a first threshold and less than a second threshold, and the second threshold is greater than the first threshold; or the second preset condition includes that the number of the non-uniform sound sources is equal to or less than the first threshold or equal to or more than the second threshold, and the second threshold exceeds the first threshold; The method of claim 12.
14. The step of determining the coding mode corresponding to the current frame based on the sound field type includes: If the sound field type is a non-uniform sound field, determining that the encoding mode is a virtual speaker selection-based HOA encoding mode; or determining that the coding mode is a HOA coding mode based on directional audio coding if the sound field type is a distributed sound field; Including, The method of claim 1.
15. The step of determining an encoding mode corresponding to the current frame based on the sound field classification result includes: determining an initial encoding mode corresponding to the current frame based on a sound field classification result of the current frame; obtaining a hangover window in which the current frame is located, the hangover window including the initial coding mode of the current frame and coding modes of N-1 frames prior to the current frame, where N is a length of the hangover window; determining an encoding mode for the current frame based on the initial encoding mode for the current frame and encoding modes of N-1 frames in the hangover time frame; Including, The method of claim 1.
16. The method of claim 1 , further comprising determining encoding parameters corresponding to the current frame based on the sound field classification result.
17. the encoding parameters include at least one of a number of channels of a virtual speaker signal, a number of channels of a residual signal, a number of coding bits of a virtual speaker signal, a number of coding bits of a residual signal, or a number of votes for searching for a best-matching speaker; the virtual speaker signals and the residual signal are generated based on the three-dimensional audio signal.
17. The method of claim 16.
18. The number of votes is: 1≦I≦d Fulfilling the relationship, I is the number of votes, and d is the number of non-uniform sound sources included in the sound field classification result.
20. The method of claim 17.
19. The sound field classification result includes a number of non-uniform sound sources and a sound field type; When the sound field type is a non-uniform sound field, the number of channels of the virtual speaker signal is F=min(S, PF) Fulfilling the relationship, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal that is preset by an encoder, or When the sound field type is a distributed sound field, the number of channels of the virtual speaker signal is F=1 Fulfilling the relationship, F is the number of channels of the virtual speaker signal; 20. The method of claim 17.
20. When the sound field type is a distributed sound field, the number of channels of the residual signal is R=max(C-1,PR) Fulfilling the relationship, R is the number of channels of the residual signal, PR is the number of channels of the residual signal preset by an encoder, and C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signals preset by the encoder, or When the sound field type is a non-uniform sound field, the number of channels of the residual signal is R = C - F Fulfilling the relationship, R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal.
20. The method of claim 17.
21. The sound field classification result includes a number of non-uniform sound sources; The number of channels of the virtual speaker signal is: F=min(S, PF) Fulfilling the relationship, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by an encoder.
20. The method of claim 17.
22. The number of channels of the residual signal is: F = C - F where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by an encoder and the number of channels of the virtual speaker signal preset by the encoder, and F is the number of channels of the virtual speaker signal.
20. The method of claim 17.
23. The sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and a sound field type, the number of coding bits of the virtual speaker signal is obtained based on a ratio of the number of coding bits of the virtual speaker signal to a number of coding bits of a transmission channel; the number of coding bits of the residual signal is obtained by a ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels, the number of coding bits of the transmission channels includes the number of coding bits of the virtual speaker signals and the number of coding bits of the residual signal, the number of non-uniform sound sources is less than or equal to the number of channels of the virtual speaker signals, and the ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels is obtained by increasing an initial ratio of the number of coding bits of the virtual speaker signals to the number of coding bits of the transmission channels.
20. The method of claim 17.
24. encoding the current frame and the sound field classification result; and writing the encoded current frame and the sound field classification result into a bitstream. The method of claim 1 further comprising:
25. A three-dimensional audio signal processing method, receiving a bitstream; Decoding the bitstream to obtain a sound field classification result of a current frame; obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result; Equipped with The step of obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result includes: determining a decoding mode for the current frame based on the sound field classification result; obtaining a three-dimensional audio signal of the decoded current frame based on the decoding mode; Including, The step of determining a decoding mode for the current frame based on the sound field classification result includes: determining the decoding mode of the current frame based on the number of non-uniform sound sources when the sound field classification result includes a number of non-uniform sound sources or when the sound field classification result includes a number of non-uniform sound sources and a sound field type; determining the decoding mode of the current frame based on a sound field type if the sound field classification result includes a sound field type, or if the sound field classification result includes a number of non-uniform sound sources and a sound field type; or determining the decoding mode of the current frame based on the number of non-uniform sound sources and the type of sound field when the sound field classification result includes a number of non-uniform sound sources and a type of sound field; Including, method.
26. The step of determining the decoding mode corresponding to the current frame based on the number of non-uniform excitation sources comprises: determining that the decoding mode is a first decoding mode if the number of the heterogeneous sound sources satisfies a preset condition; or determining that the decoding mode is a second decoding mode if the number of the non-uniform sound sources does not satisfy a preset condition; Including, The first decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional voice coding, and the second decoding mode is a HOA decoding mode based on virtual speaker selection or a HOA decoding mode based on directional voice coding, and the first decoding mode and the second decoding mode are different decoding modes.
26. The method of claim 25.
27. The preset condition includes that the number of the non-uniform sound sources is greater than a first threshold and less than a second threshold, and the second threshold is greater than the first threshold; or the preset conditions include that the number of the non-uniform sound sources is equal to or less than a first threshold or equal to or more than a second threshold, and the second threshold exceeds the first threshold; 27. The method of claim 26.
28. The step of obtaining a three-dimensional audio signal of the decoded current frame based on the sound field classification result includes: determining decoding parameters for the current frame based on the sound field classification result; obtaining a three-dimensional audio signal of the decoded current frame based on the decoding parameters; 26. The method of claim 25, comprising:
29. the decoding parameters include at least one of a number of channels of a virtual speaker signal, a number of channels of a residual signal, a number of decoding bits of a virtual speaker signal, or a number of decoding bits of a residual signal; the virtual speaker signals and the residual signal are obtained by decoding the bitstream.
30. The method of claim 28.
30. the sound field classification result includes a number of non-uniform sound sources and a sound field type; When the sound field type is a non-uniform sound field, the number of channels of the virtual speaker signal is F=min(S, PF) where F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and P is the number of channels of the virtual speaker signal that is preset by a decoder, or When the sound field type is a distributed sound field, the number of channels of the virtual speaker signal is F=1 where F is the number of channels of the virtual speaker signal.
30. The method of claim 29.
31. When the sound field type is a distributed sound field, the number of channels of the residual signal is R=max(C-1,PR) Fulfilling the relationship, R is the number of channels of the residual signal, PR is the number of channels of the residual signal preset by the decoder, and C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, or When the sound field type is a non-uniform sound field, the number of channels of the residual signal is R = C - F Fulfilling the relationship, R represents the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by the decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
31. The method of claim 30.
32. the sound field classification result includes the number of non-uniform sound sources; The number of channels of the virtual speaker signal is: F=min(S, PF) Fulfilling the relationship, F is the number of channels of the virtual speaker signal, S is the number of non-uniform sound sources, and PF is the number of channels of the virtual speaker signal preset by the decoder.
31. The method of claim 30.
33. The number of channels of the residual signal is: R = C - F where R is the number of channels of the residual signal, C is the sum of the number of channels of the residual signal preset by a decoder and the number of channels of the virtual speaker signal preset by the decoder, and F is the number of channels of the virtual speaker signal.
30. The method of claim 29.
34. The sound field classification result includes the number of non-uniform sound sources, or the sound field classification result includes the number of non-uniform sound sources and the sound field type, the number of decoded bits of the virtual speaker signal is obtained by a ratio of the number of decoded bits of the virtual speaker signal to the number of decoded bits of a transmission channel, the number of decoded bits of the residual signal is obtained by a ratio of the number of decoded bits of the virtual speaker signals to the number of decoded bits of the transmission channels, the number of decoding bits of the transmission channels includes the number of decoding bits of the virtual speaker signals and the number of decoding bits of the residual signal, and when the number of non-uniform sound sources is equal to or less than the number of channels of the virtual speaker signals, the ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels is obtained by increasing an initial ratio of the number of decoding bits of the virtual speaker signals to the number of decoding bits of the transmission channels.
31. The method of claim 30.
35. A three-dimensional audio signal processing device, a linear analysis module configured to perform a linear decomposition on the three-dimensional audio signal to obtain a linear decomposition result; a parameter generation module configured to obtain a sound field classification parameter corresponding to a current frame based on the linear decomposition result; a sound field classification module configured to determine a sound field classification result for the current frame based on the sound field classification parameters; an encoding mode decision module configured to decide an encoding mode corresponding to the current frame based on the sound field classification result; Preparation, The coding mode decision module includes: determining the coding mode corresponding to the current frame based on the number of non-uniform sound sources when the sound field classification result includes a number of non-uniform sound sources or when the sound field classification result includes the number of non-uniform sound sources and a sound field type; determining the coding mode corresponding to the current frame based on the sound field type when the sound field classification result includes the sound field type, or when the sound field classification result includes the number of non-uniform sound sources and the sound field type; or determining the encoding mode corresponding to the current frame based on the number of non-uniform sound sources and the type of sound field when the sound field classification result includes the number of non-uniform sound sources and the type of sound field; and further configured to: Three-dimensional audio signal processing device.
36. A three-dimensional audio signal processing device, a receiving module configured to receive a bitstream; a decoding module configured to decode the bitstream to obtain a sound field classification result for a current frame; a signal generation module configured to obtain a three-dimensional audio signal of the decoded current frame based on the sound field classification result; Equipped with The signal generation module includes: determining a decoding mode for the current frame based on the sound field classification result; obtaining a three-dimensional audio signal of the decoded current frame based on the decoding mode; [0023] 20. The method according to claim 1, further comprising: determining a decoding mode for the current frame based on the sound field classification result, determining the decoding mode of the current frame based on the number of non-uniform sound sources when the sound field classification result includes a number of non-uniform sound sources or when the sound field classification result includes a number of non-uniform sound sources and a sound field type; determining the decoding mode of the current frame based on a sound field type if the sound field classification result includes a sound field type, or if the sound field classification result includes a number of non-uniform sound sources and a sound field type; or determining the decoding mode of the current frame based on the number of non-uniform sound sources and the type of sound field when the sound field classification result includes a number of non-uniform sound sources and a type of sound field; and further configured to: Three-dimensional audio signal processing device.
37. 25. An apparatus for processing a three-dimensional audio signal, the apparatus comprising at least one processor coupled to a memory and configured to read and execute instructions stored in the memory to perform a method according to any one of claims 1 to 24.
38. The three-dimensional audio signal processing apparatus according to claim 37, further comprising the memory.
39. 35. An apparatus for processing a three-dimensional audio signal, the apparatus comprising at least one processor coupled to a memory and configured to read and execute instructions stored in the memory to implement a method according to any one of claims 25 to 34.
40. The three-dimensional audio signal processing apparatus according to claim 39, further comprising the memory.
41. 35. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to perform a method according to any one of claims 1 to 24 or any one of claims 25 to 34.
42. A computer-readable storage medium comprising: a bitstream generated by a computer as a result of execution of instructions corresponding to the method of claim 24, when the instructions are executed on the computer.
Citation Information
Patent Citations
Spatialised audio encoding with interpolation and quantifying of rotations
EP3706119A1
Compression of decomposed representations of sound fields
JP2016523468A
Compression of decomposed representations of sound fields
JP2016524727A
Method and apparatus for high-order ambisonics encoding and decoding using singular value decomposition
JP2017501440A
Apparatus and Method for Surround Audio Signal Processing
JP2017513383A