Coding and decoding method, device, equipment, storage medium and computer program product
Through a hybrid coding and decoding solution, combined with directional audio coding and virtual speaker selection technology, the problem of smooth transition of auditory quality of HOA signals under different sound field types is solved, achieving a smooth transition of high compression rate and auditory quality, and improving the rendering effect of audio signals.
Patent Information
- Application Number
- CN202111155384.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-09-29
AI Technical Summary
During the HOA signal encoding and decoding process for different sound field types, how to ensure a smooth transition of auditory quality when switching between encoding and decoding schemes, especially when there are few or many different sound sources in the sound field, existing technologies find it difficult to simultaneously meet high compression rates and smooth transition of auditory quality.
A hybrid coding and decoding scheme is adopted to encode the signal of the specified channel in the HOA signal into the bitstream. Combined with the coding and decoding scheme based on directional audio coding and virtual speaker selection, the virtual speaker signal and the residual signal are determined for coding and decoding to ensure a smooth transition of auditory quality.
The compression rate of the audio signal is improved, and a smooth transition of auditory quality is ensured when switching between codec schemes, thereby improving the auditory effect after rendering and playback.
Smart Images

Figure CN115881140B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio processing technology, and in particular to a coding and decoding method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Higher-order ambisonics (HOA), a 3D audio technology, has garnered widespread attention due to its enhanced flexibility in 3D audio playback. To achieve superior auditory quality, HOA requires a large amount of data to record detailed sound scene information. However, increasing the HOA order generates more data, which poses challenges in transmission and storage. Therefore, encoding and decoding HOA signals has become a key concern.
[0003] Related technologies propose two schemes for encoding and decoding HOA signals. One of the schemes is a coding and decoding scheme based on directional audio coding (DirAC). In this scheme, the encoder extracts the core layer signal and spatial parameters from the HOA signal of the current frame, and encodes the extracted core layer signal and spatial parameters into the bitstream. The decoder uses a decoding method symmetrical to the encoding to reconstruct the HOA signal of the current frame from the bitstream. Another scheme is a coding and decoding scheme based on virtual speaker selection. In this scheme, the encoder selects a target virtual speaker that matches the HOA signal of the current frame from the virtual speaker set based on the match-projection (MP) algorithm, determines the virtual speaker signal based on the HOA signal of the current frame and the target virtual speaker, determines the residual signal based on the HOA signal of the current frame and the virtual speaker signal, and encodes the virtual speaker signal and the residual signal into the bitstream. The decoder uses a decoding method symmetrical to the encoding to reconstruct the HOA signal of the current frame from the bitstream.
[0004] However, for the case where there are fewer dissimilar sound sources in the sound field, the compression rate of the codec solution selected based on the virtual speaker is higher. For the case where there are more dissimilar sound sources in the sound field, the compression rate of the codec solution based on DirAC is higher. Among them, dissimilar sound sources refer to point sound sources with different positions and / or directions. The sound field types of different audio frames (related to the dissimilar sound sources in the sound field) may be different. If you want to simultaneously meet the requirements of having a high compression rate for audio frames under different sound field types, you need to select a suitable codec solution for the corresponding audio frame according to the sound field type of each audio frame. This requires switching between different codec solutions. However, the HOA signals reconstructed based on different codec solutions have different auditory quality after rendering and playback. When switching between different codec solutions, how to ensure a smooth transition of auditory quality is a problem that needs to be considered at present. Summary of the Invention
[0005] The embodiments of the present application provide a coding method, apparatus, device, storage medium, and computer program product that can ensure a smooth transition of auditory quality when switching between different coding schemes. The technical solution is as follows:
[0006] In a first aspect, a coding method is provided, the method comprising:
[0007] The coding scheme for the current frame is determined based on the HOA signal of the current frame. The coding scheme for the current frame is one of the first, second, and third coding schemes. The first coding scheme is an HOA coding scheme based on directional audio coding (i.e., a DirAC decoding scheme), the second coding scheme is an HOA coding scheme based on virtual speaker selection (which can be simply referred to as an MP-based HOA decoding scheme), and the third coding scheme is a hybrid coding scheme. If the coding scheme for the current frame is the third coding scheme, the signal of a specified channel in the HOA signal is encoded into the bitstream, where the specified channel is a portion of all channels in the HOA signal. The hybrid coding scheme uses both technical means related to the first coding scheme (i.e., the DirAC coding scheme) and technical means related to the second coding scheme (the MP-based HOA coding scheme) during the encoding process, hence the name hybrid coding scheme.
[0008] In the embodiments of the present application, appropriate codec schemes are selected for different audio frames, thereby improving the compression rate of the audio signal. Furthermore, for certain audio frames, rather than directly adopting either the first or second coding scheme, a new codec scheme is adopted to encode and decode these audio frames. Specifically, the signals of the specified channels in the HOA signals of these audio frames are encoded into the bitstream, i.e., a compromise scheme is adopted for encoding and decoding, thereby ensuring a smooth transition in the auditory quality of the rendered and played HOA signals recovered from decoding.
[0009] Optionally, the signal of the designated channel includes a first-order ambisonics (FOA) signal, and the FOA signal includes an omnidirectional W signal, and directional X, Y, and Z signals.
[0010] Optionally, encoding a signal of a specified channel in the HOA signal into a bitstream includes: determining a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal; and encoding the virtual speaker signal and the residual signal into the bitstream.
[0011] Optionally, determining the virtual speaker signal and the residual signal based on the W signal, the X signal, the Y signal, and the Z signal includes: determining the W signal as one virtual speaker signal; determining three residual signals based on the W signal, the X signal, the Y signal, and the Z signal, or determining the X signal, the Y signal, and the Z signal as three residual signals. Optionally, determining the three residual signals as difference signals between the X signal, the Y signal, and the Z signal and the W signal, respectively.
[0012] Optionally, encoding the virtual speaker signal and the residual signal into the bitstream includes: combining the virtual speaker signal with the first preset mono signal to obtain a stereo signal; combining the three residual signals with the second preset mono signal to obtain two stereo signals; and encoding the obtained three stereo signals into the bitstream respectively through a stereo encoder.
[0013] Optionally, combining the three residual signals with a second preset mono signal to obtain two stereo signals includes: combining two residual signals with the highest correlation among the three residual signals to obtain one stereo signal among the two stereo signals; and combining one residual signal other than the two residual signals with the highest correlation among the three residual signals with the second preset mono signal to obtain the other stereo signal among the two stereo signals.
[0014] Optionally, the first preset mono signal is an all-zero signal or an all-one signal, the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency values are all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency values are all one; the second preset mono signal is an all-zero signal or an all-one signal; the first preset mono signal is the same as or different from the second preset mono signal.
[0015] Optionally, encoding the virtual speaker signal and the residual signal into the bitstream includes encoding the virtual speaker signal and each of the three residual signals into the bitstream respectively through a mono encoder.
[0016] Optionally, after determining the coding scheme of the current frame according to the HOA signal of the current frame, the method further includes: if the coding scheme of the current frame is the first coding scheme, encoding the HOA signal into the code stream according to the first coding scheme; if the coding scheme of the current frame is the second coding scheme, encoding the HOA signal into the code stream according to the second coding scheme.
[0017] Optionally, determining the coding scheme of the current frame according to the high-order ambisonics HOA signal of the current frame includes: determining an initial coding scheme of the current frame according to the HOA signal, the initial coding scheme being the first coding scheme or the second coding scheme; if the initial coding scheme of the current frame is the same as the initial coding scheme of the previous frame of the current frame, determining the coding scheme of the current frame to be the initial coding scheme of the current frame; if the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the previous frame of the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the previous frame of the current frame is the first coding scheme, determining the coding scheme of the current frame to be the third coding scheme.
[0018] Optionally, after determining the initial coding scheme of the current frame according to the HOA signal, the method further includes: encoding indication information of the initial coding scheme of the current frame into a bitstream.
[0019] Optionally, after determining the coding scheme of the current frame based on the HOA signal of the current frame, the method further includes: determining a value of a switching flag of the current frame, where when the coding scheme of the current frame is the first coding scheme or the second coding scheme, the value of the switching flag of the current frame is the first value; when the coding scheme of the current frame is the third coding scheme, the value of the switching flag of the current frame is the second value; and encoding the value of the switching flag into the bitstream. In other words, the switching flag is used to indicate whether the current frame is a switching frame.
[0020] Optionally, after determining the coding scheme of the current frame according to the HOA signal of the current frame, the method further includes: encoding indication information of the coding scheme of the current frame into a bitstream.
[0021] Optionally, the designated channel is consistent with a transmission channel preset in the first coding scheme, so as to ensure that the auditory quality of the switching frame is similar to the auditory quality of the audio frame encoded by the first coding scheme.
[0022] In a second aspect, a decoding method is provided, the method comprising:
[0023] A decoding scheme for the current frame is obtained based on the bitstream, where the decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme. The first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio decoding, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme. If the decoding scheme for the current frame is the third decoding scheme, a signal of a specified channel in the HOA signal of the current frame is determined based on the bitstream, where the specified channel is a portion of all channels in the HOA signal. Based on the signal of the specified channel, a gain of one or more remaining channels in the HOA signal other than the specified channel is determined. Based on the signal of the specified channel and the gains of the one or more remaining channels, a signal of each of the one or more remaining channels is determined. Based on the signal of the specified channel and the gains of the one or more remaining channels, a reconstructed HOA signal for the current frame is obtained. The hybrid decoding scheme utilizes both technical means associated with the first decoding scheme (i.e., the DirAC decoding scheme) and technical means associated with the second decoding scheme (the MP-based HOA decoding scheme) during the decoding process, hence the name hybrid decoding scheme.
[0024] In this embodiment of the present application, because the encoder encodes the signal of the specified channel into the bitstream when encoding the HOA signal of the current frame using the third encoding scheme, the decoder parses the signal of the specified channel from the bitstream and then reconstructs the signals of the remaining channels based on the signal of the specified channel, thereby reconstructing the HOA signal. In other words, a compromise solution is adopted, which ensures a smooth transition in auditory quality after rendering and playback of the decoded and restored HOA signal.
[0025] Optionally, determining a signal of a designated channel in the HOA signal of the current frame based on the bitstream includes: determining a virtual speaker signal and a residual signal based on the bitstream; and determining a signal of the designated channel based on the virtual speaker signal and the residual signal.
[0026] Optionally, determining the virtual speaker signal and the residual signal based on the bitstream includes: decoding the bitstream by a stereo decoder to obtain three stereo signals; and determining one virtual speaker signal and three residual signals based on the three stereo signals.
[0027] Optionally, determining one virtual speaker signal and three residual signals based on the three stereo signals includes: determining one virtual speaker signal based on one stereo signal among the three stereo signals; and determining three residual signals based on the other two stereo signals among the three stereo signals.
[0028] Optionally, determining the virtual speaker signal and the residual signal based on the bitstream includes: decoding the bitstream using a mono decoder to obtain one virtual speaker signal and three residual signals.
[0029] Optionally, the signal of the designated channel includes a first-order ambisonics (FOA) signal, and the FOA signal includes an omnidirectional W signal and directional X, Y, and Z signals. Determining the signal of the designated channel based on the virtual speaker signal and the residual signal includes: determining the W signal based on the virtual speaker signal; determining the X, Y, and Z signals based on the residual signal and the W signal, or determining the X, Y, and Z signals based on the residual signal.
[0030] Optionally, the method further includes: if the decoding scheme of the current frame is the first decoding scheme, obtaining the reconstructed HOA signal of the current frame according to the code stream according to the first decoding scheme; if the decoding scheme of the current frame is the second decoding scheme, obtaining the reconstructed HOA signal of the current frame according to the code stream according to the second decoding scheme.
[0031] Optionally, obtaining a reconstructed HOA signal for the current frame based on the bitstream according to the second decoding scheme includes: obtaining an initial HOA signal based on the bitstream according to the second decoding scheme; if the decoding scheme of the frame preceding the current frame is the third decoding scheme, performing gain adjustment on a high-order portion of the initial HOA signal based on a high-order gain of the frame preceding the current frame; and obtaining a reconstructed HOA signal based on a low-order portion of the initial HOA signal and the gain-adjusted high-order portion. Specifically, the high-order gain adjustment further smooths the transition in auditory quality.
[0032] Optionally, obtaining a decoding scheme for the current frame based on the code stream includes: parsing a value of a switching flag of the current frame from the code stream; if the value of the switching flag is a first value, parsing indication information of the decoding scheme for the current frame from the code stream, the indication information being used to indicate that the decoding scheme for the current frame is the first decoding scheme or the second decoding scheme; if the value of the switching flag is a second value, determining that the decoding scheme for the current frame is a third decoding scheme.
[0033] Optionally, obtaining the decoding scheme of the current frame based on the code stream includes: parsing indication information of the decoding scheme of the current frame from the code stream, where the indication information is used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
[0034] Optionally, obtaining a decoding scheme for the current frame based on the code stream includes: parsing an initial decoding scheme for the current frame from the code stream, where the initial decoding scheme is a first decoding scheme or a second decoding scheme; if the initial decoding scheme of the current frame is the same as the initial decoding scheme of the frame before the current frame, determining that the decoding scheme for the current frame is the initial decoding scheme for the current frame; if the initial decoding scheme for the current frame is the first decoding scheme and the initial decoding scheme for the frame before the current frame is the second decoding scheme, or if the initial decoding scheme for the current frame is the second decoding scheme and the initial decoding scheme for the frame before the current frame is the first decoding scheme, determining that the decoding scheme for the current frame is a third decoding scheme.
[0035] In a third aspect, a coding device is provided, wherein the coding device has the function of implementing the coding method described in the first aspect. The coding device includes one or more modules, wherein the one or more modules are used to implement the coding method described in the first aspect.
[0036] That is, a coding device is provided, the device comprising:
[0037] a first determining module, configured to determine a coding scheme for the current frame based on a high-order ambisonics (HOA) signal of the current frame, wherein the coding scheme for the current frame is one of a first coding scheme, a second coding scheme, and a third coding scheme; wherein the first coding scheme is an HOA coding scheme based on directional audio coding, the second coding scheme is an HOA coding scheme based on virtual speaker selection, and the third coding scheme is a hybrid coding scheme;
[0038] The first encoding module is configured to encode a signal of a specified channel in the HOA signal into a bitstream if the encoding scheme of the current frame is the third encoding scheme, where the specified channel is a portion of all channels of the HOA signal.
[0039] Optionally, the signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X, Y, and Z signals.
[0040] Optionally, the first encoding module includes:
[0041] A first determination submodule is configured to determine a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal;
[0042] The encoding submodule is used to encode the virtual loudspeaker signal and the residual signal into a bit stream.
[0043] Optionally, the first determining submodule is configured to:
[0044] Determine the W signal as a virtual loudspeaker signal;
[0045] Three residual signals are determined based on the W signal, the X signal, the Y signal, and the Z signal, or the X signal, the Y signal, and the Z signal are determined as the three residual signals.
[0046] Optionally, the encoding submodule is used to:
[0047] Combine the virtual speaker signal with the first preset mono signal to obtain a stereo signal;
[0048] Combining the three residual signals with a second preset mono signal to obtain two stereo signals;
[0049] The three stereo signals obtained are encoded into bit streams respectively through a stereo encoder.
[0050] Optionally, the encoding submodule is used to:
[0051] Combining the two residual signals with the highest correlation among the three residual signals to obtain one stereo signal among the two stereo signals;
[0052] One of the three residual signals, except the two residual signals with the highest correlation, is combined with the second preset mono signal to obtain the other stereo signal of the two stereo signals.
[0053] Optionally, the first preset mono signal is an all-zero signal or an all-one signal, the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency values are all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency values are all one; the second preset mono signal is an all-zero signal or an all-one signal; the first preset mono signal is the same as or different from the second preset mono signal.
[0054] Optionally, the encoding submodule is used to:
[0055] The virtual speaker signal and each of the three residual signals are encoded into a bitstream through a mono encoder.
[0056] Optionally, the device further comprises:
[0057] a second encoding module, configured to encode the HOA signal into a bitstream according to the first encoding scheme if the encoding scheme of the current frame is the first encoding scheme;
[0058] The third encoding module is configured to encode the HOA signal into a bitstream according to the second encoding scheme if the encoding scheme of the current frame is the second encoding scheme.
[0059] Optionally, the first determining module includes:
[0060] A second determining submodule is configured to determine an initial coding scheme for a current frame according to the HOA signal, where the initial coding scheme is the first coding scheme or the second coding scheme;
[0061] a third determining submodule, configured to determine the coding scheme of the current frame as the initial coding scheme of the current frame if the initial coding scheme of the current frame is the same as the initial coding scheme of the frame before the current frame;
[0062] The fourth determination submodule is used to determine that the coding scheme of the current frame is the third coding scheme if the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the previous frame of the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the previous frame of the current frame is the first coding scheme.
[0063] Optionally, the device further comprises:
[0064] The fourth encoding module is configured to encode indication information of an initial encoding scheme of a current frame into a bitstream.
[0065] Optionally, the device further comprises:
[0066] a second determining module, configured to determine a value of a switching flag of a current frame, wherein when the coding scheme of the current frame is the first coding scheme or the second coding scheme, the value of the switching flag of the current frame is the first value; and when the coding scheme of the current frame is the third coding scheme, the value of the switching flag of the current frame is the second value;
[0067] The fifth encoding module is used to encode the value of the switching flag into the code stream.
[0068] Optionally, the device further comprises:
[0069] The sixth encoding module is used to encode the indication information of the encoding scheme of the current frame into the bitstream.
[0070] Optionally, the designated channel is consistent with a transmission channel preset in the first coding scheme.
[0071] In a fourth aspect, a decoding device is provided, wherein the decoding device has the function of implementing the decoding method described in the second aspect. The decoding device includes one or more modules, wherein the one or more modules are used to implement the decoding method described in the second aspect.
[0072] That is, a decoding device is provided, the device comprising:
[0073] A first obtaining module is configured to obtain a decoding scheme for a current frame based on a bitstream, where the decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme; wherein the first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio decoding, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme;
[0074] a first determining module configured to determine, if the decoding scheme of the current frame is the third decoding scheme, signals of a specified channel in the HOA signal of the current frame based on the bitstream, where the specified channel is a portion of all channels of the HOA signal;
[0075] A second determination module is configured to determine, based on the signal of the designated channel, the gains of one or more remaining channels in the HOA signal except the designated channel;
[0076] a third determining module, configured to determine a signal of each of the one or more remaining channels based on the signal of the designated channel and the gains of the one or more remaining channels;
[0077] The second obtaining module is configured to obtain a reconstructed HOA signal of a current frame based on the signal of the designated channel and the signals of the one or more remaining channels.
[0078] Optionally, the first determining module includes:
[0079] A first determination submodule, configured to determine a virtual speaker signal and a residual signal based on a bit stream;
[0080] The second determining submodule is configured to determine a signal of a designated channel based on the virtual speaker signal and the residual signal.
[0081] Optionally, the first determining submodule is configured to:
[0082] Decode the code stream through a stereo decoder to obtain three-way stereo signals;
[0083] Based on the three stereo signals, one virtual speaker signal and three residual signals are determined.
[0084] Optionally, the first determining submodule is configured to:
[0085] determining a virtual loudspeaker signal based on one of the three stereo signals;
[0086] Three residual signals are determined based on the other two stereo signals of the three stereo signals.
[0087] Optionally, the first determining submodule is configured to:
[0088] The bit stream is decoded by a mono decoder to obtain one virtual speaker signal and three residual signals.
[0089] Optionally, the signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X signal, Y signal, and Z signal;
[0090] The first determination submodule is used for:
[0091] determining a W signal based on the virtual loudspeaker signal;
[0092] The X signal, the Y signal, and the Z signal are determined based on the residual signal and the W signal, or the X signal, the Y signal, and the Z signal are determined based on the residual signal.
[0093] Optionally, the device further comprises:
[0094] A first decoding module, configured to obtain a reconstructed HOA signal of the current frame according to the bitstream in accordance with the first decoding scheme if the decoding scheme of the current frame is the first decoding scheme;
[0095] The second decoding module is configured to obtain a reconstructed HOA signal of the current frame according to the bit stream in accordance with the second decoding scheme if the decoding scheme of the current frame is the second decoding scheme.
[0096] Optionally, the second decoding module includes:
[0097] A first obtaining submodule is configured to obtain an initial HOA signal according to a bit stream in accordance with a second decoding scheme;
[0098] a gain adjustment submodule, configured to adjust the gain of a high-order portion of the initial HOA signal according to the high-order gain of the frame before the current frame if the decoding scheme of the frame before the current frame is the third decoding scheme;
[0099] The second obtaining submodule is configured to obtain a reconstructed HOA signal based on the low-order portion of the initial HOA signal and the gain-adjusted high-order portion.
[0100] Optionally, the first obtaining module includes:
[0101] A first parsing submodule is used to parse the code stream to obtain the value of the switching flag of the current frame;
[0102] a second parsing submodule, configured to parse indication information of a decoding scheme for a current frame from the bitstream if the value of the switching flag is the first value, the indication information being used to indicate whether the decoding scheme for the current frame is the first decoding scheme or the second decoding scheme;
[0103] The third determining submodule is configured to determine that the decoding scheme for the current frame is a third decoding scheme if the value of the switching flag is the second value.
[0104] Optionally, the first obtaining module includes:
[0105] The third parsing submodule is configured to parse the code stream to obtain indication information of a decoding scheme for the current frame, where the indication information is used to indicate whether the decoding scheme for the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
[0106] Optionally, the first obtaining module includes:
[0107] a fourth parsing submodule, configured to parse the bitstream to obtain an initial decoding scheme for the current frame, where the initial decoding scheme is the first decoding scheme or the second decoding scheme;
[0108] a fourth determining submodule, configured to determine the decoding scheme of the current frame as the initial decoding scheme of the current frame if the initial decoding scheme of the current frame is the same as the initial decoding scheme of the frame before the current frame;
[0109] The fifth determination submodule is used to determine that the decoding scheme of the current frame is the third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme.
[0110] In a fifth aspect, an encoding terminal device is provided, comprising a processor and a memory, wherein the memory is configured to store a program for executing the encoding method provided in the first aspect, and to store data involved in implementing the encoding method provided in the first aspect. The processor is configured to execute the program stored in the memory. The operating device of the storage device may further include a communication bus configured to establish a connection between the processor and the memory.
[0111] In a sixth aspect, a decoding device is provided, comprising a processor and a memory, wherein the memory is configured to store a program for executing the decoding method provided in the second aspect, as well as data used to implement the decoding method provided in the second aspect. The processor is configured to execute the program stored in the memory. The operating device of the storage device may further include a communication bus configured to establish a connection between the processor and the memory.
[0112] In the seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer executes the encoding method described in the first aspect or the decoding method described in the second aspect.
[0113] In an eighth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the encoding method described in the first aspect or the decoding method described in the second aspect.
[0114] The technical effects obtained in the above-mentioned third, fourth, fifth, sixth, seventh and eighth aspects are similar to the technical effects obtained by the corresponding technical means in the first or second aspect, and will not be repeated here.
[0115] The technical solutions provided in the embodiments of the present application can at least bring the following beneficial effects:
[0116] In an embodiment of the present application, two schemes (i.e., a coding scheme based on virtual speaker selection and a coding scheme based on directional audio coding) are combined to encode and decode the HOA signal of the audio frame, that is, a suitable coding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to ensure a smooth transition of auditory quality when switching between different coding schemes, for some audio frames, this scheme does not directly adopt any of the above two schemes for coding and decoding, but adopts a new coding scheme to encode and decode these audio frames, that is, the signal of the specified channel in the HOA signal of these audio frames is encoded into the bit stream, that is, a compromise scheme is adopted for coding and decoding, so that the auditory quality after rendering and playing the HOA signal restored by decoding can be smoothly transitioned. BRIEF DESCRIPTION OF THE DRAWINGS
[0117] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0118] Figure 2 This is a schematic diagram of an implementation environment of a terminal scenario provided by an embodiment of the present application;
[0119] Figure 3 This is a schematic diagram of an implementation environment for a transcoding scenario of a wireless or core network device provided in an embodiment of the present application;
[0120] Figure 4 This is a schematic diagram of an implementation environment of a broadcasting and television scenario provided by an embodiment of the present application;
[0121] Figure 5 is a schematic diagram of an implementation environment of a virtual reality streaming scenario provided by an embodiment of the present application;
[0122] Figure 6 This is a flowchart of an encoding method provided in an embodiment of the present application;
[0123] Figure 7 is a schematic diagram of a switching frame coding scheme provided in an embodiment of the present application;
[0124] Figure 8 is a schematic diagram of an HOA coding scheme based on virtual speaker selection provided in an embodiment of the present application;
[0125] Figure 9 is a schematic diagram of a DirAC-based HOA coding scheme provided in an embodiment of the present application;
[0126] Figure 10 is a flowchart of another encoding method provided in an embodiment of the present application;
[0127] Figure 11 This is a flowchart of a decoding method provided by an embodiment of the present application;
[0128] Figure 12 Schematic diagram of a switching frame decoding solution provided in an embodiment of the present application;
[0129] Figure 13 is a schematic diagram of an HOA decoding solution based on virtual speaker selection provided in an embodiment of the present application;
[0130] Figure 14 is a schematic diagram of a DirAC-based HOA decoding solution provided in an embodiment of the present application;
[0131] Figure 15 is a flowchart of another decoding method provided in an embodiment of the present application;
[0132] Figure 16 This is a schematic structural diagram of an encoding device provided in an embodiment of the present application;
[0133] Figure 17 This is a schematic structural diagram of a decoding device provided in an embodiment of the present application;
[0134] Figure 18 This is a schematic block diagram of a coding and decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0135] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0136] Before explaining the encoding and decoding method provided in the embodiment of the present application in detail, the implementation environment involved in the embodiment of the present application is first introduced.
[0137] Please refer to Figure 1 , Figure 1is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. Source device 10 can generate encoded media data. Therefore, source device 10 can also be referred to as a media data encoding device. Destination device 20 can decode the encoded media data generated by source device 10. Therefore, destination device 20 can also be referred to as a media data decoding device. Link 30 can receive the encoded media data generated by source device 10 and transmit the encoded media data to destination device 20. Storage device 40 can receive the encoded media data generated by source device 10 and store the encoded media data. Under such conditions, destination device 20 can directly obtain the encoded media data from storage device 40. Alternatively, storage device 40 can correspond to a file server or another intermediate storage device that can store the encoded media data generated by source device 10. Under such conditions, destination device 20 can stream or download the encoded media data stored by storage device 40.
[0138] The source device 10 and the destination device 20 may each include one or more processors and a memory coupled to the one or more processors, which may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, any other medium that can be used to store desired program code in the form of computer-accessible instructions or data structures, etc. For example, the source device 10 and the destination device 20 may each include a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video game console, an in-vehicle computer, or the like.
[0139] Link 30 may include one or more media or devices capable of transmitting encoded media data from source device 10 to destination device 20. In one possible implementation, link 30 may include one or more communication media that enable source device 10 to send encoded media data directly to destination device 20 in real time. In an embodiment of the present application, source device 10 may modulate the encoded media data based on a communication standard, such as a wireless communication protocol, and may transmit the modulated media data to destination device 20. The one or more communication media may include wireless and / or wired communication media, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (e.g., the Internet). The one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from source device 10 to destination device 20, and the embodiments of the present application are not particularly limited in this regard.
[0140] In one possible implementation, the storage device 40 may store the received encoded media data sent by the source device 10, and the destination device 20 may directly obtain the encoded media data from the storage device 40. Under such conditions, the storage device 40 may include any of a variety of distributed or locally accessible data storage media, for example, any of the various distributed or locally accessible data storage media may be a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded media data.
[0141] In one possible implementation, storage device 40 may correspond to a file server or another intermediate storage device that can store the encoded media data generated by source device 10. Destination device 20 may stream or download the media data stored on storage device 40. The file server may be any type of server capable of storing and transmitting the encoded media data to destination device 20. In one possible implementation, the file server may include a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. Destination device 20 may obtain the encoded media data via any standard data connection, including an internet connection. Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of the two suitable for obtaining the encoded media data stored on the file server. The transmission of the encoded media data from storage device 40 may be streaming, downloading, or a combination of the two.
[0142] Figure 1 The implementation environment shown is only one possible implementation method, and the technology of the embodiment of the present application is not only applicable to Figure 1 The source device 10 that can encode media data and the destination device 20 that can decode the encoded media data shown can also be applied to other devices that can encode media data and decode encoded media data, and the embodiments of the present application do not specifically limit this.
[0143] exist Figure 1 In the illustrated embodiment, source device 10 includes a data source 120, an encoder 100, and an output interface 140. In some embodiments, output interface 140 may include a modem and / or a transmitter, where the transmitter may also be referred to as a transmitter. Data source 120 may include an image capture device (e.g., a camera), an archive containing previously captured media data, a feed interface for receiving media data from a media data content provider, and / or a computer graphics system for generating media data, or a combination of these sources of media data.
[0144] The data source 120 may send media data to the encoder 100, and the encoder 100 may encode the media data received from the data source 120 to obtain encoded media data. The encoder may send the encoded media data to an output interface. In some embodiments, the source device 10 sends the encoded media data directly to the destination device 20 via the output interface 140. In other embodiments, the encoded media data may also be stored on the storage device 40 for later retrieval by the destination device 20 for decoding and / or display.
[0145] exist Figure 1 In the illustrated implementation, the destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, the input interface 240 includes a receiver and / or a modem. The input interface 240 may receive encoded media data via the link 30 and / or from the storage device 40, and then transmit the encoded media data to the decoder 200. The decoder 200 may decode the received encoded media data to obtain decoded media data. The decoder may transmit the decoded media data to the display device 220. The display device 220 may be integrated with the destination device 20 or may be external to the destination device 20. Generally, the display device 220 displays the decoded media data. The display device 220 may be any of a variety of types of display devices, for example, the display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0146] although Figure 1 Although not shown, in some aspects, encoder 100 and decoder 200 can be integrated with an encoder and decoder, respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software for encoding both audio and video in a common data stream or in separate data streams. In some embodiments, the MUX-DEMUX units can conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP), if applicable.
[0147] The encoder 100 and the decoder 200 can each be any of the following circuits: one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. If the technology of the embodiments of the present application is implemented in part in software, the device can store instructions for the software in a suitable non-volatile computer-readable storage medium, and can use one or more processors to execute the instructions in hardware to implement the technology of the embodiments of the present application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be regarded as one or more processors. Each of the encoder 100 and the decoder 200 can be included in one or more encoders or decoders, and any of the encoders or decoders can be integrated as part of a combined encoder / decoder (encoder / decoder) in the corresponding device.
[0148] Embodiments of the present application may generally refer to encoder 100 as "signaling" or "sending" certain information to another device, such as decoder 200. The terms "signaling" or "sending" may generally refer to the transmission of syntax elements and / or other data used to decode compressed media data. This transmission may occur in real time or near real time. Alternatively, this communication may occur over time, such as when syntax elements are stored in the encoded bitstream to a computer-readable storage medium during encoding, and a decoding device may then retrieve the syntax elements at any time after they are stored to this medium.
[0149] The encoding and decoding method provided in the embodiments of the present application can be applied to a variety of scenarios. Next, taking the media data to be encoded as an HOA signal as an example, several scenarios are introduced respectively.
[0150] Please refer to Figure 2 , Figure 2 This is a schematic diagram of an implementation environment for a coding and decoding method provided in an embodiment of the present application, applied to a terminal scenario. The implementation environment includes a first terminal 101 and a second terminal 201, which are in communication with each other. This communication connection can be a wireless connection or a wired connection, which is not limited in this embodiment of the present application.
[0151] The first terminal 101 can be a transmitting device or a receiving device. Similarly, the second terminal 201 can be a receiving device or a transmitting device. For example, when the first terminal 101 is a transmitting device, the second terminal 201 is a receiving device. When the first terminal 101 is a receiving device, the second terminal 201 is a transmitting device.
[0152] The following is an introduction using the first terminal 101 as a transmitting device and the second terminal 201 as a receiving device as an example.
[0153] The first terminal 101 and the second terminal 201 each include an audio acquisition module, an audio playback module, an encoder, a decoder, a channel encoding module, and a channel decoding module. In the embodiment of the present application, the encoder is a 3D audio encoder, and the decoder is a 3D audio decoder.
[0154] The audio acquisition module in the first terminal 101 collects the HOA signal and transmits it to the encoder. The encoder encodes the HOA signal using the encoding method provided in the embodiments of the present application. This encoding method can be called source coding. Subsequently, to enable transmission of the HOA signal over the channel, the channel coding module further performs channel coding. The resulting encoded code stream is then transmitted over a digital channel via wireless or wired network communication equipment.
[0155] The second terminal 201 receives the code stream transmitted in the digital channel through a wireless or wired network communication device. The channel decoding module performs channel decoding on the code stream. Then, the decoder uses the decoding method provided in the embodiment of the present application to decode the HOA signal, which is then played through the audio playback module.
[0156] Among them, the first terminal 101 and the second terminal 201 can be any electronic product that can interact with the user through one or more methods such as keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, such as personal computer (PC), mobile phone, smart phone, personal digital assistant (PDA), wearable device, PPC (pocket PC), tablet computer, smart car machine, smart TV, smart speaker, etc.
[0157] Those skilled in the art should understand that the above-mentioned terminals are only examples, and other existing or future terminals that are applicable to the embodiments of the present application should also be included in the protection scope of the embodiments of the present application and are included here by reference.
[0158] Please refer to Figure 3 , Figure 3This is a schematic diagram illustrating an implementation environment for a codec method provided in an embodiment of the present application, applied to a transcoding scenario for wireless or core network devices. The implementation environment includes a channel decoding module, an audio decoder, an audio encoder, and a channel coding module. In this embodiment of the present application, the audio encoder is a 3D audio encoder, and the audio decoder is a 3D audio decoder.
[0159] Among them, the audio decoder can be a decoder using the decoding method provided in the embodiment of the application, or a decoder using other decoding methods. The audio encoder can be an encoder using the encoding method provided in the embodiment of the application, or an encoder using other encoding methods. In the case where the audio decoder is a decoder using the decoding method provided in the embodiment of the application, the audio encoder is an encoder using other encoding methods; in the case where the audio decoder is a decoder using other decoding methods, the audio encoder is an encoder using the encoding method provided in the embodiment of the application.
[0160] In the first case, the audio decoder is a decoder using the decoding method provided in the embodiment of the present application, and the audio encoder is an encoder using other encoding methods.
[0161] In this case, the channel decoding module is used to perform channel decoding on the received code stream. Then, the audio decoder is used to perform source decoding using the decoding method provided in the embodiment of the present application. Then, the audio encoder is used to encode the code stream using another encoding method to achieve conversion from one format to another, i.e., transcoding. After that, the code stream is transmitted after channel encoding.
[0162] In the second case, the audio decoder is a decoder that uses other decoding methods, and the audio encoder is an encoder that uses the encoding method provided in the embodiment of the present application.
[0163] In this case, the channel decoding module is used to perform channel decoding on the received code stream, and then the audio decoder is used to perform source decoding using other decoding methods. The audio encoder is then used to encode the code stream using the encoding method provided in the embodiment of the present application, thereby converting the code stream from one format to another, i.e., transcoding. The code stream is then transmitted after channel encoding.
[0164] The wireless device may be a wireless access point, a wireless router, a wireless connector, etc. The core network device may be a mobility management entity, a gateway, etc.
[0165] Those skilled in the art should understand that the above-mentioned wireless devices or core network devices are only examples. Other existing or future wireless or core network devices that are applicable to the embodiments of the present application should also be included in the scope of protection of the embodiments of the present application and are included here by reference.
[0166] Please refer to Figure 4, Figure 4 This is a schematic diagram of an implementation environment for a coding and decoding method provided in an embodiment of the present application, applied to a broadcasting and television scenario. Broadcasting and television scenarios are divided into live broadcasting and post-production scenarios. For live broadcasting, the implementation environment includes a live program 3D sound production module, a 3D sound encoding module, a set-top box, and a speaker system. The set-top box includes a 3D sound decoding module. For post-production, the implementation environment includes a post-program 3D sound production module, a 3D sound encoding module, a network receiver, a mobile terminal, headphones, and more.
[0167] In a live broadcast scenario, the live program 3D sound production module produces a 3D sound signal (such as an HOA signal). The 3D sound signal is encoded using the encoding method of the embodiment of the present application to obtain a code stream. The code stream is transmitted to the user side via the broadcasting network. The 3D sound decoder in the set-top box uses the decoding method provided in the embodiment of the present application to decode the code stream, thereby reconstructing the 3D sound signal, which is played back by the speaker group. Alternatively, the code stream is transmitted to the user side via the Internet, and the 3D sound decoder in the network receiver uses the decoding method provided in the embodiment of the present application to decode the code stream, thereby reconstructing the 3D sound signal, which is played back by the speaker group. Alternatively, the code stream is transmitted to the user side via the Internet, and the 3D sound decoder in the mobile terminal uses the decoding method provided in the embodiment of the present application to decode the code stream, thereby reconstructing the 3D sound signal, which is played back by the headphones.
[0168] In a post-production scenario, the post-production program 3D sound production module produces a 3D sound signal. This 3D sound signal is encoded using the encoding method of an embodiment of the present application to obtain a bitstream. This bitstream is transmitted to the user via a broadcasting network. The 3D sound decoder in the set-top box decodes the bitstream using the decoding method provided in an embodiment of the present application, thereby reconstructing the 3D sound signal for playback by the speaker group. Alternatively, the bitstream is transmitted to the user via the Internet. The 3D sound decoder in the network receiver decodes the bitstream using the decoding method provided in an embodiment of the present application, thereby reconstructing the 3D sound signal for playback by the speaker group. Alternatively, the bitstream is transmitted to the user via the Internet. The 3D sound decoder in the mobile terminal decodes the bitstream using the decoding method provided in an embodiment of the present application, thereby reconstructing the 3D sound signal for playback by headphones.
[0169] Please refer to Figure 5 , Figure 5 This is a schematic diagram of an implementation environment for a coding and decoding method provided in an embodiment of the present application, applied to a virtual reality streaming scenario. The implementation environment includes an encoding end and a decoding end. The encoding end includes an acquisition module, a preprocessing module, an encoding module, a packaging module, and a sending module. The decoding end includes an unpacking module, a decoding module, a rendering module, and headphones.
[0170] The acquisition module collects the HOA signal, which is then preprocessed by the preprocessing module. This includes filtering out the low-frequency portion of the HOA signal, typically at a cutoff of 20 Hz or 50 Hz, and extracting the azimuth information from the HOA signal. The encoding module then encodes the signal using the encoding method provided in the embodiments of this application. After encoding, the package module packages the signal and then transmits it to the decoder via the transmission module.
[0171] The decoding end's unpacking module first unpacks the audio, which is then decoded by the decoding module using the decoding method provided in the embodiments of the present application. The rendering module then performs binaural rendering on the decoded signal, which is then mapped to the listener's headphones. The headphones can be standalone or integrated into a virtual reality headset.
[0172] It should be noted that the system architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0173] Next, the encoding and decoding method provided in the embodiment of the present application is explained in detail. Figure 1 In the illustrated implementation environment, any of the following encoding methods may be executed by the encoder 100 in the source device 10 , and any of the following decoding methods may be executed by the decoder 200 in the destination device 20 .
[0174] Figure 6 This is a flowchart of an encoding method provided in an embodiment of the present application, which is applied to the encoding end. Figure 6 , the method includes the following steps.
[0175] Step 601: Determine a coding scheme for the current frame according to the HOA signal of the current frame.
[0176] For the HOA signals of multiple audio frames to be encoded, the encoding end encodes them frame by frame. Among them, the HOA signal of the audio frame is an audio signal obtained by the HOA acquisition technology. The HOA signal is a scene audio signal and also a three-dimensional audio signal. The HOA signal refers to the audio signal obtained by collecting the sound field at the position of the microphone in the space. The collected audio signal is called the original HOA signal. The HOA signal of the audio frame can also be an HOA signal obtained by converting three-dimensional audio signals in other formats. For example, a 5.1-channel signal is converted into an HOA signal, or a three-dimensional audio signal mixed with a 5.1-channel signal and object audio is converted into an HOA signal. Optionally, the HOA signal of the audio frame to be encoded is a time domain signal or a frequency domain signal, which can include all channels of the HOA signal or some channels of the HOA signal. For example, if the order of the HOA signal of the audio frame is 3, the number of channels of the HOA signal is 16, the frame length of the audio frame is 20 ms, and the sampling rate is 48 kHz, then the HOA signal of the audio frame to be encoded includes 16 channel signals, and each channel includes 960 sampling points.
[0177] To reduce computational complexity, if the HOA signal of the audio frame obtained by the encoder is the original HOA signal, and the original HOA signal has a large number of sampling points or frequencies, the encoder can downsample the original HOA signal to obtain the HOA signal of the audio frame to be encoded. For example, the encoder performs 1 / Q downsampling on the original HOA signal to reduce the number of sampling points or frequencies of the HOA signal to be encoded. For example, in the embodiment of the present application, each channel of the original HOA signal contains 960 sampling points. After 1 / 120 downsampling, each channel of the obtained HOA signal to be encoded contains 8 sampling points.
[0178] In this embodiment of the present application, the encoding method of the encoder is described using the encoding of the current frame as an example. The current frame is an audio frame to be encoded. Specifically, the encoder obtains the HOA signal of the current frame and encodes the HOA signal of the current frame using the encoding method provided in this embodiment of the present application.
[0179] It should be noted that in order to meet the requirements of a high compression rate for audio frames under different sound field types, it is necessary to select a suitable coding scheme for the corresponding audio frame according to the sound field type of each audio frame. In an embodiment of the present application, the encoding end first determines the initial coding scheme of the current frame based on the HOA signal of the current frame, and the initial coding scheme is the first coding scheme or the second coding scheme. The encoding end then determines whether to use the first coding scheme, the second coding scheme or the third coding scheme to encode the HOA signal of the current frame by comparing whether the initial coding scheme of the current frame is the same as the initial coding scheme of the previous frame of the current frame. Among them, if the initial coding scheme of the current frame is the same as the initial coding scheme of the previous frame of the current frame, the encoding end adopts a coding scheme consistent with the initial coding scheme of the current frame to encode the HOA signal of the current frame. If the initial coding scheme of the current frame is different from the initial coding scheme of the previous frame of the current frame, the encoding end adopts a switching frame coding scheme to encode the HOA signal of the current frame.
[0180] In an embodiment of the present application, the coding scheme of the current frame is one of the first coding scheme, the second coding scheme and the third coding scheme. Among them, the first coding scheme is an HOA coding scheme based on DirAC, the second coding scheme is an HOA coding scheme based on virtual speaker selection, and the third coding scheme is a hybrid coding scheme. Optionally, the hybrid coding scheme is also called a switching frame coding scheme. The third coding scheme is a switching frame coding scheme provided in an embodiment of the present application. The third coding scheme is for a smooth transition of auditory quality when switching between different coding and decoding schemes. The embodiment of the present application will introduce these three coding schemes in detail below. In an embodiment of the present application, the HOA coding scheme based on virtual speaker selection is also called an MP-based HOA coding scheme.
[0181] In this embodiment of the present application, the encoder determines the initial coding scheme for the current frame based on the HOA signal of the current frame. The encoder then determines the coding scheme for the current frame based on the initial coding scheme of the current frame and the initial coding scheme of the frame immediately preceding the current frame. It should be noted that this embodiment of the present application does not limit the implementation method for determining the initial coding scheme by the encoder.
[0182] Optionally, the encoder performs a sound field type analysis on the HOA signal of the current frame to obtain a sound field classification result for the current frame, and determines an initial encoding scheme for the current frame based on the sound field classification result of the current frame. It should be noted that the embodiments of the present application do not limit the method of sound field type analysis. For example, the encoder may perform a sound field type analysis by performing singular value decomposition on the HOA signal of the current frame, or perform other linear decomposition on the HOA signal to perform sound field type analysis.
[0183] Optionally, the sound field classification result includes the number of dissimilar sound sources. Taking the example of the encoder directly performing sound field type analysis on the HOA signal of the current frame, one implementation method for performing sound field type analysis on the HOA signal of the current frame to obtain the sound field classification result of the current frame is as follows: the encoder performs singular value decomposition on the HOA signal of the current frame to obtain M singular values. The encoder calculates the ratio of the i-th singular value to the i+1-th singular value of the M singular values to obtain M-1 sound field classification parameters. Where i = 1, 2, …, M. Based on these M-1 sound field classification parameters, the encoder determines the number of dissimilar sound sources corresponding to the current frame. Where M = min(L, K), where L represents the number of channels of the HOA signal of the current frame, K represents the number of signal points in each channel of the HOA signal of the current frame, and min represents a minimum value operation. If the HOA signal is a time domain signal, the number of signal points is the number of sampling points; if the HOA signal is a frequency domain signal, the number of signal points is the number of frequency points.
[0184] Optionally, assuming that the M-1 sound field type parameters are temp[i], i = 0, 1, ..., M-2, the encoder determines the number of dissimilar sound sources corresponding to the current frame based on the M-1 sound field classification parameters. One implementation method is: starting from i = 0, the encoder sequentially executes the following process: determining whether temp[i] is greater than a preset dissimilar sound source determination threshold. If temp[i] is less than the dissimilar sound source determination threshold in this round of the process, the value of i is updated to i+1, and the next round of the process is continued. If temp[i] is greater than or equal to the dissimilar sound source determination threshold in this round of the process, the number of dissimilar sound sources corresponding to the current frame is determined to be i+1, and the process ends. Optionally, the dissimilar sound source determination threshold is 30, 80, or 100, etc. The dissimilar sound source determination threshold is a preset value, which can be preset based on experience or statistics.
[0185] Accordingly, in one implementation, after determining the number of dissimilar sound sources corresponding to the current frame, if the number of dissimilar sound sources corresponding to the current frame is greater than a first threshold and less than a second threshold, the encoder determines the initial encoding scheme for the current frame to be the second encoding scheme. If the number of dissimilar sound sources corresponding to the current frame is not greater than the first threshold or not less than the second threshold, the encoder determines the initial encoding scheme for the current frame to be the first encoding scheme. The first threshold is less than the second threshold. Optionally, the first threshold is 0 or another value, and the second threshold is 3 or another value. The first and second thresholds are preset values that can be preset based on experience or statistics.
[0186] Exemplarily, assume that the number of channels L of the HOA signal of the current frame is 16, the number of frequency points K of each channel is 8, and min(L, K) = 8. Then, the encoding end performs singular value decomposition on the HOA signal of the current frame to obtain singular values v[i], where i = 0, 1, …, min(L, K) - 1. The encoding end calculates the ratio between adjacent singular values and uses the obtained ratio as the sound field classification result temp[i] of the current frame, temp[i] = v[i] / v[i + 1], where i = 0, 1, …, min(L, K) - 2. Assume that the threshold for determining dissimilar sound sources is 100. The process of determining the number n of dissimilar sound sources is as follows: Starting from i = 0, judge whether temp[i] is greater than or equal to 100. If temp[i] is greater than or equal to 100, that is, temp[i] ≥ 100, then stop the judgment; otherwise, i = i + 1 and continue the judgment. If the judgment stops, the number i plus 1 when the judgment stops is equal to the number n of dissimilar sound sources corresponding to the current frame. For example, when i = 0, if temp[0] ≥ 100, then stop the judgment, and the number n of dissimilar sound sources is equal to 1; otherwise, let i = 1 and continue the judgment for i = 1; when i = 1, if temp[1] ≥ 100, then stop the judgment, and the number n of dissimilar sound sources is equal to i + 1 = 2. Assume that the first threshold is 0 and the second threshold is 3. Then, if the number n of dissimilar sound sources corresponding to the current frame satisfies 0 < n < 3, the encoding end determines that the initial encoding scheme of the current frame is the second encoding scheme. If the number n of dissimilar sound sources corresponding to the current frame satisfies n = 0 or n ≥ 3, the encoding end determines that the initial encoding scheme of the current frame is the first encoding scheme.
[0187] Optionally, the sound field classification result includes a sound field type, and the sound field type is divided into a diffuse sound field and a dissimilar sound field. The sound field type can be determined according to the number of dissimilar sound sources obtained by the foregoing method, that is, the encoding end determines the sound field type of the current frame based on the number of dissimilar sound sources corresponding to the current frame. For example, if the number of dissimilar sound sources corresponding to the current frame is greater than the first threshold and less than the second threshold, the encoding end determines that the sound field type of the current frame is a dissimilar sound field. If the number of dissimilar sound sources corresponding to the current frame is not greater than the first threshold or not less than the second threshold, the encoding end determines that the sound field type of the current frame is a diffuse sound field. Correspondingly, if the sound field type of the current frame is a dissimilar sound field, the encoding end determines that the initial encoding scheme of the current frame is the second encoding scheme, that is, the HOA encoding scheme based on MP. If the sound field type of the current frame is a diffuse sound field type, the encoding end determines that the initial encoding scheme of the current frame is the first encoding scheme, that is, the HOA encoding scheme based on DirAC.
[0188] In certain embodiments, after determining the initial coding scheme of each audio frame (including current frame) by the above-mentioned implementation method, the situation that the initial coding scheme of each audio frame switches back and forth may occur, that is, the switching frames that ultimately need to be encoded are more. Since the problems caused by the switching between the coding schemes are more, that is, the problems that need to be solved are more, the problems caused by the switching can be reduced by reducing the number of switching frames. In order to reduce the number of switching frames, the coding end can first determine the expected coding scheme of the current frame according to the sound field classification result of the current frame, that is, the coding end will use the initial coding scheme determined according to the aforementioned method as the expected coding scheme. Then, the coding end adopts the method of sliding window to update the initial coding scheme of the current frame based on the expected coding scheme, as the coding end updates the initial coding scheme of the current frame by hangover processing.
[0189] Optionally, assuming that the length of the sliding window is N, the sliding window contains the expected coding scheme of the current frame and the updated initial coding schemes of the N-1 frames before the current frame. If the cumulative number of second coding schemes in the sliding window is not less than the first specified threshold, the encoding end updates the initial coding scheme of the current frame to the second coding scheme. If the cumulative number of second coding schemes in the sliding window is less than the first specified threshold, the encoding end updates the initial coding scheme of the current frame to the first coding scheme. Among them, the length N of the sliding window is 8, 10, 15, etc., and the first specified threshold is 5, 6, 7, etc. The embodiment of the present application does not limit the value of the length of the sliding window and the first specified threshold. An example is given below. Assume that the length of the sliding window is 10, the first specified threshold is 7, and the sliding window contains the expected coding scheme of the current frame and the updated initial coding schemes of the 9 frames before the current frame. If the cumulative number of second coding schemes in the sliding window is not less than 7, the encoder determines the initial coding scheme of the current frame as the second coding scheme. If the cumulative number of second coding schemes in the sliding window is less than 7, the encoder updates the initial coding scheme of the current frame to the first coding scheme.
[0190] Alternatively, if the cumulative number of the first coding scheme within the sliding window is not less than a second specified threshold, the encoder updates the initial coding scheme for the current frame to the first coding scheme. If the cumulative number of the first coding scheme within the sliding window is less than a second specified threshold, the encoder updates the initial coding scheme for the current frame to the second coding scheme. The second specified threshold can be a value such as 5, 6, or 7, and the embodiment of the present application does not limit the value of the second specified threshold. Optionally, the second specified threshold is different from or the same as the first specified threshold.
[0191] In addition to the implementation methods described above, the encoding end may also adopt other methods to obtain the sound field classification result of the current frame, and the method for determining the initial encoding scheme based on the sound field classification result may also adopt other methods, which are not limited in the embodiments of the present application.
[0192] In an embodiment of the present application, after the encoding end determines the initial encoding scheme of the current frame, if the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame, the encoding end determines that the encoding scheme of the current frame is the initial encoding scheme of the current frame. If the initial encoding scheme of the current frame is different from the initial encoding scheme of the previous frame of the current frame, the encoding end determines that the encoding scheme of the current frame is the third encoding scheme. That is, if the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame and is the first encoding scheme, the encoding end determines that the encoding scheme of the current frame is the first encoding scheme. If the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame and is the second encoding scheme, the encoding end determines that the encoding scheme of the current frame is the second encoding scheme. If one of the initial encoding scheme of the current frame and the initial encoding scheme of the previous frame of the current frame is the first encoding scheme and the other is the second encoding scheme, the encoding end determines that the encoding scheme of the current frame is the third encoding scheme. Among them, one of the initial coding scheme of the current frame and the initial coding scheme of the frame before the current frame is the first coding scheme, and the other is the second coding scheme, that is, the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the frame before the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the frame before the current frame is the first coding scheme. In other words, for switching frames, the encoder does not use either the first coding scheme or the second coding scheme to encode the HOA signal of the switching frame, but instead uses the switching frame coding scheme to encode the HOA signal of the switching frame. For non-switching frames, the encoder uses a coding scheme consistent with the initial coding scheme of the non-switching frame to encode the HOA signal of the switching frame. Among them, an audio frame whose initial coding scheme is different from the initial coding scheme of the previous frame is a switching frame, and an audio frame whose initial coding scheme is the same as the initial coding scheme of the previous frame is a non-switching frame.
[0193] It should be noted that, in addition to determining the encoding scheme for the current frame, the encoder also needs to encode information indicating the encoding scheme for the current frame into the bitstream, so that the decoder can determine which decoding scheme to use to decode the bitstream for the current frame. In the embodiments of the present application, there are multiple ways to implement the encoder encode information indicating the encoding scheme for the current frame into the bitstream, three of which are described below.
[0194] The first implementation method , encoding switching flag and indication information of the two encoding schemes
[0195] In this implementation, the encoder needs to determine the value of the switching flag for the current frame and encode the value of the switching flag for the current frame into the bitstream. Specifically, when the encoding scheme for the current frame is the first encoding scheme or the second encoding scheme, the value of the switching flag for the current frame is the first value. When the encoding scheme for the current frame is the third encoding scheme, the value of the switching flag for the current frame is the second value. Optionally, the first value is "0" and the second value is "1." The first and second values may also be other values.
[0196] In addition, the encoder encodes information indicating the initial coding scheme for the current frame into the bitstream. Alternatively, if the value of the switching flag for the current frame is a first value, the encoder encodes information indicating the initial coding scheme for the current frame into the bitstream; if the value of the switching flag for the current frame is a second value, the encoder encodes preset indication information into the bitstream.
[0197] Optionally, the indication information of the initial coding scheme is represented by a coding mode (coding mode) corresponding to the initial coding scheme, that is, the coding mode is used as the indication information. For example, the coding mode corresponding to the initial coding scheme is the initial coding mode, and the initial coding mode is the first coding mode (i.e., DirAC mode) or the second coding mode (i.e., MP mode). Optionally, the preset indication information is a preset coding mode, and the preset coding mode is the first coding mode or the second coding mode. In some other embodiments, the preset indication information is other coding modes, that is, it is not limited to what the indication information of the coding scheme of the switching frame encoded into the code stream is specifically.
[0198] That is, in this first implementation, the encoder uses a switching flag to indicate a switching frame, and the coding scheme information for the switching frame encoded into the bitstream is not limited. The coding scheme information for the switching frame can be an initial coding mode, a preset coding mode, a random selection from the first coding mode and the second coding mode, or other indication information. It should be noted that in this implementation, the switching flag is used to indicate whether the current frame is a switching frame. In this way, the decoder can directly determine whether the current frame is a switching frame by obtaining the switching flag in the bitstream.
[0199] Optionally, in this first implementation, the switching flag of the current frame and the indication information of the initial coding scheme each occupy one bit of the code stream. Exemplarily, the value of the switching flag of the current frame is "0" or "1", wherein the value of the switching flag is "0" indicating that the current frame is not a switching frame, that is, the value of the switching flag of the current frame is the first value. The switching flag is "1" indicating that the current frame is a switching frame, that is, the value of the switching flag of the current frame is the second value. Optionally, the indication information of the initial coding scheme is "0" or "1", wherein "0" indicates DirAC mode (that is, DirAC coding scheme) and "1" indicates MP mode (that is, MP-based coding scheme).
[0200] In some other embodiments, if the initial coding scheme of the current frame is different from the initial coding scheme of the frame immediately preceding the current frame, the encoder determines that the value of the switching flag for the current frame is a second value and encodes the value of the switching flag for the current frame into the bitstream. In other words, for a switching frame, since the switching flag in the bitstream already indicates the switching frame, there is no need to encode information indicating the coding scheme of the switching frame.
[0201] The second implementation method , encoding instructions for two encoding schemes
[0202] In this implementation, the encoder encodes information indicating the initial coding scheme for the current frame into the bitstream. Taking the coding mode as an example, the information encoded into the bitstream is essentially the coding mode consistent with the initial coding scheme, namely the initial coding mode, which can be either the first coding mode or the second coding mode. Furthermore, the encoder may not encode a switching flag.
[0203] Optionally, in this first implementation, the indication information of the initial coding scheme occupies one bit of the bitstream. For example, taking the coding mode as the indication information, the coding mode encoded into the bitstream is "0" or "1," where "0" indicates DirAC mode, indicating that the initial coding scheme of the current frame is the first coding scheme, and "1" indicates MP mode, indicating that the initial coding scheme of the current frame is the second coding scheme.
[0204] The third implementation method , encoding instructions for three encoding schemes
[0205] In this implementation, the encoder encodes information indicating the current frame's coding scheme into the bitstream. Taking the coding mode as an example, the indication information encoded into the bitstream is essentially the coding mode consistent with the current frame's coding scheme. The coding mode consistent with the current frame's coding scheme is the actual coding mode, which is the first coding mode, the second coding mode, or the third coding mode. Optionally, the third coding mode is the MP-W mode.
[0206] Optionally, in this third implementation, the information indicating the coding scheme of the current frame occupies two bits of the code stream. Exemplarily, the information indicating the coding scheme of the current frame is "00", "01", or "10". Among them, "00" indicates that the coding scheme of the current frame is the first coding scheme, "01" indicates that the coding scheme of the current frame is the second coding scheme, and "10" indicates that the coding scheme of the current frame is the third coding scheme.
[0207] As can be seen from the above, in the first implementation method mentioned above, after the encoding end determines the initial coding scheme of the current frame, it determines the value of the switching flag and encodes the value of the switching flag into the code stream. In addition, the indication information of the initial coding scheme of the current frame is encoded into the code stream, or if the current frame is a switching frame, the encoding end encodes the preset indication information into the code stream, and if the current frame is a non-switching frame, the encoding end encodes the indication information of the initial coding scheme of the current frame into the code stream. In the second implementation method mentioned above, after the encoding end determines the initial coding scheme of the current frame, it directly encodes the indication information of the initial coding scheme of the current frame into the code stream. In the third implementation method mentioned above, after the encoding end determines the initial coding scheme of the current frame, it determines the coding scheme of the current frame based on the initial coding scheme of the current frame and the initial coding scheme of the previous frame of the current frame, and encodes the indication information of the coding scheme of the current frame into the code stream.
[0208] Step 602: If the coding scheme of the current frame is the third coding scheme, signals of designated channels in the HOA signal are encoded into the bitstream, where the designated channels are some of all channels in the HOA signal.
[0209] In an embodiment of the present application, if the coding scheme of the current frame is the third coding scheme, indicating that the current frame is a switching frame, the encoding end encodes the HOA signal of the current frame according to the third coding scheme (i.e., the hybrid coding scheme). Corresponding to the first implementation method in the above step 601, if the value of the switching flag of the current frame is the second value, it indicates that the current frame is a switching frame. Corresponding to the second implementation method in the above step 601, if the initial coding scheme of the current frame is different from the initial coding scheme of the previous frame of the current frame, it indicates that the current frame is a switching frame. Corresponding to the third implementation method in the above step 601, if the coding scheme of the current frame is the third coding scheme, the coding scheme of the current frame indicates that the current frame is a switching frame. For the switching frame, the encoding end uses the third coding scheme to encode the HOA signal of the current frame. Among them, the third coding scheme indicates that the signal of the specified channel in the HOA signal of the current frame is encoded into the code stream, wherein the specified channel is a portion of all channels of the HOA signal. That is, for the switching frame, the encoding end encodes the signal of the specified channel in the HOA signal of the switching frame into the bitstream, rather than using the first encoding scheme or the second encoding scheme to encode the switching frame. That is, in order to ensure a smooth transition of the auditory quality when the encoding scheme is switched, this scheme adopts a compromise method to encode the switching frame.
[0210] Optionally, the designated channel is consistent with the transmission channel preset in the first coding scheme, that is, the designated channel is the preset channel. That is, under the premise that the third coding scheme is different from the second coding scheme, in order to make the coding effect of the third coding scheme close to that of the second coding scheme, the encoding end encodes the signal of the channel that is the same as the transmission channel preset in the first coding scheme in the HOA signal of the switching frame into the bit stream, so that the auditory quality transitions as smoothly as possible. It should be noted that different transmission channels can be preset according to different coding bandwidths, bit rates, and even different application scenarios. Optionally, the preset transmission channels can also be the same under different coding bandwidths, bit rates or application scenarios.
[0211] Optionally, the designated channel signal includes a FOA signal, which includes an omnidirectional W signal and directional X, Y, and Z signals. That is, the designated channel includes the FOA channel, and the FOA channel signal is a low-order signal. That is, if the current frame is a switching frame, the encoder encodes the low-order portion of the HOA signal of the current frame into the bitstream. The low-order portion includes the W, X, Y, and Z signals of the FOA channel.
[0212] It should be noted that in the embodiment of the present application, there are many ways to implement the encoding end to encode the signal of the specified channel in the HOA signal into the bitstream, and it is sufficient to encode the signal of the specified channel into the bitstream. The following describes some of these implementation methods.
[0213] In an embodiment of the present application, if the designated channel includes a FOA channel, the encoder determines a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal, and encodes the virtual speaker signal and the residual signal into a bitstream.
[0214] Optionally, the encoder determines the W signal as a virtual speaker signal and determines three residual signals based on the W signal, the X signal, the Y signal, and the Z signal, or alternatively, determines the X signal, the Y signal, and the Z signal as the three residual signals. Optionally, the encoder determines the difference signals between any three of the W signal, the X signal, the Y signal, and the Z signal and the remaining signal as the three residual signals. For example, the encoder determines the difference signals between the X signal, the Y signal, and the Z signal and the W signal as the three residual signals. Exemplarily, the encoder uses the difference signals X', Y', and Z' obtained from XW, YW, and ZW, respectively, as the three residual signals.
[0215] If the encoding end uses a core encoder to encode the current frame, the core encoder is a stereo encoder. Since the determined virtual speaker signal and three residual signals are all mono signals, the encoding end needs to first combine a stereo signal based on these mono signals, and then use the stereo encoder for encoding. Optionally, the encoding end combines the one virtual speaker signal with the first preset mono signal to obtain one stereo signal, and combines the three residual signals with the second preset mono signal to obtain two stereo signals. The encoding end encodes the obtained three stereo signals into the bitstream respectively through the stereo encoder.
[0216] Among them, the embodiment of the present application does not limit the specific combination method of the encoding end combining the three residual signals with a preset mono signal to obtain two stereo signals. Optionally, the encoding end combines the two residual signals with the highest correlation among the three residual signals to obtain one stereo signal in the two stereo signals, and combines the residual signal other than the two residual signals with the highest correlation among the three residual signals with the second preset mono signal to obtain the other stereo signal in the two stereo signals. That is, the encoding end combines and obtains the stereo signal according to the correlation of the signals. In some other embodiments, the encoding end may also combine any two residual signals among the three residual signals to obtain one stereo signal in the two stereo signals, and combine the remaining residual signal with the second preset mono signal to obtain the other stereo signal in the two stereo signals.
[0217] Optionally, in the embodiment of the present application, the first preset mono signal is an all-zero signal or an all-one signal, and the second preset mono signal is an all-zero signal or an all-one signal. Optionally, the first preset mono signal and the second preset mono signal are the same or different, that is, the first preset mono signal and the second preset mono signal are both all-zero signals or all-one signals, or the first preset mono signal is an all-zero signal and the second preset mono signal is an all-one signal, or the first preset mono signal is an all-one signal and the second preset mono signal is an all-zero signal. Wherein, the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency values are all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency values are all one. Wherein, if the HOA signal is a time domain signal, the all-zero signal includes a signal whose sampling point values are all zero, and the all-one signal includes a signal whose sampling point values are all one. If the HOA signal is a frequency domain signal, the all-zero signal includes a signal whose frequency values are all zero, and the all-one signal includes a signal whose frequency values are all one. In some other embodiments, the first preset mono signal and / or the second preset mono signal may also be other preset signals.
[0218] If the core encoder used by the encoding end is a mono encoder, the encoding end encodes the virtual speaker signal and each of the three residual signals into a bitstream through the mono encoder.
[0219] Figure 7 This is a schematic diagram of a switching frame coding scheme provided by an embodiment of the present application. Figure 7 The current frame to be encoded is a switching frame. The encoder obtains the HOA signal of the current frame, uses the W signal in the HOA signal as the virtual speaker signal, and determines a residual signal based on the FOA signal in the HOA signal. For example, the residual signal is determined based on the X, Y, and Z signals in the HOA signal, or based on the W signal and the X, Y, and Z signals. The encoder encodes the determined virtual speaker signal and residual signal into the bitstream through the core encoder to obtain the bitstream of the switching frame.
[0220] Alternatively, in other embodiments, the encoder determines two of the W, X, Y, and Z signals as two virtual speaker signals, and determines the remaining two signals as two residual signals. The encoder combines the two virtual speaker signals to obtain one stereo signal, and combines the two residual signals to obtain another stereo signal. The encoder encodes the obtained two stereo signals into a bitstream using a stereo encoder.
[0221] The embodiments of the present application do not limit the specific manner in which the encoder combines the W signal, X signal, Y signal, and Z signal in pairs to obtain two stereo signals. Optionally, the encoder determines the W signal as one virtual speaker signal and determines the signal with the highest correlation with the W signal among the X signal, Y signal, and Z signal as another virtual speaker signal. Specifically, the encoder combines the W signal and the signal with the highest correlation with the W signal among the four signals included in the FOA channel, and then combines the remaining two signals. Alternatively, the encoder combines any two signals among the W signal, X signal, Y signal, and Z signal to obtain one stereo signal and combines the remaining two signals to obtain another stereo signal.
[0222] It should be noted that the embodiment of the present application does not limit the specific implementation method of encoding the virtual speaker signal and the residual signal using the core encoder at the encoding end, for example, does not limit the number of encoding bits corresponding to the virtual speaker signal and the residual signal respectively.
[0223] The above describes the process by which the encoder encodes the current frame when the current frame is a switching frame. Specifically, the encoder encodes the signal of a specified channel in the HOA signal of the switching frame into the bitstream according to the third encoding scheme. The third encoding scheme is the switching frame encoding scheme. As can be seen from the above, in this embodiment of the present application, the signal of the specified channel may include a W signal, which is a core signal of the HOA signal. Thus, the switching frame encoding scheme may also be referred to as an MP-W-based encoding scheme. Next, the process by which the encoder encodes the current frame when the current frame is a non-switching frame is described.
[0224] In this embodiment of the present application, if the current frame's coding scheme is the first coding scheme, the encoder encodes the HOA signal of the current frame into the bitstream according to the first coding scheme. If the current frame's coding scheme is the second coding scheme, the encoder encodes the HOA signal of the current frame into the bitstream according to the second coding scheme. In other words, if the current frame is not a switching frame, the encoder encodes the current frame using the initial coding scheme of the current frame.
[0225] For example, see Figure 8 The encoder encodes the HOA signal of the current frame into the bitstream according to the second coding scheme as follows: the encoder selects a target virtual speaker that matches the HOA signal of the current frame from a set of virtual speakers based on the MP algorithm; determines a virtual speaker signal based on the HOA signal of the current frame and the target virtual speaker using the MP-based spatial encoder; determines a residual signal based on the HOA signal of the current frame and the virtual speaker signal using the MP-based spatial encoder; and encodes the virtual speaker signal and the residual signal into the bitstream using the core encoder. It should be noted that the MP-based HOA coding scheme and the switching frame coding scheme differ in the principles and specific methods for determining the virtual speaker signal and the residual signal, and the virtual speaker signal and the residual signal determined by the two schemes are also different. For the same frame, the MP-based HOA coding scheme encodes more effective information into the bitstream than the switching frame coding scheme. However, under the premise that the switching frame coding scheme differs from the second coding scheme, in order to achieve similar coding effects, the switching frame coding scheme also encodes the virtual speaker signal and the residual signal into the bitstream, thereby ensuring the smoothest possible transition in auditory quality.
[0226] The implementation process of encoding the HOA signal of the current frame into the bitstream according to the first encoding scheme is as follows: the encoding end extracts the core layer signal and spatial parameters from the HOA signal of the current frame, and encodes the extracted core layer signal and spatial parameters into the bitstream. For example, see Figure 9The encoding end extracts the core layer signal from the HOA signal of the current frame through the core coding signal acquisition module, extracts the spatial parameters from the HOA signal of the current frame through the DirAC-based spatial parameter extraction module, encodes the core layer signal into the bitstream through the core encoder, and encodes the spatial parameters into the bitstream through the spatial parameter encoder. The channel corresponding to the core layer signal is consistent with the specified channel in this scheme. In addition, in addition to encoding the core layer signal into the bitstream, the first encoding scheme also encodes the extracted spatial parameters into the bitstream. The spatial parameters contain rich scene information, such as direction information. It can be seen that for the same frame, the effective information encoded into the bitstream using the DirAC-based HOA coding scheme will be more than the effective information encoded into the bitstream using the switching frame coding scheme. Under the premise that the switching frame coding scheme is different from the first coding scheme, in order to make the coding effect of the switching frame coding scheme close to that of the first coding scheme, the switching frame coding scheme also encodes the signal of the transmission channel preset by the first coding scheme in the HOA signal into the bitstream, but will not encode more information in the HOA signal except the signal of the specified channel into the bitstream, that is, it will not extract spatial parameters, and will not encode spatial parameters into the bitstream, so that the auditory quality transitions as smoothly as possible.
[0227] Figure 10 This is a flowchart of another encoding method provided by the embodiment of this application. Please refer to Figure 10 , taking the encoding of the indication information of the initial coding scheme of the current frame into the bitstream as an example, the encoding method provided in the embodiment of the present application is explained again. The encoding end first obtains the HOA signal of the current frame to be encoded. Then, the encoding end performs sound field type analysis on the HOA signal to determine the initial coding scheme of the current frame, and the encoding end encodes the indication information of the initial coding scheme of the current frame into the bitstream. The encoding end determines whether the initial coding scheme of the current frame is the same as the initial coding scheme of the previous frame. If the initial coding scheme of the current frame is the same as the initial coding scheme of the previous frame, the encoding end adopts the initial coding scheme of the current frame to encode the HOA signal of the current frame to obtain the bitstream of the current frame. If the initial coding scheme of the current frame is different from the initial coding scheme of the previous frame, the encoding end adopts the switching frame coding scheme to encode the HOA signal of the current frame to obtain the bitstream of the current frame.
[0228] It should be noted that if the current frame is the first audio frame to be encoded, the initial encoding scheme of the current frame is the first encoding scheme or the second encoding scheme, and the encoding end uses the initial encoding scheme of the current frame to encode the HOA signal of the current frame into the bitstream.
[0229] In summary, in an embodiment of the present application, two schemes (i.e., a coding scheme based on virtual speaker selection and a coding scheme based on directional audio coding) are combined to encode and decode the HOA signal of the audio frame, that is, a suitable coding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to ensure a smooth transition of auditory quality when switching between different coding schemes, for some audio frames in this scheme, instead of directly using any of the above two schemes for encoding, a new coding scheme is used to encode and decode these audio frames, that is, the signals of the specified channels in the HOA signals of these audio frames are encoded into the bitstream, that is, a compromise scheme is used for encoding and decoding, so that the auditory quality after rendering and playback of the decoded and restored HOA signals can be smoothly transitioned.
[0230] Figure 11 This is a flowchart of a decoding method provided by an embodiment of the present application, which is applied to a decoding end. It should be noted that the decoding method corresponds to Figure 6 Please refer to the encoding method shown in Figure 11 , the method includes the following steps.
[0231] Step 1101: Obtain a decoding solution for the current frame based on the code stream.
[0232] The decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme. The first decoding scheme is a DirAC-based HOA decoding scheme, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme. Optionally, the hybrid decoding scheme is also called a switching frame decoding scheme.
[0233] It should be noted that, since the encoding end uses different encoding schemes to encode different audio frames, the decoding end also needs to use the corresponding decoding scheme to decode each audio frame.
[0234] Next, we will first introduce how the decoding end determines the encoding scheme of the current frame. Figure 6 Step 601 of the encoding method shown introduces three implementation methods for the encoder to encode information that can be used to indicate the encoding scheme of the current frame into the bitstream. Correspondingly, there are also three corresponding implementation methods for the decoder to determine the encoding scheme of the current frame, which will be introduced below.
[0235] The first implementation method , encoding the switching flag and the indication information of the two encoding schemes
[0236] The decoding end first parses the value of the switching flag for the current frame from the bitstream. If the value of the switching flag is the first value, the decoding end then parses the bitstream to obtain indication information of the decoding scheme for the current frame. This indication information is used to indicate whether the decoding scheme for the current frame is the first decoding scheme or the second decoding scheme. If the value of the switching flag is the second value, the decoding end determines that the decoding scheme for the current frame is the third decoding scheme. It should be noted that the indication information of the encoding scheme encoded into the bitstream by the encoding end is the indication information of the decoding scheme parsed from the bitstream by the decoding end.
[0237] In other words, if the decoder parses the current frame's switching flag and finds the value to be the first, the current frame is a non-switching frame. The decoder then parses the bitstream to find the decoding scheme indication information and determines the decoding scheme for the current frame based on the indication information. If the decoder parses the current frame's switching flag and finds the value to be the second, the current frame is a switching frame. Even if the bitstream contains the indication information, the decoder does not need to decode the indication information.
[0238] It should be noted that if the value of the switching flag is the second value, the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme, and the current frame is a switching frame. The switching frame decoding scheme is a decoding scheme different from the first decoding scheme and the second decoding scheme. The switching frame decoding scheme is for a smooth transition of auditory quality.
[0239] Optionally, in this first implementation, the indication information of the decoding scheme and the switching flag each occupy one bit of the code stream. Exemplarily, the decoding end first parses the value of the switching flag of the current frame from the code stream. If the value of the parsed switching flag is "0", that is, the value of the switching flag is the first value, then the decoding end parses the indication information of the decoding scheme of the current frame from the code stream. If the parsed indication information is "0", the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is "1", the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed switching flag is the value "1", the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme (the third decoding scheme).
[0240] The second implementation method , encoding the indication information of two encoding schemes
[0241] The decoding end parses the initial decoding scheme of the current frame from the code stream, and the initial decoding scheme is the first decoding scheme or the second decoding scheme. If the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame of the current frame, the decoding scheme of the current frame is determined to be the initial decoding scheme of the current frame. If the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame of the current frame, the decoding scheme of the current frame is determined to be the third decoding scheme, that is, the mixed decoding scheme. Among them, the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame of the current frame, which means that the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme. That is, one of the initial decoding scheme of the current frame and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, and the other is the second decoding scheme.
[0242] Optionally, in this second implementation, the indication information for indicating the initial coding scheme occupies one bit of the code stream. Taking the coding mode as the indication information as an example, the coding mode in the code stream occupies one bit. Exemplarily, the decoding end parses the indication information of the initial coding scheme of the current frame from the code stream. If the parsed indication information is "0" and the indication information of the previous frame of the current frame is also "0", the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is "1" and the indication information of the previous frame of the current frame is also "1", the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed indication information is "0" and the indication information of the previous frame of the current frame is "1", or the parsed indication information is "1" and the indication information of the previous frame of the current frame is "0", the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme.
[0243] Optionally, the indication information of the initial decoding scheme of the frame before the current frame is cached data. When decoding the current frame, the decoding end can obtain the indication information of the initial decoding scheme of the frame before the current frame from the cache.
[0244] The third implementation method , encoding the indication information of three encoding schemes
[0245] The decoding end parses the code stream to obtain indication information of a decoding scheme for the current frame, where the indication information is used to indicate whether the decoding scheme for the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
[0246] Optionally, in the third implementation, the indication information of the decoding scheme occupies two bits of the code stream. For example, assuming that the coding mode is used as the indication information, the coding mode of the current frame occupies two bits of the code stream.
[0247] Exemplarily, the decoding end parses the indication information of the decoding scheme of the current frame from the bitstream. If the parsed indication information is "00", the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is "01", the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed indication information is "10", the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme.
[0248] Step 1102: If the decoding scheme of the current frame is the third decoding scheme, determine the signals of the designated channels in the HOA signal of the current frame based on the code stream, where the designated channels are some channels in all channels of the HOA signal.
[0249] In this embodiment of the present application, after the decoder obtains the decoding scheme for the current frame, if the decoding scheme for the current frame is the third encoding scheme, indicating that the current frame is a handover frame, the decoder then determines the signal of the designated channel in the HOA signal of the current frame based on the bitstream. In other words, for handover frames, the encoder encodes the signal of the designated channel into the bitstream, so the decoder uses the handover frame decoding scheme to decode the handover frame, which means that the signal of the designated channel must first be parsed from the bitstream.
[0250] Next, the implementation process of decoding the switching frame using the switching frame decoding scheme at the decoding end is described in detail. That is, when the current frame is a switching frame, the implementation process of the decoding end determining the signal of the specified channel in the HOA signal of the current frame based on the code stream is described in detail.
[0251] It should be noted that the process by which the decoder determines the signal of a specific channel in the HOA signal of the current frame based on the bitstream is symmetrical to the process by which the encoder encodes the signal of the specific channel in the HOA signal of the current frame into the bitstream. The aforementioned encoding method embodiments describe some implementations for encoding the signal of the specific channel into the bitstream. The decoding process on the decoder is symmetrical to these implementations.
[0252] In an embodiment of the present application, if the encoding end first determines the virtual speaker signal and the residual signal based on the signal of the specified channel, and then encodes the virtual speaker signal and the residual signal into the bitstream, then, correspondingly, the decoding end first determines the virtual speaker signal and the residual signal based on the bitstream, and then determines the signal of the specified channel based on the virtual speaker signal and the residual signal.
[0253] Optionally, if the encoder encodes three stereo signals obtained by combining virtual speaker signals and residual signals into a bitstream using a stereo encoder, the decoder decodes the bitstream using a stereo decoder to obtain three stereo signals, and then determines one virtual speaker signal and three residual signals based on the three stereo signals. Optionally, the decoder determines one virtual speaker signal based on one stereo signal from the three stereo signals, and determines three residual signals based on the other two stereo signals from the three stereo signals. That is, the decoder first parses the three stereo signals from the bitstream, and then decomposes the three stereo signals to obtain one virtual speaker signal and three residual signals.
[0254] For example, the decoder parses the bitstream to generate three stereo signals, S1, S2, and S3. S1 is a combination of a virtual speaker signal and a preset mono signal, S2 is a combination of two residual signals, and S3 is a combination of the remaining residual signal and a preset mono signal. The decoder decomposes S1 into a virtual speaker signal, S2 into two residual signals, and S3 into the remaining residual signal.
[0255] Optionally, if the encoding end encodes four mono signals determined based on the virtual speaker signal and the residual signal into a bitstream through a mono encoder, then the decoding end decodes the bitstream through a mono decoder to obtain one virtual speaker signal and three residual signals, and the four mono signals include the one virtual speaker signal and the three residual signals.
[0256] Optionally, if the signal of the designated channel includes a FOA signal, and the FOA signal includes an omnidirectional W signal and directional X, Y, and Z signals, then the decoder determines the W signal based on the virtual speaker signal after determining the virtual speaker signal and the residual signal based on the bitstream. The decoder determines the X, Y, and Z signals based on the residual signal and the W signal, or alternatively, the decoder determines the X, Y, and Z signals based on the residual signal. For example, if the decoder parses three residual signals, the sum of the three residual signals and the W signal is determined as the X, Y, and Z signals, or alternatively, the three residual signals are determined as the X, Y, and Z signals. If the encoder determines the difference signals between the X, Y, and Z signals and the W signal as three residual signals, the decoder determines the sum of the three residual signals and the W signal as the X, Y, and Z signals. If the encoder determines the X signal, Y signal, and Z signal as three residual signals, the decoder determines the three residual signals as the X signal, Y signal, and Z signal, respectively. That is, the decoding process of the decoder matches the encoding process of the encoder.
[0257] If the encoder encodes two stereo signals determined based on virtual speaker signals and residual signals into a bitstream using a stereo encoder, the decoder decodes the bitstream using a stereo decoder to obtain the two stereo signals. The decoder determines two virtual speaker signals based on one of the two stereo signals and two residual signals based on the other of the two stereo signals. The two virtual speaker signals and the two residual signals include the W signal, the X signal, the Y signal, and the Z signal. Optionally, if the encoder determines the W signal and the signal with the highest correlation with the W signal among the X signal, the Y signal, and the Z signal as the two virtual speaker signals, the two virtual speaker signals determined by the decoder include the W signal and the signal with the highest correlation with the W signal among the X signal, the Y signal, and the Z signal. Assuming that the signal with the highest correlation with the W signal among the X signal, the Y signal, and the Z signal is the X signal, the two virtual speaker signals determined by the decoder include the W signal and the X signal, and the two residual signals determined by the decoder include the Y signal and the Z signal.
[0258] Step 1103: Based on the signal of the designated channel, determine the gains of one or more remaining channels except the designated channel in the HOA signal of the current frame.
[0259] In an embodiment of the present application, after the decoding end determines the signal of the designated channel in the HOA signal of the current frame based on the code stream, it determines the gain of one or more remaining channels in the HOA signal except the designated channel based on the signal of the designated channel.
[0260] Exemplarily, assuming that the designated channel is the FOA channel, the FOA channel can be called a low-order channel, the signal of the FOA channel can be called the low-order part of the HOA signal, and one or more remaining channels in the HOA signal other than the designated channel are called high-order channels, and the signals of the high-order channels can be called the high-order part of the HOA signal. Then, the decoding end determines the high-order gain of the HOA signal, that is, the gain of the high-order channel, based on the low-order part of the HOA signal.
[0261] Optionally, the decoding end first performs analysis filtering on the signal of the designated channel in the HOA signal to obtain the signal of the designated channel that has been analyzed and filtered, and determines the gain of the one or more remaining channels based on the signal of the designated channel that has been analyzed and filtered. For example, assuming that the signal of the designated channel is the low-order part of the HOA signal, the decoding end first performs analysis filtering on the low-order part of the HOA signal to obtain the low-order part of the HOA signal that has been analyzed and filtered, and then estimates the high-order gain based on the low-order part of the HOA signal that has been analyzed and filtered. Optionally, in this solution, for the switching frame, the analysis filter used by the decoding end for analysis filtering is the same as the analysis filter used in the HOA decoding solution based on DirAC, so that the decoding delay of the switching frame can be consistent with the decoding delay of the HOA decoding solution based on DirAC, that is, the delay is aligned. It should be noted that the decoding delay mentioned in this article is the end-to-end encoding and decoding delay, and the decoding delay can also be called the encoding delay.
[0262] It should be noted that in the embodiment of the present application, the process by which the decoding end determines the gain of one or more remaining channels in the HOA signal other than the designated channel based on the signal of the designated channel, i.e., the process of estimating the gain of the remaining channels based on the signal of the designated channel, is specifically implemented in the same manner as the remaining channel gain estimation method in the DirAC-based coding and decoding scheme, and is not described in detail in the embodiment of the present application. For example, in this scheme, for switching frames, the method by which the decoding end estimates the high-order gain based on the low-order part of the HOA signal is the same as the high-order gain estimation method in the DirAC-based coding and decoding scheme.
[0263] Step 1104 : Determine the signal of each of the one or more remaining channels based on the signal of the designated channel and the gains of the one or more remaining channels.
[0264] In an embodiment of the present application, the decoder determines the signal of each of the one or more remaining channels based on the signal of the designated channel and the gain of the one or more remaining channels. For example, assuming that the signal of the designated channel is the low-order portion of the HOA signal and the gain of the one or more remaining channels is the high-order gain, the decoder can determine the high-order portion of the HOA signal based on the W signal and the high-order gain in the low-order portion. Alternatively, if the decoder performs analysis filtering on the low-order portion of the HOA signal, the decoder can determine the high-order portion of the HOA signal after analysis filtering based on the W signal and the high-order gain in the low-order portion of the HOA signal after analysis filtering.
[0265] Step 1105: Obtain a reconstructed HOA signal of the current frame based on the signal of the designated channel and the signals of the one or more remaining channels.
[0266] In an embodiment of the present application, after obtaining the signal of the designated channel and the signals of the one or more remaining channels, the decoder obtains a reconstructed HOA signal for the current frame based on the signal of the designated channel and the signals of the one or more remaining channels, that is, reconstructs the HOA signal for the current frame. Exemplarily, the decoder performs synthesis filtering on the signal of the designated channel and the signals of the one or more remaining channels to obtain the reconstructed HOA signal for the current frame. For example, assuming that the signal of the designated channel is the low-order portion of the HOA signal and the signals of the one or more remaining channels are the high-order portion of the HOA signal, the decoder performs synthesis filtering on the low-order portion and the high-order portion of the HOA signal to obtain the reconstructed HOA signal for the current frame. Alternatively, if the decoder performs analysis filtering on the low-order portion of the HOA signal, the decoder performs synthesis filtering on the low-order portion of the analysis-filtered HOA signal and the high-order portion of the analysis-filtered HOA signal to obtain the reconstructed HOA signal for the current frame. Optionally, in this solution, for the switching frame, the synthesis filter used by the decoding end for synthesis filtering processing is the same as the synthesis filter used in the DirAC-based HOA encoding and decoding solution. This can make the decoding delay of the switching frame consistent with the decoding delay of the DirAC-based HOA decoding solution, that is, the delay is aligned.
[0267] Figure 12 Schematic diagram of a switching frame decoding solution provided by an embodiment of the present application. Figure 12 The current frame to be decoded is a switching frame. Assuming that the signal of the designated channel is the low-order portion of the HOA signal, during the decoding process, the decoder obtains the bitstream of the current frame to be decoded and performs core decoding on the bitstream through the core decoder to reconstruct the low-order portion of the HOA signal of the current frame. Based on this low-order portion, a high-order portion is estimated using a method similar to that used to determine the high-order portion in the DirAC-based HOA decoding scheme, thereby reconstructing the high-order portion of the HOA signal. The decoder then reconstructs the HOA signal based on the decoded low-order portion and the estimated high-order portion.
[0268] The above describes the decoding process for the current frame when it is a handover frame. Specifically, the decoder uses a handover frame decoding scheme to decode handover frames. Specifically, the decoder first decodes the signal of the specified channel in the HOA signal (e.g., the low-order portion) and then reconstructs the signals of the remaining channels (e.g., the high-order portion). Next, we will describe the decoding process for the current frame when it is not a handover frame.
[0269] In this embodiment of the present application, after the decoding end determines the decoding scheme for the current frame, if the decoding scheme for the current frame is the first decoding scheme, the decoding end obtains the reconstructed HOA signal for the current frame based on the bitstream according to the first decoding scheme. If the decoding scheme for the current frame is the second decoding scheme, the decoding end obtains the reconstructed HOA signal for the current frame based on the bitstream according to the second decoding scheme.
[0270] In the examples of this application, see Figure 13 The decoding end uses the second decoding scheme to obtain the reconstructed HOA signal of the current frame based on the bitstream. The decoding end uses the core decoder to parse the bitstream to obtain the virtual speaker signal and residual signal, and then feeds the parsed virtual speaker signal and residual signal into the MP-based spatial decoder to obtain the reconstructed HOA signal of the current frame. It should be noted that Figure 13 The decoding scheme shown is the same as Figure 8 The encoding scheme shown corresponds to .
[0271] The decoding end obtains the reconstructed HOA signal of the current frame according to the bitstream according to the first decoding scheme as follows: the decoding end parses the core layer signal and spatial parameters from the bitstream, and reconstructs the HOA signal of the current frame based on the core layer signal and spatial parameters. For example, see Figure 14 The decoding end parses the core layer signal from the bitstream through the core decoder, parses the spatial parameters from the bitstream through the spatial parameter decoder, and performs DirAC-based HOA signal synthesis processing based on the parsed core layer signal and spatial parameters to obtain the reconstructed HOA signal of the current frame. It should be noted that Figure 14 The decoding scheme shown is the same as Figure 9 The encoding scheme shown corresponds to .
[0272] Optionally, since the high-order portion of the HOA signal has a significant impact on auditory quality, in order to further ensure a smooth transition in auditory quality when switching between different coding and decoding schemes, the decoder may further perform gain adjustment on the high-order portion of the current frame while obtaining the reconstructed HOA signal of the current frame based on the bitstream according to the second decoding scheme. For example, the decoder obtains the initial HOA signal based on the bitstream according to the second decoding scheme. If the decoding scheme of the frame preceding the current frame is the third decoding scheme, that is, the frame preceding the current frame is a switching frame, the decoder performs gain adjustment on the high-order portion of the initial HOA signal based on the high-order gain of the frame preceding the current frame. The decoder then obtains the reconstructed HOA signal of the current frame based on the low-order portion of the initial HOA signal and the gain-adjusted high-order portion.
[0273] It should be noted that if the previous frame of the current frame is a switching frame, the current frame uses the high-order gain of the previous frame to adjust the gain of the high-order portion of the initial HOA signal of the current frame, so that the gain-adjusted high-order portion of the current frame is similar to the high-order portion of the previous frame. For example, the gain adjustment makes the energy of the high-order portions of the HOA signals of these two adjacent frames similar. In this way, during the subsequent rendering and playback of each audio frame by the decoding end, the auditory quality of the switching frame and the auditory quality of the next frame after the switching frame can both be smoothly transitioned.
[0274] Optionally, in addition to performing high-order gain adjustment on the audio frames whose decoding scheme is the second decoding scheme after the switching frame, for other audio frames whose decoding scheme is the second decoding scheme, the decoding end may also perform gain adjustment on the high-order parts of the HOA signals of these audio frames. The embodiment of the present application does not limit the specific implementation method of performing gain adjustment on the high-order parts of the HOA signals of these audio frames. Optionally, in addition to performing gain adjustment on the high-order parts, the decoding end may also perform gain adjustment on other parts of the HOA signals of these audio frames. That is, the embodiment of the present application does not limit which channels of the HOA signal are gain-adjusted. In other words, the decoding end may perform gain adjustment on the signals of any one or more channels in the HOA signal, and the one or more channels may include part or all of the high-order channels, or part or all of the remaining channels except the specified channels, or other channels.
[0275] Figure 15 This is a flowchart of another decoding method provided by an embodiment of the present application. Figure 15 , taking the example of the encoding end encoding the indication information of the initial encoding scheme into the bitstream, and assuming that the switching flag is not encoded in the bitstream, during the decoding process, the decoding end first parses the indication information of the initial decoding scheme of the current frame from the bitstream. Then, the decoding end determines whether the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame. If the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame, it means that the current frame is a non-switching frame, and the decoding end uses the initial decoding scheme of the current frame to decode the bitstream to obtain the reconstructed HOA signal of the current frame. If the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame, it means that the current frame is a switching frame, and the decoding end uses the switching frame decoding scheme to decode the bitstream to obtain the reconstructed HOA signal of the current frame.
[0276] In summary, in an embodiment of the present application, two schemes (i.e., a coding scheme based on virtual speaker selection and a coding scheme based on directional audio coding) are combined to encode and decode the HOA signal of the audio frame, that is, a suitable coding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to achieve a smooth transition of auditory quality when switching between different coding schemes, for some audio frames in this scheme, instead of directly adopting any of the above two schemes for coding and decoding, a new coding scheme is adopted to encode and decode these audio frames, that is, the signals of the specified channels in the HOA signals of these audio frames are encoded into the bitstream during encoding, that is, a compromise scheme is adopted for coding and decoding, so that the auditory quality after rendering and playback of the decoded and restored HOA signals can be smoothly transitioned.
[0277] Figure 16 1600 is a schematic diagram of the structure of an encoding device 1600 provided in an embodiment of the present application. The encoding device 1600 can be implemented by software, hardware, or a combination of both to form part or all of an encoding terminal device. The encoding terminal device can be any encoding terminal device in the aforementioned embodiments. Figure 16 The device 1600 includes: a first determination module 1601 and a first encoding module 1602.
[0278] A first determining module 1601 is configured to determine a coding scheme for the current frame based on a high-order ambisonics (HOA) signal of the current frame, where the coding scheme for the current frame is one of a first coding scheme, a second coding scheme, and a third coding scheme; wherein the first coding scheme is an HOA coding scheme based on directional audio coding, the second coding scheme is an HOA coding scheme based on virtual speaker selection, and the third coding scheme is a hybrid coding scheme;
[0279] The first encoding module 1602 is configured to encode signals of designated channels in the HOA signal into a bitstream if the encoding scheme of the current frame is the third encoding scheme, where the designated channels are some of all channels of the HOA signal.
[0280] Optionally, the signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X, Y, and Z signals.
[0281] Optionally, the first encoding module 1602 includes:
[0282] A first determination submodule is configured to determine a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal;
[0283] The encoding submodule is used to encode the virtual loudspeaker signal and the residual signal into a bit stream.
[0284] Optionally, the first determining submodule is configured to:
[0285] Determine the W signal as a virtual loudspeaker signal;
[0286] Three residual signals are determined based on the W signal, the X signal, the Y signal, and the Z signal, or the X signal, the Y signal, and the Z signal are determined as the three residual signals.
[0287] Optionally, the encoding submodule is used to:
[0288] Combine the virtual speaker signal with the first preset mono signal to obtain a stereo signal;
[0289] Combining the three residual signals with a second preset mono signal to obtain two stereo signals;
[0290] The three stereo signals obtained are encoded into bit streams respectively through a stereo encoder.
[0291] Optionally, the encoding submodule is used to:
[0292] Combining the two residual signals with the highest correlation among the three residual signals to obtain one stereo signal among the two stereo signals;
[0293] One of the three residual signals, except the two residual signals with the highest correlation, is combined with the second preset mono signal to obtain the other stereo signal of the two stereo signals.
[0294] Optionally, the first preset mono signal is an all-zero signal or an all-one signal, the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency values are all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency values are all one; the second preset mono signal is an all-zero signal or an all-one signal; the first preset mono signal is the same as or different from the second preset mono signal.
[0295] Optionally, the encoding submodule is used to:
[0296] The virtual speaker signal and each of the three residual signals are encoded into a bitstream through a mono encoder.
[0297] Optionally, the apparatus 1600 further includes:
[0298] a second encoding module, configured to encode the HOA signal into a bitstream according to the first encoding scheme if the encoding scheme of the current frame is the first encoding scheme;
[0299] The third encoding module is configured to encode the HOA signal into a bitstream according to the second encoding scheme if the encoding scheme of the current frame is the second encoding scheme.
[0300] Optionally, the first determining module 1601 includes:
[0301] A second determining submodule is configured to determine an initial coding scheme for a current frame according to the HOA signal, where the initial coding scheme is the first coding scheme or the second coding scheme;
[0302] a third determining submodule, configured to determine the coding scheme of the current frame as the initial coding scheme of the current frame if the initial coding scheme of the current frame is the same as the initial coding scheme of the frame before the current frame;
[0303] The fourth determination submodule is used to determine that the coding scheme of the current frame is the third coding scheme if the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the previous frame of the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the previous frame of the current frame is the first coding scheme.
[0304] Optionally, the apparatus 1600 further includes:
[0305] The fourth encoding module is configured to encode indication information of an initial encoding scheme of a current frame into a bitstream.
[0306] Optionally, the apparatus 1600 further includes:
[0307] a second determining module, configured to determine a value of a switching flag of a current frame, wherein when the coding scheme of the current frame is the first coding scheme or the second coding scheme, the value of the switching flag of the current frame is the first value; and when the coding scheme of the current frame is the third coding scheme, the value of the switching flag of the current frame is the second value;
[0308] The fifth encoding module is used to encode the value of the switching flag into the code stream.
[0309] Optionally, the apparatus 1600 further includes:
[0310] The sixth encoding module is used to encode the indication information of the encoding scheme of the current frame into the bitstream.
[0311] Optionally, the designated channel is consistent with a transmission channel preset in the first coding scheme.
[0312] In an embodiment of the present application, two schemes (i.e., a coding scheme based on virtual speaker selection and a coding scheme based on directional audio coding) are combined to encode and decode the HOA signal of the audio frame, that is, a suitable coding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to ensure a smooth transition of auditory quality when switching between different coding schemes, for some audio frames, this scheme does not directly adopt any of the above two schemes for coding and decoding, but adopts a new coding scheme to encode and decode these audio frames, that is, the signal of the specified channel in the HOA signal of these audio frames is encoded into the bit stream, that is, a compromise scheme is adopted for coding and decoding, so that the auditory quality after rendering and playing the HOA signal restored by decoding can be smoothly transitioned.
[0313] It should be noted that the encoding device provided in the above embodiment, when encoding audio frames, uses the division of the aforementioned functional modules as an example only. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the encoding device provided in the above embodiment and the encoding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0314] Figure 17 This is a schematic diagram of the structure of a decoding device 1700 provided in an embodiment of the present application. The decoding device 1700 can be implemented by software, hardware, or a combination of both to form part or all of a decoding end device. The decoding end device can be any encoding end device in the aforementioned embodiments. Figure 17 The device 1700 includes: a first obtaining module 1701, a first determining module 1702, a second determining module 1703, a third determining module 1704 and a second obtaining module 1705.
[0315] A first obtaining module 1701 is configured to obtain a decoding scheme for a current frame based on a bitstream, where the decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme; wherein the first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio decoding, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme;
[0316] A first determining module 1702 is configured to determine, based on a bitstream, signals of a specified channel in the HOA signal of the current frame if the decoding scheme of the current frame is the third decoding scheme, where the specified channel is a portion of all channels of the HOA signal;
[0317] A second determining module 1703 is configured to determine gains of one or more remaining channels in the HOA signal except the designated channel based on the signal of the designated channel;
[0318] A third determining module 1704 is configured to determine a signal of each of the one or more remaining channels based on the signal of the designated channel and the gains of the one or more remaining channels;
[0319] The second obtaining module 1705 is configured to obtain a reconstructed HOA signal of a current frame based on the signal of the designated channel and the signals of the one or more remaining channels.
[0320] Optionally, the first determining module 1702 includes:
[0321] A first determination submodule, configured to determine a virtual speaker signal and a residual signal based on a bit stream;
[0322] The second determining submodule is configured to determine a signal of a designated channel based on the virtual speaker signal and the residual signal.
[0323] Optionally, the first determining submodule is configured to:
[0324] Decode the code stream through a stereo decoder to obtain three-way stereo signals;
[0325] Based on the three stereo signals, one virtual speaker signal and three residual signals are determined.
[0326] Optionally, the first determining submodule is configured to:
[0327] determining a virtual loudspeaker signal based on one of the three stereo signals;
[0328] Three residual signals are determined based on the other two stereo signals of the three stereo signals.
[0329] Optionally, the first determining submodule is configured to:
[0330] The bit stream is decoded by a mono decoder to obtain one virtual speaker signal and three residual signals.
[0331] Optionally, the signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X signal, Y signal, and Z signal;
[0332] The first determination submodule is used for:
[0333] determining a W signal based on the virtual loudspeaker signal;
[0334] The X signal, the Y signal, and the Z signal are determined based on the residual signal and the W signal, or the X signal, the Y signal, and the Z signal are determined based on the residual signal.
[0335] Optionally, the apparatus 1700 further includes:
[0336] A first decoding module, configured to obtain a reconstructed HOA signal of the current frame according to the bitstream in accordance with the first decoding scheme if the decoding scheme of the current frame is the first decoding scheme;
[0337] The second decoding module is configured to obtain a reconstructed HOA signal of the current frame according to the bit stream in accordance with the second decoding scheme if the decoding scheme of the current frame is the second decoding scheme.
[0338] Optionally, the second decoding module includes:
[0339] A first obtaining submodule is configured to obtain an initial HOA signal according to a bit stream in accordance with a second decoding scheme;
[0340] a gain adjustment submodule, configured to adjust the gain of a high-order portion of the initial HOA signal according to the high-order gain of the frame before the current frame if the decoding scheme of the frame before the current frame is the third decoding scheme;
[0341] The second obtaining submodule is configured to obtain a reconstructed HOA signal based on the low-order portion of the initial HOA signal and the gain-adjusted high-order portion.
[0342] Optionally, the first obtaining module 1701 includes:
[0343] A first parsing submodule is used to parse the code stream to obtain the value of the switching flag of the current frame;
[0344] a second parsing submodule, configured to parse indication information of a decoding scheme for a current frame from the bitstream if the value of the switching flag is the first value, the indication information being used to indicate whether the decoding scheme for the current frame is the first decoding scheme or the second decoding scheme;
[0345] The third determining submodule is configured to determine that the decoding scheme for the current frame is a third decoding scheme if the value of the switching flag is the second value.
[0346] Optionally, the first obtaining module 1701 includes:
[0347] The third parsing submodule is configured to parse the code stream to obtain indication information of a decoding scheme for the current frame, where the indication information is used to indicate whether the decoding scheme for the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
[0348] Optionally, the first obtaining module 1701 includes:
[0349] a fourth parsing submodule, configured to parse the bitstream to obtain an initial decoding scheme for the current frame, where the initial decoding scheme is the first decoding scheme or the second decoding scheme;
[0350] a fourth determining submodule, configured to determine the decoding scheme of the current frame as the initial decoding scheme of the current frame if the initial decoding scheme of the current frame is the same as the initial decoding scheme of the frame before the current frame;
[0351] The fifth determination submodule is used to determine that the decoding scheme of the current frame is the third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme.
[0352] In an embodiment of the present application, two schemes (i.e., a coding scheme based on virtual speaker selection and a coding scheme based on directional audio coding) are combined to encode and decode the HOA signal of the audio frame, that is, a suitable coding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to ensure a smooth transition of auditory quality when switching between different coding schemes, for some audio frames, this scheme does not directly adopt any of the above two schemes for coding and decoding, but adopts a new coding scheme to encode and decode these audio frames, that is, the signal of the specified channel in the HOA signal of these audio frames is encoded into the bit stream, that is, a compromise scheme is adopted for coding and decoding, so that the auditory quality after rendering and playing the HOA signal restored by decoding can be smoothly transitioned.
[0353] It should be noted that the decoding device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the decoding of audio frames. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the decoding device provided in the above embodiment and the decoding method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0354] Figure 18 1800 is a schematic block diagram of a coding and decoding device 1800 used in an embodiment of the present application. The coding and decoding device 1800 may include a processor 1801, a memory 1802, and a bus system 1803. The processor 1801 and the memory 1802 are connected via the bus system 1803. The memory 1802 is used to store instructions, and the processor 1801 is used to execute the instructions stored in the memory 1802 to perform the various coding or decoding methods described in the embodiments of the present application. To avoid repetition, a detailed description is not given here.
[0355] In the embodiment of the present application, the processor 1801 may be a central processing unit (CPU), or may be another general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor.
[0356] The memory 1802 may include a ROM device or a RAM device. Any other suitable type of storage device may also be used as the memory 1802. The memory 1802 may include code and data 18021 accessed by the processor 1801 using the bus 1803. The memory 1802 may further include an operating system 18023 and an application 18022, which includes at least one program that allows the processor 1801 to execute the encoding or decoding method described in the embodiment of the present application. For example, the application 18022 may include applications 1 to N, which further include an encoding or decoding application (referred to as a codec application) that executes the encoding or decoding method described in the embodiment of the present application.
[0357] In addition to the data bus, the bus system 1803 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 1803 in the figure.
[0358] Optionally, the codec apparatus 1800 may further include one or more output devices, such as a display 1804. In one example, the display 1804 may be a touch-sensitive display that combines a display with a touch-sensitive unit operable to sense touch input. The display 1804 may be connected to the processor 1801 via a bus 1803.
[0359] It should be noted that the encoding and decoding device 1800 can execute the encoding method in the embodiment of the present application, and can also execute the decoding method in the embodiment of the present application.
[0360] Those skilled in the art will appreciate that the functions described in conjunction with the various illustrative logic blocks, modules, and algorithm steps disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions described in the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to tangible media, such as data storage media, or communication media including any media that facilitates the transfer of computer programs from one place to another (e.g., based on a communication protocol). In this manner, computer-readable media can generally correspond to (1) non-transitory tangible computer-readable storage media, or (2) communication media, such as signals or carrier waves. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in this application. A computer program product can include computer-readable media.
[0361] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disks and optical disks include compact discs (CDs), laser discs, optical discs, DVDs, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0362] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor," as used herein, may refer to any of the aforementioned structures or any other structures suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described by the various illustrative logic blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements. In one example, the various illustrative logic blocks, units, and modules in encoder 100 and decoder 200 may be understood as corresponding circuit devices or logic elements.
[0363] The techniques of the embodiments of the present application can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., a chipset). The various components, modules, or units described in the embodiments of the present application are intended to emphasize the functional aspects of the apparatus for performing the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. In fact, as described above, the various units can be combined in a codec hardware unit in conjunction with appropriate software and / or firmware, or provided by interoperable hardware units (including one or more processors as described above).
[0364] That is, in the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0365] It should be understood that the "at least one" mentioned herein refers to one or more, and "a plurality of" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0366] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A coding method, characterized in that: The method comprises: Determining a coding scheme for the current frame according to a high-order ambisonics (HOA) signal of the current frame, where the coding scheme for the current frame is one of a first coding scheme, a second coding scheme, and a third coding scheme; wherein the first coding scheme is an HOA coding scheme based on directional audio coding, the second coding scheme is an HOA coding scheme based on virtual speaker selection, and the third coding scheme is a hybrid coding scheme, which refers to a scheme that uses both the first coding scheme and the second coding scheme during an encoding process; If the encoding scheme of the current frame is the third encoding scheme, encoding a signal of a specified channel in the HOA signal into a bitstream, the specified channel being some channels among all channels of the HOA signal; The step of determining the encoding scheme of the current frame according to the high-order ambisonics (HOA) signal of the current frame includes: determining an initial coding scheme for the current frame according to the HOA signal, where the initial coding scheme is the first coding scheme or the second coding scheme; If the initial coding scheme of the current frame is the same as the initial coding scheme of the frame before the current frame, determining the coding scheme of the current frame as the initial coding scheme of the current frame; If the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the previous frame of the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the previous frame of the current frame is the first coding scheme, then the coding scheme of the current frame is determined to be the third coding scheme.
2. The method according to claim 1, wherein The signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X, Y, and Z signals.
3. The method according to claim 2, wherein The encoding of the signal of the designated channel in the HOA signal into the bitstream includes: determining a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal; The virtual speaker signal and the residual signal are encoded into the bitstream.
4. The method according to claim 3, wherein The determining of a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal includes: Determine the W signal as one of the virtual loudspeaker signals; Three residual signals are determined based on the W signal, the X signal, the Y signal, and the Z signal, or the X signal, the Y signal, and the Z signal are determined as the three residual signals.
5. The method according to claim 4, wherein The step of encoding the virtual speaker signal and the residual signal into the bitstream includes: Combining the virtual loudspeaker signal with the first preset mono signal to obtain a stereo signal; Combining the three residual signals with a second preset mono signal to obtain two stereo signals; The obtained three-way stereo signals are respectively encoded into the bit stream through a stereo encoder.
6. The method according to claim 5, wherein The step of combining the three residual signals with a second preset mono signal to obtain two stereo signals includes: Combining two residual signals with the highest correlation among the three residual signals to obtain one stereo signal among the two stereo signals; Combining one of the three residual signals except the two residual signals with the highest correlation with the second preset mono signal to obtain another stereo signal of the two stereo signals.
7. The method according to claim 5, wherein The first preset mono signal is an all-zero signal or an all-one signal, wherein the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency value is all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency value is all one; The second preset mono signal is an all-zero signal or an all-one signal; The first preset mono signal and the second preset mono signal are the same as or different from each other.
8. The method according to claim 4, wherein The step of encoding the virtual speaker signal and the residual signal into the bitstream includes: The one-way virtual speaker signal and each residual signal of the three-way residual signal are respectively encoded into the bit stream through a mono encoder.
9. The method according to any one of claims 1 to 8, wherein: After determining the encoding scheme of the current frame according to the high-order ambisonics (HOA) signal of the current frame, the method further includes: If the coding scheme of the current frame is the first coding scheme, encoding the HOA signal into the bitstream according to the first coding scheme; If the coding scheme of the current frame is the second coding scheme, the HOA signal is encoded into the code stream according to the second coding scheme.
10. The method according to claim 1, wherein After determining the initial coding scheme of the current frame according to the HOA signal, the method further includes: Encode indication information of the initial coding scheme of the current frame into the bitstream.
11. The method according to any one of claims 1 to 8, wherein: After determining the encoding scheme of the current frame according to the high-order ambisonics (HOA) signal of the current frame, the method further includes: determining a value of a switching flag of the current frame, where when the coding scheme of the current frame is the first coding scheme or the second coding scheme, the value of the switching flag of the current frame is a first value; and when the coding scheme of the current frame is the third coding scheme, the value of the switching flag of the current frame is a second value; The value of the switching flag is encoded into the code stream.
12. The method according to any one of claims 1 to 8, wherein: After determining the coding scheme of the current frame according to the HOA signal of the current frame, the method further includes: Encode indication information of the coding scheme of the current frame into the bitstream.
13. The method according to any one of claims 1 to 8, wherein: The designated channel is consistent with a transmission channel preset in the first coding scheme.
14. A decoding method, characterized in that: The method comprises: Obtaining a decoding scheme for a current frame based on a bitstream, where the decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme; wherein the first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio decoding, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme, which refers to a scheme that uses both the first decoding scheme and the second decoding scheme during a decoding process; If the decoding scheme of the current frame is the third decoding scheme, determining a signal of a designated channel in the HOA signal of the current frame based on the bitstream, where the designated channel is a portion of all channels of the HOA signal; determining, based on the signal of the designated channel, gains of one or more remaining channels in the HOA signal except the designated channel; determining a signal for each of the one or more remaining channels based on the signal of the designated channel and the gains of the one or more remaining channels; Obtaining a reconstructed HOA signal of the current frame based on the signal of the designated channel and the signals of the one or more remaining channels; The decoding scheme for obtaining the current frame based on the code stream includes: Parsing the value of the switching flag of the current frame from the code stream; If the value of the switching flag is the first value, parsing indication information of the decoding scheme of the current frame from the code stream, the indication information being used to indicate whether the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme; and If the value of the switching flag is the second value, determining that the decoding scheme for the current frame is the third decoding scheme; The decoding scheme for obtaining the current frame based on the code stream includes: Parsing an initial decoding scheme for the current frame from the code stream, where the initial decoding scheme is the first decoding scheme or the second decoding scheme; If the initial decoding scheme of the current frame is the same as the initial decoding scheme of the frame before the current frame, determining the decoding scheme of the current frame as the initial decoding scheme of the current frame; and If the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the frame before the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the frame before the current frame is the first decoding scheme, then the decoding scheme of the current frame is determined to be the third decoding scheme.
15. The method according to claim 14, wherein The determining, based on the code stream, a signal of a designated channel in the HOA signal of the current frame includes: determining a virtual speaker signal and a residual signal based on the bit stream; The signal of the designated channel is determined based on the virtual speaker signal and the residual signal.
16. The method according to claim 15, wherein The determining of the virtual speaker signal and the residual signal based on the code stream includes: Decoding the code stream by a stereo decoder to obtain three-way stereo signals; Based on the three stereo signals, one virtual speaker signal and three residual signals are determined.
17. The method according to claim 16, wherein The determining, based on the three-way stereo signal, one-way virtual speaker signal and three-way residual signal, comprises: Determining the one-way virtual loudspeaker signal based on one-way stereo signal among the three-way stereo signals; The three residual signals are determined based on the other two stereo signals of the three stereo signals.
18. The method according to claim 15, wherein The determining of the virtual speaker signal and the residual signal based on the code stream includes: The code stream is decoded by a mono decoder to obtain one channel of the virtual speaker signal and three channels of the residual signals.
19. The method according to claim 15, wherein The signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X signal, Y signal, and Z signal; The determining the signal of the designated channel based on the virtual speaker signal and the residual signal includes: determining the W signal based on the virtual speaker signal; The X signal, the Y signal, and the Z signal are determined based on the residual signal and the W signal, or the X signal, the Y signal, and the Z signal are determined based on the residual signal.
20. The method according to any one of claims 14 to 19, wherein: The method further comprises: If the decoding scheme of the current frame is the first decoding scheme, obtaining a reconstructed HOA signal of the current frame according to the bit stream according to the first decoding scheme; If the decoding scheme of the current frame is the second decoding scheme, the reconstructed HOA signal of the current frame is obtained according to the code stream in accordance with the second decoding scheme.
21. The method according to claim 20, wherein Obtaining the reconstructed HOA signal of the current frame according to the bit stream according to the second decoding scheme includes: Obtaining an initial HOA signal according to the bit stream according to the second decoding scheme; If the decoding scheme of the frame before the current frame is the third decoding scheme, performing gain adjustment on the high-order part of the initial HOA signal according to the high-order gain of the frame before the current frame; The reconstructed HOA signal is obtained based on the low-order part and the gain-adjusted high-order part of the initial HOA signal.
22. The method according to any one of claims 14 to 19, wherein: The decoding scheme for obtaining the current frame based on the code stream includes: Parse the code stream to obtain indication information of the decoding scheme of the current frame, where the indication information is used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
23. An encoding device, characterized in that The device comprises: a first determining module, configured to determine a coding scheme for the current frame based on a high-order ambisonics (HOA) signal of the current frame, where the coding scheme for the current frame is one of a first coding scheme, a second coding scheme, and a third coding scheme; wherein the first coding scheme is an HOA coding scheme based on directional audio coding, the second coding scheme is an HOA coding scheme based on virtual speaker selection, and the third coding scheme is a hybrid coding scheme, where the hybrid coding scheme uses both the first coding scheme and the second coding scheme during an encoding process; a first encoding module, configured to encode signals of designated channels in the HOA signal into a bitstream if the encoding scheme of the current frame is the third encoding scheme, where the designated channels are some of all channels of the HOA signal; The first determining module includes: A second determining submodule, configured to determine an initial coding scheme for the current frame according to the HOA signal, where the initial coding scheme is the first coding scheme or the second coding scheme; a third determining submodule, configured to determine, if the initial coding scheme of the current frame is the same as the initial coding scheme of a frame preceding the current frame, the coding scheme of the current frame as the initial coding scheme of the current frame; The fourth determination submodule is used to determine that the coding scheme of the current frame is the third coding scheme if the initial coding scheme of the current frame is the first coding scheme and the initial coding scheme of the frame before the current frame is the second coding scheme, or the initial coding scheme of the current frame is the second coding scheme and the initial coding scheme of the frame before the current frame is the first coding scheme.
24. The device according to claim 23, wherein The signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X, Y, and Z signals.
25. The device according to claim 24, wherein The first encoding module includes: a first determination submodule, configured to determine a virtual speaker signal and a residual signal based on the W signal, the X signal, the Y signal, and the Z signal; The encoding submodule is configured to encode the virtual speaker signal and the residual signal into the bitstream.
26. The device according to claim 25, characterized in that The first determining submodule is used for: Determine the W signal as one of the virtual loudspeaker signals; The difference signals between the X signal, the Y signal, and the Z signal and the W signal are determined as three residual signals, or the X signal, the Y signal, and the Z signal are determined as three residual signals.
27. The device according to claim 26, wherein The encoding submodule is used for: Combining the virtual loudspeaker signal with the first preset mono signal to obtain a stereo signal; Combining the three residual signals with a second preset mono signal to obtain two stereo signals; The obtained three-way stereo signals are respectively encoded into the bit stream through a stereo encoder.
28. The device according to claim 27, wherein The encoding submodule is used for: Combining two residual signals with the highest correlation among the three residual signals to obtain one stereo signal among the two stereo signals; Combining one of the three residual signals except the two residual signals with the highest correlation with the second preset mono signal to obtain another stereo signal of the two stereo signals.
29. The device according to claim 27, wherein The first preset mono signal is an all-zero signal or an all-one signal, wherein the all-zero signal includes a signal whose sampling point values are all zero or a signal whose frequency value is all zero, and the all-one signal includes a signal whose sampling point values are all one or a signal whose frequency value is all one; The second preset mono signal is an all-zero signal or an all-one signal; The first preset mono signal and the second preset mono signal are the same as or different from each other.
30. The device according to claim 26, wherein The encoding submodule is used for: The one-way virtual speaker signal and each residual signal of the three-way residual signal are respectively encoded into the bit stream through a mono encoder.
31. The device according to any one of claims 23 to 30, characterized in that The device further comprises: a second encoding module, configured to encode the HOA signal into the bitstream according to the first encoding scheme if the encoding scheme of the current frame is the first encoding scheme; a third encoding module, configured to encode the HOA signal into the bitstream according to the second encoding scheme if the encoding scheme of the current frame is the second encoding scheme.
32. The device according to claim 23, wherein The device further comprises: The fourth encoding module is configured to encode indication information of the initial encoding scheme of the current frame into the bitstream.
33. The device according to any one of claims 23 to 30, characterized in that The device further comprises: a second determining module, configured to determine a value of a switching flag of the current frame, wherein when the coding scheme of the current frame is the first coding scheme or the second coding scheme, the value of the switching flag of the current frame is a first value; and when the coding scheme of the current frame is the third coding scheme, the value of the switching flag of the current frame is a second value; The fifth encoding module is configured to encode the value of the switching flag into the bitstream.
34. The device according to any one of claims 23 to 30, characterized in that The device further comprises: The sixth encoding module is configured to encode the indication information of the encoding scheme of the current frame into the bitstream.
35. The device according to any one of claims 23 to 30, characterized in that The designated channel is consistent with a transmission channel preset in the first coding scheme.
36. A decoding device, characterized in that: The device comprises: A first obtaining module is configured to obtain a decoding scheme for a current frame based on a bitstream, where the decoding scheme for the current frame is one of a first decoding scheme, a second decoding scheme, and a third decoding scheme; wherein the first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio decoding, the second decoding scheme is an HOA decoding scheme based on virtual speaker selection, and the third decoding scheme is a hybrid decoding scheme, which refers to a scheme that uses both the first decoding scheme and the second decoding scheme during a decoding process; a first determining module configured to determine, if the decoding scheme of the current frame is the third decoding scheme, signals of designated channels in the HOA signal of the current frame based on a bitstream, where the designated channels are some of all channels of the HOA signal; a second determining module, configured to determine, based on the signal of the designated channel, gains of one or more remaining channels in the HOA signal except the designated channel; a third determining module, configured to determine a signal of each of the one or more remaining channels based on the signal of the designated channel and the gains of the one or more remaining channels; A second obtaining module is configured to obtain a reconstructed HOA signal of the current frame based on the signal of the designated channel and the signals of the one or more remaining channels; Wherein, the first obtaining module includes: A first parsing submodule, configured to parse the bitstream to obtain a value of the switching flag of the current frame; a second parsing submodule, configured to parse indication information of a decoding scheme for the current frame from the bitstream if the value of the switching flag is a first value, the indication information being used to indicate whether the decoding scheme for the current frame is the first decoding scheme or the second decoding scheme; and a third determining submodule, configured to determine that the decoding scheme for the current frame is the third decoding scheme if the value of the switching flag is the second value; Wherein, the first obtaining module includes: a fourth parsing submodule, configured to parse the bitstream to obtain an initial decoding scheme for the current frame, where the initial decoding scheme is the first decoding scheme or the second decoding scheme; a fourth determining submodule, configured to determine, if the initial decoding scheme of the current frame is the same as the initial decoding scheme of a frame previous to the current frame, the decoding scheme of the current frame as the initial decoding scheme of the current frame; and The fifth determination submodule is used to determine that the decoding scheme of the current frame is the third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the frame before the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the frame before the current frame is the first decoding scheme.
37. The device according to claim 36, wherein The first determining module includes: A first determining submodule, configured to determine a virtual speaker signal and a residual signal based on the bit stream; The second determining submodule is configured to determine the signal of the designated channel based on the virtual speaker signal and the residual signal.
38. The device according to claim 37, wherein The first determining submodule is used for: Decoding the code stream by a stereo decoder to obtain three-way stereo signals; Based on the three stereo signals, one virtual speaker signal and three residual signals are determined.
39. The device according to claim 38, wherein The first determining submodule is used for: Determining the one-way virtual loudspeaker signal based on one-way stereo signal among the three-way stereo signals; The three residual signals are determined based on the other two stereo signals of the three stereo signals.
40. The device according to claim 37, wherein The first determining submodule is used for: The code stream is decoded by a mono decoder to obtain one channel of the virtual speaker signal and three channels of the residual signals.
41. The device according to claim 37, wherein The signal of the designated channel includes a first-order ambisonics FOA signal, and the FOA signal includes an omnidirectional W signal, and directional X signal, Y signal, and Z signal; The first determining submodule is used for: determining the W signal based on the virtual speaker signal; The X signal, the Y signal, and the Z signal are determined based on the residual signal and the W signal, or the X signal, the Y signal, and the Z signal are determined based on the residual signal.
42. The device according to any one of claims 36 to 41, characterized in that The device further comprises: a first decoding module, configured to obtain a reconstructed HOA signal of the current frame according to the bitstream in accordance with the first decoding scheme if the decoding scheme of the current frame is the first decoding scheme; A second decoding module is configured to obtain a reconstructed HOA signal of the current frame according to the bit stream in accordance with the second decoding scheme if the decoding scheme of the current frame is the second decoding scheme.
43. The device according to claim 42, wherein The second decoding module includes: A first obtaining submodule, configured to obtain an initial HOA signal according to the bit stream in accordance with the second decoding scheme; a gain adjustment submodule, configured to, if the decoding scheme of the frame preceding the current frame is the third decoding scheme, perform gain adjustment on the high-order portion of the initial HOA signal according to the high-order gain of the frame preceding the current frame; The second obtaining submodule is configured to obtain the reconstructed HOA signal based on the low-order part and the gain-adjusted high-order part of the initial HOA signal.
44. The device according to any one of claims 36 to 41, wherein: The first obtaining module includes: The third parsing submodule is configured to parse the code stream to obtain indication information of the decoding scheme of the current frame, where the indication information is used to indicate whether the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme.
45. A coding terminal device, characterized in that: The encoding end device includes a memory and a processor; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the encoding method according to any one of claims 1 to 13.
46. A decoding terminal device, characterized in that: The decoding end device includes a memory and a processor; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the decoding method according to any one of claims 14 to 22.
47. A computer-readable storage medium, characterized in that The storage medium stores instructions, and when the instructions are executed on the computer, the computer executes the steps of the method according to any one of claims 1 to 22.
48. A computer program product, characterized in that The computer program product comprises instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 22.
Citation Information
Patent Citations
Coded HOA data frame representation that includes non-differential gain values associated with channel signals of specific ones of the data frames of an HOA data frame representation
CN107077852A
Audio scene encoder, audio scene decoder and related methods using hybrid encoder / decoder spatial analysis
CN112074902A