Decoding method, apparatus, device, storage medium and computer program product
By determining the decoding scheme based on the bitstream and performing alignment processing at the HOA signal decoding end, the problem of inconsistent time delay between different encoding and decoding schemes is solved, ensuring the consistency and efficiency of HOA signal decoding delay and realizing smooth switching between different encoding and decoding schemes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2026-04-10
AI Technical Summary
When switching between different codec schemes, the decoding delay of HOA signals varies, resulting in inconsistent decoding processes. In particular, the decoding delay of DirAC-based codec schemes is relatively large, affecting the decoding efficiency and quality of audio frames.
By determining the decoding scheme of the current frame based on the bitstream at the decoding end, and by using alignment processing to ensure consistent decoding latency between different encoding and decoding schemes, including techniques such as increasing latency, performing analysis and synthesis filtering, and circular buffering, the decoding latency of all frames is ensured to be consistent.
It enables smooth switching between different codec schemes, ensures consistent decoding latency of audio frames, and improves decoding efficiency and quality.
Smart Images

Figure CN115881138B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of audio processing, and in particular, to a decoding method and device, equipment, storage medium and computer program product. BACKGROUND
[0002] As a three-dimensional audio technology, higher order ambisonics (HOA) technology has attracted extensive attention due to its higher flexibility in three-dimensional audio playback. In order to achieve better hearing effect, HOA technology needs to record detailed sound scene information with a large amount of data. However, more data will be generated with the increase of HOA order, and a large amount of data will cause difficulties in transmission and storage. Therefore, how to encode and decode HOA signals has become a focus.
[0003] The related technology proposes two schemes for encoding and decoding HOA signals. One of the schemes is a directional audio coding (DirAC) based encoding and decoding scheme. In this scheme, the encoding end extracts a core layer signal and spatial parameters from the HOA signal of the current frame, and encodes the extracted core layer signal and spatial parameters into a code stream. The decoding end decodes the core layer signal and spatial parameters from the code stream, and performs analysis-synthesis filtering processing on the core layer signal and spatial parameters to reconstruct the HOA signal of the current frame. The other scheme is a virtual loudspeaker selection based encoding and decoding scheme. In this scheme, the encoding end selects a target virtual loudspeaker matching the HOA signal of the current frame from a virtual loudspeaker set based on a match-projection (MP) algorithm, determines a virtual loudspeaker signal based on the HOA signal of the current frame and the target virtual loudspeaker, determines a residual signal based on the HOA signal of the current frame and the virtual loudspeaker signal, and encodes the virtual loudspeaker signal and the residual signal into a code stream. The decoding end reconstructs the HOA signal of the current frame from the code stream by using a decoding method symmetrical to the encoding.
[0004] However, for the case that there are less different sound sources in the sound field, the compression rate of the codec scheme based on the virtual loudspeaker selection is higher, and for the case that there are more different sound sources in the sound field, the compression rate of the codec scheme based on DirAC is higher. The different sound sources refer to point sound sources with different positions and / or directions. The sound field types (related to the different sound sources in the sound field) of different audio frames can be different. If it is desired to have high compression rates for audio frames in different sound field types, a suitable codec scheme needs to be selected for the corresponding audio frame according to the sound field type of each audio frame, which requires switching between different codec schemes. However, the decoding delays of different codec schemes are different. For example, the decoding delay of the codec scheme based on DirAC is higher than that of the codec scheme based on the virtual loudspeaker selection, because the analysis-synthesis filtering process needs to be performed in the codec scheme based on DirAC. In the case of switching between different codec schemes, how to solve the problem of different decoding delays is the focus of current research. SUMMARY
[0005] Embodiments of the present application provide a decoding method, device, equipment, storage medium and computer program product, which can solve the problem of different decoding delays in the case of switching between different codec schemes. The technical solution is as follows:
[0006] In a first aspect, a decoding method is provided, which includes:
[0007] determining a decoding scheme of a current frame according to a bitstream, the decoding scheme of the current frame being a first decoding scheme or a non-first decoding scheme, the first decoding scheme being a HOA decoding scheme based on DirAC; if the decoding scheme of the current frame is the first decoding scheme, reconstructing, by a decoding end, a first audio signal according to the bitstream in a manner of the first decoding scheme, the reconstructed first audio signal being a reconstructed HOA signal of the current frame; if the decoding scheme of the current frame is the non-first decoding scheme, reconstructing, by the decoding end, a second audio signal according to the bitstream in a manner of the non-first decoding scheme, and performing alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, the alignment processing making the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme.
[0008] That is, since the decoding delay of the DirAC-based HOA decoding scheme is large, for the current frame encoded by the first encoding scheme, the current frame can be decoded according to the first decoding scheme. For the current frame not encoded by the first encoding scheme, the decoding delay of the current frame needs to be aligned with the decoding delay of the first decoding scheme. Since the decoding delay of the DirAC decoding scheme is fixed, the decoding delay of the current frame can be aligned with the decoding delay of the first decoding scheme (i.e., the DirAC decoding scheme) through the alignment processing. Generally, the decoding delay of the current frame can be aligned with the decoding delay of the first decoding scheme (i.e., the DirAC decoding scheme) by increasing the delay in the alignment processing. The first encoding scheme corresponds to the first decoding scheme, that is, if the first decoding scheme is the DirAC decoding scheme, the first encoding scheme is the DirAC encoding scheme; correspondingly, the second encoding scheme corresponds to the second decoding scheme, and the third encoding scheme also corresponds to the third decoding scheme.
[0009] Optionally, the decoding end determines the decoding scheme of the current frame according to the code stream, including: parsing the value of the switching flag of the current frame from the code stream; if the value of the switching flag is the first value, parsing the indication information of the decoding scheme of the current frame from the code stream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme, and the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection (which can be referred to as a HOA decoding scheme based on MP); if the value of the switching flag is the second value, determining that the decoding scheme of the current frame is the third decoding scheme, and the third decoding scheme being a hybrid decoding scheme. It should be noted that the hybrid decoding scheme is a scheme designed by the embodiments of the present application for the switching frame, and the encoding and decoding schemes of the previous frame and the next frame of the switching frame are different. The code stream contains the switching flag, and the value of the switching flag is the first value, indicating that the current frame is a non-switching frame, and the value of the switching flag is the second value, indicating that the current frame is a switching frame. The decoding end first parses the value of the switching flag from the code stream, and then parses the indication information of the decoding scheme of the current frame from the code stream to determine whether the first decoding scheme or the second decoding scheme in the case of determining that the current frame is not a switching frame based on the value of the switching flag. It can be seen that the decoding end can directly determine whether the current frame is a switching frame based on the switching flag, and the decoding efficiency is high. The hybrid decoding scheme refers to the use of both the technical means related to the first decoding scheme (i.e., the DirAC decoding scheme) and the technical means related to the second decoding scheme (the HOA decoding scheme based on MP) in the decoding process, so it is called the hybrid decoding scheme.
[0010] Optionally, the decoding end determines the decoding scheme of the current frame according to the code stream, comprising: parsing the indication information of the decoding scheme of the current frame from the code stream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme or the third decoding scheme, the second decoding scheme being the HOA decoding scheme based on the virtual speaker selection, and the third decoding scheme being the hybrid decoding scheme. That is, the indication information of the decoding scheme is directly included in the code stream, so that the decoding end directly determines the decoding scheme of the current frame based on the indication information, and the decoding efficiency is also higher.
[0011] Optionally, the decoding end determines the decoding scheme of the current frame according to the code stream, comprising: parsing the initial decoding scheme of the current frame from the code stream, the initial decoding scheme being the first decoding scheme or the second decoding scheme, and the second decoding scheme being the HOA decoding scheme based on the virtual speaker selection; if the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame of the current frame, determining the decoding scheme of the current frame as the initial decoding scheme of the current frame; if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, determining the decoding scheme of the current frame as the third decoding scheme, and the third decoding scheme being the hybrid decoding scheme. That is, the indication information of the initial decoding scheme is included in the code stream, and the decoding end determines whether the current frame is a switching frame by comparing the initial decoding scheme of the current frame with the initial decoding scheme of the previous frame, the decoding scheme of the switching frame being the third decoding scheme, and the decoding scheme of the non-switching frame being the initial decoding scheme of the non-switching frame.
[0012] Optionally, the non-first decoding scheme is the second decoding scheme or the third decoding scheme, the second decoding scheme being the HOA decoding scheme based on the virtual speaker selection, and the third decoding scheme being the hybrid decoding scheme; if the decoding scheme of the current frame is the third decoding scheme, reconstructing the second audio signal according to the code stream, comprising: reconstructing the signal of the specified channel according to the code stream, the reconstructed signal of the specified channel being the reconstructed second audio signal, and the specified channel being part of the channels of the HOA signal of the current frame. That is, for the switching frame decoded by using the third decoding scheme, the decoding end reconstructs the signal of the specified channel according to the code stream, rather than the complete HOA signal.
[0013] Optionally, if the decoding scheme of the current frame is the third decoding scheme, the decoding end performs alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, including: performing analysis filtering processing on the reconstructed signal of the specified channel; determining the gain of one or more remaining channels of the HOA signal of the current frame except the specified channel based on the analysis-filtered signal of the specified channel; determining the signal of the one or more remaining channels based on the gain of the one or more remaining channels and the analysis-filtered signal of the specified channel; and performing synthesis filtering processing on the analysis-filtered signal of the specified channel and the signal of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame. That is, for the switching frame, the decoding end needs to reconstruct the signal of the remaining channel except the specified channel, and through the analysis-synthesis filtering processing, the decoding delay of the current frame is increased to be consistent with the decoding delay of the first decoding scheme.
[0014] Optionally, the non-first decoding scheme is the second decoding scheme or the third decoding scheme, the second decoding scheme is an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme is a hybrid decoding scheme. If the decoding scheme of the current frame is the second decoding scheme, the decoding end reconstructs the second audio signal according to the code stream, including: reconstructing the first HOA signal according to the code stream according to the second decoding scheme, and the reconstructed first HOA signal is the reconstructed second audio signal. That is, for the audio frame encoded by the second encoding scheme, the decoding end first reconstructs the first HOA signal according to the second decoding scheme.
[0015] Optionally, the decoding end performs alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, including: performing analysis-synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame. That is, after the decoding end reconstructs the first HOA signal according to the second decoding scheme, the analysis-synthesis filtering processing is performed for time delay alignment.
[0016] Optionally, the decoding end performs analysis-synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame, including: performing analysis filtering processing on the reconstructed first HOA signal to obtain a second HOA signal; performing gain adjustment on the signal of one or more remaining channels of the second HOA signal to obtain gain-adjusted signal of the one or more remaining channels, the one or more remaining channels being channels of the HOA signal except the specified channel; and performing synthesis filtering processing on the signal of the specified channel in the second HOA signal and the gain-adjusted signal of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame. That is, for the audio frame encoded by the second encoding scheme, in the process of time delay alignment through the analysis-synthesis filtering processing, the gain adjustment is also performed to smoothly transition the auditory quality.
[0017] Optionally, the gain adjustment of the one or more remaining channel signals in the second HOA signal by the decoding end to obtain the gain-adjusted one or more remaining channel signals comprises: if the decoding scheme of the previous frame of the current frame is the third decoding scheme, gain-adjusting the one or more remaining channel signals in the second HOA signal according to the gain of the one or more remaining channels of the previous frame of the current frame to obtain the gain-adjusted one or more remaining channel signals. That is, if the previous frame of the current frame is a switching frame, the decoding end adjusts the signals of the remaining channels of the current frame according to the remaining channel gain of the switching frame, so that the auditory quality of the current frame is similar to that of the previous frame, to achieve smooth transition.
[0018] Optionally, the specified channel comprises a first-order ambisonics (FOA) channel. Optionally, the specified channel is consistent with the preset channel in the first decoding scheme.
[0019] Optionally, the decoding scheme of the previous frame of the current frame is the second decoding scheme; and the alignment processing of the reconstructed second audio signal by the decoding end to obtain the reconstructed HOA signal of the current frame comprises: cyclic buffering processing of the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame. That is, if the decoding scheme of the current frame is the second decoding scheme but the previous frame of the current frame is a non-switching frame, the decoding end can also achieve time delay alignment through cyclic buffering processing.
[0020] Optionally, the cyclic buffering processing of the reconstructed first HOA signal by the decoding end to obtain the reconstructed HOA signal of the current frame comprises: obtaining first data, the first data being data in the HOA signal of the previous frame of the current frame between a first time and an end time of the HOA signal of the previous frame, a time length between the first time and the end time being a first time length, the first time length being equal to a coding time delay difference between the first decoding scheme and the second decoding scheme; and merging the first data and second data to obtain the reconstructed HOA signal of the current frame, the second data being data in the reconstructed first HOA signal between a second time and a start time of the reconstructed first HOA signal, a time length between the second time and the start time being a second time length, a sum of the first time length and the second time length being equal to a frame length of the current frame. That is, the cyclic buffering processing is essentially a time delay alignment through data buffering.
[0021] Optionally, the method further comprises: buffering third data, the third data being data in the reconstructed first HOA signal other than the second data. That is, the buffering of the third data is for decoding of a next frame of the current frame.
[0022] In a second aspect, a decoding apparatus is provided, which has the function of implementing the behavior of the decoding method in the first aspect. The decoding apparatus comprises one or more modules for implementing the decoding method provided in the first aspect.
[0023] The first determining module is configured to determine, according to the bitstream, a decoding scheme of the current frame, the decoding scheme of the current frame being a first decoding scheme or a non-first decoding scheme, the first decoding scheme being a high-order ambisonic (HOA) decoding scheme based on directional audio coding (DirAC);
[0024] The first decoding module is configured to, if the decoding scheme of the current frame is the first decoding scheme, reconstruct the first audio signal according to the bitstream in the first decoding scheme, the reconstructed first audio signal being a reconstructed HOA signal of the current frame.
[0025] The second decoding module is configured to, if the decoding scheme of the current frame is the non-first decoding scheme, reconstruct the second audio signal according to the bitstream in the non-first decoding scheme, and perform alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, the alignment processing making the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme.
[0026] Optionally, the non-first decoding scheme is a second decoding scheme or a third decoding scheme, the second decoding scheme being an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme being a hybrid decoding scheme.
[0027] The second decoding module comprises:
[0028] The first reconstruction submodule is configured to, if the decoding scheme of the current frame is the third decoding scheme, reconstruct the signal of the specified channel according to the bitstream, the reconstructed signal of the specified channel being the reconstructed second audio signal, and the specified channel being part of all channels of the HOA signal of the current frame.
[0029] Optionally, the second decoding module comprises:
[0030] The analysis filtering submodule is configured to perform analysis filtering processing on the reconstructed signal of the specified channel.
[0031] The first determination submodule is configured to determine, based on the analysis-filtered signal of the specified channel, the gain of one or more remaining channels of the HOA signal of the current frame other than the specified channel.
[0032] The second determination submodule is configured to determine, based on the gain of the one or more remaining channels and the analysis-filtered signal of the specified channel, the signal of the one or more remaining channels.
[0033] a synthesis filtering sub-module, configured to perform a synthesis filtering process on the signals of the specified channel and the signals of the one or more remaining channels of the second HOA signal to obtain the reconstructed HOA signal of the current frame.
[0034] Optionally, the non-first decoding scheme is a second decoding scheme or a third decoding scheme, the second decoding scheme is an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme is a hybrid decoding scheme.
[0035] The second decoding module includes:
[0036] The second reconstruction sub-module is configured to, if the decoding scheme of the current frame is the second decoding scheme, reconstruct the first HOA signal according to the bitstream in the second decoding scheme, and the reconstructed first HOA signal is the reconstructed second audio signal.
[0037] Optionally, the second decoding module includes:
[0038] The analysis and synthesis filtering sub-module is configured to perform an analysis and synthesis filtering process on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
[0039] Optionally, the analysis and synthesis filtering sub-module is configured to:
[0040] perform an analysis filtering process on the reconstructed first HOA signal to obtain a second HOA signal;
[0041] perform a gain adjustment on the signals of one or more remaining channels of the second HOA signal to obtain gain-adjusted signals of the one or more remaining channels, the one or more remaining channels being channels other than the specified channel in the HOA signal;
[0042] perform a synthesis filtering process on the signals of the specified channel in the second HOA signal and the gain-adjusted signals of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
[0043] Optionally, the analysis and synthesis filtering sub-module is configured to:
[0044] perform a gain adjustment on the signals of the one or more remaining channels of the second HOA signal according to the gain of the one or more remaining channels of the previous frame of the current frame to obtain gain-adjusted signals of the one or more remaining channels, if the decoding scheme of the previous frame of the current frame is the third decoding scheme.
[0045] Optionally, the specified channel includes a first-order ambisonic (FOA) channel.
[0046] Optionally, the decoding scheme of the previous frame of the current frame is the second decoding scheme.
[0047] The second decoding module includes:
[0048] a circular buffer submodule configured to perform a circular buffer process on the reconstructed first HOA signal to obtain a reconstructed HOA signal of the current frame.
[0049] Optionally, the circular buffer submodule is configured to:
[0050] obtain first data, the first data being data in a previous frame HOA signal of the current frame and located between a first time and an end time of the previous frame HOA signal, a time length between the first time and the end time being a first time length, the first time length being equal to a coding time delay difference between the first decoding scheme and the second decoding scheme;
[0051] merge the first data and second data to obtain the reconstructed HOA signal of the current frame, the second data being data in the reconstructed first HOA signal and located between a start time of the reconstructed first HOA signal and a second time, a time length between the second time and the start time being a second time length, a sum of the first time length and the second time length being equal to a frame length of the current frame.
[0052] Optionally, the circular buffer submodule is configured to:
[0053] buffer third data, the third data being data in the reconstructed first HOA signal and excluding the second data.
[0054] Optionally, the first determining module comprises:
[0055] a first parsing submodule configured to parse a value of a switching flag of the current frame from the bitstream;
[0056] a second parsing submodule configured to, if the value of the switching flag is a first value, parse indication information of a decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection;
[0057] a third determining submodule configured to, if the value of the switching flag is a second value, determine that the decoding scheme of the current frame is a third decoding scheme, the third decoding scheme being a hybrid decoding scheme.
[0058] Optionally, the first determining module comprises:
[0059] a third parsing submodule configured to parse indication information of the decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme or the third decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection, the third decoding scheme being a hybrid decoding scheme.
[0060] Optionally, the first determining module comprises:
[0061] a fourth parsing submodule, configured to parse an initial decoding scheme of the current frame from the bitstream, the initial decoding scheme being the first decoding scheme or the second decoding scheme, the second decoding scheme being an HOA decoding scheme selected based on the virtual loudspeaker;
[0062] a fourth determining submodule, configured to determine the decoding scheme of the current frame as the initial decoding scheme of the current frame if the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame of the current frame;
[0063] a fifth determining submodule, configured to determine the decoding scheme of the current frame as a third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, the third decoding scheme being a hybrid decoding scheme.
[0064] In a third aspect, a decoding-side device is provided. The decoding-side device includes a processor and a memory. The memory is configured to store a program for executing the decoding method provided in the first aspect, and store data used in implementing the decoding method provided in the first aspect. The processor is configured to execute the program stored in the memory. The operation apparatus of the storage device can further include a communication bus for establishing a connection between the processor and the memory.
[0065] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to perform the decoding method provided in the first aspect.
[0066] In a fifth aspect, a computer program product is provided. The computer program product includes instructions, which, when executed on a computer, cause the computer to perform the decoding method provided in the first aspect.
[0067] The second aspect, the third aspect, the fourth aspect and the fifth aspect have similar technical effects to those of the first aspect, and thus details are not repeated here.
[0068] The technical solutions provided in the embodiments of the present application can at least bring the following beneficial effects:
[0069] In this embodiment, since the decoding delay of the HOA decoding scheme based on directional audio coding is relatively large, for the current frame encoded by the first encoding scheme, the bitstream of the current frame can be decoded according to the first decoding scheme. For the current frame not encoded by the first encoding scheme, the second audio signal is first reconstructed from the bitstream, and then the reconstructed second audio signal is aligned to obtain the reconstructed HOA signal of the current frame. That is, the alignment process makes the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme. In this way, this scheme can make the decoding delay of each audio frame consistent, that is, ensure delay alignment so that different encoding and decoding schemes can be switched well. Attached Figure Description
[0070] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0071] Figure 2 This is a schematic diagram of an implementation environment for a terminal scenario provided in an embodiment of this application;
[0072] Figure 3 This is a schematic diagram illustrating the implementation environment of a transcoding scenario for a wireless or core network device provided in an embodiment of this application;
[0073] Figure 4 This is a schematic diagram of an implementation environment for a broadcast television scenario provided in an embodiment of this application;
[0074] Figure 5 This is a schematic diagram of an implementation environment for a virtual reality streaming scene provided in an embodiment of this application;
[0075] Figure 6 This is a flowchart of an encoding method provided in an embodiment of this application;
[0076] Figure 7 This is a flowchart of another encoding method provided in the embodiments of this application;
[0077] Figure 8 This is a flowchart of a decoding method provided in an embodiment of this application;
[0078] Figure 9 This is an encoding diagram illustrating an encoding scheme switching method provided in an embodiment of this application;
[0079] Figure 10 This is a decoding diagram illustrating an encoding scheme switching method provided in an embodiment of this application;
[0080] Figure 11 This is a decoding diagram illustrating another encoding scheme switching provided in an embodiment of this application;
[0081] Figure 12is a structural schematic diagram of a decoding apparatus provided by an embodiment of the present application.
[0082] Figure 13 is a schematic block diagram of a coding and decoding apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION
[0083] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0084] Before the coding method provided by the embodiments of the present application is explained in detail, the implementation environment related to the embodiments of the present application will be introduced.
[0085] Please refer to Figure 1 , Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application. The implementation environment includes a source apparatus 10, a destination apparatus 20, a link 30 and a storage apparatus 40. The source apparatus 10 can generate encoded media data. Therefore, the source apparatus 10 can also be referred to as a media data encoding apparatus. The destination apparatus 20 can decode the encoded media data generated by the source apparatus 10. Therefore, the destination apparatus 20 can also be referred to as a media data decoding apparatus. The link 30 can receive the encoded media data generated by the source apparatus 10, and can transmit the encoded media data to the destination apparatus 20. The storage apparatus 40 can receive the encoded media data generated by the source apparatus 10, and can store the encoded media data. In this case, the destination apparatus 20 can directly obtain the encoded media data from the storage apparatus 40. Alternatively, the storage apparatus 40 can correspond to a file server or another intermediate storage apparatus that can save the encoded media data generated by the source apparatus 10. In this case, the destination apparatus 20 can obtain the encoded media data stored by the storage apparatus 40 via streaming or downloading.
[0086] Source device 10 and destination device 20 each can include one or more processors and a memory coupled to the one or more processors, the memory can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, any other memory suitable for storing information used by the computer, etc. Source device 10 and destination device 20 each can include a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a television, a camera, a display device, a digital media player, a video gaming console, an automobile computer, or the like.
[0087] Link 30 can include one or more media or devices capable of communicating encoded media data from source device 10 to destination device 20. In one possible implementation, link 30 can include one or more communication media that enable source device 10 to transmit encoded media data directly to destination device 20 in real-time. In embodiments of the present application, source device 10 can modulate the encoded media data based on a communication standard, which can be a wireless communication protocol, and transmit the modulated media data to destination device 20. The one or more communication media can include wireless and / or wired communication media, such as the one or more communication media can include one or more physical transmission lines, a wireless link, and / or the like. The one or more communication media can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet, and / or the like. The one or more communication media can include routers, switches, base stations, or other equipment that facilitates communication from source device 10 to destination device 20, and / or the like, which embodiments of the present application do not limit.
[0088] In one possible implementation, storage device 40 can store the encoded media data received by source device 10, and destination device 20 can access the encoded media data directly from storage device 40. In this regard, storage device 40 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded media data and the like.
[0089] In one possible implementation, storage device 40 can correspond to a file server or another intermediate storage device that stores the encoded media data generated by source device 10, and destination device 20 can access the encoded media data stored in storage device 40 via streaming or download. The file server can be any type of server capable of storing encoded media data and transmitting the encoded media data to destination device 20. In one possible implementation, the file server can include a web server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive, among others. Destination device 20 can access the encoded media data through any standard data connection, including an Internet connection. The
[0090] Figure 1 The illustrated implementation environment is merely one possible implementation, and the techniques of embodiments of the present application can be applicable to other implementation environments Figure 1 The illustrated source device 10 that encodes media data and the destination device 20 that decodes encoded media data can also be applicable to other devices that encode media data and decode encoded media data, which are not limited herein by embodiments of the present application.
[0091] In Figure 1In the illustrated implementation environment, source device 10 includes a data source 120, an encoder 100, and an output interface 140. In some embodiments, output interface 140 can include a modulator / demodulator (modem) and / or a transmitter, which can also be referred to as a transmitter. Data source 120 can include an image capture device (e.g., a video camera, etc.), an archive containing previously captured media data, a feed interface to receive media data from a media data content provider, and / or a computer graphics system to generate media data, or a combination of these sources of media data.
[0092] Data source 120 can send media data to encoder 100, which can encode the media data received by data source 120 to produce encoded media data. The encoder can send the encoded media data to the output interface. In some embodiments, source device 10 sends the encoded media data directly to destination device 20 via output interface 140. In other embodiments, the encoded media data can also be stored onto storage device 40 for later retrieval and use for decoding and / or display by destination device 20.
[0093] In Figure 1 In the illustrated implementation environment, destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, input interface 240 includes a receiver and / or a modem. Input interface 240 can receive encoded media data via link 30 and / or from storage device 40, and then send the encoded media data to decoder 200, which can decode the received encoded media data to produce decoded media data. The decoder can send the decoded media data to display device 220. Display device 220 can be integrated with destination device 20 or can be external to destination device 20. In general, display device 220 displays the decoded media data. Display device 220 can be any of a variety of types of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other type of display device.
[0094] Although Figure 1Although not shown, in some aspects, the encoder 100 and the decoder 200 can each be integrated with an encoder and a decoder, respectively, and can include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software, for encoding both audio and video in a common data stream or separate data streams. In some embodiments, the MUX-DEMUX units can comply with the ITU H.223 multiplexer protocol, or other protocols, such as the user datagram protocol (UDP), if applicable.
[0095] The encoder 100 and the decoder 200 can each be any of the following: one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuitry, hardware, or any combination thereof. If the techniques of this disclosure are implemented in software, the devices can store instructions for the software in a suitable, non- volatile computer-readable storage medium and can execute the instructions in hardware using one or more processors to implement the techniques of this disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) can be considered to be one or more processors. Each of the encoder 100 and the decoder 200 can be included in one or more encoders or decoders, respectively, any of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device.
[0096] Embodiments of the disclosure can generally refer to the encoder 100 as "signaling" or "sending" certain information to another device, such as the decoder 200. The term "signaling" or "sending" can generally refer to the communication of syntax elements and / or other data used for decoding the compressed media data. This communication can occur in real-time or near real-time. Alternatively, this communication can occur after a period of time, such as when syntax elements are stored in a computer-readable storage medium in an encoded bitstream at the time of encoding, which a decoding device can then retrieve at any time after the syntax elements are stored to this medium.
[0097] The coding method provided by the embodiments of the disclosure can be applied to various scenes. Next, taking the media data to be encoded as an HOA signal as an example, several of the scenes are introduced respectively.
[0098] Please refer to Figure 2 , Figure 2 is a schematic diagram of an implementation environment of a coding method provided by an embodiment of the present application applied to a terminal scenario. The implementation environment includes a first terminal 101 and a second terminal 201, and the first terminal 101 is in communication connection with the second terminal 201. The communication connection can be a wireless connection or a wired connection, and the present application embodiment does not limit this.
[0099] Among them, the first terminal 101 can be a sending end device or a receiving end device, and similarly, the second terminal 201 can be a receiving end device or a sending end device. For example, in the case of the first terminal 101 as a sending end device, the second terminal 201 is a receiving end device, and in the case of the first terminal 101 as a receiving end device, the second terminal 201 is a sending end device.
[0100] Next, take the first terminal 101 as a sending end device and the second terminal 201 as a receiving end device as an example for introduction.
[0101] The first terminal 101 and the second terminal 201 both include an audio acquisition module, an audio playback module, an encoder, a decoder, a channel encoding module, and a channel decoding module. In the present application embodiment, the encoder is a kind of three-dimensional audio encoder, and the decoder is a kind of three-dimensional audio decoder.
[0102] The audio acquisition module in the first terminal 101 acquires the HOA signal and transmits it to the encoder, and the encoder encodes the HOA signal by using the encoding method provided by the present application embodiment. This encoding can be called source encoding. Then, in order to realize the transmission of the HOA signal in the channel, the channel encoding module needs to perform channel encoding again, and then transmits the encoded stream in the digital channel through a wireless or wired network communication device.
[0103] The second terminal 201 receives the stream transmitted in the digital channel through a wireless or wired network communication device, and the channel decoding module decodes the stream. Then, the decoder decodes the HOA signal by using the decoding method provided by the present application embodiment, and then plays it through the audio playback module.
[0104] Among them, the first terminal 101 and the second terminal 201 can be any kind of electronic product that can interact with the user through one or more ways such as keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device, such as personal computer (PC), mobile phone, smart phone, personal digital assistant (PDA), wearable device, pocket PC (pocket PC), tablet computer, smart car machine, smart TV, smart speaker, etc.
[0105] Those skilled in the art shall understand that the terminal described above is only an example, and other existing or future terminals, such as those applicable to the embodiments of the present application, shall also be included in the protection scope of the embodiments of the present application, and are hereby incorporated by reference.
[0106] Please refer to Figure 3 , Figure 3 is a schematic diagram of an implementation environment of a transcoding scene of a coding and decoding method provided by the embodiments of the present application applied to a wireless or core network device. The implementation environment includes a channel decoding module, an audio decoder, an audio encoder and a channel encoding module. In the embodiments of the present application, the audio encoder is a three-dimensional audio encoder, and the audio decoder is a three-dimensional audio decoder.
[0107] The audio decoder can be a decoder using the decoding method provided by the embodiments of the present application, or a decoder using other decoding methods. The audio encoder can be an encoder using the encoding method provided by the embodiments of the present application, or an encoder using other encoding methods. In the case that the audio decoder is a decoder using the decoding method provided by the embodiments of the present application, the audio encoder is an encoder using other encoding methods, and in the case that the audio decoder is a decoder using other decoding methods, the audio encoder is an encoder using the encoding method provided by the embodiments of the present application.
[0108] The first case is that the audio decoder is a decoder using the decoding method provided by the embodiments of the present application, and the audio encoder is an encoder using other encoding methods.
[0109] At this time, the channel decoding module is used to perform channel decoding on the received code stream, then the audio decoder is used to perform source decoding using the decoding method provided by the embodiments of the present application, and then the audio encoder is used to perform encoding according to other encoding methods, to realize the conversion of one format to another format, i.e. transcoding. After that, it is sent again after channel encoding.
[0110] The second case is that the audio decoder is a decoder using other decoding methods, and the audio encoder is an encoder using the encoding method provided by the embodiments of the present application.
[0111] At this time, the channel decoding module is used to perform channel decoding on the received code stream, then the audio decoder is used to perform source decoding using other decoding methods, and then the audio encoder is used to perform encoding using the encoding method provided by the embodiments of the present application, to realize the conversion of one format to another format, i.e. transcoding. After that, it is sent again after channel encoding.
[0112] The wireless device can be a wireless access point, a wireless router, a wireless connector, etc. The core network device can be a mobility management entity, a gateway, etc.
[0113] Those skilled in the art shall understand that the wireless device or the core network device described above are only examples, and other existing or future wireless or core network devices, such as those applicable to the embodiments of the present application, shall also be included in the protection scope of the embodiments of the present application, and are hereby incorporated by reference.
[0114] Please refer to Figure 4 , Figure 4 is a schematic diagram of an implementation environment of a coding method provided by the embodiments of the present application applied to a broadcast television scenario. The broadcast television scenario is divided into a live broadcast scenario and a post-production scenario. For the live broadcast scenario, the implementation environment includes a live broadcast program three-dimensional sound production module, a three-dimensional sound encoding module, a set-top box, and a speaker group, and the set-top box includes a three-dimensional sound decoding module. For the post-production scenario, the implementation environment includes a post-production program three-dimensional sound production module, a three-dimensional sound encoding module, a network receiver, a mobile terminal, a headset, and the like.
[0115] In the live broadcast scenario, the live broadcast program three-dimensional sound production module produces a three-dimensional sound signal (such as an HOA signal), the three-dimensional sound signal is subjected to the encoding method of the embodiments of the present application to obtain a code stream, the code stream is transmitted to the user side through a broadcast network, the three-dimensional sound decoder in the set-top box decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the speaker group plays back the three-dimensional sound signal. Alternatively, the code stream is transmitted to the user side through the Internet, the three-dimensional sound decoder in the network receiver decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the speaker group plays back the three-dimensional sound signal. Alternatively, the code stream is transmitted to the user side through the Internet, the three-dimensional sound decoder in the mobile terminal decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the headset plays back the three-dimensional sound signal.
[0116] In the post-production scenario, the post-production program three-dimensional sound production module produces a three-dimensional sound signal, the three-dimensional sound signal is subjected to the encoding method of the embodiments of the present application to obtain a code stream, the code stream is transmitted to the user side through a broadcast network, the three-dimensional sound decoder in the set-top box decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the speaker group plays back the three-dimensional sound signal. Alternatively, the code stream is transmitted to the user side through the Internet, the three-dimensional sound decoder in the network receiver decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the speaker group plays back the three-dimensional sound signal. Alternatively, the code stream is transmitted to the user side through the Internet, the three-dimensional sound decoder in the mobile terminal decodes the code stream by using the decoding method provided by the embodiments of the present application, thereby reconstructing the three-dimensional sound signal, and the headset plays back the three-dimensional sound signal.
[0117] Please refer to Figure 5 , Figure 5is a schematic diagram of an implementation environment in which a coding and decoding method provided by an embodiment of the present application is applied to a virtual reality streaming scenario. The implementation environment includes an encoding end and a decoding end. The encoding end includes a collection module, a preprocessing module, an encoding module, a packing module, and a sending module. The decoding end includes an unpacking module, a decoding module, a rendering module, and a headset.
[0118] The collection module collects an HOA signal. The preprocessing module then performs preprocessing operations on the HOA signal, including filtering out low-frequency parts of the HOA signal, usually using 20 Hz or 50 Hz as a dividing point, extracting azimuth information in the HOA signal, and the like. The encoding module then performs encoding processing using the encoding method provided by an embodiment of the present application. After encoding, the packing module packs the signal, and the sending module sends the signal to the decoding end.
[0119] The unpacking module of the decoding end first performs unpacking. The decoding module then performs decoding using the decoding method provided by an embodiment of the present application. The rendering module then performs binaural rendering processing on the decoded signal, and the processed signal is mapped to a listener's headset. The headset can be a separate headset or a headset on a virtual reality-based glasses device.
[0120] It should be noted that the system architecture and business scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, as the system architecture evolves and new business scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0121] Next, the coding and decoding method provided by an embodiment of the present application is explained in detail. It should be noted that, in combination with the implementation environment shown in Figure 1 , any of the encoding methods in the following can be performed by the encoder 100 in the source device 10. Any of the decoding methods in the following can be performed by the decoder 200 in the destination device 20.
[0122] Figure 6 is a flowchart of an encoding method provided by an embodiment of the present application, which is applied to an encoding end. Please refer to Figure 6 , which includes the following steps.
[0123] Step 601: Determine an encoding scheme for a current frame according to an HOA signal of the current frame.
[0124] The HOA signal of the audio frame to be encoded is obtained by using the HOA acquisition technology. The HOA signal is a kind of scene audio signal and a kind of three-dimensional audio signal. The HOA signal refers to an audio signal obtained by collecting a sound field at a position of a microphone in space. The obtained audio signal is referred to as an original HOA signal. The HOA signal of the audio frame can also be an HOA signal obtained by converting a three-dimensional audio signal in another format. For example, converting a 5.1 channel signal into an HOA signal, or converting a three-dimensional audio signal mixed with an object audio into an HOA signal. Optionally, the HOA signal of the audio frame to be encoded is a time-domain signal or a frequency-domain signal, and can include all channels of the HOA signal or part of the channels of the HOA signal. For example, if the order of the HOA signal of the audio frame is 3, the number of channels of the HOA signal is 16, the frame length of the audio frame is 20 ms, and the sampling rate is 48 KHz, the HOA signal of the audio frame to be encoded includes 16 channels of signals, and each channel includes 960 sampling points.
[0125] In order to reduce the calculation complexity, if the HOA signal of the audio frame obtained by the encoding end is an original HOA signal, the number of sampling points or the number of frequency points of the original HOA signal is large, and then the encoding end can down-sample the original HOA signal to obtain the HOA signal of the audio frame to be encoded. For example, the encoding end down-samples the original HOA signal by 1 / Q to reduce the number of sampling points or the number of frequency points of the HOA signal to be encoded. For example, each channel of the original HOA signal includes 960 sampling points, and after down-sampling by 1 / 120, each channel of the HOA signal to be encoded includes 8 sampling points.
[0126] In the embodiment of the present application, the encoding method of the encoding end is introduced by taking the encoding of the current frame by the encoding end as an example. The current frame is an audio frame to be encoded. That is, the encoding end obtains the HOA signal of the current frame, and encodes the HOA signal of the current frame by using the encoding method provided in the embodiment of the present application.
[0127] It should be noted that, in order to achieve a high compression rate for audio frames under different sound field types, a suitable encoding / decoding scheme needs to be selected for each audio frame based on its sound field type. In this embodiment, the encoder first determines the initial encoding scheme for the current frame based on its sound field type. The initial encoding scheme can be either a first encoding scheme or a second encoding scheme. The encoder then determines whether to use the first, second, or third encoding scheme to encode the HOA signal of the current frame by comparing the initial encoding scheme of the current frame with that of the previous frame. Specifically, if the initial encoding scheme of the current frame is the same as that of the previous frame, the encoder uses the same encoding scheme to encode the HOA signal of the current frame. If the initial encoding scheme of the current frame is different from that of the previous frame, the encoder uses the third encoding scheme to encode the HOA signal of the current frame.
[0128] In this embodiment, the encoding scheme of the current frame is one of a first encoding scheme, a second encoding scheme, and a third encoding scheme. The first encoding scheme is a DirAC-based HOA encoding scheme, the second encoding scheme is a virtual speaker selection-based HOA encoding scheme, and the third encoding scheme is a hybrid encoding scheme. Optionally, the hybrid encoding scheme is also called a switching frame encoding scheme. The third encoding scheme is a switching frame encoding scheme provided in this embodiment, designed to ensure a smooth transition in auditory quality when switching between different encoding / decoding schemes. These three encoding schemes will be described in detail below. In this embodiment, the virtual speaker selection-based HOA encoding scheme is also called an MP-based HOA encoding scheme.
[0129] In this embodiment, the encoder determines the initial encoding scheme for the current frame based on the HOA signal of the current frame. Then, the encoder determines the encoding scheme for the current frame based on the initial encoding scheme of the current frame and the initial encoding scheme of the previous frame. It should be noted that this embodiment does not limit the implementation method of the encoder determining the initial encoding scheme.
[0130] Optionally, the encoding end performs sound field type analysis on the HOA signal of the current frame to obtain the sound field classification result of the current frame, and determines the initial encoding scheme of the current frame based on the sound field classification result of the current frame. It should be noted that the embodiments of this application do not limit the method of sound field type analysis. For example, the encoding end performs sound field type analysis by performing singular value decomposition on the HOA signal of the current frame.
[0131] Optionally, the sound field classification result includes the number of dissimilarity sound sources, and the application embodiments do not limit the determination method of the number of dissimilarity sound sources. After determining the number of dissimilarity sound sources corresponding to the current frame, if the number of dissimilarity sound sources corresponding to the current frame is greater than a first threshold and less than a second threshold, the encoding end determines that the initial encoding scheme of the current frame is the second encoding scheme. If the number of dissimilarity sound sources corresponding to the current frame is not greater than the first threshold or not less than the second threshold, the encoding end determines that the initial encoding scheme of the current frame is the first encoding scheme. The first threshold is less than the second threshold. Optionally, the first threshold is 0 or other values, and the second threshold is 3 or other values.
[0132] In the case of determining the initial encoding scheme of each audio frame (including the current frame) by the above method, the initial encoding scheme of each audio frame may be switched back and forth, so that more switching frames need to be encoded finally. Since there are many problems caused by the switching between encoding schemes, i.e., there are many problems to be solved, the number of switching frames can be reduced to reduce the problems caused by the switching. The switching frame refers to an audio frame whose initial encoding scheme is different from that of the previous frame. Optionally, in order to reduce the number of switching frames, the encoding end can first determine the expected encoding scheme of the current frame according to the sound field classification result of the current frame, i.e., the encoding end takes the initial encoding scheme determined by the foregoing method as the expected encoding scheme. Then, the encoding end updates the initial encoding scheme of the current frame based on the expected encoding scheme of the current frame by using a sliding window method, such as updating the initial encoding scheme of the current frame by using hangover processing.
[0133] Optionally, it is assumed that the length of the sliding window is N, and the sliding window includes the expected encoding scheme of the current frame and the updated initial encoding schemes of the previous N-1 frames of the current frame. If the cumulative number of the second encoding schemes in the sliding window is not less than a first specified threshold, the encoding end updates the initial encoding scheme of the current frame to the second encoding scheme. If the cumulative number of the second encoding schemes in the sliding window is less than the first specified threshold, the encoding end updates the initial encoding scheme of the current frame to the first encoding scheme. The length N of the sliding window is 8, 10, 15, etc., and the first specified threshold is 5, 6, 7, etc. The application embodiments do not limit the values of the length of the sliding window and the first specified threshold. For example, it is assumed that the length of the sliding window is 10, the first specified threshold is 7, the sliding window includes the expected encoding scheme of the current frame and the updated initial encoding schemes of the previous 9 frames of the current frame, and the cumulative number of the second encoding schemes in the sliding window is not less than 7. In this case, the encoding end updates the initial encoding scheme of the current frame to the second encoding scheme. If the cumulative number of the second encoding schemes in the sliding window is less than 7, the encoding end updates the initial encoding scheme of the current frame to the first encoding scheme.
[0134] Alternatively, if the number of the first encoding scheme in the sliding window is accumulated to be not less than the second specified threshold, the encoding end updates the initial encoding scheme of the current frame to the first encoding scheme. If the number of the first encoding scheme in the sliding window is accumulated to be less than the second specified threshold, the encoding end updates the initial encoding scheme of the current frame to the second encoding scheme. The second specified threshold is 5, 6, 7, etc., and the embodiments of the present application do not limit the value of the second specified threshold. Alternatively, the second specified threshold is different from or the same as the first specified threshold.
[0135] In addition to the above-mentioned implementation manners, the encoding end can also use other methods to obtain the sound field classification result of the current frame, and the method for determining the initial encoding scheme based on the sound field classification result can also be other methods, and the embodiments of the present application do not limit this.
[0136] In the embodiments of the present application, after the encoding end determines the initial encoding scheme of the current frame, if the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame, the encoding end determines the encoding scheme of the current frame to be the initial encoding scheme of the current frame. If the initial encoding scheme of the current frame is different from the initial encoding scheme of the previous frame of the current frame, the encoding end determines the encoding scheme of the current frame to be the third encoding scheme. That is, if the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame and is the first encoding scheme, the encoding end determines the encoding scheme of the current frame to be the first encoding scheme. If the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame and is the second encoding scheme, the encoding end determines the encoding scheme of the current frame to be the second encoding scheme. If one of the initial encoding scheme of the current frame and the initial encoding scheme of the previous frame of the current frame is the first encoding scheme and the other is the second encoding scheme, the encoding end determines the encoding scheme of the current frame to be the third encoding scheme. That is, one of the initial encoding scheme of the current frame and the initial encoding scheme of the previous frame of the current frame is the first encoding scheme and the other is the second encoding scheme, that is, the initial encoding scheme of the current frame is the first encoding scheme and the initial encoding scheme of the previous frame of the current frame is the second encoding scheme, or the initial encoding scheme of the current frame is the second encoding scheme and the initial encoding scheme of the previous frame of the current frame is the first encoding scheme. That is, for the switching frame, the encoding end does not use the first encoding scheme or the second encoding scheme to encode the HOA signal of the switching frame, but uses the switching frame encoding scheme to encode the HOA signal of the switching frame. For the non-switching frame, the encoding end uses the encoding scheme consistent with the initial encoding scheme of the non-switching frame to encode the HOA signal of the switching frame. The audio frame whose initial encoding scheme is different from the initial encoding scheme of the previous frame is the switching frame, and the audio frame whose initial encoding scheme is the same as the initial encoding scheme of the previous frame is the non-switching frame.
[0137] It should be noted that, in addition to determining the encoding scheme of the current frame, the encoding end also needs to encode information capable of indicating the encoding scheme of the current frame into the code stream, so as to determine which decoding scheme is used to decode the code stream of the current frame by the decoding end. In the embodiments of the present application, the encoding end has multiple implementation manners of encoding information capable of indicating the encoding scheme of the current frame into the code stream, and three implementation manners are introduced as follows.
[0138] First implementation , the encoding switching flag, and the indication information of the two encoding schemes
[0139] In the implementation manner, the encoding end needs to determine the value of the switching flag of the current frame, and encode the value of the switching flag of the current frame into the code stream. When the encoding scheme of the current frame is the first encoding scheme or the second encoding scheme, the value of the switching flag of the current frame is a first value. When the encoding scheme of the current frame is the third encoding scheme, the value of the switching flag of the current frame is a second value. Optionally, the first value is “0”, and the second value is “1”. The first value and the second value can also be other values.
[0140] In addition, the encoding end encodes the indication information of the initial encoding scheme of the current frame into the code stream. Alternatively, if the value of the switching flag of the current frame is the first value, the encoding end encodes the indication information of the initial encoding scheme of the current frame into the code stream, and if the value of the switching flag of the current frame is the second value, the encoding end encodes preset indication information into the code stream.
[0141] Optionally, the indication information of the initial encoding scheme is represented by a coding mode corresponding to the initial encoding scheme, that is, the coding mode is used as the indication information. For example, the coding mode corresponding to the initial encoding scheme is an initial coding mode, and the initial coding mode is a first coding mode (that is, a DirAC coding mode, that is, a DirAC encoding scheme) or a second coding mode (that is, an MP coding mode, that is, an MP encoding scheme). Optionally, the preset indication information is a preset coding mode, and the preset coding mode is the first coding mode or the second coding mode. In other embodiments, the preset indication information is other coding modes, that is, the indication information of the encoding scheme of the switching frame encoded into the code stream is not limited to be specific.
[0142] That is, in the first implementation manner, the encoding end uses the switching flag to indicate the switching frame, and can not limit the indication information of the encoding scheme of the switching frame encoded into the code stream. The indication information of the encoding scheme of the switching frame can be the initial coding mode of the switching frame, can be the preset coding mode, or can be randomly selected from the first coding mode and the second coding mode. It should be noted that, in this implementation manner, the switching flag is used to indicate whether the current frame is a switching frame, so that the decoding end can directly determine whether the current frame is a switching frame by acquiring the switching flag in the code stream.
[0143] Optionally, in the first implementation, the switching flag of the current frame and the indication information of the initial coding scheme each occupies one bit of the bitstream. For example, the switching flag of the current frame has a value of "0" or "1", wherein the value of "0" of the switching flag indicates that the current frame is not a switching frame, i.e., the switching flag of the current frame has a first value. The value of "1" of the switching flag indicates that the current frame is a switching frame, i.e., the switching flag of the current frame has a second value. Optionally, the indication information of the initial coding scheme is "0" or "1", wherein "0" represents the DirAC mode and "1" represents the MP mode.
[0144] In some other embodiments, if the initial coding scheme of the current frame is different from the initial coding scheme of the previous frame of the current frame, the encoding end determines that the value of the switching flag of the current frame is the second value and encodes the value of the switching flag of the current frame into the bitstream. That is, for the switching frame, since the switching flag in the bitstream can indicate the switching frame, the indication information of the coding scheme of the switching frame does not need to be encoded.
[0145] Second implementation , encoding indication information of two coding schemes
[0146] The encoding end encodes the indication information of the initial coding scheme of the current frame into the bitstream. For example, taking the encoding mode as the indication information, the indication information encoded into the bitstream is essentially the encoding mode consistent with the initial coding scheme, i.e., the DirAC mode or the MP mode.
[0147] Optionally, in the first implementation, the indication information of the initial coding scheme occupies one bit of the bitstream. For example, taking the encoding mode as the indication information, the indication information is "0" or "1", wherein "0" represents the DirAC mode, indicating that the initial coding scheme of the current frame is the first coding scheme, and "1" represents the MP mode, indicating that the initial coding scheme of the current frame is the second coding scheme.
[0148] Third implementation , encoding indication information of three coding schemes
[0149] In this implementation, the encoding end encodes the indication information of the coding scheme of the current frame into the bitstream. For example, taking the encoding mode as the indication information, the indication information encoded into the bitstream is essentially the encoding mode consistent with the coding scheme of the current frame, i.e., the DirAC mode, the MP mode or the MP-W mode. The MP-W mode is the encoding mode corresponding to the coding scheme of the switching frame. If the indication information is the MP-W mode, it indicates that the current frame is a switching frame, and if the indication information is the DirAC mode or the MP mode, it indicates that the current frame is a non-switching frame.
[0150] Optionally, in the third implementation, the indication information of the encoding scheme of the current frame occupies two bits of the bitstream. For example, the indication information encoded in the bitstream is "00", "01" or "10". Wherein, "00" indicates that the encoding scheme of the current frame is the first encoding scheme, "01" indicates that the encoding scheme of the current frame is the second encoding scheme, and "10" indicates that the encoding scheme of the current frame is the third encoding scheme.
[0151] Step 602: If the encoding scheme of the current frame is the third encoding scheme, encode the signals of the specified channels of the HOA signal into the bitstream, and the specified channels are part of the channels of the HOA signal.
[0152] In the embodiments of the present application, the third encoding scheme indicates that only the signals of the specified channels of the HOA signal of the current frame are encoded into the bitstream. Wherein, the specified channels are part of the channels of the HOA signal. That is, for the switching frame, the encoding end encodes the signals of the specified channels of the HOA signal of the switching frame into the bitstream, instead of encoding the switching frame by using the first encoding scheme or the second encoding scheme, that is, the present scheme adopts a compromise way to encode the switching frame in order to smoothly transit the auditory quality when the encoding scheme is switched.
[0153] Optionally, the specified channels are consistent with the preset transmission channels in the first encoding scheme, that is, the specified channels are the preset channels. That is, under the premise that the third encoding scheme is different from the second encoding scheme, in order to make the encoding effect of the third encoding scheme close to that of the first encoding scheme, the encoding end encodes the signals of the channels of the HOA signal of the switching frame which are the same as the preset transmission channels in the first encoding scheme into the bitstream, so as to make the auditory quality as smoothly transit as possible when the encoding scheme is switched. It should be noted that, according to the different encoding bandwidth, code rate, or even the different application scenarios, different transmission channels can be preset. Optionally, the preset transmission channels can be the same under different encoding bandwidth, code rate or application scenarios.
[0154] It should be noted that in the embodiments of the present application, there are many implementation manners for the encoding end to encode the signals of the specified channels in the HOA signal into the code stream, and the encoding end can encode the signals of the specified channels into the code stream, and the embodiments of the present application do not limit this. Optionally, the signals of the specified channels include FOA signals, and the FOA signals include omnidirectional W signals and directional X signals, Y signals and Z signals. That is, the specified channels include FOA channels, and the signals of the FOA channels are low-order signals, that is, if the current frame is a switching frame, the encoding end only encodes the low-order part of the HOA signal of the current frame into the code stream, and the low-order part includes the W signals, the X signals, the Y signals and the Z signals of the FOA channels. For example, the encoding end determines the virtual loudspeaker signals and the residual signals based on the signals of the specified channels, and encodes the virtual loudspeaker signals and the residual signals into the code stream. For example, if the specified channels include FOA channels, the encoding end determines the W signals as one virtual loudspeaker signal, and determines the difference signals between the X signals, the Y signals and the Z signals and the W signals as three residual signals, or determines the X signals, the Y signals and the Z signals as three residual signals. The encoding end encodes the one virtual loudspeaker signal and the three residual signals into the code stream through the core encoder. Optionally, the core encoder is a stereo encoder or a single-channel encoder.
[0155] The above introduces the process that the encoding end encodes the current frame by using the switching frame encoding scheme in the case that the current frame is a switching frame, that is, the encoding end encodes the signals of the specified channels in the HOA signal of the current frame into the code stream based on the third encoding scheme. In the embodiments of the present application, the switching frame encoding scheme can also be referred to as an encoding scheme based on MP-W. Next, the process that the encoding end encodes the current frame in the case that the current frame is a non-switching frame is introduced.
[0156] In the embodiments of the present application, if the encoding scheme of the current frame is the first encoding scheme, the encoding end encodes the HOA signal of the current frame into the code stream according to the first encoding scheme. If the encoding scheme of the current frame is the second encoding scheme, the encoding end encodes the HOA signal of the current frame into the code stream according to the second encoding scheme. That is, if the current frame is not a switching frame, the encoding end encodes the current frame by using the initial encoding scheme of the current frame.
[0157] In the embodiment of the present application, the implementation process of the encoding end for encoding the HOA signal of the current frame into the code stream according to the first encoding scheme is as follows: the encoding end extracts the core layer signal and the spatial parameter from the HOA signal of the current frame, and encodes the extracted core layer signal and spatial parameter into the code stream. For example, the encoding end extracts the core layer signal from the HOA signal of the current frame through a core encoding signal acquisition module, extracts the spatial parameter from the HOA signal of the current frame through a spatial parameter extraction module based on DirAC, encodes the core layer signal into the code stream through a core encoder, and encodes the spatial parameter into the code stream through a spatial parameter encoder. The channel corresponding to the core layer signal is consistent with the designated channel in the present scheme. In addition, it is emphasized that, in addition to encoding the core layer signal into the code stream, the first encoding scheme also encodes the extracted spatial parameter into the code stream. The spatial parameter contains rich scene information, such as direction information. The switching frame encoding scheme provided in the present application only encodes the signal of the designated channel into the code stream. It can be seen that, for the same frame, the effective information encoded into the code stream by the DirAC-based HOA encoding scheme is more than that encoded into the code stream by the switching frame encoding scheme. Under the premise that the switching frame encoding scheme is different from the first encoding scheme, in order to make the encoding effect of the switching frame encoding scheme close to that of the first encoding scheme, the switching frame encoding scheme also encodes the signal of the designated channel in the HOA signal which is the same as the preset transmission channel in the first encoding scheme into the code stream, but does not encode more information in the HOA signal into the code stream, that is, does not extract the spatial parameter, and does not encode the spatial parameter into the code stream, so as to make the auditory quality transition as smoothly as possible.
[0158] The implementation process of the encoding end for encoding the HOA signal of the current frame into the code stream according to the second encoding scheme is as follows: the encoding end selects a target virtual loudspeaker matching the HOA signal of the current frame from the virtual loudspeaker set based on the MP algorithm, determines the virtual loudspeaker signal based on the HOA signal of the current frame and the target virtual loudspeaker through an MP-based spatial encoder, determines the residual signal based on the HOA signal of the current frame and the virtual loudspeaker signal through the MP-based spatial encoder, and encodes the virtual loudspeaker signal and the residual signal into the code stream through a core encoder. It is emphasized that the principle and specific manner of determining the virtual loudspeaker signal and the residual signal in the MP-based HOA encoding scheme are different from those in the switching frame encoding scheme, and the virtual loudspeaker signal and the residual signal determined by the two schemes are also different. For the same frame, the effective information encoded into the code stream by the MP-based HOA encoding scheme is more than that encoded into the code stream by the switching frame encoding scheme. Under the premise that the switching frame encoding scheme is different from the second encoding scheme, in order to make the encoding effect of the switching frame encoding scheme close to that of the first encoding scheme, the switching frame encoding scheme also adopts the manner of encoding the virtual loudspeaker signal and the residual signal, so as to make the auditory quality transition as smoothly as possible.
[0159] Figure 7 is a flowchart of another encoding method provided by an embodiment of the present application. Please refer to Figure 7 to take the example of encoding the indication information of the initial encoding scheme of the current frame into the code stream, the encoding method provided by an embodiment of the present application is explained again. The encoding end first acquires the HOA signal of the current frame to be encoded. Then, the encoding end analyzes the sound field type of the HOA signal to determine the initial encoding scheme of the current frame. The encoding end judges whether the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame. If the initial encoding scheme of the current frame is the same as the initial encoding scheme of the previous frame of the current frame, the encoding end encodes the HOA signal of the current frame by using the initial encoding scheme of the current frame to obtain the code stream of the current frame. If the initial encoding scheme of the current frame is different from the initial encoding scheme of the previous frame of the current frame, the encoding end encodes the HOA signal of the current frame by using the switching frame encoding scheme to obtain the code stream of the current frame.
[0160] It should be noted that if the current frame is the first audio frame to be encoded, the initial encoding scheme of the current frame is the first encoding scheme or the second encoding scheme, and the encoding end encodes the HOA signal of the current frame into the code stream by using the initial encoding scheme of the current frame.
[0161] In summary, in the embodiment of the present application, the HOA signal of the audio frame is encoded and decoded in combination with two schemes (i.e. the encoding and decoding scheme based on virtual loudspeaker selection and the encoding and decoding scheme based on DirAC), that is, a suitable encoding and decoding scheme is selected for different audio frames, which can improve the compression rate of the audio signal. At the same time, in order to make the smooth transition of the auditory quality when switching between different encoding and decoding schemes, for the switching frame in the present scheme, instead of directly using any of the above two schemes to encode the switching frame, the signal of the specified channel in the HOA signal of the switching frame is encoded into the code stream, that is, a compromise scheme is used to encode and decode the switching frame, so that the auditory quality after the rendering and playing of the HOA signal of the switching frame recovered by decoding can be smoothly transitioned.
[0162] Figure 8 is a flowchart of a decoding method provided by an embodiment of the present application, which is applied to the decoding end. It should be noted that the decoding method corresponds to the encoding method shown in Figure 6 . Please refer to Figure 8 , which includes the following steps.
[0163] Step 801: determining the decoding scheme of the current frame according to the code stream, the decoding scheme of the current frame being the first decoding scheme or the non-first decoding scheme, and the first decoding scheme being the HOA decoding scheme based on DirAC.
[0164] It should be noted that since the encoding end adopts different encoding schemes to encode different audio frames, the decoding end also needs to use the corresponding decoding scheme to decode each audio frame. Next, how the decoding end determines the decoding scheme of the current frame will be introduced.
[0165] As can be seen from the foregoing, in Figure 6 The three implementation manners of the encoding end encoding the information capable of indicating the encoding scheme of the current frame into the code stream are introduced in step 601 of the encoding method shown in the table. Correspondingly, there are also three implementation manners for the decoding end to determine the encoding scheme of the current frame. Next, this will be introduced.
[0166] First implementation , the switching flag and the indication information of the two encoding schemes are encoded
[0167] The decoding end first parses the value of the switching flag of the current frame from the code stream. If the value of the switching flag is the first value, the decoding end then parses the indication information of the decoding scheme of the current frame from the code stream. The indication information is used to indicate that the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme. If the value of the switching flag is the second value, the decoding end determines that the encoding scheme of the current frame is the third encoding scheme. It should be noted that the indication information of the encoding scheme encoded into the code stream by the encoding end is the indication information of the decoding scheme parsed from the code stream by the decoding end.
[0168] In other words, if the decoding end parses that the value of the switching flag of the current frame is the first value, it means that the current frame is a non-switching frame. The decoding end then parses the indication information of the decoding scheme from the code stream and determines the decoding scheme of the current frame based on the indication information. If the value of the switching flag is the second value, the decoding end determines that the decoding scheme of the current frame is the third decoding scheme, and the current frame is a switching frame. In this case, even if the code stream contains the indication information, the decoding end does not need to decode the indication information. The third decoding scheme is a hybrid decoding scheme, i.e., the switching frame decoding scheme.
[0169] It should be noted that if the value of the switching flag is the second value, the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme, and the current frame is a switching frame. The switching frame decoding scheme is a hybrid decoding scheme different from the first decoding scheme and the second decoding scheme, and the switching frame decoding scheme is for smooth transition of auditory quality and time delay alignment.
[0170] Optionally, in the first implementation, the indication information of the decoding scheme and the switch flag each occupies one bit of the bitstream. Illustratively, the decoding end first parses the value of the switch flag of the current frame from the bitstream, and if the parsed value of the switch flag is "0", i.e., the value of the switch flag is the first value, the decoding end then parses the indication information of the decoding scheme of the current frame from the bitstream, and if the parsed indication information is "0", the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is "1", the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed value of the switch flag is "1", i.e., the value of the switch flag is the second value, the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme, i.e., the third decoding scheme.
[0171] Optionally, in the case that the current frame is a switching frame, the decoding end can determine the switching state of the current frame based on the switch flag of the current frame and the decoding scheme of the previous frame of the current frame. For example, if the value of the switch flag of the current frame is the first value and the decoding scheme of the previous frame of the current frame is the first decoding scheme, the decoding end determines that the switching state of the current frame is the first switching state, which refers to the state of switching from the DirAC-based HOA decoding scheme to the MP-based HOA decoding scheme. If the value of the switch flag of the current frame is the second value and the decoding scheme of the previous frame of the current frame is the second decoding scheme, the decoding end determines that the switching state of the current frame is the second switching state, which refers to the state of switching from the MP-based HOA decoding scheme to the DirAC-based HOA decoding scheme.
[0172] Second implementation , the indication information of the two encoding schemes
[0173] The decoding end parses the initial decoding scheme of the current frame from the bitstream, and the initial decoding scheme is the first decoding scheme or the second decoding scheme. If the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame of the current frame, the decoding end determines that the decoding scheme of the current frame is the initial decoding scheme of the current frame. If the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame of the current frame, the decoding end determines that the decoding scheme of the current frame is the third decoding scheme, i.e., the hybrid decoding scheme. Wherein, the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame of the current frame refers to that the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme. That is, one of the initial decoding scheme of the current frame and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, and the other is the second decoding scheme.
[0174] Optionally, taking the coding mode as an example of the indication information of the initial coding scheme coded into the bitstream, the indication information parsed from the bitstream is referred to as the coding mode. It should be noted that if the initial decoding scheme of the current frame is different from the initial decoding scheme of the previous frame of the current frame, it indicates that the current frame is a switching frame. If the initial decoding scheme of the current frame is the same as the initial decoding scheme of the previous frame of the current frame, it indicates that the current frame is a non-switching frame.
[0175] Optionally, in the second implementation mode, the indication information for indicating the initial decoding scheme occupies one bit of the bitstream. Taking the coding mode as an example of the indication information, the coding mode in the bitstream occupies one bit. Illustratively, the decoding end parses the indication information of the current frame from the bitstream. If the parsed indication information is “0” and the indication information of the previous frame of the current frame is also “0”, the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is “1” and the indication information of the previous frame of the current frame is also “1”, the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed indication information is “0” and the indication information of the previous frame of the current frame is “1”, or the parsed indication information is “1” and the indication information of the previous frame of the current frame is “0”, the decoding end determines that the decoding scheme of the current frame is the third decoding scheme.
[0176] Optionally, the indication information of the initial decoding scheme of the previous frame of the current frame is buffered data. When decoding to the current frame, the decoding end can obtain the indication information of the initial decoding scheme of the previous frame of the current frame from the buffer.
[0177] Optionally, in the case that the current frame is a switching frame, the decoding end can determine the switching state of the current frame based on the initial decoding scheme of the previous frame of the current frame. For example, if the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, the decoding end determines that the switching state of the current frame is the first switching state, and the first switching state refers to the state of switching from the DirAC-based HOA decoding scheme to the MP-based HOA decoding scheme. If the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, the decoding end determines that the switching state of the current frame is the second switching state, and the second switching state refers to the state of switching from the MP-based HOA decoding scheme to the DirAC-based HOA decoding scheme.
[0178] Third implementation , the indication information of the three coding schemes is coded
[0179] The decoding end parses the indication information of the decoding scheme of the current frame from the bitstream, and the indication information is used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme or the third decoding scheme.
[0180] For example, assuming that the encoding mode is used as the indication information, the decoding end parses the encoding mode of the current frame from the bitstream. If the encoding mode of the current frame is the DirAC mode, the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the encoding mode of the current frame is the MP mode, the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the encoding mode of the current frame is the MP-W mode, the decoding end determines that the decoding scheme of the current frame is the third decoding scheme.
[0181] Optionally, in the third implementation, the indication information of the decoding scheme occupies two bits of the bitstream. For example, assuming that the encoding mode is used as the indication information, the encoding mode of the current frame occupies two bits of the bitstream. For example, the decoding end parses the indication information of the decoding scheme of the current frame from the bitstream. If the parsed indication information is “00”, the decoding end determines that the decoding scheme of the current frame is the first decoding scheme. If the parsed indication information is “01”, the decoding end determines that the decoding scheme of the current frame is the second decoding scheme. If the parsed indication information is “10”, the decoding end determines that the decoding scheme of the current frame is the switching frame decoding scheme.
[0182] Optionally, in the case that the current frame is a switching frame, the decoding end can determine the switching state of the current frame based on the decoding scheme of the previous frame of the current frame. For example, if the decoding scheme of the previous frame of the current frame is the first decoding scheme, the decoding end determines that the switching state of the current frame is the first switching state, which is the state of switching from the DirAC-based HOA decoding scheme to the MP-based HOA decoding scheme. If the decoding scheme of the previous frame of the current frame is the second decoding scheme, the decoding end determines that the switching state of the current frame is the second switching state, which is the state of switching from the MP-based HOA decoding scheme to the DirAC-based HOA decoding scheme.
[0183] Step 802: If the encoding scheme of the current frame is the first encoding scheme, reconstruct the first audio signal according to the bitstream in the first encoding scheme, and the reconstructed first audio signal is the reconstructed HOA signal of the current frame.
[0184] In the embodiments of the present application, because the decoding delay of the DirAC-based HOA decoding scheme is large, if the decoding scheme of the current frame is the first decoding scheme, the decoding end decodes the bitstream by using the first decoding scheme, and the reconstructed HOA signal of the current frame can be obtained. That is, if the decoding scheme of the current frame is the first decoding scheme, the decoding end reconstructs the first audio signal according to the bitstream in the first decoding scheme, and the reconstructed first audio signal is the reconstructed HOA signal of the current frame.
[0185] The implementation process of reconstructing the first audio signal according to the bitstream by the decoding end according to the first decoding scheme is as follows: the decoding end parses the core layer signal and the spatial parameter from the bitstream, and reconstructs the HOA signal of the current frame based on the core layer signal and the spatial parameter. Illustratively, the decoding end parses the core layer signal from the bitstream through a core decoder, parses the spatial parameter from the bitstream through a spatial parameter decoder, and performs DirAC-based HOA signal synthesis processing based on the parsed core layer signal and the spatial parameter, to reconstruct the first audio signal. The reconstructed first audio signal is the reconstructed HOA signal of the current frame.
[0186] Step 803: If the encoding scheme of the current frame is the non-first encoding scheme, the second audio signal is reconstructed according to the bitstream according to the non-first decoding scheme, and the reconstructed second audio signal is aligned to obtain the reconstructed HOA signal of the current frame. The alignment processing makes the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme.
[0187] As known from the foregoing, if the encoding scheme of the current frame is the first encoding scheme, the decoding end can obtain the reconstructed HOA signal of the current frame by decoding the bitstream according to the first encoding scheme, without other processing. To solve the problem that the decoding delays of different encoding and decoding schemes are different, if the decoding scheme of the current frame is the non-first decoding scheme, the decoding end first reconstructs the second audio signal according to the bitstream, and then needs to perform alignment processing on the second audio signal, or perform alignment processing based on the second audio signal, to obtain the reconstructed HOA signal of the current frame. The alignment processing makes the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme. It should be noted that the decoding delay referred to herein is the end-to-end encoding and decoding delay, which can also be regarded as the encoding delay. The encoding delays of the three encoding schemes are consistent, and the decoding delays need to be aligned according to the decoding method provided in the embodiments of the present application.
[0188] In the embodiments of the present application, the decoding scheme of the current frame being the non-first encoding scheme includes two cases, i.e., the decoding scheme of the current frame being the second decoding scheme, and the decoding scheme of the current frame being the third decoding scheme, i.e., the non-first decoding scheme being the second decoding scheme or the third decoding scheme. Next, the decoding processes of the two cases will be introduced.
[0189] In the embodiments of the present application, if the decoding scheme of the current frame is the third decoding scheme, i.e., the current frame is a switching frame, the decoding end reconstructs the signal of the specified channel according to the bitstream, and takes the reconstructed signal of the specified channel as the reconstructed second audio signal. The specified channel is part of all channels of the HOA signal of the current frame. Correspondingly, the decoding end performs alignment processing on the reconstructed signal of the specified channel to obtain the reconstructed HOA signal of the current frame.
[0190] In the embodiments of the present application, the process of reconstructing the signal of the specified channel at the decoding end is symmetrical with, i.e. matches, the process of encoding the signal of the specified channel into the bitstream at the encoding end. Assuming that the encoding end determines the virtual loudspeaker signal and the residual signal based on the signal of the specified channel in the HOA signal of the current frame, and encodes the virtual loudspeaker signal and the residual signal into the bitstream, then the decoding end determines the virtual loudspeaker signal and the residual signal from the bitstream, and reconstructs the signal of the specified channel based on the virtual loudspeaker signal and the residual signal. Illustratively, the decoding end parses the virtual loudspeaker signal and the residual signal from the bitstream through a core decoder, which can be a stereo decoder or a monaural decoder.
[0191] For the switching frame, after the decoding end reconstructs the signal of the specified channel, the decoding end performs analysis filtering on the reconstructed signal of the specified channel, and determines the gains of one or more remaining channels other than the specified channel in the HOA signal of the current frame based on the analysis-filtered signal of the specified channel. The decoding end determines the signals of the one or more remaining channels based on the gains of the one or more remaining channels and the analysis-filtered signal of the specified channel. The decoding end performs synthesis filtering on the analysis-filtered signal of the specified channel and the signals of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame. That is, for the switching frame, the alignment processing includes reconstructing the signals of the remaining channels and the time-delay alignment processing based on the analysis-synthesis filtering. The decoding end increases the decoding time delay of the switching frame through the analysis-synthesis filtering processing, so that the decoding time delay of the switching frame is consistent with the decoding time delay of the first decoding scheme. The analysis-synthesis filtering processing includes the analysis filtering processing and the synthesis filtering processing.
[0192] Illustratively, assuming that the signal of the specified channel is the lower-order part of the HOA signal of the current frame, and the signals of the one or more remaining channels are the higher-order part of the HOA signal of the current frame, the decoding end performs analysis filtering on the reconstructed lower-order part of the HOA signal, and determines the higher-order gains of the current frame based on the analysis-filtered lower-order part of the HOA signal. The higher-order gains include the gains of the channels included in the higher-order part of the HOA signal. The decoding end determines the higher-order part of the HOA signal of the current frame based on the analysis-filtered lower-order part of the HOA signal and the higher-order gains. The decoding end performs synthesis filtering on the analysis-filtered lower-order part of the HOA signal and the higher-order part to obtain the reconstructed HOA signal of the current frame. That is, in the case where the signal of the specified channel is the lower-order part of the HOA signal of the current frame, the alignment processing corresponding to the switching frame includes reconstructing the higher-order part and the time-delay alignment processing based on the analysis-synthesis filtering.
[0193] Optionally, the specified channel is consistent with a preset transmission channel in the first decoding scheme (or the first codec scheme or the first encoding scheme). Optionally, the specified channel includes a first-order ambisonics (FOA) channel, the signal of the specified channel includes a FOA signal, and the FOA signal includes an omni-directional W signal and directional X, Y and Z signals. The FOA signal is a low-order part of the HOA signal.
[0194] Exemplarily, the decoder inputs the reconstructed low-order part of the HOA signal into an analysis filter to perform analysis filtering on the reconstructed low-order part of the HOA signal through the analysis filter, so as to obtain an analysis-filtered low-order part of the HOA signal. Based on the analysis-filtered low-order part of the HOA signal, a high-order gain of the current frame is determined, and an analysis-filtered high-order part is determined based on the analysis-filtered low-order part of the HOA signal and the high-order gain. The analysis-filtered low-order part and the high-order part of the HOA signal are subjected to a synthesis filtering through a synthesis filter to obtain a reconstructed HOA signal of the current frame output by the synthesis filter. That is, the analysis synthesis filter adds a time delay for the current frame. The analysis synthesis filter is the same as the analysis synthesis filter used in the DirAC-based HOA codec scheme, so that the time delay added by the same analysis synthesis filter to the first HOA signal of the current frame is consistent with the processing time delay of the analysis synthesis filter in the DirAC-based HOA decoding scheme, and thus the decoding time delay of the current frame is consistent with the decoding time delay of the DirAC-based HOA decoding scheme. For example, the analysis synthesis filtering adds a time delay of 5 ms, and thus the HOA signal of the current frame is output 5 ms later than in the case without analysis synthesis filtering, so as to achieve the purpose of time delay alignment.
[0195] The analysis synthesis filter can be a complex domain low delay filter bank (CLDFB) or other filter with a time delay characteristic.
[0196] The above describes the process of alignment processing of the switching frame to achieve time delay alignment. Next, the process of alignment processing of the audio frame with the second decoding scheme to achieve time delay alignment is described.
[0197] In the embodiments of the present application, if the decoding scheme of the current frame is the second decoding scheme, the decoder first reconstructs the first HOA signal according to the code stream in the second decoding scheme, and the reconstructed first HOA signal is the reconstructed second audio signal. Then, the decoder performs alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame.
[0198] The implementation process of the decoding end reconstructing the first HOA signal according to the code stream according to the second decoding scheme is as follows: the decoding end parses the virtual loudspeaker signal and the residual signal from the code stream through a core decoder, and inputs the parsed virtual loudspeaker signal and the residual signal into an MP-based spatial decoder to reconstruct the first HOA signal. It should be noted that the process of the decoding end reconstructing the first HOA signal according to the code stream according to the second decoding scheme corresponds to the process of the encoding end encoding the HOA signal of the current frame into the code stream according to the second encoding scheme, and the virtual loudspeaker signal and the residual signal in the second encoding and decoding scheme are different from the virtual loudspeaker signal and the residual signal in the switching frame encoding scheme.
[0199] Optionally, the decoding end performs alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame. There are various ways to perform mode alignment, such as performing mode alignment through analysis-synthesis filtering processing to align the time delay, or performing mode alignment through circular buffer processing to align the time delay. Next, the time delay alignment processing based on analysis-synthesis filtering and the time delay alignment processing based on circular buffer will be introduced respectively.
[0200] First, the implementation process of the time delay alignment processing based on analysis-synthesis filtering for the current frame with the decoding scheme being the second decoding scheme is introduced. In the embodiment of the present application, after the decoding end reconstructs the first HOA signal, the decoding end performs analysis-synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame. That is, for the current frame encoded by the MP-based HOA encoding scheme, the decoding end first reconstructs the HOA signal of the current frame based on the code stream according to the second decoding scheme, that is, reconstructs the first HOA signal, and then performs time delay alignment through analysis-synthesis filtering processing.
[0201] For example, the decoding end inputs the reconstructed first HOA signal into an analysis-synthesis filter to obtain the reconstructed HOA signal of the current frame output by the analysis-synthesis filter. That is, the analysis-synthesis filter adds a time delay to the current frame. The analysis-synthesis filter is the same as the analysis-synthesis filter used in the DirAC-based HOA decoding scheme, so that the time delay added to the first HOA signal of the current frame through the same analysis-synthesis filter is consistent with the processing time delay of the analysis-synthesis filter in the DirAC-based HOA decoding scheme, and thus the decoding time delay of the current frame is consistent with the decoding time delay of the DirAC-based HOA decoding scheme. The analysis-synthesis filter can be a complex domain low delay filter bank (CLDFB) or other filters with time delay characteristics.
[0202] The energy of the high order part of the HOA signal decoded by the DirAC-based HOA decoding scheme is large, and the energy of the high order part of the HOA signal decoded by the MP-based HOA decoding scheme is small. Based on this, in the embodiments of the present application, in order to make the energy of the high order part of the reconstructed HOA signals of adjacent audio frames differ less, so as to make the auditory quality transition smoothly, the decoding end can also perform gain adjustment on the high order part of the HOA signal decoded by the MP-based HOA decoding scheme, so as to make the energy of the gain-adjusted high order part increased.
[0203] Optionally, the decoding end performs analysis filtering processing on the reconstructed first HOA signal, to obtain a second HOA signal. The decoding end performs gain adjustment on the high order part of the second HOA signal, to obtain a gain-adjusted high order part. The decoding end performs synthesis filtering processing on the low order part of the second HOA signal and the gain-adjusted high order part, to obtain the reconstructed HOA signal of the current frame. It should be noted that in this case, the alignment processing can be considered to include the high order gain adjustment and the time delay alignment processing based on analysis and synthesis filtering.
[0204] Optionally, if the decoding scheme of the current frame is the second decoding scheme, and the decoding scheme of the previous frame of the current frame is the third decoding scheme, that is, the previous frame of the current frame is the switching frame, the decoding end performs gain adjustment on the high order part of the second HOA signal according to the high order gain of the previous frame of the current frame, to obtain a gain-adjusted high order part. That is, for an audio frame adjacent to the switching frame and decoded by the MP-based HOA decoding scheme after the switching frame, the decoding end can use the high order gain of the switching frame before the audio frame to adjust the high order part of the HOA signal of the audio frame, so as to make the energy of the high order part of the reconstructed HOA signal of the audio frame finally obtained close to the energy of the high order part of the reconstructed HOA signal of the switching frame, and realize the smooth transition of the auditory quality. Optionally, the audio frame located after the switching frame and adjacent to the switching frame in the decoding process can be referred to as an MP decoding high order gain adjustment frame, and the decoding end needs to perform high order gain adjustment and time delay alignment processing based on analysis and synthesis filtering on the MP decoding high order gain adjustment frame. Optionally, for the MP decoding high order gain adjustment frame, the high order gain used for performing the high order gain adjustment can be the high order gain of the previous frame, or can be a high order gain obtained according to other manners, and the embodiments of the present application do not limit this.
[0205] Optionally, if the coding scheme of the current frame is the second decoding scheme and the decoding scheme of the previous frame of the current frame is the second decoding scheme, i.e., the previous frame of the current frame is not a switching frame, the decoding end can also perform gain adjustment on the high-order part of the second HOA signal of the current frame by using a high-order gain to obtain a gain-adjusted high-order part. It should be noted that the embodiments of the present application do not limit the method for obtaining the high-order gain. The high-order gain can be the high-order gain of the previous frame of the current frame, can be determined according to the high-order gain of the previous frame and a preset gain adjustment function, or can be determined by other methods.
[0206] Optionally, in addition to gain adjustment on the high-order part, the decoding end can also perform gain adjustment on other parts of the HOA signal of the audio frame with the second decoding scheme. That is, the embodiments of the present application do not limit which channels of the HOA signal are gain-adjusted. In other words, the decoding end can perform gain adjustment on the signals of any one or more channels of the HOA signal, for example, the gain-adjusted channels can include all or part of the high-order channels, or all or part of the remaining channels except the specified channels, or other channels.
[0207] Taking gain adjustment on the signals of one or more remaining channels except the specified channels as an example, after the decoding end performs analysis filtering processing on the reconstructed first HOA signal to obtain a second HOA signal, the decoding end performs gain adjustment on the signals of one or more remaining channels of the second HOA signal to obtain gain-adjusted signals of the one or more remaining channels. The one or more remaining channels are channels of the HOA signal except the specified channels. The decoding end performs synthesis filtering processing on the signals of the specified channels of the second HOA signal and the gain-adjusted signals of the one or more remaining channels to obtain a reconstructed HOA signal of the current frame. Optionally, if the decoding scheme of the previous frame of the current frame is the third decoding scheme, the decoding end performs gain adjustment on the signals of the one or more remaining channels of the second HOA signal according to the gain of the one or more remaining channels of the previous frame of the current frame to obtain gain-adjusted signals of the one or more remaining channels. That is, for the HOA signal of the audio frame coded by using the second decoding scheme, the decoding end performs gain adjustment on the signals of the remaining channels except the specified channels. If the previous frame of the current frame is a switching frame, the decoding end performs gain adjustment on the signals of the remaining channels of the current frame based on the gain of the remaining channels of the switching frame, so that the signal strength of the remaining channels of the current frame is close to the signal strength of the remaining channels of the switching frame, and the transition of the auditory quality is smoother.
[0208] It should be noted that in the embodiments of the present application, for the current frame with the second decoding scheme, the decoding end can perform time delay alignment by using time delay alignment processing based on analysis-synthesis filtering.
[0209] Then the implementation process of the delay alignment processing based on the circular buffer for the current frame with the second decoding scheme is introduced. In the embodiment of the present application, after the first HOA signal is reconstructed at the decoding end, if the decoding scheme of the current frame is the second decoding scheme and the decoding scheme of the previous frame of the current frame is the second decoding scheme, that is, the previous frame of the current frame is a non-switching frame, the decoding end performs the circular buffer processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame. That is, for the current frame with the second decoding scheme and the previous frame being a non-switching frame, the decoding end can also perform the delay alignment based on the circular buffer to perform the delay alignment. For the current frame with the second decoding scheme and the previous frame being a switching frame, the decoding end still performs the delay alignment based on the analysis-synthesis filtering to perform the delay alignment.
[0210] Optionally, the implementation process of the decoding end performing the circular buffer processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame is as follows: the decoding end acquires first data, merges the first data and second data to obtain the reconstructed HOA signal of the current frame. The first data is the data in the previous frame HOA signal of the current frame between a first time and the end time of the previous frame HOA signal, the time length between the first time and the end time is a first time length, that is, the first time is a time point before the end time and a first time length away from the end time, and the first time length is equal to the encoding delay difference between the first decoding scheme and the second decoding scheme. The second data is the data in the reconstructed first HOA signal between a start time of the reconstructed first HOA signal and a second time, the time length between the second time and the start time is a second time length, that is, the second time is a time point after the start time and a second time length away from the start time, and the sum of the first time length and the second time length is equal to the frame length of the current frame. It should be noted that in this case, the previous frame of the current frame is also an audio frame encoded based on the MP-based HOA encoding scheme, that is, the decoding scheme of the previous frame of the current frame is also the second decoding scheme, and in the decoding process of the previous frame of the current frame, a first HOA signal also needs to be reconstructed, and the previous frame HOA signal in the circular buffer processing refers to the reconstructed first HOA signal of the previous frame.
[0211] Optionally, after the decoding end merges the first data and the second data to obtain the reconstructed HOA signal of the current frame, the third data is buffered, and the third data is the data in the reconstructed first HOA signal except the second data. The third data is used for decoding of the next frame of the current frame.
[0212] For example, assuming that the difference in encoding delay between the first encoding scheme and the second encoding scheme is 5 ms (milliseconds), the frame length of the current frame is 20 ms, the first data is the buffered 5 ms data, and the 5 ms data is the last 5 ms data of the HOA signal of the previous frame of the current frame, the decoding end obtains the buffered 5 ms data, merges the 5 ms data with the first 15 ms data of the reconstructed first HOA signal of the current frame to obtain the reconstructed HOA signal of the current frame. In addition, the decoding end buffers the last 5 ms data of the reconstructed first HOA signal of the current frame for decoding of the next frame of the current frame. For example, assuming that the current buffered data is the last 5 ms data corresponding to the i-th frame, i is a positive integer, if the decoding scheme of the (i+1)-th frame is the second decoding scheme, when the (i+1)-th frame is decoded, the decoding end reconstructs the first HOA signal of the (i+1)-th frame, obtains the buffered 5 ms data, and merges the obtained 5 ms data with the first 15 ms data of the reconstructed first HOA signal of the (i+1)-th frame to obtain the reconstructed HOA signal of the (i+1)-th frame. If the decoding scheme of the (i+1)-th frame is the switching frame decoding scheme, when the (i+1)-th frame is decoded, the decoding end obtains the buffered 5 ms data, and in the process of decoding the switching frame based on the analysis-synthesis filtering processing, the 5 ms data is processed in the analysis-synthesis filter and merged with the first 15 ms data corresponding to the (i+1)-th frame to be the reconstructed HOA signal of the current frame.
[0213] As can be seen from the above description, in the embodiments of the present application, for the switching frame, the decoding end decodes the switching frame according to the switching frame decoding scheme, that is, the switching frame needs to be subjected to residual channel signal reconstruction (for example, high-order part reconstruction) and delay alignment processing based on analysis-synthesis filtering. For the audio frame whose decoding scheme is the second decoding scheme, the decoding end performs the delay alignment processing based on analysis-synthesis filtering, and optionally, high-order gain adjustment can also be performed.
[0214] Alternatively, for the switching frame, the decoding end decodes the switching frame according to the switching frame decoding scheme. For the audio frame whose decoding scheme is the second decoding scheme and whose previous frame is a switching frame, the decoding end performs the delay alignment processing based on analysis-synthesis filtering, and optionally, high-order gain adjustment can also be performed. For the audio frame whose decoding scheme is the second decoding scheme and whose previous frame is not a switching frame, the decoding end performs the delay alignment processing based on circular buffering. For the first audio frame to be decoded, if the decoding scheme of the first audio frame is the second decoding scheme, the decoding end performs the delay alignment processing based on analysis-synthesis filtering or the delay alignment processing based on circular buffering.
[0215] Figure 9 is a kind of encoding schematic diagram of the encoding scheme switching provided by the embodiments of the present application. Referring to Figure 9, the current frame is a switching frame, and the current frame is encoded based on the MP-W HOA encoding scheme (i.e., a switching frame encoding scheme). A previous frame of the current frame is a DirAC encoded frame, and the previous frame is encoded based on the DirAC HOA encoding scheme. A next frame of the current frame is an MP encoded frame, and the next frame is encoded based on the MP HOA encoding scheme. Figure 9 The switching state of the switching frame is a first switching state, and the first switching state refers to a state of switching from the DirAC HOA encoding scheme to the MP HOA encoding scheme. The DirAC encoded frame refers to an audio frame with a first encoding scheme, and the MP encoded frame refers to an audio frame with a second encoding scheme.
[0216] Figure 10 is a decoding diagram of the encoding scheme switching provided by an embodiment of the present application. Figure 10 The decoding process when the switching state of the switching frame is the first switching state as shown in Figure 9 is shown. Referring to Figure 10 , the current frame is a switching frame, and the current frame is decoded based on the MP-W HOA decoding scheme. A previous frame of the current frame is a DirAC decoded frame, and the previous frame is decoded based on the DirAC HOA decoding scheme. A next frame of the current frame is an MP decoded high-order gain adjustment frame, and the next frame is decoded based on the MP HOA decoding scheme, and is subjected to delay alignment processing based on analysis-synthesis filtering and high-order gain adjustment. MP decoded frames located between the MP decoded high-order gain adjustment frame and a next switching frame, i.e., subsequent MP decoded frames, are decoded based on the MP HOA decoding scheme, and are subjected to delay alignment processing based on analysis-synthesis filtering. The DirAC decoded frame refers to an audio frame with a first decoding scheme, and the MP decoded frame refers to an audio frame with a second decoding scheme.
[0217] Figure 11 is another decoding diagram of the encoding scheme switching provided by an embodiment of the present application. Figure 11 The decoding process when the switching state of the switching frame is the first switching state as shown in Figure 9 is shown. Referring to Figure 11 , Figure 11 The decoding process shown in Figure 10 is different from the decoding process shown in that MP decoded frames located between the MP decoded high-order gain adjustment frame and a next switching frame, i.e., subsequent MP decoded frames, are decoded based on the MP HOA decoding scheme, and are subjected to delay alignment processing based on circular buffering.
[0218] From the above, in the embodiments of the present application, in the case that the switching state of the switching frame is the first switching state, that is, it is required to switch from the DirAC-based HOA encoding scheme to the MP-based HOA encoding scheme, that is, it is required to switch from a large delay to a small delay, since the MP-based HOA decoding scheme itself has a small decoding delay and does not include delay alignment processing, then it is required to perform delay alignment processing on the MP-decoded frame after the switching frame. For the switching frame, the switching frame encoding scheme provided by the present solution itself can be considered to include delay alignment processing. In the case that the switching state of the switching frame is the second switching state, that is, it is required to switch from a small delay to a large delay, since the DirAC-based HOA decoding scheme itself has a large delay, then it is not required to perform additional processing on the DirAC-decoded frame after the switching frame.
[0219] In summary, in the embodiments of the present application, since the decoding delay of the HOA decoding scheme based on directional audio coding is large, for the current frame encoded by the first encoding scheme, the code stream of the current frame can be decoded according to the first decoding scheme. For the current frame not encoded by the first encoding scheme, the second audio signal is first reconstructed according to the code stream, and then the reconstructed second audio signal is aligned to obtain the reconstructed HOA signal of the current frame, that is, the decoding delay of the current frame is made consistent with the decoding delay of the first decoding scheme through the alignment processing. In this way, the decoding delay of each audio frame can be made consistent by using the present solution, that is, the delay alignment is ensured to enable the different encoding and decoding schemes to be switched well.
[0220] Figure 12 FIG. 12 is a structural schematic diagram of a decoding device 1200 provided by the embodiments of the present application. The decoding device 1200 can be realized by software, hardware or a combination of both to become part or all of a decoding end device. The decoding end device can be any decoding end device in the above embodiments. Figure 12 The decoding device 1200 includes a first determination module 1201, a first decoding module 1202 and a second decoding module 1203.
[0221] The first determination module 1201 is configured to determine the decoding scheme of the current frame according to the code stream. The decoding scheme of the current frame is the first decoding scheme or a non-first decoding scheme. The first decoding scheme is a high-order ambisonics (HOA) decoding scheme based on directional audio coding (DirAC).
[0222] The first decoding module 1202 is configured to, if the decoding scheme of the current frame is the first decoding scheme, reconstruct a first audio signal according to the code stream according to the first decoding scheme. The reconstructed first audio signal is a reconstructed HOA signal of the current frame.
[0223] The second decoding module 1203 is configured to, if the decoding scheme of the current frame is not the first decoding scheme, reconstruct the second audio signal according to the bitstream according to the non-first decoding scheme, perform alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, and the alignment processing makes the decoding delay of the current frame consistent with the decoding delay of the first decoding scheme.
[0224] Optionally, the non-first decoding scheme is a second decoding scheme or a third decoding scheme, the second decoding scheme is an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme is a hybrid decoding scheme.
[0225] The second decoding module 1203 includes:
[0226] The first reconstruction submodule is configured to, if the decoding scheme of the current frame is the third decoding scheme, reconstruct the signal of the specified channel according to the bitstream, and the reconstructed signal of the specified channel is the reconstructed second audio signal, and the specified channel is part of the channels of the HOA signal of the current frame.
[0227] Optionally, the second decoding module 1203 includes:
[0228] The analysis filtering submodule is configured to perform analysis filtering processing on the reconstructed signal of the specified channel.
[0229] The first determination submodule is configured to determine the gain of one or more remaining channels of the HOA signal of the current frame other than the specified channel based on the analysis-filtered signal of the specified channel.
[0230] The second determination submodule is configured to determine the signal of the one or more remaining channels based on the gain of the one or more remaining channels and the analysis-filtered signal of the specified channel.
[0231] The synthesis filtering submodule is configured to perform synthesis filtering processing on the analysis-filtered signal of the specified channel and the signal of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
[0232] Optionally, the non-first decoding scheme is a second decoding scheme or a third decoding scheme, the second decoding scheme is an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme is a hybrid decoding scheme.
[0233] The second decoding module 1203 includes:
[0234] The second reconstruction submodule is configured to, if the decoding scheme of the current frame is the second decoding scheme, reconstruct the first HOA signal according to the bitstream according to the second decoding scheme, and the reconstructed first HOA signal is the reconstructed second audio signal.
[0235] Optionally, the second decoding module 1203 includes:
[0236] an analysis-synthesis filtering sub-module, configured to perform analysis-synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
[0237] Optionally, the analysis-synthesis filtering sub-module is configured to:
[0238] perform analysis filtering processing on the reconstructed first HOA signal to obtain a second HOA signal;
[0239] perform gain adjustment on the signals of one or more residual channels in the second HOA signal to obtain gain-adjusted signals of the one or more residual channels, the one or more residual channels being channels in the HOA signal other than the designated channels;
[0240] perform synthesis filtering processing on the signals of the designated channels in the second HOA signal and the gain-adjusted signals of the one or more residual channels to obtain the reconstructed HOA signal of the current frame.
[0241] Optionally, the analysis-synthesis filtering sub-module is configured to:
[0242] if the decoding scheme of the previous frame of the current frame is the third decoding scheme, perform gain adjustment on the signals of the one or more residual channels in the second HOA signal according to the gain of the one or more residual channels of the previous frame of the current frame to obtain gain-adjusted signals of the one or more residual channels.
[0243] Optionally, the designated channels include first-order ambisonic (FOA) channels.
[0244] Optionally, the decoding scheme of the previous frame of the current frame is the second decoding scheme.
[0245] the second decoding module 1203 comprises:
[0246] a circular buffering sub-module, configured to perform circular buffering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
[0247] Optionally, the circular buffering sub-module is configured to:
[0248] obtain first data, the first data being data in the HOA signal of the previous frame of the current frame located between a first time and an end time of the HOA signal of the previous frame, a time length between the first time and the end time being a first time length, the first time length being equal to a coding time delay difference between the first decoding scheme and the second decoding scheme;
[0249] merge the first data and second data to obtain a reconstructed HOA signal of the current frame, the second data being data of the reconstructed first HOA signal located between a start time of the reconstructed first HOA signal and a second time, a time length between the second time and the start time being a second time length, a sum of the first time length and the second time length being equal to a frame length of the current frame.
[0250] Optionally, the circular buffer submodule is configured to:
[0251] buffer third data, the third data being data of the reconstructed first HOA signal other than the second data.
[0252] Optionally, the first determining module 1201 comprises:
[0253] a first parsing submodule configured to parse a value of a switching flag of the current frame from the bitstream;
[0254] a second parsing submodule configured to, if the value of the switching flag is a first value, parse indication information of a decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is a first decoding scheme or a second decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection;
[0255] a third determining submodule configured to, if the value of the switching flag is a second value, determine that the decoding scheme of the current frame is a third decoding scheme, the third decoding scheme being a hybrid decoding scheme.
[0256] Optionally, the first determining module 1201 comprises:
[0257] a third parsing submodule configured to parse indication information of a decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme or the third decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection, the third decoding scheme being a hybrid decoding scheme.
[0258] Optionally, the first determining module 1201 comprises:
[0259] a fourth parsing submodule configured to parse an initial decoding scheme of the current frame from the bitstream, the initial decoding scheme being the first decoding scheme or the second decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection;
[0260] a fourth determining submodule configured to, if the initial decoding scheme of the current frame is the same as an initial decoding scheme of a previous frame of the current frame, determine that the decoding scheme of the current frame is the initial decoding scheme of the current frame.
[0261] The fifth determining sub-module is configured to determine the decoding scheme of the current frame as a third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme, the third decoding scheme being a hybrid decoding scheme.
[0262] In the embodiments of the present application, since the decoding delay of the DirAC-based HOA decoding scheme is large, for the current frame encoded by the first encoding scheme, the code stream of the current frame is decoded according to the first decoding scheme. For the current frame not encoded by the first encoding scheme, the second audio signal is reconstructed according to the code stream, and the reconstructed second audio signal is aligned to obtain the reconstructed HOA signal of the current frame, that is, the decoding delay of the current frame is made consistent with the decoding delay of the first decoding scheme through the alignment processing. In this way, the decoding delay of each audio frame is consistent by using the present scheme, that is, the delay alignment is ensured to enable the different encoding and decoding schemes to be switched well.
[0263] It should be noted that the decoding apparatus provided in the above embodiments is only used for example to divide the above functions into different functional modules when decoding the audio frame, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the above described functions. In addition, the decoding apparatus and the decoding method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be described here.
[0264] Figure 13 A schematic block diagram of a coding and decoding apparatus 1300 for the embodiments of the present application is shown. The coding and decoding apparatus 1300 can include a processor 1301, a memory 1302 and a bus system 1303. The processor 1301 and the memory 1302 are connected through the bus system 1303. The memory 1302 is configured to store instructions, and the processor 1301 is configured to execute the instructions stored in the memory 1302 to perform various encoding or decoding methods described in the embodiments of the present application. To avoid repetition, the detailed description will not be given here.
[0265] In the embodiments of the present application, the processor 1301 can be a central processing unit (CPU), and the processor 1301 can also be other general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0266] The memory 1302 can include a ROM device or a RAM device. Any other suitable type of memory device can also be used as the memory 1302. The memory 1302 can include code and data 13021 that is accessed by the processor 1301 using a bus 1303. The memory 1302 can further include an operating system 13023 and application programs 13022, including at least one program that permits the processor 1301 to perform the encoding or decoding methods described in the embodiments herein. For example, the application programs 13022 can include applications 1 through N, which further include an encoding or decoding application (codec application for short) that performs the encoding or decoding methods described in the embodiments herein.
[0267] The bus system 1303 can include, in addition to the data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, all buses are referred to as the bus system 1303 in the figure.
[0268] Optionally, the codec apparatus 1300 can also include one or more output devices, such as a display 1304. In one example, the display 1304 can be a touch-sensitive display that combines a display with a touch-sensitive unit that is operable to sense touch inputs. The display 1304 can be connected to the processor 1301 via the bus 1303.
[0269] It is noted that the codec apparatus 1300 can perform the encoding methods in the embodiments herein, and also perform the decoding methods in the embodiments herein.
[0270] Those skilled in the art will appreciate that the functions described with reference to the various illustrative logical blocks, modules, and algorithm steps described in this specification can be implemented as hardware, software, firmware, or any combination thereof. If implemented in software, the functions described with reference to the various illustrative logical blocks, modules, and steps described in this specification can be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally can correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this specification. A computer program product can include a computer-readable medium.
[0271] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code means in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to non-transient, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, DVD, and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0272] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein can refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein as being performed by various illustrative logical blocks, modules, and steps can be implemented in hardware and / or software implemented within a special- purpose hardware and / or software module associated with a combined codec, or incorporated into a combined codec. Moreover, the techniques can be embodied completely within one or more circuits or logic elements. In one example, the various illustrative logical blocks, units, modules, and steps described in the encoder 100 and the decoder 200 can be understood as corresponding to respective circuitry or logic elements.
[0273] The techniques of this disclosure can be implemented in a wide variety of over-the-top (OTT) devices or apparatuses, including wireless handsets, integrated circuit (IC) or set of ICs (e.g., a chip set). Various components, modules, or units described in the embodiments of this disclosure can be enabled by hardware, software, firmware or any combination thereof. Moreover, alternate implementations having more or fewer features than those depicted in the figures are also possible. For example, various components, modules, or units described herein can be enabled by one or more processors configured with software instructions. In another example, various components, modules, or units described herein can be enabled by one or more circuits, such as one or more analog or digital circuits.
[0274] That is, in the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website site, computer, server or data center to another website site, computer, server or data center through wired (for example: coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example: infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (for example: floppy disk, hard disk, magnetic tape), an optical medium (for example: digital versatile disc (DVD)) or a semiconductor medium (for example: solid state disk (SSD)) and the like. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.
[0275] It should be understood that "at least one" mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, "first", "second" and the like are used to distinguish the same items or similar items with basically the same function and role in the embodiments of the present application. Those skilled in the art can understand that "first", "second" and the like do not limit the number and execution order, and "first", "second" and the like do not necessarily mean different.
[0276] The above describes the embodiments provided by the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A decoding method, comprising: The method comprises: determining a decoding scheme of a current frame according to a bitstream, the decoding scheme of the current frame being a first decoding scheme or a non-first decoding scheme, the first decoding scheme being a high-order ambisonic (HOA) decoding scheme based on directional audio coding (DirAC), the non-first decoding scheme being a second decoding scheme or a third decoding scheme, the second decoding scheme being an HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme being a hybrid decoding scheme, which refers to a scheme of using both the first decoding scheme and the second decoding scheme in a decoding process; if the decoding scheme of the current frame is the first decoding scheme, reconstructing a first audio signal according to the bitstream according to the first decoding scheme, the reconstructed first audio signal being a reconstructed HOA signal of the current frame; if the decoding scheme of the current frame is the non-first decoding scheme, reconstructing a second audio signal according to the bitstream according to the non-first decoding scheme, and performing alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, the alignment processing making a decoding time delay of the current frame consistent with a decoding time delay of the first decoding scheme; wherein the determining of the decoding scheme of the current frame according to the bitstream comprises: parsing a value of a switching flag of the current frame from the bitstream, parsing indication information of the decoding scheme of the current frame from the bitstream if the value of the switching flag is a first value, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme, and determining that the decoding scheme of the current frame is the third decoding scheme if the value of the switching flag is a second value; or parsing indication information of the decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme or the third decoding scheme; or parsing an initial decoding scheme of the current frame from the bitstream, the initial decoding scheme being the first decoding scheme or the second decoding scheme, determining that the decoding scheme of the current frame is the initial decoding scheme of the current frame if the initial decoding scheme of the current frame is the same as an initial decoding scheme of a previous frame of the current frame, and determining that the decoding scheme of the current frame is the third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or if the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme.
2. The method of claim 1, wherein, if the decoding scheme of the current frame is the third decoding scheme, the reconstructing of the second audio signal according to the bitstream comprises: reconstructing a signal of a specified channel according to the bitstream, the reconstructed signal of the specified channel being the reconstructed second audio signal, and the specified channel being part of all channels of the HOA signal of the current frame.
3. The method of claim 2, wherein, The alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame comprises: analyzing filtering processing on the reconstructed specified channel signal; determining gains of one or more remaining channels of the HOA signal of the current frame except the specified channel based on the analyzed filtered specified channel signal; determining signals of the one or more remaining channels based on the gains of the one or more remaining channels and the analyzed filtered specified channel signal; synthesis filtering processing on the analyzed filtered specified channel signal and the signals of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
4. The method of claim 1, wherein, If the decoding scheme of the current frame is the second decoding scheme, the reconstructing a second audio signal according to the bitstream comprises: reconstructing a first HOA signal according to the bitstream in the second decoding scheme, and the reconstructed first HOA signal is the reconstructed second audio signal.
5. The method of claim 4, wherein, The alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame comprises: analyzing synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
6. The method of claim 5, wherein, The analyzing synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame comprises: analyzing filtering processing on the reconstructed first HOA signal to obtain a second HOA signal; gain adjustment on signals of one or more remaining channels of the second HOA signal to obtain gain-adjusted signals of the one or more remaining channels, the one or more remaining channels being channels of the HOA signal except a specified channel; synthesis filtering processing on a signal of the specified channel of the second HOA signal and the gain-adjusted signals of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
7. The method of claim 6, wherein, The gain adjustment on signals of one or more remaining channels of the second HOA signal to obtain gain-adjusted signals of the one or more remaining channels comprises: if the decoding scheme of a previous frame of the current frame is the third decoding scheme, gain adjustment on the signals of the one or more remaining channels of the second HOA signal according to the gains of the one or more remaining channels of the previous frame of the current frame to obtain the gain-adjusted signals of the one or more remaining channels.
8. The method of any one of claims 2-3, 6-7, wherein, The specified channel comprises a first-order ambisonic (FOA) channel.
9. The method of claim 4, wherein, The decoding scheme of a previous frame of the current frame is the second decoding scheme. The alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame comprises: circular buffer processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
10. The method of claim 9, wherein, The circular buffer processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame comprises: obtaining first data, the first data being data in a previous frame HOA signal of the current frame and located between a first time and an end time of the previous frame HOA signal, a time length between the first time and the end time being a first time length, the first time length being equal to a coding time delay difference between the first decoding scheme and the second decoding scheme; merging the first data and second data to obtain a reconstructed HOA signal of the current frame, the second data being data in the reconstructed first HOA signal and located between a start time and a second time of the reconstructed first HOA signal, a time length between the second time and the start time being a second time length, a sum of the first time length and the second time length being equal to a frame length of the current frame.
11. The method of claim 10, wherein, The method further includes: buffering third data, the third data being data in the reconstructed first HOA signal and excluding the second data.
12. A decoding apparatus, characterized by comprising: The apparatus includes: a first determining module configured to determine a decoding scheme of a current frame according to a bitstream, the decoding scheme of the current frame being a first decoding scheme or a non-first decoding scheme, the first decoding scheme being a high-order ambisonics (HOA) decoding scheme based on directional audio coding (DirAC), the non-first decoding scheme being a second decoding scheme or a third decoding scheme, the second decoding scheme being a HOA decoding scheme based on virtual loudspeaker selection, and the third decoding scheme being a hybrid decoding scheme, the hybrid decoding scheme being a scheme in which both the first decoding scheme and the second decoding scheme are used in a decoding process; a first decoding module configured to, if the decoding scheme of the current frame is the first decoding scheme, reconstruct a first audio signal according to the bitstream and according to the first decoding scheme, the reconstructed first audio signal being a reconstructed HOA signal of the current frame; a second decoding module configured to, if the decoding scheme of the current frame is the non-first decoding scheme, reconstruct a second audio signal according to the bitstream and according to the non-first decoding scheme, and perform alignment processing on the reconstructed second audio signal to obtain the reconstructed HOA signal of the current frame, the alignment processing making a decoding time delay of the current frame consistent with a decoding time delay of the first decoding scheme; wherein the first determining module includes: a first parsing sub-module configured to parse a value of a switching flag of the current frame from the bitstream; a second parsing sub-module configured to, if the value of the switching flag is a first value, parse indication information of the decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme or the second decoding scheme; a third determining sub-module configured to, if the value of the switching flag is a second value, determine that the decoding scheme of the current frame is the third decoding scheme; or, the first determining module includes: a third parsing sub-module configured to parse indication information of the decoding scheme of the current frame from the bitstream, the indication information being used to indicate that the decoding scheme of the current frame is the first decoding scheme, the second decoding scheme, or the third decoding scheme; Alternatively, the first determining module comprises: a fourth parsing sub-module, configured to parse an initial decoding scheme of the current frame from the code stream, the initial decoding scheme being the first decoding scheme or the second decoding scheme; a fourth determining sub-module, configured to determine the decoding scheme of the current frame as the initial decoding scheme of the current frame if the initial decoding scheme of the current frame is the same as an initial decoding scheme of a previous frame of the current frame; a fifth determining sub-module, configured to determine the decoding scheme of the current frame as the third decoding scheme if the initial decoding scheme of the current frame is the first decoding scheme and the initial decoding scheme of the previous frame of the current frame is the second decoding scheme, or the initial decoding scheme of the current frame is the second decoding scheme and the initial decoding scheme of the previous frame of the current frame is the first decoding scheme.
13. The apparatus of claim 12, wherein, The second decoding module comprises: a first reconstruction sub-module, configured to reconstruct a signal of a specified channel according to the code stream if the decoding scheme of the current frame is the third decoding scheme, the reconstructed signal of the specified channel being the reconstructed second audio signal, the specified channel being part of all channels of the HOA signal of the current frame.
14. The apparatus of claim 13, wherein, The second decoding module comprises: an analysis filtering sub-module, configured to perform analysis filtering processing on the reconstructed signal of the specified channel; a first determining sub-module, configured to determine a gain of one or more remaining channels of the HOA signal of the current frame other than the specified channel based on the analysis-filtered signal of the specified channel; a second determining sub-module, configured to determine a signal of the one or more remaining channels based on the gain of the one or more remaining channels and the analysis-filtered signal of the specified channel; a synthesis filtering sub-module, configured to perform synthesis filtering processing on the analysis-filtered signal of the specified channel and the signal of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
15. The apparatus of claim 12, wherein, The second decoding module comprises: a second reconstruction sub-module, configured to reconstruct a first HOA signal according to the code stream according to the second decoding scheme if the decoding scheme of the current frame is the second decoding scheme, the reconstructed first HOA signal being the reconstructed second audio signal.
16. The apparatus of claim 15, wherein, The second decoding module comprises: an analysis-synthesis filtering sub-module, configured to perform analysis-synthesis filtering processing on the reconstructed first HOA signal to obtain the reconstructed HOA signal of the current frame.
17. The apparatus of claim 16, wherein, The analysis-synthesis filtering sub-module is configured to: perform analysis filtering processing on the reconstructed first HOA signal to obtain a second HOA signal; perform gain adjustment on a signal of one or more remaining channels of the second HOA signal to obtain gain-adjusted signals of the one or more remaining channels, the one or more remaining channels being channels of the HOA signal other than a specified channel; perform synthesis filtering processing on a signal of the specified channel of the second HOA signal and the gain-adjusted signals of the one or more remaining channels to obtain the reconstructed HOA signal of the current frame.
18. The apparatus of claim 17, wherein, The analysis-synthesis filtering sub-module is configured to: if a decoding scheme of a previous frame of the current frame is the third decoding scheme, gain adjusting signals of the one or more residual channels in the second HOA signal according to gains of the one or more residual channels of the previous frame of the current frame to obtain the gain adjusted signals of the one or more residual channels.
19. The apparatus of any one of claims 13, 14, 17, 18, wherein, The specified channels include first-order ambisonic (FOA) channels.
20. The apparatus of claim 15, wherein, The decoding scheme of the previous frame of the current frame is the second decoding scheme. The second decoding module includes: The circular buffer submodule is configured to perform circular buffer processing on the reconstructed first HOA signal to obtain a reconstructed HOA signal of the current frame.
21. The apparatus of claim 20, wherein, The circular buffer submodule is configured to: obtain first data, the first data being data in a previous frame HOA signal of the current frame between a first time and an end time of the previous frame HOA signal, a time length between the first time and the end time being a first time length, the first time length being equal to a coding delay difference between the first decoding scheme and the second decoding scheme; merge the first data and second data to obtain a reconstructed HOA signal of the current frame, the second data being data in the reconstructed first HOA signal between a start time of the reconstructed first HOA signal and a second time, a time length between the second time and the start time being a second time length, a sum of the first time length and the second time length being equal to a frame length of the current frame.
22. The apparatus of claim 21, wherein, The circular buffer submodule is configured to: buffer third data, the third data being data in the reconstructed first HOA signal other than the second data.
23. A decoding-side device, comprising: The decoding end device includes a memory and a processor. The memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to implement the decoding method in any one of claims 1-11.
24. A computer-readable storage medium, characterized in that, The storage medium has instructions stored therein, and when the instructions are executed on the computer, the computer executes the steps of the method in any one of claims 1-11.
25. A computer program product, characterised in that, The computer program product includes instructions executed by a processor to implement the method in any one of claims 1-11. The computer program product includes instructions executed by a processor to implement the method in any one of claims 1-11.
Citation Information
Patent Citations
Signal reconstruction method and device during stereo signal coding
CN109427337A
Audio scene encoder, audio scene decoder and related methods using hybrid encoder / decoder spatial analysis
CN112074902A