Scene audio composite method and electronic device
The scene audio decoding method combines multiple encoding schemes to address the challenges of HOA signal encoding, reducing bitrate overhead and complexity while maintaining quality, enhancing the flexibility and efficiency of HOA signal processing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-01-11
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional techniques have poor encoding performance for Higher-Order Ambisonics (HOA) signals, leading to challenges in data transmission and storage due to increased data amounts with higher dimensions, necessitating improved encoding methods that balance bitrate overhead and encoding quality.
A scene audio decoding method that combines multiple encoding schemes, including a direct encoding scheme, spatial encoding scheme, and a decorrelation encoding scheme, to reduce bitrate overhead while maintaining encoding quality, allowing for flexible adaptation to various scenes and network conditions.
The combined encoding approach reduces bitrate overhead and encoding complexity while ensuring a certain level of encoding quality, improving the versatility and efficiency of HOA signal processing.
Smart Images

Figure 2026513587000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments of this application relate to the field of audio encoding and decoding, and more particularly to a scene audio decoding method and electronic device. [Background technology]
[0002] This application was filed with the China National Intellectual Property Administration on April 13, 2023, and claims priority to China Patent Application No. 202310428890.5, entitled “Scene Audio Combination Method and Electronic Device,” which is incorporated herein by reference in its entirety.
[0003] [background] Three-dimensional audio technology is an audio technology that uses computers, signal processing, and similar technologies to acquire, process, transmit, render, and reproduce acoustic events and three-dimensional sound field information in the real world. Three-dimensional audio makes it possible to impart a strong sense of spatial awareness, envelopment, and immersion to sound, providing an extraordinary "immersive" auditory experience. In Higher-Order Ambisonics (HOA) technology, the recording, encoding, and playback stages are independent of speaker placement, and data in the HOA format is rotatably reproduced. Therefore, HOA technology has greater flexibility in the reproduction of three-dimensional audio and is attracting more widespread attention and exploration.
[0004] The number of channels corresponding to the Nth-order HOA signal is (N+1). 2 As the number of dimensions of the HOA increases, the amount of information used to record a more detailed sound scene in the HOA signal also increases accordingly. However, the amount of data in the HOA signal also increases accordingly, creating challenges in both transmission and storage. Therefore, it is necessary to encode and decode the HOA signal. However, conventional techniques have poor encoding performance for HOA signals. [Overview of the Initiative]
[0005] With this in mind, this application provides a scene audio decoding method and an electronic device.
[0006] In accordance with a first aspect, one embodiment of the present application provides a scene audio coding method. The method includes: acquiring a scene audio signal; and coding the scene audio signal based on a combination of coding schemes. The combination of coding schemes includes at least one of the following combinations: a combination of a first coding scheme, a second coding scheme, and a third coding scheme, and a combination of the first coding scheme and the third coding scheme. The first coding scheme is a signal coding scheme, the second coding scheme is a spatial coding scheme, and the third coding scheme is a coding scheme other than the first and second coding schemes.
[0007] Encoding based on the first encoding scheme can improve encoding quality, but requires high bitrate overhead. Encoding based on another encoding scheme (second or third encoding scheme) can reduce bitrate overhead, but degrades encoding quality. Therefore, in this application, encoding is performed based on a combination of the first and other encoding schemes, thereby reducing bitrate overhead and encoding complexity while ensuring a certain level of encoding quality.
[0008] For example, the scene audio signal in this embodiment of the present application may be a signal used to describe a sound field. The scene audio signal may include an HOA signal (the HOA signal may include a three-dimensional HOA signal and a two-dimensional HOA signal (sometimes also called a planar HOA signal)) and a three-dimensional audio signal. The three-dimensional audio signal may be an audio signal other than the HOA signal among the scene audio signals.
[0009] For example, a scene audio signal may include an audio signal with C channels, where C is a positive integer.
[0010] For example, if the scene audio signal is an HOA signal, the HOA signal is an Nth-order HOA signal, that is, the one in equation (3) where m is truncated to the Nth term.
[0011]
number
[0012] It is possible.
[0013] For example, an Nth-order HOA signal may contain an audio signal with C channels. C = (N+1) 2 For example, if N=3, the Nth-order HOA signal contains an audio signal with 16 channels. If N=4, the Nth-order HOA signal contains an audio signal with 25 channels.
[0014] For example, the first encoding scheme, sometimes called a direct encoding scheme, specifically performs processes such as time-frequency conversion, preprocessing, bit allocation, quantization, and entropy coding on the signal. In this way, the signal is encoded. It should be noted here that encoding a signal means encoding the signal (or data) to be encoded. For example, if the signal to be encoded is a scene audio signal, encoding the scene audio signal based on the first encoding scheme may mean encoding an audio signal that has some or all of the channels within the scene audio signal.
[0015] For example, a scene audio signal may include one or more frames.
[0016] In the first embodiment, the spatial coding scheme is an coding scheme in which the attribute information of the target virtual speaker is coded, and the attribute information of the target virtual speaker is determined based on the scene audio signal.
[0017] The following points need to be noted. That is, the position of the target virtual speaker coincides with the position of the sound source in the scene audio signal. Based on the attribute information of the target virtual speaker and the audio signal having some channels in the scene audio signal, a virtual speaker signal corresponding to the target virtual speaker can be generated. And based on that virtual speaker signal, the scene audio signal can be reconstructed. Therefore, the encoder side encodes the audio signal having some channels in the scene audio signal and the attribute information of the target virtual speaker, and then transmits the encoded audio signal and the encoded attribute information to the decoder side. The decoder side can reconstruct the scene audio signal based on the reconstructed audio signal having some channels and the attribute information in the target virtual speaker obtained through decoding.
[0018] The data amount of the attribute information of the target virtual speaker is much less than the data amount of the audio signal having one channel. Therefore, compared with the encoding performed based on the first encoding method, the encoding performed based on the second encoding method requires less bitrate overhead.
[0019] The attribute information of the target virtual speaker includes at least one of the following. That is, the position information of the target virtual speaker, the position index corresponding to the position information of the target virtual speaker, or the virtual speaker index of the target virtual speaker.
[0020] For example, in the spherical coordinate system, the position information of the target virtual speaker may be, for example,
[0021]
Number
[0022] It may be. In this specification, θ s3 is the horizontal angle information of the target virtual speaker,
[0023] [Number]
[0024] is the pitch angle information of the target virtual speaker.
[0025] For example, the position index is used to uniquely identify the position of the virtual speaker. The position index may include a horizontal angle index (used to uniquely identify one horizontal angle information) and a pitch angle index (used to uniquely identify one pitch angle information). The position index of the virtual speaker has a one-to-one correspondence with the position information of the virtual speaker.
[0026] For example, the virtual speaker index may be used to uniquely identify the virtual speaker, and the position information / position index of the virtual speaker has a one-to-one correspondence with the virtual speaker index.
[0027] According to the first aspect, or any one of the implementations of the first aspect, the third encoding method includes one or more encoding methods.
[0028] It should be understood that the types and quantities of the encoding methods included in the third encoding method are not limited in this application. In this way, the third encoding method can be determined based on the current scene (such as the complexity of the content in the scene audio signal, the network bandwidth, or the same), and the encoding performance can be further improved.
[0029] According to the first aspect, or any one of the implementations of the first aspect, the third encoding method includes a channel copy encoding method. In this way, the degree of variation of the reconstructed scene audio signal obtained through decoding can be reduced, and the smoothness of the audio can be improved.
[0030] In the first embodiment, or any one of its implementations, the channel copy coding scheme is a decorrelation coding scheme. In this way, the bitrate overhead required by the third coding scheme is lower than the bitrate overhead required by the second coding scheme. Compared to coding a scene audio signal based on a combination of the first, second, and third coding schemes, coding a scene audio signal based on a combination of the first and third coding schemes reduces bitrate overhead and coding complexity.
[0031] According to the first embodiment, or any one implementation thereof, the step of encoding a scene audio signal based on a combination of encoding schemes includes: that is, the step of encoding one frame of the scene audio signal based on a plurality of encoding schemes in the combination of encoding schemes. In this way, the encoding quality of each frame of the scene audio signal can be improved.
[0032] According to the first embodiment, or any one implementation thereof, a scene audio signal comprises X frames, where X frames comprise the i-th frame and the j-th frame, and the step of encoding the scene audio signal based on a combination of encoding schemes includes: the step of encoding the i-th frame in the scene audio signal based on a plurality of encoding schemes in the combination of encoding schemes; and the step of encoding the j-th frame in the scene audio signal based on a first encoding scheme in the combination of encoding schemes. Herein, i and j are integers from 1 to X, where i is not equal to j, and X is a positive integer.
[0033] In other words, a portion of the scene audio signal frame is encoded based on multiple encoding schemes in the encoding scheme combination, while the other portion of the scene audio signal frame is encoded based on a first encoding scheme in the encoding scheme combination. This allows for flexible adaptation to encoding in various scenes (e.g., complexity of the scene audio signal content in different scenes, network bandwidth, or similar factors), thereby improving the versatility of the scene audio encoding method in this application.
[0034] According to the first embodiment, or any one implementation thereof, a scene audio signal comprises X frames, each of which comprises an i-th frame and a j-th frame, and the step of encoding the scene audio signal based on a combination of encoding schemes includes: the step of encoding the i-th frame of the scene audio signal based on a plurality of encoding schemes in a k1-th combination of encoding schemes; and the step of encoding the j-th frame of the scene audio signal based on a plurality of encoding schemes in a k2-th combination of encoding schemes. Herein, i and j are positive integers from 1 to X, i is not equal to j, and k1 and k2 are positive integers, k1 is not equal to k2.
[0035] In other words, different frames of a scene audio signal are encoded based on multiple encoding schemes in a combination of different encoding schemes. This allows for flexible adaptation to encoding in various scenes (e.g., complexity of the content of the scene audio signal in different scenes, network bandwidth, or similar factors), thereby improving the versatility of the scene audio encoding method in this application.
[0036] According to the first embodiment, or any one implementation of the first embodiment, if a frame of a scene audio signal includes an audio signal having C channels, and the combination of encoding schemes includes a combination of a first encoding scheme, a second encoding scheme, and a third encoding scheme, the step of encoding a frame of a scene audio signal based on the multiple encoding schemes in the combination of encoding schemes includes: that is, for a frame of a scene audio signal, the step of encoding C1 first channels based on the first encoding scheme; the step of encoding C2 second channels based on the second encoding scheme; and the step of encoding C3 third channels based on the third encoding scheme. C is equal to the sum of C1, C2, and C3, and C, C1, C2, and C3 are positive integers.
[0037] According to the first embodiment, or any one implementation of the first embodiment, a scene audio signal in a frame includes an audio signal having C channels, and an audio signal having one channel has Y bandwidths, where C and Y are positive integers. The step of encoding a frame of a scene audio signal based on a plurality of encoding schemes in a combination of encoding schemes includes the step of encoding Y bandwidths of one channel in a frame of a scene audio signal based on a plurality of encoding schemes in a combination of encoding schemes.
[0038] For example, the frequency range of each band may be set according to requirements. This is not limited to the present application.
[0039] According to the first embodiment, or any one of the implementations of the first embodiment, the scene audio signal is an Nth-order higher-order ambisonic HOA signal. The first C1 channel is one of the C1 channels contained in the zeroth to Mth order signals of the Nth order HOA signal. M is an integer less than N, C is equal to the square of (N+1), and C1 is less than or equal to the square of (M+1). The second set of two C channels includes the other four C channels included in the zero- to M-th order signals of the N-th order HOA signal, and five C channels other than those included in the zero- to M-th order signals of the N-th order HOA signal. The three third channels C3 include the other C6 channels included in the zeroth to Mth order signals of the Nth-order HOA signal, and the other C5 channels not included in the zeroth to Mth order signals of the Nth-order HOA signal. Here, C2 is equal to the sum of C4 and C5, C3 is equal to the sum of C6 and C7, the square of (M+1) is equal to the sum of C1, C4, and C6, and C4, C5, C6, and C7 are integers.
[0040] According to the first embodiment, or any one implementation of the first embodiment, the step of encoding C3 third channels based on a third encoding scheme includes: that is, the step of encoding a first preset identifier corresponding to the C3 third channels; the first preset identifier indicates that the third encoding scheme for the C3 third channels is a time-domain decorrelation coding scheme or a frequency-domain decorrelation coding scheme.
[0041] If the third encoding scheme is a time-domain decorrelation coding scheme, the first preset identifier is the first preset value. If the third encoding scheme is a frequency-domain decorrelation coding scheme, the first preset identifier is the second preset value. In this way, the decoder recognizes whether a time-domain decorrelation decoding scheme or a frequency-domain decorrelation decoding scheme is being used for the three third channels C.
[0042] Indeed, the encoder and decoder may also agree in advance whether to use the time-domain decorrelation encoding scheme or the frequency-domain decorrelation encoding scheme for the three third channels C3. This is not limited to the present application.
[0043] According to the first embodiment, or any one implementation thereof, if the combination of encoding schemes is a combination of a first encoding scheme, a second encoding scheme, and a third encoding scheme, the method further includes the step of encoding feature information corresponding to C2 second channels.
[0044] For example, in a scene audio signal C2 second channels The corresponding feature information can be determined based on information such as the energy and intensity of the scene audio signal. For example, the feature information could be gain information.
[0045] In this way, after obtaining feature information from the bitstream through decoding, the decoder can correct the reconstructed scene audio signal, which has C2 second channels, thereby improving the audio quality of the reconstructed audio signal with C2 second channels.
[0046] According to the first embodiment, or any one implementation thereof, the method further includes: namely, the step of encoding a second preset identifier; the second preset identifier indicates the type of encoding scheme combination; and in this way the decoder uses a specific revenge Recognizes combinations of numbering schemes.
[0047] Certainly, the encoder and decoder may agree in advance on the types of decoding scheme combinations. This is not limited to the present application.
[0048] In accordance with a second aspect, one embodiment of the present application provides a scene audio decoding method. The decoding method includes the steps of: receiving a bitstream; and decoding the bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal. The combination of decoding schemes includes at least one of the following combinations: a combination of a first decoding scheme, a second decoding scheme, and a third decoding scheme, and a combination of the first decoding scheme and the third decoding scheme. The first decoding scheme decodes encoded data obtained by encoding a signal, the second decoding scheme is a spatial decoding scheme, and the third decoding scheme is a decoding scheme other than the first and second decoding schemes.
[0049] Corresponding to the decoder side, this application performs decoding based on a combination of a first decoding method and another decoding method, thereby reducing the complexity of decoding while ensuring a certain level of audio quality.
[0050] In the second aspect, the spatial decoding method is a decoding method in which reconstruction is performed based on the attribute information of the target virtual speaker, attribute The information is obtained by decoding the bitstream.
[0051] According to the second embodiment, or any one implementation of the second embodiment, the third decoding scheme includes one or more decoding schemes.
[0052] According to the second embodiment, or any one of the implementations of the second embodiment, the third decoding scheme includes a channel copy decoding scheme.
[0053] According to the second embodiment, or any one of the implementations of the second embodiment, the channel copy decoding scheme is a correlation-removal decoding scheme.
[0054] In this way, if the third decoding method is a channel copy decoding method, the degree of variation in the reconstructed scene audio signal obtained through decoding can be reduced, and the smoothness of the audio can be improved.
[0055] According to the second embodiment, or any one implementation of the second embodiment, the step of decoding a bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal includes: A step of decoding one frame of a bitstream based on multiple decoding schemes in a combination of decoding schemes to obtain one frame of a reconstructed scene audio signal.
[0056] According to the second embodiment, or any one implementation of the second embodiment, the bitstream includes X frames, each of which includes the i-th and j-th frames, and the step of decoding the bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal includes: A step of decoding the i-th frame of the bitstream based on multiple decoding schemes in a combination of decoding schemes to obtain the i-th frame of the reconstructed scene audio signal. and, A step of decoding the j-th frame of the bitstream based on the first decoding scheme in the combination of decoding schemes to obtain the j-th frame of the reconstructed scene audio signal.
[0057] In this specification, i and j are integers from 1 to X, where i is not equal to j and X is a positive integer.
[0058] According to the second embodiment, or any one implementation of the second embodiment, the bitstream includes X frames, each of which includes the i-th and j-th frames, and the step of decoding the bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal includes: A step of decoding the i-th frame of the bitstream based on multiple decoding schemes in the k1-th decoding scheme combination to obtain the i-th frame of the reconstructed scene audio signal. A step of decoding the j-th frame of the bitstream based on multiple decoding schemes in a k2th decoding scheme combination to obtain the j-th frame of the reconstructed scene audio signal.
[0059] In this specification, i and j are positive integers from 1 to X, i is not equal to j, and k1 and k2 are positive integers, k1 is not equal to k2.
[0060] According to the second embodiment, or any one implementation of the second embodiment, if the reconstructed scene audio signal of a frame includes a reconstructed audio signal having C channels, and the combination of decoding schemes is a combination of a first decoding scheme, a second decoding scheme, and a third decoding scheme, the step of decoding frames of a bitstream based on the multiple decoding schemes in the combination of decoding schemes to obtain frames of the reconstructed scene audio signal includes the following: A step of obtaining one frame of a reconstructed scene audio signal by decoding C1 first channels based on a first decoding scheme, C2 second channels based on a second decoding scheme, and C3 third channels based on a third decoding scheme, based on a bitstream of one frame.
[0061] C is equal to the sum of C1, C2, and C3, where C, C1, C2, and C3 are positive integers.
[0062] According to the second embodiment, or any one implementation of the second embodiment, the reconstructed scene audio signal of a frame includes a reconstructed audio signal having C channels, and the reconstructed audio signal having one channel has Y bandwidths, where C and Y are positive integers. The step of decoding one frame of the bitstream based on multiple decoding schemes in the combination of decoding schemes to obtain one frame of the reconstructed scene audio signal includes the following: A step of obtaining a frame of the reconstructed scene audio signal by decoding one frame of the bitstream based on multiple decoding schemes in a combination of decoding schemes, for Y bandwidths of one channel within one frame of the reconstructed scene audio signal.
[0063] According to the second embodiment, or any one of the implementations of the second embodiment, the reconstructed scene audio signal is an Nth-order higher-order ambisonic HOA signal. The first C1 channel is one of the C1 channels contained in the zeroth to Mth order signals of the Nth order HOA signal. M is an integer less than N, C is equal to the square of (N+1), and C1 is less than or equal to the square of (M+1). The second set of two C channels includes the other four C channels included in the zero- to M-th order signals of the N-th order HOA signal, and five C channels other than those included in the zero- to M-th order signals of the N-th order HOA signal. The third channel C3 includes the other C6 channels included in the zero- to M-th order signals of the N-th order HOA signal, and the other C7 channels of the N-th order HOA signal that are not included in the zero- to M-th order signals.
[0064] C2 is equal to the sum of C4 and C5, C3 is equal to the sum of C6 and C7, the square of (M+1) is equal to the sum of C1, C4, and C6, and C4, C5, C6, and C7 are integers.
[0065] According to the second embodiment, or any one implementation of the second embodiment, the step of decoding C3 third channels based on a third decoding scheme includes: namely, the step of decoding C3 third channels based on a third decoding scheme corresponding to a first preset identifier obtained by analyzing a bitstream; the third decoding scheme corresponding to the first preset identifier is a time-domain decorrelation decoding scheme or a frequency-domain decorrelation decoding scheme.
[0066] According to the second embodiment, or any one implementation of the second embodiment, if the combination of decoding schemes is a combination of the first decoding scheme, the second decoding scheme, and the third decoding scheme, the method further includes the following: A step of correcting a reconstructed audio signal having two second channels in the reconstructed scene audio signal, based on feature information corresponding to the second channel of C2, which is obtained by analyzing the bitstream.
[0067] According to the second embodiment, or any one implementation of the second embodiment, the method further includes: namely, A step to analyze a second preset identifier from the bitstream. The step of decoding the bitstream based on the combination of decoding schemes to obtain a reconstructed scene audio signal includes the following: A step of decoding the bitstream based on a combination of decoding schemes corresponding to a second preset identifier to obtain a reconstructed scene audio signal.
[0068] Any of the second embodiment and any one of its implementations corresponds to any one of the first embodiment and any one of its implementations. For technical effects corresponding to any of the second embodiment and any one of its implementations, please refer to the technical effects corresponding to any one of the first embodiment and any one of its implementations. Further details will not be described in this specification.
[0069] In a third aspect, one embodiment of the present application provides a bitstream generation method. In this method, a bitstream can be generated according to either the first aspect or an implementation of the first aspect.
[0070] Any third aspect and any one implementation of the third aspect correspond to any one of the first aspect and any one implementation of the first aspect. For technical effects corresponding to any one of the third aspect and any one implementation of the third aspect, please refer to the technical effects corresponding to any one of the first aspect and any one implementation of the first aspect. Further details will not be described in this specification.
[0071] According to a fourth aspect, one embodiment of the present application provides a scene audio coding device. The device includes the following: A signal acquisition module configured to acquire the scene audio signal to be encoded, and, An encoding module configured to encode a scene audio signal based on a combination of encoding schemes. The combination of encoding schemes includes at least one of the following combinations: a combination of a first encoding scheme, a second encoding scheme, and a third encoding scheme, and a combination of the first encoding scheme and the third encoding scheme. The first encoding scheme encodes a signal, the second encoding scheme is a spatial encoding scheme, and the third encoding scheme is an encoding scheme other than the first and second encoding schemes.
[0072] The scene audio encoding device in the fourth embodiment may perform the steps in the first embodiment and any one of the implementations of the first embodiment. Further details are not described herein.
[0073] Furthermore, the scene audio encoding device in the fourth embodiment may further include a communication module.
[0074] Any of the fourth aspect and any one of its implementations corresponds to any one of the first aspect and any one of its implementations. For technical effects corresponding to any of the fourth aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the first aspect and any one of its implementations. Further details will not be described in this specification.
[0075] According to a fifth aspect, one embodiment of the present application provides a scene audio decoding device. The device includes the following: A bitstream receiver module configured to receive a bitstream, and A decoding module configured to decode a bitstream and obtain a reconstructed scene audio signal based on a combination of decoding schemes.
[0076] The combination of decoding methods includes at least one of the following combinations: namely, a combination of the first decoding method, the second decoding method, and the third decoding method, and a combination of the first decoding method and the third decoding method. The first decoding method is used to decode encoded data obtained by encoding a signal, the second decoding method is a spatial decoding method, and the third decoding method is a decoding method other than the first and second decoding methods.
[0077] The scene audio decoding device in the fifth embodiment may perform the steps in the second embodiment and any one of the implementations of the second embodiment. Further details are not described herein.
[0078] Furthermore, the scene audio decoding device in the fifth embodiment may further include a communication module.
[0079] Any fifth aspect and any one implementation of the fifth aspect correspond to any one implementation of the second aspect. For technical effects corresponding to any one implementation of the fifth aspect and any one implementation of the fifth aspect, please refer to the technical effects corresponding to any one implementation of the second aspect and any one implementation of the second aspect. Further details will not be described in this specification.
[0080] According to the sixth aspect, one embodiment of the present application provides an electronic device including a memory and a processor. The memory is connected to the processor and stores program instructions, and when the program instructions are executed by the processor, the electronic device is capable of performing a scene audio coding method according to the first aspect or any one of the possible implementations thereof.
[0081] Any of the sixth aspect and any one of its implementations corresponds to any one of the first aspect and any one of its implementations. For the technical effects corresponding to any of the sixth aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the first aspect and any one of its implementations. Further details will not be explained in this specification.
[0082] In accordance with the seventh aspect, one embodiment of the present application provides an electronic device including a memory and a processor. The memory is connected to the processor and stores program instructions, and when the program instructions are executed by the processor, the electronic device is capable of performing a scene audio decoding method according to the second aspect or any possible implementation thereof.
[0083] Any seventh aspect and any one implementation of the seventh aspect correspond to any one implementation of the second aspect and any one implementation of the second aspect. For technical effects corresponding to any seventh aspect and any one implementation of the seventh aspect, please refer to the technical effects corresponding to any one implementation of the second aspect and any one implementation of the second aspect. Further details will not be described in this specification.
[0084] According to the eighth aspect, one embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors. The one or more processors receive or transmit data through the one or more interface circuits. When the one or more processors execute computer instructions, the electronic device is able to perform a scene audio coding method according to the first aspect or any one of the possible implementations of the first aspect.
[0085] Any eighth aspect and any one implementation of the eighth aspect correspond to any one of the first aspect and any one implementation of the first aspect. For the technical effects corresponding to any eighth aspect and any one implementation of the eighth aspect, please refer to the technical effects corresponding to any one of the first aspect and any one implementation of the first aspect. Further details will not be explained in this specification.
[0086] In accordance with the ninth aspect, one embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors. The one or more processors receive or transmit data through the one or more interface circuits. When the one or more processors execute computer instructions, the electronic device is able to perform a scene audio decoding method according to the second aspect, or any one of the possible implementations of the second aspect.
[0087] Any one of the ninth aspect and its implementation corresponds to any one of the second aspect and its implementation. For the technical effects corresponding to any one of the ninth aspect and its implementation, please refer to the technical effects corresponding to any one of the second aspect and its implementation. Further details will not be described in this specification.
[0088] In accordance with the tenth aspect, one embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When this computer program is executed on a computer or processor, the computer or processor is able to perform a scene speech coding method according to the first aspect or any one of a possible implementation of the first aspect.
[0089] Any one of the tenth aspect and the tenth implementation corresponds to any one of the first aspect and the first implementation. For the technical effects corresponding to any one of the tenth aspect and the tenth implementation, please refer to the technical effects corresponding to any one of the first aspect and the first implementation. Further details will not be described in this specification.
[0090] In accordance with the eleventh aspect, one embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When this computer program is executed on a computer or processor, the computer or processor is able to perform a scene audio decoding method according to the second aspect, or any one of a possible implementation of the second aspect.
[0091] Any one of the eleventh aspect and any one of its implementations corresponds to any one of the second aspect and any one of its implementations. For the technical effects corresponding to any one of the eleventh aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the second aspect and any one of its implementations. Further details will not be described in this specification.
[0092] In accordance with the twelfth aspect, one embodiment of the present application provides a computer program product. The computer program product includes a software program. When this software program is executed by a computer or processor, the computer or processor is able to perform a scene audio coding method according to the first aspect or any one of a possible implementation of the first aspect.
[0093] Any one of the twelfth aspect and any one of its implementations corresponds to any one of the first aspect and any one of its implementations. For the technical effects corresponding to any one of the twelfth aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the first aspect and any one of its implementations. Further details will not be described in this specification.
[0094] In accordance with the thirteenth aspect, one embodiment of the present application provides a computer program product, which includes a software program. When this software program is executed by a computer or processor, the computer or processor is able to perform a scene audio decoding method according to the second aspect or any one of a possible implementation of the second aspect.
[0095] The thirteenth aspect, and any one of the thirteenth aspects, corresponds to the second aspect, and any one of the implementations of the second aspect. For the technical effects corresponding to the thirteenth aspect, and any one of the thirteenth aspects, please refer to the technical effects corresponding to the second aspect, and any one of the implementations of the second aspect. Further details will not be described in this specification.
[0096] According to the fourteenth aspect, one embodiment of the present application provides a bitstream storage device. The device includes a receiver and at least one storage medium. The receiver is configured to receive a bitstream. The at least one storage medium is configured to store the bitstream. The bitstream is generated according to either the first aspect or an implementation of the first aspect.
[0097] Any one of the fourteenth aspect and any one of its implementations corresponds to any one of the first aspect and any one of its implementations. For the technical effects corresponding to any one of the fourteenth aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the first aspect and any one of its implementations. Further details will not be described in this specification.
[0098] According to the fifteenth aspect, one embodiment of the present application provides a bitstream transmission device. The device includes a transmitter and at least one storage medium, the storage medium configured to store a bitstream. The bitstream is generated according to either the first aspect or an implementation of the first aspect. The transmitter is configured to receive the bitstream from the storage medium and transmit the bitstream to a terminal device via a transmission medium.
[0099] Any one of the fifteenth aspect and any one of its implementations corresponds to any one of the first aspect and any one of its implementations. For the technical effects corresponding to any one of the fifteenth aspect and any one of its implementations, please refer to the technical effects corresponding to any one of the first aspect and any one of its implementations. Further details will not be described in this specification.
[0100] According to the sixteenth aspect, one embodiment of the present application provides a bitstream distribution system. The system includes: at least one storage medium configured to store at least one bitstream, wherein at least one bitstream is generated according to either a first aspect or an implementation of the first aspect; and a streaming media device configured to retrieve a target bitstream from at least one storage medium and transmit the target bitstream to a terminal device, wherein the streaming media device includes a content server or a content distribution server.
[0101] Any sixteenth aspect and any one implementation of the sixteenth aspect correspond to any one of the first aspect and any one implementation of the first aspect. For the technical effects corresponding to any sixteenth aspect and any one implementation of the sixteenth aspect, please refer to the technical effects corresponding to any one of the first aspect and any one implementation of the first aspect. Further details will not be described in this specification. [Brief explanation of the drawing]
[0102] [Figure 1a] This is a diagram illustrating an exemplary application scenario. [Figure 1b] This is a diagram illustrating an exemplary application scenario. [Figure 2] This diagram illustrates the encoding process for an exemplary scene audio signal. [Figure 3] This diagram illustrates the decoding process of an exemplary scene audio signal. [Figure 4a] This diagram illustrates the encoding process for an exemplary scene audio signal. [Figure 4b] This figure shows the distribution of exemplary candidate virtual speakers. [Figure 5] This diagram illustrates the decoding process of an exemplary scene audio signal. [Figure 6] This diagram illustrates the encoding process for an exemplary scene audio signal. [Figure 7]This diagram illustrates the decoding process of an exemplary scene audio signal. [Figure 8] This diagram illustrates the encoding process for an exemplary scene audio signal. [Figure 9] This diagram illustrates the decoding process of an exemplary scene audio signal. [Figure 10] This is a diagram illustrating the configuration of an exemplary device. [Modes for carrying out the invention]
[0103] The technical solutions in the embodiments of this application are described below clearly and completely with reference to the accompanying drawings of the embodiments. It will be apparent that the embodiments described are part of, but not all, the embodiments of this application. All other embodiments that can be obtained without creative effort by those skilled in the art based on the embodiments of this application are included in the scope of protection of this application.
[0104] In this specification, the term "and / or" describes only the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may represent the following three cases: when only A exists, when both A and B exist, and when only B exists.
[0105] In the specification and claims of the embodiments of this application, terms such as “first,” “second,” etc., are intended to distinguish different objects and not to indicate a particular order of objects. For example, “first object of subject,” “second object of subject,” and similar terms are used to distinguish different objects of subject, but not to describe a particular order of objects of subject.
[0106] In embodiments of this application, phrases such as “example” or “for example” indicate that they are providing illustrations, demonstrations, or explanations. In embodiments of this application, none of the embodiments or design schemes described as “example” or “for example” should be described as being preferable or having more advantages than other embodiments or design schemes. More precisely, the use of phrases such as “example” or “for example” is intended to present relative concepts in a particular manner.
[0107] In the description of embodiments of this application, "multiple" means two or more unless otherwise specified. For example, "multiple processing units" means two or more processing units, and "multiple systems" means two or more systems.
[0108] To clearly and concisely describe the following embodiments, we will first briefly explain the related technologies.
[0109] Sound is a continuous wave produced by the vibration of an object. An object that vibrates and emits sound waves is called a sound source. As sound waves propagate through a medium (such as air, a solid, or a liquid), the auditory organs of humans or animals can perceive the sound.
[0110] The characteristics of sound waves include tone, intensity, and tamber. Tone indicates the level of the sound. Intensity indicates the loudness of the sound. This intensity is sometimes called volume or sound level. The unit of intensity is decibel (dB). Tamber is sometimes called sound quality.
[0111] The frequency of a sound wave determines its tone level. It has been shown that the higher the frequency, the higher the tone. The number of times an object vibrates per second is called its frequency, and the unit of frequency is Hertz (Hz). The range of acoustic frequencies that the human ear can perceive is from 20 Hz to 20,000 Hz.
[0112] The amplitude of a sound wave determines its intensity. It has been shown that the larger the amplitude, the stronger the intensity. It has also been shown that the shorter the distance from the sound source, the stronger the intensity.
[0113] The waveform of a sound wave determines the tamper. Sound wave waveforms include square waves, sawtooth waves, sine waves, pulse waves, and similar types.
[0114] Acoustics can be classified into regular and irregular sounds based on the characteristics of sound waves. Irregular sounds are sounds emitted through irregular vibrations from a sound source. Irregular sounds are, for example, noises that affect people's work, study, rest, and similar activities. Regular sounds are sounds emitted through regular vibrations from a sound source. Regular sounds include the sounds of speech and music. Electrically, regular sounds are analog signals that change continuously in the time-frequency domain. These analog signals are sometimes called audio signals. Audio signals are information carriers that transmit speech, music, and sound effects.
[0115] Human hearing has the ability to distinguish the spatial distribution of sound sources; therefore, when listening to sounds in space, we can perceive the direction and location of the sound, in addition to its tone, intensity, and tampering.
[0116] As interest in and quality demands for the auditory system experience increase, three-dimensional audio technologies have emerged to enhance the sense of depth, presence, and spatiality of sound. As a result, listeners not only perceive sounds emanating from sound sources in front, behind, and to the sides, but also feel as if their surroundings are enveloped by the spatial sound field (abbreviated as "sound field") created by these sound sources, and as if the sound is spreading outwards. This creates an "immersive" sound effect, similar to that experienced when listening in a theater or concert hall.
[0117] In embodiments of this application, the scene audio signal may be a signal used to describe a sound field. The scene audio signal may include an HOA signal (the HOA signal may include a three-dimensional HOA signal and a two-dimensional HOA signal (sometimes called a planar HOA signal)), as well as a three-dimensional audio signal. The three-dimensional audio signal may be an audio signal other than the HOA signal among the scene audio signals. An explanation is provided below using the HOA signal as an example.
[0118] It is well known that sound waves propagate in an ideal medium, with wavenumber k = w / c and angular frequency w = 2πf. In this specification, f is the frequency of the sound wave and c is the speed of sound. The sound pressure P satisfies equation (1). In this specification, ∇ 2 This is the Laplacian operator.
[0119]
number
[0120] We assume that the spatial system outside the human ear is a sphere, and that the listener is at the center of this sphere. Sound transmitted from outside the sphere is projected onto the spherical surface, and sound outside that surface is removed. We assume that sound sources are distributed on the spherical surface, and that the sound field generated by the sound sources on the spherical surface fits the sound field generated by the original sound sources. In other words, three-dimensional audio technology is a sound field fitting method. Specifically, equation (1) is solved in a spherical coordinate system. In a passive spherical region, the solution to equation (1) becomes equation (2).
[0121]
number
[0122] In this specification, r represents the radius of the sphere, and θ represents the horizontal angle information (also called the azimuth angle information).
[0123]
number
[0124] k represents pitch angle information (also called elevation angle information), k represents the wavenumber, S represents the amplitude of the ideal plane wave, and m represents the sequential number of dimensions of the HOA signal.
[0125]
number
[0126] This represents a spherical Bessel function, which is also called a radial basis function. First, "j" is the imaginary unit,
[0127]
number
[0128] It does not change with angle.
[0129]
number
[0130] θ and
[0131]
number
[0132] This represents the spherical harmonics in the direction,
[0133]
number
[0134] This represents the spherical harmonics in the direction of the sound source. The HOA signal satisfies equation (3).
[0135]
number
[0136] Substituting equation (3) into equation (2), equation (2) can be transformed into equation (4).
[0137]
number
[0138] In this specification, m is truncated to the Nth term, i.e., m = N.
[0139]
number
[0140] It is used as an approximate description of the sound field. In this case,
[0141]
number
[0142] This is sometimes called the HOA coefficient (which can be used to represent an Nth-order HOA signal). The sound field is the region in a medium where sound waves exist. N is an integer greater than or equal to 1.
[0143] The scene audio signal is an information carrier that carries spatial positional information of the sound source in the sound field and describes the listener's sound field in space. Equation (4) shows that the sound field can be extended onto a sphere based on spherical harmonics. In other words, the sound field can be decomposed into a superposition of multiple plane waves. Therefore, the sound field described by the HOA signal can be represented through a superposition of multiple plane waves, and the sound field is reconstructed based on the HOA coefficients.
[0144] The HOA signal to be encoded in the embodiments of this application may be an Nth-order HOA signal and can be represented using HOA coefficients or ambisonic coefficients, where N is an integer greater than or equal to 1 (when N is equal to 1, a first-order HOA signal is sometimes called an FOA (first-order ambisonics) signal). An Nth-order HOA signal is (N+1) 2 Includes an audio signal having a number of channels.
[0145] Figure 1a shows an example of an application scenario. Figure 1a illustrates a scenario for decoding and decoding a scene audio signal.
[0146] As shown in Figure 1a, for example, the first electronic device may include a first audio recording module, a first scene audio encoding module, a first channel encoding module, a first channel decoding module, a first scene audio decoding module, and a first audio playback module. It should be understood that the first electronic device may include more or fewer modules than those shown in Figure 1a. This is not limited to the present application.
[0147] As shown in Figure 1a, for example, the second electronic device may include a second audio recording module, a second scene audio coding module, a second channel coding module, a second channel decoding module, a second scene audio decoding module, and a second audio playback module. It should be understood that the second electronic device may include more or fewer modules than those shown in Figure 1a. This is not limited to the present application.
[0148] For example, the process in which a first electronic device encodes a scene audio signal, transmits the encoded scene audio signal to a second electronic device, and the second electronic device performs decoding and audio playback may be as follows: The first audio recording module may perform audio recording and output a scene audio signal to a first scene audio encoding module. The first scene audio encoding module may then encode the scene audio signal and output a bitstream to a first channel encoding module. The first channel encoding module may then perform channel coding on the bitstream and transmit the bitstream obtained through channel coding to the second electronic device via a wireless or wired network communication device. The second channel decoding module in the second electronic device may then perform channel decoding on the received data to obtain a bitstream and output that bitstream to a second scene audio decoding module. The second scene audio decoding module may then decode the bitstream to obtain a reconstructed scene audio signal. The second scene audio decoding module may then output the reconstructed scene audio signal to a second audio playback module. The second audio playback module then performs audio playback.
[0149] The second audio playback module performs post-processing on the reconstructed scene audio signal (e.g., audio rendering (e.g., (N+1) 2 It should be noted that a reconstructed scene audio signal containing an audio signal with a number of channels can be converted into an audio signal with the same number of channels as the number of speakers in a second electronic device, or that the reconstructed scene audio signal can be converted into an audio signal suitable for playback by the speakers in the second electronic device by performing processes such as loudness normalization, user interaction, audio format conversion, or noise reduction.
[0150] It should be understood that the process in which the second electronic device encodes the scene audio signal, transmits the encoded scene audio signal to the first electronic device, and the first electronic device performs decoding and audio playback is similar to the process described above in which the first electronic device transmits the scene audio signal to the second electronic device, and the second electronic device performs audio playback. Further details will not be described again in this specification.
[0151] For example, the first electronic device and the second electronic device may include, but are not limited to, personal computers, computer workstations, smartphones, tablet computers, servers, smart cameras, intelligent vehicles, other types of mobile phones, media consumer devices, wearable devices, set-top boxes, game consoles, and similar devices.
[0152] For example, this application may be particularly applicable to VR (Virtual Reality) / AR (Augmented Reality) scenarios. In one possible embodiment, the first electronic device is a server and the second electronic device is a VR / AR device. In another possible embodiment, the second electronic device is a server and the first electronic device is a VR / AR device.
[0153] For example, the first scene audio encoding module and the second scene audio encoding module could be scene audio encoders. The first scene audio decoding module and the second scene audio decoding module could be scene audio decoders.
[0154] For example, if a first electronic device encodes a scene audio signal and a second electronic device reconstructs the scene audio signal, the first electronic device may be called the encoder side, and the second electronic device may be called the decoder side. Also, if a second electronic device encodes a scene audio signal and a first electronic device reconstructs the scene audio signal, the second electronic device may be called the encoder side, and the first electronic device may be called the decoder side.
[0155] Figure 1b shows an example of an application scenario. Figure 1b shows a scene audio signal transcoding scenario.
[0156] For example, as shown in (1) in Figure 1b, a wireless or core network device may include, for example, a channel decoding module, another audio decoding module, a scene audio coding module, and a channel coding module. The wireless or core network device may be configured to perform audio transcoding.
[0157] For example, a specific application scenario in (1) in Figure 1b could be as follows: The first electronic device does not have a scene audio encoding module, but only another audio encoding module. The second electronic device has only a scene audio decoding module, but no other audio decoding module. A wireless or core network device may be used for transcoding, so that the second electronic device can decode and reproduce the scene audio signal encoded by the first electronic device using another audio encoding module.
[0158] Specifically, the first electronic device encodes the scene audio signal using another audio encoding module to obtain a first bitstream. The first electronic device then performs channel encoding on the first bitstream and transmits the encoded first bitstream to a wireless or core network device. The channel decoding module of the wireless or core network device then performs channel decoding and may output the first bitstream obtained through channel decoding to another audio decoding module. The other audio decoding module then decodes the first bitstream to obtain the scene audio signal and outputs the scene audio signal to the scene audio encoding module. The scene audio encoding module then encodes the scene audio signal to obtain a second bitstream and may output the second bitstream to the channel encoding module. After performing channel encoding on the second bitstream, the channel encoding module transmits the encoded second bitstream to the second electronic device. In this way, the second electronic device can call the scene audio decoding module to decode the second bitstream obtained through channel decoding and obtain a reconstructed scene audio signal. Subsequently, the second electronic device may perform audio playback based on the reconstructed scene audio signal.
[0159] For example, as shown in (2) in Figure 1b, a wireless or core network device may include, for example, a channel decoding module, a scene audio decoding module, another audio encoding module, and a channel encoding module. The wireless or core network device may be configured to perform audio transcoding.
[0160] For example, a specific application scenario in (2) in Figure 1b could be as follows: The first electronic device is equipped only with a scene audio encoding module and no other audio encoding module. The second electronic device is equipped only with a different audio decoding module and no scene audio decoding module. A wireless or core network device may be used for transcoding, thereby allowing the second electronic device to decode and reproduce the scene audio signal encoded by the first electronic device using the scene audio encoding module.
[0161] Specifically, the first electronic device encodes the scene audio signal using a scene audio encoding module to obtain a first bitstream. The first electronic device then performs channel encoding on the first bitstream and transmits the encoded first bitstream to a wireless or core network device. The channel decoding module of the wireless or core network device then performs channel decoding and outputs the first bitstream obtained through channel decoding to the scene audio decoding module. The scene audio decoding module then decodes the first bitstream to obtain a scene audio signal and outputs the scene audio signal to another audio encoding module. The other audio encoding module then encodes the scene audio signal to obtain a second bitstream and outputs the second bitstream to the channel encoding module. After performing channel encoding on the second bitstream, the channel encoding module transmits the encoded second bitstream to the second electronic device. In this way, the second electronic device calls another audio decoding module to decode the second bitstream obtained through channel decoding and obtain a reconstructed scene audio signal. Subsequently, the second electronic device may perform audio playback based on the reconstructed scene audio signal.
[0162] The encoding and decoding processes for scene audio signals are described below.
[0163] Figure 2 shows an example of the encoding process for a scene audio signal.
[0164] S201: Acquire the scene audio signal.
[0165] For example, a scene audio signal to be encoded can be obtained. This scene audio signal may contain an audio signal with C channels, where C is a positive integer.
[0166] For example, if the scene audio signal is an HOA signal, the HOA signal is an Nth-order HOA signal, that is, the one in equation (3) where m is truncated to the Nth term.
[0167]
number
[0168] It is possible.
[0169] For example, an Nth-order HOA signal may contain an audio signal with C channels. C = (N+1) 2 For example, if N=3, the Nth-order HOA signal contains an audio signal with 16 channels. If N=4, the Nth-order HOA signal contains an audio signal with 25 channels.
[0170] For example, a scene audio signal may include one or more frames.
[0171] S202: Encodes the scene audio signal based on a combination of encoding schemes.
[0172] For example, a scene audio signal may be encoded based on a combination of encoding schemes, including multiple encoding schemes, in order to obtain a bitstream of the scene audio signal.
[0173] In one possible embodiment, a first encoding scheme, a second encoding scheme, and a third encoding scheme may form a combination of encoding schemes.
[0174] In one possible embodiment, a first encoding scheme and a third encoding scheme may form a combination of encoding schemes.
[0175] For example, the first encoding scheme encodes a signal, and specifically, it may perform operations on the signal such as time-frequency conversion, preprocessing, bit allocation, quantization, and entropy coding. The first encoding scheme is sometimes also called a direct encoding scheme.
[0176] For example, the second encoding scheme could be a spatial encoding scheme, which encodes attribute information belonging to the target virtual speaker, and which is determined based on the scene audio signal. The process of determining the attribute information of the target virtual speaker based on the scene audio signal is described below.
[0177] For example, the third encoding scheme may include one or more encoding schemes other than the first and second encoding schemes.
[0178] In one possible form, the third encoding scheme is a channel copy (or HOA copy) encoding scheme. Optionally, the third encoding scheme is a correlation-removal encoding scheme. It should be understood that the number and types of encoding schemes included in the third encoding scheme are not limited in this application.
[0179] It should be understood that the combination of encoding schemes includes at least one of the following combinations: namely, a combination of the first encoding scheme and the third encoding scheme, or a combination of the first encoding scheme, the second encoding scheme, and the third encoding scheme.
[0180] Encoding based on a direct encoding scheme (e.g., the first encoding scheme) can improve encoding quality, but requires a high bitrate overhead. Encoding based on another encoding scheme (the second or third encoding scheme) can reduce bitrate overhead, but degrades encoding quality. Therefore, in this application, encoding is performed based on a combination of the direct encoding scheme and another encoding scheme, thereby reducing bitrate overhead and encoding complexity while ensuring a certain level of encoding quality.
[0181] If the bitrate overhead required by the third encoding scheme is lower than the bitrate overhead required by the second encoding scheme, for example, if the third encoding scheme is a channel copy encoding scheme, then, in this application, encoding the scene audio signal based on a combination of the first and third encoding schemes reduces bitrate overhead and encoding complexity compared to encoding the scene audio signal based on a combination of the first, second, and third encoding schemes.
[0182] In one possible embodiment, the encoder and decoder may agree in advance on the type of encoding scheme combination (or decoding scheme combination).
[0183] In one possible embodiment, the encoder and decoder do not agree in advance on the type of encoding scheme combination (or decoding scheme combination). In this case, the encoder may further encode a second preset identifier, which may indicate the type of encoding scheme combination (or decoding scheme combination). Thus, after the second preset identifier is transmitted to the decoder using a bitstream, the decoder may recognize the type of encoding scheme combination (or decoding scheme combination) and perform decoding based on the decoding scheme combination corresponding to the encoding scheme combination on the encoder side to obtain a reconstructed scene audio signal.
[0184] Figure 3 shows an example of scene audio decoding. The embodiment in Figure 3 is a decoding process that corresponds to the encoding process included in the embodiment in Figure 2.
[0185] S301: Receives a bitstream.
[0186] S302: Based on the combination of decoding methods, the bitstream is decoded to obtain the reconstructed scene audio signal.
[0187] For example, in order to obtain a reconstructed scene audio signal, the bitstream may be decoded based on a combination of decoding schemes, including multiple decoding methods.
[0188] In one possible embodiment, a first decoding scheme, a second decoding scheme, and a third decoding scheme may form a combination of decoding schemes.
[0189] In one possible embodiment, a first decoding scheme and a third decoding scheme may form a combination of decoding schemes.
[0190] For example, the first decoding method decodes encoded data obtained by encoding a signal, and specifically, it may perform operations such as entropy decoding, inverse quantization, bit allocation, post-processing, and time-frequency conversion on the encoded data obtained by encoding a signal. The first decoding method is sometimes also called the direct decoding method.
[0191] For example, the second decoding method could be a spatial decoding method, which performs reconstruction based on the attribute information of the target virtual speaker. The process of reconstruction based on the attribute information of the target virtual speaker is described below.
[0192] For example, the third decoding method may include one or more decoding methods other than the first and second decoding methods. In one possible embodiment, the third decoding method may be a channel copy (HOA copy) decoding method. Optionally, the third decoding method is a correlation-removal decoding method. The number and types of decoding methods included in the third decoding method are not limited in this application.
[0193] It should be understood that the combination of decoding methods includes at least one of the following combinations: namely, a combination of the first decoding method and the third decoding method, or a combination of the first decoding method, the second decoding method, and the third decoding method.
[0194] For example, if the encoder and decoder agree in advance on the types of encoding schemes (decoding schemes), the decoder can decode the bitstream based on the pre-agreed decoding scheme combination to obtain a reconstructed scene audio signal.
[0195] For example, if the encoder and decoder have not agreed in advance on the type of encoding scheme combination (or decoding scheme combination), the decoder may parse a second preset identifier from the bitstream and then decode the bitstream based on the decoding scheme combination corresponding to the second preset identifier to obtain a reconstructed scene audio signal.
[0196] In this way, if the third decoding method is a channel copy decoding method, the degree of variation in the reconstructed scene audio signal obtained through decoding can be reduced, and the smoothness of the audio can be improved.
[0197] In one possible embodiment, if the scene audio signal to be encoded contains a single frame, that frame of the scene audio signal may be encoded based on multiple encoding schemes in a combination of encoding schemes. An explanation may be provided with reference to the embodiment in Figure 4a or the embodiment in Figure 6. Accordingly, the decoder receives a bitstream of a single frame and decodes the frame of that bitstream based on multiple decoding schemes in a combination of decoding schemes to obtain a single frame of the reconstructed scene audio signal. References may be made to the following explanation relating to the embodiment in Figure 5 or the embodiment in Figure 7.
[0198] In one possible embodiment, if the scene audio signal to be encoded contains multiple frames, each frame of the scene audio signal may be encoded based on multiple encoding schemes in the combination of encoding schemes. An explanation may be provided by referring to the embodiment in Figure 4a or the embodiment in Figure 6. Accordingly, the decoder receives a bitstream of multiple frames and decodes each frame of the bitstream based on multiple decoding schemes in the combination of decoding schemes to obtain one frame of the reconstructed scene audio signal. A reference may be made to the following explanation relating to the embodiment in Figure 5 or the embodiment in Figure 7.
[0199] The process of encoding and decoding a single frame of a scene audio signal is described below, using an example where the third encoding scheme is correlation-rejection coding.
[0200] Figure 4a shows an example of scene audio encoding processing. In the embodiment shown in Figure 4a, multiple channels within a single frame of a scene audio signal are encoded based on multiple encoding schemes in a combination of encoding schemes. The combination of encoding schemes is a combination of a first encoding scheme, a second encoding scheme, and a third encoding scheme.
[0201] S401: Acquire the scene audio signal. Here, the scene audio signal includes an audio signal having C channels.
[0202] S402: For the scene audio signal, one first channel C is encoded based on the first encoding scheme, two second channels C is encoded based on the second encoding scheme, and three third channels C is encoded based on the third encoding scheme.
[0203] For example, for a single frame of a scene audio signal, it can be determined that of the C channels contained in that frame of the scene audio signal, C1 first channels are encoded according to a first encoding scheme, C2 second channels are encoded according to a second encoding scheme, and C3 third channels are encoded according to a third encoding scheme. C is equal to the sum of C1, C2, and C3, C, C1, C2, and C3 are positive integers, and the first, second, and third channels are distinct channels.
[0204] Next, in order to obtain the bitstream of the scene audio signal for that frame, the first channel C1 is encoded based on the first encoding scheme, the second channel C2 is encoded based on the second encoding scheme, and the third channel C3 is encoded based on the third encoding scheme.
[0205] For example, the process of encoding C1 first channels based on a first encoding scheme may be as follows: namely, the step of encoding an audio signal having C1 first channels based on a direct encoding scheme; specifically, the step of performing operations such as time-frequency conversion, preprocessing, bit allocation, quantization, and entropy coding on the audio signal having C1 first channels to obtain encoded data of the audio signal having C1 first channels.
[0206] For example, the process of encoding C2 second channels based on a second encoding scheme may be as follows: a step of determining the attribute information of the target virtual speaker based on the scene audio signal, and a step of encoding the attribute information of the target virtual speaker.
[0207] For example, a virtual speaker is a speaker that is virtual, not a speaker that actually exists.
[0208] For example, as can be seen from the explanation above, a scene audio signal may be represented through the superposition of multiple plane waves, and a target virtual speaker may be determined to be used to simulate the sound source in the scene audio signal. In this way, in the subsequent decoding process, the virtual speaker signal corresponding to the target virtual speaker is used to reconstruct the scene audio signal.
[0209] In one possible embodiment, multiple candidate virtual speakers located at different positions may be arranged on a spherical surface. Then, a target virtual speaker whose position coincides with the position of the sound source in the scene audio signal may be selected from the multiple candidate virtual speakers.
[0210] Figure 4b shows an example of the distribution of candidate virtual speakers. In Figure 4b, multiple candidate virtual speakers can be evenly distributed on a sphere, and each point on the sphere represents one candidate virtual speaker.
[0211] It should be noted that the number of candidate virtual speakers and their distribution are not limited in this application and can be set as needed. Further details will be provided later.
[0212] For example, a target virtual speaker whose position corresponds to the position of a sound source in the scene audio signal may be selected from a plurality of candidate virtual speakers based on the scene audio signal. There may be one target virtual speaker or a plurality of them; this is not limited in this application. See S11 to S13 for details.
[0213] S11: Obtain multiple virtual speaker coefficient sets corresponding to multiple candidate virtual speakers. Here, the multiple virtual speaker coefficient sets have a one-to-one correspondence with the multiple candidate virtual speakers.
[0214] For example, first configuration information of an encoding module (e.g., a scene audio encoding module) can be obtained. Based on the first configuration information of the encoding module, second configuration information of a candidate virtual speaker is determined. Then, based on the second configuration information of the candidate virtual speaker, multiple candidate virtual speakers are generated.
[0215] For example, the first configuration information includes, but is not limited to, the encoding bitrate, user-defined information (e.g., the number of HOA dimensions corresponding to the encoding module (the number of dimensions of the HOA signal that can be encoded by the encoding module), the number of dimensions of the reconstructed scene audio signal (the expected number of dimensions of the reconstructed HOA signal obtained through decoding by the decoder), and the format of the reconstructed scene audio signal (the expected format of the reconstructed HOA signal obtained through decoding by the decoder). This is not limited to these.
[0216] For example, the second configuration information includes, but is not limited to, the total number of candidate virtual speakers, the number of HOA dimensions for each candidate virtual speaker, and the position information for each candidate virtual speaker. This is not limited to this application.
[0217] For example, the second configuration information of a candidate virtual speaker may be determined in several embodiments based on the first configuration information of the encoding module. For example, if the encoding bitrate is low, the number of candidate virtual speakers may be configured to be small. Or, if the encoding bitrate is high, the number of candidate virtual speakers may be configured to be large. As another example, the number of HOA dimensions of a virtual speaker may be configured as the number of HOA dimensions of the encoding module. In this embodiment of the application, in addition to determining the second configuration information of a candidate virtual speaker based on the first configuration information of the encoding module, the second configuration information of a candidate virtual speaker may also be determined based on user-defined information (for example, information such as the total number of candidate virtual speakers, the number of HOA dimensions of each candidate virtual speaker, and the location information of each candidate virtual speaker, which may be customized by the user). This is not limited to these.
[0218] For example, a configuration table may be pre-configured. The configuration table contains the relationship between the number of candidate virtual speakers and the location information of each candidate virtual speaker. In this way, after the total number of candidate virtual speakers is determined, the location information of each candidate virtual speaker can be determined by searching the configuration table.
[0219] For example, after the second configuration information of a candidate virtual speaker is determined, multiple candidate virtual speakers may be generated based on the second configuration information of the candidate virtual speaker. For example, a corresponding number of candidate virtual speakers may be generated based on the total number of candidate virtual speakers, the HOA dimension of each candidate virtual speaker may be set based on the HOA dimension of each candidate virtual speaker, and the position of each candidate virtual speaker may be set based on the position information of each candidate virtual speaker.
[0220] For example, if each candidate virtual speaker functions as a virtual sound source, the virtual speaker signal generated by the virtual sound source is a plane wave, and this plane wave can be unfolded in spherical coordinates. The amplitude is S and the direction is
[0221]
number
[0222] For an ideal plane wave, the form obtained through an expansion based on spherical harmonics can be shown in equation (3). The HOA dimension of a candidate virtual speaker is the truncated value of m in equation (3).
[0223] Next, based on the HOA dimension of each candidate virtual speaker, the virtual speaker coefficients corresponding to each candidate virtual speaker can be determined (each candidate virtual speaker corresponds to one set of virtual speaker coefficients). For example, referring to equation (3) for a candidate virtual speaker, the truncated value of m in equation (3) is set to the HOA dimension of the candidate virtual speaker, and in equation (3)
[0224]
number
[0225] However, the location information of the candidate virtual speaker
[0226]
number
[0227] It is set to this. In this case, in equation (3)
[0228]
number
[0229] This represents a pair of virtual speaker coefficients (and the virtual speaker coefficients are also HOA coefficients. As can be seen from equation (3), if the position of a candidate virtual speaker differs from the position of the sound source in the scene audio signal, the virtual speaker coefficients of the candidate virtual speaker and the scene audio signal will have different HOA coefficients). In this way, a pair of virtual speaker coefficients corresponding to each candidate virtual speaker can be determined.
[0230] S12: Based on the scene audio signal and multiple virtual speaker coefficient groups, a target virtual speaker is selected from multiple candidate virtual speakers.
[0231] For example, to obtain multiple dot products, the dot product of the scene audio signal and each of the multiple virtual speaker coefficient groups is obtained. There is a one-to-one correspondence between the multiple dot products and the multiple virtual speaker coefficient groups. For example, to obtain the corresponding dot product, the dot product of the scene audio signal and one set of virtual speaker coefficient groups corresponding to each of the multiple candidate virtual speakers may be obtained.
[0232] Next, the target virtual speaker may be selected from multiple candidate virtual speakers based on multiple dot product values. In one possible embodiment, G candidate virtual speakers (G being a positive integer) having the largest dot product value may be selected first as the target virtual speaker. In another possible embodiment, the candidate virtual speaker having the largest dot product may be selected first as the target virtual speaker. To obtain the projection vector, the scene audio signal is projected and superimposed onto a linear combination of virtual speaker coefficients corresponding to the candidate virtual speaker having the largest dot product. Then, to obtain the difference, the projection vector is subtracted from the scene audio signal. Next, in order to perform iterative calculations, the above process is repeated on this difference, and one target virtual speaker is generated in each iteration.
[0233] S13: Retrieve attribute information of the target virtual speaker.
[0234] In one possible embodiment, attribute information of the target virtual speaker is generated based on the location information of the target virtual speaker. In one possible embodiment, the location information of the target virtual speaker (including pitch angle information and horizontal angle information) may be used as attribute information of the target virtual speaker. In one possible embodiment, a location index corresponding to the location information of the target virtual speaker (including a pitch angle index (which may be used to uniquely identify the pitch angle information) and a horizontal angle index (which may be used to uniquely identify the horizontal angle information)) is used as attribute information of the target virtual speaker.
[0235] In one possible embodiment, the virtual speaker index of the target virtual speaker (e.g., a virtual speaker identifier) can be used as attribute information of the target virtual speaker. The virtual speaker index has a one-to-one correspondence with location information.
[0236] In one possible embodiment, the virtual speaker coefficients of the target virtual speaker can be used as attribute information of the target virtual speaker. For example, C virtual speaker coefficients are determined for the target virtual speaker, and these C virtual speaker coefficients are used as attribute information of the target virtual speaker. The C virtual speaker coefficients for the target virtual speaker are , re The configuration scene has a one-to-one correspondence with the audio signal containing C channels.
[0237] It should be noted that the amount of data for the virtual speaker coefficient is much larger than the amount of data for the location information, the location information index, and the virtual speaker index. Certain information among the location information, location information index, virtual speaker index, and virtual speaker coefficient used as attribute information for the target virtual speaker may be determined based on bandwidth. For example, if the bandwidth is wide, the virtual speaker coefficient may be used as attribute information for the target virtual speaker. In this way, the decoder does not need to calculate the virtual speaker coefficient for the target virtual speaker, and the decoder's computational power can be saved. If the bandwidth is narrow, any one of the location information, location information index, and virtual speaker index may be used as attribute information for the target virtual speaker. In this way, the bitrate can be reduced. Alternatively, certain information among the location information, location information index, virtual speaker index, and virtual speaker coefficient used as attribute information for the target virtual speaker may be pre-set. This is not limited to this application.
[0238] In one possible embodiment, the target virtual speaker can be pre-configured.
[0239] The method for determining the target virtual speaker is not limited in this application, nor is the method for determining the attribute information of the target virtual speaker limited in this application.
[0240] For example, if the third encoding scheme is a correlation-removal encoding scheme, encoding C3 third channels based on the third encoding scheme may not be the same as processing an audio signal with C3 third channels. Instead, the decoder performs correlation-removal decoding on the C3 third channels to determine the reconstructed audio signal with C3 third channels.
[0241] In one possible embodiment, the decorrelation coding scheme may include a time-domain decorrelation coding scheme and a frequency-domain decorrelation coding scheme. When the third coding scheme is a decorrelation coding scheme, the step of coding C3 third channels based on the third coding scheme may be as follows: that is, for C3 third channels, determine whether the third coding scheme is a time-domain decorrelation coding scheme or a frequency-domain decorrelation coding scheme. When the third coding scheme is a time-domain decorrelation coding scheme, the audio signal having C3 third channels may not be processed. Instead, the value of the first preset identifier is set to the first preset value, and the first preset identifier corresponding to the C3 third channels is coded. The first preset identifier set to the first preset value indicates that the third coding scheme for C3 third channels is a time-domain decorrelation coding scheme. When the third coding scheme is a frequency-domain decorrelation coding scheme, the audio signal having C3 third channels may not be processed. Instead, the value of the first preset identifier is set to the second preset value, and the first preset identifier corresponding to the three third channels is encoded. The first preset identifier set to the second preset value indicates that the third encoding scheme having three third channels is a frequency-domain decorrelation coding scheme. In this way, the decoder recognizes whether the third encoding scheme is a time-domain decorrelation coding scheme or a frequency-domain decorrelation coding scheme, and then performs decoding based on the algorithm of the corresponding decorrelation decoding scheme.
[0242] The following describes how to determine the first channel (C1), the second channel (C2), and the third channel (C3) using an example where the scene audio signal is an Nth-order HOA signal.
[0243] For example, the first C1 channel can be one of the C1 channels included in the zeroth to Mth order signals of an Nth-order HOA signal, where M is an integer less than N, C is equal to the square of (N+1), and C1 is less than or equal to the square of (M+1).
[0244] For example, the second set of two C channels may include four other C channels included in the zero- to M-th order signals of the N-th order HOA signal, and five C channels other than those included in the zero- to M-th order signals of the N-th order HOA signal.
[0245] For example, the third channel C3 may include the other C6 channels included in the zero- to M-th order signals of the N-th order HOA signal, and the other C7 channels other than those included in the zero- to M-th order signals of the N-th order HOA signal.
[0246] C2 = C4 + C5, C3 = C6 + C7, and C1 + C4 + C6 = the square of (M + 1), where C4, C5, C6, and C7 are integers.
[0247] When N=3, C=16. Specifically, the third-order HOA signal contains 16 channels.
[0248] If the value of n in equation (3) is 0, a single monomial can be obtained by expanding equation (3), as shown in equation (5). In this case, an audio signal with one channel (hereinafter referred to as channel 1) can be obtained. If the value of n in equation (3) is 1, three monomials can be obtained by expanding equation (3), as shown in equation (6). In this case, an audio signal with three channels (referred to as channel 2, channel 3, and channel 4, respectively) can be obtained. 3 If the value of n in () is 2, then five monomials can be obtained by expanding equation (3), as shown in equation (7). In this case, an audio signal with five channels (referred to as channel 5, channel 6, channel 7, channel 8, and channel 9, respectively) can be obtained. 3When the value of n in (0) is 3, as shown in Equation (8), by expanding Equation (3), seven monomials can be obtained. In this case, an audio signal having seven channels (which are called Channel 10, Channel 11, Channel 12, Channel 13, Channel 14, Channel 15, and Channel 16 in sequence) can be obtained. Each monomial corresponds to one channel, and each monomial can represent an audio signal having one channel.
[0249]
Number
[0250]
Number
[0251]
Number
[0252]
Number
[0253] is.
[0254] In this specification,
[0255]
Number
[0256] is the position information of the sound source in the scene audio signal.
[0257] For example, if M=0, that is, if m in equation (3) is equal to 0, then the value of n can be 0. By expanding equation (3), a monomial can be obtained. In other words, the zero-order signal in the cubic HOA signal contains one channel, i.e., channel 1. In this case, C1=1. Specifically, the first channel is channel 1, and C4=0 and C6=0. Accordingly, in addition to the channel contained in the zero-order signal, the cubic HOA signal further contains 15 channels from channel 2 to channel 16. From these 15 channels, C5 channels can be selected as the second channel, and C7 channels can be selected as the third channel.
[0258] Here is an example.
[0259] If C5=14 and C7=1, the second channel with 2 Cs will be channels 2 through 15, and the third channel with 3 Cs will be channel 16. Alternatively, the second channel with 2 Cs will be channels 2 through 13 and channel 16, and the third channel with 3 Cs will be channel 15. This can be specifically determined according to the requirements. This is not limited to this application.
[0260] For example, if M=1, that is, if m in equation (3) is equal to 1, then the value of n can be either 0 or 1. By expanding equation (3), four monomials can be obtained. In other words, the zero and primary signals in the cubic HOA signal contain four channels, i.e., channels 1 through 4. In this case, from the four channels, C1 channels may be selected as the first channels, the other C4 channels may be selected as the second channels, and the other C6 channels may be selected as the third channels. Accordingly, in addition to the channels contained in the zero and primary signals, the cubic HOA signal further contains 12 other channels, i.e., channels 5 through 16. Of the 12 channels, C5 channels may be selected as the second channels, and C7 channels may be selected as the third channels.
[0261] Here is an example.
[0262] If C1=4, C4=0, C6=0, C5=10, and C7=2, then the first channel with C1 will be channels 1 through 4, the second channel with C2 will be channels 5 through 14, and the third channel with C3 will be channels 16 and 15. Alternatively, the first channel with C1 will be channels 1 through 4, the second channel with C2 will be channels 5 through 10 and channels 13 through 16, and the third channel with C3 will be channels 11 and 12. This can be specifically determined as needed. This is not limited to the present application.
[0263] For example, if M=2, that is, if m in equation (3) is equal to 2, then the value of n can be 0, 1, or 2. By expanding equation (3), nine monomials can be obtained. In other words, the zero- to secondary signals in the cubic HOA signal contain nine channels, i.e., channels 1 through 9. In this case, from the nine channels, C1 channels may be selected as the first channels, another C4 channels may be selected as the second channels, and another C6 channels may be selected as the third channels. Accordingly, in addition to the channels contained in the zero- to secondary signals, the cubic HOA signal further contains seven other channels, i.e., channels 10 through 17. Of the seven channels, C5 channels may be selected as the second channels, and C7 channels may be selected as the third channels.
[0264] Here is an example.
[0265] For example, when C1 = 8, C4 = 1, C6 = 0, C5 = 5, and C7 = 2, the C1 first channels include channels 1 to 5 and channels 7 to 9, the C2 second channels include channel 6 and channels 11 to 15, and the C3 third channels include channel 10 and channel 16. Correspondingly, the encoding method corresponding to the 16 channels can be shown as "sub - combination 2" in the third column of Table 1.
[0266] For example, when C1 = 9, C4 = 0, C6 = 0, C5 = 5, and C7 = 2, the C1 first channels include channels 1 to 9, the C2 second channels include channels 11 to 15, and the C3 third channels include channel 10 and channel 16. Correspondingly, the encoding method corresponding to the 16 channels can be shown in the four nth column of Table 1 as "sub - combination 3".
[0267] It should be understood that C1, C2, and C3 may be set to other values, and the first channels, second channels, and third channels may be selected in other manners. In this way, for the 16 channels within one frame of the scene audio signal, r (r is a positive integer greater than or equal to 2) different sub - combinations can be obtained. This is not limited in this application.
[0268]
Table 1
[0269] For example, when the encoding combination method is a combination of the first encoding method and the third encoding method, the C8 fourth channels can be encoded based on the first encoding method, and the C9 fifth channels can be encoded based on the third encoding method.
[0270] The fourth set of C8 channels consists of C8 channels included in the zeroth to Mth order signals within the Nth order HOA signal. M is an integer less than N, C is equal to the square of (N+1), and C8 is less than or equal to the square of (M+1). The fifth set of C9 channels consists of the other C10 channels included in the zeroth to Mth order signals within the Nth order HOA signal, plus the channels not included in the zeroth to Mth order signals within the Nth order HOA signal. C8 + C10 is equal to the square of (M+1).
[0271] For example, if M=0, that is, if m in equation (3) is equal to 0, then the value of n can be 0. By expanding equation (3), a monomial can be obtained. In other words, the zero-order signal in the cubic HOA signal contains one channel, i.e., channel 1. In this case, C8=1. Specifically, the fourth channel is channel 1, and C10=0. Accordingly, in addition to the channel contained in the zero-order signal, the cubic HOA signal further contains 15 other channels, i.e., channels 2 through 15, which are all fifth channels. In this case, the encoding scheme corresponding to the 16 channels can be shown as "Subcombination 1" in the second column of Table 1.
[0272] For example, when M=1, that is, when m in equation (3) is equal to 1, the value of n can be either 0 or 1. By expanding equation (3), four monomials can be obtained. In other words, the zero and primary signals in the cubic HOA signal contain four channels, namely channels 1 through 4. In this case, from the four channels, C8 channels can be selected as the third channel, and the remaining C10 channels can be selected as the fifth channel. Accordingly, in addition to the channels contained in the zero and primary signals, the cubic HOA signal further contains 12 channels, namely channels 5 through 16, all of which are the fifth channel.
[0273] Here is an example.
[0274] If C8=4, C10=0, and C9=12, then C8 The nine fourth channels become channels 1 through 4, and the nine fifth channels become channels 5 through 16.
[0275] If C8=2, C10=2, and C9=12, then C 8 The fourth channel becomes channel 1 or channel 2, and the fifth channel (C9) becomes channel 3 through channel 16.
[0276] For example, when M=2, that is, when m in equation (3) is equal to 2, the value of n can be 0, 1, or 2. By expanding equation (3), nine monomials can be obtained. In other words, the zero- and secondary signals in the cubic HOA signal contain nine channels, i.e., channels 1 through 9. In this case, of the nine channels, C8 channels can be selected as the fourth channel, and the remaining C10 channels can be selected as the fifth channel. Accordingly, in addition to the channels contained in the zero- and secondary signals, the cubic HOA signal further contains seven other channels, i.e., channels 10 through 16, all of which are the fifth channel.
[0277] Here is an example.
[0278] For example, if C8=8, C10=0, and C9=8, then the C8 fourth channels include channels 1 through 8, and the C9 fifth channels include channels 9 through 16.
[0279] The values of M, C1 through C10, and the channels specifically included in the first channel (C1), the second channel (C2), the third channel (C3), the fourth channel (C8), and the fifth channel (C9) can be determined based on the characteristics of the audio signals having different channels. The characteristics of the audio signals having different channels may include direction, and the directions of the audio signals having different channels may be shown in Table 2.
[0280] [Table 2]
[0281] For example, in an Nth-order HOA signal, C7 channels whose vertical components are greater than the threshold, excluding channels included in the zeroth to Mth-order signals, are selected as the third channel, and these C7 channels are included in the audio signal. For example, if the vertical components of an audio signal having channels 10 and 16 are greater than the threshold, then, as shown in "Sub-combination 2" and "Sub-combination 3" in Table 1, two channels, namely channels 10 and 16, are selected as the third channel.
[0282] In one possible embodiment, the decoder and encoder agree in advance on a specific encoding scheme (or a specific decoding scheme) to be used for encoding each channel within a single frame of the scene audio signal.
[0283] In one possible embodiment, the decoder and encoder do not agree in advance on the specific encoding scheme (or specific decoding scheme) to be used for encoding each channel within a single frame of the scene audio signal. In this case, the encoder may encode a fourth preset identifier, which may indicate the encoding scheme corresponding to each channel. Optionally, the encoder may... fourIt may not always be necessary to encode the preset identifier. In this way, after the fourth preset identifier is sent to the decoder, the decoder performs the decoding.
[0284] Figure 5 shows an example of scene audio decoding. The embodiment in Figure 5 is the decoding process corresponding to the encoding process in the embodiment in Figure 4a.
[0285] S501: Receives a bitstream.
[0286] S502: Based on the bitstream, one first channel is decoded based on the first decoding method, two second channels are decoded based on the second decoding method, and three third channels are decoded based on the third decoding method to obtain one frame of the reconstructed scene audio signal.
[0287] For example, a reconstructed scene audio signal may contain a reconstructed audio signal having C channels. For example, for one frame of the reconstructed scene audio signal, it can be determined that of the C channels contained in the frame of the reconstructed scene audio signal, C1 first channels are decoded according to a first decoding scheme, C2 second channels are decoded according to a second decoding scheme, and C3 third channels are decoded according to a third decoding scheme. C is equal to the sum of C1, C2, and C3, C, C1, C2, and C3 are positive integers, and the first channel, second channel, and third channel are distinct channels.
[0288] Next, in order to obtain the bitstream of the scene audio signal frames, based on the bitstream, C1 first channels are decoded based on the first decoding scheme, C2 second channels are decoded based on the second decoding scheme, and C3 third channels are decoded based on the third decoding scheme.
[0289] For example, based on a bitstream, the process of decoding C1 first channels according to a first decoding method may be as follows. That is, the step of analyzing the bitstream to determine the encoded data corresponding to an audio signal having C1 first channels. And then, performing operations such as entropy decoding, inverse quantization, bit allocation, post-processing, and time-frequency conversion on the encoded data corresponding to the audio signal having C1 first channels to obtain a reconstructed audio signal having C1 first channels.
[0290] For example, based on a bitstream, the process of decoding C2 second channels according to a second decoding method may include S21 to S24. C1 is equal to (M + 1) 2 is equal to.
[0291] S21: Determine a first virtual speaker coefficient corresponding to the target virtual speaker based on the attribute information of the target virtual speaker.
[0292] For example, the encoder side may write M into the first bitstream, and further, M may be obtained from the first bitstream through decoding (certainly, the encoder side and the decoder side may agree on M in advance, which is not limited in this application). For example, when the attribute information of the target virtual speaker is position information, the position information of the target virtual speaker may be substituted into Equation (3), and m in Equation (3) becomes equal to M, whereby the first virtual speaker coefficient corresponding to the target virtual speaker may be obtained. The first virtual speaker coefficient includes (M + 1) 2 virtual speaker coefficients, and the (M + 1) 2 virtual speaker coefficients correspond to C1 first channels.
[0293] For example, if the attribute information of the target virtual speaker is a location index of location information, the location information of the target virtual speaker can be determined based on the relationship between the location information and the location index. Then, in the embodiment described above, the first virtual speaker coefficient can be determined. This will not be explained again in this specification.
[0294] For example, if the attribute information of the target virtual speaker is a virtual speaker index, the location information of the target virtual speaker can be determined based on the relationship between the location information of the target virtual speaker and the virtual speaker index. Then, in the embodiment described above, the first virtual speaker coefficient is determined. This will not be explained again in this specification.
[0295] For example, if the attribute information of the target virtual speaker is virtual speaker coefficients, as can be seen from the explanation above, the group of virtual speaker coefficients corresponding to the target virtual speaker includes C virtual speaker coefficients. In this case, the virtual speaker coefficients corresponding to the C1 first channels included in the reconstructed audio signal may be selected as the first virtual speaker coefficients.
[0296] S22: A virtual speaker signal is generated based on a reconstructed audio signal having one first channel and a first virtual speaker coefficient.
[0297] For example, a virtual speaker signal may be generated based on a reconstructed audio signal having one first channel and a first virtual speaker coefficient.
[0298] For example, suppose a matrix A of size (Y1 × P) represents the first virtual speaker coefficient of the target virtual speaker. In this specification, Y1 (where Y1 is a positive integer) is the number of target virtual speakers, and P is the number of audio signal channels (M+1) contained in the reconstructed audio signal having C1 first channels. 2Furthermore, a matrix X of size (L×P) represents a reconstructed audio signal having C1 first channels, where L is the number of sampling points in the reconstructed audio signal having C1 first channels. The theoretical optimal solution w is obtained by the least squares method, and as shown in equation (9), w represents a virtual speaker signal.
[0299]
number
[0300] matrix A -1 This is the inverse matrix of matrix A.
[0301] S23: Based on the attribute information of the target virtual speaker, the second virtual speaker coefficient corresponding to the target virtual speaker is determined.
[0302] For example, based on the expected number of dimensions N of the reconstructed scene audio signal, it can be determined that m in equation (3) is equal to N. Then, if the attribute information of the target virtual speaker is location information, the location information of the target virtual speaker can be substituted into equation (3), and m in equation (3) becomes equal to N. This allows a second virtual speaker coefficient to be obtained. The second virtual speaker coefficient contains C virtual speaker coefficients, and the C virtual speaker coefficients correspond to the C channels in the reconstructed scene audio signal.
[0303] For example, if the attribute information of the target virtual speaker is a location index of location information, the location information of the target virtual speaker can be determined based on the relationship between the location information and the location index. Then, in the embodiment described above, the first virtual speaker coefficient is determined. This will not be explained again in this specification.
[0304] For example, if the attribute information of the target virtual speaker is a virtual speaker index, the location information of the target virtual speaker can be determined based on the relationship between the location information and the virtual speaker index. Then, in the embodiment described above, the first virtual speaker coefficient is determined. This will not be explained again in this specification.
[0305] For example, if the attribute information of the target virtual speaker is a virtual speaker coefficient, then the attribute information of the target virtual speaker can be directly used as a second virtual speaker coefficient.
[0306] S24: Based on the virtual speaker signal and the second virtual speaker coefficient, a reconstructed audio signal with C2 second channels is obtained.
[0307] For example, suppose a matrix A with size (Y1 × C) represents the second virtual speaker coefficient, where Y1 is the number of target virtual speakers and C is the number of channels in the reconstructed scene audio signal. Furthermore, a matrix B with size (L × Y1) represents the virtual speaker signal, where L is the number of sampling points in the reconstructed scene audio signal. In this case... , re The constituent scene audio signal can be represented by H, as shown in equation (10).
[0308]
number
[0309] Next , re From the constructed scene audio signal, a reconstructed audio signal having two second channels can be selected.
[0310] In one possible embodiment, during the encoding process, feature information corresponding to the two second channels C in the scene audio signal may be further extracted, encoded, and transmitted to the decoder. After receiving the bitstream, the decoder may correct the reconstructed audio signal having the two second channels C in the reconstructed scene audio signal based on the feature information, thereby improving the sound quality of the reconstructed audio signal having the two second channels C in the reconstructed scene audio signal.
[0311] For example, the gain information Gain(i) corresponding to the two second channels C in the scene audio signal can be calculated by referring to equation (11).
[0312]
number
[0313] In this specification, j is the channel number of C2 channels, E(i) is the energy of the i-th channel, and E(1) is the energy of the audio signal having C channels in the scene audio signal.
[0314] For example, if the feature information is gain information, correction may be performed by referring to equation (12).
[0315]
number
[0316] In this specification, j is the channel number of the two second channels, E(i) is the energy of the i-th channel, and E(l) is the energy of the reconstructed audio signal having C channels in the reconstructed scene audio signal.
[0317]
number
[0318] This is gain information corresponding to the second channel of C2 in the scene audio signal.
[0319] In one possible embodiment, the process of decoding C3 third channels based on a bitstream and a third decoding scheme may be as follows: a step of processing a reconstructed audio signal having one or more channels in a first channel C1 by using an all-pass filter to obtain a reconstructed audio signal having C3 third channels.
[0320] For example, Table 1 is pre-synchronized between the decoder and encoder sides, and the corresponding decoder side stores Table 3, which corresponds to Table 1.
[0321] [Table 3]
[0322] In one possible embodiment, if the decoder and encoder have agreed in advance on a specific encoding scheme (or a specific decoding scheme) to be used for encoding each channel within a frame of the scene audio signal, the decoder may perform decoding for each channel within a frame of the scene audio signal based on the pre-agreed decoding scheme.
[0323] In one possible embodiment, if the decoder and encoder have not previously agreed on a specific encoding scheme (or a specific decoding scheme) to be used for encoding each channel within a single frame of the scene audio signal, the decoder may parse a fourth preset identifier from the bitstream. The decoder may then determine the decoding scheme for each channel based on the fourth preset identifier and decode each frame of the bitstream based on the decoding scheme for each channel to obtain each frame of the reconstructed scene audio signal. In other words, the encoder and decoder encode and decode the same channel within the same frame based on the corresponding encoding and decoding schemes.
[0324] In one possible embodiment, the i-th frame of the scene audio signal is encoded based on multiple encoding schemes in a k1-th combination of encoding schemes. Then, the j-th frame of the scene audio signal is encoded based on multiple encoding schemes in a k2-th combination of encoding schemes. Herein, i and j are positive integers between 1 and X, i is not equal to j, and k1 and k2 are positive integers, k1 is not equal to k2. That is, different frames of the scene audio signal are encoded based on multiple encoding scheme combinations in different types of encoding scheme combinations.
[0325] Accordingly, after receiving the bitstream, the decoder decodes the i-th frame of the bitstream using one of the multiple decoding schemes in the k1-th decoding scheme combination to obtain the i-th frame of the reconstructed scene audio signal. Then, the decoder decodes the j-th frame of the bitstream using one of the multiple decoding schemes in the k2-th decoding scheme combination to obtain the j-th frame of the reconstructed scene audio signal.
[0326] Figure 6 shows an example of scene audio encoding processing. In the embodiment shown in Figure 6, multiple bandwidths of the same channel within a single frame are encoded based on multiple encoding schemes in a combination of encoding schemes.
[0327] S601: Acquire the scene audio signal. Here, the scene audio signal includes an audio signal having C channels, and each audio signal having one channel has Y bandwidths.
[0328] For example, an audio signal having each channel within a single frame of a scene audio signal can be divided into Y bandwidths, where Y can be set as needed. This is not limited to the present application.
[0329] For example, Y=2. Specifically, an audio signal with one channel may contain band 1 and band 2. The frequencies of band 1 are below the first frequency threshold, and the frequencies of band 2 are above the first frequency threshold.
[0330] For example, Y=3. Specifically, an audio signal with one channel may include band 1, band 2, and band 3. The frequency of band 1 is less than the first frequency threshold, the frequency of band 2 is greater than the first frequency threshold but less than the second frequency threshold, and the frequency of band 3 is greater than the second frequency threshold.
[0331] It should be noted that the number of bandwidths Y may differ for different channels, and the thresholds used for bandwidth division may also differ. This is not limited to this application. The first and second frequency thresholds may be set as needed. Further details are not described herein.
[0332] S602: Encodes Y bandwidths of one channel in a scene audio signal based on multiple encoding schemes in a combination of encoding schemes.
[0333] The following explanation is provided using one channel as an example.
[0334] For example, if Y=2 and the encoding scheme combination is a combination of the first and third encoding schemes, then, as shown in subcombination 1 in Table 4, the audio signal in channel band 1 may be encoded based on the first encoding scheme, and the audio signal in channel band 2 may be encoded based on the third encoding scheme. Alternatively, the audio signal in channel band 1 may be encoded based on the third encoding scheme, and the audio signal in channel band 2 may be encoded based on the first encoding scheme.
[0335] For example, if Y=2 and the combination of encoding schemes is the first, second, and third encoding schemes, then, as shown in subcombination 2 in Table 4, the audio signal in channel band 1 may be encoded based on the first encoding scheme, and the audio signal in channel band 2 may be encoded based on the second encoding scheme. Alternatively, the audio signal in channel band 1 may be encoded based on the second encoding scheme, and the audio signal in channel band 2 may be encoded based on the third encoding scheme. Alternatively, the audio signal in channel band 1 may be encoded based on the first encoding scheme, and the audio signal in channel band 2 may be encoded based on the third encoding scheme.
[0336] For example, if Y=3 and the encoding scheme combination is a combination of the first and third encoding schemes, then, as shown in subcombination 3 in Table 4, the audio signals in channel band 1 and band 2 may be encoded based on the first encoding scheme, and the audio signal in channel band 3 may be encoded based on the third encoding scheme. Alternatively, the audio signals in channel band 1 and band 3 may be encoded based on the first encoding scheme, and the audio signal in channel band 2 may be encoded based on the third encoding scheme. Alternatively, the audio signals in channel band 1 and band 2 may be encoded based on the third encoding scheme, and the audio signal in channel band 3 may be encoded based on the first encoding scheme. Or similar.
[0337] For example, if Y=3 and the combination of encoding schemes is the first, second, and third encoding schemes, then, as shown in subcombination 4 in Table 4, the audio signal in channel band 1 may be encoded based on the first encoding scheme, the audio signal in channel band 2 may be encoded based on the second encoding scheme, and the audio signal in channel band 3 may be encoded based on the third encoding scheme. Alternatively, the audio signal in channel band 1 may be encoded based on the second encoding scheme, the audio signal in channel band 2 may be encoded based on the first encoding scheme, and the audio signal in channel band 3 may be encoded based on the third encoding scheme. Or, similarly.
[0338] In this way, for a single channel, w (where w is a positive integer) subcombinations can be generated, and different coding scheme combinations may contain one or more subcombinations.
[0339] [Table 4]
[0340] In one possible embodiment, the decoder and encoder may pre-agree on a specific encoding scheme (or a specific decoding scheme) to be used for encoding Y bandwidths of one channel within a single frame of the scene audio signal.
[0341] In one possible embodiment, the decoder and encoder do not agree in advance on the specific encoding scheme (or specific decoding scheme) to be used for encoding the Y bandwidths of a single channel within a single frame of the scene audio signal. In this case, the encoder may encode a fifth preset identifier, which may indicate the encoding scheme corresponding to each bandwidth of each channel. Optionally, the encoder may... Five It may not be necessary to encode the preset identifier. In this way, after the fifth preset identifier is sent to the decoder, the decoder performs the decoding process.
[0342] It should be understood that multiple bandwidths of different channels within a single frame of a scene audio signal may be encoded based on multiple encoding schemes in combinations of different types of decoding schemes. This is not limited to the present application.
[0343] Figure 7 shows an example of scene audio decoding. The embodiment in Figure 7 is a decoding process that corresponds to the encoding process in the embodiment in Figure 6.
[0344] S701: Receives a bitstream.
[0345] For example, a reconstructed audio signal having each channel within a single frame of a reconstructed scene audio signal can be divided into Y bandwidths, where Y can be set as needed. This is not limited to the present application.
[0346] For example, Y=2. Specifically, a reconstructed audio signal with one channel may contain band 1 and band 2. The frequencies of band 1 are below the first frequency threshold, and the frequencies of band 2 are above the first frequency threshold.
[0347] For example, Y=3. Specifically, a reconstructed audio signal with one channel may include band 1, band 2, and band 3. The frequency of band 1 is less than the first frequency threshold, the frequency of band 2 is greater than the first frequency threshold but less than the second frequency threshold, and the frequency of band 3 is greater than the second frequency threshold.
[0348] It should be noted that the number of bandwidths Y may differ for different channels, and the thresholds used for bandwidth division may also differ. This is not limited to this application. The first and second frequency thresholds may be set as needed. Further details are not described herein.
[0349] It is important to understand that the manner in which the decoder divides each channel into bandwidths is the same as the manner in which the encoder divides each channel into bandwidths.
[0350] S702: For Y bandwidths of one channel within a single frame, one frame of the bitstream is decoded based on multiple decoding schemes in a combination of decoding schemes, and one frame of the reconstructed scene audio signal is obtained.
[0351] The following explanation is provided using one channel as an example.
[0352] For example, if Y=2 and the decoding scheme combination is a combination of the first decoding scheme and the third decoding scheme, then, as shown in subcombination 1 in Table 5, the audio signal in channel band 1 may be decoded based on the first decoding scheme, and the audio signal in channel band 2 may be decoded based on the third decoding scheme. Alternatively, the audio signal in channel band 1 may be decoded based on the third decoding scheme, and the audio signal in channel band 2 may be decoded based on the first decoding scheme.
[0353] For example, if Y=2 and the combination of decoding schemes is the first decoding scheme, the second decoding scheme, and the third decoding scheme, then, as shown in subcombination 2 in Table 5, the audio signal in channel band 1 may be decoded based on the first decoding scheme, and the audio signal in channel band 2 may be decoded based on the second decoding scheme. Alternatively, the audio signal in channel band 1 may be decoded based on the second decoding scheme, and the audio signal in channel band 2 may be decoded based on the third decoding scheme. Alternatively, the audio signal in channel band 1 may be... one The audio signal in channel bandwidth 2 is decoded based on the first decoding scheme, and the audio signal in channel bandwidth 2 is decoded based on the third decoding scheme.
[0354] For example, if Y=3 and the decoding scheme combination is a combination of the first and third decoding schemes, then, as shown in subcombination 3 in Table 5, the audio signals in channel band 1 and channel band 2 may be decoded based on the first decoding scheme, and the audio signal in channel band 3 may be decoded based on the third decoding scheme. Alternatively, the audio signals in channel band 1 and channel band 3 may be decoded based on the first decoding scheme, and the audio signal in channel band 2 may be decoded based on the third decoding scheme. Alternatively, the audio signals in channel band 1 and channel band 2 may be decoded based on the third decoding scheme, and the audio signal in channel band 3 may be decoded based on the first decoding scheme. Or similar.
[0355] For example, if Y=3 and the combination of decoding schemes is the first, second, and third decoding schemes, then, as shown in subcombination 4 in Table 5, the audio signal in channel band 1 may be decoded based on the first decoding scheme, the audio signal in channel band 2 may be decoded based on the second decoding scheme, and the audio signal in channel band 3 may be decoded based on the third decoding scheme. Alternatively, the audio signal in channel band 1 may be decoded based on the second decoding scheme, the audio signal in channel band 2 may be decoded based on the first decoding scheme, and the audio signal in channel band 3 may be decoded based on the third decoding scheme. Or similar.
[0356] In this way, for a single channel, w (where w is a positive integer) subcombinations can be generated, and different decoding scheme combinations may include one or more subcombinations.
[0357] [Table 5]
[0358] In one possible embodiment, if the decoder and encoder have previously agreed on a specific encoding scheme (or a specific decoding scheme) to be used for encoding Y bands of one channel within a frame of a scene audio signal, the decoder may decode the corresponding bands based on the previously agreed encoding scheme and obtain a reconstructed audio signal having that channel.
[0359] In one possible embodiment, if the decoder and encoder have not previously agreed on a specific encoding scheme (or a specific decoding scheme) to be used for encoding Y frequency bands of one channel within a single frame of the scene audio signal, after receiving the bitstream, the decoder parses a fifth preset identifier from the bitstream, and then, based on the fifth preset identifier, determines the Y frequency bands of each channel. band The decoding scheme for each band within the region can be determined. Then, in order to obtain a reconstructed audio signal having that channel, the bitstream is decoded based on the decoding scheme determined for the Y bands of each channel based on the fifth preset identifier.
[0360] Figure 8 shows an example of scene audio encoding processing. In the embodiment shown in Figure 8, a portion of the frame is encoded based on multiple encoding schemes in a combination of encoding schemes, and another portion of the frame is encoded based on a first encoding scheme in a combination of encoding schemes.
[0361] S801: Acquires the scene audio signal.
[0362] S802: Encodes the i-th frame of the scene audio signal based on multiple encoding schemes in the combination of encoding schemes.
[0363] For example, please refer to the explanation above regarding S802. Further details will not be provided in this specification.
[0364] S803: Encode the j-th frame of the scene audio signal based on the first encoding scheme in the combination of encoding schemes.
[0365] In this specification, i and j are positive integers between 1 and X, i is not equal to j, and k1 and k2 are positive integers, k1 is not equal to k2.
[0366] For example, regardless of whether the combination of encoding schemes is a combination of the first and third encoding schemes, or a combination of the first, second, and third encoding schemes, the j-th frame of the scene audio signal is encoded based on the first encoding scheme in the combination of encoding schemes.
[0367] In this application, the number of frames of the scene audio signal encoded based on the first encoding scheme in the combination of encoding schemes, and the number of frames of the scene audio signal encoded using multiple encoding schemes in the combination of encoding schemes, are not limited. Furthermore, in this application, the specific frames encoded based on the first encoding scheme in the combination of encoding schemes are not limited, nor are the specific frames encoded based on multiple encoding schemes in the combination of encoding schemes.
[0368] In one possible embodiment, the encoder and decoder may agree in advance on specific frames to be encoded based on a first encoding scheme in the combination of encoding schemes, and may also agree in advance on specific frames to be encoded based on multiple encoding schemes in the combination of encoding schemes.
[0369] In one possible embodiment, the encoder and decoder do not agree in advance on the encoding and decoding schemes for each frame. In this case, if the i-th frame of the scene audio signal is encoded based on multiple encoding schemes in the encoding scheme combination, a second preset identifier (which indicates the type of encoding scheme combination corresponding to each frame) is written to the i-th frame of the bitstream. Then, if the j-th frame of the scene audio signal is encoded based on the first encoding scheme in the encoding scheme combination, a sixth preset identifier is written to the j-th frame of the bitstream. The sixth preset identifier indicates that the current frame is encoded based on the first encoding scheme.
[0370] Figure 9 shows an example of scene audio decoding. Figure 9 shows the decoding process corresponding to the encoding process in Figure 8.
[0371] S901: Receives a bitstream.
[0372] S902: Based on multiple decoding schemes in the combination of decoding schemes, the i-th frame of the bitstream is decoded to obtain the i-th frame of the reconstructed scene audio signal.
[0373] For example, please refer to the explanation above regarding S902. Further details will not be provided in this specification.
[0374] S903: Based on the decoding scheme in the combination of decoding schemes, the j-th frame of the bitstream is decoded and the j-th frame of the reconstructed scene audio signal is obtained.
[0375] In this specification, i and j are positive integers between 1 and X, i is not equal to j, and k1 and k2 are positive integers, k1 is not equal to k2.
[0376] For example, regardless of whether the combination of decoding schemes is a combination of the first and third decoding schemes, or a combination of the first, second, and third decoding schemes, the j-th frame of the scene audio signal is decoded based on the first decoding scheme in the combination of decoding schemes.
[0377] In this application, the number of frames of the scene audio signal decoded based on the first decoding scheme in the combination of decoding schemes, and the number of frames of the scene audio signal decoded by the multiple decoding schemes in the combination of decoding schemes, are not limited.
[0378] In one possible embodiment, if the decoder and encoder have previously agreed on a particular frame to be encoded based on multiple schemes in the combination of encoding schemes and a particular frame to be encoded in the first encoding scheme in the combination of encoding schemes, the decoder may decode the bitstream in the previously agreed encoding scheme to obtain a reconstructed scene audio signal.
[0379] In one possible embodiment, if the decoder and encoder have not previously agreed on a particular frame to be encoded in multiple schemes in the combination of encoding schemes and a particular frame to be encoded based on the first encoding scheme in the combination of encoding schemes, after receiving the bitstream, the decoder will then consider multiple decoding scheme combinations corresponding to a second preset identifier. revenge Based on the coding scheme, it is possible to decode the frames of the bitstream to obtain one frame of the reconstructed scene audio signal. Once a sixth preset identifier is parsed from a particular frame of the bitstream, the frames of the bitstream are first... revenge Decryption is performed based on the encoding scheme.
[0380] Figure 10 is a block diagram showing an apparatus 1000 according to one embodiment of the present invention. The apparatus 1000 may include a processor 1001 and a transceiver / transceiver pin 1002, and optionally further include a memory 1003.
[0381] Each component of the device 1000 is interconnected via bus 1004. In addition to the data bus, bus 1004 further includes a power bus, a control bus, and a status signal bus. However, for clarity of explanation, in the diagram, the various buses are referred to as bus 1004.
[0382] Optionally, memory 1003 may be configured to store instructions in the embodiments of the method described above. The processor 1001 may be configured to execute instructions in memory 1003, control the receive pin to receive signals, and control the transmit pin to transmit signals. When the processor 1001 is configured to execute instructions in memory 1003, This device This makes it possible to implement the method in the embodiment described above by performing the steps of the related method described above.
[0383] The device 1000 may be an electronic device in the embodiment of the method described above, or it may be a chip of an electronic device.
[0384] The electronic device may be a terminal device or a server.
[0385] All relevant details of the steps in the embodiments of the method described above can be referenced in the functional description of the corresponding functional module. Further details will not be described further in this specification.
[0386] One embodiment of this application further provides a chip comprising one or more interface circuits and one or more processors. One or more processors receive or transmit data through one or more interface circuits. When one or more processors execute computer instructions, the electronic device can perform the steps of the associated method described above to implement the method in the embodiment described above. The interface circuit is a transceiver / transmit pin / receive pin 1002.
[0387] Furthermore, one embodiment provides a computer-readable storage medium. This computer-readable storage medium stores computer instructions. When these computer instructions are executed on an electronic device, the electronic device can perform the steps of the related method described above to implement the method in the embodiment described above.
[0388] Furthermore, one embodiment provides a computer program product. This computer program product includes computer instructions, and when these instructions are executed by a computer or processor, the computer is able to perform the associated steps described above and carry out the method in the embodiment described above.
[0389] In addition, one embodiment of the present application further provides an apparatus, which may specifically be a chip, component, or module. The apparatus may include an attached processor and memory. The memory is configured to store computer executable instructions. When the apparatus is in operation, the processor executes the computer executable instructions stored in memory, enabling the chip to perform the method in the embodiment of the method described above.
[0390] An electronic device, computer-readable storage medium, computer program product, or chip provided in one embodiment is configured to perform the corresponding method provided above. Therefore, for the beneficial effects that can be achieved, please refer to the beneficial effects of the corresponding method provided above. Further details are not described herein.
[0391] Based on the above-described embodiments, those skilled in the art will understand that, for the sake of convenience and brevity of explanation, the division into functional modules described above is used as an example. In actual applications, the functions described above may be assigned to different functional modules and implemented as required. In other words, the internal configuration of the device is divided into different functional modules in order to implement all or some of the functions described above.
[0392] In some embodiments provided in this application, it should be understood that the disclosed apparatus and methods may be implemented in other embodiments. For example, the embodiments of the described apparatus are merely examples. For example, the division into modules or units is merely a logical functional division, and other divisions may be in actual implementation. For example, multiple units or components may be combined or integrated into another apparatus, or some functions may be ignored or not performed. Furthermore, the mutual or direct coupling or communication connection shown or described may be implemented through some interface. Indirect coupling or communication connection between apparatus or units may be implemented in an electrical, mechanical, or other form.
[0393] Units described as separate parts may or may not be physically separated, and parts shown as units may be one or more physical units, may be located in one place, or may be distributed in different locations. Some or all units may be selected depending on the actual requirements in order to achieve the objectives of the solutions of the embodiments.
[0394] Furthermore, the functional units in the embodiments of this application may be integrated into a single processing unit, each unit may exist physically independently, or two or more units may be integrated into a single unit. The integrated unit may be implemented in hardware form or in the form of a software functional unit.
[0395] Any content in any embodiment of this application, and any content in the same embodiment, can be freely combined. Any combination of the above-described content is included within the scope of this application.
[0396] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit may be stored in a readable storage medium. Based on this understanding, the essence of the technical solution of the embodiments of this application, or the portion that contributes to the prior art, or all or part of the technical solution, may be implemented in the form of a software product. The software product is stored in a storage medium and includes a number of instructions for instructing a device (single-chip microcomputer, chip, or similar) or processor to perform all or part of the steps of the method described in the embodiments of this application. The storage medium includes various media capable of storing program code, such as USB flash drives, removable hard disk drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0397] Embodiments of this application have been described above with reference to the attached drawings. However, this application is not limited to the specific implementations described above. The specific implementations described above are merely examples, not limitations. A person skilled in the art, inspired by this application, may make further modifications without departing from the purpose and scope of the claims of this application, and all such modifications shall be within the scope of the application.
[0398] Steps of methods or algorithms described in conjunction with those disclosed in these embodiments of this application may be implemented by hardware or by a processor executing software instructions. Software instructions may include corresponding software modules. Software modules may be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art. For example, the storage medium may be coupled to a processor so that the processor can read information from and write information to the storage medium. Indeed, the storage medium may be a component of the processor. The processor and the storage medium may be located within an ASIC.
[0399] Those skilled in the art should recognize that, in one or more of the examples described above, the functions described in the embodiments of this application may be implemented by hardware, software, firmware, or any combination thereof. If these functions are implemented by software, they may be stored in a computer-readable medium or transmitted as one or more instructions or codes within a computer-readable medium. Computer-readable mediums include computer-readable storage media and communication media, where communication media include any medium that enables the transmission of a computer program from one location to another. Storage media may be any available medium accessible to a general-purpose computer or a dedicated computer.
[0400] Embodiments of this application have been described above with reference to the attached drawings. However, this application is not limited to the specific implementations described above. The specific implementations described above are merely examples, not limitations. A person skilled in the art, inspired by this application, may make further modifications without departing from the purpose and scope of the claims of this application, and all such modifications shall be within the scope of the application.
Claims
1. A scene audio decoding method, The steps include receiving a bitstream, A step of decoding the bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal, wherein the combination of decoding schemes includes at least one of the following combinations: a combination of a first decoding scheme, a second decoding scheme, and a third decoding scheme, and a combination of the first decoding scheme and the third decoding scheme. Equipped with, The first decoding method is a method for decoding encoded data obtained by encoding a signal, the second decoding method is a spatial decoding method, and the third decoding method is a decoding method other than the first decoding method and the second decoding method. method.
2. The method according to claim 1, wherein the spatial decoding method is a decoding method in which reconstruction is performed based on attribute information of the target virtual speaker, and information of the target virtual speaker is obtained by decoding the bitstream.
3. The method according to claim 1 or 2, wherein the third decoding method includes one or more decoding methods.
4. The third decoding method is a correlation-removal decoding method according to any one of claims 1 to 3, including a channel copy decoding method.
5. The method according to claim 4, wherein the channel copy decoding method is a correlation removal decoding method.
6. The step of decoding the bitstream based on the combination of decoding methods and obtaining the reconstructed scene audio signal is: The step of decoding one frame of the bitstream within the bitstream based on the multiple decoding methods in the combination of decoding methods to obtain one frame of the reconstructed scene audio signal. The method according to any one of claims 1 to 5, including
7. The bitstream includes X frames, each of which includes the i-th and j-th frames, where X is a positive integer, and the step of decoding the bitstream based on the combination of decoding schemes to obtain a reconstructed scene audio signal is: The steps include: decoding the i-th frame of the bitstream based on the plurality of decoding methods in the combination of decoding methods to obtain the i-th frame of the reconstructed scene audio signal; The steps include: decoding the j-th frame of the bitstream based on the first decoding method in the combination of decoding methods to obtain the j-th frame of the reconstructed scene audio signal; Includes, i and j are integers from 1 to X, and i is not equal to j. The method according to any one of claims 1 to 5.
8. The bitstream includes X frames, each of which includes the i-th and j-th frames, where X is a positive integer, and the step of decoding the bitstream based on the combination of decoding schemes to obtain a reconstructed scene audio signal is: The steps include decoding the i-th frame of the bitstream based on multiple decoding schemes within the k1-th decoding scheme combination and obtaining the i-th frame of the reconstructed scene audio signal, The steps include: decoding the j-th frame of the bitstream based on multiple decoding schemes within a combination of k2 decoding schemes to obtain the j-th frame of the reconstructed scene audio signal; Includes, i and j are integers between 1 and X, i is not equal to j, k1 is not equal to k2, and k1 and k2 are positive integers. The method according to any one of claims 1 to 5.
9. If one frame of the reconstructed scene audio signal includes a reconstructed audio signal having a C channel, and the combination of decoding schemes is a combination of the first decoding scheme, the second decoding scheme, and the third decoding scheme, the step of decoding the bitstream of the one frame based on the plurality of decoding schemes in the combination of decoding schemes to obtain the reconstructed scene audio signal of the one frame is: The steps involve obtaining a reconstructed scene audio signal for the frame by decoding one first channel based on the first decoding method, two second channels based on the second decoding method, and three third channels based on the third decoding method, based on the bitstream of the frame. Includes, C is equal to the sum of C1, C2, and C3, where C1, C2, and C3 are positive integers. The method according to claim 6.
10. The reconstructed scene audio signal of the aforementioned one frame includes a reconstructed audio signal having C channels, and each channel of the reconstructed audio signal has Y bandwidths, where C and Y are positive integers. The step of decoding the bitstream of one frame based on the plurality of decoding methods within the combination of the plurality of decoding methods and obtaining the reconstructed scene audio signal of the one frame is: The step of decoding the bitstream of the frame based on the plurality of decoding schemes in the combination of decoding schemes for Y bandwidths of one channel in the reconstructed scene audio signal of the frame, in order to obtain the reconstructed scene audio signal of the frame. including, The method according to claim 6.
11. The reconstructed scene audio signal is an Nth-order higher-order ambisonic HOA signal. The aforementioned C1 first channel is a channel C1 included in the zero- to M-th order signal of the N-th order HOA signal, where M is an integer less than N, C is equal to the square of (N+1), and C1 is less than or equal to the square of (M+1). The two second channels C include the other four channels C included in the zero-th to M-th order signals of the N-th order HOA signal, and the five channels C in the N-th order HOA signal other than the channels included in the zero-th to M-th order signals. The three third channels C include the other six channels C included in the zero-th to M-th order signals of the N-th order HOA signal, and the other seven channels C in the N-th order HOA signal other than the channels included in the zero-th to M-th order signals. C2 is equal to the sum of C4 and C5, C3 is equal to the sum of C6 and C7, the square of (M+1) is equal to the sum of C1, C4, and C6, and C4, C5, C6, and C7 are integers. The method according to claim 9.
12. The step of decoding the three third channels C based on the third decoding method is as follows: A step of decoding the three third channels C3 based on a third decoding scheme corresponding to a first preset identifier obtained by analyzing the bitstream, The third decoding method corresponding to the first preset identifier is a time-domain decorrelation decoding method or a frequency-domain decorrelation decoding method. Step The method according to claim 9, including the method described in claim 9.
13. If the combination of decoding methods is a combination of the first decoding method, the second decoding method, and the third decoding method, The step of correcting the reconstructed audio signal having the second channel of C2 in the reconstructed scene audio signal based on feature information corresponding to the two second channels of C2 obtained by analyzing the bitstream. The method according to claim 9 or 12, further comprising:
14. Steps to analyze a second preset identifier from the bitstream. Furthermore, The step of decoding the bitstream based on the combination of decoding methods and obtaining the reconstructed scene audio signal is: The step of decoding the bitstream based on a combination of decoding schemes corresponding to the second preset identifier to obtain the reconstructed scene audio signal. including, The method according to any one of claims 1 to 13.
15. A scene audio decoding device, A bitstream receiver module configured to receive a bitstream, A decoding module configured to decode the bitstream based on a combination of decoding schemes to obtain a reconstructed scene audio signal, Equipped with, The aforementioned combination of decoding methods includes at least one of the following combinations: a combination of the first decoding method, the second decoding method, and the third decoding method, and a combination of the first decoding method and the third decoding method. The first decoding method is a method for decoding encoded data obtained by encoding a signal, the second decoding method is a spatial decoding method, and the third decoding method is a decoding method other than the first decoding method and the second decoding method. Scene audio decoding device.
16. It is an electronic device, Memory and processor, wherein the memory is connected to the processor. Equipped with, The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device becomes capable of executing the scene audio decoding method according to any one of claims 1 to 14. electronic equipment.
17. A chip comprising one or more interface circuits and one or more processors, wherein the interface circuits are configured to receive signals from the memory of an electronic device and to transmit the signals to the processor, the signals include computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device is able to perform the scene audio decoding method according to any one of claims 1 to 14.
18. A computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed on a computer or processor, the computer or processor is able to execute the scene audio decoding method according to any one of claims 1 to 14.
19. A computer program product, wherein the computer program product includes a software program, and when the software program is executed by a computer or processor, a step of the method according to any one of claims 1 to 14 is performed.