Audio encoding and decoding method and device
By selecting the target virtual speaker and generating a virtual speaker signal for encoding, the problem of large amount of HOA signal data and difficult transmission is solved, efficient encoding and codec is achieved, and encoding quality is maintained.
Patent Information
- Application Number
- CN202011377320.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-11-30
AI Technical Summary
Advanced stereo reverb (HOA) technology has nothing to do with speaker layout in the recording, encoding and playback stages, but as the HOA order increases, the amount of data is large and transmission and storage is difficult, so HOA signals need to be encoded and coded. The existing multi-channel encoding and decoding methods require adapting the codec according to the number of channels, with large data volume and high bandwidth occupancy.
By selecting the target virtual speaker in the preset virtual speaker set, a virtual speaker signal is generated based on the current scene audio signal and the attribute information of the target virtual speaker and encoded it to reduce the amount of data.
The amount of data for encoding and decoding is reduced, the encoding and decoding efficiency is improved, and the encoding quality is maintained high, which can effectively represent the sound field of the listening person in the space.
Smart Images

Figure CN114582356B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio coding and decoding, and in particular to an audio coding and decoding method and device. Background Art
[0002] Three-dimensional audio technology is an audio technology that acquires, processes, transmits, renders and plays back sound events and three-dimensional sound field information in the real world. Three-dimensional audio technology gives sound a strong sense of space, envelopment and immersion, giving people an extraordinary auditory experience of "being there". Higher-order ambisonics (HOA) technology has the property of being independent of speaker layout during recording, encoding and playback, and the rotatable playback characteristics of HOA format data. It has higher flexibility in three-dimensional audio playback, and has therefore received more extensive attention and research.
[0003] In order to achieve better audio hearing effects, HOA technology requires a large amount of data to record more detailed sound scene information. Although this scene-based three-dimensional audio signal sampling and storage is more conducive to the preservation and transmission of audio signal spatial information, as the HOA order increases, more data will be generated. A large amount of data makes transmission and storage difficult, so HOA signals need to be encoded and decoded.
[0004] Currently, there is a method for encoding and decoding multi-channel data, including: at the encoding end, directly encoding each channel of the original scene audio signal through a core encoder (for example, a 16-channel encoder), and then outputting a bit stream. At the decoding end, decoding the bit stream through a core decoder (for example, a 16-channel decoder) to obtain each channel of the decoded scene audio signal.
[0005] The above multi-channel encoding and decoding method needs to adapt the corresponding codec according to the number of channels of the original scene audio signal, and as the number of channels increases, the compressed code stream has the problems of large data volume and high bandwidth occupancy. Summary of the invention
[0006] The embodiments of the present application provide an audio coding method and device for reducing the amount of coding and decoding data to improve coding and decoding efficiency.
[0007] To solve the above technical problems, the present application provides the following technical solutions:
[0008] In a first aspect, an embodiment of the present application provides an audio encoding method, including:
[0009] Selecting a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal;
[0010] Generate a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker;
[0011] The first virtual speaker signal is encoded to obtain a code stream.
[0012] In an embodiment of the present application, a first target virtual speaker is selected from a preset virtual speaker set according to the current scene audio signal; a first virtual speaker signal is generated according to the attribute information of the current scene audio signal and the first target virtual speaker; and the first virtual speaker signal is encoded to obtain a code stream. Since the first virtual speaker signal can be generated according to the attribute information of the first scene audio signal and the first target virtual speaker in the embodiment of the present application, the audio encoding end encodes the first virtual speaker signal instead of directly encoding the first scene audio signal. In the embodiment of the present application, the first target virtual speaker is selected according to the first scene audio signal. The first virtual speaker signal generated based on the first target virtual speaker can represent the sound field at the position of the listener in the space. The sound field at the position is as close as possible to the original sound field when the first scene audio signal was recorded, which ensures the encoding quality of the audio encoding end, and the first virtual speaker signal and the residual signal are encoded to obtain a code stream. The amount of encoded data of the first virtual speaker signal is related to the first target virtual speaker, but has nothing to do with the number of channels of the first scene audio signal, which reduces the amount of encoded data and improves the encoding efficiency.
[0013] In a possible implementation, the method further includes:
[0014] Acquire main sound field components from the current scene audio signal according to the virtual speaker set;
[0015] The selecting a first target virtual speaker from a preset virtual speaker set according to the current scene audio signal comprises:
[0016] The first target virtual speaker is selected from the virtual speaker set according to the main sound field component.
[0017] In the above scheme, each virtual speaker in the virtual speaker set corresponds to a sound field component, and the first target virtual speaker is selected from the virtual speaker set according to the main sound field component. For example, the virtual speaker corresponding to the main sound field component is the first target virtual speaker selected by the encoding end. In the embodiment of the present application, the encoding end can select the first target virtual speaker according to the main sound field component, which solves the problem that the encoding end needs to determine the first target virtual speaker.
[0018] In a possible implementation manner, the selecting the first target virtual speaker from the virtual speaker set according to the main sound field component includes:
[0019] Selecting, according to the main sound field component, an HOA coefficient corresponding to the main sound field component from a high-order ambisonic reverberation HOA coefficient set, wherein the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set;
[0020] A virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component in the virtual speaker set is determined as the first target virtual speaker.
[0021] In the above scheme, the encoding end pre-configures the HOA coefficient set according to the virtual speaker set, and there is a one-to-one correspondence between the HOA coefficients in the HOA coefficient set and the virtual speakers in the virtual speaker set. Therefore, after the HOA coefficients are selected according to the main sound field components, the target virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component is searched from the virtual speaker set according to the above one-to-one correspondence. The searched target virtual speaker is the first target virtual speaker, which solves the problem that the encoding end needs to determine the first target virtual speaker.
[0022] In a possible implementation manner, the selecting the first target virtual speaker from the virtual speaker set according to the main sound field component includes:
[0023] Acquire configuration parameters of the first target virtual speaker according to the main sound field components;
[0024] Generate an HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker;
[0025] Determine a virtual speaker corresponding to the HOA coefficient corresponding to the first target virtual speaker in the virtual speaker set as the target virtual speaker.
[0026] In the above scheme, after the encoding end obtains the main sound field component, the configuration parameters of the first target virtual speaker can be determined based on the main sound field component. For example, the main sound field component is one or several sound field components with the largest value among multiple sound field components, or the main sound field component can be one or several sound field components with dominant direction among multiple sound field components. The main sound field component can be used to determine the first target virtual speaker that matches the current scene audio signal. The first target virtual speaker is configured with corresponding attribute information. The configuration parameters of the first target virtual speaker can be used to generate the HOA coefficient of the first target virtual speaker. The generation process of the HOA coefficient can be achieved by the HOA algorithm, which is not described in detail here. Each virtual speaker in the virtual speaker set corresponds to an HOA coefficient, so the first target virtual speaker can be selected from the virtual speaker set according to the HOA coefficient corresponding to each virtual speaker, which solves the problem that the encoding end needs to determine the first target virtual speaker.
[0027] In a possible implementation manner, acquiring configuration parameters of the first target virtual speaker according to the main sound field component includes:
[0028] Determining configuration parameters of multiple virtual speakers in the virtual speaker set according to configuration information of the audio encoder;
[0029] The configuration parameters of the first target virtual speaker are selected from the configuration parameters of the plurality of virtual speakers according to the main sound field components.
[0030] In the above scheme, the configuration parameters of multiple virtual speakers can be pre-stored in the audio encoder, and the configuration parameters of each virtual speaker can be determined by the configuration information of the audio encoder. The audio encoder refers to the aforementioned encoding end, and the configuration information of the audio encoder includes but is not limited to: HOA order, encoding bit rate, etc. The configuration information of the audio encoder can be used to determine the number of virtual speakers and the position parameters of each virtual speaker, which solves the problem that the encoding end needs to determine the configuration parameters of the virtual speakers. An example is as follows. If the encoding bit rate is low, a smaller number of virtual speakers can be configured, and if the encoding bit rate is high, a plurality of virtual speakers can be configured. For example, the HOA order of the virtual speaker can be equal to the HOA order of the audio encoder. It is not limited that in the embodiment of the present application, in addition to determining the configuration parameters of multiple virtual speakers by the configuration information of the audio encoder, the configuration parameters of multiple virtual speakers can also be determined according to user-defined information. For example, the user can customize the position of the virtual speaker, the HOA order, the number of virtual speakers, etc.
[0031] In a possible implementation, the configuration parameters of the first target virtual speaker include: position information and HOA order information of the first target virtual speaker;
[0032] The generating the HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker includes:
[0033] An HOA coefficient corresponding to the first target virtual speaker is determined according to the position information and the HOA order information of the first target virtual speaker.
[0034] In the above scheme, the HOA coefficient of each virtual speaker can be generated using the position information and HOA order information of the virtual speaker. The generation process of the HOA coefficient can be implemented by the HOA algorithm, which solves the problem that the encoding end needs to determine the HOA coefficient of the first target virtual speaker.
[0035] In a possible implementation, the method further includes:
[0036] The attribute information of the first target virtual speaker is encoded and written into the bitstream.
[0037] In the above scheme, in addition to encoding the virtual speaker, the encoding end may also encode the attribute information of the first target virtual speaker, and write the encoded attribute information of the first target virtual speaker into the bitstream, and the obtained bitstream may include: the encoded virtual speaker and the encoded attribute information of the first target virtual speaker. In the embodiment of the present application, the bitstream may carry the encoded attribute information of the first target virtual speaker, so that the decoding end can determine the attribute information of the first target virtual speaker by decoding the bitstream, which is convenient for audio decoding at the decoding end.
[0038] In a possible implementation, the current scene audio signal includes: a high-order ambisonic reverberation HOA signal to be encoded; the attribute information of the first target virtual speaker includes the HOA coefficient of the first target virtual speaker;
[0039] The generating a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker comprises:
[0040] The HOA signal to be encoded and the HOA coefficients are linearly combined to obtain the first virtual speaker signal.
[0041] In the above scheme, taking the current scene audio signal as the HOA signal to be encoded as an example, the encoding end first determines the HOA coefficient of the first target virtual speaker. For example, the encoding end selects the HOA coefficient from the HOA coefficient set according to the main sound field components. The selected HOA coefficient is the HOA coefficient of the first target virtual speaker. After the encoding end obtains the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, the first virtual speaker signal can be generated according to the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, wherein the HOA signal to be encoded can be obtained by linearly combining the HOA coefficients of the first target virtual speaker, and the solution of the first virtual speaker signal can be converted into a problem of solving the linear combination.
[0042] In a possible implementation, the current scene audio signal includes: a high-order ambisonic reverberation (HOA) signal to be encoded; the attribute information of the first target virtual speaker includes position information of the first target virtual speaker;
[0043] The generating a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker comprises:
[0044] Acquire the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker;
[0045] The HOA signal to be encoded and the HOA coefficients are linearly combined to obtain the first virtual speaker signal.
[0046] In the above scheme, the attribute information of the first target virtual speaker may include: the position information of the first target virtual speaker, the encoder pre-stores the HOA coefficient of each virtual speaker in the virtual speaker set, and the encoder also stores the position information of each virtual speaker. There is a corresponding relationship between the position information of the virtual speaker and the HOA coefficient of the virtual speaker, so the encoder can determine the HOA coefficient of the first target virtual speaker through the position information of the first target virtual speaker. If the attribute information includes the HOA coefficient, the encoder can obtain the HOA coefficient of the first target virtual speaker by decoding the attribute information of the first target virtual speaker.
[0047] In a possible implementation, the method further includes:
[0048] Selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0049] Generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0050] The second virtual speaker signal is encoded and written into the bit stream.
[0051] In the above scheme, the second target virtual speaker is another target virtual speaker selected by the encoding end that is different from the first target virtual encoder. The first scene audio signal is the original scene audio signal to be encoded, and the second target virtual speaker can be a virtual speaker in the virtual speaker set. For example, a pre-configured target virtual speaker selection strategy can be used to select the second target virtual speaker from the preset virtual speaker set. The target virtual speaker selection strategy is a strategy for selecting a target virtual speaker that matches the first scene audio signal from the virtual speaker set, for example, selecting the second target virtual speaker according to the sound field components obtained by each virtual speaker from the first scene audio signal.
[0052] In a possible implementation, the method further includes:
[0053] Performing alignment processing on the first virtual loudspeaker signal and the second virtual loudspeaker signal to obtain an aligned first virtual loudspeaker signal and an aligned second virtual loudspeaker signal;
[0054] Accordingly, encoding the second virtual loudspeaker signal comprises:
[0055] encoding the aligned second virtual loudspeaker signal;
[0056] Accordingly, encoding the first virtual loudspeaker signal includes:
[0057] The aligned first virtual loudspeaker signal is encoded.
[0058] In the above scheme, after the encoding end obtains the aligned first virtual speaker signal, the aligned first virtual speaker signal can be encoded. In the embodiment of the present application, the correlation between channels is enhanced by realigning the channels of the first virtual speaker signal, which is beneficial to the encoding processing of the first virtual speaker signal by the core encoder.
[0059] In a possible implementation, the method further includes:
[0060] Selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0061] Generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0062] Accordingly, encoding the first virtual loudspeaker signal includes:
[0063] obtaining a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, wherein the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal;
[0064] The downmix signal and the side information are encoded.
[0065] In the above scheme, after the encoding end obtains the first virtual speaker signal and the second virtual speaker signal, the encoding end may also perform downmix processing according to the first virtual speaker signal and the second virtual speaker signal to generate a downmix signal, for example, downmixing the first virtual speaker signal and the second virtual speaker signal in amplitude to obtain a downmix signal. In addition, side information may be generated according to the first virtual speaker signal and the second virtual speaker signal, the side information being used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal, the relationship having multiple implementation methods, and the side information may be used by the decoding end to perform upmixing on the downmix signal to restore the first virtual speaker signal and the second virtual speaker signal. For example, the side information includes a signal information loss analysis parameter, so that the decoding end restores the first virtual speaker signal and the second virtual speaker signal through the signal information loss analysis parameter.
[0066] In a possible implementation, the method further includes:
[0067] Performing alignment processing on the first virtual loudspeaker signal and the second virtual loudspeaker signal to obtain an aligned first virtual loudspeaker signal and an aligned second virtual loudspeaker signal;
[0068] Correspondingly, obtaining the downmix signal and the side information according to the first virtual speaker signal and the second virtual speaker signal includes:
[0069] Obtain the downmix signal and the side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal;
[0070] Correspondingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
[0071] In the above scheme, before generating the downmix signal, the encoder can first perform an alignment operation on the virtual speaker signal, and after completing the alignment operation, generate the downmix signal and the side information. In the embodiment of the present application, by realigning the first virtual speaker signal and the channels of the second virtual speaker, the correlation between the channels is enhanced, which is beneficial to the encoding processing of the first virtual speaker signal by the core encoder.
[0072] In a possible implementation manner, before selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal, the method further includes:
[0073] Determine whether it is necessary to obtain a target virtual speaker other than the first target virtual speaker according to the encoding rate and / or the signal type information of the current scene audio signal;
[0074] If it is necessary to obtain a target virtual speaker other than the first target virtual speaker, a second target virtual speaker is selected from the virtual speaker set according to the current scene audio signal.
[0075] In the above scheme, the encoding end can also perform signal selection to determine whether it is necessary to obtain the second target virtual speaker. In the case where the second target virtual speaker needs to be obtained, the encoding end can generate the second virtual speaker signal. In the case where the second target virtual speaker does not need to be obtained, the encoding end may not generate the second virtual speaker signal. Among them, the encoder can make a decision based on the configuration information of the audio encoder and / or the signal type information of the first scene audio signal to determine whether it is necessary to select another target virtual speaker in addition to selecting the first target virtual speaker. For example, if the encoding rate is higher than the preset threshold, it is determined that the target virtual speakers corresponding to the two main sound field components need to be obtained. In addition to determining the first target virtual speaker, the second target virtual speaker can also be determined. For another example, according to the signal type information of the first scene audio signal, it is determined that the target virtual speakers corresponding to the two main sound field components with dominant sound source directions need to be obtained. In addition to determining the first target virtual speaker, the second target virtual speaker can also be determined. On the contrary, if it is determined that only one target virtual speaker needs to be obtained according to the encoding rate and / or the signal type information of the first scene audio signal, after determining the first target virtual speaker, it is determined that no target virtual speakers other than the first target virtual speaker will be obtained. In the embodiment of the present application, signal selection can be used to reduce the amount of data encoded by the encoding end and improve the encoding efficiency.
[0076] In a second aspect, an embodiment of the present application further provides an audio decoding method, comprising:
[0077] Receive code stream;
[0078] Decoding the bit stream to obtain a virtual speaker signal;
[0079] A reconstructed scene audio signal is obtained according to the property information of the target virtual speaker and the virtual speaker signal.
[0080] In an embodiment of the present application, a code stream is first received, then the code stream is decoded to obtain a virtual speaker signal, and finally a reconstructed scene audio signal is obtained according to the property information of the target virtual speaker and the virtual speaker signal. In an embodiment of the present application, a virtual speaker signal can be decoded from the code stream, and a reconstructed scene audio signal is obtained by using the property information of the target virtual speaker and the virtual speaker signal. In an embodiment of the present application, the obtained code stream carries a virtual speaker signal and a residual signal, which reduces the amount of decoded data and improves decoding efficiency.
[0081] In a possible implementation, the method further includes:
[0082] The code stream is decoded to obtain the attribute information of the target virtual speaker.
[0083] In the above scheme, in addition to encoding the virtual speaker, the encoding end may also encode the attribute information of the target virtual speaker, and write the encoded attribute information of the target virtual speaker into the bitstream, for example, the attribute information of the first target virtual speaker may be obtained through the bitstream. In the embodiment of the present application, the bitstream may carry the encoded attribute information of the first target virtual speaker, so that the decoding end may determine the attribute information of the first target virtual speaker by decoding the bitstream, which is convenient for audio decoding at the decoding end.
[0084] In a possible implementation, the attribute information of the target virtual speaker includes a high-order ambisonic reverberation HOA coefficient of the target virtual speaker;
[0085] The step of obtaining the reconstructed scene audio signal according to the property information of the target virtual speaker and the virtual speaker signal comprises:
[0086] The virtual speaker signal and the HOA coefficient of the target virtual speaker are synthesized to obtain the reconstructed scene audio signal.
[0087] In the above scheme, the decoder first determines the HOA coefficient of the target virtual speaker. For example, the decoder may pre-store the HOA coefficient of the target virtual speaker. After the decoder obtains the virtual speaker signal and the HOA coefficient of the target virtual speaker, the reconstructed scene audio signal may be obtained according to the virtual speaker signal and the HOA coefficient of the target virtual speaker. This improves the quality of the reconstructed scene audio signal.
[0088] In a possible implementation manner, the attribute information of the target virtual speaker includes position information of the target virtual speaker;
[0089] The step of obtaining the reconstructed scene audio signal according to the property information of the target virtual speaker and the virtual speaker signal comprises:
[0090] Determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker;
[0091] The virtual speaker signal and the HOA coefficient of the target virtual speaker are synthesized to obtain the reconstructed scene audio signal.
[0092] In the above scheme, the attribute information of the target virtual speaker may include: the position information of the target virtual speaker. The decoding end pre-stores the HOA coefficient of each virtual speaker in the virtual speaker set, and the decoding end also stores the position information of each virtual speaker. For example, the decoding end can determine the HOA coefficient corresponding to the position information of the target virtual speaker based on the correspondence between the position information of the virtual speaker and the HOA coefficient of the virtual speaker, or the decoding end can calculate the HOA coefficient of the target virtual speaker based on the position information of the target virtual speaker. Therefore, the decoding end can determine the HOA coefficient of the target virtual speaker through the position information of the target virtual speaker. The problem that the decoding end needs to determine the HOA coefficient of the target virtual speaker is solved.
[0093] In a possible implementation manner, the virtual loudspeaker signal is a downmixed signal obtained by downmixing the first virtual loudspeaker signal and the second virtual loudspeaker signal, and the method further includes:
[0094] Decoding the bitstream to obtain side information, where the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal;
[0095] Obtain the first virtual speaker signal and the second virtual speaker signal according to the side information and the downmix signal;
[0096] Accordingly, obtaining the reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal includes:
[0097] The reconstructed scene audio signal is obtained according to the property information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
[0098] In the above scheme, the encoder generates a downmix signal when performing downmix processing based on the first virtual speaker signal and the second virtual speaker signal. The encoder can also perform signal compensation on the downmix signal to generate side information, which can be written into the bit stream. The decoder can obtain the side information through the bit stream. The decoder can perform signal compensation based on the side information to obtain the first virtual speaker signal and the second virtual speaker signal. Therefore, when reconstructing the signal, the first virtual speaker signal and the second virtual speaker signal, as well as the aforementioned attribute information of the target virtual speaker, can be used, thereby improving the quality of the decoded signal at the decoder.
[0099] In a third aspect, an embodiment of the present application provides an audio encoding device, including:
[0100] An acquisition module, configured to select a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal;
[0101] A signal generating module, configured to generate a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker;
[0102] The encoding module is used to encode the first virtual speaker signal to obtain a code stream.
[0103] In a possible implementation, the acquisition module is configured to acquire a main sound field component from the current scene audio signal according to the virtual speaker set; and select the first target virtual speaker from the virtual speaker set according to the main sound field component.
[0104] In the third aspect of the present application, the constituent modules of the audio encoding device may also execute the steps described in the aforementioned first aspect and various possible implementations. For details, please refer to the aforementioned description of the first aspect and various possible implementations.
[0105] In a possible implementation, the acquisition module is used to select HOA coefficients corresponding to the main sound field components from the high-order stereo reverberation HOA coefficient set according to the main sound field components, and the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set; and determine that the virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component in the virtual speaker set is the first target virtual speaker.
[0106] In a possible implementation, the acquisition module is used to acquire the configuration parameters of the first target virtual speaker according to the main sound field components; generate the HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker; and determine that the virtual speaker corresponding to the HOA coefficient corresponding to the first target virtual speaker in the virtual speaker set is the target virtual speaker.
[0107] In a possible implementation, the acquisition module is used to determine configuration parameters of multiple virtual speakers in the virtual speaker set according to configuration information of the audio encoder; and select configuration parameters of the first target virtual speaker from the configuration parameters of the multiple virtual speakers according to the main sound field components.
[0108] In a possible implementation, the configuration parameters of the first target virtual speaker include: position information and HOA order information of the first target virtual speaker;
[0109] The acquisition module is used to determine the HOA coefficient corresponding to the first target virtual speaker according to the position information and HOA order information of the first target virtual speaker.
[0110] In a possible implementation manner, the encoding module is further configured to encode the attribute information of the first target virtual speaker and write the attribute information into the bit stream.
[0111] In a possible implementation, the current scene audio signal includes: a to-be-encoded HOA signal; the attribute information of the first target virtual speaker includes an HOA coefficient of the first target virtual speaker;
[0112] The signal generating module is used to linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
[0113] In a possible implementation, the current scene audio signal includes: a high-order ambisonic reverberation (HOA) signal to be encoded; the attribute information of the first target virtual speaker includes position information of the first target virtual speaker;
[0114] The signal generating module is used to obtain the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker; and linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
[0115] In a possible implementation, the acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0116] The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0117] The encoding module is used to encode the second virtual speaker signal and write it into the code stream.
[0118] In a possible implementation, the signal generating module is configured to perform alignment processing on the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal;
[0119] Correspondingly, the encoding module is used to encode the aligned second virtual loudspeaker signal;
[0120] Correspondingly, the encoding module is used to encode the aligned first virtual loudspeaker signal.
[0121] In a possible implementation, the acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0122] The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0123] Correspondingly, the encoding module is used to obtain a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, wherein the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal; and encode the downmix signal and the side information.
[0124] In a possible implementation, the signal generating module is configured to perform alignment processing on the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal;
[0125] Correspondingly, the encoding module is used to obtain the downmix signal and the side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal;
[0126] Correspondingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
[0127] In one possible implementation, the acquisition module is used to determine whether it is necessary to acquire a target virtual speaker other than the first target virtual speaker based on the encoding rate and / or the signal type information of the current scene audio signal before selecting the second target virtual speaker from the virtual speaker set based on the current scene audio signal; if it is necessary to acquire a target virtual speaker other than the first target virtual speaker, the second target virtual speaker is selected from the virtual speaker set based on the current scene audio signal.
[0128] In a fourth aspect, an embodiment of the present application provides an audio decoding device, including:
[0129] A receiving module, used for receiving a code stream;
[0130] A decoding module, used for decoding the code stream to obtain a virtual speaker signal;
[0131] The reconstruction module is used to obtain a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal.
[0132] In a possible implementation manner, the decoding module is further configured to decode the bit stream to obtain property information of the target virtual speaker.
[0133] In a possible implementation, the attribute information of the target virtual speaker includes a high-order ambisonic reverberation HOA coefficient of the target virtual speaker;
[0134] The reconstruction module is used to perform synthesis processing on the virtual speaker signal and the HOA coefficient of the target virtual speaker to obtain the reconstructed scene audio signal.
[0135] In a possible implementation manner, the attribute information of the target virtual speaker includes position information of the target virtual speaker;
[0136] The reconstruction module is used to determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker; and synthesize the virtual speaker signal and the HOA coefficient of the target virtual speaker to obtain the reconstructed scene audio signal.
[0137] In a possible implementation, the virtual speaker signal is a downmixed signal obtained by downmixing the first virtual speaker signal and the second virtual speaker signal, and the device further includes: a signal compensation module, wherein:
[0138] The decoding module is used to decode the bit stream to obtain side information, where the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal;
[0139] The signal compensation module is configured to obtain the first virtual speaker signal and the second virtual speaker signal according to the side information and the downmix signal;
[0140] Correspondingly, the reconstruction module is used to obtain the reconstructed scene audio signal according to the attribute information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
[0141] In the fourth aspect of the present application, the constituent modules of the audio decoding device may also execute the steps described in the aforementioned second aspect and various possible implementations. For details, please refer to the aforementioned description of the second aspect and various possible implementations.
[0142] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect or the second aspect above.
[0143] In a sixth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect or the second aspect above.
[0144] In the seventh aspect, an embodiment of the present application provides a communication device, which may include entities such as a terminal device or a chip, and the communication device includes: a processor, and optionally, the communication device also includes a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, so that the communication device performs a method as described in any one of the first or second aspects above.
[0145] In an eighth aspect, the present application provides a chip system, which includes a processor for supporting an audio encoding device or an audio decoding device to implement the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the audio encoding device or the audio decoding device. The chip system can be composed of a chip, or it can include a chip and other discrete devices.
[0146] In a ninth aspect, the present application provides a computer-readable storage medium, comprising a code stream generated by the method described in any one of the first aspects above. BRIEF DESCRIPTION OF THE DRAWINGS
[0147] Figure 1 A schematic diagram of the structure of the audio processing system provided in the embodiment of the present application;
[0148] Figure 2a A schematic diagram of an audio encoder and an audio decoder provided in an embodiment of the present application being applied to a terminal device;
[0149] Figure 2b A schematic diagram of an audio encoder provided in an embodiment of the present application being applied to a wireless device or a core network device;
[0150] Figure 2c A schematic diagram of an audio decoder provided in an embodiment of the present application being applied to a wireless device or a core network device;
[0151] Figure 3a A schematic diagram of a multi-channel encoder and a multi-channel decoder provided in an embodiment of the present application applied to a terminal device;
[0152] Figure 3b A schematic diagram of a multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device;
[0153] Figure 3c A schematic diagram of a multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device;
[0154] Figure 4 A schematic diagram of an interaction process between an audio encoding device and an audio decoding device in an embodiment of the present application;
[0155] Figure 5 A schematic diagram of the structure of the encoding end provided in an embodiment of the present application;
[0156] Figure 6 A schematic diagram of the structure of a decoding end provided in an embodiment of the present application;
[0157] Figure 7 A schematic diagram of the structure of the encoding end provided in an embodiment of the present application;
[0158] Figure 8 A schematic diagram of virtual speakers approximately evenly distributed on a sphere provided in an embodiment of the present application;
[0159] Fig. 9 A schematic diagram of the structure of the encoding end provided in an embodiment of the present application;
[0160] Fig.10 A schematic diagram of the structure of an audio encoding device provided in an embodiment of the present application;
[0161] Fig.11 A schematic diagram of the structure of an audio decoding device provided in an embodiment of the present application;
[0162] Fig.12 A schematic diagram of the structure of another audio encoding device provided in an embodiment of the present application;
[0163] Fig.13 A schematic diagram of the composition structure of another audio decoding device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0164] The embodiments of the present application provide an audio coding and decoding method and device for reducing the data volume of a coded scene audio signal and improving coding and decoding efficiency.
[0165] The embodiments of the present application are described below in conjunction with the accompanying drawings.
[0166] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0167] The technical solution of the embodiment of the present application can be applied to various audio processing systems, such as Figure 1 As shown, it is a schematic diagram of the composition structure of the audio processing system provided in an embodiment of the present application. The audio processing system 100 may include: an audio encoding device 101 and an audio decoding device 102. Among them, the audio encoding device 101 can be used to generate a code stream, and then the audio encoding code stream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the code stream and then perform the audio decoding function of the audio decoding device 102 to finally obtain a reconstructed signal.
[0168] In an embodiment of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio encoding device can be an audio encoder of the above-mentioned terminal device or wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio decoding device can be an audio decoder of the above-mentioned terminal device or wireless device or core network device. For example, the audio encoder can include a wireless access network, a media gateway of the core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc. The audio encoder can also be an audio codec used in virtual reality (VR) streaming media services.
[0169] In the application embodiment, taking the audio encoding and decoding module (audio encoding and audio decoding) suitable for virtual reality streaming media (VR streaming) service as an example, the end-to-end processing flow of the audio signal includes: the audio signal A passes through the acquisition module (acquisition) and then performs a preprocessing operation (audio preprocessing), the preprocessing operation includes filtering out the low-frequency part of the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the orientation information in the signal, and then performing encoding processing (audio encoding) and packaging (file / segment encapsulation) and sending (delivery) to the decoding end, the decoding end first unpacks (file / segment decapsulation), and then decodes (audio decoding), and performs binaural rendering (audio rendering) on the decoded signal. The rendered signal is mapped to the listener's headphones (headphones), which can be independent headphones or headphones on eyewear devices.
[0170] like Figure 2a As shown, it is a schematic diagram of the audio encoder and audio decoder provided by the embodiment of the present application being applied to a terminal device. For each terminal device, it can include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used to perform channel encoding on the audio signal, and the channel decoder is used to perform channel decoding on the audio signal. For example, in the first terminal device 20, it can include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. In the second terminal device 21, it can include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and the wireless or wired second network communication device 23 are connected through a digital channel, and the second terminal device 21 is connected to the wireless or wired second network communication device 23. Among them, the above-mentioned wireless or wired network communication device can generally refer to a signal transmission device, such as a communication base station, a data exchange device, etc.
[0171] In audio communication, the terminal device at the sending end first collects audio, encodes the collected audio signal, and then transmits it in a digital channel through a wireless network or core network after channel encoding. The terminal device at the receiving end performs channel decoding based on the received signal to obtain a bit stream, and then restores the audio signal through audio decoding, which is then played back by the terminal device at the receiving end.
[0172] like Figure 2b As shown, it is a schematic diagram of the audio encoder provided in the embodiment of the present application applied to a wireless device or a core network device. Among them, the wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, an audio encoder 253 provided in the embodiment of the present application, and a channel encoder 254, wherein other audio decoders 252 refer to other audio decoders other than the audio decoder. In the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, and then the other audio decoder 252 is used to perform audio decoding, and then the audio encoder 253 provided in the embodiment of the present application is used to perform audio encoding, and finally the channel encoder 254 is used to perform channel encoding on the audio signal, and then the channel encoding is completed before it is transmitted. Among them, the other audio decoder 252 performs audio decoding on the code stream decoded by the channel decoder 251.
[0173] like Figure 2c As shown, it is a schematic diagram of the audio decoder provided in the embodiment of the present application being applied to a wireless device or a core network device. Among them, the wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in the embodiment of the present application, other audio encoders 256, and a channel encoder 254, wherein other audio encoders 256 refer to other audio encoders other than the audio encoder. In the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, and then the received audio coding stream is decoded using the audio decoder 255, and then the other audio encoder 256 is used for audio encoding, and finally the channel encoder 254 is used to channel-encode the audio signal, and then the channel encoding is completed before it is transmitted. In the wireless device or core network device, if transcoding needs to be implemented, the corresponding audio codec processing needs to be performed. Among them, the wireless device refers to the radio frequency-related device in the communication, and the core network device refers to the core network-related device in the communication.
[0174] In some embodiments of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio encoding device can be a multi-channel encoder of the above terminal device or wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio decoding device can be a multi-channel decoder of the above terminal device or wireless device or core network device.
[0175] like Figure 3a As shown, it is a schematic diagram of a multi-channel encoder and a multi-channel decoder provided in an embodiment of the present application applied to a terminal device, and each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in an embodiment of the present application, and the multi-channel decoder can execute the audio decoding method provided in an embodiment of the present application. Specifically, the channel encoder is used to perform channel encoding on a multi-channel signal, and the channel decoder is used to perform channel decoding on a multi-channel signal. For example, in the first terminal device 30, it may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. In the second terminal device 31, it may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and the wireless or wired second network communication device 33 are connected through a digital channel, and the second terminal device 31 is connected to a wireless or wired second network communication device 33. The wireless or wired network communication equipment mentioned above may generally refer to signal transmission equipment, such as communication base stations, data exchange equipment, etc. In audio communication, the terminal device as the transmitter performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and then transmits it in a digital channel through a wireless network or a core network. The terminal device as the receiving end performs channel decoding according to the received signal to obtain a multi-channel signal encoding stream, and then recovers the multi-channel signal through multi-channel decoding, and then plays it back by the terminal device as the receiving end.
[0176] like Figure 3b As shown, it is a schematic diagram of a multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or the core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, which are the same as the aforementioned Figure 2b Similar, no further description is given here.
[0177] like Figure 3cAs shown, it is a schematic diagram of a multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or the core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, which are the same as the aforementioned Figure 2c Similar, no further description is given here.
[0178] Among them, the audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the collected multi-channel signal can be to obtain an audio signal after processing the collected multi-channel signal, and then encode the obtained audio signal according to the method provided in the embodiment of the present application; the decoding end encodes the code stream according to the multi-channel signal, decodes the audio signal, and restores the multi-channel signal after upmixing. Therefore, the embodiment of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding needs to be implemented, corresponding multi-channel encoding and decoding processing is required.
[0179] The audio encoding and decoding method provided in the embodiment of the present application may include: an audio encoding method and an audio decoding method, wherein the audio encoding method is executed by an audio encoding device, and the audio decoding method is executed by an audio decoding device, and the audio encoding device and the audio decoding device can communicate with each other. Next, based on the aforementioned system architecture and the audio encoding device and the audio decoding device, the audio encoding method and the audio decoding method provided in the embodiment of the present application are described. Figure 4 As shown, it is a schematic diagram of an interaction process between the audio encoding device and the audio decoding device in an embodiment of the present application, wherein the following steps 401 to 403 can be performed by the audio encoding device (hereinafter referred to as the encoding end), and the following steps 411 to 413 can be performed by the audio decoding device (hereinafter referred to as the decoding end), which mainly includes the following process:
[0180] 401. Select a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal.
[0181] The encoder obtains the current scene audio signal, which refers to an audio signal obtained by collecting the sound field at the position of the microphone in the space, and the current scene audio signal can also be called the original scene audio signal. For example, the current scene audio signal can be an audio signal obtained by higher order ambisonics (HOA) technology.
[0182] In an embodiment of the present application, the encoding end can pre-configure a virtual speaker set, which may include multiple virtual speakers. When the scene audio signal is actually played back, it can be played back through headphones or through multiple speakers arranged in the room. When using speakers for playback, the basic method is to superimpose the signals of multiple speakers so that the sound field at a certain point in the space (the location of the listener) is as close as possible to the original sound field when the scene audio signal was recorded under a certain standard. In an embodiment of the present application, a virtual speaker is used to calculate the playback signal corresponding to the scene audio signal, and the playback signal is used as a transmission signal to generate a compressed signal. The virtual speaker represents a speaker that exists virtually in the spatial sound field, and the virtual speaker can realize the playback of the scene audio signal at the encoding end.
[0183] In an embodiment of the present application, the virtual speaker set includes multiple virtual speakers, and each of the multiple virtual speakers corresponds to a virtual speaker configuration parameter (referred to as configuration parameter). The virtual speaker configuration parameters include but are not limited to: the number of virtual speakers, the HOA order of the virtual speaker, the position coordinates of the virtual speaker and other information. After the encoding end obtains the above-mentioned virtual speaker set, the first target virtual speaker is selected from the preset virtual speaker set according to the current scene audio signal. The current scene audio signal is the original scene audio signal to be encoded. The first target virtual speaker can be a virtual speaker in the virtual speaker set. For example, the first target virtual speaker can be selected from the preset virtual speaker set using a pre-configured target virtual speaker selection strategy. The target virtual speaker selection strategy is a strategy for selecting a target virtual speaker that matches the current scene audio signal from the virtual speaker set, for example, according to the sound field components obtained by each virtual speaker from the current scene audio signal to select the first target virtual speaker. For example, the first target virtual speaker is selected from the current scene audio signal according to the position information of each virtual speaker. The first target virtual speaker is a virtual speaker in the virtual speaker set used to play back the current scene audio signal, that is, the encoder can select a target virtual encoder that can play back the current scene audio signal from the virtual speaker set.
[0184] It is not limited that, in the embodiment of the present application, after the first target virtual speaker is selected through step 401, a subsequent processing process for the first target virtual speaker can be performed, such as subsequent steps 402 to 403. In the embodiment of the present application, not only the first target virtual speaker can be selected, but also more target virtual speakers can be selected, for example, a second target virtual speaker can be selected, and for the second target virtual speaker, a process similar to subsequent steps 402 to 403 also needs to be performed, see the description of the subsequent embodiment for details.
[0185] In an embodiment of the present application, after the encoding end selects the first target virtual speaker, the encoding end can also obtain the attribute information of the first target virtual speaker. The attribute information of the first target virtual speaker includes information related to the attributes of the first target virtual speaker. The attribute information can be set according to the specific application scenario. For example, the attribute information of the first target virtual speaker includes: the position information of the first target virtual speaker, or the HOA coefficient of the first target virtual speaker. Among them, the position information of the first target virtual speaker can be the distribution position of the first target virtual speaker in space, or it can be the information of the position of the first target virtual speaker relative to other virtual speakers in the virtual speaker set, which is not limited here. Each virtual speaker in the virtual speaker set corresponds to an HOA coefficient, which can also be called an Ambisonic coefficient. The HOA coefficient corresponding to the virtual speaker is explained below.
[0186] For example, the HOA order can be one of the orders from 2 to 10, the signal sampling rate when recording the audio signal is 48 to 192 kHz, the sampling depth is 16 or 24 bits, and the HOA signal can be generated by the HOA coefficient of the virtual speaker and the scene audio signal. The HOA signal is characterized by carrying the spatial information of the sound field. The HOA signal is information that describes the sound field signal of a certain point in the space with a certain accuracy. Therefore, it is possible to consider using another representation to describe the sound field signal of a certain position point. This description method can use less data to achieve the same accuracy of the signal of the spatial position point, thereby achieving the purpose of signal compression. The spatial sound field can be decomposed into the superposition of multiple plane waves. Therefore, in theory, the sound field expressed by the HOA signal can be expressed by reusing the superposition of multiple plane waves, and each plane wave is represented by an audio signal of one channel and a direction vector. The representation form of plane wave superposition can accurately express the original sound field using a smaller number of channels to achieve the purpose of signal compression.
[0187] In some embodiments of the present application, in addition to performing the aforementioned step 401, the audio encoding method provided by the embodiment of the present application further includes the following steps:
[0188] A1. Obtaining main sound field components from the current scene audio signal according to the virtual speaker set.
[0189] The main sound field component in step A1 may also be referred to as a first main sound field component.
[0190] In the scenario of executing step A1, the aforementioned step 401 selects a first target virtual speaker from a preset virtual speaker set according to the current scene audio signal, including:
[0191] B1. Select a first target virtual speaker from the virtual speaker set according to the main sound field components.
[0192] Among them, the encoding end obtains a virtual speaker set, and the encoding end uses the virtual speaker set to perform signal decomposition on the current scene audio signal to obtain the main sound field component corresponding to the current scene audio signal. Among them, the main sound field component represents the audio signal corresponding to the main sound field in the current scene audio signal. For example, the virtual speaker set includes multiple virtual speakers, and multiple sound field components can be obtained from the current scene audio signal according to the multiple virtual speakers, that is, each virtual speaker can obtain a sound field component from the current scene audio signal, and then the main sound field component is selected from the multiple sound field components. For example, the main sound field component can be one or several sound field components with the largest value among the multiple sound field components, or the main sound field component can be one or several sound field components with dominant directions among the multiple sound field components. Each virtual speaker in the virtual speaker set corresponds to a sound field component, and the first target virtual speaker is selected from the virtual speaker set according to the main sound field component. For example, the virtual speaker corresponding to the main sound field component is the first target virtual speaker selected by the encoding end. In the embodiment of the present application, the encoding end can select the first target virtual speaker through the main sound field component, which solves the problem that the encoding end needs to determine the first target virtual speaker.
[0193] It is not limited that, in an embodiment of the present application, the encoding end has multiple ways to select the first target virtual speaker. For example, the encoding end can preset a virtual speaker at a specified position as the first target virtual speaker, that is, according to the position of each virtual speaker in the virtual speaker set, a virtual speaker that meets the specified position is selected as the first target virtual speaker.
[0194] In some embodiments of the present application, the aforementioned step B1 selects a first target virtual speaker from a set of virtual speakers according to the main sound field components, including:
[0195] According to the main sound field components, HOA coefficients corresponding to the main sound field components are selected from the high-order ambisonic reverberation HOA coefficient set, and the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set;
[0196] The virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component in the virtual speaker set is determined as the first target virtual speaker.
[0197] Among them, the encoding end pre-configures the HOA coefficient set according to the virtual speaker set, and there is a one-to-one correspondence between the HOA coefficients in the HOA coefficient set and the virtual speakers in the virtual speaker set. Therefore, after the HOA coefficients are selected according to the main sound field components, the target virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component is searched from the virtual speaker set according to the above one-to-one correspondence. The target virtual speaker found is the first target virtual speaker, which solves the problem that the encoding end needs to determine the first target virtual speaker. An example is as follows: the HOA coefficient set includes HOA coefficient 1, HOA coefficient 2, and HOA coefficient 3, and the virtual speaker set includes virtual speaker 1, virtual speaker 2, and virtual speaker 3, wherein the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set, for example: HOA coefficient 1 corresponds to virtual speaker 1, HOA coefficient 2 corresponds to virtual speaker 2, and HOA coefficient 3 corresponds to virtual speaker 3. If HOA coefficient 3 is selected from the HOA coefficient set according to the main sound field component, it can be determined that the first target virtual speaker is virtual speaker 3.
[0198] In some embodiments of the present application, the aforementioned step B1 selects a first target virtual speaker from a set of virtual speakers according to the main sound field components, and further includes:
[0199] C1. Obtain configuration parameters of the first target virtual speaker according to the main sound field components;
[0200] C2. Generate an HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker;
[0201] C3. Determine that the virtual speaker corresponding to the HOA coefficient corresponding to the first target virtual speaker in the virtual speaker set is the first target virtual speaker.
[0202] Among them, after the encoding end obtains the main sound field component, the configuration parameters of the first target virtual speaker can be determined based on the main sound field component. For example, the main sound field component is one or several sound field components with the largest value among multiple sound field components, or the main sound field component can be one or several sound field components with dominant direction among multiple sound field components. The main sound field component can be used to determine the first target virtual speaker that matches the current scene audio signal. The first target virtual speaker is configured with corresponding attribute information. The configuration parameters of the first target virtual speaker can be used to generate the HOA coefficient of the first target virtual speaker. The generation process of the HOA coefficient can be achieved by the HOA algorithm, which is not described in detail here. Each virtual speaker in the virtual speaker set corresponds to an HOA coefficient, so the first target virtual speaker can be selected from the virtual speaker set according to the HOA coefficient corresponding to each virtual speaker, which solves the problem that the encoding end needs to determine the first target virtual speaker.
[0203] In some embodiments of the present application, step C1 obtains configuration parameters of the first target virtual speaker according to the main sound field components, including:
[0204] Determine configuration parameters of multiple virtual speakers in the virtual speaker set according to configuration information of the audio encoder;
[0205] The configuration parameters of the first target virtual speaker are selected from the configuration parameters of the plurality of virtual speakers according to the main sound field components.
[0206] Among them, the configuration parameters of multiple virtual speakers can be pre-stored in the audio encoder, and the configuration parameters of each virtual speaker can be determined by the configuration information of the audio encoder. The audio encoder refers to the aforementioned encoding end, and the configuration information of the audio encoder includes but is not limited to: HOA order, encoding bit rate, etc. The configuration information of the audio encoder can be used to determine the number of virtual speakers and the position parameters of each virtual speaker, which solves the problem that the encoding end needs to determine the configuration parameters of the virtual speakers. An example is as follows. If the encoding bit rate is low, a smaller number of virtual speakers can be configured, and if the encoding bit rate is high, a plurality of virtual speakers can be configured. For example, the HOA order of the virtual speaker can be equal to the HOA order of the audio encoder. It is not limited that, in the embodiment of the present application, in addition to determining the configuration parameters of multiple virtual speakers by the configuration information of the audio encoder, the configuration parameters of multiple virtual speakers can also be determined according to user-defined information. For example, the user can customize the position of the virtual speaker, the HOA order, the number of virtual speakers, etc.
[0207] The encoder obtains the configuration parameters of multiple virtual speakers from the virtual speaker set. For each virtual speaker, there are corresponding virtual speaker configuration parameters. Each virtual speaker configuration parameter includes but is not limited to: the HOA order of the virtual speaker, the position coordinates of the virtual speaker and other information. The HOA coefficient of the virtual speaker can be generated using the configuration parameters of each virtual speaker. The generation process of the HOA coefficient can be achieved by the HOA algorithm, which is not described in detail here. An HOA coefficient is generated for each virtual speaker in the virtual speaker set. The HOA coefficients configured for all virtual speakers in the virtual speaker set constitute an HOA coefficient set, which solves the problem that the encoder needs to determine the HOA coefficient of each virtual speaker in the virtual speaker set.
[0208] Wherein, in some embodiments of the present application, the configuration parameters of the first target virtual speaker include: position information and HOA order information of the first target virtual speaker;
[0209] The aforementioned step C2 generates the HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker, including:
[0210] The HOA coefficient corresponding to the first target virtual speaker is determined according to the position information and the HOA order information of the first target virtual speaker.
[0211] Among them, the configuration parameters of each virtual speaker in the virtual speaker set may include the position information of the virtual speaker and the HOA order information of the virtual speaker. Similarly, the configuration parameters of the first target virtual speaker include: the position information and HOA order information of the first target virtual speaker. For example, the position information of each virtual speaker in the virtual speaker set can be determined according to the local equidistant virtual speaker spatial distribution method. The local equidistant virtual speaker spatial distribution method refers to the distribution of multiple virtual speakers in space in a locally equidistant manner. For example, local equidistant may include: uniform distribution or uneven distribution. The HOA coefficient of the virtual speaker can be generated using the position information and HOA order information of each virtual speaker. The generation process of the HOA coefficient can be implemented by the HOA algorithm, which solves the problem that the encoding end needs to determine the HOA coefficient of the first target virtual speaker.
[0212] In addition, in the embodiment of the present application, a set of HOA coefficients is generated for each virtual speaker in the virtual speaker set, and multiple sets of HOA coefficients constitute the aforementioned HOA coefficient set. The HOA coefficients configured for all virtual speakers in the virtual speaker set constitute the HOA coefficient set, which solves the problem that the encoding end needs to determine the HOA coefficient of each virtual speaker in the virtual speaker set.
[0213] 402. Generate a first virtual speaker signal according to the current scene audio signal and attribute information of the first target virtual speaker.
[0214] Among them, after the encoding end obtains the attribute information of the current scene audio signal and the first target virtual speaker, the encoding end can play back the current scene audio signal. The encoding end generates a first virtual speaker signal according to the attribute information of the current scene audio signal and the first target virtual speaker. The first virtual speaker signal is the playback signal of the current scene audio signal. The attribute information of the first target virtual speaker describes the information related to the attributes of the first target virtual speaker. The first target virtual speaker is a virtual speaker selected by the encoding end that can play back the current scene audio signal. Therefore, the current scene audio signal is played back through the attribute information of the first target virtual speaker to obtain the first virtual speaker signal. The data volume of the first virtual speaker signal has nothing to do with the number of channels of the current scene audio signal, and the data volume of the first virtual speaker signal is related to the first target virtual speaker. For example, in an embodiment of the present application, the first virtual speaker signal is represented by fewer channels than the current scene audio signal. For example, the current scene audio signal is a third-order HOA signal, and the HOA signal has 16 channels. In an embodiment of the present application, the 16 channels can be compressed into 2 channels, that is, the virtual speaker signal generated by the encoding end is 2 channels. For example, the virtual speaker signal generated by the encoding end may include the aforementioned first virtual speaker signal and the second virtual speaker signal, etc. The number of channels of the virtual speaker signal generated by the encoding end is independent of the number of channels of the first scene audio signal. It can be seen from the description of the subsequent steps that the first virtual speaker signal with 2 channels can be carried in the bitstream. Accordingly, the decoding end receives the bitstream, and the virtual speaker signal obtained by decoding the bitstream is 2 channels. The decoding end can reconstruct a 16-channel scene audio signal through the 2-channel virtual speaker signal, and ensures that the reconstructed scene audio signal has the same subjective and objective quality as the original scene audio signal.
[0215] It can be understood that the aforementioned steps 401 and 402 can be implemented by a spatial encoder such as a moving picture experts group (MPEG) spatial encoder.
[0216] In some embodiments of the present application, the current scene audio signal may include: an HOA signal to be encoded; the property information of the first target virtual speaker includes an HOA coefficient of the first target virtual speaker;
[0217] Step 402 generates a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker, including:
[0218] The HOA signal to be encoded and the HOA coefficient of the first target virtual speaker are linearly combined to obtain a first virtual speaker signal.
[0219] Among them, taking the current scene audio signal as the HOA signal to be encoded as an example, the encoding end first determines the HOA coefficient of the first target virtual speaker. For example, the encoding end selects the HOA coefficient from the HOA coefficient set according to the main sound field components. The selected HOA coefficient is the HOA coefficient of the first target virtual speaker. After the encoding end obtains the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, the first virtual speaker signal can be generated according to the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker. Among them, the HOA signal to be encoded can be obtained by linearly combining the HOA coefficients of the first target virtual speaker, and the solution of the first virtual speaker signal can be converted into a problem of solving the linear combination.
[0220] For example, the attribute information of the first target virtual speaker may include: the HOA coefficient of the first target virtual speaker. The encoding end can obtain the HOA coefficient of the first target virtual speaker by decoding the attribute information of the first target virtual speaker. The encoding end linearly combines the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, that is, the encoding end combines the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker together to obtain a linear combination matrix. Next, the encoding end can find the optimal solution for the linear combination matrix, and the optimal solution obtained is the first virtual speaker signal. Among them, the optimal solution is related to the algorithm used when solving the linear combination matrix. The embodiment of the present application solves the problem that the encoding end needs to generate a first virtual speaker signal.
[0221] In some embodiments of the present application, the current scene audio signal includes: a high-order ambisonic reverberation HOA signal to be encoded; the attribute information of the first target virtual speaker includes the position information of the first target virtual speaker;
[0222] Step 402 generates a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker, including:
[0223] Acquire the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker;
[0224] The HOA signal to be encoded and the HOA coefficients corresponding to the first target virtual speaker are linearly combined to obtain a first virtual speaker signal.
[0225] Among them, the attribute information of the first target virtual speaker may include: the position information of the first target virtual speaker, the encoder pre-stores the HOA coefficient of each virtual speaker in the virtual speaker set, and the encoder also stores the position information of each virtual speaker. There is a corresponding relationship between the position information of the virtual speaker and the HOA coefficient of the virtual speaker, so the encoder can determine the HOA coefficient of the first target virtual speaker through the position information of the first target virtual speaker. If the attribute information includes the HOA coefficient, the encoder can obtain the HOA coefficient of the first target virtual speaker by decoding the attribute information of the first target virtual speaker.
[0226] After the encoder obtains the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, the encoder linearly combines the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker, that is, the encoder combines the HOA signal to be encoded and the HOA coefficient of the first target virtual speaker to obtain a linear combination matrix. Next, the encoder can find the optimal solution for the linear combination matrix, and the optimal solution obtained is the first virtual speaker signal.
[0227] An example is given below. The HOA coefficient of the first target virtual speaker is represented by a matrix A. The matrix A can be used to linearly combine the HOA signal to be encoded. The least square method can be used to obtain the theoretical optimal solution w, which is the first virtual speaker signal. For example, the following calculation formula can be used:
[0228] w=A -1 X,
[0229] Among them, A -1 represents the inverse matrix of matrix A, the size of matrix A is (M×C), C is the number of first target virtual speakers, M is the number of channels of the HOA coefficient of order N, and a represents the HOA coefficient of the first target virtual speaker, for example,
[0230]
[0231] Where X represents the HOA signal to be encoded, the size of the matrix X is (M×L), M is the number of channels of the N-order HOA coefficient, L is the number of sampling points, and x represents the coefficient of the HOA signal to be encoded, for example,
[0232]
[0233] 403. Encode the virtual speaker signal to obtain a bit stream.
[0234] In an embodiment of the present application, after the encoding end generates the first virtual speaker signal, the encoding end may encode the first virtual speaker signal to obtain a code stream. For example, the encoding end may specifically be a core encoder, and the core encoder encodes the first virtual speaker signal to obtain a code stream. The code stream may also be referred to as an audio signal encoding code stream. In an embodiment of the present application, the encoding end encodes the first virtual speaker signal, and no longer encodes the scene audio signal. By selecting the first target virtual speaker, the sound field at the position of the listener in the space is as close as possible to the original sound field when the scene audio signal was recorded, thereby ensuring the encoding quality of the encoding end, and the amount of encoded data of the first virtual speaker signal is independent of the number of channels of the scene audio signal, thereby reducing the amount of data of the encoded scene audio signal and improving the encoding and decoding efficiency.
[0235] In some embodiments of the present application, after the encoding end performs the above steps 401 to 403, the audio encoding method provided in the embodiment of the present application further includes the following steps:
[0236] The attribute information of the first target virtual speaker is encoded and written into a bitstream.
[0237] In addition to encoding the virtual speaker, the encoding end may also encode the attribute information of the first target virtual speaker, and write the encoded attribute information of the first target virtual speaker into the bitstream. At this time, the obtained bitstream may include: the encoded virtual speaker and the encoded attribute information of the first target virtual speaker. In the embodiment of the present application, the bitstream may carry the encoded attribute information of the first target virtual speaker, so that the decoding end can determine the attribute information of the first target virtual speaker by decoding the bitstream, which is convenient for audio decoding at the decoding end.
[0238] It should be noted that the aforementioned steps 401 to 403 describe the process of generating a first virtual speaker signal based on the first target virtual speaker and encoding the signal according to the first virtual speaker when the first target speaker is selected from the virtual speaker set. It is not limited that, in the embodiment of the present application, the encoding end can not only select the first target virtual speaker, but also select more target virtual speakers, for example, the second target virtual speaker can also be selected. For the second target virtual speaker, it is also necessary to perform a process similar to the aforementioned steps 402 to 403, which will be described in detail below.
[0239] In some embodiments of the present application, in addition to performing the aforementioned steps, the encoding end may further include:
[0240] D1, selecting a second target virtual speaker from the virtual speaker set according to the first scene audio signal;
[0241] D2. generating a second virtual speaker signal according to the first scene audio signal and the attribute information of the second target virtual speaker;
[0242] D3. Encode the second virtual speaker signal and write it into a bit stream.
[0243] Among them, the implementation method of step D1 is similar to the aforementioned step 401, and the second target virtual speaker is another target virtual speaker selected by the encoding end that is different from the first target virtual encoder. The first scene audio signal is the original scene audio signal to be encoded, and the second target virtual speaker can be a virtual speaker in the virtual speaker set. For example, a pre-configured target virtual speaker selection strategy can be used to select the second target virtual speaker from the preset virtual speaker set. The target virtual speaker selection strategy is a strategy for selecting a target virtual speaker that matches the first scene audio signal from the virtual speaker set, for example, selecting the second target virtual speaker according to the sound field components obtained by each virtual speaker from the first scene audio signal.
[0244] In some embodiments of the present application, the audio encoding method provided in the embodiments of the present application further includes the following steps:
[0245] E1. Acquire a second main sound field component from the first scene audio signal according to the virtual speaker set.
[0246] In the scenario where step E1 is executed, the aforementioned step D1 selects a second target virtual speaker from a preset virtual speaker set according to the first scene audio signal, including:
[0247] F1. Select a second target virtual speaker from the virtual speaker set according to the second main sound field component.
[0248] Among them, the encoding end obtains a virtual speaker set, and the encoding end uses the virtual speaker set to perform signal decomposition on the first scene audio signal to obtain the second main sound field component corresponding to the first scene audio signal. Among them, the second main sound field component represents the audio signal corresponding to the main sound field in the first scene audio signal. For example, the virtual speaker set includes multiple virtual speakers, and multiple sound field components can be obtained from the first scene audio signal according to the multiple virtual speakers, that is, each virtual speaker can obtain a sound field component from the first scene audio signal, and then the second main sound field component is selected from the multiple sound field components. For example, the second main sound field component can be one or several sound field components with the largest value among the multiple sound field components, or the second main sound field component can be one or several sound field components with dominant directions among the multiple sound field components. According to the second main sound field component, the second target virtual speaker is selected from the virtual speaker set, for example, the virtual speaker corresponding to the second main sound field component is the second target virtual speaker selected by the encoding end. In the embodiment of the present application, the encoding end can select the second target virtual speaker through the main sound field component, which solves the problem that the encoding end needs to determine the second target virtual speaker.
[0249] In some embodiments of the present application, the aforementioned step F1 selects a second target virtual speaker from the virtual speaker set according to the second main sound field component, including:
[0250] Selecting an HOA coefficient corresponding to the second main sound field component from the HOA coefficient set according to the second main sound field component, wherein the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set;
[0251] The virtual speaker corresponding to the HOA coefficient corresponding to the second main sound field component in the virtual speaker set is determined as the second target virtual speaker.
[0252] The above implementation is similar to the process of determining the first target virtual speaker in the aforementioned embodiment, and will not be described in detail here.
[0253] In some embodiments of the present application, the aforementioned step F1 selects a second target virtual speaker from the virtual speaker set according to the second main sound field component, and further includes:
[0254] G1. Acquire configuration parameters of a second target virtual speaker according to a second main sound field component;
[0255] G2. Generate an HOA coefficient corresponding to the second target virtual speaker according to the configuration parameters of the second target virtual speaker;
[0256] G3. Determine the virtual speaker corresponding to the HOA coefficient corresponding to the second target virtual speaker in the virtual speaker set as the second target virtual speaker.
[0257] The above implementation is similar to the process of determining the first target virtual speaker in the aforementioned embodiment, and will not be described in detail here.
[0258] The above implementation is similar to the process of determining the first target virtual speaker in the aforementioned embodiment, and will not be described in detail here.
[0259] In some embodiments of the present application, step G1 obtains configuration parameters of the second target virtual speaker according to the second main sound field component, including:
[0260] Determine configuration parameters of multiple virtual speakers in the virtual speaker set according to configuration information of the audio encoder;
[0261] The configuration parameters of the second target virtual speaker are selected from the configuration parameters of the plurality of virtual speakers according to the second main sound field component.
[0262] The above implementation is similar to the process of determining the configuration parameters of the first target virtual speaker in the aforementioned embodiment, and will not be described in detail here.
[0263] Wherein, in some embodiments of the present application, the configuration parameters of the second target virtual speaker include: position information and HOA order information of the second target virtual speaker;
[0264] The aforementioned step G2 generates the HOA coefficient corresponding to the second target virtual speaker according to the configuration parameters of the second target virtual speaker, including:
[0265] The HOA coefficient corresponding to the second target virtual speaker is determined according to the position information and the HOA order information of the second target virtual speaker.
[0266] The above implementation is similar to the process of determining the HOA coefficient corresponding to the first target virtual speaker in the aforementioned embodiment, and will not be repeated here.
[0267] In some embodiments of the present application, the first scene audio signal includes: a to-be-encoded HOA signal; the property information of the second target virtual speaker includes an HOA coefficient of the second target virtual speaker;
[0268] Step D2 generates a second virtual speaker signal according to the first scene audio signal and the attribute information of the second target virtual speaker, including:
[0269] The HOA signal to be encoded and the HOA coefficient of the second target virtual speaker are linearly combined to obtain a second virtual speaker signal.
[0270] In some embodiments of the present application, the first scene audio signal includes: a high-order ambisonic reverberation HOA signal to be encoded; the attribute information of the second target virtual speaker includes position information of the second target virtual speaker;
[0271] Step D2 generates a second virtual speaker signal according to the first scene audio signal and the attribute information of the second target virtual speaker, including:
[0272] Acquire the HOA coefficient corresponding to the second target virtual speaker according to the position information of the second target virtual speaker;
[0273] The HOA signal to be encoded and the HOA coefficients corresponding to the second target virtual speaker are linearly combined to obtain a second virtual speaker signal.
[0274] The above implementation is similar to the process of determining the first virtual speaker signal in the aforementioned embodiment, and will not be described in detail here.
[0275] In the embodiment of the present application, after the encoding end generates the second virtual speaker signal, the encoding end may further perform step D3 to encode the second virtual speaker signal and write it into the bitstream. The encoding method adopted by the encoding end is similar to step 403, so that the bitstream can carry the encoding result of the second virtual speaker signal.
[0276] In some embodiments of the present application, the audio encoding method performed by the encoding end may further include the following steps:
[0277] I1. Align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal.
[0278] In the scenario of executing step I1, correspondingly, step D3 of encoding the second virtual speaker signal includes:
[0279] encoding the aligned second virtual loudspeaker signal;
[0280] Accordingly, step 403 encodes the first virtual loudspeaker signal, including:
[0281] The aligned first virtual loudspeaker signal is encoded.
[0282] Among them, the encoding end can generate a first virtual speaker signal and a second virtual speaker signal, and the encoding end can align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal. An example is given as follows. There are two virtual speaker signals. The channel order of the virtual speaker signal of the current frame is 1 and 2, which correspond to the virtual speaker signals generated by the target virtual speakers P1 and P2 respectively. The channel order of the virtual speaker signal of the previous frame is 1 and 2, which correspond to the virtual speaker signals generated by the target virtual speakers P2 and P1 respectively. Then, the channel order of the virtual speaker signal of the current frame can be adjusted according to the order of the target virtual speakers of the previous frame. For example, the channel order of the virtual speaker signal of the current frame is adjusted to 2 and 1, so that the virtual speaker signals generated by the same target virtual speakers are on the same channel.
[0283] After the encoding end obtains the aligned first virtual speaker signal, the aligned first virtual speaker signal can be encoded. In the embodiment of the present application, the correlation between channels is enhanced by realigning the channels of the first virtual speaker signal, which is beneficial to the encoding processing of the first virtual speaker signal by the core encoder.
[0284] In some embodiments of the present application, in addition to performing the aforementioned steps, the encoding end may further include:
[0285] D1, selecting a second target virtual speaker from the virtual speaker set according to the first scene audio signal;
[0286] D2. Generate a second virtual speaker signal according to the first scene audio signal and the attribute information of the second target virtual speaker.
[0287] Accordingly, in the scenario where the encoding end performs steps D1 to D2, step 403 encodes the first virtual speaker signal, including:
[0288] J1. Obtain a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, where the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal;
[0289] J2. Encode the downmix signal and the side information.
[0290] Among them, after the encoding end obtains the first virtual speaker signal and the second virtual speaker signal, the encoding end can also perform downmix processing according to the first virtual speaker signal and the second virtual speaker signal to generate a downmix signal, for example, the first virtual speaker signal and the second virtual speaker signal are downmixed in amplitude to obtain a downmix signal. In addition, side information can be generated according to the first virtual speaker signal and the second virtual speaker signal, and the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal. The relationship has multiple implementation methods. The side information can be used by the decoding end to perform upmixing on the downmix signal to restore the first virtual speaker signal and the second virtual speaker signal. For example, the side information includes a signal information loss analysis parameter, so that the decoding end restores the first virtual speaker signal and the second virtual speaker signal through the signal information loss analysis parameter. For another example, the side information can specifically be a correlation parameter of the first virtual speaker signal and the second virtual speaker signal, for example, it can be an energy ratio parameter of the first virtual speaker signal and the second virtual speaker signal. So that the decoding end can restore the first virtual speaker signal and the second virtual speaker signal through the above correlation parameter or energy ratio parameter.
[0291] In some embodiments of the present application, when the encoding end performs steps D1 to D2, the encoding end may further perform the following steps:
[0292] I1. Align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal.
[0293] In the scenario of executing step I1, correspondingly, step J1 obtains the downmix signal and the side information according to the first virtual speaker signal and the second virtual speaker signal, including:
[0294] Obtain a downmix signal and side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal;
[0295] Accordingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
[0296] Before generating the downmix signal, the encoder may first perform an alignment operation on the virtual speaker signal, and after completing the alignment operation, generate the downmix signal and the side information. In the embodiment of the present application, by realigning the first virtual speaker signal and the channels of the second virtual speaker, the correlation between the channels is enhanced, which is beneficial to the encoding processing of the first virtual speaker signal by the core encoder.
[0297] It should be noted that in the above embodiments of the present application, the second scene audio signal can be obtained based on the first virtual speaker signal before alignment and the second virtual speaker signal before alignment, or it can be obtained based on the first virtual speaker signal after alignment and the second virtual speaker signal after alignment. The specific implementation method depends on the application scenario and is not limited here.
[0298] In some embodiments of the present application, before selecting a second target virtual speaker from a virtual speaker set according to the first scene audio signal in step D1, the audio signal encoding method provided in the embodiment of the present application further includes:
[0299] K1. Determine whether it is necessary to obtain a target virtual speaker other than the first target virtual speaker according to the coding rate and / or the signal type information of the first scene audio signal;
[0300] K2. If it is necessary to obtain a target virtual speaker other than the first target virtual speaker, a second target virtual speaker is selected from the virtual speaker set according to the first scene audio signal.
[0301] Among them, the encoding end can also perform signal selection to determine whether it is necessary to obtain the second target virtual speaker. In the case where the second target virtual speaker needs to be obtained, the encoding end can generate a second virtual speaker signal. In the case where the second target virtual speaker does not need to be obtained, the encoding end may not generate the second virtual speaker signal. Among them, the encoder can make a decision based on the configuration information of the audio encoder and / or the signal type information of the first scene audio signal to determine whether it is necessary to select another target virtual speaker in addition to selecting the first target virtual speaker. For example, if the encoding rate is higher than the preset threshold, it is determined that the target virtual speakers corresponding to the two main sound field components need to be obtained. In addition to determining the first target virtual speaker, the second target virtual speaker can also be determined. For another example, according to the signal type information of the first scene audio signal, it is determined that the target virtual speakers corresponding to the two main sound field components with dominant sound source directions need to be obtained. In addition to determining the first target virtual speaker, the second target virtual speaker can also be determined. On the contrary, if it is determined that only one target virtual speaker needs to be obtained according to the encoding rate and / or the signal type information of the first scene audio signal, after determining the first target virtual speaker, it is determined that no target virtual speakers other than the first target virtual speaker will be obtained. In the embodiment of the present application, signal selection can be used to reduce the amount of data encoded by the encoding end and improve the encoding efficiency.
[0302] When the encoder performs signal selection, it can determine whether it is necessary to generate a second virtual speaker signal. Since the encoder performs signal selection, information loss will occur, so it is necessary to perform signal compensation on the virtual speaker signal that is not transmitted. Signal compensation can be selected but not limited to information loss analysis, energy compensation, envelope compensation, noise compensation, etc. The compensation method can be selected as linear compensation or nonlinear compensation, etc. After signal compensation, side information can be generated, and the side information can be written into the bit stream, so that the decoder can obtain the side information through the bit stream, and the decoder can perform signal compensation based on the side information, thereby improving the quality of the decoded signal at the decoder.
[0303] Through the examples of the aforementioned embodiments, in the embodiments of the present application, a first virtual speaker signal can be generated according to the attribute information of the first scene audio signal and the first target virtual speaker, and the audio encoding end encodes the first virtual speaker signal instead of directly encoding the first scene audio signal. In the embodiments of the present application, a first target virtual speaker is selected according to the first scene audio signal, and the first virtual speaker signal generated based on the first target virtual speaker can represent the position sound field of the listener in the space, and the position sound field is as close as possible to the original sound field when the first scene audio signal is recorded, thereby ensuring the encoding quality of the audio encoding end, and the first virtual speaker signal and the residual signal are encoded to obtain a bit stream, and the amount of encoded data of the first virtual speaker signal is related to the first target virtual speaker, but has nothing to do with the number of channels of the first scene audio signal, thereby reducing the amount of encoded data and improving encoding efficiency.
[0304] In the embodiment of the application, the encoding end encodes the virtual speaker signal to generate a code stream. The encoding end can then output the code stream and send it to the decoding end through the audio transmission channel. The decoding end performs subsequent steps 411 to 413.
[0305] 411. Receive code stream.
[0306] The decoding end receives a bitstream from the encoding end. The bitstream may carry the encoded first virtual speaker signal. It is not limited that the bitstream may also carry the attribute information of the encoded first target virtual speaker. It should be noted that the bitstream may not carry the attribute information of the first target virtual speaker, and the decoding end may determine the attribute information of the first target virtual speaker by pre-configuration.
[0307] In addition, in some embodiments of the present application, when the encoding end generates a second virtual speaker signal, the code stream may also carry the second virtual speaker signal. It is not limited that the code stream may also carry the attribute information of the encoded second target virtual speaker. It should be noted that the code stream may not carry the attribute information of the second target virtual speaker, and the decoding end may determine the attribute information of the second target virtual speaker by pre-configuration.
[0308] 412. Decode the bit stream to obtain a virtual speaker signal.
[0309] After receiving the code stream from the encoding end, the decoding end decodes the code stream and obtains a virtual speaker signal from the code stream.
[0310] It should be noted that the virtual speaker signal may specifically be the aforementioned first virtual speaker signal, or may also be the aforementioned first virtual speaker signal and the second virtual speaker signal, which is not limited here.
[0311] In some embodiments of the present application, after the decoding end executes the above steps 411 to 412, the audio decoding method provided in the embodiment of the present application further includes the following steps:
[0312] Decode the bitstream to obtain the property information of the target virtual speaker.
[0313] In addition to encoding the virtual speaker, the encoding end may also encode the attribute information of the target virtual speaker, and write the encoded attribute information of the target virtual speaker into the bitstream, for example, the attribute information of the first target virtual speaker may be obtained through the bitstream. In the embodiment of the present application, the bitstream may carry the encoded attribute information of the first target virtual speaker, so that the decoding end can determine the attribute information of the first target virtual speaker by decoding the bitstream, which is convenient for audio decoding at the decoding end.
[0314] 413. Obtain a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal.
[0315] Among them, the decoding end can obtain the attribute information of the target virtual speaker, which is a virtual speaker in the virtual speaker set used to play back the reconstructed scene audio signal. The attribute information of the target virtual speaker may include the position information of the target virtual speaker and the HOA coefficient of the target virtual speaker. After the decoding end obtains the virtual speaker signal, the decoding end uses the attribute information of the target virtual speaker to reconstruct the signal, and the reconstructed scene audio signal can be output through signal reconstruction.
[0316] In some embodiments of the present application, the attribute information of the target virtual speaker includes the HOA coefficient of the target virtual speaker;
[0317] Step 413 obtains a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal, including:
[0318] The virtual speaker signal and the HOA coefficient of the target virtual speaker are synthesized to obtain a reconstructed scene audio signal.
[0319] The decoding end first determines the HOA coefficient of the target virtual speaker. For example, the HOA coefficient of the target virtual speaker can be pre-stored in the decoding end. After the decoding end obtains the virtual speaker signal and the HOA coefficient of the target virtual speaker, the reconstructed scene audio signal can be obtained according to the virtual speaker signal and the HOA coefficient of the target virtual speaker. Thus, the quality of the reconstructed scene audio signal is improved.
[0320] An example is given below. The HOA coefficient of the target virtual speaker is represented by the matrix A'. The size of the matrix A' is (M×C), where C is the number of target virtual speakers and M is the number of channels of the HOA coefficient of order N. The virtual speaker signal is represented by the matrix W'. The size of the matrix W' is (C×L), where L is the number of signal sampling points. The reconstructed HOA signal is obtained by the following calculation formula:
[0321] H=A'W',
[0322] The H obtained by the above calculation formula is the reconstructed HOA signal.
[0323] In some embodiments of the present application, the attribute information of the target virtual speaker includes location information of the target virtual speaker;
[0324] Step 413 obtains a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal, including:
[0325] Determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker;
[0326] The virtual speaker signal and the HOA coefficient of the target virtual speaker are synthesized to obtain a reconstructed scene audio signal.
[0327] Among them, the attribute information of the target virtual speaker may include: the position information of the target virtual speaker. The decoding end pre-stores the HOA coefficient of each virtual speaker in the virtual speaker set, and the decoding end also stores the position information of each virtual speaker. For example, the decoding end can determine the HOA coefficient corresponding to the position information of the target virtual speaker based on the correspondence between the position information of the virtual speaker and the HOA coefficient of the virtual speaker, or the decoding end can calculate the HOA coefficient of the target virtual speaker based on the position information of the target virtual speaker. Therefore, the decoding end can determine the HOA coefficient of the target virtual speaker through the position information of the target virtual speaker. The problem that the decoding end needs to determine the HOA coefficient of the target virtual speaker is solved.
[0328] In some embodiments of the present application, it can be seen from the method description of the encoding end that the virtual speaker signal is a downmixed signal obtained by downmixing the first virtual speaker signal and the second virtual speaker signal. In this implementation scenario, the audio decoding method provided by the embodiment of the present application also includes:
[0329] Decoding the bitstream to obtain side information, where the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal;
[0330] A first virtual speaker signal and a second virtual speaker signal are obtained according to the side information and the downmix signal.
[0331] In an embodiment of the present invention, the relationship between the first virtual speaker signal and the second virtual speaker signal may be a direct relationship or an indirect relationship; for example, when the relationship between the first virtual speaker signal and the second virtual speaker signal is a direct relationship, the first side information may include a correlation parameter between the first virtual speaker signal and the second virtual speaker signal, for example, an energy ratio parameter between the first virtual speaker signal and the second virtual speaker signal; for example, when the relationship between the first virtual speaker signal and the second virtual speaker signal is an indirect relationship, the first side information may include a correlation parameter between the first virtual speaker signal and the downmix signal, and a correlation parameter between the second virtual speaker signal and the downmix signal, for example, an energy ratio parameter between the first virtual speaker signal and the downmix signal, and an energy ratio parameter between the second virtual speaker signal and the downmix signal.
[0332] When the relationship between the first virtual speaker signal and the second virtual speaker signal may be a direct relationship, the decoder may determine the first virtual speaker signal and the second virtual speaker signal based on the downmix signal, the method for obtaining the downmix signal and the direct relationship; when the relationship between the first virtual speaker signal and the second virtual speaker signal may be an indirect relationship, the decoder may determine the first virtual speaker signal and the second virtual speaker signal based on the downmix signal and the indirect relationship.
[0333] Correspondingly, step 413 obtains a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal, including:
[0334] A reconstructed scene audio signal is obtained according to the property information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
[0335] Among them, the encoding end generates a downmix signal when downmixing the first virtual speaker signal and the second virtual speaker signal. The encoding end can also perform signal compensation on the downmix signal to generate side information. The side information can be written into the bit stream. The decoding end can obtain the side information through the bit stream. The decoding end can perform signal compensation according to the side information to obtain the first virtual speaker signal and the second virtual speaker signal. Therefore, when reconstructing the signal, the first virtual speaker signal and the second virtual speaker signal, as well as the aforementioned attribute information of the target virtual speaker, can be used, thereby improving the quality of the decoded signal at the decoding end.
[0336] Through the examples of the aforementioned embodiments, in the embodiments of the present application, a virtual speaker signal can be decoded from the bit stream, and the virtual speaker signal is used as the playback signal of the scene audio signal. The reconstructed scene audio signal is obtained through the attribute information of the target virtual speaker and the virtual speaker signal. In the embodiments of the present application, the obtained bit stream carries the virtual speaker signal and the residual signal, which reduces the amount of decoded data and improves the decoding efficiency.
[0337] An example is given below. In an embodiment of the present application, the first virtual speaker signal is represented by fewer channels than the first scene audio signal. For example, the first scene audio signal is a third-order HOA signal, and the HOA signal has 16 channels. In an embodiment of the present application, the 16 channels can be compressed into 2 channels, that is, the virtual speaker signal generated by the encoding end is 2 channels. For example, the virtual speaker signal generated by the encoding end may include the aforementioned first virtual speaker signal and the second virtual speaker signal, etc. The number of channels of the virtual speaker signal generated by the encoding end is independent of the number of channels of the first scene audio signal. It can be seen from the description of the subsequent steps that a virtual speaker signal of 2 channels can be carried in the bitstream. Accordingly, the decoding end receives the bitstream, and the virtual speaker signal obtained by decoding the bitstream is 2 channels. The decoding end can reconstruct a scene audio signal of 16 channels through the virtual speaker signal of 2 channels, and ensures that the reconstructed scene audio signal has the same subjective and objective quality as the original scene audio signal.
[0338] In order to better understand and implement the above-mentioned solutions in the embodiments of the present application, the following examples are given for specific explanation by using corresponding application scenarios.
[0339] In the embodiment of the present application, the scene audio signal is taken as an HOA signal as an example. The sound wave propagates in an ideal medium, the wave number is k=w / c, the angular frequency w=2πf, f is the sound wave frequency, and c is the sound speed. Then the sound pressure p satisfies the following calculation formula, where is the Laplace operator:
[0340]
[0341] Solving the above equation in spherical coordinates, in the passive spherical region, the equation is solved as follows:
[0342]
[0343] In the above calculation formula, r represents the radius of the sphere, θ represents the horizontal angle, represents the elevation angle, k represents the wave number, s represents the amplitude of the ideal plane wave, and m represents the HOA order number. is the spherical Bessel function, also known as the radial basis function, where the first j is the imaginary unit. Does not vary with angle. That is θ, Spherical harmonics of the direction, is the spherical harmonic function of the sound source direction.
[0344] The HOA coefficient can be expressed as:
[0345] Then the following calculation formula is given:
[0346]
[0347] The above formula shows that the sound field can be expanded on the sphere according to spherical harmonics, using the coefficient Alternatively, if the coefficients are known, The sound field can be reconstructed. The above formula is truncated to the Nth term with coefficients As an approximate description of the sound field, it is called the N-order HOA coefficient, which can also be called the Ambisonic coefficient. The N-order HOA coefficient has a total of (N+1) 2 Channels. Among them, Ambisonic signals above the first order are also called HOA signals. By superimposing the coefficients of the spherical harmonics corresponding to a sampling point of the HOA signal, the spatial sound field at the time corresponding to the sampling point can be reconstructed.
[0348] For example, in one configuration, the HOA order can be 2 to 6, the signal sampling rate is 48 to 192kHz when recording scene audio, and the sampling depth is 16 or 24 bits. The characteristic of HOA signal is that it carries the spatial information of the sound field, and is a description of the sound field signal at a certain point in space with a certain accuracy. Therefore, it is possible to consider using another representation form to describe the sound field signal at that point. If this description method can achieve the same accuracy of the description of the signal at that point using less data, the purpose of signal compression can be achieved.
[0349] The spatial sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field expressed by the HOA signal can be re-expressed by the superposition of multiple plane waves, each plane wave is represented by an audio signal of one channel and a direction vector. If the representation of the superposition of plane waves can better express the original sound field with fewer channels, the purpose of signal compression can be achieved.
[0350] When the HOA signal is actually played back, it can be played back through headphones or through multiple speakers arranged in the room. When using speakers for playback, the basic method is to superimpose the sound fields of multiple speakers so that the sound field at a certain point in the space (the location of the listener) is as close as possible to the original sound field when the HOA signal was recorded under a certain standard. The embodiment of the present application assumes a virtual speaker array, and then calculates the playback signal of the virtual speaker array, uses the playback signal as the transmission signal, and then generates a compressed signal. The decoding end obtains the playback signal by decoding the bit stream, and reconstructs the scene audio signal through the playback signal.
[0351] The embodiment of the present application provides an encoding end suitable for encoding scene audio signals, and a decoding end suitable for decoding scene audio signals. The encoding end encodes the original HOA signal into a compressed code stream, and the encoding end sends the compressed code stream to the decoding end, and then the decoding end restores the compressed code stream to reconstruct the HOA signal. In the embodiment of the present application, the amount of data after compression by the encoding end is as small as possible, or the quality of the HOA signal obtained after reconstruction by the decoding end is higher at the same bit rate.
[0352] The embodiment of the present application can solve the problems of large data volume, high bandwidth occupancy, low compression efficiency and low encoding quality when encoding HOA signals. 2 Channels, directly transmitting the HOA signal consumes a large bandwidth, so an effective multi-channel encoding scheme is needed.
[0353] The embodiment of the present application adopts a different channel extraction method, and the embodiment of the present application does not limit the assumption of the sound source, does not rely on the single sound source assumption in the time-frequency domain, and can more effectively process complex scenes such as multiple sound source signals. The codec of the embodiment of the present application provides a spatial coding method that uses fewer channels to represent the original HOA signal. Figure 5 As shown in FIG. 1 , a schematic diagram of the structure of the encoding end provided in an embodiment of the present application is shown. The encoding end includes a spatial encoder and a core encoder. The spatial encoder can perform channel extraction on the HOA signal to be encoded to generate a virtual speaker signal. The core encoder can encode the virtual speaker signal to obtain a bit stream. The encoding end sends the bit stream to the decoding end. Figure 6 As shown, it is a structural schematic diagram of a decoding end provided in an embodiment of the present application, and the decoding end includes: a core decoder and a spatial decoder, wherein the core decoder first receives a code stream from the encoding end, and then decodes a virtual speaker signal from the code stream, and then the spatial decoder reconstructs the virtual speaker signal to obtain a reconstructed HOA signal.
[0354] Next, examples are given from the encoding end and the decoding end respectively.
[0355] like Figure 7 As shown, first, the encoding end provided by the embodiment of the present application is described, and the encoding end may include: a virtual speaker configuration unit, a coding analysis unit, a virtual speaker set generation unit, a virtual speaker selection unit, a virtual speaker signal generation unit, and a core encoder processing unit. Next, the functions of each component unit of the encoding end are described respectively. In the embodiment of the present application, Figure 7 The encoding end shown in the figure can generate one virtual speaker signal or multiple virtual speaker signals, wherein the generation process of multiple virtual speaker signals can be based on Figure 7The encoder structure shown is generated multiple times, and the following takes the generation process of a virtual speaker signal as an example.
[0356] The virtual speaker configuration unit is used to configure the virtual speakers in the virtual speaker set to obtain a plurality of virtual speakers.
[0357] The virtual speaker configuration unit outputs virtual speaker configuration parameters according to the encoder configuration information. The encoder configuration information includes but is not limited to: HOA order, encoding bit rate, user-defined information, etc. The virtual speaker configuration parameters include but are not limited to: the number of virtual speakers, the HOA order of the virtual speakers, the position coordinates of the virtual speakers, etc.
[0358] The virtual speaker configuration parameters output by the virtual speaker configuration unit serve as inputs of the virtual speaker set generation unit.
[0359] The coding analysis unit is used to perform coding analysis on the HOA signal to be coded, for example, analyzing the sound field distribution of the HOA signal to be coded, including the number of sound sources, directivity, dispersion and other characteristics of the HOA signal to be coded, as one of the judgment conditions for determining how to select the target virtual speaker.
[0360] It is not limited that, in the embodiment of the present application, the encoding end may not include a coding analysis unit, that is, the encoding end may not analyze the input signal, and a default configuration is used to determine how to select the target virtual speaker.
[0361] Among them, the encoding end obtains the HOA signal to be encoded, for example, the HOA signal recorded from the actual acquisition device or the HOA signal synthesized by using an artificial audio object can be used as the input of the encoder, and the HOA signal to be encoded input by the encoder can be a time domain HOA signal or a frequency domain HOA signal.
[0362] The virtual speaker set generating unit is used to generate a virtual speaker set. The virtual speaker set may include: a plurality of virtual speakers. The virtual speakers in the virtual speaker set may also be referred to as "candidate virtual speakers".
[0363] The virtual speaker set generation unit generates the specified candidate virtual speaker HOA coefficients. The generation of candidate virtual speaker HOA coefficients requires the coordinates of the candidate virtual speakers (i.e., position coordinates or position information) and the HOA order of the candidate virtual speakers. The coordinate determination method of the candidate virtual speakers includes but is not limited to generating K virtual speakers according to the equidistant rule and generating K candidate virtual speakers with non-uniform distribution according to the principle of auditory perception. The following is an example of a method for generating a fixed number of uniformly distributed virtual speakers.
[0364] The coordinates of evenly distributed candidate virtual speakers are generated according to the number of candidate virtual speakers, for example, a numerical iterative calculation method is used to give an approximately even speaker arrangement. Figure 8 As shown in the figure, it is a schematic diagram of virtual speakers that are approximately uniformly distributed on a sphere. It is assumed that some particles are distributed on the unit sphere, and a quadratic repulsion is set between these particles, which is similar to the electrostatic repulsion between charges of the same kind. Let these particles move freely under the repulsive force, and it can be expected that when they reach a steady state, the distribution of the particles should tend to be uniform. In the calculation, the actual physical laws are simplified, and the moving distance of the particles is directly equal to the force. Then for the i-th particle, its moving distance at a certain step of the iterative calculation, that is, the virtual force it is subjected to, is calculated as follows:
[0365]
[0366] in, represents the displacement vector, represents the force vector, r ij represents the distance between the i-th mass point and the j-th mass point, represents the direction vector from the jth particle to the ith particle. The parameter k controls the size of the single step, and the initial position of the particle can be randomly specified.
[0367] The particle is displaced by a vector After the movement, it will generally deviate from the unit sphere. Before the next iteration, the distance between the particle and the center of the sphere can be normalized and moved back to the unit sphere. This can be obtained as follows Figure 8 As shown in the schematic diagram of virtual speaker distribution, multiple virtual speakers are approximately evenly distributed on the sphere.
[0368] Next, generate candidate virtual speaker HOA coefficients. The amplitude is s and the speaker position coordinates are The ideal plane wave of , after expansion using spherical harmonics, is given by the following calculation formula:
[0369]
[0370] The HOA coefficient for a plane wave is Satisfies the following calculation formula:
[0371]
[0372] The HOA coefficients of the candidate virtual speakers output by the virtual speaker set generation unit serve as input to the virtual speaker selection unit.
[0373] The virtual speaker selection unit is used to select a target virtual speaker from multiple candidate virtual speakers in the virtual speaker set according to the HOA signal to be encoded. The target virtual speaker can be called a "virtual speaker matching the HOA signal to be encoded" or simply a matching virtual speaker.
[0374] The virtual speaker selection unit matches the HOA signal to be encoded with the candidate virtual speaker HOA coefficients output by the virtual speaker set generation unit, and selects a designated matching virtual speaker.
[0375] Next, an example of a method for selecting a virtual speaker is given. In one embodiment, after obtaining a candidate virtual speaker, the HOA signal to be encoded is matched with the candidate virtual speaker HOA coefficient output by the virtual speaker set generation unit to find the best match of the HOA signal to be encoded on the candidate virtual speaker. The goal is to use the candidate virtual speaker HOA coefficient to match the combined HOA signal to be encoded. In one embodiment, the candidate virtual speaker HOA coefficient is used to make an inner product with the HOA signal to be encoded, and the candidate virtual speaker with the largest absolute value of the inner product is selected as the target virtual speaker, that is, the matching virtual speaker, and the projection of the HOA signal to be encoded on the candidate virtual speaker is superimposed on the linear combination of the candidate virtual speaker HOA coefficient, and then the projection vector is subtracted from the HOA signal to be encoded to obtain the difference, and the above process is repeated for the difference to implement iterative calculation, and a matching virtual speaker is generated each iteration, and the matching virtual speaker coordinates and the matching virtual speaker HOA coefficient are output. It can be understood that multiple matching virtual speakers will be selected, and a matching virtual speaker will be generated each iteration.
[0376] The coordinates of the target virtual speaker and the HOA coefficient of the target virtual speaker output by the virtual speaker selection unit are used as inputs of the virtual speaker signal generation unit.
[0377] In some embodiments of the present application, the encoding end includes Figure 7 In addition to the component units shown, a side information generating unit may also be included. It is not limited that the encoding end may not include the side information generating unit, which is only an example.
[0378] The coordinates of the target virtual speaker and / or the HOA coefficient of the target virtual speaker output by the virtual speaker selection unit are used as input of the side information generation unit.
[0379] The side information generation unit converts the HOA coefficients of the target virtual speakers or the coordinates of the target virtual speakers into side information, which is convenient for processing and transmission by the core encoder.
[0380] The output of the side information generation unit serves as the input of the core encoder processing unit.
[0381] The virtual speaker signal generating unit is used to generate a virtual speaker signal according to the HOA signal to be encoded and the attribute information of the target virtual speaker.
[0382] The virtual speaker signal generating unit calculates a virtual speaker signal by using the HOA signal to be encoded and the HOA coefficient of the target virtual speaker.
[0383] The HOA coefficients of the matched virtual speakers are represented by a matrix A. The matrix A can be used to linearly combine the HOA signals to be encoded. The theoretical optimal solution w can be obtained by the least square method, which is the virtual speaker signal. For example, the following calculation formula can be used:
[0384] w=A -1 X,
[0385] Among them, A -1 represents the inverse matrix of matrix A, the size of matrix A is (M×C), C is the number of target virtual speakers, M is the number of channels of the HOA coefficient of order N, and a represents the HOA coefficient of the target virtual speaker, for example,
[0386]
[0387] Where X represents the HOA signal to be encoded, the size of the matrix X is (M×L), M is the number of channels of the N-order HOA coefficient, L is the number of sampling points, and x represents the coefficient of the HOA signal to be encoded, for example,
[0388]
[0389] The virtual speaker signal output by the virtual speaker signal generating unit is used as the input of the core encoder processing unit.
[0390] In some embodiments of the present application, the encoding end includes Figure 7 In addition to the component units shown, a signal alignment unit may also be included. It is not limited that the encoding end may not include the signal alignment unit, which is only an example.
[0391] The virtual speaker signal output by the virtual speaker signal generating unit serves as the input of the signal aligning unit.
[0392] The signal alignment unit is used to realign the channels of the virtual speaker signal to enhance the correlation between the channels and facilitate the processing of the core encoder.
[0393] The aligned virtual speaker signal output by the signal alignment unit is the input of the core encoder processing unit.
[0394] The core encoder processing unit is used for performing core encoder processing on the side information and the aligned virtual speaker signal to obtain a transmission code stream.
[0395] The core encoder processing includes but is not limited to transformation, quantization, psychoacoustic model, bitstream generation, etc. It can process frequency domain channels as well as time domain channels, which is not limited here.
[0396] like Fig. 9 As shown, the decoding end provided in the embodiment of the present application may include: a core decoder processing unit and an HOA signal reconstruction unit.
[0397] The core decoder processing unit is used for performing core decoder processing on the transport code stream to obtain a virtual speaker signal.
[0398] It is not limited that if the encoding end carries side information in the bit stream, the decoding end also needs to include: a side information decoding unit.
[0399] The side information decoding unit is used to decode the decoded side information output by the core decoder processing unit to obtain decoded side information.
[0400] The core decoder processing may include transformation, code stream analysis, inverse quantization, etc., and may process the frequency domain channel or the time domain channel, which is not limited here.
[0401] The virtual loudspeaker signal output by the core decoder processing unit is the input of the HOA signal reconstruction unit, and the decoded side information output by the core decoder processing unit is the input of the side information decoding unit.
[0402] The side information decoding unit converts the decoded side information into the HOA coefficients of the target virtual speakers.
[0403] The HOA coefficients of the target virtual speakers output by the side information decoding unit are input to the HOA signal reconstruction unit.
[0404] The HOA signal reconstruction unit is used to reconstruct the HOA signal through the virtual speaker signal and the HOA coefficient of the target virtual speaker.
[0405] The HOA coefficients of the target virtual speakers are represented by the matrix A', the size of which is (M×C), denoted as A', where C is the number of target virtual speakers and M is the number of channels of the HOA coefficients of order N. The virtual speaker signals form a (C×L) matrix, denoted as W', where L is the number of signal sampling points. The reconstructed HOA signal H is obtained by the following calculation formula:
[0406] H=A'W',
[0407] The reconstructed HOA signal output by the HOA signal reconstruction unit is the output of the decoding end.
[0408] In the embodiment of the present application, the encoding end can use a spatial encoder to represent the original HOA signal with fewer channels. For example, the original third-order HOA signal can be compressed from 16 channels to 4 channels by using the spatial encoder of the embodiment of the present application, and it is ensured that there is no obvious difference in subjective hearing. Among them, subjective hearing test is an evaluation standard in audio coding and decoding, and no obvious difference is a level of subjective evaluation.
[0409] In other embodiments of the present application, the virtual speaker selection unit of the encoding end selects a target virtual speaker from a virtual speaker set, and may also use a virtual speaker in a specified orientation as the target virtual speaker. The virtual speaker signal generation unit directly projects onto each target virtual speaker to obtain a virtual speaker signal.
[0410] In the above manner, by specifying a virtual speaker in a certain position as a target virtual speaker, the virtual speaker selection process can be simplified and the encoding and decoding speed can be improved.
[0411] In some other embodiments of the present application, the encoder side may not include a signal alignment unit, and the output of the virtual speaker signal generation unit is directly encoded by the core encoder. In the above manner, the signal alignment process is reduced and the complexity of the encoder side is reduced.
[0412] It can be seen from the above examples that the embodiment of the present application applies the selected target virtual speaker to the HOA signal encoding and decoding. The embodiment of the present application can obtain accurate HOA signal sound source positioning, reconstruct the HOA signal direction more accurately, have higher coding efficiency, and the complexity of the decoding end is very low, which is beneficial to mobile applications and can improve the performance of encoding and decoding.
[0413] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0414] In order to better implement the above-mentioned solution of the embodiment of the present application, relevant devices for implementing the above-mentioned solution are also provided below.
[0415] See also Fig.10 As shown, an audio encoding device 1000 provided in an embodiment of the present application may include: an acquisition module 1001, a signal generation module 1002 and an encoding module 1003, wherein:
[0416] An acquisition module, configured to select a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal;
[0417] A signal generating module, configured to generate a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker;
[0418] The encoding module is used to encode the first virtual speaker signal to obtain a code stream.
[0419] In some embodiments of the present application, the acquisition module is used to acquire a main sound field component from the current scene audio signal according to the virtual speaker set; and select the first target virtual speaker from the virtual speaker set according to the main sound field component.
[0420] In some embodiments of the present application, the acquisition module is used to select HOA coefficients corresponding to the main sound field components from the high-order stereo reverberation HOA coefficient set according to the main sound field components, and the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set; and determine that the virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component in the virtual speaker set is the first target virtual speaker.
[0421] In some embodiments of the present application, the acquisition module is used to acquire the configuration parameters of the first target virtual speaker according to the main sound field components; generate the HOA coefficient corresponding to the first target virtual speaker according to the configuration parameters of the first target virtual speaker; and determine that the virtual speaker corresponding to the HOA coefficient corresponding to the first target virtual speaker in the virtual speaker set is the target virtual speaker.
[0422] In some embodiments of the present application, the acquisition module is used to determine the configuration parameters of multiple virtual speakers in the virtual speaker set according to the configuration information of the audio encoder; and select the configuration parameters of the first target virtual speaker from the configuration parameters of the multiple virtual speakers according to the main sound field components.
[0423] In some embodiments of the present application, the configuration parameters of the first target virtual speaker include: location information and HOA order information of the first target virtual speaker;
[0424] The acquisition module is used to determine the HOA coefficient corresponding to the first target virtual speaker according to the position information and HOA order information of the first target virtual speaker.
[0425] In some embodiments of the present application, the encoding module is further used to encode the attribute information of the first target virtual speaker and write it into the bit stream.
[0426] In some embodiments of the present application, the current scene audio signal includes: a to-be-encoded HOA signal; the attribute information of the first target virtual speaker includes an HOA coefficient of the first target virtual speaker;
[0427] The signal generating module is used to linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
[0428] In some embodiments of the present application, the current scene audio signal includes: a high-order ambisonic reverberation HOA signal to be encoded; the attribute information of the first target virtual speaker includes the position information of the first target virtual speaker;
[0429] The signal generating module is used to obtain the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker; and linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
[0430] In some embodiments of the present application, the acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0431] The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0432] The encoding module is used to encode the second virtual speaker signal and write it into the code stream.
[0433] In some embodiments of the present application, the signal generating module is used to align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal;
[0434] Correspondingly, the encoding module is used to encode the aligned second virtual loudspeaker signal;
[0435] Correspondingly, the encoding module is used to encode the aligned first virtual loudspeaker signal.
[0436] In some embodiments of the present application, the acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal;
[0437] The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker;
[0438] Correspondingly, the encoding module is used to obtain a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, wherein the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal; and encode the downmix signal and the side information.
[0439] In some embodiments of the present application, the signal generating module is used to align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal;
[0440] Correspondingly, the encoding module is used to obtain the downmix signal and the side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal;
[0441] Correspondingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
[0442] In some embodiments of the present application, the acquisition module is used to determine whether it is necessary to acquire a target virtual speaker other than the first target virtual speaker based on the encoding rate and / or the signal type information of the current scene audio signal before selecting the second target virtual speaker from the virtual speaker set based on the current scene audio signal; if it is necessary to acquire a target virtual speaker other than the first target virtual speaker, the second target virtual speaker is selected from the virtual speaker set based on the current scene audio signal.
[0443] See also Fig.11 As shown, an audio decoding device 1100 provided in an embodiment of the present application may include: a receiving module 1101, a decoding module 1102, and a reconstruction module 1103, wherein:
[0444] A receiving module, used for receiving a code stream;
[0445] A decoding module, used for decoding the code stream to obtain a virtual speaker signal;
[0446] The reconstruction module is used to obtain a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal.
[0447] In some embodiments of the present application, the decoding module is further used to decode the code stream to obtain the property information of the target virtual speaker.
[0448] In some embodiments of the present application, the attribute information of the target virtual speaker includes a high-order ambisonic reverberation HOA coefficient of the target virtual speaker;
[0449] The reconstruction module is used to perform synthesis processing on the virtual speaker signal and the HOA coefficient of the target virtual speaker to obtain the reconstructed scene audio signal.
[0450] In some embodiments of the present application, the attribute information of the target virtual speaker includes location information of the target virtual speaker;
[0451] The reconstruction module is used to determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker; and synthesize the virtual speaker signal and the HOA coefficient of the target virtual speaker to obtain the reconstructed scene audio signal.
[0452] In some embodiments of the present application, the virtual speaker signal is a downmixed signal obtained by downmixing the first virtual speaker signal and the second virtual speaker signal, and the device further includes: a signal compensation module, wherein:
[0453] The decoding module is used to decode the bit stream to obtain side information, where the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal;
[0454] The signal compensation module is configured to obtain the first virtual speaker signal and the second virtual speaker signal according to the side information and the downmix signal;
[0455] Correspondingly, the reconstruction module is used to obtain the reconstructed scene audio signal according to the attribute information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
[0456] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.
[0457] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a program, and the program executes some or all of the steps recorded in the above method embodiment.
[0458] Next, another audio encoding device provided by the embodiment of the present application is introduced. Fig.12 As shown, the audio encoding device 1200 includes:
[0459] The receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 (wherein the number of the processor 1203 in the audio encoding device 1200 can be one or more, Fig.12 In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 may be connected via a bus or other means, wherein: Fig.12 The example of connecting through bus is taken in the following.
[0460] The memory 1204 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1203. A portion of the memory 1204 may also include a non-volatile random access memory (NVRAM). The memory 1204 stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.
[0461] The processor 1203 controls the operation of the audio encoding device, and the processor 1203 may also be referred to as a central processing unit (CPU). In a specific application, the various components of the audio encoding device are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, various buses are referred to as bus systems in the figure.
[0462] The method disclosed in the above embodiment of the present application can be applied to the processor 1203, or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1203. The above processor 1203 can be a general processor, a digital signal processor (digital signal processing, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and completes the steps of the above method in combination with its hardware.
[0463] The receiver 1201 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the audio encoding device. The transmitter 1202 may include a display device such as a display screen. The transmitter 1202 can be used to output digital or character information through an external interface.
[0464] In the embodiment of the present application, the processor 1203 is used to execute the above-mentioned embodiment. Figure 4 The audio encoding method shown is performed by the audio encoding device.
[0465] Next, another audio decoding device provided by the embodiment of the present application is introduced. Fig.13 As shown, the audio decoding device 1300 includes:
[0466] The receiver 1301, the transmitter 1302, the processor 1303 and the memory 1304 (wherein the number of the processor 1303 in the audio decoding device 1300 can be one or more, Fig.13 In some embodiments of the present application, the receiver 1301, the transmitter 1302, the processor 1303 and the memory 1304 may be connected via a bus or other means, wherein: Fig.13 The example of connecting through bus is taken in the following.
[0467] The memory 1304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1303. A portion of the memory 1304 may also include an NVRAM. The memory 1304 stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.
[0468] The processor 1303 controls the operation of the audio decoding device, and the processor 1303 can also be called a CPU. In a specific application, the various components of the audio decoding device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.
[0469] The method disclosed in the above embodiment of the present application can be applied to the processor 1303, or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1303. The above processor 1303 can be a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be combined and executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1304, and the processor 1303 reads the information in the memory 1304 and completes the steps of the above method in combination with its hardware.
[0470] In the embodiment of the present application, the processor 1303 is used to execute the above embodiment. Figure 4 The audio decoding method shown is performed by the audio decoding device.
[0471] In another possible design, when the audio encoding device or the audio decoding device is a chip in a terminal, the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute computer-executable instructions stored in the storage unit, so that the chip in the terminal executes the audio encoding method of any one of the first aspects or the audio decoding method of any one of the second aspects. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the terminal, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0472] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect or second aspect method.
[0473] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0474] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0475] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0476] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. An audio encoding method, It is characterized in that include: Obtaining main sound field components from the current scene audio signal according to the virtual speaker set; Selecting a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal; Generate a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker; Encoding the first virtual speaker signal to obtain a bit stream; Wherein, selecting a first target virtual speaker from a preset virtual speaker set according to the current scene audio signal comprises: selecting the first target virtual speaker from the virtual speaker set according to the main sound field component; The step of selecting the first target virtual speaker from the virtual speaker set according to the main sound field component comprises: Selecting, according to the main sound field component, an HOA coefficient corresponding to the main sound field component from a high-order ambisonic reverberation HOA coefficient set, wherein the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set; Determine a virtual speaker in the virtual speaker set that corresponds to the HOA coefficient corresponding to the main sound field component as the first target virtual speaker.
2. The method according to claim 1, It is characterized in that The method further comprises: The attribute information of the first target virtual speaker is encoded and written into the bitstream.
3. The method according to any one of claims 1 to 2, It is characterized in that The current scene audio signal includes: a high-order ambisonic reverberation HOA signal to be encoded; the attribute information of the first target virtual speaker includes the HOA coefficient of the first target virtual speaker; The generating a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker comprises: The HOA signal to be encoded and the HOA coefficients are linearly combined to obtain the first virtual speaker signal.
4. The method according to any one of claims 1 to 2, It is characterized in that The current scene audio signal includes: a high-order ambisonic reverberation (HOA) signal to be encoded; the attribute information of the first target virtual speaker includes the position information of the first target virtual speaker; The generating a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker comprises: Acquire the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker; The HOA signal to be encoded and the HOA coefficients are linearly combined to obtain the first virtual speaker signal.
5. The method according to any one of claims 1 to 2, It is characterized in that The method further comprises: Selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal; Generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker; The second virtual speaker signal is encoded and written into the bit stream.
6. The method according to claim 5, It is characterized in that The method further comprises: Performing alignment processing on the first virtual loudspeaker signal and the second virtual loudspeaker signal to obtain an aligned first virtual loudspeaker signal and an aligned second virtual loudspeaker signal; Accordingly, encoding the second virtual loudspeaker signal comprises: encoding the aligned second virtual loudspeaker signal; Accordingly, encoding the first virtual loudspeaker signal includes: The aligned first virtual loudspeaker signal is encoded.
7. The method according to any one of claims 1 to 2, It is characterized in that The method further comprises: Selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal; Generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker; Accordingly, encoding the first virtual loudspeaker signal includes: obtaining a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, wherein the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal; The downmix signal and the side information are encoded.
8. The method according to claim 7, It is characterized in that The method further comprises: Performing alignment processing on the first virtual loudspeaker signal and the second virtual loudspeaker signal to obtain an aligned first virtual loudspeaker signal and an aligned second virtual loudspeaker signal; Correspondingly, obtaining the downmix signal and the side information according to the first virtual speaker signal and the second virtual speaker signal includes: Obtain the downmix signal and the side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal; Correspondingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
9. The method according to claim 5, It is characterized in that Before selecting a second target virtual speaker from the virtual speaker set according to the current scene audio signal, the method further includes: Determine whether it is necessary to obtain a target virtual speaker other than the first target virtual speaker according to the encoding rate and / or the signal type information of the current scene audio signal; If it is necessary to obtain a target virtual speaker other than the first target virtual speaker, a second target virtual speaker is selected from the virtual speaker set according to the current scene audio signal.
10. An audio decoding method, It is characterized in that include: Receive code stream; Decoding the bit stream to obtain a virtual speaker signal; Decoding the bit stream to obtain property information of the target virtual speaker; Obtaining a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal; Wherein, the attribute information of the target virtual speaker includes the position information of the target virtual speaker; The step of obtaining the reconstructed scene audio signal according to the property information of the target virtual speaker and the virtual speaker signal comprises: Determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker; The virtual speaker signal and the HOA coefficient of the target virtual speaker are synthesized to obtain the reconstructed scene audio signal.
11. The method according to claim 10, It is characterized in that The virtual speaker signal is a downmixed signal obtained by downmixing the first virtual speaker signal and the second virtual speaker signal, and the method further includes: Decoding the bitstream to obtain side information, where the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal; Obtain the first virtual speaker signal and the second virtual speaker signal according to the side information and the downmix signal; Correspondingly, obtaining the reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal includes: The reconstructed scene audio signal is obtained according to the property information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
12. An audio encoding device, It is characterized in that include: An acquisition module, configured to select a first target virtual speaker from a preset virtual speaker set according to a current scene audio signal; A signal generating module, configured to generate a first virtual speaker signal according to the current scene audio signal and the attribute information of the first target virtual speaker; An encoding module, used for encoding the first virtual speaker signal to obtain a code stream; The acquisition module is used to acquire a main sound field component from the current scene audio signal according to the virtual speaker set; and select the first target virtual speaker from the virtual speaker set according to the main sound field component; The acquisition module is used to select the HOA coefficient corresponding to the main sound field component from the high-order stereo reverberation HOA coefficient set according to the main sound field component, and the HOA coefficients in the HOA coefficient set correspond one-to-one to the virtual speakers in the virtual speaker set; determine that the virtual speaker corresponding to the HOA coefficient corresponding to the main sound field component in the virtual speaker set is the first target virtual speaker.
13. The device according to claim 12, It is characterized in that The encoding module is further used to encode the attribute information of the first target virtual speaker and write it into the code stream.
14. The device according to any one of claims 12 to 13, It is characterized in that The current scene audio signal includes: a to-be-encoded HOA signal; the attribute information of the first target virtual speaker includes an HOA coefficient of the first target virtual speaker; The signal generating module is used to linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
15. The device according to any one of claims 12 to 13, It is characterized in that The current scene audio signal includes: a high-order ambisonic reverberation (HOA) signal to be encoded; the attribute information of the first target virtual speaker includes the position information of the first target virtual speaker; The signal generating module is used to obtain the HOA coefficient corresponding to the first target virtual speaker according to the position information of the first target virtual speaker; and linearly combine the HOA signal to be encoded and the HOA coefficient to obtain the first virtual speaker signal.
16. The device according to any one of claims 12 to 13, It is characterized in that The acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal; The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker; The encoding module is used to encode the second virtual speaker signal and write it into the code stream.
17. The device according to claim 16, It is characterized in that The signal generating module is used to align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal; Correspondingly, the encoding module is used to encode the aligned second virtual loudspeaker signal; Correspondingly, the encoding module is used to encode the aligned first virtual loudspeaker signal.
18. The device according to any one of claims 12 to 13, It is characterized in that The acquisition module is used to select a second target virtual speaker from the virtual speaker set according to the current scene audio signal; The signal generating module is used to generate a second virtual speaker signal according to the current scene audio signal and the attribute information of the second target virtual speaker; Correspondingly, the encoding module is used to obtain a downmix signal and side information according to the first virtual speaker signal and the second virtual speaker signal, wherein the side information is used to indicate a relationship between the first virtual speaker signal and the second virtual speaker signal; The downmix signal and the side information are encoded.
19. The device according to claim 18, It is characterized in that The signal generating module is used to align the first virtual speaker signal and the second virtual speaker signal to obtain an aligned first virtual speaker signal and an aligned second virtual speaker signal; Correspondingly, the encoding module is used to obtain the downmix signal and the side information according to the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal; Correspondingly, the side information is used to indicate the relationship between the aligned first virtual loudspeaker signal and the aligned second virtual loudspeaker signal.
20. The device according to claim 16, It is characterized in that The acquisition module is used to determine whether it is necessary to acquire a target virtual speaker other than the first target virtual speaker according to the encoding rate and / or the signal type information of the current scene audio signal before selecting the second target virtual speaker from the virtual speaker set according to the current scene audio signal; if it is necessary to acquire a target virtual speaker other than the first target virtual speaker, the second target virtual speaker is selected from the virtual speaker set according to the current scene audio signal.
21. An audio decoding device, It is characterized in that include: A receiving module, used for receiving a code stream; A decoding module, used for decoding the code stream to obtain a virtual speaker signal; The decoding module is further used to decode the code stream to obtain the property information of the target virtual speaker; A reconstruction module, used for obtaining a reconstructed scene audio signal according to the attribute information of the target virtual speaker and the virtual speaker signal; The attribute information of the target virtual speaker includes the position information of the target virtual speaker; The reconstruction module is used to determine the HOA coefficient of the target virtual speaker according to the position information of the target virtual speaker; and synthesize the virtual speaker signal and the HOA coefficient of the target virtual speaker to obtain the reconstructed scene audio signal.
22. The device according to claim 21, It is characterized in that The virtual speaker signal is a downmixed signal obtained by downmixing the first virtual speaker signal and the second virtual speaker signal. The device further includes: a signal compensation module, wherein: The decoding module is used to decode the bit stream to obtain side information, where the side information is used to indicate the relationship between the first virtual speaker signal and the second virtual speaker signal; The signal compensation module is configured to obtain the first virtual speaker signal and the second virtual speaker signal according to the side information and the downmix signal; Correspondingly, the reconstruction module is used to obtain the reconstructed scene audio signal according to the attribute information of the target virtual speaker, the first virtual speaker signal and the second virtual speaker signal.
23. An audio encoding device, It is characterized in that The audio encoding device comprises at least one processor, and the at least one processor is used to be coupled to a memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 9.
24. The audio encoding device according to claim 23, It is characterized in that The audio encoding device further includes: the memory.
25. An audio decoding device, It is characterized in that The audio decoding device comprises at least one processor, and the at least one processor is used to be coupled to a memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 10 to 11.
26. The audio decoding device according to claim 25, It is characterized in that The audio decoding device further includes: the memory.
27. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 9, or 10 to 11.
28. A computer-readable storage medium, comprising a code stream generated by the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Efficient rendering of virtual soundfields
US20190379992A1