Three-dimensional audio signal encoding method, apparatus, and encoder
By adjusting the voting value of the virtual loudspeaker and using the representative coefficient, the problems of instability between frames and high computational complexity in 3D audio signal encoding were solved, resulting in more stable audio-visual representation and a more efficient encoding process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-05-17
- Publication Date
- 2026-05-22
AI Technical Summary
In existing 3D audio signal encoding technologies, the frequent switching of virtual speakers between frames leads to unstable sound images in the reconstructed 3D audio signal, affecting sound quality, and also results in high computational complexity.
By inheriting the voting values of the representative virtual loudspeakers from previous frames in the encoder, adjusting the initial voting values of the current frame, selecting a more stable virtual loudspeaker, reducing the jumps between frames, enhancing the signal orientation continuity, and using a smaller number of representative coefficients for virtual loudspeaker search, the computational complexity is reduced.
It improves the acoustic stability and sound quality of the reconstructed 3D audio signal, reduces the computational burden on the encoder, and improves the efficiency of compression coding.
Smart Images

Figure CN115376530B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia, and in particular to a three-dimensional audio signal encoding method, apparatus and encoder. Background Technology
[0002] With the rapid development of high-performance computers and signal processing technology, listeners have placed increasingly higher demands on voice and audio experiences, and immersive audio can meet these needs. For example, 3D audio technology has been widely used in wireless communication (such as 4G / 5G), virtual reality / augmented reality, and media audio. 3D audio technology is an audio technology that acquires, processes, transmits, renders, and plays back real-world sounds and 3D sound field information, giving sound a strong sense of space, immersion, and presence, providing listeners with an extraordinary auditory experience that makes them feel as if they are actually there.
[0003] Typically, acquisition devices (e.g., microphones) collect large amounts of data to record 3D sound field information and transmit 3D audio signals to playback devices (e.g., speakers, headphones) for playback. Due to the large amount of 3D sound field information, significant storage space is required, and the bandwidth demand for transmitting the 3D audio signals is high. To address these issues, the 3D audio signals can be compressed, and the compressed data can be stored or transmitted. Currently, the encoder first iterates through the set of candidate virtual speakers and uses the selected virtual speakers to compress the 3D audio signal. However, if the results of selecting virtual speakers in consecutive frames differ significantly, the reconstructed 3D audio signal becomes unstable, reducing its sound quality. Summary of the Invention
[0004] This application provides a three-dimensional audio signal encoding method, apparatus, and encoder, which can enhance the positional continuity between frames, improve the stability of the sound image of the reconstructed three-dimensional audio signal, and ensure the sound quality of the reconstructed three-dimensional audio signal.
[0005] Firstly, this application provides a three-dimensional audio signal encoding method, which can be executed by an encoder and specifically includes the following steps: After the encoder obtains a first number of initial voting values for the current frame of the three-dimensional audio signal, it obtains a seventh number of final voting values for the current frame corresponding to a seventh number of virtual speakers, based on the first number of initial voting values for the current frame and the sixth number of final voting values for the sixth number of virtual speakers corresponding to the previous frames of the three-dimensional audio signal. Wherein, the virtual speakers correspond one-to-one with the initial voting values for the current frame; the first number of virtual speakers includes a first virtual speaker, and the initial voting value for the first virtual speaker in the current frame is used to characterize the priority of using the first virtual speaker when encoding the current frame; the seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number of virtual speakers includes the sixth number of virtual speakers. Then, the encoder selects representative virtual speakers for the second number of current frames from the seventh number of virtual speakers based on the final voting values of the seventh number of current frames. The second number is less than the seventh number, which means that the representative virtual speakers for the second number of current frames are a portion of the virtual speakers in the seventh number of virtual speakers. The current frame is then encoded based on the representative virtual speakers of the second number of current frames to obtain the bitstream.
[0006] During the virtual speaker search process, the locations of real sound sources and virtual speakers may not coincide, leading to a one-to-one correspondence between virtual speakers and real sound sources. Furthermore, in complex real-world scenarios, a limited set of virtual speakers may not be able to represent all sound sources in the sound field. In such cases, frequent jumps in the virtual speakers found between frames can significantly impact the listener's auditory experience, resulting in noticeable discontinuities and noise in the decoded and reconstructed 3D audio signal. The virtual speaker selection method provided in this application inherits the representative virtual speaker from previous frames. Specifically, for virtual speakers with the same number, the initial voting value of the current frame is adjusted using the final voting value of the previous frame. This makes the encoder more inclined to select the representative virtual speaker from the previous frame, thereby reducing frequent jumps in virtual speakers between frames, enhancing the continuity of signal orientation between frames, improving the stability of the sound image in the reconstructed 3D audio signal, and ensuring the sound quality of the reconstructed 3D audio signal.
[0007] For example, if the sixth number of virtual speakers includes the first virtual speaker, obtaining the seventh number of final voting values of the seventh number of virtual speakers corresponding to the current frame based on the first number of initial voting values of the current frame and the sixth number of voting values of the sixth number of virtual speakers and the previous frame of the three-dimensional audio signal includes: updating the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.
[0008] In one possible implementation, if a first number of virtual speakers includes a second virtual speaker and a sixth number of virtual speakers does not include a second virtual speaker, then the final vote value of the second virtual speaker in the current frame is equal to the initial vote value of the second virtual speaker in the current frame; or, if a sixth number of virtual speakers includes a third virtual speaker and a first number of virtual speakers does not include a third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.
[0009] In another possible implementation, updating the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame includes: the encoder adjusting the final voting value of the first virtual speaker in the previous frame according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame; and updating the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame.
[0010] The first adjustment parameter is determined based on at least one of the following: the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type. Therefore, the encoder uses the first adjustment parameter to adjust the final voting value of the first virtual loudspeaker in the previous frame, making the encoder more inclined to select the representative virtual loudspeaker in the previous frame. This enhances the directional continuity between frames, improves the stability of the sound image of the reconstructed 3D audio signal, and ensures the sound quality of the reconstructed 3D audio signal.
[0011] In another possible implementation, updating the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame includes: the encoder adjusts the initial voting value of the first virtual speaker in the current frame according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame; and updates the adjusted voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame.
[0012] The second adjustment parameter is determined based on the adjusted voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame. Therefore, the encoder uses the second adjustment parameter to adjust the initial voting value of the first virtual speaker in the current frame, reducing frequent jumps in the initial voting value. This makes the encoder more inclined to select the representative virtual speaker in the previous frame, thereby enhancing the positional continuity between frames, improving the stability of the sound image of the reconstructed 3D audio signal, and ensuring the sound quality of the reconstructed 3D audio signal.
[0013] The second quantity represents the number of representative virtual speakers for the current frame selected by the encoder. A larger second quantity indicates a larger number of representative virtual speakers for the current frame, resulting in more sound field information in the 3D audio signal; a smaller second quantity indicates a smaller number of representative virtual speakers for the current frame, resulting in less sound field information in the 3D audio signal. Therefore, the number of representative virtual speakers for the current frame selected by the encoder can be controlled by setting the second quantity. For example, the second quantity can be preset, or it can be determined based on the current frame. For instance, the value of the second quantity can be 1, 2, 4, or 8.
[0014] In another possible implementation, obtaining the first number of initial voting values for the current frame corresponding to the first number of virtual speakers and the current frame of the 3D audio signal includes: the encoder determining the first number of virtual speakers and the first number of initial voting values for the current frame based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds. The candidate virtual speaker set includes a fifth number of virtual speakers, the fifth number of virtual speakers includes the first number of virtual speakers, the first number is less than or equal to the fifth number, and the number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.
[0015] Currently, in the virtual speaker search process, the encoder uses the correlation calculation results between the 3D audio signal to be encoded and the virtual speaker as the selection metric. Furthermore, if the encoder transmits a virtual speaker for each coefficient, efficient data compression cannot be achieved, placing a heavy computational burden on the encoder. The virtual speaker selection method provided in this application involves the encoder using a smaller number of representative coefficients to vote on each virtual speaker in the candidate virtual speaker set, replacing all coefficients of the current frame, and selecting the representative virtual speaker for the current frame based on the voting values. Furthermore, the encoder uses the representative virtual speaker of the current frame to compress and encode the 3D audio signal to be encoded, effectively improving the compression ratio of the 3D audio signal and reducing the computational complexity of the encoder's virtual speaker search, thereby reducing the computational complexity of the 3D audio signal compression and encoding and alleviating the encoder's computational burden.
[0016] In another possible implementation, before determining the first number of virtual speakers and the first number of initial voting values for the current frame based on the third number of representative coefficients, the candidate virtual speaker set, and the number of voting rounds, the method further includes: the encoder obtaining the fourth number of coefficients for the current frame and the frequency domain feature values of the fourth number of coefficients; and selecting the third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients, wherein the third number is less than the fourth number, indicating that the third number of representative coefficients are a subset of the fourth number of coefficients.
[0017] The current frame of the three-dimensional audio signal is a higher-order ambisonics (HOA) signal; the frequency domain characteristic values of the coefficients are determined based on the coefficients of the HOA signal.
[0018] Thus, since the encoder selects a portion of the coefficients from all coefficients in the current frame as representative coefficients, and uses a smaller number of representative coefficients to replace all coefficients in the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder searching for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden on the encoder.
[0019] In addition, the encoder encodes the current frame based on the representative virtual speakers of the second number of current frames to obtain the bitstream, including: the encoder generates a virtual speaker signal based on the representative virtual speakers of the second number of current frames and the current frame; and encodes the virtual speaker signal to obtain the bitstream.
[0020] In another possible implementation, the method further includes: the encoder obtaining a first correlation between the current frame and the representative virtual speaker set of previous frames; if the first correlation does not satisfy the reuse condition, obtaining a fourth number of coefficients of the current frame of the 3D audio signal, and the frequency domain feature values of the fourth number of coefficients. The representative virtual speaker set of previous frames includes a sixth number of virtual speakers, which are the representative virtual speakers of the previous frames used to encode the 3D audio signal; the first correlation is used to determine whether to reuse the representative virtual speaker set of previous frames when encoding the current frame.
[0021] In this way, the encoder can first determine whether the representative virtual speaker set of the previous frame can be reused to encode the current frame. If the encoder reuses the representative virtual speaker set of the previous frame to encode the current frame, it avoids the encoder from performing the virtual speaker search process again, effectively reducing the computational complexity of the encoder searching for virtual speakers. Therefore, it reduces the computational complexity of compressing and encoding the 3D audio signal and alleviates the computational burden on the encoder. In addition, it can also reduce the frequent jumps of virtual speakers between frames, enhance the continuity of the orientation between frames, improve the stability of the sound image of the reconstructed 3D audio signal, and ensure the sound quality of the reconstructed 3D audio signal. If the encoder cannot reuse the representative virtual speaker set of the previous frame to encode the current frame, the encoder then selects representative coefficients and uses the representative coefficients of the current frame to vote on each virtual speaker in the candidate virtual speaker set. Based on the voting value, the representative virtual speaker of the current frame is selected, thereby achieving the purpose of reducing the computational complexity of compressing and encoding the 3D audio signal and alleviating the computational burden on the encoder.
[0022] Optionally, the method further includes: the encoder can also acquire the current frame of the three-dimensional audio signal so as to compress and encode the current frame of the three-dimensional audio signal to obtain a bitstream, and transmit the bitstream to the decoding end.
[0023] Secondly, this application provides a three-dimensional audio signal encoding apparatus, the apparatus comprising modules for performing the three-dimensional audio signal encoding method of the first aspect or any possible design of the first aspect. For example, the three-dimensional audio signal encoding apparatus includes a virtual speaker selection module and an encoding module. The virtual speaker selection module is used to obtain a first number of initial voting values for the current frame corresponding to a first number of virtual speakers and the current frame of the 3D audio signal. Each virtual speaker corresponds one-to-one with the initial voting value for the current frame. The first number of virtual speakers includes a first virtual speaker, and the initial voting value of the first virtual speaker for the current frame is used to characterize the priority of using the first virtual speaker when encoding the current frame. The virtual speaker selection module is also used to obtain a seventh number of final voting values for the current frame corresponding to a seventh number of virtual speakers, based on the first number of initial voting values for the current frame and the sixth number of final voting values for the previous frames corresponding to a sixth number of virtual speakers and the 3D audio signal. The seventh number of virtual speakers includes the first number of virtual speakers and also includes the sixth number of virtual speakers. The virtual speaker selection module is also used to select representative virtual speakers for a second number of current frames from the seventh number of virtual speakers, based on the seventh number of final voting values for the current frames. The second number is less than the seventh number. The encoding module is used to encode the current frame based on the representative virtual speakers of the second number of current frames to obtain a bitstream. These modules can perform the corresponding functions in the method examples in the first aspect above. Please refer to the detailed description in the method examples for details, which will not be repeated here.
[0024] Thirdly, this application provides an encoder including at least one processor and a memory, wherein the memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, it performs the operation steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation of the first aspect.
[0025] Fourthly, this application provides a system comprising an encoder as described in the third aspect and a decoder, wherein the encoder is configured to perform operational steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation thereof, and the decoder is configured to decode the bitstream generated by the encoder.
[0026] Fifthly, this application provides a computer-readable storage medium, comprising: computer software instructions; when the computer software instructions are executed in an encoder, causing the encoder to perform the operation steps of the method as described in the first aspect or any possible implementation thereof.
[0027] Sixthly, this application provides a computer program product that, when run on an encoder, causes the encoder to perform the operation steps of the method as described in the first aspect or any possible implementation thereof.
[0028] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the structure of an audio encoding and decoding system provided in an embodiment of this application;
[0030] Figure 2 This is a schematic diagram of a scenario for an audio encoding / decoding system provided in an embodiment of this application;
[0031] Figure 3 This is a schematic diagram of the structure of an encoder provided in an embodiment of this application;
[0032] Figure 4 A flowchart illustrating a three-dimensional audio signal encoding and decoding method provided in an embodiment of this application;
[0033] Figure 5 A flowchart illustrating a method for selecting a virtual speaker, provided as an embodiment of this application;
[0034] Figure 6 A flowchart illustrating a three-dimensional audio signal encoding method provided in an embodiment of this application;
[0035] Figure 7 A flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application;
[0036] Figure 8 A flowchart illustrating a method for adjusting voting values provided in an embodiment of this application;
[0037] Figure 9 A flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application;
[0038] Figure 10 A schematic diagram of the structure of an encoding device provided in this application;
[0039] Figure 11 This is a schematic diagram of the structure of an encoder provided in this application. Detailed Implementation
[0040] To ensure clarity and brevity in the description of the following embodiments, a brief introduction to the relevant technologies is given first.
[0041] Sound is a continuous wave produced by the vibration of an object. The object that produces vibrations and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solids, or liquids), the auditory organs of humans or animals can perceive the sound.
[0042] Sound waves are characterized by pitch, intensity, and timbre. Pitch indicates the highness or lowness of a sound. Intensity indicates the loudness or volume of a sound. The unit of intensity is the decibel (dB). Timbre is also known as tone color.
[0043] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is hertz (Hz). The human ear can distinguish sounds with frequencies between 20 Hz and 20,000 Hz.
[0044] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.
[0045] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, and pulse waves, among others.
[0046] Based on the characteristics of sound waves, sound can be divided into regular sound and irregular sound. Irregular sound refers to sound emitted by the irregular vibration of a sound source. Irregular sound is, for example, noise that affects people's work, study, and rest. Regular sound refers to sound emitted by the regular vibration of a sound source. Regular sound includes speech and musical tones. When sound is represented electronically, regular sound is an analog signal that varies continuously in the time and frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.
[0047] Because human hearing has the ability to distinguish the location of sound sources in space, when a listener hears a sound in space, in addition to being able to perceive the pitch, intensity, and timbre of the sound, they can also perceive the location of the sound.
[0048] As people pay increasing attention to and demand higher quality in their auditory experience, three-dimensional audio technology has emerged to enhance the depth, presence, and spatial feel of sound. This allows listeners to not only perceive sounds from front, back, left, and right sources, but also to feel surrounded by the spatial sound field created by these sources, and to experience the sound spreading outwards, creating an immersive audio experience as if the listener were in a cinema or concert hall.
[0049] Three-dimensional audio technology refers to the concept of the space outside the human ear as a system, where the signal received at the eardrum is a three-dimensional audio signal output after the sound emitted from the sound source has been filtered by this external system. For example, the system outside the human ear can be defined as the system impulse response h(n), any sound source can be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal described in this application's embodiments may refer to a higher-order ambisonics (HOA) signal. Three-dimensional audio can also be called three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio, or binaural audio, etc.
[0050] As is well known, sound waves propagate in an ideal medium with a wave number of k = w / c and an angular frequency of w = 2πf, where f is the sound wave frequency and c is the speed of sound. The sound pressure p satisfies formula (1). For the Laplace operator.
[0051]
[0052] Assuming the spatial system outside the human ear is a sphere, with the listener at the center, the sound from outside the sphere has a projection on the sphere's surface. Filtering out sounds from outside the sphere, and assuming the sound sources are distributed on this sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound source. That is, three-dimensional audio technology is a method of fitting a sound field. Specifically, in spherical coordinates, equation (1) is solved. In the passive spherical region, the solution to equation (1) is as follows: equation (2).
[0053]
[0054] Where r represents the radius of the sphere, and θ represents the horizontal angle. denoted by pitch angle, k by wave number, s by amplitude of ideal plane wave, and m by order number of three-dimensional audio signal (or HOA signal). Let represent the spherical Bessel function, also known as the radial basis function, where the first 'j' represents the imaginary unit. It does not change with the angle. express spherical harmonic function of direction, The spherical harmonic function represents the direction of the sound source. The coefficients of the three-dimensional audio signal satisfy formula (3).
[0055]
[0056] Substituting formula (3) into formula (2), formula (2) can be transformed into formula (4).
[0057]
[0058] in, The coefficients of the three-dimensional audio signal of order N are used to approximate the sound field. A sound field refers to the region in a medium where sound waves exist. N is an integer greater than or equal to 1. For example, the value of N ranges from 2 to 6. The coefficients of the three-dimensional audio signal described in the embodiments of this application may refer to HOA coefficients or ambisonic coefficients.
[0059] A three-dimensional audio signal is an information carrier that carries the spatial location information of the sound source in the sound field, describing the sound field of the listener in space. Equation (4) shows that the sound field can be expanded on a sphere according to the spherical harmonic function, that is, the sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional audio signal can be expressed by the superposition of multiple plane waves, and the sound field can be reconstructed through the coefficients of the three-dimensional audio signal.
[0060] Compared to a 5.1 channel audio signal or a 7.1 channel audio signal, an Nth-order HOA signal has (N+1) 2 With multiple channels, the HOA signal contains a large amount of data describing the spatial information of the sound field. If the acquisition device (e.g., a microphone) transmits this 3D audio signal to the playback device (e.g., a speaker), it consumes a significant amount of bandwidth. Currently, encoders can use spatial squeezed surround audio coding (S3AC) or directional audio coding (DirAC) to compress and encode the 3D audio signal to obtain a bitstream, which is then transmitted to the playback device. The playback device decodes the bitstream, reconstructs the 3D audio signal, and plays the reconstructed 3D audio signal. This reduces the amount of data transmitted to the playback device and the bandwidth usage. However, the computational complexity of compressing and encoding 3D audio signals is high, consuming excessive computational resources. Therefore, how to reduce the computational complexity of compressing and encoding 3D audio signals is a problem that urgently needs to be solved.
[0061] This application provides an audio encoding and decoding technology, particularly a three-dimensional audio encoding and decoding technology for three-dimensional audio signals. Specifically, it provides an encoding and decoding technology that uses fewer channels to represent three-dimensional audio signals, thereby improving traditional audio encoding and decoding systems. Audio encoding (or commonly referred to as encoding) includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and typically includes processing (e.g., compressing) the raw audio to reduce the amount of data required to represent the raw audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and typically includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are also collectively referred to as encoding and decoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0062] Figure 1 This is a schematic diagram of an audio encoding / decoding system provided in an embodiment of this application. The audio encoding / decoding system 100 includes a source device 110 and a destination device 120. The source device 110 is used to compress and encode a three-dimensional audio signal to obtain a bitstream, and transmit the bitstream to the destination device 120. The destination device 120 decodes the bitstream, reconstructs the three-dimensional audio signal, and plays the reconstructed three-dimensional audio signal.
[0063] Specifically, the source device 110 includes an audio acquisition unit 111, a preprocessor 112, an encoder 113, and a communication interface 114.
[0064] Audio acquirer 111 is used to acquire raw audio. Audio acquirer 111 can be any type of audio acquisition device for capturing real-world sounds, and / or any type of audio generation device. Audio acquirer 111 is, for example, a computer audio processor for generating computer audio. Audio acquirer 111 can also be any type of memory or storage device for storing audio. Audio includes real-world sounds, virtual scene sounds (e.g., VR or augmented reality (AR) sounds), and / or any combination thereof.
[0065] The preprocessor 112 receives the raw audio acquired by the audio acquirer 111 and preprocesses the raw audio to obtain a three-dimensional audio signal. For example, the preprocessing performed by the preprocessor 112 includes channel conversion, audio format conversion, or noise reduction.
[0066] Encoder 113 receives the three-dimensional audio signal generated by preprocessor 112 and compresses and encodes the three-dimensional audio signal to obtain a bitstream. For example, encoder 113 may include spatial encoder 1131 and core encoder 1132. Spatial encoder 1131 selects (or searches for) virtual speakers from a set of candidate virtual speakers based on the three-dimensional audio signal, and generates a virtual speaker signal based on the three-dimensional audio signal and the virtual speakers. The virtual speaker signal can also be called a playback signal. Core encoder 1132 encodes the virtual speaker signal to obtain a bitstream.
[0067] The communication interface 114 is used to receive the code stream generated by the encoder 113 and send the code stream to the destination device 120 through the communication channel 130 so that the destination device 120 can reconstruct the three-dimensional audio signal based on the code stream.
[0068] The target device 120 includes a player 121, a post-processor 122, a decoder 123, and a communication interface 124.
[0069] Communication interface 124 is used to receive the bitstream sent by communication interface 114 and transmit the bitstream to decoder 123 so that decoder 123 can reconstruct the three-dimensional audio signal based on the bitstream.
[0070] Communication interfaces 114 and 124 can be used to send or receive relevant data of the original audio through a direct communication link between the source device 110 and the destination device 120, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.
[0071] Both communication interface 114 and communication interface 124 can be configured as follows: Figure 1 The arrow pointing from the source device 110 to the corresponding communication channel 130 of the destination device 120 indicates a one-way or two-way communication interface, which can be used to send and receive messages, establish connections, acknowledge and exchange any other information related to the communication link and / or data transmission such as encoded bitstream transmission, etc.
[0072] Decoder 123 is used to decode the bitstream and reconstruct the three-dimensional audio signal. For example, decoder 123 includes a core decoder 1231 and a spatial decoder 1232. Core decoder 1231 decodes the bitstream to obtain a virtual speaker signal. Spatial decoder 1232 reconstructs the three-dimensional audio signal based on a set of candidate virtual speakers and the virtual speaker signal, obtaining the reconstructed three-dimensional audio signal.
[0073] The post-processor 122 receives the reconstructed 3D audio signal generated by the decoder 123 and performs post-processing on the reconstructed 3D audio signal. For example, the post-processing performed by the post-processor 122 includes audio rendering, loudness normalization, user interaction, audio format conversion, or noise reduction.
[0074] Player 121 is used to play the reconstructed sound based on the reconstructed 3D audio signal.
[0075] It should be noted that the audio acquirer 111 and encoder 113 can be integrated into a single physical device or located on different physical devices; there is no limitation on this. For example, such as... Figure 1 The source device 110 shown includes an audio acquirer 111 and an encoder 113, indicating that the audio acquirer 111 and encoder 113 are integrated into a single physical device. Therefore, the source device 110 can also be referred to as a capture device. The source device 110 can be, for example, a media gateway in a wireless access network, a media gateway in a core network, a transcoding device, a media resource server, an AR device, a VR device, a microphone, or other audio capture devices. If the source device 110 does not include the audio acquirer 111, it means that the audio acquirer 111 and encoder 113 are two different physical devices, and the source device 110 can acquire raw audio from other devices (such as audio capture devices or audio storage devices).
[0076] Furthermore, the player 121 and decoder 123 can be integrated into a single physical device or located on different physical devices; there is no limitation on this. For example, such as... Figure 1 The destination device 120 shown includes a player 121 and a decoder 123, indicating that the player 121 and decoder 123 are integrated into a single physical device. Therefore, the destination device 120 can also be referred to as a playback device. The destination device 120 has the function of decoding and playing back the reconstructed audio. The destination device 120 can be, for example, a speaker, headphones, or other audio playback device. If the destination device 120 does not include the player 121, it means that the player 121 and decoder 123 are two different physical devices. After decoding and reconstructing the three-dimensional audio signal from the bitstream, the destination device 120 transmits the reconstructed three-dimensional audio signal to other playback devices (such as speakers or headphones) for playback.
[0077] also, Figure 1 It is shown that the source device 110 and the destination device 120 can be integrated into one physical device or set on different physical devices, without limitation.
[0078] For example, such as Figure 2As shown in (a), source device 110 can be a microphone in a recording studio, and destination device 120 can be a speaker. Source device 110 can acquire the original audio of various musical instruments, transmit the original audio to an encoding / decoding device, the encoding / decoding device performs encoding and decoding processing on the original audio to obtain a reconstructed three-dimensional audio signal, and the destination device 120 plays back the reconstructed three-dimensional audio signal. As another example, source device 110 can be a microphone in a terminal device, and destination device 120 can be headphones. Source device 110 can acquire external sounds or audio synthesized by the terminal device.
[0079] For example, such as Figure 2 As shown in (b), if the source device 110 and the destination device 120 are integrated into a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, or an extended reality (XR) device, then the VR / AR / MR / XR device has the functions of capturing raw audio, playing back audio, and encoding / decoding. The source device 110 can capture the sound emitted by the user and the sound emitted by virtual objects in the virtual environment in which the user is located.
[0080] In these embodiments, the source device 110 or its corresponding functions and the destination device 120 or its corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. As described, Figure 1 The presence and division of different units or functions in the source device 110 and / or destination device 120 shown may vary depending on the actual device and application, which is obvious to those skilled in the art.
[0081] The structure of the audio codec system described above is only illustrative. In some possible implementations, the audio codec system may also include other devices, such as end-side devices or cloud-side devices. After the source device 110 acquires the raw audio, it preprocesses the raw audio to obtain a three-dimensional audio signal; and then transmits the three-dimensional audio to the end-side device or cloud-side device, which performs the encoding and decoding function on the three-dimensional audio signal.
[0082] The audio signal encoding and decoding method provided in this application is mainly applied at the encoding end. Combined with... Figure 3 The structure of the encoder is described in detail. For example... Figure 3 As shown, the encoder 300 includes a virtual speaker configuration unit 310, a virtual speaker set generation unit 320, an encoding analysis unit 330, a virtual speaker selection unit 340, a virtual speaker signal generation unit 350, and an encoding unit 360.
[0083] The virtual speaker configuration unit 310 generates virtual speaker configuration parameters based on encoder configuration information to obtain multiple virtual speakers. Encoder configuration information includes, but is not limited to: the order of the three-dimensional audio signal (or commonly referred to as the HOA order), the encoding bit rate, user-defined information, etc. Virtual speaker configuration parameters include, but are not limited to: the number of virtual speakers, the order of the virtual speakers, the position coordinates of the virtual speakers, etc. The number of virtual speakers can be, for example, 2048, 1669, 1343, 1024, 530, 512, 256, 128, or 64. The order of the virtual speakers can be any from 2nd to 6th order. The position coordinates of the virtual speakers include the horizontal angle and the pitch angle.
[0084] The virtual speaker configuration parameters output by the virtual speaker configuration unit 310 are used as input to the virtual speaker set generation unit 320.
[0085] The virtual speaker set generation unit 320 is used to generate a candidate virtual speaker set based on virtual speaker configuration parameters. The candidate virtual speaker set includes multiple virtual speakers. Specifically, the virtual speaker set generation unit 320 determines the multiple virtual speakers included in the candidate virtual speaker set based on the number of virtual speakers, and determines the coefficients of the virtual speakers based on the position information (e.g., coordinates) and order of the virtual speakers. For example, the method for determining the coordinates of the virtual speakers includes, but is not limited to: generating multiple virtual speakers according to an equidistant rule, or generating multiple non-uniformly distributed virtual speakers based on the principle of auditory perception; then, the coordinates of the virtual speakers are generated based on the number of virtual speakers.
[0086] Based on the above principle of generating three-dimensional audio signals, the coefficients of a virtual speaker can also be generated. The coefficients of the virtual speaker can be calculated using θ in formula (3). s and Set the position coordinates of the virtual speaker respectively. This represents the coefficients of an Nth-order virtual loudspeaker. The coefficients of a virtual loudspeaker can also be called ambisonic coefficients.
[0087] The encoding analysis unit 330 is used to perform encoding analysis on three-dimensional audio signals, such as analyzing the sound field distribution characteristics of three-dimensional audio signals, that is, the number of sound sources, the directionality of sound sources, and the dispersion of sound sources.
[0088] The coefficients of multiple virtual speakers included in the candidate virtual speaker set output by the virtual speaker set generation unit 320 are used as input to the virtual speaker selection unit 340.
[0089] The sound field distribution characteristics of the three-dimensional audio signal output by the encoding analysis unit 330 are used as the input of the virtual speaker selection unit 340.
[0090] The virtual speaker selection unit 340 is used to determine a representative virtual speaker that matches the three-dimensional audio signal based on the three-dimensional audio signal to be encoded, the sound field distribution characteristics of the three-dimensional audio signal, and the coefficients of multiple virtual speakers.
[0091] Not limited to this, the encoder 300 in this embodiment may also exclude the encoding analysis unit 330, that is, the encoder 300 may not analyze the input signal, and the virtual speaker selection unit 340 may use a default configuration to determine the representative virtual speaker. For example, the virtual speaker selection unit 340 may determine the representative virtual speaker that matches the three-dimensional audio signal only based on the coefficients of the three-dimensional audio signal and multiple virtual speakers.
[0092] The encoder 300 can use a three-dimensional audio signal acquired from the acquisition device or a three-dimensional audio signal synthesized using artificial audio objects as its input. Furthermore, the three-dimensional audio signal input to the encoder 300 can be either a time-domain three-dimensional audio signal or a frequency-domain three-dimensional audio signal, without limitation.
[0093] The virtual speaker selection unit 340 outputs the position information and coefficients representing the virtual speaker as inputs to the virtual speaker signal generation unit 350 and the encoding unit 360.
[0094] The virtual speaker signal generation unit 350 generates a virtual speaker signal based on a three-dimensional audio signal and attribute information representing a virtual speaker. The attribute information representing the virtual speaker includes at least one of position information representing the virtual speaker, coefficients representing the virtual speaker, and coefficients of the three-dimensional audio signal. If the attribute information is position information representing the virtual speaker, the coefficients representing the virtual speaker are determined based on the position information; if the attribute information includes coefficients of the three-dimensional audio signal, the coefficients representing the virtual speaker are obtained based on the coefficients of the three-dimensional audio signal. Specifically, the virtual speaker signal generation unit 350 calculates the virtual speaker signal based on the coefficients of the three-dimensional audio signal and the coefficients representing the virtual speaker.
[0095] For example, assume matrix A represents the coefficients of the virtual loudspeaker, and matrix X represents the HOA coefficients of the HOA signal. Matrix X is the inverse of matrix A. The theoretical optimal solution w is obtained using the least squares method, where w represents the virtual loudspeaker signal. The virtual loudspeaker signal satisfies formula (5).
[0096] w = A -1 Formula X (5)
[0097] Among them, A -1Let represent the inverse matrix of matrix A. Matrix A has a size of (M×C), where C represents the number of virtual loudspeakers, M represents the number of channels in the Nth-order HOA signal, and 'a' represents the coefficients of the virtual loudspeakers. Matrix X has a size of (M×L), where L represents the number of coefficients in the HOA signal, and 'x' represents the coefficients of the HOA signal. The coefficients representing virtual loudspeakers can refer to either the HOA coefficients or the ambisonic coefficients of the virtual loudspeakers. For example,
[0098] The virtual speaker signal output by the virtual speaker signal generation unit 350 is used as the input of the encoding unit 360.
[0099] The encoding unit 360 is used to perform core encoding processing on the virtual speaker signal to obtain a bitstream. Core encoding processing includes, but is not limited to: transformation, quantization, psychoacoustic modeling, noise shaping, bandwidth expansion, downmixing, arithmetic coding, and bitstream generation.
[0100] It is worth noting that the spatial encoder 1131 may include a virtual loudspeaker configuration unit 310, a virtual loudspeaker set generation unit 320, an encoding analysis unit 330, a virtual loudspeaker selection unit 340, and a virtual loudspeaker signal generation unit 350. That is, the virtual loudspeaker configuration unit 310, the virtual loudspeaker set generation unit 320, the encoding analysis unit 330, the virtual loudspeaker selection unit 340, and the virtual loudspeaker signal generation unit 350 implement the functions of the spatial encoder 1131. The core encoder 1132 may include an encoding unit 360. That is, the encoding unit 360 implements the functions of the core encoder 1132.
[0101] Figure 3 The encoder shown can generate one virtual speaker signal or multiple virtual speaker signals. Multiple virtual speaker signals can be generated by... Figure 3 The encoder shown can be obtained by multiple executions, or it can be obtained by... Figure 3 The encoder shown is obtained by executing it once.
[0102] Next, the encoding and decoding process of three-dimensional audio signals will be explained with reference to the accompanying drawings. Figure 4 This is a flowchart illustrating a three-dimensional audio signal encoding and decoding method provided in an embodiment of this application. Figure 1 The process of encoding and decoding three-dimensional audio signals, using the source device 110 and the destination device 120 as an example, will be explained below. Figure 4 As shown, the method includes the following steps.
[0103] S410, source device 110 acquires the current frame of the three-dimensional audio signal.
[0104] As described in the above embodiments, if the source device 110 carries an audio acquisition device 111, the source device 110 can acquire raw audio through the audio acquisition device 111. Optionally, the source device 110 can also receive raw audio acquired by other devices; or acquire raw audio from the memory or other memory in the source device 110. The raw audio may include at least one of real-world sounds acquired in real time, audio stored in the device, and audio synthesized from multiple audio sources. This embodiment does not limit the method of acquiring raw audio or the type of raw audio.
[0105] After acquiring the original audio, the source device 110 generates a three-dimensional audio signal based on three-dimensional audio technology and the original audio, so as to provide the listener with an "immersive" sound effect when playing back the original audio. The specific method for generating the three-dimensional audio signal can be found in the description of the preprocessor 112 in the above embodiments and in the description of existing technologies.
[0106] Furthermore, an audio signal is a continuous analog signal. In audio signal processing, the audio signal can first be sampled to generate a digital signal of frame sequences. A frame can include multiple sample points. A frame can also refer to the sample points obtained from sampling. A frame can also include subframes obtained by dividing a frame. For example, if a frame has a length of L sample points and is divided into N subframes, then each subframe corresponds to L / N sample points. Audio encoding and decoding typically refer to processing audio frame sequences containing multiple sample points.
[0107] An audio frame may include the current frame or a previous frame. In various embodiments of this application, the current frame or previous frame can refer to a frame or a subframe. The current frame refers to the frame undergoing encoding / decoding processing at the current moment. A previous frame refers to a frame that underwent encoding / decoding processing at a time prior to the current moment. A previous frame can be a frame from the moment before the current moment or from several previous moments. In embodiments of this application, the current frame of a 3D audio signal refers to a frame of 3D audio signal undergoing encoding / decoding processing at the current moment. A previous frame refers to a frame of 3D audio signal that underwent encoding / decoding processing at a time prior to the current moment. The current frame of a 3D audio signal can refer to the current frame of the 3D audio signal to be encoded. The current frame of a 3D audio signal can be simply referred to as the current frame. The previous frame of a 3D audio signal can be simply referred to as the previous frame.
[0108] S420 and source device 110 determine the candidate set of virtual speakers.
[0109] In one scenario, the source device 110 has a pre-configured set of candidate virtual speakers in its memory. The source device 110 can read the set of candidate virtual speakers from the memory. The set of candidate virtual speakers includes multiple virtual speakers. Virtual speakers represent speakers that are virtually present in the spatial sound field. The virtual speakers are used to calculate virtual speaker signals based on the three-dimensional audio signal so that the destination device 120 can play back the reconstructed three-dimensional audio signal.
[0110] In another scenario, the source device 110 has virtual speaker configuration parameters pre-configured in its memory. The source device 110 generates a candidate set of virtual speakers based on the virtual speaker configuration parameters. Optionally, the source device 110 generates the candidate set of virtual speakers in real time based on its own computing resources (e.g., processor) capabilities and the characteristics of the current frame (e.g., channel and data volume).
[0111] The specific method for generating a candidate set of virtual speakers can be found in the prior art, as well as in the description of the virtual speaker configuration unit 310 and the virtual speaker set generation unit 320 in the above embodiments.
[0112] S430 and source device 110 select the representative virtual speaker of the current frame from the candidate virtual speaker set based on the current frame of the three-dimensional audio signal.
[0113] The source device 110 votes on the virtual speakers based on the coefficients of the current frame and the coefficients of the virtual speakers, and selects the representative virtual speaker for the current frame from the candidate virtual speaker set based on the voting values of the virtual speakers. A limited number of representative virtual speakers for the current frame are searched from the candidate virtual speaker set as the best matching virtual speaker for the current frame to be encoded, thereby achieving the purpose of data compression of the three-dimensional audio signal to be encoded.
[0114] Figure 5 This is a flowchart illustrating a method for selecting a virtual speaker, as provided in an embodiment of this application. Figure 5 The method described is... Figure 4 This section describes the specific operational procedures included in the S430. (Hereinafter provided is a summary of the details.) Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Specifically, the function of the virtual speaker selection unit 340 is implemented. For example... Figure 5 As shown, the method includes the following steps.
[0115] S510, encoder 113 obtains the representative coefficients of the current frame.
[0116] Representative coefficients can refer to frequency domain representative coefficients or time domain representative coefficients. Frequency domain representative coefficients can also be called frequency domain representative frequency points or spectrum representative coefficients. Time domain representative coefficients can also be called time domain representative sampling points. For specific methods to obtain the representative coefficients of the current frame, please refer to the following: Figure 7 Explanation of S6101 and S6102.
[0117] S520, encoder 113 selects the representative virtual speaker for the current frame from the candidate virtual speaker set based on the voting values of the virtual speakers in the candidate virtual speaker set according to the representative coefficient of the current frame. Execute S440 to S460.
[0118] Encoder 113 votes on the virtual speakers in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker, and selects (searches) the representative virtual speaker of the current frame from the candidate virtual speaker set based on the final voting value of the virtual speaker in the current frame. The specific method for selecting the representative virtual speaker of the current frame can be found below. Figure 8 and Figure 7 The explanation of S6103.
[0119] It should be noted that the encoder first traverses the virtual speakers included in the candidate virtual speaker set, and then compresses the current frame using the representative virtual speaker of the current frame selected from the candidate virtual speaker set. However, if the results of the virtual speakers selected in consecutive frames differ significantly, it will lead to unstable sound image of the reconstructed 3D audio signal, reducing the sound quality of the reconstructed 3D audio signal. In the embodiments of this application, the encoder 113 can update the initial voting value of the virtual speakers in the current frame included in the candidate virtual speaker set based on the final voting value of the representative virtual speaker of the previous frame, to obtain the final voting value of the virtual speaker in the current frame, and then select the representative virtual speaker of the current frame from the candidate virtual speaker set based on the final voting value of the virtual speaker in the current frame. Thus, by referring to the representative virtual speaker of the previous frame to select the representative virtual speaker of the current frame, the encoder tends to select the same virtual speaker as the representative virtual speaker of the previous frame when selecting the representative virtual speaker of the current frame, increasing the directional continuity between consecutive frames and overcoming the problem of large differences in the results of the virtual speakers selected in consecutive frames. Therefore, the embodiments of this application may also include S530.
[0120] S530, encoder 113 adjusts the initial voting value of the virtual speaker in the current frame of the candidate virtual speaker set according to the final voting value of the representative virtual speaker in the previous frame, and obtains the final voting value of the virtual speaker in the current frame.
[0121] Encoder 113 votes on the virtual speakers in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker, obtaining the initial voting value of the virtual speakers in the current frame. Then, it adjusts the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame based on the final voting value of the representative virtual speaker in the previous frame, thus obtaining the final voting value of the virtual speakers in the current frame. The representative virtual speaker in the previous frame is the virtual speaker used by encoder 113 when encoding the previous frame. The specific method for adjusting the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame can be found below. Figure 6 S620 to S630, and Figure 8 The explanation of S810 to S840 in the document.
[0122] In some embodiments, if the current frame is the first frame in the original audio, encoder 113 executes steps S510 to S520. If the current frame is any frame above the second frame in the original audio, encoder 113 may first determine whether to reuse the representative virtual speaker of the previous frame to encode the current frame or determine whether to perform a virtual speaker search, to ensure the continuity of orientation between consecutive frames and reduce encoding complexity. Embodiments of this application may also include step S540.
[0123] S540 and encoder 113 determine whether to perform a virtual speaker search based on the representative virtual speaker in the previous frame and the current frame.
[0124] If encoder 113 determines to perform a virtual speaker search, it executes steps S510 to S530. Optionally, encoder 113 may first execute step S510, which involves encoder 113 obtaining the representative coefficients of the current frame. Encoder 113 then determines whether to perform a virtual speaker search based on the representative coefficients of the current frame and the coefficients of the representative virtual speakers in previous frames. If encoder 113 determines to perform a virtual speaker search, it then executes steps S520 to S530.
[0125] If encoder 113 determines that it will not perform a virtual speaker search, execute S550.
[0126] S550, encoder 113 determines the representative virtual speaker of the previous frame to encode the current frame.
[0127] Encoder 113 multiplexes the virtual speaker of the previous frame and the current frame to generate a virtual speaker signal, encodes the virtual speaker signal to obtain a bit stream, and sends the bit stream to the destination device 120, i.e., executes S450 and S460.
[0128] For specific methods on determining whether to perform a virtual speaker search, please refer to the following. Figure 9 Explanation of S650 to S680.
[0129] S440 and source device 110 generate a virtual speaker signal based on the current frame of the three-dimensional audio signal and the representative virtual speaker of the current frame.
[0130] The source device 110 generates a virtual speaker signal based on the coefficients of the current frame and the coefficients representing the virtual speaker in the current frame. Specific methods for generating the virtual speaker signal can be found in existing technologies and the description of the virtual speaker signal generation unit 350 in the above embodiments.
[0131] S450 and source device 110 encode the virtual speaker signal to obtain a bitstream.
[0132] The source device 110 can perform encoding operations such as transformation or quantization on the virtual speaker signal to generate a bitstream, thereby achieving the purpose of data compression of the three-dimensional audio signal to be encoded. Specific methods for generating the bitstream can be found in existing technologies and the description of the encoding unit 360 in the above embodiments.
[0133] S460, source device 110 sends a bitstream to destination device 120.
[0134] Source device 110 can send the original audio bitstream to destination device 120 after encoding the entire original audio. Alternatively, source device 110 can encode the three-dimensional audio signal in real time, frame by frame, and send the bitstream of each frame after encoding. Specific methods for sending the bitstream can be found in existing technologies and the descriptions of communication interfaces 114 and 124 in the above embodiments.
[0135] S470, the destination device 120 decodes the bitstream sent by the source device 110, reconstructs the three-dimensional audio signal, and obtains the reconstructed three-dimensional audio signal.
[0136] After receiving the bitstream, the target device 120 decodes it to obtain the virtual speaker signal. Then, based on the candidate virtual speaker set and the virtual speaker signal, it reconstructs the three-dimensional audio signal to obtain the reconstructed three-dimensional audio signal. The target device 120 plays back the reconstructed three-dimensional audio signal. Alternatively, the target device 120 transmits the reconstructed three-dimensional audio signal to other playback devices, which then play it, making the "immersive" sound effect in places like cinemas, concert halls, or virtual scenes even more realistic.
[0137] To increase the continuity of orientation between consecutive frames and overcome the problem of large differences in the results of virtual speakers selected in consecutive frames, encoder 113 adjusts the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame based on the final voting value of the representative virtual speaker in the previous frame, thus obtaining the final voting value of the virtual speaker for the current frame. Figure 6The diagram shown is a flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application. Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Figure 6 The method described is... Figure 5 The specific operational procedures included in S530 are explained. For example... Figure 6 As shown, the method includes the following steps.
[0138] S610, encoder 113 acquires the first number of initial vote values of the current frame of the three-dimensional audio signal.
[0139] The encoder 113 can use the representative coefficient of the current frame to vote for each virtual speaker in the candidate virtual speaker set, obtain the initial voting value of the virtual speaker in the current frame, and select the representative virtual speaker of the current frame based on the voting value, thereby reducing the computational complexity of virtual speaker search and alleviating the computational burden of the encoder.
[0140] Figure 7 This is a flowchart illustrating another three-dimensional audio signal encoding method provided in an embodiment of this application. Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Figure 7 The method described is... Figure 5 The specific operational procedures included in S510 and S520 are described. For example... Figure 7 As shown, the method includes the following steps.
[0141] S6101, encoder 113 acquires the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients.
[0142] Assuming the 3D audio signal is a HOA signal, encoder 113 can sample the current frame of the HOA signal to obtain L·(N+1). 2 The number of sampling points yields the fourth set of coefficients. N represents the order of the HOA signal. For example, assuming the duration of the current frame of the HOA signal is 20 milliseconds, encoder 113 samples the current frame at a frequency of 48 kHz, obtaining 960·(N+1) coefficients in the time domain. 2 Each sampling point can also be called a time-domain coefficient.
[0143] The frequency domain coefficients of the current frame of the 3D audio signal can be obtained by time-frequency transformation based on the time domain coefficients of the current frame of the 3D audio signal. The method of time-domain to frequency-domain transformation is not limited. For example, a modified discrete cosine transform (MDCT) can be used to obtain 960·(N+1) in the frequency domain. 2 Frequency domain coefficients. Frequency domain coefficients can also be called spectral coefficients or frequency points.
[0144] The frequency domain eigenvalues of the sampling points satisfy p(j) = norm(x(j)), where j = 1, 2, ..., L, L represents the number of sampling times, x represents the frequency domain coefficients of the current frame of the three-dimensional audio signal, such as MDCT coefficients, norm is used to calculate the L2 norm, and x(j) represents the (N+1)-th sampling time. 2 Frequency domain coefficients of each sampling point.
[0145] S6102 and encoder 113 select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients.
[0146] Encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least one sub-band. Specifically, dividing the spectral range indicated by the fourth number of coefficients into one sub-band means that the spectral range of this sub-band is equal to the spectral range indicated by the fourth number of coefficients, which is equivalent to encoder 113 not dividing the spectral range indicated by the fourth number of coefficients.
[0147] If encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two frequency band sub-bands, in one case, encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two sub-bands, each of the at least two sub-bands containing the same number of coefficients.
[0148] In another scenario, encoder 113 unequally divides the spectral range indicated by the fourth number of coefficients, such that at least two sub-bands contain different numbers of coefficients, or each of the at least two sub-bands contains a different number of coefficients. For example, encoder 113 can unequally divide the spectral range indicated by the fourth number of coefficients based on a low-frequency range, a mid-frequency range, and a high-frequency range, such that each spectral range includes at least one sub-band. Each sub-band in the at least one low-frequency range contains the same number of coefficients. Each sub-band in the at least one mid-frequency range contains the same number of coefficients. Each sub-band in the at least one high-frequency range contains the same number of coefficients. Sub-bands in the three spectral ranges (low-frequency, mid-frequency, and high-frequency) may contain different numbers of coefficients.
[0149] Furthermore, encoder 113 selects representative coefficients from at least one sub-band included in the spectral range indicated by the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients, thus obtaining a third number of representative coefficients. The third number is less than the fourth number, and the fourth number of coefficients includes the third number of representative coefficients.
[0150] For example, encoder 113 selects Z representative coefficients from each sub-band according to the descending order of the frequency domain characteristic values of the coefficients in at least one sub-band included in the spectral range indicated by the fourth number of coefficients, and combines the Z representative coefficients in at least one sub-band to obtain a third number of representative coefficients, where Z is a positive integer.
[0151] For example, when at least one sub-band includes at least two sub-bands, encoder 113 determines the weight of each sub-band based on the frequency domain feature values of the first candidate coefficients within each of the at least two sub-bands; and adjusts the frequency domain feature values of the second candidate coefficients within each sub-band according to their respective weights, obtaining the adjusted frequency domain feature values of the second candidate coefficients within each sub-band, where the first and second candidate coefficients are partial coefficients within the sub-band. Encoder 113 determines a third number of representative coefficients based on the adjusted frequency domain feature values of the second candidate coefficients within the at least two sub-bands, and the frequency domain feature values of the coefficients within the at least two sub-bands excluding the second candidate coefficients.
[0152] Since the encoder selects a portion of the coefficients from all coefficients in the current frame as representative coefficients, and uses a smaller number of representative coefficients to replace all coefficients in the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder searching for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden on the encoder.
[0153] S6103, encoder 113 determines the first number of virtual speakers and the first number of voting values based on the third number of representative coefficients of the current frame, the candidate virtual speaker set and the number of voting rounds.
[0154] The voting rounds limit the number of times a virtual speaker is voted on. The voting rounds are integers greater than or equal to 1, and are less than or equal to the number of virtual speakers included in the candidate virtual speaker set, and less than or equal to the number of virtual speaker signals transmitted by the encoder. For example, the candidate virtual speaker set may include a fifth number of virtual speakers, which includes a first number of virtual speakers, where the first number is less than or equal to the fifth number, and the voting rounds are integers greater than or equal to 1, and less than or equal to the fifth number. A virtual speaker signal also refers to the transmission channel representing the virtual speaker in the current frame corresponding to the current frame. Typically, the number of virtual speaker signals is less than or equal to the number of virtual speakers.
[0155] In one possible implementation, the number of voting rounds can be pre-configured or determined based on the encoder's computing power. For example, the number of voting rounds can be determined based on the encoder's encoding rate and / or the encoding application scenario.
[0156] In another possible implementation, the number of voting rounds is determined based on the number of directional sound sources in the current frame. For example, when there are 2 directional sound sources in the sound field, the number of voting rounds is set to 2.
[0157] This application provides three possible implementations for determining a first number of virtual speakers and a first number of voting values. The three methods are described in detail below.
[0158] In the first possible implementation, the number of voting rounds is equal to 1. After the encoder 113 samples multiple representative coefficients, it obtains the voting value of each representative coefficient of the current frame for all virtual speakers in the candidate virtual speaker set, and accumulates the voting values of virtual speakers with the same number to obtain a first number of virtual speakers and a first number of voting values. Understandably, the candidate virtual speaker set includes a first number of virtual speakers. The first number is equal to the number of virtual speakers included in the candidate virtual speaker set. Assuming the candidate virtual speaker set includes a fifth number of virtual speakers, then the first number is equal to the fifth number. The first number of voting values includes the voting values of all virtual speakers in the candidate virtual speaker set. The encoder 113 can use the first number of voting values as the initial voting values for the first number of virtual speakers in the current frame and execute steps S620 to S640.
[0159] In this system, each virtual speaker corresponds one-to-one with a voting value. For example, a first number of virtual speakers includes a first virtual speaker, and a first number of voting values includes the voting value of the first virtual speaker; the first virtual speaker corresponds to the voting value of the first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of using the first virtual speaker when encoding the current frame. Priority can also be described as preference; that is, the voting value of the first virtual speaker is used to characterize the preference for using the first virtual speaker when encoding the current frame. Understandably, the larger the voting value of the first virtual speaker, the higher the priority or preference of the first virtual speaker. Compared to virtual speakers in the candidate virtual speaker set with a smaller voting value than the first virtual speaker, the encoder 113 is more inclined to select the first virtual speaker to encode the current frame.
[0160] In the second possible implementation, the difference from the first possible implementation is that after the encoder 113 obtains the voting values of each representative coefficient for all virtual speakers in the candidate virtual speaker set in the current frame, it selects a portion of the voting values from each representative coefficient for all virtual speakers in the candidate virtual speaker set, and accumulates the voting values of the virtual speakers with the same number corresponding to the portion of the voting values, to obtain a first number of virtual speakers and a first number of voting values. Understandably, the candidate virtual speaker set includes the first number of virtual speakers. The first number is less than or equal to the number of virtual speakers included in the candidate virtual speaker set. The first number of voting values includes the voting values of a portion of the virtual speakers included in the candidate virtual speaker set, or the first number of voting values includes the voting values of all the virtual speakers included in the candidate virtual speaker set.
[0161] In the third possible implementation, the difference from the second possible implementation is that the number of voting rounds is an integer greater than or equal to 2. For each representative coefficient of the current frame, the encoder 113 performs at least two rounds of voting on all virtual speakers in the candidate virtual speaker set, selecting the virtual speaker with the largest vote value in each round. After performing at least two rounds of voting on all virtual speakers for each representative coefficient of the current frame, the vote values of virtual speakers with the same number are accumulated to obtain a first number of virtual speakers and a first number of vote values.
[0162] S620, encoder 113 obtains the seventh number of final votes of the current frame corresponding to the seventh number of virtual speakers based on the first number of initial votes of the current frame and the sixth number of final votes of the previous frame.
[0163] The encoder 113 can use the method described in S610 above to determine a first number of virtual speakers and a first number of voting values based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the number of voting rounds, and then use the first number of voting values as the initial voting values of the first number of virtual speakers for the current frame.
[0164] Each virtual speaker corresponds one-to-one with the initial voting value of the current frame; that is, one virtual speaker corresponds to one initial voting value of the current frame. For example, a first number of virtual speakers includes a first virtual speaker, and a first number of initial voting values of the current frame include the initial voting value of the first virtual speaker. The first virtual speaker corresponds to the initial voting value of the first virtual speaker in the current frame. The initial voting value of the first virtual speaker in the current frame is used to characterize the priority of using the first virtual speaker when encoding the current frame.
[0165] The sixth set of virtual speakers can be the representative virtual speakers of the preceding frames used by encoder 113 to encode the preceding frames of the three-dimensional audio signal. In S650, when encoder 113 obtains the first correlation between the current frame of the three-dimensional audio signal and the set of representative virtual speakers of the preceding frames, the set of representative virtual speakers of the preceding frames includes the sixth set of virtual speakers.
[0166] Specifically, encoder 113 updates the initial voting values of the first number of current frames based on the final voting values of the sixth number of prior frames. That is, encoder 113 calculates the sum of the initial voting values of the current frames and the final voting values of the prior frames for the first number of virtual speakers and the virtual speakers with the same number among the sixth number of virtual speakers, and obtains the final voting values of the seventh number of current frames corresponding to the seventh number of virtual speakers.
[0167] In the first possible scenario, the first number of virtual speakers includes the sixth number of virtual speakers, and the first number equals the sixth number. The numbering of the first number of virtual speakers is the same as the numbering of the sixth number of virtual speakers. Understandably, the first number of virtual speakers acquired by encoder 113 is the sixth number of virtual speakers, and the final vote value of the sixth number of virtual speakers in the previous frame is the final vote value of the first number of virtual speakers in the previous frame. Encoder 113 can use the final vote value of the sixth number of virtual speakers in the previous frame to update the initial vote value of the first number of virtual speakers in the current frame. Therefore, the seventh number of virtual speakers is also the first number of virtual speakers, and the final vote value of the seventh number of virtual speakers in the current frame is the sum of the final vote value of the first number of virtual speakers in the previous frame and the initial vote value of the first number of virtual speakers in the current frame.
[0168] For example, suppose the sixth set of virtual speakers includes the first virtual speaker, the first set of virtual speakers includes the first virtual speaker, and neither the sixth set of virtual speakers nor the first set of virtual speakers includes any other virtual speakers. Encoder 113 can update the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame, to obtain the final voting value of the first virtual speaker in the current frame. The final voting value of the first virtual speaker in the current frame is the sum of the final voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame.
[0169] In the second possible scenario, the first number of virtual speakers includes the sixth number of virtual speakers, and the first number is greater than the sixth number. Understandably, the first number of virtual speakers also includes other virtual speakers besides the sixth number. Encoder 113 can update the initial voting value of the virtual speakers in the first number of virtual speakers that have the same number as the sixth number of virtual speakers in the previous frame using the final voting value of the sixth number of virtual speakers in the previous frame. Therefore, the seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number is equal to the first number, and the number of the seventh number of virtual speakers is the same as the number of the first number of virtual speakers. The final voting value of the seventh number of virtual speakers in the current frame includes the final voting value of the virtual speakers in the first number of virtual speakers that have the same number as the sixth number of virtual speakers in the previous frame, and the final voting value of the virtual speakers in the first number of virtual speakers that have different numbers than the sixth number of virtual speakers in the previous frame.
[0170] The final vote value for the current frame of the virtual speaker with the same number as the sixth virtual speaker in the first set of virtual speakers is the sum of the final vote value of the sixth virtual speaker in the previous frame and the initial vote value of the first set of virtual speakers in the current frame. The final vote value for the current frame of the virtual speaker with a different number than the sixth virtual speaker in the first set of virtual speakers is the initial vote value of the current frame of the virtual speaker with a different number than the sixth virtual speaker in the first set of virtual speakers.
[0171] For example, assuming a first set of virtual speakers includes a first virtual speaker and a second virtual speaker, a sixth set of virtual speakers includes the first virtual speaker, and the sixth set of virtual speakers does not include the second virtual speaker, then the final vote value of the second virtual speaker in the current frame is equal to its initial vote value in the current frame. Encoder 113 can update the initial vote value of the first virtual speaker in the current frame based on the final vote value of the first virtual speaker in a previous frame to obtain the final vote value of the first virtual speaker in the current frame. The final vote value of the first virtual speaker in the current frame is the sum of the final vote value of the first virtual speaker in a previous frame and the initial vote value of the first virtual speaker in the current frame.
[0172] In a third possible scenario, the first number of virtual speakers includes some of the virtual speakers in the sixth number of virtual speakers, and the sixth number of virtual speakers also includes other virtual speakers with numbers different from those in the first number of virtual speakers. Therefore, the seventh number of virtual speakers includes the first number of virtual speakers, as well as the virtual speakers in the sixth number of virtual speakers with numbers different from those in the first number of virtual speakers. The seventh number of final votes for the current frame includes the final votes for the first number of virtual speakers in the current frame, and the final votes for the virtual speakers in the sixth number of virtual speakers with numbers different from those in the first number of virtual speakers in the current frame.
[0173] The final vote value of the first set of virtual speakers in the current frame includes the final vote value of the virtual speakers in the current frame that have the same number as the sixth set of virtual speakers. Optionally, the final vote value of the first set of virtual speakers in the current frame may also include the final vote value of the virtual speakers in the current frame that have different numbers from the sixth set of virtual speakers.
[0174] The final vote value of the virtual speaker in the current frame that has a different number than the virtual speaker in the first number of virtual speakers is the final vote value of the virtual speaker in the previous frame that has a different number than the virtual speaker in the first number of virtual speakers.
[0175] For example, suppose the sixth set of virtual speakers includes the first and third virtual speakers, the first set of virtual speakers includes the first virtual speaker, and the first set of virtual speakers does not include the third virtual speaker. Then, the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame. Encoder 113 can update the initial vote value of the first virtual speaker in the current frame based on the final vote value of the first virtual speaker in the previous frame to obtain the final vote value of the first virtual speaker in the current frame. The final vote value of the first virtual speaker in the current frame is the sum of the final vote value of the first virtual speaker in the previous frame and the initial vote value of the first virtual speaker in the current frame.
[0176] In some of the embodiments, such as Figure 8 The diagram shown is a flowchart illustrating a method for updating the initial voting value of a virtual speaker in the current frame, provided in an embodiment of this application.
[0177] S810, encoder 113 adjusts the final voting value of the first virtual speaker in the previous frame according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame.
[0178] The first adjustment parameter is determined based on at least one of the following: the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type. The voting value of the first virtual loudspeaker after adjustment in the previous frame satisfies the following formula (6).
[0179] VOTE_f g ′=VOTE_f g ·w1·w2·w3 Formula (6)
[0180] Among them, VOTE_f g ' represents the set of vote values after frame adjustment, VOTE_f g `g` represents the set of final votes for the preceding frames, and `w1` represents a parameter related to the coding rate, `w2` represents a parameter related to the frame type, and `w3` represents a parameter related to the number of directional sound sources. Frame types include transient frames and non-transient frames.
[0181] For example, if the encoding rate is less than or equal to 128kbps, w1 = 1; if the encoding rate is greater than 128kbps, w1 = 0. If the preceding frame is a transient frame, w2 = 1; if the preceding frame is a non-transient frame, w2 = 0. If the number of directional sound sources is greater than the preset number of virtual speaker signals, w3 = 0.8; if the number of directional sound sources is less than or equal to the preset number of virtual speaker signals, w3 = 0.5.
[0182] S820, encoder 113 updates the initial voting value of the first virtual speaker in the current frame according to the voting value of the first virtual speaker after adjustment in the previous frame, and obtains the final voting value of the first virtual speaker in the current frame.
[0183] The final vote value of the first virtual speaker in the current frame is the sum of the adjusted vote value of the first virtual speaker in the previous frame and the initial vote value of the first virtual speaker in the current frame. The final vote value of the first virtual speaker in the current frame satisfies the following formula (7).
[0184] VOTE_M g =VOTE_f g ′+VOTE g Formula (7)
[0185] Among them, VOTE_M g VOTE_f represents the final set of votes for the current frame. g ' represents the set of vote values after frame adjustment, VOTE g This represents the initial set of vote values for the current frame.
[0186] Optionally, the encoder 113 updates the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame, specifically including the following steps.
[0187] S830 and encoder 113 adjust the initial voting value of the first virtual speaker in the current frame according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame.
[0188] The adjusted voting value of the first virtual speaker in the current frame satisfies the following formula (8).
[0189] VOTE g ′=VOTE g ·w4 formula (8)
[0190] Among them, VOTE g ' represents the set of adjusted voting values for the current frame, and w4 represents the second adjustment parameter. For example, if norm(VOTE) g >norm(VOTE_f g ′), Understandably, if the initial voting value of the current frame is greater than the voting value after adjustment in the previous frame, w4 is used to represent amplifying the voting value after adjustment in the previous frame.
[0191] If norm(VOTE) g )≤norm(VOTE_f gw4 = 1. Understandably, if the initial vote value of the current frame is less than or equal to the vote value after the previous frame adjustment, there is no need to use w4 to represent the amplification of the vote value after the previous frame adjustment.
[0192] The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.
[0193] S840, encoder 113 updates the current frame adjusted voting value of the first virtual speaker according to the previous frame adjusted voting value of the first virtual speaker, and obtains the current frame final voting value of the first virtual speaker.
[0194] The final vote value of the first virtual speaker in the current frame is the sum of the adjusted vote value of the first virtual speaker in the previous frame and the adjusted vote value of the first virtual speaker in the current frame. The final vote value of the first virtual speaker in the current frame satisfies the following formula (9).
[0195] VOTE_M g =VOTE_f g ′+VOTE g ′ Formula (9)
[0196] Among them, VOTE_M g VOTE_f represents the final set of votes for the current frame. g ' represents the set of vote values after frame adjustment, VOTE g ′ represents the set of voting values after adjustment in the current frame.
[0197] S630 and encoder 113 select representative virtual speakers from the seventh number of virtual speakers for the second number of current frames based on the final voting values of the seventh number of current frames.
[0198] The encoder 113 selects a representative virtual speaker for a second number of current frames from the seventh number of virtual speakers based on the final voting value of the seventh number of current frames, and the final voting value of the representative virtual speaker for the second number of current frames is greater than a preset threshold.
[0199] The encoder 113 can also select representative virtual speakers for the second number of current frames from the seventh number of virtual speakers based on the final voting values of the seventh number of current frames. For example, the final voting values of the second number of current frames can be determined from the final voting values of the seventh number of current frames in descending order, and the virtual speakers corresponding to the final voting values of the second number of current frames can be used as representative virtual speakers for the second number of current frames.
[0200] Optionally, if the voting values of virtual speakers with different numbers among the seventh number of virtual speakers are the same, and the voting values of the virtual speakers with different numbers are greater than a preset threshold, then the encoder 113 can use the virtual speakers with different numbers as the representative virtual speakers of the current frame.
[0201] It should be noted that the second number is less than the seventh number. The seventh number of virtual speakers includes the second number of representative virtual speakers for the current frame. The second number can be preset, or it can be determined based on the number of sound sources in the sound field of the current frame. For example, the second number can be directly equal to the number of sound sources in the sound field of the current frame, or it can be obtained by processing the number of sound sources in the sound field of the current frame according to a preset algorithm, and using the processed number as the second number. The preset algorithm can be designed as needed. For example, the preset algorithm can be: second number = number of sound sources in the sound field of the current frame + 1, or second number = number of sound sources in the sound field of the current frame - 1, etc.
[0202] In addition, before encoding the next frame of the current frame, if the encoder 113 determines to reuse the representative virtual speakers of the previous frame to encode the next frame, the encoder 113 can use the second number of representative virtual speakers of the current frame as the second number of representative virtual speakers of the previous frame, and use the second number of representative virtual speakers of the previous frame to encode the next frame of the current frame.
[0203] S640 and encoder 113 encode the current frame according to the second number of representative virtual speakers of the current frame to obtain the bit stream.
[0204] Encoder 113 generates a virtual speaker signal based on the representative virtual speakers of the second number of current frames and the current frame; and encodes the virtual speaker signal to obtain a bitstream.
[0205] During the virtual speaker search process, the positions of the real sound sources and virtual speakers may not coincide, leading to a one-to-one correspondence between virtual speakers and real sound sources. Furthermore, in complex real-world scenarios, virtual speakers may fail to represent independent sound sources within the sound field. This can result in frequent jumps between frames, significantly impacting the listener's auditory experience and causing noticeable noise in the decoded and reconstructed 3D audio signal. The virtual speaker selection method provided in this application inherits the representative virtual speaker from previous frames. Specifically, for virtual speakers with the same number, the initial voting value of the current frame is adjusted using the final voting value of the previous frame. This makes the encoder more inclined to select the representative virtual speaker from the previous frame, thereby enhancing the continuity of orientation between frames. Additionally, adjusting parameters ensures that the final voting value of the previous frame is not inherited too far back, preventing the algorithm from being unable to adapt to scenarios involving sound source movement and other changes in the sound field.
[0206] Furthermore, this application also provides a method for selecting virtual speakers. The encoder can first determine whether it can reuse the representative virtual speaker set of the previous frame to encode the current frame. If the encoder reuses the representative virtual speaker set of the previous frame to encode the current frame, it avoids performing the virtual speaker search process again, effectively reducing the computational complexity of the encoder searching for virtual speakers. Therefore, it reduces the computational complexity of compressing and encoding 3D audio signals and alleviates the computational burden on the encoder. If the encoder cannot reuse the representative virtual speaker set of the previous frame to encode the current frame, the encoder selects representative coefficients and uses the representative coefficients of the current frame to vote on each virtual speaker in the candidate virtual speaker set. Based on the voting value, the representative virtual speaker of the current frame is selected, thereby achieving the purpose of reducing the computational complexity of compressing and encoding 3D audio signals and alleviating the computational burden on the encoder. Figure 9 This is a flowchart illustrating a method for selecting virtual speakers according to an embodiment of this application. Before encoder 113 acquires the initial voting values of the first number of virtual speakers and the current frame corresponding to the current frame of the three-dimensional audio signal, i.e., before S610, as follows... Figure 9 As shown, the method includes the following steps.
[0207] S650, encoder 113 acquires the first correlation between the current frame of the three-dimensional audio signal and the representative set of virtual speakers in the previous frame.
[0208] The sixth set of representative virtual loudspeakers in the preceding frame contains a number of virtual loudspeakers, which are the representative virtual loudspeakers of the preceding frame used to encode the 3D audio signal. A first relevance is used to characterize the priority of reusing the representative virtual loudspeaker set of the preceding frame when encoding the current frame. Priority can also be described as preference; that is, the first relevance is used to determine whether to reuse the representative virtual loudspeaker set of the preceding frame when encoding the current frame. Understandably, the higher the first relevance of the representative virtual loudspeaker set of the preceding frame, the higher the priority or preference of the representative virtual loudspeaker set of the preceding frame, and the more likely the encoder 113 is to select the representative virtual loudspeakers of the preceding frame for encoding the current frame.
[0209] S660 and encoder 113 determine whether the first correlation degree meets the reuse condition.
[0210] If the first relevance does not meet the reuse condition, it means that the encoder 113 is more inclined to perform virtual speaker search. The current frame is encoded according to the representative virtual speaker of the current frame. S610 is executed, and the encoder 113 obtains the first number of initial voting values of the first number of virtual speakers and the current frame of the three-dimensional audio signal.
[0211] Optionally, encoder 113 may select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients, and then use the largest representative coefficient among the third number of representative coefficients as the coefficient of the current frame for obtaining the first correlation. Then encoder 113 obtains the first correlation between the largest representative coefficient among the third number of representative coefficients of the current frame and the set of representative virtual speakers in the previous frame. If the first correlation does not meet the reuse condition, S6103 is executed, that is, encoder 113 selects a second number of representative virtual speakers of the current frame from the first number of virtual speakers based on the first number of voting values.
[0212] If the first relevance satisfies the reuse condition, it means that encoder 113 is more inclined to select the representative virtual speaker of the previous frame to encode the current frame, and encoder 113 executes S670 and S680.
[0213] S670, encoder 113 generates a virtual speaker signal based on the representative set of virtual speakers in the previous frame and the current frame.
[0214] The S680 and encoder 113 encode the virtual speaker signal to obtain a bitstream.
[0215] The method for selecting virtual speakers provided in this application uses the representative coefficient of the current frame and the correlation between the representative virtual speakers of the previous frame to determine whether to perform a virtual speaker search. While ensuring the accuracy of the selection of the correlation of the representative virtual speakers of the current frame, it effectively reduces the complexity of the encoding end.
[0216] It is understood that, in order to achieve the functions in the above embodiments, the encoder includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0217] The above text combines Figures 1 to 9 This document describes in detail the three-dimensional audio signal encoding method provided according to this embodiment. The following will combine... Figure 10 and Figure 11 This describes the three-dimensional audio signal encoding apparatus and encoder provided according to this embodiment.
[0218] Figure 10This is a schematic diagram of a possible three-dimensional audio signal encoding device provided in this embodiment. These three-dimensional audio signal encoding devices can be used to implement the function of encoding three-dimensional audio signals in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In this embodiment, the three-dimensional audio signal encoding device can be as follows: Figure 1 The encoder 113 shown, or as... Figure 3 The encoder 300 shown can also be a module (such as a chip) applied to terminal devices or servers.
[0219] like Figure 10 As shown, the three-dimensional audio signal encoding device 1000 includes a communication module 1010, a coefficient selection module 1020, a virtual speaker selection module 1030, an encoding module 1040, and a storage module 1050. The three-dimensional audio signal encoding device 1000 is used to implement the above-mentioned... Figures 6 to 9 The function of encoder 113 in the method embodiment shown.
[0220] The communication module 1010 is used to acquire the current frame of the three-dimensional audio signal. Optionally, the communication module 1010 can also receive the current frame of the three-dimensional audio signal acquired by other devices; or acquire the current frame of the three-dimensional audio signal from the storage module 1050. The current frame of the three-dimensional audio signal is a HOA signal; the frequency domain characteristic values of the coefficients are determined based on the coefficients of the HOA signal.
[0221] The virtual speaker selection module 1030 is used to obtain a first number of initial voting values of the current frame of the three-dimensional audio signal. The first number of virtual speakers correspond one-to-one with the initial voting values of the current frame. The first number of virtual speakers includes a first virtual speaker. The initial voting value of the first virtual speaker in the current frame is used to characterize the priority of using the first virtual speaker when encoding the current frame.
[0222] The virtual speaker selection module 1030 is further configured to obtain the seventh number of final voting values of the current frame corresponding to the seventh number of virtual speakers based on the first number of initial voting values of the current frame and the sixth number of final voting values of the previous frame. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The sixth number of virtual speakers corresponds one-to-one with the sixth number of final voting values of the previous frame. The sixth number of virtual speakers are virtual speakers used when encoding the previous frames of the three-dimensional audio signal.
[0223] If the first number of virtual speakers includes the second virtual speaker and the sixth number of virtual speakers does not include the second virtual speaker, then the final vote value of the second virtual speaker in the current frame is equal to the initial vote value of the second virtual speaker in the current frame; or, if the sixth number of virtual speakers includes the third virtual speaker and the first number of virtual speakers does not include the third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.
[0224] When the three-dimensional audio signal encoding device 1000 is used to implement Figures 6 to 9 In the method embodiment shown, when the encoder 113 functions, the virtual speaker selection module 1030 is used to implement the related functions of S610 to S630 and S650 to S680.
[0225] For example, when the virtual speaker selection module 1030 updates the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame, it is specifically used to: adjust the final voting value of the first virtual speaker in the previous frame based on the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame; and update the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame.
[0226] For example, when the virtual speaker selection module 1030 updates the initial voting value of the first virtual speaker in the current frame according to the voting value of the first virtual speaker after adjustment in the previous frame, it is specifically used to: adjust the initial voting value of the first virtual speaker in the current frame according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame; and update the adjusted voting value of the first virtual speaker in the current frame according to the voting value of the first virtual speaker after adjustment in the previous frame.
[0227] The first adjustment parameter is determined based on at least one of the following: the number of directional sound sources in the previous frame, the coding rate for encoding the current frame, and the frame type.
[0228] The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.
[0229] When the three-dimensional audio signal encoding device 1000 is used to implement Figure 7 In the method embodiment shown, when the encoder 113 functions, the coefficient selection module 1020 is used to implement the related functions of S6101 and S6102. Specifically, when the coefficient selection module 1020 obtains the third number of representative coefficients of the current frame, it is specifically used to: obtain the fourth number of coefficients of the current frame and the frequency domain feature values of the fourth number of coefficients; and select the third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients, wherein the third number is less than the fourth number.
[0230] The encoding module 1140 is used to encode the current frame according to the second number of representative virtual speakers of the current frame to obtain a bit stream.
[0231] When the three-dimensional audio signal encoding device 1000 is used to implement Figures 6 to 9 In the method embodiment shown, when encoder 113 functions, encoding module 1140 is used to implement the related functions of S630. For example, encoding module 1140 is specifically used to generate a virtual speaker signal based on a second number of representative virtual speakers of the current frame and the current frame; and to encode the virtual speaker signal to obtain a bitstream.
[0232] The storage module 1050 is used to store coefficients related to the three-dimensional audio signal, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and selected coefficients and virtual speakers, so that the encoding module 1040 can encode the current frame to obtain a bitstream and transmit the bitstream to the decoder.
[0233] It should be understood that the three-dimensional audio signal encoding device 1000 of this application embodiment can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can also be implemented by software. Figures 6 to 9 In the three-dimensional audio signal encoding method shown, the three-dimensional audio signal encoding device 1000 and its various modules can also be software modules.
[0234] For a more detailed description of the aforementioned communication module 1010, coefficient selection module 1020, virtual speaker selection module 1030, encoding module 1040, and storage module 1050, please refer to [the relevant documentation / reference]. Figures 6 to 9 The relevant descriptions in the method embodiments shown are directly obtained and will not be repeated here.
[0235] Figure 11 This is a schematic diagram of the structure of an encoder 1100 provided in this embodiment. Figure 11 As shown, the encoder 1100 includes a processor 1110, a bus 1120, a memory 1130, and a communication interface 1140.
[0236] It should be understood that in this embodiment, the processor 1110 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0237] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0238] The communication interface 1140 is used to enable communication between the encoder 1100 and external devices or components. In this embodiment, the communication interface 1140 is used to receive three-dimensional audio signals.
[0239] Bus 1120 may include a pathway for transmitting information between the aforementioned components (such as processor 1110 and memory 1130). In addition to a data bus, bus 1120 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 1120 in the figure.
[0240] As an example, encoder 1100 may include multiple processors. A processor may be a multi-CPU processor. Here, "processor" can refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions). Processor 1110 may access coefficients associated with the three-dimensional audio signal stored in memory 1130, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and selected coefficients and virtual speakers, etc.
[0241] It is worth noting that, Figure 11 Taking encoder 1100 as an example, which includes one processor 1110 and one memory 1130, the processor 1110 and the memory 1130 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0242] The memory 1130 may correspond to the storage medium used in the above method embodiment for storing coefficients related to the three-dimensional audio signal, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and information such as selected coefficients and virtual speakers, for example, a disk, such as a mechanical hard disk or a solid-state drive.
[0243] The encoder 1100 described above can be a general-purpose device or a special-purpose device. For example, the encoder 1100 can be an x86 or ARM-based server, or other special-purpose servers, such as a policy control and charging (PCC) server. This application does not limit the type of encoder 1100.
[0244] It should be understood that the encoder 1100 according to this embodiment can correspond to the three-dimensional audio signal encoding device 1100 in this embodiment, and can correspond to the execution according to Figures 6 to 9 The corresponding subject in any of the methods, and the above and other operations and / or functions of each module in the three-dimensional audio signal encoding device 1100 are respectively for implementing Figures 6 to 9 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.
[0245] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a network device or terminal device. Of course, the processor and storage medium can also exist as discrete components in the network device or terminal device.
[0246] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0247] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A three-dimensional audio signal encoding method, characterized in that, include: Obtain a first number of initial voting values for the current frame of the three-dimensional audio signal, wherein the first number of virtual speakers correspond one-to-one with the initial voting values for the current frame, the first number of virtual speakers include a first virtual speaker, and the initial voting value for the current frame of the first virtual speaker is used to characterize the priority of the first virtual speaker; Based on the first number of initial voting values of the current frame and the sixth number of final voting values of the preceding frames, a seventh number of virtual speakers are obtained, corresponding to the seventh number of final voting values of the current frame. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The sixth number of virtual speakers corresponds one-to-one with the sixth number of final voting values of the preceding frames. The sixth number of virtual speakers are used to encode the preceding frames of the three-dimensional audio signal. Based on the final voting values of the seventh number of current frames, a second number of representative virtual speakers for the current frames are selected from the seventh number of virtual speakers, where the second number is less than the seventh number; The current frame is encoded using the representative virtual speakers of the second number of current frames to obtain a bitstream.
2. The method according to claim 1, characterized in that, If the first number of virtual speakers includes the second virtual speaker, and the sixth number of virtual speakers does not include the second virtual speaker, then the final voting value of the second virtual speaker in the current frame is equal to the initial voting value of the second virtual speaker in the current frame; or If the sixth number of virtual speakers includes the third virtual speaker, and the first number of virtual speakers does not include the third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.
3. The method according to claim 1 or 2, characterized in that, If the sixth number of virtual speakers includes the first virtual speaker, obtaining the seventh number of final voting values of the seventh virtual speaker corresponding to the current frame based on the first number of initial voting values of the current frame and the sixth number of prior frame voting values of the sixth number of virtual speakers and the prior frame of the three-dimensional audio signal includes: The initial voting value of the first virtual speaker in the current frame is updated based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.
4. The method according to claim 3, characterized in that, The step of updating the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame includes: The final voting value of the first virtual speaker in the previous frame is adjusted according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame. The initial voting value of the first virtual speaker in the current frame is updated based on the adjusted voting value of the first virtual speaker in the previous frame.
5. The method according to claim 4, characterized in that, The step of updating the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame includes: The initial voting value of the first virtual speaker in the current frame is adjusted according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame. The current frame adjusted voting value of the first virtual speaker is updated based on the previous frame adjusted voting value of the first virtual speaker.
6. The method according to claim 4 or 5, characterized in that, The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type of the current frame.
7. The method according to claim 5, characterized in that, The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.
8. The method according to claim 1 or 2, characterized in that, The second quantity is preset, or the second quantity is determined based on the current frame.
9. The method according to claim 1 or 2, characterized in that, The first number of initial voting values for the current frame of the current frame used to acquire the three-dimensional audio signal include: The first number of virtual speakers and the first number of initial voting values for the current frame are determined based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds. The candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.
10. The method according to claim 9, characterized in that, Before determining the first number of virtual speakers and the first number of initial voting values for the current frame based on the third number of representative coefficients, the candidate virtual speaker set, and the number of voting rounds of the current frame, the method further includes: Obtain the fourth number of coefficients of the current frame, and the frequency domain feature values of the fourth number of coefficients; Based on the frequency domain characteristic values of the fourth number of coefficients, a third number of representative coefficients are selected from the fourth number of coefficients, wherein the third number is less than the fourth number.
11. The method according to claim 10, characterized in that, The method further includes: Obtain the first correlation between the current frame and the representative virtual speaker set of the previous frame, wherein the representative virtual speaker set of the previous frame includes the sixth number of virtual speakers, wherein the sixth number of virtual speakers are the representative virtual speakers of the previous frame used to encode the previous frame, and the first correlation is used to determine whether the representative virtual speaker set of the previous frame is reused when encoding the current frame; If the first correlation does not meet the reuse condition, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients.
12. The method according to claim 1 or 2, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
13. A three-dimensional audio signal encoding device, characterized in that, include: A virtual speaker selection module is used to obtain a first number of initial voting values for the current frame of a three-dimensional audio signal, wherein the first number of virtual speakers corresponds one-to-one with the initial voting values for the current frame, the first number of virtual speakers includes a first virtual speaker, and the initial voting value for the current frame of the first virtual speaker is used to characterize the priority of the first virtual speaker. The virtual speaker selection module is further configured to obtain, based on the first number of initial voting values of the current frames and the sixth number of final voting values of the preceding frames, the seventh number of virtual speakers corresponding to the current frame and the current frame, wherein the seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers, and the sixth number of virtual speakers corresponds one-to-one with the sixth number of final voting values of the preceding frames, and the sixth number of virtual speakers are virtual speakers used when encoding the preceding frames of the three-dimensional audio signal; The virtual speaker selection module is further configured to select representative virtual speakers for a second number of current frames from the seventh number of virtual speakers based on the final voting values of the seventh number of current frames, wherein the second number is less than the seventh number; An encoding module is used to encode the current frame according to the second number of representative virtual speakers of the current frame to obtain a bitstream.
14. The apparatus according to claim 13, characterized in that, If the first number of virtual speakers includes the second virtual speaker, and the sixth number of virtual speakers does not include the second virtual speaker, then the final voting value of the second virtual speaker in the current frame is equal to the initial voting value of the second virtual speaker in the current frame; or If the sixth number of virtual speakers includes the third virtual speaker, and the first number of virtual speakers does not include the third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.
15. The apparatus according to claim 13 or 14, characterized in that, If the sixth number of virtual speakers includes the first virtual speaker, when the virtual speaker selection module obtains the seventh number of virtual speakers and the seventh number of final voting values of the current frame corresponding to the current frame based on the first number of initial voting values of the current frame and the sixth number of previous frame voting values of the sixth number of virtual speakers and the previous frame corresponding to the three-dimensional audio signal, it is specifically used for: The initial voting value of the first virtual speaker in the current frame is updated based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.
16. The apparatus according to claim 15, characterized in that, When the virtual speaker selection module updates the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame, it is specifically used for: The final voting value of the first virtual speaker in the previous frame is adjusted according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame. The initial voting value of the first virtual speaker in the current frame is updated based on the adjusted voting value of the first virtual speaker in the previous frame.
17. The apparatus according to claim 16, characterized in that, When the virtual speaker selection module updates the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame, it is specifically used for: The initial voting value of the first virtual speaker in the current frame is adjusted according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame. The current frame adjusted voting value of the first virtual speaker is updated based on the previous frame adjusted voting value of the first virtual speaker.
18. The apparatus according to claim 16 or 17, characterized in that, The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type of the current frame.
19. The apparatus according to claim 17, characterized in that, The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.
20. The apparatus according to claim 13 or 14, characterized in that, The second quantity is preset, or the second quantity is determined based on the current frame.
21. The apparatus according to claim 13 or 14, characterized in that, When the virtual speaker selection module obtains the first number of initial voting values for the current frame of the three-dimensional audio signal, it is specifically used for: The first number of virtual speakers and the first number of initial voting values for the current frame are determined based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds. The candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.
22. The apparatus according to claim 21, characterized in that, The device also includes a coefficient selection module; The coefficient selection module is used to obtain the fourth number of coefficients of the current frame and the frequency domain feature values of the fourth number of coefficients. The coefficient selection module is further configured to select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients, wherein the third number is less than the fourth number.
23. The apparatus according to claim 22, characterized in that, The virtual speaker selection module is also used for: Obtain a first correlation between the current frame and the representative virtual speaker set of the previous frame, wherein the representative virtual speaker set of the previous frame includes the sixth number of virtual speakers, and the virtual speakers included in the sixth number of virtual speakers are the representative virtual speakers of the previous frame used to encode the previous frame. The first correlation is used to determine whether the representative virtual speaker set of the previous frame is reused when encoding the current frame. If the first correlation does not meet the reuse condition, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients.
24. The apparatus according to claim 13 or 14, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
25. An encoder, characterized in that, The encoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the three-dimensional audio signal encoding method as described in any one of claims 1-12.
26. A system, characterized in that, The system includes an encoder as described in claim 25, and a decoder, wherein the encoder is used to perform the operational steps of the method according to any one of claims 1-12, and the decoder is used to decode the bitstream generated by the encoder.
27. A computer-readable storage medium, characterized in that, Includes computer software instructions; when the computer software instructions are executed in the encoder, the encoder causes the encoder to perform the three-dimensional audio signal encoding method as described in any one of claims 1-12.
28. A computer-readable storage medium, characterized in that, The bitstream obtained by the three-dimensional audio signal encoding method as described in any one of claims 1-12.