Three-dimensional audio signal coding method and apparatus, and encoder

ZA202310451BActive Publication Date: 2026-08-26HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202310451
Authority / Receiving Office
ZA · ZA
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2023-11-09
Publication Date
2026-08-26
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

Existing three-dimensional audio signal coding technology has large differences in virtual speaker selection between frames, resulting in unstable reconstructed sound images, reduced sound quality, high computational complexity, and difficulty in achieving efficient data compression.

Method used

By obtaining the voting values ​​of the current frame and the previous frame in the encoder, adjusting the initial voting value of the virtual speaker, selecting the representative virtual speaker, reducing the frequent jumps of the virtual speaker between frames, and using fewer representative coefficients for voting, reducing calculations complexity.

Benefits of technology

It improves the audio-visual stability and sound quality of the three-dimensional audio signal, reduces the coding complexity and computational burden, and achieves more efficient data compression.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A three-dimensional audio signal encoding method and apparatus, and an encoder (113) are provided, relate to the multimedia field. The methods includes: The encoder (113) obtains a first quantity of current-frame initial vote values for a current frame of a three-dimensional audio signal (S610). Then, the encoder (113) obtains, based on the first quantity of current-frame initial vote values and sixth quantity of previous-frame final vote values, a seventh of current-frame final vote values that are of a seventh quantity of virtual loudspeaker and that correspond to the current frame (S620). Further, the encoder (113) selects a second quantity of current-frame representative virtual loudspeaker from the seventh quantity of virtual loudspeaker based on the seventh quantity of current-frame final vote values (S630). The encoder (113) encodes the current frame based on the second quantity of current-frame representative virtual loudspeakers, to obtain a bitstream (640). In this way, signal directional continuity between frames is enhanced, stability of a spatial image of the reconstructed three-dimensional audio signal is improved, and sound quality of the reconstructed three-dimensional audio signal is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional audio signal encoding method, device and encoder

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on May 17, 2021, with application number 202110536634.9 and application name “Three-dimensional audio signal encoding method, device and encoder”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of multimedia, and in particular to a three-dimensional audio signal encoding method, device, and encoder. Background Art

[0003] With the rapid development of high-performance computers and signal processing technologies, listeners have increasingly higher expectations for voice and audio experiences, and immersive audio can meet these needs. For example, three-dimensional audio technology has been widely used in wireless communications (such as 4G / 5G), voice, virtual reality / augmented reality, and media audio. Three-dimensional audio technology captures, processes, transmits, and renders and plays back real-world sounds and three-dimensional sound field information. This technology imparts a strong sense of space, envelopment, and immersion, giving listeners an extraordinary, "immersive" auditory experience.

[0004] Typically, an acquisition device (such as a microphone) collects a large amount of data to record three-dimensional sound field information, and transmits a three-dimensional audio signal to a playback device (such as a speaker, headphones, etc.) so that the playback device can play the three-dimensional audio. Due to the large amount of data in the three-dimensional sound field information, a large amount of storage space is required to store the data, and the bandwidth requirement for transmitting the three-dimensional audio signal is high. In order to solve the above problems, the three-dimensional audio signal can be compressed, and the compressed data can be stored or transmitted. At present, the encoder first traverses the virtual speakers in the candidate virtual speaker set, and uses the selected virtual speakers to compress the three-dimensional audio signal. However, if the results of the virtual speakers selected in consecutive frames are quite different, the sound image of the reconstructed three-dimensional audio signal will be unstable, which will reduce the sound quality of the reconstructed three-dimensional audio signal.

[0005] Summary of the Invention

[0006] The present application provides a three-dimensional audio signal encoding method, apparatus, and encoder, thereby enhancing the azimuth continuity between frames, improving the stability of the sound image of the reconstructed three-dimensional audio signal, and ensuring the sound quality of the reconstructed three-dimensional audio signal.

[0007] In a first aspect, the present application provides a three-dimensional audio signal encoding method, which can be executed by an encoder and specifically includes the following steps: after the encoder obtains the first number of current frame initial voting values ​​of the current frame of the three-dimensional audio signal, based on the first number of current frame initial voting values ​​and the sixth number of previous frame final voting values ​​corresponding to the sixth number of virtual speakers and the previous frame of the three-dimensional audio signal, obtains the seventh number of current frame final voting values ​​corresponding to the current frame of the seventh number of virtual speakers. The virtual speakers correspond one-to-one to the initial voting values ​​of the current frame, the first number of virtual speakers includes the first virtual speaker, and the current frame initial voting value of the first virtual speaker is used to represent the priority of using the first virtual speaker when encoding the current frame. The seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number of virtual speakers includes the sixth number of virtual speakers. Furthermore, the encoder selects a second number of representative virtual speakers for the current frames from the seventh number of virtual speakers based on the final voting values ​​of the seventh number of current frames. The second number is less than the seventh number, indicating that the representative virtual speakers for the second number of current frames are part of the seventh number of virtual speakers. The current frame is encoded according to the representative virtual speakers for the second number of current frames to obtain a bit stream.

[0008] During the virtual speaker search process, since the position of the real sound source and the position of the virtual speaker do not necessarily coincide, the virtual speaker may not be able to form a one-to-one correspondence with the real sound source. In addition, in actual complex scenarios, a limited number of virtual speaker sets may not be able to represent all sound sources in the sound field. At this time, the virtual speakers searched between frames may frequently jump, which will significantly affect the listener's auditory perception and cause obvious discontinuity and noise in the three-dimensional audio signal after decoding and reconstruction. The method for selecting virtual speakers provided in the embodiment of the present application inherits the representative virtual speakers of the previous frame, that is, for virtual speakers with the same number, the final voting value of the previous frame is used to adjust the initial voting value of the current frame, so that the encoder is more inclined to select the representative virtual speaker of the previous frame, thereby reducing the frequent jumps of virtual speakers between frames, enhancing the continuity of signal orientation between frames, improving the stability of the sound image of the reconstructed three-dimensional audio signal, and ensuring the sound quality of the reconstructed three-dimensional audio signal.

[0009] For example, if the sixth number of virtual speakers includes the first virtual speaker, obtaining the seventh number of current frame final voting values ​​corresponding to the seventh number of virtual speakers and the current frame based on the first number of current frame initial voting values ​​and the sixth number of previous frame voting values ​​corresponding to the sixth number of virtual speakers and the previous frame of the three-dimensional audio signal includes: updating the current frame initial voting value of the first virtual speaker based on the previous frame final voting value of the first virtual speaker to obtain the current frame final voting value of the first virtual speaker.

[0010] In one possible implementation, if the first number of virtual speakers includes the second virtual speaker and the sixth number of virtual speakers does not include the second virtual speaker, the current frame final voting value of the second virtual speaker is equal to the current frame initial voting value of the second virtual speaker; or, if the sixth number of virtual speakers includes the third virtual speaker and the first number of virtual speakers does not include the third virtual speaker, the current frame final voting value of the third virtual speaker is equal to the previous frame final voting value of the third virtual speaker.

[0011] In another possible implementation, updating the initial voting value of the current frame of the first virtual speaker according to the final voting value of the previous frame of the first virtual speaker includes: the encoder adjusts the final voting value of the previous frame of the first virtual speaker according to the first adjustment parameter to obtain the adjusted voting value of the previous frame of the first virtual speaker; and updates the initial voting value of the current frame of the first virtual speaker according to the adjusted voting value of the previous frame of the first virtual speaker.

[0012] The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, the encoding rate used to encode the current frame, and the frame type. The encoder then uses the first adjustment parameter to adjust the final voting value of the first virtual speaker in the previous frame, making it more likely to select the representative virtual speaker in the previous frame. This enhances the continuity of orientation between frames, improves the stability of the sound image of the reconstructed 3D audio signal, and ensures the sound quality of the reconstructed 3D audio signal.

[0013] In another possible implementation, updating the current frame initial voting value of the first virtual speaker according to the adjusted voting value of the first virtual speaker in the previous frame includes: the encoder adjusts the current frame initial voting value of the first virtual speaker according to the second adjustment parameter to obtain the current frame adjusted voting value of the first virtual speaker; and updates the current frame adjusted voting value of the first virtual speaker according to the adjusted voting value of the first virtual speaker in the previous frame.

[0014] The second adjustment parameter is determined based on the adjusted voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame. Thus, the encoder uses the second adjustment parameter to adjust the initial voting value of the first virtual speaker in the current frame, reducing frequent jumps in the initial voting value in the current frame. This makes the encoder more likely to select the representative virtual speaker in the previous frame, thereby enhancing the positional continuity between frames, improving the stability of the sound image of the reconstructed 3D audio signal, and ensuring the sound quality of the reconstructed 3D audio signal.

[0015] The second number is used to characterize the number of representative virtual speakers for the current frame selected by the encoder. A larger second number indicates a larger number of representative virtual speakers for the current frame, and more sound field information of the three-dimensional audio signal; a smaller second number indicates a smaller number of representative virtual speakers for the current frame, and less sound field information of the three-dimensional audio signal. Therefore, the number of representative virtual speakers for the current frame selected by the encoder can be controlled by setting the second number. For example, the second number can be preset, or, as another example, the second number can be determined based on the current frame. For example, the value of the second number can be 1, 2, 4, or 8.

[0016] In another possible implementation, obtaining a first number of current-frame initial voting values ​​corresponding to the first number of virtual speakers and the current frame of the three-dimensional audio signal includes: an encoder determining the first number of virtual speakers and the first number of current-frame initial voting values ​​based on a third number of representative coefficients of the current frame, a set of candidate virtual speakers, and a number of voting rounds. The set of candidate virtual speakers includes a fifth number of virtual speakers, the fifth number of virtual speakers includes the first number of virtual speakers, the first number is less than or equal to the fifth number, the number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.

[0017] At present, in the process of searching for virtual speakers, the encoder uses the result of the correlation calculation between the three-dimensional audio signal to be encoded and the virtual speaker as a measurement index for selecting the virtual speaker. Moreover, if the encoder transmits a virtual speaker for each coefficient, the purpose of efficient data compression cannot be achieved, and a heavy computational burden will be imposed on the encoder. In the method for selecting virtual speakers provided in the embodiment of the present application, the encoder uses a smaller number of representative coefficients instead of all the coefficients of the current frame to vote for each virtual speaker in the candidate virtual speaker set, and selects the representative virtual speaker of the current frame based on the voting value. Furthermore, the encoder uses the representative virtual speaker of the current frame to compress and encode the three-dimensional audio signal to be encoded, which not only effectively improves the compression rate of the compression encoding of the three-dimensional audio signal, but also reduces the computational complexity of the encoder's search for virtual speakers, thereby reducing the computational complexity of the compression encoding of the three-dimensional audio signal and alleviating the computational burden of the encoder.

[0018] In another possible implementation, before determining the first number of virtual speakers and the first number of initial voting values ​​of the current frame based on the third number of representative coefficients of the current frame, the set of candidate virtual speakers and the number of voting rounds, the method also includes: the encoder obtains the fourth number of coefficients of the current frame, and the frequency domain eigenvalues ​​of the fourth number of coefficients; based on the frequency domain eigenvalues ​​of the fourth number of coefficients, selects a third number of representative coefficients from the fourth number of coefficients, the third number being less than the fourth number, indicating that the third number of representative coefficients is part of the fourth number of coefficients.

[0019] The current frame of the three-dimensional audio signal is a higher order ambisonics (HOA) signal; and the frequency domain eigenvalues ​​of the coefficients are determined based on the coefficients of the HOA signal.

[0020] In this way, since the encoder selects part of the coefficients from all the coefficients of the current frame as representative coefficients, and uses a smaller number of representative coefficients instead of all the coefficients of the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder's search for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden of the encoder.

[0021] In addition, the encoder encodes the current frame according to the second number of representative virtual speakers of the current frames to obtain the code stream, including: the encoder generates a virtual speaker signal according to the second number of representative virtual speakers of the current frames and the current frame; and encodes the virtual speaker signal to obtain the code stream.

[0022] In another possible implementation, the method further includes: the encoder obtaining a first correlation between a current frame and a representative virtual speaker set of a previous frame; if the first correlation does not meet a reuse condition, obtaining a fourth number of coefficients of the current frame of the three-dimensional audio signal and frequency-domain eigenvalues ​​of the fourth number of coefficients; the representative virtual speaker set of the previous frame includes a sixth number of virtual speakers, the virtual speakers included in the sixth number of virtual speakers being representative virtual speakers of the previous frame used to encode the previous frame of the three-dimensional audio signal; and the first correlation is used to determine whether to reuse the representative virtual speaker set of the previous frame when encoding the current frame.

[0023] In this way, the encoder can first determine whether the representative virtual speaker set of the previous frame can be reused to encode the current frame. If the encoder reuses the representative virtual speaker set of the previous frame to encode the current frame, the encoder avoids the need to perform the virtual speaker search process again, effectively reducing the computational complexity of the encoder's search for virtual speakers, thereby reducing the computational complexity of compressing and encoding the three-dimensional audio signal and alleviating the computational burden of the encoder. In addition, the frequent jumps of virtual speakers between frames can be reduced, the continuity of the orientation between frames can be enhanced, the stability of the sound image of the reconstructed three-dimensional audio signal can be improved, and the sound quality of the reconstructed three-dimensional audio signal can be ensured. If the encoder cannot reuse the representative virtual speaker set of the previous frame to encode the current frame, the encoder selects a representative coefficient and uses the representative coefficient of the current frame to vote for each virtual speaker in the candidate virtual speaker set. The representative virtual speaker of the current frame is selected based on the voting value, thereby reducing the computational complexity of compressing and encoding the three-dimensional audio signal and alleviating the computational burden of the encoder.

[0024] Optionally, the method further includes: the encoder may further collect a current frame of the 3D audio signal, so as to compress and encode the current frame of the 3D audio signal to obtain a code stream, and transmit the code stream to the decoding end.

[0025] In a second aspect, the present application provides a 3D audio signal encoding apparatus, the apparatus comprising modules for executing the 3D audio signal encoding method of the first aspect or any possible design of the first aspect. For example, the 3D audio signal encoding apparatus comprises a virtual speaker selection module and an encoding module. The virtual speaker selection module is configured to obtain a first number of current frame initial voting values ​​corresponding to a first number of virtual speakers for a current frame of a three-dimensional audio signal, wherein the virtual speakers correspond to the current frame initial voting values ​​in a one-to-one manner, the first number of virtual speakers including the first virtual speaker, and the current frame initial voting value of the first virtual speaker is used to indicate a priority of using the first virtual speaker when encoding the current frame. The virtual speaker selection module is further configured to obtain a seventh number of current frame final voting values ​​corresponding to a seventh number of virtual speakers for the current frame based on the first number of current frame initial voting values ​​and a sixth number of previous frame final voting values ​​corresponding to a sixth number of virtual speakers for a previous frame of the three-dimensional audio signal, wherein the seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number of virtual speakers includes the sixth number of virtual speakers. The virtual speaker selection module is further configured to select a second number of representative virtual speakers for the current frame from the seventh number of virtual speakers based on the seventh number of final voting values ​​for the current frame, wherein the second number is less than the seventh number. The encoding module is configured to encode the current frame based on the second number of representative virtual speakers for the current frame to obtain a bitstream. These modules can perform the corresponding functions in the above-mentioned first aspect method example. Please refer to the detailed description in the method example for details, which will not be repeated here.

[0026] In a third aspect, the present application provides an encoder comprising at least one processor and a memory, wherein the memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, the operating steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation of the first aspect are performed.

[0027] In a fourth aspect, the present application provides a system, comprising an encoder as described in the third aspect, and a decoder, wherein the encoder is configured to execute the operating steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation of the first aspect, and the decoder is configured to decode the bitstream generated by the encoder.

[0028] In a fifth aspect, the present application provides a computer-readable storage medium, comprising: computer software instructions; when the computer software instructions are executed in an encoder, the encoder executes the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0029] In a sixth aspect, the present application provides a computer program product. When the computer program product runs on an encoder, it enables the encoder to perform the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0030] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] FIG1 is a schematic structural diagram of an audio encoding and decoding system provided in an embodiment of the present application;

[0032] FIG2 is a schematic diagram of a scenario of an audio encoding and decoding system provided in an embodiment of the present application;

[0033] FIG3 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;

[0034] FIG4 is a schematic diagram of a flow chart of a three-dimensional audio signal encoding and decoding method provided in an embodiment of the present application;

[0035] FIG5 is a flow chart of a method for selecting a virtual speaker provided in an embodiment of the present application;

[0036] FIG6 is a schematic flow chart of a three-dimensional audio signal encoding method provided in an embodiment of the present application;

[0037] FIG7 is a flow chart of another method for selecting a virtual speaker provided in an embodiment of the present application;

[0038] FIG8 is a flow chart of a method for adjusting a voting value according to an embodiment of the present application;

[0039] FIG9 is a flow chart of another method for selecting a virtual speaker provided in an embodiment of the present application;

[0040] FIG10 is a schematic structural diagram of an encoding device provided by the present application;

[0041] FIG11 is a schematic structural diagram of an encoder provided in this application. DETAILED DESCRIPTION

[0042] In order to make the description of the following embodiments clear and concise, a brief introduction to the relevant technology is first given.

[0043] Sound is a continuous wave produced by the vibration of an object. The object that vibrates and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solid, or liquid), they are detected by the human or animal auditory system.

[0044] The characteristics of sound waves include pitch, intensity, and timbre. Pitch indicates the highness of a sound. Intensity indicates how loud a sound is. Intensity can also be called loudness or volume. The unit of intensity is decibel (dB). Timbre is also called timbre quality.

[0045] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The frequency is the number of times an object vibrates in one second, measured in hertz (Hz). The human ear can detect sounds with frequencies between 20Hz and 20,000Hz.

[0046] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.

[0047] The waveform of the sound wave determines the timbre, which includes square wave, sawtooth wave, sine wave and pulse wave.

[0048] Based on the characteristics of sound waves, sounds can be divided into regular sounds and irregular sounds. Irregular sounds are sounds produced by the irregular vibration of a sound source. For example, irregular sounds can disrupt people's work, study, and rest. Regular sounds are sounds produced by the regular vibration of a sound source. Regular sounds include speech and music. When represented electrically, regular sounds are analog signals that continuously vary in the time-frequency domain. These analog signals are called audio signals. Audio signals are information carriers that carry speech, music, and sound effects.

[0049] Since human hearing has the ability to distinguish the positional distribution of sound sources in space, when listeners hear sounds in space, in addition to being able to feel the pitch, intensity and timbre of the sounds, they can also feel the direction of the sounds.

[0050] As people's attention to and demand for higher quality in their auditory systems grows, three-dimensional audio technology has emerged to enhance the sense of depth, presence, and spatiality. This allows listeners to perceive not only sounds from sources in front, behind, left, and right, but also a sense of being enveloped by the spatial sound field (or "sound field") generated by these sources, with the sound spreading outward in all directions. This creates an immersive sound experience that transports listeners to places like a movie theater or concert hall.

[0051] Three-dimensional audio technology refers to assuming that the space outside the human ear is a system, and the signal received at the eardrum is a three-dimensional audio signal output by filtering the sound emitted by the sound source through the system outside the ear. For example, the system outside the human ear can be defined as the system impulse response h(n), any sound source can be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal described in the embodiment of the present application may refer to a higher order ambisonics (HOA) signal. Three-dimensional audio can also be called three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio or binaural audio, etc.

[0052] As we all know, when sound waves propagate in an ideal medium, the wave number is k = w / c and the angular frequency is w = 2πf, where f is the sound wave frequency and c is the sound speed. The sound pressure p satisfies formula (1), is the Laplace operator.

[0053]

[0054] Assume that the spatial system outside the human ear is a sphere, with the listener at the center of the sphere. Sounds from outside the sphere have a projection on the sphere, filtering out sounds outside the sphere. Assuming that the sound sources are distributed on this sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound source. In other words, three-dimensional audio technology is a method for fitting sound fields. Specifically, equation (1) is solved in a spherical coordinate system. Within the passive spherical region, equation (1) is solved as follows (2).

[0055]

[0056] Where r is the radius of the sphere, θ is the horizontal angle, represents the pitch angle, k represents the wave number, s represents the amplitude of the ideal plane wave, and m represents the order number of the three-dimensional audio signal (or the order number of the HOA signal). represents the spherical Bessel function, which is also called the radial basis function, where the first j represents the imaginary unit, Does not change with angle. express Spherical harmonics of the direction, The spherical harmonics representing the direction of the sound source. The three-dimensional audio signal coefficients satisfy formula (3).

[0057]

[0058] Substituting formula (3) into formula (2), formula (2) can be transformed into formula (4).

[0059]

[0060] in, Represents the Nth-order three-dimensional audio signal coefficients, used to approximate the sound field. A sound field is the area within a medium where sound waves exist. N is an integer greater than or equal to 1. For example, N ranges from 2 to 6. The three-dimensional audio signal coefficients described in the embodiments of the present application may refer to HOA coefficients or ambisonic coefficients.

[0061] A three-dimensional audio signal is an information carrier that carries the spatial position information of the sound source in the sound field and describes the sound field of the listener in space. Formula (4) shows that the sound field can be expanded on a sphere using spherical harmonics, that is, the sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional audio signal can be expressed as the superposition of multiple plane waves, and the sound field can be reconstructed using the three-dimensional audio signal coefficients.

[0062] Compared with 5.1-channel audio signals or 7.1-channel audio signals, the N-order HOA signal has (N+1) 2 If there are multiple channels, the HOA signal includes a large amount of data for describing the spatial information of the sound field. If the acquisition device (such as a microphone) transmits the three-dimensional audio signal to the playback device (such as a speaker), a large bandwidth is consumed. At present, the encoder can use spatial squeezed surround audio coding (S3AC) or directional audio coding (DirAC) to compress and encode the three-dimensional audio signal to obtain a bit stream, and transmit the bit stream to the playback device. The playback device decodes the bit stream, reconstructs the three-dimensional audio signal, and plays the reconstructed three-dimensional audio signal. This reduces the amount of data transmitted to the playback device and the bandwidth occupied. However, the computational complexity of the encoder's compression encoding of the three-dimensional audio signal is high, which occupies too many computing resources of the encoder. Therefore, how to reduce the computational complexity of compression encoding of the three-dimensional audio signal is an urgent problem to be solved.

[0063] An embodiment of the present application provides an audio coding and decoding technology, and in particular provides a three-dimensional audio coding and decoding technology for three-dimensional audio signals, and specifically provides a coding and decoding technology that uses fewer channels to represent three-dimensional audio signals to improve traditional audio coding and decoding systems. Audio coding (or commonly referred to as coding) includes two parts: audio coding and audio decoding. Audio coding is performed on the source side and generally includes processing (for example, compressing) the original audio to reduce the amount of data required to represent the original audio, thereby more efficiently storing and / or transmitting. Audio decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the original audio. The coding part and the decoding part are also collectively referred to as coding and decoding. The implementation of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0064] Figure 1 is a schematic diagram of the structure of an audio codec system provided in an embodiment of the present application. Audio codec system 100 includes a source device 110 and a destination device 120. Source device 110 compresses and encodes a 3D audio signal to generate a bitstream, which is then transmitted to destination device 120. Destination device 120 decodes the bitstream, reconstructs the 3D audio signal, and plays the reconstructed 3D audio signal.

[0065] Specifically, the source device 110 includes an audio acquirer 111 , a pre-processor 112 , an encoder 113 , and a communication interface 114 .

[0066] The audio acquirer 111 is used to acquire raw audio. The audio acquirer 111 can be any type of audio acquisition device for capturing real-world sounds, and / or any type of audio generation device. The audio acquirer 111 is, for example, a computer audio processor for generating computer audio. The audio acquirer 111 can also be any type of memory or storage for storing audio. The audio includes real-world sounds, virtual scene (such as VR or augmented reality (AR)) sounds, and / or any combination thereof.

[0067] The preprocessor 112 is configured to receive the original audio collected by the audio acquirer 111 and preprocess the original audio to obtain a three-dimensional audio signal. For example, the preprocessing performed by the preprocessor 112 includes channel conversion, audio format conversion, or noise removal.

[0068] Encoder 113 is configured to receive the 3D audio signal generated by preprocessor 112 and compress and encode the 3D audio signal to obtain a bitstream. For example, encoder 113 may include a spatial encoder 1131 and a core encoder 1132. Spatial encoder 1131 is configured to select (or search for) a virtual speaker from a set of candidate virtual speakers based on the 3D audio signal and generate a virtual speaker signal based on the 3D audio signal and the virtual speaker. The virtual speaker signal may also be referred to as a playback signal. Core encoder 1132 is configured to encode the virtual speaker signal to obtain a bitstream.

[0069] The communication interface 114 is configured to receive the code stream generated by the encoder 113 and send the code stream to the destination device 120 through the communication channel 130 , so that the destination device 120 can reconstruct a 3D audio signal based on the code stream.

[0070] The destination device 120 includes a player 121 , a post-processor 122 , a decoder 123 , and a communication interface 124 .

[0071] The communication interface 124 is configured to receive the code stream sent by the communication interface 114 and transmit the code stream to the decoder 123 so that the decoder 123 can reconstruct the three-dimensional audio signal according to the code stream.

[0072] The communication interface 114 and the communication interface 124 may be configured to send or receive data related to the original audio via a direct communication link between the source device 110 and the destination device 120, such as a direct wired or wireless connection, or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network, a public network, or any combination thereof.

[0073] Both the communication interface 114 and the communication interface 124 may be configured as a unidirectional communication interface, as indicated by the arrow pointing from the source device 110 to the corresponding communication channel 130 of the destination device 120 in FIG1 , or a bidirectional communication interface, and may be used to send and receive messages, etc., to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission, such as the transmission of an encoded bit stream, etc.

[0074] Decoder 123 is used to decode the bitstream and reconstruct the 3D audio signal. For example, decoder 123 includes a core decoder 1231 and a spatial decoder 1232. Core decoder 1231 is used to decode the bitstream to obtain a virtual speaker signal. Spatial decoder 1232 is used to reconstruct the 3D audio signal based on the candidate virtual speaker set and the virtual speaker signal, thereby obtaining a reconstructed 3D audio signal.

[0075] The post-processor 122 is configured to receive the reconstructed 3D audio signal generated by the decoder 123 and perform post-processing on the reconstructed 3D audio signal. For example, the post-processing performed by the post-processor 122 may include audio rendering, loudness normalization, user interaction, audio format conversion, or noise removal.

[0076] The player 121 is configured to play the reconstructed sound according to the reconstructed 3D audio signal.

[0077] It should be noted that the audio acquirer 111 and the encoder 113 can be integrated into one physical device, or can be set on different physical devices, without limitation. For example, the source device 110 shown in Figure 1 includes an audio acquirer 111 and an encoder 113, indicating that the audio acquirer 111 and the encoder 113 are integrated into one physical device, and the source device 110 can also be called an acquisition device. The source device 110 is, for example, a media gateway of a wireless access network, a media gateway of a core network, a transcoding device, a media resource server, an AR device, a VR device, a microphone or other audio acquisition device. If the source device 110 does not include the audio acquirer 111, it means that the audio acquirer 111 and the encoder 113 are two different physical devices, and the source device 110 can obtain the original audio from other devices (such as: an audio acquisition device or an audio storage device).

[0078] In addition, the player 121 and the decoder 123 can be integrated into one physical device, or can be set on different physical devices, without limitation. For example, the destination device 120 shown in Figure 1 includes a player 121 and a decoder 123, which means that the player 121 and the decoder 123 are integrated into one physical device, then the destination device 120 can also be called a playback device, and the destination device 120 has the function of decoding and playing reconstructed audio. The destination device 120 is, for example, a speaker, a headset or other device that plays audio. If the destination device 120 does not include the player 121, it means that the player 121 and the decoder 123 are two different physical devices. After the destination device 120 decodes the code stream to reconstruct the three-dimensional audio signal, it transmits the reconstructed three-dimensional audio signal to other playback devices (such as speakers or headphones), and the other playback devices play back the reconstructed three-dimensional audio signal.

[0079] In addition, FIG. 1 shows that the source device 110 and the destination device 120 may be integrated into one physical device, or may be set on different physical devices, which is not limited.

[0080] For example, as shown in Figure 2 (a), source device 110 can be a microphone in a recording studio, and destination device 120 can be a speaker. Source device 110 can capture raw audio from various instruments and transmit the raw audio to a codec device. The codec device encodes and decodes the raw audio to generate a reconstructed 3D audio signal, which is then played back by destination device 120. For another example, source device 110 can be a microphone in a terminal device, and destination device 120 can be headphones. Source device 110 can capture external sounds or audio synthesized by the terminal device.

[0081] As another example, as shown in FIG2(b), source device 110 and destination device 120 are integrated into a virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, or extended reality (XR) device. The VR / AR / MR / XR device has the functions of capturing original audio, playing back audio, and encoding and decoding. Source device 110 can capture sounds emitted by the user and sounds emitted by virtual objects in the user's virtual environment.

[0082] In these embodiments, the source device 110 or its corresponding functions and the destination device 120 or its corresponding functions may be implemented using the same hardware and / or software or through separate hardware and / or software or any combination thereof. As described, the presence and division of different units or functions in the source device 110 and / or destination device 120 shown in FIG1 may vary depending on the actual device and application, which will be apparent to those skilled in the art.

[0083] The structure of the audio codec system described above is merely illustrative. In some possible implementations, the audio codec system may also include other devices, such as end-side devices or cloud-side devices. After capturing raw audio, source device 110 preprocesses the raw audio to generate a 3D audio signal. The 3D audio is then transmitted to the end-side device or cloud-side device, which then encodes and decodes the 3D audio signal.

[0084] The audio signal encoding and decoding method provided in the embodiments of this application is primarily used on the encoding side. The structure of the encoder is described in detail with reference to FIG3 . As shown in FIG3 , the encoder 300 includes a virtual speaker configuration unit 310, a virtual speaker set generation unit 320, a coding analysis unit 330, a virtual speaker selection unit 340, a virtual speaker signal generation unit 350, and an encoding unit 360.

[0085] The virtual speaker configuration unit 310 is used to generate virtual speaker configuration parameters based on the encoder configuration information to obtain multiple virtual speakers. The encoder configuration information includes but is not limited to: the order of the three-dimensional audio signal (or commonly referred to as the HOA order), the encoding bit rate, user-defined information, etc. The virtual speaker configuration parameters include but are not limited to: the number of virtual speakers, the order of the virtual speakers, the position coordinates of the virtual speakers, etc. The number of virtual speakers can be, for example, 2048, 1669, 1343, 1024, 530, 512, 256, 128, or 64. The order of the virtual speakers can be any one of 2 to 6 orders. The position coordinates of the virtual speakers include the horizontal angle and the pitch angle.

[0086] The virtual speaker configuration parameters output by the virtual speaker configuration unit 310 serve as input to the virtual speaker set generation unit 320 .

[0087] The virtual speaker set generation unit 320 is configured to generate a candidate virtual speaker set based on the virtual speaker configuration parameters, where the candidate virtual speaker set includes a plurality of virtual speakers. Specifically, the virtual speaker set generation unit 320 determines the plurality of virtual speakers included in the candidate virtual speaker set based on the number of virtual speakers, and determines the coefficients of the virtual speakers based on the position information (e.g., coordinates) of the virtual speakers and the order of the virtual speakers. For example, methods for determining the coordinates of the virtual speakers include, but are not limited to: generating a plurality of virtual speakers according to an equidistant rule, or generating a plurality of non-uniformly distributed virtual speakers based on the principle of auditory perception; and then generating the coordinates of the virtual speakers based on the number of virtual speakers.

[0088] According to the above principle of generating three-dimensional audio signals, the coefficients of virtual speakers can also be generated. s and are set as the position coordinates of the virtual speakers, Indicates the coefficients of a virtual speaker of order N. The coefficients of a virtual speaker are also called ambisonics coefficients.

[0089] The coding analysis unit 330 is used to perform coding analysis on the 3D audio signal, for example, analyzing the sound field distribution characteristics of the 3D audio signal, that is, the number of sound sources, the directionality of the sound sources, and the dispersion of the sound sources.

[0090] The coefficients of the plurality of virtual speakers included in the candidate virtual speaker set output by the virtual speaker set generation unit 320 serve as input to the virtual speaker selection unit 340 .

[0091] The sound field distribution feature of the three-dimensional audio signal output by the encoding analysis unit 330 serves as an input to the virtual speaker selection unit 340 .

[0092] The virtual speaker selection unit 340 is configured to determine a representative virtual speaker that matches the 3D audio signal according to the 3D audio signal to be encoded, the sound field distribution characteristics of the 3D audio signal, and coefficients of multiple virtual speakers.

[0093] Without limitation, the encoder 300 of the embodiment of the present application may not include the encoding analysis unit 330. That is, the encoder 300 may not analyze the input signal, and the virtual speaker selection unit 340 may use a default configuration to determine the representative virtual speaker. For example, the virtual speaker selection unit 340 may determine the representative virtual speaker that matches the 3D audio signal based solely on the 3D audio signal and the coefficients of the multiple virtual speakers.

[0094] The encoder 300 may use a 3D audio signal acquired from an acquisition device or a 3D audio signal synthesized using an artificial audio object as input to the encoder 300. Furthermore, the 3D audio signal input to the encoder 300 may be a time-domain 3D audio signal or a frequency-domain 3D audio signal, without limitation.

[0095] The position information representing the virtual speaker and the coefficient representing the virtual speaker output by the virtual speaker selection unit 340 are input to the virtual speaker signal generation unit 350 and the encoding unit 360 .

[0096] The virtual speaker signal generation unit 350 is configured to generate a virtual speaker signal based on a 3D audio signal and attribute information representing a virtual speaker. The attribute information representing the virtual speaker includes at least one of position information representing the virtual speaker, coefficients representing the virtual speaker, and coefficients of the 3D audio signal. If the attribute information includes position information representing the virtual speaker, the coefficients representing the virtual speaker are determined based on the position information. If the attribute information includes coefficients of the 3D audio signal, the coefficients representing the virtual speaker are obtained based on the coefficients of the 3D audio signal. Specifically, the virtual speaker signal generation unit 350 calculates the virtual speaker signal based on the coefficients of the 3D audio signal and the coefficients representing the virtual speaker.

[0097] For example, assume that matrix A represents the coefficients of the virtual speaker and matrix X represents the HOA coefficients of the HOA signal. Matrix X is the inverse matrix of matrix A. The theoretical optimal solution w is obtained using the least squares method, where w represents the virtual speaker signal. The virtual speaker signal satisfies formula (5).

[0098] w=A -1 X Formula (5)

[0099] Among them, A -1Represents the inverse matrix of matrix A. The size of matrix A is (M×C), C represents the number of virtual speakers, M represents the number of channels of the HOA signal of order N, a represents the coefficient representing the virtual speaker, and the size of matrix X is (M×L), L represents the number of coefficients of the HOA signal, and x represents the coefficient of the HOA signal. The coefficient representing the virtual speaker can refer to the HOA coefficient representing the virtual speaker or the ambisonics coefficient representing the virtual speaker. For example,

[0100] The virtual speaker signal output by the virtual speaker signal generating unit 350 serves as an input to the encoding unit 360 .

[0101] The encoding unit 360 is used to perform core encoding processing on the virtual speaker signal to obtain a bitstream. The core encoding processing includes but is not limited to: transformation, quantization, psychoacoustic modeling, noise shaping, bandwidth extension, downmixing, arithmetic coding, bitstream generation, etc.

[0102] It is worth noting that the spatial encoder 1131 may include a virtual speaker configuration unit 310, a virtual speaker set generation unit 320, a coding analysis unit 330, a virtual speaker selection unit 340, and a virtual speaker signal generation unit 350. That is, the virtual speaker configuration unit 310, the virtual speaker set generation unit 320, the coding analysis unit 330, the virtual speaker selection unit 340, and the virtual speaker signal generation unit 350 implement the functions of the spatial encoder 1131. The core encoder 1132 may include an encoding unit 360. That is, the encoding unit 360 implements the functions of the core encoder 1132.

[0103] The encoder shown in Figure 3 can generate one virtual speaker signal or multiple virtual speaker signals. The multiple virtual speaker signals can be obtained by executing the encoder shown in Figure 3 multiple times or by executing the encoder shown in Figure 3 once.

[0104] Next, the encoding and decoding process of a 3D audio signal will be described with reference to the accompanying drawings. Figure 4 is a schematic flow chart of a 3D audio signal encoding and decoding method provided in an embodiment of the present application. This description will be based on an example of the 3D audio signal encoding and decoding process performed by source device 110 and destination device 120 in Figure 1. As shown in Figure 4, the method includes the following steps.

[0105] S410: The source device 110 obtains a current frame of a 3D audio signal.

[0106] As described in the above embodiment, if source device 110 carries audio acquirer 111, source device 110 can capture raw audio through audio acquirer 111. Alternatively, source device 110 can receive raw audio captured by other devices, or acquire raw audio from a memory or other storage device within source device 110. Raw audio can include at least one of real-world sounds captured in real time, audio stored on the device, and audio synthesized from multiple audio sources. This embodiment does not limit the method for acquiring raw audio or the type of raw audio.

[0107] After receiving the original audio, source device 110 generates a 3D audio signal based on 3D audio technology and the original audio. This generates an immersive audio experience for the listener when the original audio is played back. For details on how to generate the 3D audio signal, see the description of preprocessor 112 in the above embodiment and the prior art.

[0108] Furthermore, audio signals are continuous analog signals. During audio signal processing, the audio signal can be sampled to generate a digital signal consisting of a sequence of frames. A frame can include multiple sampling points. A frame can also refer to a sampling point obtained by sampling. A frame can also include subframes obtained by dividing a frame. A frame can also refer to a subframe obtained by dividing a frame. For example, if a frame is L sampling points long and is divided into N subframes, then each subframe corresponds to L / N sampling points. Audio encoding and decoding generally refers to processing a sequence of audio frames containing multiple sampling points.

[0109] An audio frame may include a current frame or a previous frame. The current frame or previous frame described in various embodiments of the present application may refer to a frame or a subframe. The current frame refers to a frame that is being coded and decoded at the current moment. The previous frame refers to a frame that has been coded and decoded at a moment before the current moment. The previous frame may be a frame at a moment before or multiple moments before the current moment. In an embodiment of the present application, the current frame of a three-dimensional audio signal refers to a frame of a three-dimensional audio signal that is being coded and decoded at the current moment. The previous frame refers to a frame of a three-dimensional audio signal that has been coded and decoded at a moment before the current moment. The current frame of a three-dimensional audio signal may refer to the current frame of the three-dimensional audio signal to be encoded. The current frame of a three-dimensional audio signal may be referred to as the current frame. The previous frame of a three-dimensional audio signal may be referred to as the previous frame.

[0110] S420 : The source device 110 determines a candidate virtual speaker set.

[0111] In one scenario, a candidate virtual speaker set is preconfigured in the memory of source device 110. Source device 110 can read the candidate virtual speaker set from the memory. The candidate virtual speaker set includes multiple virtual speakers. Virtual speakers represent speakers that exist virtually in a spatial sound field. The virtual speakers are used to calculate virtual speaker signals based on the 3D audio signal, so that destination device 120 can playback the reconstructed 3D audio signal.

[0112] In another scenario, virtual speaker configuration parameters are pre-configured in the memory of source device 110. Source device 110 generates a candidate virtual speaker set based on the virtual speaker configuration parameters. Alternatively, source device 110 generates the candidate virtual speaker set in real time based on its own computing resources (e.g., processor) capabilities and characteristics of the current frame (e.g., channels and data volume).

[0113] The specific method for generating the candidate virtual speaker set may refer to the prior art and the description of the virtual speaker configuration unit 310 and the virtual speaker set generation unit 320 in the above embodiments.

[0114] S430 : The source device 110 selects a representative virtual speaker for the current frame from the candidate virtual speaker set according to the current frame of the 3D audio signal.

[0115] Source device 110 votes for virtual speakers based on the coefficients of the current frame and the coefficients of the virtual speakers. A representative virtual speaker for the current frame is selected from a set of candidate virtual speakers based on the votes. A limited number of representative virtual speakers for the current frame are searched from the set of candidate virtual speakers to be selected as the best-matching virtual speakers for the current frame to be encoded, thereby achieving data compression of the 3D audio signal to be encoded.

[0116] Figure 5 is a schematic flow chart of a method for selecting virtual speakers provided in an embodiment of the present application. The method flow shown in Figure 5 illustrates the specific operations included in S430 in Figure 4 . This description uses the example of encoder 113 executing virtual speaker selection in source device 110 shown in Figure 1 as an example. Specifically, the functions of virtual speaker selection unit 340 are implemented. As shown in Figure 5 , the method includes the following steps.

[0117] S510 : The encoder 113 obtains representative coefficients of the current frame.

[0118] The representative coefficients may refer to frequency domain representative coefficients or time domain representative coefficients. Frequency domain representative coefficients may also be referred to as frequency domain representative frequency points or spectrum representative coefficients. Time domain representative coefficients may also be referred to as time domain representative sampling points. For a specific method of obtaining the representative coefficients for the current frame, refer to S6101 and S6102 in FIG. 7 .

[0119] S520: The encoder 113 selects a representative virtual speaker for the current frame from the candidate virtual speaker set according to the voting values ​​of the representative coefficients of the current frame for the virtual speakers in the candidate virtual speaker set. Execute S440 to S460.

[0120] The encoder 113 votes for the virtual speakers in the candidate virtual speaker set based on the representative coefficients of the current frame and the coefficients of the virtual speakers. The encoder 113 then selects (searches) a representative virtual speaker for the current frame from the candidate virtual speaker set based on the final vote value of the virtual speaker for the current frame. The specific method for selecting the representative virtual speaker for the current frame can be found in FIG. 8 and the description of S6103 in FIG. 7 .

[0121] It should be noted that the encoder first traverses the virtual speakers included in the candidate virtual speaker set and compresses the current frame using the representative virtual speaker selected from the candidate virtual speaker set for the current frame. However, if the results of the virtual speakers selected in consecutive frames differ significantly, this can lead to unstable sound images in the reconstructed 3D audio signal and reduce the sound quality of the reconstructed 3D audio signal. In an embodiment of the present application, the encoder 113 can update the initial voting values ​​of the virtual speakers included in the candidate virtual speaker set in the current frame based on the final voting values ​​of the representative virtual speakers in the previous frame, obtaining the final voting values ​​of the virtual speakers in the current frame. The encoder then selects the representative virtual speaker for the current frame from the candidate virtual speaker set based on the final voting values ​​of the virtual speakers in the current frame. Thus, by referencing the representative virtual speakers in the previous frame to select the representative virtual speakers for the current frame, the encoder tends to select the same virtual speakers as the representative virtual speakers in the previous frame when selecting the representative virtual speakers for the current frame, thereby increasing the continuity of the orientations between consecutive frames and overcoming the problem of large differences in the results of the virtual speakers selected in consecutive frames. Therefore, an embodiment of the present application can also include S530.

[0122] S530 : The encoder 113 adjusts the current frame initial voting value of the virtual speaker in the candidate virtual speaker set according to the previous frame final voting value of the representative virtual speaker in the previous frame to obtain the current frame final voting value of the virtual speaker.

[0123] The encoder 113 votes for the virtual speakers in the candidate virtual speaker set based on the representative coefficients of the current frame and the coefficients of the virtual speakers. After obtaining the initial voting values ​​of the virtual speakers in the current frame, the encoder 113 adjusts the initial voting values ​​of the virtual speakers in the current frame according to the final voting values ​​of the representative virtual speakers in the previous frame to obtain the final voting values ​​of the virtual speakers in the current frame. The representative virtual speakers in the previous frame are the virtual speakers used by the encoder 113 when encoding the previous frame. The specific method of adjusting the initial voting values ​​of the virtual speakers in the current frame in the candidate virtual speaker set can be referred to S620 to S630 in Figure 6 and S810 to S840 in Figure 8.

[0124] In some embodiments, if the current frame is the first frame in the original audio, encoder 113 executes S510 to S520. If the current frame is any frame above the second frame in the original audio, encoder 113 may first determine whether to reuse the representative virtual speaker of the previous frame to encode the current frame or determine whether to perform a virtual speaker search to ensure position continuity between consecutive frames and reduce encoding complexity. Embodiments of the present application may also include S540.

[0125] S540 : The encoder 113 determines whether to perform a virtual speaker search based on the representative virtual speaker of the previous frame and the current frame.

[0126] If the encoder 113 determines to perform a virtual speaker search, S510 to S530 are executed. Alternatively, the encoder 113 may first execute S510, i.e., the encoder 113 obtains the representative coefficients of the current frame, and then determines whether to perform a virtual speaker search based on the representative coefficients of the current frame and the coefficients of the representative virtual speakers of the previous frame. If the encoder 113 determines to perform a virtual speaker search, S520 to S530 are executed.

[0127] If the encoder 113 determines not to perform a virtual speaker search, step S550 is executed.

[0128] S550 : The encoder 113 determines to reuse the representative virtual speaker of the previous frame to encode the current frame.

[0129] The encoder 113 multiplexes the representative virtual speaker of the previous frame and the current frame to generate a virtual speaker signal, encodes the virtual speaker signal to obtain a code stream, and sends the code stream to the destination device 120, that is, executing S450 and S460.

[0130] The specific method of determining whether to perform a virtual speaker search may refer to the description of S650 to S680 in FIG. 9 below.

[0131] S440 : The source device 110 generates a virtual speaker signal according to the current frame of the 3D audio signal and the representative virtual speaker of the current frame.

[0132] The source device 110 generates a virtual speaker signal based on the coefficients of the current frame and the coefficients representing the virtual speaker of the current frame. The specific method of generating the virtual speaker signal can refer to the existing technology and the description of the virtual speaker signal generating unit 350 in the above embodiment.

[0133] S450 : The source device 110 encodes the virtual speaker signal to obtain a bit stream.

[0134] The source device 110 can perform encoding operations such as transformation or quantization on the virtual speaker signal to generate a bitstream, thereby achieving the purpose of data compression of the encoded 3D audio signal. The specific method of generating the bitstream can refer to the existing technology and the description of the encoding unit 360 in the above embodiment.

[0135] S460 : The source device 110 sends a code stream to the destination device 120 .

[0136] After encoding the original audio, source device 110 can send the original audio code stream to destination device 120. Alternatively, source device 110 can encode the 3D audio signal in real time, frame by frame, and send the code stream for each frame after encoding. For specific methods of sending the code stream, reference can be made to existing technologies and the description of communication interface 114 and communication interface 124 in the above embodiments.

[0137] S470 : The destination device 120 decodes the code stream sent by the source device 110 , reconstructs the 3D audio signal, and obtains a reconstructed 3D audio signal.

[0138] After receiving the bitstream, the destination device 120 decodes the bitstream to obtain virtual speaker signals. It then reconstructs a 3D audio signal based on the candidate virtual speaker set and the virtual speaker signals to obtain a reconstructed 3D audio signal. The destination device 120 then plays back the reconstructed 3D audio signal. Alternatively, the destination device 120 transmits the reconstructed 3D audio signal to another playback device, which then plays the reconstructed 3D audio signal, making the listener's "immersive" sound experience, such as being in a theater, concert hall, or virtual scene, even more realistic.

[0139] In order to increase the continuity of the orientation between consecutive frames and overcome the problem of large differences in the results of virtual speakers selected in consecutive frames, the encoder 113 adjusts the initial voting value of the current frame of the virtual speaker in the candidate virtual speaker set according to the final voting value of the previous frame of the representative virtual speaker of the previous frame to obtain the final voting value of the current frame of the virtual speaker. As shown in Figure 6, a flow chart of another method for selecting virtual speakers provided in an embodiment of the present application is provided. Here, the process of selecting virtual speakers executed by the encoder 113 in the source device 110 in Figure 1 is used as an example for explanation. Among them, the method flow described in Figure 6 is an explanation of the specific operation process included in S530 in Figure 5. As shown in Figure 6, the method includes the following steps.

[0140] S610 : The encoder 113 obtains a first number of current frame initial voting values ​​of a current frame of a 3D audio signal.

[0141] The encoder 113 can use the representative coefficient of the current frame to vote for each virtual speaker in the candidate virtual speaker set, obtain the initial voting value of the virtual speaker in the current frame, and select the representative virtual speaker of the current frame based on the voting value, thereby reducing the computational complexity of the virtual speaker search and alleviating the computational burden of the encoder.

[0142] Figure 7 is a schematic flow chart of another 3D audio signal encoding method provided in an embodiment of the present application. This method is illustrated using the virtual speaker selection process performed by encoder 113 in source device 110 in Figure 1 as an example. The method flow shown in Figure 7 is an elaboration of the specific operations included in S510 and S520 in Figure 5 . As shown in Figure 7 , the method includes the following steps.

[0143] S6101: The encoder 113 obtains a fourth number of coefficients of a current frame of a three-dimensional audio signal and frequency domain eigenvalues ​​of the fourth number of coefficients.

[0144] Assuming that the three-dimensional audio signal is an HOA signal, the encoder 113 may sample the current frame of the HOA signal to obtain L·(N+1) 2 sampling points, that is, the fourth number of coefficients is obtained. N represents the order of the HOA signal. For example, assuming that the duration of the current frame of the HOA signal is 20 milliseconds, the encoder 113 samples the current frame according to the 48KHz frequency, and obtains 960·(N+1) in the time domain. 2 Sampling points. Sampling points can also be called time domain coefficients.

[0145] The frequency domain coefficients of the current frame of the three-dimensional audio signal can be obtained by performing time-frequency conversion based on the time domain coefficients of the current frame of the three-dimensional audio signal. The method of converting the time domain to the frequency domain is not limited. The method of converting the time domain to the frequency domain is, for example, a modified discrete cosine transform (MDCT), which can obtain 960·(N+1) in the frequency domain. 2 Frequency domain coefficients. Frequency domain coefficients can also be called spectrum coefficients or frequency points.

[0146] The frequency domain eigenvalues ​​of the sampling points satisfy p(j)=norm(x(j)), where j=1, 2…L, L represents the number of sampling moments, x represents the frequency domain coefficients of the current frame of the three-dimensional audio signal, such as MDCT coefficients, and norm is the bi-norm operation; x(j) represents the (N+1) at the j-th sampling moment. 2 The frequency domain coefficients of the sampling points.

[0147] S6102: The encoder 113 selects a third number of representative coefficients from the fourth number of coefficients according to the frequency domain eigenvalues ​​of the fourth number of coefficients.

[0148] The encoder 113 divides the frequency spectrum range indicated by the fourth number of coefficients into at least one subband. The encoder 113 divides the frequency spectrum range indicated by the fourth number of coefficients into one subband. It is understandable that the frequency spectrum range of the one subband is equal to the frequency spectrum range indicated by the fourth number of coefficients, which is equivalent to the encoder 113 not dividing the frequency spectrum range indicated by the fourth number of coefficients.

[0149] If the encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two frequency band sub-bands, in one case, the encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two sub-bands, and each of the at least two sub-bands contains the same number of coefficients.

[0150] In another scenario, encoder 113 may unequally divide the spectral range indicated by the fourth number of coefficients, so that the at least two subbands obtained by the division have different numbers of coefficients, or each of the at least two subbands obtained by the division has a different number of coefficients. For example, encoder 113 may unequally divide the spectral range indicated by the fourth number of coefficients based on the low-frequency range, the mid-frequency range, and the high-frequency range in the spectral range indicated by the fourth number of coefficients, so that each of the low-frequency range, the mid-frequency range, and the high-frequency range includes at least one subband. Each subband in the at least one subband in the low-frequency range contains the same number of coefficients. Each subband in the at least one subband in the mid-frequency range contains the same number of coefficients. Each subband in the at least one subband in the high-frequency range contains the same number of coefficients. The subbands in the low-frequency range, the mid-frequency range, and the high-frequency range may contain different numbers of coefficients.

[0151] Furthermore, the encoder 113 selects representative coefficients from at least one subband within the spectrum range indicated by the fourth number of coefficients based on the frequency domain eigenvalues ​​of the fourth number of coefficients to obtain a third number of representative coefficients. The third number is smaller than the fourth number, and the fourth number of coefficients includes the third number of representative coefficients.

[0152] For example, the encoder 113 selects Z representative coefficients from each subband according to the descending order of the frequency domain eigenvalues ​​of the coefficients in at least one subband contained in the spectral range indicated by the fourth number of coefficients, and combines the Z representative coefficients in at least one subband to obtain a third number of representative coefficients, where Z is a positive integer.

[0153] For another example, when the at least one subband includes at least two subbands, the encoder 113 determines a weight for each subband based on the frequency domain eigenvalues ​​of the first candidate coefficients in each of the at least two subbands; and adjusts the frequency domain eigenvalues ​​of the second candidate coefficients in each subband based on the weights of each subband to obtain adjusted frequency domain eigenvalues ​​of the second candidate coefficients in each subband, where the first candidate coefficients and the second candidate coefficients are some of the coefficients in the subband. The encoder 113 determines a third number of representative coefficients based on the adjusted frequency domain eigenvalues ​​of the second candidate coefficients in the at least two subbands and the frequency domain eigenvalues ​​of the coefficients in the at least two subbands other than the second candidate coefficients.

[0154] Since the encoder selects some coefficients from all the coefficients of the current frame as representative coefficients, and uses a smaller number of representative coefficients instead of all the coefficients of the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder's search for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden of the encoder.

[0155] S6103: The encoder 113 determines a first number of virtual speakers and a first number of voting values ​​according to the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds.

[0156] The number of voting rounds is used to limit the number of times a virtual speaker is voted. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the number of virtual speakers included in the candidate virtual speaker set, and the number of voting rounds is less than or equal to the number of virtual speaker signals transmitted by the encoder. For example, the candidate virtual speaker set includes a fifth number of virtual speakers, the fifth number of virtual speakers includes a first number of virtual speakers, the first number is less than or equal to the fifth number, the number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number. The virtual speaker signal also refers to the transmission channel representing the virtual speaker of the current frame corresponding to the current frame. Normally, the number of virtual speaker signals is less than or equal to the number of virtual speakers.

[0157] In a possible implementation, the number of voting rounds may be preconfigured or determined according to the computing capability of the encoder. For example, the number of voting rounds is determined according to the encoding rate and / or encoding application scenario of the encoder.

[0158] In another possible implementation, the number of voting rounds is determined based on the number of directional sound sources in the current frame. For example, when the number of directional sound sources in the sound field is 2, the number of voting rounds is set to 2.

[0159] The embodiment of the present application provides three possible implementation methods for determining the first number of virtual speakers and the first number of voting values. The three methods are described in detail below.

[0160] In a first possible implementation, the number of voting rounds is equal to 1. After the encoder 113 samples multiple representative coefficients, it obtains the voting value of each representative coefficient of the current frame for all virtual speakers in the candidate virtual speaker set, accumulates the voting values ​​of the virtual speakers with the same number, and obtains a first number of virtual speakers and a first number of voting values. It can be understood that the candidate virtual speaker set includes a first number of virtual speakers. The first number is equal to the number of virtual speakers included in the candidate virtual speaker set. Assuming that the candidate virtual speaker set includes a fifth number of virtual speakers, the first number is equal to the fifth number. The first number of voting values ​​includes the voting values ​​of all virtual speakers in the candidate virtual speaker set. The encoder 113 can use the first number of voting values ​​as the initial voting values ​​of the current frame of the first number of virtual speakers and execute S620 to S640.

[0161] Among them, the virtual speakers correspond to the voting values ​​one by one, that is, one virtual speaker corresponds to one voting value. For example, the first number of virtual speakers includes the first virtual speaker, the first number of voting values ​​includes the voting value of the first virtual speaker, and the first virtual speaker corresponds to the voting value of the first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of using the first virtual speaker when encoding the current frame. Priority can also be replaced by tendency, that is, the voting value of the first virtual speaker is used to characterize the tendency to use the first virtual speaker when encoding the current frame. It can be understood that the larger the voting value of the first virtual speaker, the higher the priority or tendency of the first virtual speaker. Compared with the virtual speakers in the candidate virtual speaker set with smaller voting values ​​than the first virtual speaker, the encoder 113 is more inclined to select the first virtual speaker to encode the current frame.

[0162] In a second possible implementation, the difference from the first possible implementation is that after the encoder 113 obtains the voting value of each representative coefficient of the current frame for all virtual speakers in the candidate virtual speaker set, it selects a portion of the voting values ​​from the voting values ​​of each representative coefficient for all virtual speakers in the candidate virtual speaker set, accumulates the voting values ​​of the virtual speakers with the same number in the virtual speakers corresponding to the partial voting values, and obtains a first number of virtual speakers and a first number of voting values. It can be understood that the candidate virtual speaker set includes a first number of virtual speakers. The first number is less than or equal to the number of virtual speakers included in the candidate virtual speaker set. The first number of voting values ​​includes the voting values ​​of some virtual speakers included in the candidate virtual speaker set, or the first number of voting values ​​includes the voting values ​​of all virtual speakers included in the candidate virtual speaker set.

[0163] In a third possible implementation, which differs from the second possible implementation, the number of voting rounds is an integer greater than or equal to 2. For each representative coefficient of the current frame, the encoder 113 performs at least two rounds of voting on all virtual speakers in the candidate virtual speaker set, selecting the virtual speaker with the maximum voting value in each round. After performing at least two rounds of voting on all virtual speakers for each representative coefficient of the current frame, the voting values ​​of the virtual speakers with the same number are accumulated to obtain a first number of virtual speakers and a first number of voting values.

[0164] S620 : The encoder 113 obtains a seventh number of final voting values ​​of the current frame corresponding to a seventh number of virtual speakers and the current frame according to the first number of initial voting values ​​of the current frame and the sixth number of final voting values ​​of the previous frame.

[0165] The encoder 113 can use the method described in S610 above to determine a first number of virtual speakers and a first number of voting values ​​based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set and the number of voting rounds, and then use the first number of voting values ​​as the initial voting values ​​of the first number of virtual speakers for the current frame.

[0166] The virtual speakers correspond to the current frame initial voting values ​​in a one-to-one manner, i.e., one virtual speaker corresponds to one current frame initial voting value. For example, the first number of virtual speakers includes the first virtual speaker, the first number of current frame initial voting values ​​includes the current frame initial voting value of the first virtual speaker, and the first virtual speaker corresponds to the current frame initial voting value of the first virtual speaker. The current frame initial voting value of the first virtual speaker is used to indicate the priority of using the first virtual speaker when encoding the current frame.

[0167] The sixth number of virtual speakers may be representative virtual speakers of a previous frame used by encoder 113 to encode the previous frame of the three-dimensional audio signal. In S650, when encoder 113 obtains a first correlation between the current frame of the three-dimensional audio signal and the representative virtual speaker set of the previous frame, the representative virtual speaker set of the previous frame includes the sixth number of virtual speakers.

[0168] Specifically, the encoder 113 updates the first number of initial voting values ​​of the current frame based on the sixth number of final voting values ​​of the previous frame, that is, the encoder 113 calculates the sum of the initial voting values ​​of the current frame and the final voting values ​​of the previous frame of the first number of virtual speakers and the virtual speakers with the same numbers in the sixth number of virtual speakers, and obtains the seventh number of final voting values ​​of the current frame corresponding to the seventh number of virtual speakers and the current frame.

[0169] In the first possible scenario, the first number of virtual speakers includes the sixth number of virtual speakers, and the first number is equal to the sixth number, and the number of the first number of virtual speakers is the same as the number of the sixth number of virtual speakers. It can be understood that the first number of virtual speakers obtained by the encoder 113 is the sixth number of virtual speakers, and the final voting value of the sixth number of virtual speakers in the previous frame is the final voting value of the first number of virtual speakers in the previous frame. The encoder 113 can use the final voting value of the sixth number of virtual speakers in the previous frame to update the initial voting value of the first number of virtual speakers in the current frame. Therefore, the seventh number of virtual speakers is also the first number of virtual speakers, and the final voting value of the seventh number of virtual speakers in the current frame is the sum of the final voting value of the first number of virtual speakers in the previous frame and the initial voting value of the first number of virtual speakers in the current frame.

[0170] For example, assuming that the sixth number of virtual speakers includes the first virtual speaker, the first number of virtual speakers includes the first virtual speaker, and neither the sixth number of virtual speakers nor the first number of virtual speakers includes any other virtual speakers. The encoder 113 may update the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame. The final voting value of the first virtual speaker in the current frame is the sum of the final voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame.

[0171] In the second possible scenario, the first number of virtual speakers includes the sixth number of virtual speakers, and the first number is greater than the sixth number. It is understandable that the first number of virtual speakers also includes other virtual speakers in addition to the sixth number of virtual speakers. The encoder 113 can use the final voting value of the previous frame of the sixth number of virtual speakers to update the initial voting value of the current frame of the virtual speakers with the same number as the sixth number of virtual speakers in the first number of virtual speakers. Therefore, the seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number is equal to the first number, and the number of the seventh number of virtual speakers is the same as the number of the first number of virtual speakers. The final voting values ​​of the seventh number of current frames include the final voting values ​​of the current frame of the virtual speakers with the same number as the sixth number of virtual speakers in the first number of virtual speakers, and the final voting values ​​of the current frame of the virtual speakers with different numbers from the sixth number of virtual speakers in the first number of virtual speakers.

[0172] The final voting value of the current frame for the virtual speaker in the first number of virtual speakers that has the same number as the sixth number of virtual speakers is the sum of the final voting value of the sixth number of virtual speakers in the previous frame and the initial voting value of the first number of virtual speakers in the current frame. The final voting value of the current frame for the virtual speaker in the first number of virtual speakers that has a different number than the sixth number of virtual speakers is the initial voting value of the current frame for the virtual speaker in the first number of virtual speakers that has a different number than the sixth number of virtual speakers.

[0173] For example, assuming that the first number of virtual speakers includes the first virtual speaker and the second virtual speaker, the sixth number of virtual speakers includes the first virtual speaker, and the sixth number of virtual speakers does not include the second virtual speaker, then the final voting value of the current frame of the second virtual speaker is equal to the initial voting value of the current frame of the second virtual speaker. The encoder 113 can update the initial voting value of the current frame of the first virtual speaker based on the final voting value of the previous frame of the first virtual speaker to obtain the final voting value of the current frame of the first virtual speaker. The final voting value of the current frame of the first virtual speaker is the sum of the final voting value of the previous frame of the first virtual speaker and the initial voting value of the current frame of the first virtual speaker.

[0174] In a third possible scenario, the first number of virtual speakers includes some of the virtual speakers in the sixth number of virtual speakers, and the sixth number of virtual speakers also includes other virtual speakers with numbers different from those in the first number of virtual speakers. Therefore, the seventh number of virtual speakers includes the first number of virtual speakers and the virtual speakers in the sixth number of virtual speakers with numbers different from those in the first number of virtual speakers. The seventh number of final voting values ​​for the current frame includes the final voting values ​​for the current frame of the first number of virtual speakers and the final voting values ​​for the current frame of the virtual speakers in the sixth number of virtual speakers with numbers different from those in the first number of virtual speakers.

[0175] The final voting values ​​of the first number of virtual speakers in the current frame include the final voting values ​​of the virtual speakers in the first number of virtual speakers that have the same number as the sixth number of virtual speakers. Optionally, the final voting values ​​of the first number of virtual speakers in the current frame may also include the final voting values ​​of the virtual speakers in the first number of virtual speakers that have different number than the sixth number of virtual speakers.

[0176] The current frame final voting value of the virtual speaker with a different number from the first number of virtual speakers among the sixth number of virtual speakers is the previous frame final voting value of the virtual speaker with a different number from the first number of virtual speakers among the sixth number of virtual speakers.

[0177] For example, assuming that the sixth number of virtual speakers includes the first virtual speaker and the third virtual speaker, the first number of virtual speakers includes the first virtual speaker, and the first number of virtual speakers does not include the third virtual speaker, then the final voting value of the current frame of the third virtual speaker is equal to the final voting value of the previous frame of the third virtual speaker. The encoder 113 can update the initial voting value of the current frame of the first virtual speaker based on the final voting value of the previous frame of the first virtual speaker to obtain the final voting value of the current frame of the first virtual speaker. The final voting value of the current frame of the first virtual speaker is the sum of the final voting value of the previous frame of the first virtual speaker and the initial voting value of the current frame of the first virtual speaker.

[0178] In some embodiments, as shown in FIG8 , a flowchart of a method for updating the initial voting value of a virtual speaker in the current frame provided by an embodiment of the present application is shown.

[0179] S810 : The encoder 113 adjusts the final voting value of the first virtual speaker in the previous frame according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame.

[0180] The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, a coding rate for encoding the current frame, and a frame type. The vote value of the first virtual speaker after adjustment in the previous frame satisfies the following formula (6).

[0181] VOTE_f′ g =VOTE_f g ·w1·w2·w3 Formula (6)

[0182] Among them, VOTE_f′ g Represents the voting value set after the previous frame adjustment, VOTE_f g represents the final voting value set of the previous frame, g represents the representative virtual speaker set of the previous frame, w1 represents a parameter related to the coding rate, w2 represents a parameter related to the frame type, and w3 represents a parameter related to the number of directional sound sources. Frame types include transient frames and non-transient frames.

[0183] For example, if the coding rate is less than or equal to 128 kbps, w1 = 1; if the coding rate is greater than 128 kbps, w1 = 0. If the previous frame is a transient frame, w2 = 1; if the previous frame is a non-transient frame, w2 = 0. If the number of directional sound sources is greater than the number of preset virtual speaker signals, w3 = 0.8; if the number of directional sound sources is less than or equal to the number of preset virtual speaker signals, w3 = 0.5.

[0184] S820: The encoder 113 updates the initial voting value of the first virtual speaker in the current frame according to the adjusted voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.

[0185] The final voting value of the first virtual speaker in the current frame is the sum of the adjusted voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame. The final voting value of the first virtual speaker in the current frame satisfies the following formula (7).

[0186] VOTE_M g =VOTE_f′ g +VOTE g Formula (7)

[0187] Among them, VOTE_M g Represents the final voting value set of the current frame, VOTE_f′ g Represents the voting value set after the previous frame adjustment, VOTE g Represents the initial voting value set for the current frame.

[0188] Optionally, the encoder 113 updates the initial voting value of the first virtual speaker in the current frame according to the adjusted voting value of the first virtual speaker in the previous frame, specifically including the following steps.

[0189] S830: The encoder 113 adjusts the initial voting value of the first virtual speaker in the current frame according to the second adjustment parameter to obtain an adjusted voting value of the first virtual speaker in the current frame.

[0190] The adjusted voting value of the first virtual speaker in the current frame satisfies the following formula (8).

[0191] VOTE′ g =VOTE g ·w4 formula (8)

[0192] Among them, VOTE g represents the voting value set after adjustment for the current frame, and w4 represents the second adjustment parameter. For example, if norm(VOTE g )>norm(VOTE_f′ g ), It can be understood that when the initial voting value of the current frame is greater than the adjusted voting value of the previous frame, w4 is used to represent the amplification of the adjusted voting value of the previous frame.

[0193] If norm(VOTE g )≤norm(VOTE_f′ g), w4 = 1. It is understandable that when the initial voting value of the current frame is less than or equal to the adjusted voting value of the previous frame, there is no need to use w4 to represent the amplification of the adjusted voting value of the previous frame.

[0194] The second adjustment parameter is determined according to the adjusted voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame.

[0195] S840: The encoder 113 updates the adjusted voting value of the first virtual speaker in the current frame according to the adjusted voting value of the first virtual speaker in the previous frame to obtain a final voting value of the first virtual speaker in the current frame.

[0196] The final voting value of the first virtual speaker in the current frame is the sum of the adjusted voting value of the first virtual speaker in the previous frame and the adjusted voting value of the first virtual speaker in the current frame. The final voting value of the first virtual speaker in the current frame satisfies the following formula (9).

[0197] VOTE_M g =VOTE_f′ g +VOTE′ g Formula (9)

[0198] Among them, VOTE_M g Represents the final voting value set of the current frame, VOTE_f′ g Represents the voting value set after adjustment in the previous frame, VOTE′ g Represents the adjusted voting value set for the current frame.

[0199] S630 : The encoder 113 selects a second number of representative virtual speakers for the current frames from the seventh number of virtual speakers according to the final voting values ​​of the seventh number of current frames.

[0200] The encoder 113 selects a second number of representative virtual speakers for the current frames from the seventh number of virtual speakers according to the seventh number of final voting values ​​of the current frames, and the current frame final voting values ​​of the second number of representative virtual speakers for the current frames are greater than a preset threshold.

[0201] The encoder 113 may also select representative virtual speakers for the second number of current frames from the seventh number of virtual speakers based on the final voting values ​​of the seventh number of current frames. For example, the encoder 113 may determine the second number of current frame final voting values ​​from the seventh number of current frames in descending order of the final voting values ​​of the seventh number of current frames, and select the virtual speakers in the seventh number of virtual speakers corresponding to the second number of current frame final voting values ​​as the representative virtual speakers for the second number of current frames.

[0202] Optionally, if the voting values ​​of virtual speakers with different numbers in the seventh number of virtual speakers are the same and the voting values ​​of the virtual speakers with different numbers are greater than a preset threshold, the encoder 113 may use the virtual speakers with different numbers as representative virtual speakers for the current frame.

[0203] It should be noted that the second number is smaller than the seventh number. The seventh number of virtual speakers includes the second number of representative virtual speakers of the current frame. The second number may be preset, or the second number may be determined based on the number of sound sources in the sound field of the current frame. For example, the second number may be directly equal to the number of sound sources in the sound field of the current frame, or the number of sound sources in the sound field of the current frame may be processed according to a preset algorithm, and the number obtained by processing may be used as the second number. The preset algorithm may be designed as needed. For example, the preset algorithm may be: the second number = the number of sound sources in the sound field of the current frame + 1, or the second number = the number of sound sources in the sound field of the current frame - 1, and so on.

[0204] In addition, before the encoder 113 encodes the next frame of the current frame, if the encoder 113 determines to reuse the representative virtual speakers of the previous frame to encode the next frame, the encoder 113 can use the second number of representative virtual speakers of the current frame as the second number of representative virtual speakers of the previous frame, and use the second number of representative virtual speakers of the previous frame to encode the next frame of the current frame.

[0205] S640 : The encoder 113 encodes the current frame according to the second number of representative virtual speakers of the current frame to obtain a code stream.

[0206] The encoder 113 generates a virtual speaker signal according to the second number of representative virtual speakers of the current frames and the current frame; and encodes the virtual speaker signal to obtain a bit stream.

[0207] During the virtual speaker search process, since the position of the real sound source and the position of the virtual speaker do not necessarily coincide, the virtual speaker may not be able to form a one-to-one correspondence with the real sound source. In addition, in actual complex scenarios, the virtual speaker may not be able to represent the independent sound source in the sound field. At this time, the virtual speaker searched between frames may frequently jump. This frequent jump will significantly affect the listener's auditory experience, resulting in obvious noise in the three-dimensional audio signal after decoding and reconstruction. The method for selecting a virtual speaker provided in an embodiment of the present application inherits the representative virtual speaker of the previous frame, that is, for virtual speakers with the same number, the final voting value of the previous frame is used to adjust the initial voting value of the current frame, so that the encoder is more inclined to select the representative virtual speaker of the previous frame, thereby enhancing the continuity of the orientation between frames. In addition, the parameters are adjusted to ensure that the final voting value of the previous frame is not inherited too long ago, so as to avoid the algorithm being unable to adapt to the scene of sound field changes such as sound source movement.

[0208] In addition, an embodiment of the present application provides a method for selecting virtual speakers. The encoder can first determine whether the representative virtual speaker set of the previous frame can be reused to encode the current frame. If the encoder reuses the representative virtual speaker set of the previous frame to encode the current frame, the encoder avoids performing the virtual speaker search process again, effectively reducing the computational complexity of the encoder's virtual speaker search, thereby reducing the computational complexity of compressing and encoding the three-dimensional audio signal and alleviating the computational burden of the encoder. If the encoder cannot reuse the representative virtual speaker set of the previous frame to encode the current frame, the encoder selects a representative coefficient and uses the representative coefficient of the current frame to vote for each virtual speaker in the candidate virtual speaker set. The representative virtual speaker of the current frame is selected based on the vote value, thereby reducing the computational complexity of compressing and encoding the three-dimensional audio signal and alleviating the computational burden of the encoder. Figure 9 is a flow diagram of a method for selecting virtual speakers provided by an embodiment of the present application. Before the encoder 113 obtains the first number of initial voting values ​​for the current frame corresponding to the first number of virtual speakers and the current frame of the three-dimensional audio signal, that is, before S610, as shown in Figure 9, the method includes the following steps.

[0209] S650 : The encoder 113 obtains a first correlation between a current frame and a representative virtual speaker set of a previous frame of the 3D audio signal.

[0210] The representative virtual speaker set of the previous frame includes a sixth number of virtual speakers, and the virtual speakers included in the sixth number of virtual speakers are the representative virtual speakers of the previous frame used to encode the previous frame of the three-dimensional audio signal. The first correlation is used to represent the priority of reusing the representative virtual speaker set of the previous frame when encoding the current frame. The priority can also be replaced by a preference, that is, the first correlation is used to determine whether to reuse the representative virtual speaker set of the previous frame when encoding the current frame. It can be understood that the greater the first correlation of the representative virtual speaker set of the previous frame, the higher the priority or preference of the representative virtual speaker set of the previous frame, and the encoder 113 is more inclined to select the representative virtual speaker of the previous frame to encode the current frame.

[0211] S660: The encoder 113 determines whether the first correlation satisfies a multiplexing condition.

[0212] If the first correlation does not meet the multiplexing condition, it indicates that the encoder 113 prefers to perform a virtual speaker search and encode the current frame according to the representative virtual speaker of the current frame. S610 is executed, and the encoder 113 obtains a first number of initial voting values ​​of the current frame corresponding to the first number of virtual speakers and the current frame of the three-dimensional audio signal.

[0213] Optionally, the encoder 113 may also select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain eigenvalues ​​of the fourth number of coefficients, and then use the largest representative coefficient among the third number of representative coefficients as the coefficient of the current frame for obtaining the first correlation. The encoder 113 then obtains the first correlation between the largest representative coefficient among the third number of representative coefficients of the current frame and the representative virtual speaker set of the previous frame. If the first correlation does not meet the multiplexing condition, S6103 is executed, that is, the encoder 113 selects a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on the first number of voting values.

[0214] If the first correlation satisfies the multiplexing condition, it means that the encoder 113 prefers to select the representative virtual speaker of the previous frame to encode the current frame, and the encoder 113 executes S670 and S680.

[0215] S670 : The encoder 113 generates a virtual speaker signal according to the representative virtual speaker set of the previous frame and the current frame.

[0216] S680: The encoder 113 encodes the virtual speaker signal to obtain a code stream.

[0217] The method for selecting a virtual speaker provided in an embodiment of the present application uses the correlation between the representative coefficient of the current frame and the representative virtual speaker of the previous frame to determine whether to perform a virtual speaker search. While ensuring the accuracy of the selection of the correlation of the representative virtual speaker of the current frame, it effectively reduces the complexity of the encoding end.

[0218] It is understood that in order to implement the functions in the above embodiments, the encoder includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily appreciate that, in conjunction with the various exemplary units and method steps described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0219] The above describes in detail the 3D audio signal encoding method provided by this embodiment in conjunction with FIG. 1 to FIG. 9 . The following describes the 3D audio signal encoding apparatus and encoder provided by this embodiment in conjunction with FIG. 10 and FIG. 11 .

[0220] Figure 10 is a schematic diagram of the structure of a possible 3D audio signal encoding device provided in this embodiment. These 3D audio signal encoding devices can be used to implement the function of encoding 3D audio signals in the above-mentioned method embodiments, thereby also achieving the beneficial effects of the above-mentioned method embodiments. In this embodiment, the 3D audio signal encoding device can be encoder 113 as shown in Figure 1, or encoder 300 as shown in Figure 3, or a module (such as a chip) applied to a terminal device or server.

[0221] As shown in FIG10 , a 3D audio signal encoding apparatus 1000 includes a communication module 1010, a coefficient selection module 1020, a virtual speaker selection module 1030, an encoding module 1040, and a storage module 1050. The 3D audio signal encoding apparatus 1000 is configured to implement the functions of the encoder 113 in the method embodiments shown in FIG6 to FIG9 .

[0222] The communication module 1010 is configured to obtain a current frame of a 3D audio signal. Optionally, the communication module 1010 may also receive the current frame of a 3D audio signal from another device, or obtain the current frame of a 3D audio signal from the storage module 1050. The current frame of the 3D audio signal is an HOA signal; the frequency domain eigenvalues ​​of the coefficients are determined based on the coefficients of the HOA signal.

[0223] The virtual speaker selection module 1030 is used to obtain a first number of current frame initial voting values ​​of the current frame of the three-dimensional audio signal. The first number of virtual speakers corresponds one-to-one to the current frame initial voting values. The first number of virtual speakers includes the first virtual speaker. The current frame initial voting value of the first virtual speaker is used to represent the priority of using the first virtual speaker when encoding the current frame.

[0224] The virtual speaker selection module 1030 is further used to obtain a seventh number of virtual speakers and a seventh number of final voting values ​​of the current frame corresponding to the current frame based on the first number of initial voting values ​​of the current frames and the sixth number of final voting values ​​of the previous frames, the seventh number of virtual speakers including the first number of virtual speakers, and the seventh number of virtual speakers including the sixth number of virtual speakers, the sixth number of virtual speakers corresponding one-to-one to the sixth number of final voting values ​​of the previous frame, and the sixth number of virtual speakers are virtual speakers used when encoding the previous frame of the three-dimensional audio signal.

[0225] If the first number of virtual speakers includes the second virtual speaker and the sixth number of virtual speakers does not include the second virtual speaker, the current frame final voting value of the second virtual speaker is equal to the current frame initial voting value of the second virtual speaker; or, if the sixth number of virtual speakers includes the third virtual speaker and the first number of virtual speakers does not include the third virtual speaker, the current frame final voting value of the third virtual speaker is equal to the previous frame final voting value of the third virtual speaker.

[0226] When the 3D audio signal encoding apparatus 1000 is used to implement the function of the encoder 113 in the method embodiments shown in FIG. 6 to FIG. 9 , the virtual speaker selection module 1030 is used to implement the related functions of S610 to S630 and S650 to S680 .

[0227] For example, when the virtual speaker selection module 1030 updates the initial voting value of the current frame of the first virtual speaker according to the final voting value of the previous frame of the first virtual speaker, it is specifically used to: adjust the final voting value of the previous frame of the first virtual speaker according to the first adjustment parameter to obtain the adjusted voting value of the previous frame of the first virtual speaker; and update the initial voting value of the current frame of the first virtual speaker according to the adjusted voting value of the previous frame of the first virtual speaker.

[0228] For another example, when the virtual speaker selection module 1030 updates the initial voting value of the current frame of the first virtual speaker according to the adjusted voting value of the previous frame of the first virtual speaker, it is specifically used to: adjust the initial voting value of the current frame of the first virtual speaker according to the second adjustment parameter to obtain the adjusted voting value of the current frame of the first virtual speaker; and update the adjusted voting value of the current frame of the first virtual speaker according to the adjusted voting value of the previous frame of the first virtual speaker.

[0229] The first adjustment parameter is determined based on at least one of the number of directional sound sources in a previous frame, a coding rate for encoding the current frame, and a frame type.

[0230] The second adjustment parameter is determined according to the adjusted voting value of the first virtual speaker in the previous frame and the initial voting value of the first virtual speaker in the current frame.

[0231] When the three-dimensional audio signal encoding apparatus 1000 is used to implement the functions of the encoder 113 in the method embodiment shown in FIG7 , the coefficient selection module 1020 is used to implement the related functions of S6101 and S6102 . Specifically, when the coefficient selection module 1020 obtains the third number of representative coefficients of the current frame, it is specifically configured to: obtain a fourth number of coefficients of the current frame and frequency domain eigenvalues ​​of the fourth number of coefficients; and select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain eigenvalues ​​of the fourth number of coefficients, where the third number is less than the fourth number.

[0232] The encoding module 1140 is configured to encode the current frame according to a second number of representative virtual speakers of the current frame to obtain a code stream.

[0233] When 3D audio signal encoding apparatus 1000 is configured to implement the functions of encoder 113 in the method embodiments shown in Figures 6 to 9 , encoding module 1140 is configured to implement the functions related to step S630 . For example, encoding module 1140 is configured to generate virtual speaker signals based on the second number of representative virtual speakers of the current frame and the current frame, and to encode the virtual speaker signals to generate a bitstream.

[0234] The storage module 1050 is used to store coefficients related to the three-dimensional audio signal, a set of candidate virtual speakers, a representative set of virtual speakers in the previous frame, and the selected coefficients and virtual speakers, etc., so that the encoding module 1040 can encode the current frame to obtain a bitstream and transmit the bitstream to the decoder.

[0235] It should be understood that the 3D audio signal encoding apparatus 1000 of the embodiment of the present application can be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Alternatively, when the 3D audio signal encoding method shown in FIG. 6 to FIG. 9 is implemented using software, the 3D audio signal encoding apparatus 1000 and its various modules can also be software modules.

[0236] A more detailed description of the communication module 1010 , coefficient selection module 1020 , virtual speaker selection module 1030 , encoding module 1040 and storage module 1050 can be directly obtained by referring to the relevant descriptions in the method embodiments shown in FIG. 6 to FIG. 9 , and is not repeated here.

[0237] FIG11 is a schematic diagram of the structure of an encoder 1100 provided in this embodiment. As shown in FIG11 , the encoder 1100 includes a processor 1110 , a bus 1120 , a memory 1130 , and a communication interface 1140 .

[0238] It should be understood that in this embodiment, the processor 1110 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0239] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits for controlling the execution of the program of the present application.

[0240] The communication interface 1140 is used to implement communication between the encoder 1100 and an external device or component. In this embodiment, the communication interface 1140 is used to receive a three-dimensional audio signal.

[0241] Bus 1120 may include a path for transmitting information between the aforementioned components (e.g., processor 1110 and memory 1130). In addition to a data bus, bus 1120 may also include a power bus, a control bus, and a status signal bus. However, for clarity, various buses are collectively labeled bus 1120 in the figure.

[0242] As an example, encoder 1100 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions). Processor 1110 may access coefficients related to a three-dimensional audio signal, candidate virtual speaker sets, a representative virtual speaker set from a previous frame, and selected coefficients and virtual speakers stored in memory 1130.

[0243] It is worth noting that Figure 11 only takes the encoder 1100 including 1 processor 1110 and 1 memory 1130 as an example. Here, the processor 1110 and the memory 1130 are respectively used to indicate a type of device or equipment. In a specific embodiment, the number of each type of device or equipment can be determined according to business requirements.

[0244] The memory 1130 may correspond to a storage medium for storing information such as coefficients related to three-dimensional audio signals, a set of candidate virtual speakers, a representative virtual speaker set of a previous frame, and selected coefficients and virtual speakers in the above-mentioned method embodiment, for example, a disk such as a mechanical hard disk or a solid-state drive.

[0245] The encoder 1100 can be a general-purpose device or a dedicated device. For example, the encoder 1100 can be an X86 or ARM-based server, or can be another dedicated server, such as a policy control and charging (PCC) server. The embodiment of the present application does not limit the type of the encoder 1100.

[0246] It should be understood that the encoder 1100 according to this embodiment may correspond to the 3D audio signal encoding apparatus 1100 in this embodiment, and may correspond to executing the corresponding subject in any of the methods in FIG. 6 to FIG. 9 , and the above-mentioned and other operations and / or functions of each module in the 3D audio signal encoding apparatus 1100 are for implementing the corresponding process of each method in FIG. 6 to FIG. 9 , respectively, and for the sake of brevity, they are not further described here.

[0247] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a network device or a terminal device. Of course, the processor and storage medium can also exist as discrete components in a network device or a terminal device.

[0248] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).

[0249] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A three-dimensional audio signal encoding method, characterized in that, include: Obtain a first number of initial voting values ​​for the current frame of the three-dimensional audio signal, wherein the first number of virtual speakers correspond one-to-one with the initial voting values ​​for the current frame, the first number of virtual speakers include a first virtual speaker, and the initial voting value for the current frame of the first virtual speaker is used to characterize the priority of the first virtual speaker; Based on the first number of initial voting values ​​of the current frame and the sixth number of final voting values ​​of the preceding frames, a seventh number of virtual speakers are obtained, corresponding to the seventh number of final voting values ​​of the current frame. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The sixth number of virtual speakers corresponds one-to-one with the sixth number of final voting values ​​of the preceding frames. The sixth number of virtual speakers are used to encode the preceding frames of the three-dimensional audio signal. Based on the final voting values ​​of the seventh number of current frames, a second number of representative virtual speakers for the current frames are selected from the seventh number of virtual speakers, where the second number is less than the seventh number; The current frame is encoded using the representative virtual speakers of the second number of current frames to obtain a bitstream.

2. The method according to claim 1, characterized in that, If the first number of virtual speakers includes the second virtual speaker, and the sixth number of virtual speakers does not include the second virtual speaker, then the final voting value of the second virtual speaker in the current frame is equal to the initial voting value of the second virtual speaker in the current frame; or If the sixth number of virtual speakers includes the third virtual speaker, and the first number of virtual speakers does not include the third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.

3. The method according to claim 1 or 2, characterized in that, If the sixth number of virtual speakers includes the first virtual speaker, obtaining the seventh number of final voting values ​​of the seventh virtual speaker corresponding to the current frame based on the first number of initial voting values ​​of the current frame and the sixth number of prior frame voting values ​​of the sixth number of virtual speakers and the prior frame of the three-dimensional audio signal includes: The initial voting value of the first virtual speaker in the current frame is updated based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.

4. The method according to claim 3, characterized in that, The step of updating the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame includes: The final voting value of the first virtual speaker in the previous frame is adjusted according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame. The initial voting value of the first virtual speaker in the current frame is updated based on the adjusted voting value of the first virtual speaker in the previous frame.

5. The method according to claim 4, characterized in that, The step of updating the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame includes: The initial voting value of the first virtual speaker in the current frame is adjusted according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame. The current frame adjusted voting value of the first virtual speaker is updated based on the previous frame adjusted voting value of the first virtual speaker.

6. The method according to claim 4 or 5, characterized in that, The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type of the current frame.

7. The method according to claim 5, characterized in that, The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.

8. The method according to any one of claims 1-7, characterized in that, The second quantity is preset, or the second quantity is determined based on the current frame.

9. The method according to any one of claims 1-8, characterized in that, The process of obtaining the first number of initial voting values ​​for the current frame corresponding to the first number of virtual speakers and the current frame of the three-dimensional audio signal includes: The first number of virtual speakers and the first number of initial voting values ​​for the current frame are determined based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds. The candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.

10. The method according to claim 9, characterized in that, Before determining the first number of virtual speakers and the first number of initial voting values ​​for the current frame based on the third number of representative coefficients, the candidate virtual speaker set, and the number of voting rounds of the current frame, the method further includes: Obtain the fourth number of coefficients of the current frame, and the frequency domain feature values ​​of the fourth number of coefficients; Based on the frequency domain characteristic values ​​of the fourth number of coefficients, a third number of representative coefficients are selected from the fourth number of coefficients, wherein the third number is less than the fourth number.

11. The method according to claim 10, characterized in that, The method further includes: Obtain the first correlation between the current frame and the representative virtual speaker set of the previous frame, wherein the representative virtual speaker set of the previous frame includes the sixth number of virtual speakers, wherein the sixth number of virtual speakers are the representative virtual speakers of the previous frame used to encode the previous frame, and the first correlation is used to determine whether the representative virtual speaker set of the previous frame is reused when encoding the current frame; If the first correlation does not meet the reuse condition, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values ​​of the fourth number of coefficients.

12. The method according to any one of claims 1-11, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values ​​of the coefficients of the current frame are determined based on the coefficients of the HOA signal.

13. A three-dimensional audio signal encoding device, characterized in that, include: A virtual speaker selection module is used to obtain a first number of initial voting values ​​for the current frame of a three-dimensional audio signal, wherein the first number of virtual speakers corresponds one-to-one with the initial voting values ​​for the current frame, the first number of virtual speakers includes a first virtual speaker, and the initial voting value for the current frame of the first virtual speaker is used to characterize the priority of the first virtual speaker. The virtual speaker selection module is further configured to obtain, based on the first number of initial voting values ​​of the current frames and the sixth number of final voting values ​​of the preceding frames, the seventh number of virtual speakers corresponding to the current frame and the current frame, wherein the seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers, and the sixth number of virtual speakers corresponds one-to-one with the sixth number of final voting values ​​of the preceding frames, wherein the sixth number of virtual speakers are virtual speakers used when encoding the preceding frames of the three-dimensional audio signal; The virtual speaker selection module is further configured to select representative virtual speakers for a second number of current frames from the seventh number of virtual speakers based on the final voting values ​​of the seventh number of current frames, wherein the second number is less than the seventh number; An encoding module is used to encode the current frame according to the second number of representative virtual speakers of the current frame to obtain a bitstream.

14. The apparatus according to claim 13, characterized in that, If the first number of virtual speakers includes the second virtual speaker, and the sixth number of virtual speakers does not include the second virtual speaker, then the final voting value of the second virtual speaker in the current frame is equal to the initial voting value of the second virtual speaker in the current frame; or If the sixth number of virtual speakers includes the third virtual speaker, and the first number of virtual speakers does not include the third virtual speaker, then the final vote value of the third virtual speaker in the current frame is equal to the final vote value of the third virtual speaker in the previous frame.

15. The apparatus according to claim 13 or 14, characterized in that, If the sixth number of virtual speakers includes the first virtual speaker, when the virtual speaker selection module obtains the seventh number of virtual speakers and the seventh number of final voting values ​​of the current frame corresponding to the current frame based on the first number of initial voting values ​​of the current frame and the sixth number of previous frame voting values ​​of the sixth number of virtual speakers and the previous frame corresponding to the three-dimensional audio signal, it is specifically used for: The initial voting value of the first virtual speaker in the current frame is updated based on the final voting value of the first virtual speaker in the previous frame to obtain the final voting value of the first virtual speaker in the current frame.

16. The apparatus according to claim 15, characterized in that, When the virtual speaker selection module updates the initial voting value of the first virtual speaker in the current frame based on the final voting value of the first virtual speaker in the previous frame, it is specifically used for: The final voting value of the first virtual speaker in the previous frame is adjusted according to the first adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the previous frame. The initial voting value of the first virtual speaker in the current frame is updated based on the adjusted voting value of the first virtual speaker in the previous frame.

17. The apparatus according to claim 16, characterized in that, When the virtual speaker selection module updates the initial voting value of the first virtual speaker in the current frame based on the adjusted voting value of the first virtual speaker in the previous frame, it is specifically used for: The initial voting value of the first virtual speaker in the current frame is adjusted according to the second adjustment parameter to obtain the adjusted voting value of the first virtual speaker in the current frame. The current frame adjusted voting value of the first virtual speaker is updated based on the previous frame adjusted voting value of the first virtual speaker.

18. The apparatus according to claim 16 or 17, characterized in that, The first adjustment parameter is determined based on at least one of the number of directional sound sources in the previous frame, the encoding rate for encoding the current frame, and the frame type of the current frame.

19. The apparatus according to claim 17, characterized in that, The second adjustment parameter is determined based on the voting value of the first virtual speaker after adjustment in the previous frame and the initial voting value of the first virtual speaker in the current frame.

20. The apparatus according to any one of claims 13-19, characterized in that, The second quantity is preset, or the second quantity is determined based on the current frame.

21. The apparatus according to any one of claims 13-20, characterized in that, When the virtual speaker selection module obtains the initial voting values ​​of the first number of virtual speakers and the current frame corresponding to the current frame of the 3D audio signal, it is specifically used for: The first number of virtual speakers and the first number of initial voting values ​​for the current frame are determined based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds. The candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.

22. The apparatus according to claim 21, characterized in that, The device also includes a coefficient selection module; The coefficient selection module is used to obtain the fourth number of coefficients of the current frame and the frequency domain feature values ​​of the fourth number of coefficients. The coefficient selection module is further configured to select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain characteristic values ​​of the fourth number of coefficients, wherein the third number is less than the fourth number.

23. The apparatus according to claim 22, characterized in that, The virtual speaker selection module is also used for: Obtain a first correlation between the current frame and the representative virtual speaker set of the previous frame, wherein the representative virtual speaker set of the previous frame includes the sixth number of virtual speakers, and the virtual speakers included in the sixth number of virtual speakers are the representative virtual speakers of the previous frame used to encode the previous frame. The first correlation is used to determine whether the representative virtual speaker set of the previous frame is reused when encoding the current frame. If the first correlation does not meet the reuse condition, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values ​​of the fourth number of coefficients.

24. The apparatus according to any one of claims 13-23, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values ​​of the coefficients of the current frame are determined based on the coefficients of the HOA signal.

25. An encoder, characterized in that, The encoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the three-dimensional audio signal encoding method as described in any one of claims 1-12.

26. A system, characterized in that, The system includes an encoder as described in claim 25, and a decoder, wherein the encoder is used to perform the operational steps of the method according to any one of claims 1-12, and the decoder is used to decode the bitstream generated by the encoder.

27. A computer program, characterized in that, When the computer program is executed, it implements the three-dimensional audio signal encoding method as described in any one of claims 1-12.

28. A computer-readable storage medium, characterized in that, Includes computer software instructions; when the computer software instructions are executed in the encoder, the encoder causes the encoder to perform the three-dimensional audio signal encoding method as described in any one of claims 1-12.

29. A computer-readable storage medium, characterized in that, The bitstream obtained by the three-dimensional audio signal encoding method as described in any one of claims 1-12.