Three-dimensional audio signal coding method and apparatus, and encoder
Patent Information
- Application Number
- JP2023571255
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-17
- Filing Date
- 2022-05-07
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2042-05-07
AI Technical Summary
も実装することができる。本実施形態において、三次元オーディオ信号符号化装置は、図1に示されるエンコーダ113、もしくは図3に示されるエンコーダ300であってよく、または、端末デバイスもしくはサーバに対して適用されるモジュール(チップなど)であってよい。
Smart Images

Figure 0007915767000016 
Figure 0007915767000017 
Figure 0007915767000018
Abstract
Description
[Technical Field]
[0001] This application relates to the multimedia field, and more particularly to a method and apparatus for coding three-dimensional audio signals, as well as an encoder. [Background technology]
[0002] This application claims priority to Chinese Patent Application No. 202110536631.5, filed with the National Intellectual Property Administration of China on 17 May 2021, titled "THREE-DIMENSIONAL AUDIO SIGNAL CODING METHOD AND APPARATUS, AND ENCODER," which is incorporated herein by reference in its entirety.
[0003] With the rapid development of high-performance computers and signal processing technologies, listeners are increasingly demanding higher standards for voice and audio experiences. Immersive audio can satisfy these requirements. For example, three-dimensional audio technology is widely used in wireless communication (e.g., 4G / 5G) voice, virtual reality / augmented reality, media audio, and other applications. Three-dimensional audio technology is an audio technology that acquires, processes, transmits, renders, and reproduces real-world sound and three-dimensional sound field information to provide sound with a strong sense of space, immersion, and presence. This provides listeners with an extraordinary "immersive" auditory experience.
[0004] Generally, a data acquisition device (e.g., a microphone) collects a large amount of data, records three-dimensional sound field information, and transmits a three-dimensional audio signal to a playback device (e.g., a speaker or earphone), which then reproduces the three-dimensional audio. Because the amount of data for three-dimensional sound field information is large, a large amount of storage space is required to store the data, and high bandwidth is needed to transmit the three-dimensional audio signal. To solve the aforementioned problems, the three-dimensional audio signal can be compressed, and the compressed data can be stored or transmitted. Currently, encoders can compress three-dimensional audio signals by using multiple pre-configured virtual speakers. However, the computational complexity of performing compression coding on a three-dimensional audio signal by an encoder is high. Therefore, reducing the computational complexity of performing compression coding on a three-dimensional audio signal is an urgent issue that needs to be addressed. [Overview of the project]
[0005] This application provides a method and apparatus for coding a three-dimensional audio signal, as well as an encoder, for reducing the computational complexity of performing compression coding on a three-dimensional audio signal.
[0006] According to a first aspect, the present application provides a method for encoding a three-dimensional audio signal. The method may be performed by an encoder and specifically includes the following steps: After determining a first quantity of virtual speakers and a first quantity of voting values based on the current frame of the three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, the encoder selects a second quantity of representative virtual speakers for the current frame from the first quantity of virtual speakers based on the first quantity of voting values, and further encodes the current frame based on the second quantity of representative virtual speakers for the current frame to obtain a bitstream. The second quantity is less than the first quantity, which indicates that the second quantity of representative virtual speakers for the current frame are some virtual speakers in the candidate virtual speaker set. It can be understood that the virtual speakers correspond one-to-one with the voting values. For example, a first quantity virtual speaker includes the first virtual speaker, the first quantity vote value includes the first virtual speaker vote value, and the first virtual speaker corresponds to the first virtual speaker vote value. The first virtual speaker vote value represents the priority of using the first virtual speaker when the current frame is encoded. A candidate virtual speaker set includes a fifth quantity virtual speaker, the fifth quantity virtual speaker includes the first quantity virtual speaker, the first quantity is less than or equal to the fifth quantity, the voting round quantity is an integer greater than or equal to 1, and the voting round quantity is less than or equal to the fifth quantity.
[0007] Currently, in the process of searching for a virtual speaker, the encoder uses the result of a related calculation between the three-dimensional audio signal to be encoded and the virtual speaker as a selection measurement indicator for the virtual speaker. Furthermore, if the encoder transmits a virtual speaker for each coefficient, efficient data compression cannot be achieved, and a heavy computational load is imposed on the encoder. According to the method for selecting a virtual speaker provided in this embodiment of the present application, the encoder votes for each virtual speaker in the candidate virtual speaker set by substituting all the coefficients in the current frame with a small number of representative coefficients, and selects a representative virtual speaker for the current frame based on the vote value. Furthermore, the encoder performs compression coding on the three-dimensional audio signal to be encoded using the representative virtual speaker for the current frame, which not only effectively improves the compression rate for compressing or coding the three-dimensional audio signal but also reduces the computational complexity of searching for a virtual speaker by the encoder, thereby reducing the computational complexity of performing compression coding on the three-dimensional audio signal and reducing the computational load on the encoder.
[0008] The second quantity represents the number of representative virtual speakers for the current frame selected by the encoder. A larger second quantity indicates a larger number of representative virtual speakers and more sound field information in the three-dimensional audio signal for the current frame, while a smaller second quantity indicates a smaller number of representative virtual speakers and less sound field information in the three-dimensional audio signal for the current frame. Therefore, the second quantity can be set to control the number of representative virtual speakers for the current frame selected by the encoder. For example, the second quantity may be preset. Alternatively, the second quantity may be determined based on the current frame. For example, the value of the second quantity may be 1, 2, 4, or 8.
[0009] Specifically, the encoder may select a representative virtual speaker for a second quantity for the current frame using one of the following two methods:
[0010] Method 1: The encoder selecting a representative virtual speaker for a second quantity for the current frame from the virtual speakers for a first quantity based on the voting value of a first quantity particularly includes selecting a representative virtual speaker for a second quantity for the current frame from the virtual speakers for a first quantity based on the voting value of a first quantity and a preset threshold.
[0011] Method 2: The encoder selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers for the first quantity based on the voting value of the first quantity particularly includes determining the voting value of the second quantity from the voting value of the first quantity based on the voting value of the first quantity, and using the virtual speaker for the second quantity within the virtual speakers for the first quantity, which corresponds to the voting value of the second quantity, as the representative virtual speaker for the second quantity for the current frame.
[0012] Furthermore, the voting round quantity may be determined based on at least one of the following: the number of directional sound sources in the current frame of the three-dimensional audio signal, the coding rate at which the current frame is encoded, and the coding complexity at which the current frame is encoded. A larger voting round quantity indicates that the encoder can use a representative coefficient of a smaller quantity to perform multiple iterative votes against virtual speakers in a candidate virtual speaker set, and select a representative virtual speaker for the current frame based on the voting values in multiple voting rounds, thereby improving the accuracy of selecting a representative virtual speaker for the current frame.
[0013] In a possible implementation, the encoder may determine a first quantity of virtual speakers and a first quantity of votes based on the votes of all virtual speakers in the candidate virtual speaker set.
[0014] Specifically, when the first quantity is equal to the fifth quantity, that the encoder determines the first quantity of virtual speakers and the first quantity of voting values based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the number of voting rounds particularly comprises the following. Assuming that the encoder obtains a third quantity of representative coefficients of the current frame, and the third quantity of representative coefficients includes a first representative coefficient and a second representative coefficient, the encoder obtains a fifth quantity of first voting values corresponding to the fifth quantity of virtual speakers, the fifth quantity of first voting values being obtained by performing voting rounds of the number of voting rounds by using the first representative coefficient, and obtains a fifth quantity of second voting values corresponding to the fifth quantity of virtual speakers, the fifth quantity of second voting values being obtained by performing voting rounds of the number of voting rounds by using the second representative coefficient. The fifth quantity of first voting values includes a first voting value of a first virtual speaker, and the fifth quantity of second voting values includes a second voting value of the first virtual speaker. Further, the encoder obtains a respective voting value for each of the fifth quantity of virtual speakers based on the fifth quantity of first voting values and the fifth quantity of second voting values. The voting value of the first virtual speaker is obtained based on a sum of the first voting value of the first virtual speaker and the second voting value of the first virtual speaker, and it can be understood that the fifth quantity is equal to the first quantity. Therefore, for each coefficient of the current frame, the encoder votes for the fifth quantity of virtual speakers included in the candidate virtual speaker set, and comprehensively covers the fifth quantity of virtual speakers by using the voting values of the fifth quantity of virtual speakers included in the candidate virtual speaker set as a selection criterion, thereby ensuring the accuracy of the representative virtual speaker for the current frame that is selected by the encoder.
[0015] For example, the encoder obtaining the first vote value of the fifth quantity, which is the first vote value of the fifth quantity of the virtual speaker of the fifth quantity, obtained by performing a voting round of the voting round quantity by using a first representative coefficient, includes determining the first vote value of the fifth quantity based on the coefficient of the virtual speaker of the fifth quantity and the first representative coefficient.
[0016] In another possible implementation, the encoder may determine a first quantity of virtual speakers and a first quantity of votes based on the votes of several virtual speakers in a candidate set of virtual speakers.
[0017] Specifically, when the first quantity is less than or equal to the fifth quantity, and the first quantity of virtual speakers and the first quantity of voting values are determined based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the number of voting rounds, the differences from the foregoing possible implementation are as follows. After the encoder obtains the fifth quantity of first voting values and the fifth quantity of second voting values, the encoder selects an eighth quantity of virtual speakers from the fifth quantity of virtual speakers based on the fifth quantity of first voting values, where the eighth quantity is less than the fifth quantity, which indicates that the eighth quantity of virtual speakers are a part of the fifth quantity of virtual speakers; the encoder selects a ninth quantity of virtual speakers from the fifth quantity of virtual speakers based on the fifth quantity of second voting values, where the ninth quantity is less than the fifth quantity, which indicates that the ninth quantity of virtual speakers are a part of the fifth quantity of virtual speakers. Further, the encoder obtains a tenth quantity of third voting values for a tenth quantity of virtual speakers based on the first voting values of the eighth quantity of virtual speakers and the second voting values of the ninth quantity of virtual speakers, that is, the encoder obtains, through accumulation, the voting values of virtual speakers with the same index among the eighth quantity of virtual speakers and the ninth quantity of virtual speakers. Therefore, the encoder obtains the first quantity of virtual speakers and the first quantity of voting values based on the eighth quantity of first voting values, the ninth quantity of second voting values, and the tenth quantity of third voting values. It can be understood that the first quantity of virtual speakers include the eighth quantity of virtual speakers and the ninth quantity of virtual speakers. The eighth quantity of virtual speakers include the tenth quantity of virtual speakers, and the ninth quantity of virtual speakers include the tenth quantity of virtual speakers. The tenth quantity of virtual speakers include a second virtual speaker, and a third voting value of the second virtual speaker is obtained based on a sum of a first voting value of the second virtual speaker and a second voting value of the second virtual speaker, the tenth quantity is less than or equal to the eighth quantity, and the tenth quantity is less than or equal to the ninth quantity. Further, the tenth quantity may be an integer greater than or equal to 1.
[0018] Optionally, there are no virtual speakers with the same number for the virtual speaker of the eighth quantity and the virtual speaker of the ninth quantity; that is, the tenth quantity can be equal to 0. The encoder obtains the virtual speaker of the first quantity and the vote value of the first quantity based on the first vote value of the eighth quantity and the second vote value of the ninth quantity.
[0019] In this way, for each coefficient of the current frame, the encoder selects a voting value that is larger from the voting values of the fifth quantity virtual speaker included in the candidate virtual speaker set, and uses the voting value that is larger to determine the virtual speaker of the first quantity and the voting value of the first quantity, thereby reducing the computational complexity of searching for a virtual speaker by the encoder, while ensuring the accuracy of the representative virtual speaker of the current frame, which is the representative virtual speaker selected by the encoder.
[0020] Furthermore, the encoder obtaining a representative coefficient of the third quantity in the current frame includes obtaining the coefficient of the fourth quantity in the current frame, and the frequency domain feature value of the coefficient of the fourth quantity, and selecting a representative coefficient of the third quantity from the coefficient of the fourth quantity based on the frequency domain feature value of the coefficient of the fourth quantity, wherein the third quantity is less than the fourth quantity, and this indicates that the representative coefficient of the third quantity is part of the coefficient of the fourth quantity. The current frame of the three-dimensional audio signal may be a higher-order ambisonics (HOA) signal, and the frequency domain feature value of the coefficient of the current frame is determined based on the coefficient of the HOA signal.
[0021] In this way, the encoder selects several coefficients from all the coefficients in the current frame as representative coefficients, and uses a small number of representative coefficients to replace all the coefficients in the current frame, thereby selecting a representative virtual speaker from a candidate set of virtual speakers. This effectively reduces the computational complexity of searching for virtual speakers by the encoder, thereby reducing the computational complexity of performing compression coding on the three-dimensional audio signal and reducing the computational load on the encoder.
[0022] The encoder encoding the current frame based on a second quantity of representative virtual speakers for the current frame to obtain a bitstream includes: the encoder generating a virtual speaker signal based on the second quantity of representative virtual speakers for the current frame and the current frame, encoding the virtual speaker signal to obtain a bitstream.
[0023] Since the frequency domain feature values of the coefficients of the current frame represent the sound field features of the three-dimensional audio signal, the encoder selects a representative coefficient for the current frame, which has representative sound field components, based on the frequency domain feature values of the coefficients of the current frame. By using this representative coefficient, the representative virtual speaker for the current frame, selected from a candidate virtual speaker set, can fully represent the sound field features of the three-dimensional audio signal. This further improves the accuracy of the virtual speaker signal generated when the encoder compresses or encodes the three-dimensional audio signal to be encoded by using the representative virtual speaker for the current frame. In this way, the compression rate for compressing or encoding the three-dimensional audio signal is improved, thereby reducing the bandwidth occupied by the encoder to transmit the bitstream.
[0024] Optionally, before the encoder selects a representative coefficient of the third quantity from the coefficient of the fourth quantity based on the frequency domain feature value of the coefficient of the fourth quantity, the method further includes the steps of obtaining a first correlation between the current frame and a representative virtual speaker set for a previous frame, and, if the first correlation does not satisfy the reuse condition, obtaining the coefficient of the fourth quantity of the current frame of the three-dimensional audio signal and the frequency domain feature value of the coefficient of the fourth quantity. The representative virtual speaker set for a previous frame includes a virtual speaker of a sixth quantity, the virtual speaker included in the virtual speaker of the sixth quantity being a representative virtual speaker for a previous frame used to encode the previous frame of the three-dimensional audio signal, and the first correlation is used to determine whether to reuse the representative virtual speaker set for a previous frame when the current frame is encoded.
[0025] In this way, the encoder can first determine whether it is possible to reuse a representative set of virtual speakers set up for a previous frame to encode the current frame. If the encoder can reuse a representative set of virtual speakers set up for a previous frame to encode the current frame, the encoder does not perform the process of searching for virtual speakers, which effectively reduces the computational complexity of searching for virtual speakers by the encoder, thereby reducing the computational complexity of compressing the three-dimensional audio signal and reducing the computational load on the encoder. Furthermore, frequent changes in virtual speakers between different frames may be reduced, thereby reducing orientation continuity between frames, improving the audio stability of the reconstructed three-dimensional audio signal and ensuring the sound quality of the reconstructed three-dimensional audio signal. If the encoder cannot reuse a representative set of virtual speakers set up for a previous frame to encode the current frame, the encoder selects representative coefficients and uses the representative coefficients for the current frame to vote for each virtual speaker in the candidate set of virtual speakers, and selects a representative virtual speaker for the current frame based on the votes, thereby reducing the computational complexity of compressing the three-dimensional audio signal and reducing the computational load on the encoder.
[0026] Optionally, the encoder may select a representative virtual speaker for the current frame of the second quantity from the virtual speakers of the first quantity based on the vote value of the first quantity, which includes obtaining the final vote value of the seventh quantity for the current frame and the current frame, corresponding to the virtual speaker of the seventh quantity, based on the vote value of the first quantity and the final vote value of the sixth quantity for the previous frame, and selecting a representative virtual speaker for the current frame of the second quantity from the virtual speakers of the seventh quantity based on the final vote value of the seventh quantity for the current frame, wherein the second quantity is less than the seventh quantity, and this indicates that the representative virtual speaker of the second quantity for the current frame is part of the virtual speaker of the seventh quantity. The virtual speaker of the seventh quantity includes the virtual speaker of the first quantity, and the virtual speaker of the seventh quantity includes the virtual speaker of the sixth quantity, and the virtual speaker included in the virtual speaker of the sixth quantity is a representative virtual speaker for the previous frame, used to encode the previous frame of the three-dimensional audio signal. The virtual speaker for the sixth quantity included in the representative virtual speaker set for the previous frame corresponds one-to-one with the final vote value of the sixth quantity in the previous frame.
[0027] In the process of searching for virtual speakers, the positions of actual sound sources unnecessarily overlap with the positions of virtual speakers, so virtual speakers may not be able to form a one-to-one correspondence with actual sound sources. Furthermore, in complex real-world scenarios, a set with a limited number of virtual speakers may not be able to represent all sound sources in the sound field. In this case, the virtual speakers found in different frames may change frequently, and this change clearly affects the listener's auditory perception, resulting in noticeable discontinuities and noise in the three-dimensional audio signal obtained after decoding and reconstruction. According to the method for selecting a virtual speaker provided in this embodiment of the present application, a representative virtual speaker for a previous frame is inherited, specifically, for virtual speakers having the same number, the initial vote value for the current frame is adjusted by using the final vote value for the previous frame, resulting in the encoder having a greater tendency to select a representative virtual speaker for a previous frame, thereby reducing frequent changes in virtual speakers across different frames, increasing signal orientation continuity between frames, improving the audio stability of the reconstructed three-dimensional audio signal, and ensuring the sound quality of the reconstructed three-dimensional audio signal.
[0028] Optionally, this method further includes: the encoder may further collect the current frame of the three-dimensional audio signal, compress and encode the current frame of the three-dimensional audio signal to obtain a bitstream, and transmit the bitstream to the decoder.
[0029] According to a second aspect, the present application provides a three-dimensional audio signal coding apparatus comprising a module configured to perform a three-dimensional audio signal coding method according to either the first aspect or a possible design thereof. For example, the three-dimensional audio signal coding apparatus comprises a virtual speaker selection module and an coding module. The virtual speaker selection module is configured to determine a first quantity virtual speaker and a first quantity voting value based on the current frame of a three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, wherein a virtual speaker corresponds one-to-one with a voting value, the first quantity virtual speaker includes the first virtual speaker, the first quantity voting value includes the voting value of the first virtual speaker, the first virtual speaker corresponds to the voting value of the first virtual speaker, the voting value of the first virtual speaker represents the priority of using the first virtual speaker when the current frame is encoded, the candidate virtual speaker set includes a fifth quantity virtual speaker, the fifth quantity virtual speaker includes the first quantity virtual speaker, the voting round quantity is an integer greater than or equal to 1, and the voting round quantity is less than or equal to the fifth quantity. The virtual speaker selection module is further configured to select a representative virtual speaker for a second quantity for the current frame from among the virtual speakers of a first quantity, based on a voting value of a first quantity, where the second quantity is less than the first quantity. The encoding module is configured to encode the current frame based on the representative virtual speaker of the second quantity for the current frame to obtain a bitstream. These modules may perform the corresponding functions in the method example in the first embodiment. For further details, please refer to the detailed description in the method example. Details are not described again here.
[0030] According to a third aspect, the present application provides an encoder, the encoder comprising at least one processor and a memory, the memory being configured to store a group of computer instructions, and when executing a group of computer instructions, the processor performs the operational steps of a three-dimensional audio signal coding method according to the first aspect or one of possible implementations thereof.
[0031] According to a fourth aspect, the present application provides a system comprising an encoder according to a third aspect and a decoder. The encoder is configured to perform the operational steps of a three-dimensional audio signal coding method according to either the first aspect or a possible implementation thereof, and the decoder is configured to decode the bitstream generated by the encoder.
[0032] According to a fifth aspect, the application provides a computer-readable storage medium containing computer software instructions. When the computer software instructions are executed on an encoder, the encoder is enabled to perform operational steps of the method according to the first aspect or any possible implementation thereof.
[0033] According to a sixth aspect, the present application provides a computer program product. When the computer program product is executed on an encoder, the encoder is enabled to perform operational steps of the method according to the first aspect or any possible implementation thereof.
[0034] Based on the implementations provided in the embodiments described above, the implementations may be further combined to provide more implementations. [Brief explanation of the drawing]
[0035] [Figure 1] This is a schematic diagram of the structure of an audio coding system according to one embodiment of this application. [Figure 2] This is a schematic diagram illustrating a scenario of an audio coding system according to one embodiment of this application. [Figure 3] This is a schematic diagram of the structure of an encoder according to one embodiment of this application. [Figure 4] This is a schematic flowchart of a three-dimensional audio signal coding method according to one embodiment of this application. [Figure 5]This is a schematic flowchart of a method for selecting a virtual speaker according to one embodiment of this application. [Figure 6] This is a schematic flowchart of a three-dimensional audio signal coding method according to one embodiment of this application. [Figure 7A] This is a schematic flowchart of another method for selecting a virtual speaker according to one embodiment of this application. [Figure 7B] This is a schematic flowchart of another method for selecting a virtual speaker according to one embodiment of this application. [Figure 8] This is a schematic flowchart of another method for selecting a virtual speaker according to one embodiment of this application. [Figure 9] This is a schematic flowchart of another method for selecting a virtual speaker according to one embodiment of this application. [Figure 10] This is a schematic diagram of the structure of the encoding device according to this application. [Figure 11] This is a schematic diagram of the encoder structure according to this application. [Modes for carrying out the invention]
[0036] For a clear and concise explanation of the following embodiments, the relevant technologies will first be briefly described.
[0037] Sound is a continuous wave produced through the vibration of an object. The object that generates vibrations and emits sound waves is called a sound source. In the process by which sound waves propagate through a medium (such as air, a solid, or a liquid), the auditory organs of humans or animals can perceive sound.
[0038] The characteristics of sound waves include pitch, intensity, and timbre. Pitch indicates the height of a sound. Intensity indicates the volume of a sound, and may also be called loudness or volume, and is measured in decibels (dB). Timbre is also called sound quality.
[0039] The frequency of a sound wave determines its pitch, with higher frequencies indicating higher pitches. The number of times an object vibrates per second is called its frequency, which is measured in Hertz (Hz). The range of sound frequencies perceptible to the human ear is from 20 Hz to 20,000 Hz.
[0040] The amplitude of a sound wave determines its intensity; a larger amplitude indicates a louder sound. A shorter distance to the sound source also indicates a louder sound.
[0041] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, pulse waves, and others.
[0042] Sound can be classified into regular and irregular sounds based on the characteristics of its sound waves. Irregular sounds are sounds emitted through irregular vibrations of a sound source. Irregular sounds are, for example, noises that affect people's work, study, rest, etc. Regular sounds are sounds emitted through regular vibrations of a sound source. Regular sounds include speech and music. When sound is represented electrically, regular sounds are analog signals that change continuously in the time-frequency domain. Analog signals may also be called audio signals. Audio signals are information carriers that carry speech, music, and acoustic effects.
[0043] Human hearing has the ability to perceive the spatial dispersion of sound sources. Therefore, when listening to sound in space, listeners can perceive the direction of the sound in addition to sensing its pitch, intensity, and timbre.
[0044] As people pay increasing attention to their auditory system experience and demand higher quality requirements to enhance the sense of depth, presence, and spatiality of sound, three-dimensional audio technology is emerging. As a result, listeners not only perceive sounds emitted from sound sources in front, behind, to the left, and to the right, but also perceive that the space in which they are located is surrounded by a spatial sound field (or "sound field") generated by these sound sources, and that sound spreads around them. This creates an "immersive" sound effect that makes listeners feel as if they are in a movie theater, concert hall, or similar venue.
[0045] Three-dimensional audio technology assumes that the space outside the human ear is a system, and the signal received at the eardrum is a three-dimensional audio signal output after the sound emitted by the sound source has been filtered by the system outside the ear. For example, the system outside the human ear may be defined as a system impulse response h(n), an arbitrary sound source may be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal in the embodiments of this application may be a higher-order ambisonics (HOA) signal. Three-dimensional audio may also be referred to as three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio, binaural audio, etc.
[0046] It is well known that when sound waves propagate in an ideal medium, the wave quantity is k = w / c, and the angular frequency is w = 2πf, where f is the sound wave frequency and c is the speed of sound. The sound pressure p satisfies equation (1), and ∇² is the Laplace operator. ∇ 2 p+k 2 p=0 Equation (1)
[0047] The spatial system outside the human ear is assumed to be a sphere, the listener is located at the center of the sphere, and sound transmitted from outside the sphere is projected onto the sphere, filtering out sounds from outside the sphere. Assuming that sound sources are dispersed on the sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound sources. In other words, three-dimensional audio technology is a method for fitting the sound field. Specifically, the equation in equation (1) is solved in a spherical coordinate system. In the passive spherical region, the equation in equation (1) is solved as equation (2) below.
[0048]
number
[0049] However, r represents the radius of the sphere, θ represents the horizontal angle, φ represents the pitch angle, k represents the wave magnitude, s represents the amplitude of the ideal plane wave, and m represents the sequence number of the three-dimensional audio signal (or the sequence number of the HOA signal).
[0050]
number
[0051] This represents the spherical Bessel function, which is also called the radial basis function, and the first j represents the imaginary unit.
[0052]
number
[0053] It does not change with angle.
[0054]
number
[0055] This represents the spherical harmonics in the directions of θ and φ,
[0056]
number
[0057] This represents the spherical harmonics in the direction of the sound source, and the three-dimensional audio signal coefficients satisfy equation (3).
[0058]
number
[0059] Equation (3) can be substituted into equation (2), and equation (2) can be transformed into equation (4).
[0060]
number
[0061]
number
[0062] This represents an Nth-order three-dimensional audio signal coefficient and is used to approximately describe the sound field. The sound field is the region in which sound waves exist within a medium. N is an integer greater than or equal to 1, for example, the value of N is an integer ranging from 2 to 6. The three-dimensional audio signal coefficient in the embodiments of this application may be an HOA coefficient or an ambisonic coefficient.
[0063] A three-dimensional audio signal is an information carrier that carries spatial positional information of the sound source in the sound field and describes the sound field for the listener in space. Equation (4) shows that the sound field can be expanded onto a sphere according to spherical harmonics, that is, the sound field can be decomposed into a superposition of multiple plane waves. Therefore, the sound field described by a three-dimensional audio signal can be represented by a superposition of multiple plane waves, and the sound field can be reconstructed using three-dimensional audio signal coefficients.
[0064] Compared to a 5.1-channel audio signal or a 7.1-channel audio signal, an N-th order HOA signal is (N+1) 2 Because it has 3D channels, this HOA signal contains a large amount of data used to describe the spatial information of the sound field. When a collecting device (e.g., a microphone) transmits a three-dimensional audio signal to a playback device (e.g., a speaker), a large bandwidth must be consumed. Currently, encoders can compress the three-dimensional audio signal by using spatially squeezed surround audio coding (S3AC) or directional audio coding (DirAC) to obtain a bitstream and transmit the bitstream to the playback device. The playback device decodes the bitstream, reconstructs the three-dimensional audio signal, and plays the reconstructed three-dimensional audio signal. As a result, the amount of data in the three-dimensional audio signal transmitted to the playback device is reduced, and the occupied bandwidth is reduced. However, the computational complexity of compressing the three-dimensional audio signal by the encoder is high, and it occupies excessive computing resources of the encoder. Therefore, how to reduce the computational complexity of compressing the three-dimensional audio signal is an urgent issue that needs to be addressed.
[0065] Embodiments of this application provide audio coding techniques, and in particular three-dimensional audio coding techniques adapted to three-dimensional audio signals, and more specifically, coding techniques that enable fewer channels to represent three-dimensional audio signals to improve upon conventional audio coding systems. Video coding (or commonly referred to as coding) comprises two parts: video coding and video decoding. When performed at the source, audio coding typically involves processing (e.g., compressing) the original audio to reduce the amount of data required to represent the original audio, thereby enabling more efficient storage and / or transmission of the original audio. When performed at the destination, audio decoding typically involves inverse processing of the encoder to reconstruct the original audio. The coding and decoding parts may together be referred to as coding. The implementation of embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0066] Figure 1 is a schematic diagram of the structure of an audio coding system according to one embodiment of the present application. The audio coding system 100 includes a source device 110 and a destination device 120. The source device 110 is configured to perform compression coding on a three-dimensional audio signal to obtain a bitstream and transmit the bitstream to the destination device 120. The destination device 120 decodes the bitstream, reconstructs the three-dimensional audio signal, and plays back the reconstructed three-dimensional audio signal.
[0067] Specifically, the source device 110 includes an audio acquisition device 111, a preprocessor 112, an encoder 113, and a communication interface 114.
[0068] The audio acquisition device 111 is configured to acquire original audio. The audio acquisition device 111 may be any type of audio acquisition device and / or any type of audio generation device configured to collect sound in the real world. The audio acquisition device 111 is, for example, a computer audio processor configured to generate computer audio. Alternatively, the audio acquisition device 111 may be any type of memory or memory for storing audio. The audio includes sound in the real world, sound in a virtual scene (e.g., VR or augmented reality, AR) and / or any combination thereof.
[0069] The preprocessor 112 is configured to receive the original audio collected by the audio acquisition device 111, preprocess the original audio, and acquire a three-dimensional audio signal. For example, the preprocessing performed by the preprocessor 112 may include channel conversion, audio format conversion, noise reduction, etc.
[0070] Encoder 113 is configured to receive a three-dimensional audio signal generated by the preprocessor 112, compress the three-dimensional audio signal, and obtain a bitstream. For example, encoder 113 may include a spatial encoder 1131 and a core encoder 1132. The spatial encoder 1131 is configured to select (or "search") a virtual speaker from a candidate virtual speaker set based on the three-dimensional audio signal, and to generate a virtual speaker signal based on the three-dimensional audio signal and the virtual speaker. The virtual speaker signal may be called the playback signal. The core encoder 1132 is configured to encode the virtual speaker signal and obtain a bitstream.
[0071] The communication interface 114 is configured to receive the bitstream generated by the encoder 113 and send the bitstream to the destination device 120 via the communication channel 130, so that the destination device 120 can reconstruct a three-dimensional audio signal based on the bitstream.
[0072] The destination device 120 includes a player 121, a post-processor 122, a decoder 123, and a communication interface 124.
[0073] The communication interface 124 is configured to receive the bitstream sent by the communication interface 114 and transmit the bitstream to the decoder 123, which then reconstructs a three-dimensional audio signal based on the bitstream.
[0074] Communication interfaces 114 and 124 may be configured to send or receive the data associated with the original audio by using a direct communication link between the source device 110 and the destination device 120, for example, a direct wired connection or a direct wireless connection, or by using any type of network, for example, a wired network, a wireless network, or any combination thereof, any type of private network and a public network, or any combination of these types.
[0075] Both communication interface 114 and communication interface 124 may be configured as one-way or two-way communication interfaces, indicated by the corresponding communication channel 130 arrows in Figure 1, pointing from source device 110 to destination device 120, and may be configured to send and receive messages, etc., to establish a connection and to send and receive messages, etc., to acknowledge and exchange any other information related to the communication link and / or data transmission, e.g., coded bitstream transmission.
[0076] Decoder 123 is configured to decode a bitstream and reconstruct a three-dimensional audio signal. For example, decoder 123 includes a core decoder 1231 and a spatial decoder 1232. Core decoder 1231 is configured to decode a bitstream and obtain a virtual speaker signal. Spatial decoder 1232 is configured to reconstruct a three-dimensional audio signal based on a candidate virtual speaker set and virtual speaker signals to obtain a reconstructed three-dimensional audio signal.
[0077] The post-processor 122 receives the reconstructed three-dimensional audio signal generated by the decoder 123 and is configured to perform post-processing on the reconstructed three-dimensional audio signal. For example, post-processing performed by the post-processor 122 may include audio rendering, loudness normalization, user interaction, audio format conversion, noise reduction, and the like.
[0078] Player 121 is configured to play the reconstructed sound based on the reconstructed three-dimensional audio signal.
[0079] It should be noted that the audio acquisition device 111 and encoder 113 may be integrated into one physical device or disposed in different physical devices. This is not limited to this. For example, the source device 110 shown in Figure 1 includes the audio acquisition device 111 and encoder 113, which indicates that the audio acquisition device 111 and encoder 113 are integrated into one physical device. In this case, the source device 110 may be referred to as the acquisition device. For example, the source device 110 is a media gateway in a wireless access network, a media gateway in a core network, a transcoding device, a media resource server, an AR device, a VR device, a microphone, or another audio acquisition device. If the source device 110 does not include the audio acquisition device 111, it indicates that the audio acquisition device 111 and encoder 113 are two different physical devices, and the source device 110 may acquire the original audio from another device (e.g., an audio acquisition device or an audio storage device).
[0080] Furthermore, the player 121 and decoder 123 may be integrated into a single physical device or disposed on different physical devices. This is not limited to this. For example, the destination device 120 shown in Figure 1 includes the player 121 and decoder 123, indicating that the player 121 and decoder 123 are integrated on a single physical device. In this case, the destination device 120 may also be referred to as a playback device, which has the function of decoding and playing back the reconstructed audio. For example, the destination device 120 is a speaker, earphones, or another device that plays back audio. If the destination device 120 does not include the player 121, it indicates that the player 121 and decoder 123 are two different physical devices. After decoding the bitstream and reconstructing the three-dimensional audio signal, the destination device 120 sends the reconstructed three-dimensional audio signal to another playback device (e.g., a speaker or earphones), which then plays back the reconstructed three-dimensional audio signal.
[0081] Furthermore, Figure 1 shows that the source device 110 and the destination device 120 may be integrated into a single physical device, and alternatively, the source device 110 and the destination device 120 may be located on different physical devices. This is not limited to these.
[0082] For example, as shown in Figure 2(a), the source device 110 may be a microphone in a recording studio, and the destination device 120 may be a speaker. The source device 110 can collect the original audio of various instruments and transmit the original audio to a coding device. The coding device encodes and decodes the original audio to obtain a reconstructed three-dimensional audio signal, and the destination device 120 plays back the reconstructed three-dimensional audio signal. In another example, the source device 110 may be a microphone in a terminal device, and the destination device 120 may be an earphone. The source device 110 can collect ambient sound or audio synthesized by the terminal device.
[0083] As another example, as shown in Figure 2(b), the source device 110 and destination device 120 are integrated into a virtual reality (VR), augmented reality (AR), mixed reality (MR), or extended reality (XR) device. In this case, the VR / AR / MR / XR device has the ability to collect, play, and code the original audio. The source device 110 may collect sounds emitted by the user and sounds emitted by virtual objects in the virtual environment where the user is located.
[0084] In these embodiments, the source device 110 or its corresponding function, and the destination device 120 or its corresponding function, may be implemented by using the same hardware and / or software, by using separate hardware and / or software, or by using any combination thereof. As will be apparent to those skilled in the art, the presence and division of different units or functions in the source device 110 and / or destination device 120 shown in Figure 1 may vary depending on the actual device and application.
[0085] The structure of the audio coding system is merely an example for illustrative purposes. In some possible implementations, the audio coding system may further include other devices. For example, the audio coding system may further include an end-side device or a cloud-side device. After collecting the original audio, the source device 110 preprocesses the original audio to obtain a three-dimensional audio signal and transmits the three-dimensional audio to the end-side device or cloud-side device, which implements the functionality to code and decode the three-dimensional audio signal.
[0086] The audio signal coding method provided in the embodiments of this application is primarily applied to the encoder side. The structure of the encoder will be described in detail with reference to Figure 3. As shown in Figure 3, the encoder 300 includes a virtual speaker setting unit 310, a virtual speaker set generation unit 320, a coding analysis unit 330, a virtual speaker selection unit 340, a virtual speaker signal generation unit 350, and an encoding unit 360.
[0087] The virtual speaker setting unit 310 is configured to generate virtual speaker setting parameters based on encoder setting information and to acquire multiple virtual speakers. The encoder setting information includes, but is not limited to, the sequence of the three-dimensional audio signals (or usually referred to as the HOA sequence), the coding bitrate, and user-defined information. The virtual speaker setting parameters include, but are not limited to, the number of virtual speakers, the sequence of virtual speakers, and the position coordinates of the virtual speakers. For example, the number of virtual speakers may be 2048, 1669, 1343, 1024, 530, 512, 256, 128, or 64. The sequence of virtual speakers may be any one of sequences 2 through 6. The position coordinates of the virtual speakers include the horizontal angle and the pitch angle.
[0088] The virtual speaker setting parameters output by the virtual speaker setting unit 310 are used as inputs to the virtual speaker set generation unit 320.
[0089] The virtual speaker set generation unit 320 is configured to generate candidate virtual speaker sets based on virtual speaker setting parameters, wherein the candidate virtual speaker set includes multiple virtual speakers. Specifically, the virtual speaker set generation unit 320 determines the number of virtual speakers to be included in the candidate virtual speaker set based on the number of virtual speakers, and determines the coefficients of the virtual speakers based on the positional information (e.g., coordinates) and order of the virtual speakers. For example, a method for determining the coordinates of virtual speakers includes, but is not limited to, the following: multiple virtual speakers are generated according to an equidistant rule, or multiple virtual speakers that are not evenly distributed are generated based on the auditory principle, and then the coordinates of the virtual speakers are generated based on the number of virtual speakers.
[0090] The coefficients of the virtual speaker may be generated based on the aforementioned principle for generating three-dimensional audio signals. θ in equation (3) s and φ s These are set to the position coordinates of the virtual speaker,
[0091]
number
[0092] This represents the coefficients of a virtual speaker of order N. These virtual speaker coefficients may also be called ambisonic coefficients.
[0093] The coding analysis unit 330 is configured to perform coding analysis on a three-dimensional audio signal, for example, by analyzing the sound field dispersion features of the three-dimensional audio signal, i.e., features such as the number of sound sources, the directivity of the sound sources, and the dispersion of the sound sources.
[0094] The coefficients of multiple virtual speakers included in the candidate virtual speaker set output by the virtual speaker set generation unit 320 are used as input to the virtual speaker selection unit 340.
[0095] The sound field dispersion features of the three-dimensional audio signal, which are output by the coding analysis unit 330, are used as input to the virtual speaker selection unit 340.
[0096] The virtual speaker selection unit 340 is configured to determine a representative virtual speaker that matches the three-dimensional audio signal based on the three-dimensional audio signal to be encoded, the sound field dispersion characteristics of the three-dimensional audio signal, and the coefficients of multiple virtual speakers.
[0097] Without limitation, the encoder 300 in this embodiment of the present application may, alternatively, not include the coding analysis unit 330, specifically, the encoder 300 may not analyze the input signal, and the virtual speaker selection unit 340 may determine a representative virtual speaker through default settings. For example, the virtual speaker selection unit 340 may determine a representative virtual speaker that matches the three-dimensional audio signal based only on the three-dimensional audio signal and the coefficients of a plurality of virtual speakers.
[0098] The encoder 300 may use a three-dimensional audio signal acquired from a data acquisition device, or a three-dimensional audio signal synthesized using an artificial audio object, as its input. Furthermore, the three-dimensional audio signal input to the encoder 300 may be a time-domain three-dimensional audio signal or a frequency-domain three-dimensional audio signal. This is not limited to these.
[0099] The position information of a representative virtual speaker and the coefficients of a representative virtual speaker, output by the virtual speaker selection unit 340, are used as inputs to the virtual speaker signal generation unit 350 and the encoding unit 360.
[0100] The virtual speaker signal generation unit 350 is configured to generate a virtual speaker signal based on a three-dimensional audio signal and attribute information of a representative virtual speaker. The attribute information of a representative virtual speaker includes at least one of the following: the position information of the representative virtual speaker, the coefficients of the representative virtual speaker, and the coefficients of the three-dimensional audio signal. If the attribute information is the position information of the representative virtual speaker, the coefficients of the representative virtual speaker are determined based on the position information of the representative virtual speaker. If the attribute information includes the coefficients of the three-dimensional audio signal, the coefficients of the representative virtual speaker are determined based on the coefficients of the three-dimensional audio signal. Specifically, the virtual speaker signal generation unit 350 calculates the virtual speaker signal based on the coefficients of the three-dimensional audio signal and the coefficients of the representative virtual speaker.
[0101] For example, it is assumed that matrix A represents the coefficients of the virtual speaker and matrix X represents the coefficients of the HOA signal. Matrix X is the inverse of matrix A. The theoretical optimal solution w is obtained by using the least squares method, where w represents the virtual speaker signal. The virtual speaker signal satisfies equation (5). w=A -1 X formula (5)
[0102] A -1 represents the inverse matrix of matrix A. The size of matrix A is (M × C), where C represents the number of virtual speakers, M represents the number of sound channels in the Nth-order HOA signal, a represents the coefficients of the virtual speakers, the size of matrix X is (M × L), where L represents the number of coefficients of the HOA signal, and x represents the coefficients of the HOA signal. The coefficients of a typical virtual speaker may be the HOA coefficients of a typical virtual speaker, or the ambisonic coefficients of a typical virtual speaker. For example,
[0103]
number
[0104] And,
[0105]
number
[0106] That is the case.
[0107] The virtual speaker signal output by the virtual speaker signal generation unit 350 is used as input to the encoding unit 360.
[0108] The encoding unit 360 is configured to perform core encoding processing on the virtual speaker signal to obtain a bitstream. encoding The processing includes, but is not limited to, transformation, quantization, psychoacoustic modeling, noise shaping, bandwidth expansion, downmixing, arithmetic coding, and bitstream generation.
[0109] It should be noted that the spatial encoder 1131 may include a virtual speaker setting unit 310, a virtual speaker set generation unit 320, a coding analysis unit 330, a virtual speaker selection unit 340, and a virtual speaker signal generation unit 350, that is, the virtual speaker setting unit 310, the virtual speaker set generation unit 320, the coding analysis unit 330, the virtual speaker selection unit 340, and the virtual speaker signal generation unit 350 implement the functions of the spatial encoder 1131. The core encoder 1132 may include an encoding unit 360, that is, the encoding unit 360 implements the functions of the core encoder 1132.
[0110] The encoder shown in Figure 3 may generate one virtual speaker signal or multiple virtual speaker signals. Multiple virtual speaker signals may be acquired by the encoder shown in Figure 3 over multiple executions or in a single execution.
[0111] The process for coding a three-dimensional audio signal will be described below with reference to the attached drawings. Figure 4 is a schematic flowchart of a three-dimensional audio signal coding method according to one embodiment of this application. In this specification, the description is provided by using an example in Figure 1 in which the source device 110 and destination device 120 perform the three-dimensional audio signal coding process. As shown in Figure 4, the method includes the following steps.
[0112] S410: Source device 110 acquires the current frame of the three-dimensional audio signal.
[0113] As described in the embodiments above, if the source device 110 is equipped with an audio acquisition device 111, the source device 110 may acquire the original audio by using the audio acquisition device 111. Optionally, the source device 110 may instead receive the original audio acquired by another device, or acquire the original audio from memory in the source device 110 or from another memory. The original audio may include at least one of the following: real-world sounds acquired in real time, audio stored in the device, and audio synthesized from multiple audio sources. The method for acquiring the original audio and the type of original audio are not limited to these embodiments.
[0114] After acquiring the original audio, the source device 110 generates a three-dimensional audio signal based on the three-dimensional audio technology and the original audio to provide the listener with an "immersive" sound effect during the playback of the original audio. For specific methods of generating the three-dimensional audio signal, please refer to the description of the preprocessor 112 and the prior art in the embodiments described above.
[0115] Furthermore, audio signals are continuous analog signals. In the audio signal processing process, the audio signal may first be sampled in order to generate a digital signal of a frame sequence. A frame may contain multiple sampling points, or alternatively, a frame may be a sampling point obtained through sampling, or alternatively, a frame may contain subframes obtained by dividing a frame, or alternatively, a frame may be a subframe obtained by dividing a frame. For example, if the length of a frame is L sampling points and the frame is divided into N subframes, then each subframe corresponds to L / N sampling points. Audio coding typically means processing an audio frame sequence containing multiple sampling points.
[0116] An audio frame may include the current frame or a previous frame. As described in the embodiments of this application, the current frame or a previous frame may be a frame or a subframe. The current frame is the frame on which coding is performed at the present moment. A previous frame is a frame on which coding was performed at a moment prior to the present moment, and a previous frame may be a frame at one moment prior to the present moment or a series of frames at moments prior to the present moment. In this embodiment of this application, the current frame of a three-dimensional audio signal is the frame of the three-dimensional audio signal on which coding is performed at the present moment, and a previous frame is a frame of the three-dimensional audio signal on which coding was performed at a moment prior to the present time. The current frame of a three-dimensional audio signal may be the current frame to be encoded in the three-dimensional audio signal. The current frame of a three-dimensional audio signal may be abbreviated as the current frame, and a previous frame of a three-dimensional audio signal may be abbreviated as a previous frame.
[0117] S420: Source device 110 determines the candidate virtual speaker set.
[0118] In one case, a candidate virtual speaker set is pre-configured in the memory of the source device 110. The source device 110 can read the candidate virtual speaker set from memory. The candidate virtual speaker set includes multiple virtual speakers. The virtual speakers are , sky Virtually existing in the intersound field Represents a speaker The virtual speaker is configured to calculate a virtual speaker signal based on the three-dimensional audio signal, and as a result, the destination device 120 plays the reconstructed three-dimensional audio signal.
[0119] In another case, the virtual speaker configuration parameters are pre-configured in the memory of the source device 110. The source device 110 generates candidate virtual speaker sets based on the virtual speaker configuration parameters. Optionally, the source device 110 generates candidate virtual speaker sets in real time based on the capabilities of its computing resources (e.g., processor) and the characteristics of the current frame (e.g., channels and data volume).
[0120] For specific methods for generating candidate virtual speaker sets, please refer to the prior art and the descriptions of the virtual speaker setting unit 310 and virtual speaker set generation unit 320 in the embodiments described above.
[0121] S430: The source device 110 selects a representative virtual speaker for the current frame from a candidate virtual speaker set based on the current frame of the three-dimensional audio signal.
[0122] The source device 110 votes for a virtual speaker based on the coefficients of the current frame and the virtual speaker coefficients, and selects a representative virtual speaker for the current frame from the candidate virtual speaker set based on the vote value of the virtual speaker. The candidate virtual speaker set is searched for a limited number of representative virtual speakers for the current frame, and the limited number of representative virtual speakers are used as the virtual speaker that best matches the current frame to be encoded, thereby performing data compression on the three-dimensional audio signal to be encoded.
[0123] Figure 5 is a schematic flowchart of a method for selecting a virtual speaker according to one embodiment of the present application. The method procedure in Figure 5 illustrates the specific calculation process included in S430 in Figure 4. In this specification, the explanation is provided by using an example in which an encoder 113 in the source device 110 shown in Figure 1 performs virtual speaker selection processing. Specifically, the functionality of a virtual speaker selection unit 340 is implemented. As shown in Figure 5, the method includes the following steps.
[0124] S510: The encoder 113 obtains a representative coefficient for the current frame.
[0125] A representative coefficient may be a representative coefficient in the frequency domain or a representative coefficient in the time domain. A representative coefficient in the frequency domain may also be called a representative frequency or a representative coefficient of the spectrum in the frequency domain. A representative coefficient in the time domain may also be called a representative sampling point in the time domain. For a specific method of obtaining a representative coefficient for the current frame, please refer to the explanation of S6101 in Figure 7A.
[0126] S520: The encoder 113 selects a representative virtual speaker for the current frame from the candidate virtual speaker set, based on the voting value for the representative coefficient of the virtual speaker in the candidate virtual speaker set for the current frame, i.e., performs S440 to S460.
[0127] The encoder 113 votes for virtual speakers in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker, and selects (searches for) a representative virtual speaker for the current frame from the candidate virtual speaker set based on the final vote value of the virtual speaker for the current frame. For a specific method of selecting a representative virtual speaker for the current frame, please refer to the explanation of S610 and S620 in Figures 6 and 7A and 7B.
[0128] It should be noted that the encoder first scans the virtual speakers included in the candidate virtual speaker set and compresses the current frame by using a representative virtual speaker for the current frame selected from the candidate virtual speaker set. However, if the selection of virtual speakers for consecutive frames changes significantly, the sound image of the reconstructed three-dimensional audio signal is unstable and the sound quality of the reconstructed three-dimensional audio signal deteriorates. In this embodiment of the present application, the encoder 113 may obtain the final vote value for the virtual speaker for the current frame by updating the initial vote value for the virtual speaker included in the candidate virtual speaker set, which is the final vote value for the virtual speaker for the previous frame, based on the final vote value for the representative virtual speaker for the previous frame, and then select a representative virtual speaker for the current frame from the candidate virtual speaker set based on the final vote value for the virtual speaker for the current frame. In this way, the representative virtual speaker for the current frame is selected based on the representative virtual speaker for the previous frame. Therefore, when selecting a representative virtual speaker for the current frame, the encoder is more likely to select the same virtual speaker as the representative virtual speaker for the previous frame. This increases directional continuity between consecutive frames and overcomes the problem of the results of selecting virtual speakers for consecutive frames varying significantly. For this reason, this embodiment of the present application may further include S530.
[0129] S530: The encoder 113 adjusts the initial vote value of the virtual speaker in the candidate virtual speaker set for the current frame based on the final vote value of the representative virtual speaker for the previous frame, and obtains the final vote value of the virtual speaker for the current frame.
[0130] After the encoder 113 votes for a virtual speaker in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker to obtain the initial vote value for the virtual speaker for the current frame, the encoder 113 adjusts the initial vote value for the virtual speaker in the candidate virtual speaker set for the current frame based on the final vote value of the representative virtual speaker for the previous frame to obtain the final vote value for the virtual speaker for the current frame. The representative virtual speaker for the previous frame is the virtual speaker used by the encoder 113 when encoding the previous frame. For a specific method of adjusting the initial vote value for the virtual speaker in the candidate virtual speaker set for the current frame, please refer to the explanation of S6201 and S6202 in Figure 8.
[0131] In some embodiments, if the current frame is the first frame in the original audio, the encoder 113 performs S510 and S520. If the current frame is any frame after the second frame in the original audio, the encoder 113 may first determine whether to reuse a representative virtual speaker for a previous frame or to search for a virtual speaker in order to encode the current frame, thereby ensuring orientation continuity between consecutive frames and reducing coding complexity. This embodiment of the present application may further include S540.
[0132] S540: The encoder 113 determines whether to search for a virtual speaker based on the current frame and representative virtual speakers for previous frames.
[0133] If it is decided to search for a virtual speaker, encoder 113 performs S510 to S530. Optionally, encoder 113 may first perform S510. Encoder 113 obtains representative coefficients for the current frame. Based on the representative coefficients for the current frame and the representative virtual speaker coefficients for previous frames, encoder 113 decides whether to search for a virtual speaker. If it is decided to search for a virtual speaker, encoder 113 performs S520 to S530.
[0134] If it is decided not to search for a virtual speaker, encoder 113 performs S550.
[0135] S550: The encoder 113 decides to encode the current frame by reusing a representative virtual speaker for the previous frame.
[0136] The encoder 113 reuses a representative virtual speaker for a previous frame and the current frame to generate a virtual speaker signal, encodes the virtual speaker signal to obtain a bitstream, and sends the bitstream to the destination device 120, i.e., performs S450 and S460.
[0137] For specific methods on how to determine whether or not to search for a virtual speaker, please refer to the explanation of S640 to S670 in Figure 9.
[0138] S440: The source device 110 generates a virtual speaker signal based on the current frame of the three-dimensional audio signal and a representative virtual speaker for the current frame.
[0139] The source device 110 generates a virtual speaker signal based on the coefficients of the current frame and the coefficients of a representative virtual speaker for the current frame. For specific methods of generating the virtual speaker signal, please refer to the prior art and the description of the virtual speaker signal generation unit 350 in the embodiments described above.
[0140] S450: Source device 110 encodes the virtual speaker signal to obtain a bitstream.
[0141] The source device 110 can perform encoding operations, such as conversion or quantization, on the virtual speaker signal to generate a bitstream and perform data compression on the three-dimensional audio signal to be encoded. For specific methods of generating the bitstream, please refer to the prior art and the description of the encoding unit 360 in the above-described embodiment.
[0142] S460: Source device 110 sends a bitstream to destination device 120.
[0143] Source device 110 may encode the entire original audio and then send the bitstream of the original audio to destination device 120. Alternatively, source device 110 may encode the three-dimensional audio signal frame by frame in real time and then send the bitstream of the frames after encoding them. For specific methods of sending bitstreams, please refer to the prior art and the descriptions of communication interfaces 114 and 124 in the embodiments described above.
[0144] S470: The destination device 120 decodes the bitstream sent by the source device 110, reconstructs the three-dimensional audio signal, and obtains the reconstructed three-dimensional audio signal.
[0145] After receiving the bitstream, the destination device 120 decodes the bitstream to obtain the virtual speaker signal, and then, based on the candidate virtual speaker set and virtual speaker signal, reconstructs the three-dimensional audio signal to obtain the reconstructed three-dimensional audio signal. The destination device 120 plays the reconstructed three-dimensional audio signal. Alternatively, the destination device 120 sends the reconstructed three-dimensional audio signal to another playback device, which then plays the reconstructed three-dimensional audio signal to achieve a more vivid "immersive" sound effect that makes the listener feel as if they are in a movie theater, concert hall, virtual scene, etc.
[0146] Currently, in the process of searching for a virtual speaker, the encoder uses the result of a related calculation between the three-dimensional audio signal to be encoded and the virtual speaker as a selection measurement indicator for the virtual speaker. If the encoder transmits a virtual speaker for each coefficient, data compression cannot be achieved, and a heavy computational load is imposed on the encoder. One embodiment of the present application provides a method for selecting a virtual speaker. The encoder uses a representative coefficient of the current frame to vote for each virtual speaker in a set of candidate virtual speakers, and based on the vote value, selects a representative virtual speaker for the current frame, thereby reducing the computational complexity of searching for a virtual speaker and reducing the computational load on the encoder.
[0147] Referring to the attached drawings, the process for selecting a virtual speaker will be described in detail below. Figure 6 is a schematic flowchart of a three-dimensional audio signal coding method according to one embodiment of the present application. In this specification, the description is provided by using an example in which the encoder 113 in the source device 110 in Figure 1 performs the virtual speaker selection process. The method procedure in Figure 6 illustrates the specific calculation process included in S520 in Figure 5. As shown in Figure 6, the method includes the following steps.
[0148] S610: The encoder 113 determines the first quantity of virtual speakers and the first quantity of voting values based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity.
[0149] The voting round quantity is used to limit the number of votes a virtual speaker can receive. The voting round quantity is an integer greater than or equal to 1, and is less than or equal to the number of virtual speakers included in the candidate virtual speaker set, and is less than or equal to the number of virtual speaker signals transmitted by the encoder. For example, a candidate virtual speaker set includes a fifth quantity of virtual speakers, the fifth quantity of virtual speakers includes a first quantity of virtual speakers, the first quantity is less than or equal to the fifth quantity, the voting round quantity is an integer greater than or equal to 1, and the voting round quantity is less than or equal to the fifth quantity. The virtual speaker signal also refers to the transmission channel of the representative virtual speaker for the current frame, corresponding to the current frame. Generally, the number of virtual speaker signals is less than or equal to the number of virtual speakers.
[0150] In possible implementations, the voting round quantity may be pre-configured or determined based on the encoder's computing power. For example, the voting round quantity may be determined based on the coding rate and / or coding application scenario at which the encoder encodes the current frame.
[0151] For example, if the encoder's coding rate is low (e.g., the third-order HOA signal is encoded and transmitted at a rate of 128kbps or less), the number of voting rounds is 1; if the encoder's coding rate is medium (e.g., the third-order HOA signal is encoded and transmitted at a rate ranging from 192kbps to 512 kbps), the number of voting rounds is 4; or if the encoder's coding rate is high (e.g., the third-order HOA signal is encoded and transmitted at a rate of 768kbps or higher), the number of voting rounds is 7.
[0152] As another example, if the encoder is used for real-time communication, the coding complexity should be low and the number of voting rounds should be 1; if the encoder is used to broadcast streaming media, the coding complexity should be medium and the number of voting rounds should be 2; or if the encoder is used for high-quality data storage, the coding complexity should be high and the number of voting rounds should be 6.
[0153] As another example, if the encoder's coding rate is 128kbps and the coding complexity requirement is low, the voting round quantity is 1.
[0154] In another possible implementation, the voting round quantity is determined based on the number of directional sources in the current frame. For example, if the number of directional sources in the sound field is 2, the voting round quantity is set to 2.
[0155] This embodiment of the present application provides three possible implementations for determining a first quantity virtual speaker and a first quantity voting value. The three methods are described separately in detail below.
[0156] In the first possible implementation, the voting round quantity is equal to 1, and after sampling several representative coefficients, the encoder 113 obtains the voting values of all virtual speakers in the candidate virtual speaker set for each representative coefficient of the current frame, and accumulates the voting values of virtual speakers having the same number to obtain a first quantity of virtual speakers and a first quantity of voting values. For example, see the following explanation of S6101 to S6105 in Figure 7A.
[0157] It can be understood that the candidate virtual speaker set includes a first quantity of virtual speakers. The first quantity of virtual speakers is equal to the number of virtual speakers included in the candidate virtual speaker set. Assuming that the candidate virtual speaker set includes a fifth quantity of virtual speakers, the first quantity is equal to the fifth quantity. The voting value of the first quantity includes the voting values of all virtual speakers in the candidate virtual speaker set. The encoder 113 may perform S620 using the voting value of the first quantity as the final voting value of the first quantity of virtual speakers, which corresponds to the current frame, specifically, the encoder 113 selects a representative virtual speaker of the second quantity for the current frame from the first quantity of virtual speakers based on the voting value of the first quantity.
[0158] A virtual speaker has a one-to-one correspondence with a vote value; that is, one virtual speaker corresponds to one vote value. For example, the first quantity's virtual speaker includes the first virtual speaker, the first quantity's vote value includes the first virtual speaker's vote value, and the first virtual speaker corresponds to the first virtual speaker's vote value. The first virtual speaker's vote value represents the priority of using the first virtual speaker when the current frame is encoded. Priority may be replaced with a tendency, specifically, the first virtual speaker's vote value represents the tendency to use the first virtual speaker when the current frame is encoded. A higher vote value for the first virtual speaker indicates a higher priority or greater tendency towards the first virtual speaker, and it can be understood that the encoder 113 is more likely to select the first virtual speaker and encode the current frame compared to virtual speakers in the candidate virtual speaker set whose vote value is less than that of the first virtual speaker.
[0159] In the second possible implementation, the difference from the first possible implementation is as follows: After obtaining the vote values of all virtual speakers in the candidate virtual speaker set for each representative coefficient of the current frame, the encoder 113 selects some vote values from the vote values of all virtual speakers in the candidate virtual speaker set for each representative coefficient, and for the virtual speakers corresponding to those vote values, accumulates the vote values of virtual speakers that have the same numbers to obtain a first quantity of virtual speakers and a first quantity of vote values. It can be understood that the first quantity is less than or equal to the number of virtual speakers included in the candidate virtual speaker set. The vote values of the first quantity include the vote values of some virtual speakers included in the candidate virtual speaker set, or the vote values of the first quantity include the vote values of all virtual speakers included in the candidate virtual speaker set. For example, see the explanation of S6101 to S6104 and S6106 to S6110 in Figures 7A and 7B.
[0160] In the third possible implementation, the difference from the second possible implementation is as follows: The voting round quantity is an integer greater than or equal to 2, and for each representative coefficient of the current frame, encoder 113 casts at least two rounds of votes for all virtual speakers in the candidate virtual speaker set, selecting the virtual speaker with the highest vote value in each round. After at least two rounds of voting have been performed for all virtual speakers for each representative coefficient of the current frame, the vote values of virtual speakers with the same number are accumulated to obtain a first quantity of virtual speakers and a first quantity of vote values.
[0161] The voting round quantity is assumed to be 2, the fifth quantity of virtual speakers is assumed to include the first virtual speaker, the second virtual speaker, and the third virtual speaker, and the representative coefficient of the current frame is assumed to include the first representative coefficient and the second representative coefficient.
[0162] The encoder 113 first casts two rounds of votes for the three virtual speakers based on a first representative coefficient. In the first voting round, the encoder 113 casts votes for the three virtual speakers based on the first representative coefficient. Assuming that the maximum vote value is for the first virtual speaker, the first virtual speaker is selected. In the second voting round, the encoder 113 casts votes separately for the second and third virtual speakers based on the first representative coefficient. Assuming that the maximum vote value is for the second virtual speaker, the second virtual speaker is selected.
[0163] Furthermore, the encoder 113 casts two rounds of votes for the three virtual speakers based on a second representative coefficient. In the first voting round, the encoder 113 votes for the three virtual speakers based on the second representative coefficient. Assuming that the maximum vote value is for the second virtual speaker, the second virtual speaker is selected. In the second voting round, the encoder 113 votes separately for the first and third virtual speakers based on the second representative coefficient. Assuming that the maximum vote value is for the third virtual speaker, the third virtual speaker is selected.
[0164] Finally, the first quantity of virtual speakers includes a first virtual speaker, a second virtual speaker, and a third virtual speaker. The vote value of the first virtual speaker is equal to the vote value of the first virtual speaker for the first representative coefficient in the first voting round. The vote value of the second virtual speaker is equal to the sum of the vote value of the second virtual speaker for the first representative coefficient in the second voting round and the vote value of the second virtual speaker for the second representative coefficient in the first voting round. The vote value of the third virtual speaker is equal to the vote value of the third virtual speaker for the second representative coefficient in the second voting round.
[0165] S620: The encoder 113 selects a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity, based on the voting value of the first quantity.
[0166] The encoder 113 selects a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity, based on the voting value of the first quantity. Furthermore, the voting value of the representative virtual speaker for the second quantity for the current frame is greater than a preset threshold.
[0167] The encoder 113 may, as an alternative, select a representative virtual speaker for the second quantity for the current frame from the virtual speakers for the first quantity based on the voting value of the first quantity. For example, the voting value of the second quantity is determined from the voting values of the first quantity in descending order, and the virtual speaker in the virtual speakers for the first quantity that corresponds to the voting value of the second quantity is used as the representative virtual speaker for the second quantity for the current frame.
[0168] If, at the discretion of the first quantity of virtual speakers, virtual speakers with different numbers have the same voting value, and the voting value of the virtual speakers with different numbers is greater than a preset threshold, the encoder 113 may use all virtual speakers with different numbers as representative virtual speakers for the current frame.
[0169] It should be noted that the second quantity is less than the first quantity. The virtual speaker for the first quantity includes a representative virtual speaker for the second quantity for the current frame. The second quantity may be predetermined, or it may be determined based on the quantity of sound sources in the sound field of the current frame. For example, the second quantity may be directly equal to the quantity of sound sources in the sound field of the current frame, or the quantity of sound sources in the sound field of the current frame may be processed based on a predetermined algorithm, and the quantity obtained through the processing may be used as the second quantity. The predetermined algorithm may be designed based on requirements. For example, the predetermined algorithm may be second quantity = quantity of sound sources in the sound field of current frame + 1, or second quantity = quantity of sound sources in the sound field of current frame - 1.
[0170] S630: The encoder 113 encodes the current frame based on a second quantity of representative virtual speakers for the current frame to obtain a bitstream.
[0171] The encoder 113 generates a virtual speaker signal based on a second quantity of representative virtual speakers for the current frame and the current frame, encodes the virtual speaker signal, and obtains a bitstream.
[0172] The encoder selects several coefficients as representative coefficients from all the coefficients in the current frame, and uses a small number of representative coefficients to replace all the coefficients in the current frame, thereby selecting a representative virtual speaker from a candidate set of virtual speakers. This effectively reduces the computational complexity of searching for virtual speakers by the encoder, thereby reducing the computational complexity of compressing the three-dimensional audio signal and lowering the computational load on the encoder. For example, the frame of an N-th order HOA signal is 960·(N+1) 2has coefficients. In this embodiment, the first 10% of coefficients may be selected to participate in the search for virtual speakers. In this case, the coding complexity is reduced by 90% compared with the coding complexity generated when all coefficients participate in the search for virtual speakers.
[0173] FIG. 7A and FIG. 7B are schematic flowcharts of another method for selecting a virtual speaker according to an embodiment of the present application. The method steps in FIG. 7A and FIG. 7B illustrate a specific calculation process included in S610 in FIG. 6. It is assumed that the candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first virtual speaker.
[0174] S6101: an encoder 113 obtains a fourth number of coefficients of a current frame, and frequency domain feature values of the fourth number of coefficients.
[0175] the three-dimensional audio signal is a higher-order ambisonics (HOA) signal, the encoder 113 samples the current frame of the HOA signal to obtain L·(N+1) 2 sampling points, that is, it is assumed that the fourth number of coefficients are obtained. N is the order of the HOA signal. For example, the duration of the current frame of the HOA signal is 20 milliseconds, the encoder 113 samples the current frame at a frequency of 48 kHz to obtain 960·(N+1) 2 sampling points. The sampling points may also be referred to as time domain coefficients.
[0176] The frequency-domain coefficients of the current frame of a three-dimensional audio signal can be obtained by performing a time-frequency transformation based on the time-domain coefficients of the current frame of the three-dimensional audio signal. The method for the time-domain to frequency-domain transformation is not limited. For example, a method for the time-domain to frequency-domain transformation is the Modified Discrete Cosine Transform (MDCT), which is 960·(N+1) in the frequency domain. 2 A number of frequency-domain coefficients can be obtained. These frequency-domain coefficients may also be referred to as spectral coefficients or frequencies.
[0177] The frequency domain feature value of a sampling point satisfies p(j) = norm(x(j)), where j = 1, 2, ..., and L, where L represents the quantity at the sampling moment, x represents the frequency domain coefficient of the current frame of the three-dimensional audio signal, e.g., the MDCT coefficient, "norm" is the operation of solving the 2-norm, and x(j) is (N+1) at the j-th sampling moment. 2 This represents the frequency domain coefficients for each sampling point.
[0178] S6102: The encoder 113 selects a representative coefficient of the third quantity from the coefficients of the fourth quantity based on the frequency domain feature values of the coefficients of the fourth quantity.
[0179] Encoder 113 divides the spectral range indicated by the coefficient of the fourth quantity into at least one subband. Encoder 113 divides the spectral range indicated by the coefficient of the fourth quantity into one subband. It can be understood that the spectral range of the subband is equal to the spectral range indicated by the coefficient of the fourth quantity, which is equivalent to encoder 113 not dividing the spectral range indicated by the coefficient of the fourth quantity.
[0180] When encoder 113 divides the spectral range indicated by the coefficient of the fourth quantity into at least two sub-frequency bands, in one case encoder 113 divides the spectral range indicated by the coefficient of the fourth quantity equally into at least two sub-bands, wherein all sub-bands in at least two sub-bands contain the same quantity of coefficient.
[0181] In another case, encoder 113 unevenly divides the spectral range indicated by the coefficient of the fourth quantity, so that at least two subbands obtained through the division contain different quantities of the coefficient, or all subbands in at least two subbands obtained through the division contain different quantities of the coefficient. For example, encoder 113 may unevenly divide the spectral range indicated by the coefficient of the fourth quantity based on a low-frequency range, an intermediate-frequency range, and a high-frequency range within the spectral range indicated by the coefficient of the fourth quantity, so that each spectral range in the low-frequency range, intermediate-frequency range, and high-frequency range contains at least one subband. All subbands in at least one subband in the low-frequency range contain the same quantity of the coefficient, all subbands in at least one subband in the intermediate-frequency range contain the same quantity of the coefficient, and all subbands in at least one subband in the high-frequency range contain the same quantity of the coefficient. The subbands in the three spectral ranges, i.e., the low-frequency range, the intermediate-frequency range, and the high-frequency range, may contain different quantities of the coefficient.
[0182] Furthermore, based on the frequency domain feature values of the coefficients of the fourth quantity, the encoder 113 selects a representative coefficient from at least one subband included in the spectral range represented by the coefficients of the fourth quantity to obtain a representative coefficient of the third quantity. The third quantity is less than the fourth quantity, and the coefficients of the fourth quantity include a representative coefficient of the third quantity.
[0183] For example, the encoder 113 selects Z representative coefficients from each subband in descending order of frequency domain feature values of the coefficients in the subband within at least one subband that is included in the spectral range indicated by the coefficient of the fourth quantity, and combines the Z representative coefficients from at least one subband to obtain a representative coefficient of the third quantity, where Z is a positive integer.
[0184] As another example, if at least one subband includes at least two subbands, encoder 113 determines the weight of each of the at least two subbands based on the frequency domain feature values of the first candidate coefficients in that subband, and adjusts the frequency domain feature values of the second candidate coefficients in each subband based on the weight of that subband to obtain the adjusted frequency domain feature values of the second candidate coefficients in each subband, where the first and second candidate coefficients are partial coefficients in the subband. Encoder 113 determines a representative coefficient of a third quantity based on the adjusted frequency domain feature values of the second candidate coefficients in at least two subbands and the frequency domain feature values of the coefficients other than the second candidate coefficients in at least two subbands.
[0185] The encoder selects several coefficients from all the coefficients in the current frame as representative coefficients, and uses a small number of these representative coefficients to replace all the coefficients in the current frame, thereby selecting a representative virtual speaker from a candidate set of virtual speakers. This effectively reduces the computational complexity of searching for virtual speakers by the encoder, thereby reducing the computational complexity of compressing the three-dimensional audio signal and lowering the encoder's computational load.
[0186] Assuming that the representative coefficient of the third quantity includes the first representative coefficient and the second representative coefficient, steps S6103 to S6110 are performed.
[0187] S6103: The encoder 113 obtains the first vote value of the fifth quantity, which is the first vote value of the fifth quantity of the virtual speaker of the fifth quantity, and is obtained by performing a voting round of the voting round quantity by using the first representative coefficient.
[0188] The encoder 113 uses a first representative coefficient to represent the current frame and votes that the current frame will be encoded by using a fifth quantity virtual speaker, and determines a first vote value for the fifth quantity based on the coefficient of the fifth quantity virtual speaker and the first representative coefficient. The first vote value for the fifth quantity includes the first vote value for the first virtual speaker.
[0189] S6104: The encoder 113 obtains the second vote value of the fifth quantity, which is the second vote value of the fifth quantity of the virtual speaker of the fifth quantity, obtained by performing a voting round of the voting round quantity using a second representative coefficient.
[0190] The encoder 113 uses a second representative coefficient to represent the current frame and votes that the current frame will be encoded by using a fifth quantity virtual speaker, and determines a second vote value for the fifth quantity based on the coefficient of the fifth quantity virtual speaker and the second representative coefficient. The second vote value for the fifth quantity includes the second vote value for the first virtual speaker.
[0191] S6105: The encoder 113 obtains the respective vote values for the virtual speaker of the fifth quantity and the vote values for the first quantity, based on the first vote value of the fifth quantity and the second vote value of the fifth quantity.
[0192] For the fifth virtual speaker, the encoder 113 stores the first and second vote values for virtual speakers that have the same number. The vote value for the first virtual speaker is equal to the sum of the first vote value and the second vote value of the first virtual speaker. For example, the first vote value for the first virtual speaker is 10, the second vote value for the first virtual speaker is 15, and the total vote value for the first virtual speaker is 25.
[0193] The fifth quantity is equal to the first quantity, and it can be understood that the virtual speaker of the first quantity obtained after the encoder 113 has cast its vote is the virtual speaker of the fifth quantity. The vote value of the first quantity is the vote value of the virtual speaker of the fifth quantity.
[0194] Therefore, the encoder votes for each coefficient of the current frame on a fifth quantity virtual speaker included in the candidate virtual speaker set, and uses the vote values of the fifth quantity virtual speakers included in the candidate virtual speaker set as selection criteria to comprehensively cover all fifth quantity virtual speakers, thereby ensuring the accuracy of the representative virtual speaker for the current frame, which is the representative virtual speaker selected by the encoder.
[0195] In some other embodiments, the encoder may determine a first quantity of virtual speakers and a first quantity of voting values based on the voting values of several virtual speakers in a candidate set of virtual speakers. Following S6103 and S6104, this embodiment of the present application may further include S6106 to S6110.
[0196] S6106: The encoder 113 selects a virtual speaker for the 8th quantity from the virtual speakers for the 5th quantity based on the first vote value for the 5th quantity.
[0197] Encoder 113 sorts the first vote values of the fifth quantity and selects the virtual speaker for the eighth quantity from the virtual speakers of the fifth quantity in descending order of the first vote values of the fifth quantity, starting with the largest first vote value. The eighth quantity is less than the fifth quantity. The first vote value of the fifth quantity includes the first vote value of the eighth quantity. The eighth quantity is an integer greater than or equal to 1.
[0198] S6107: The encoder 113 selects a virtual speaker for the 9th quantity from the virtual speakers for the 5th quantity based on the second vote value of the 5th quantity.
[0199] Encoder 113 sorts the second vote values of the fifth quantity and selects the virtual speaker for the ninth quantity from the virtual speakers of the fifth quantity in descending order of the second vote values of the fifth quantity, starting with the largest second vote value. The ninth quantity is less than the fifth quantity. The second vote value of the fifth quantity includes the second vote value of the ninth quantity. The ninth quantity is an integer greater than or equal to 1.
[0200] S6108: The encoder 113 obtains the third vote value for the 10th quantity of the 10th virtual speaker based on the first vote value for the 8th quantity virtual speaker and the second vote value for the 9th quantity virtual speaker.
[0201] If virtual speakers with the same number exist for the 8th quantity virtual speaker and the 9th quantity virtual speaker, the encoder 113 accumulates the first and second vote values of the same virtual speaker to obtain the third vote value for the 10th quantity of the 10th quantity virtual speaker. For example, it is assumed that the 8th quantity virtual speaker includes the 2nd virtual speaker, and the 9th quantity virtual speaker includes that 2nd virtual speaker. The third vote value of the 2nd virtual speaker is equal to the sum of the first vote value of the 1st virtual speaker and the second vote value of the 1st virtual speaker.
[0202] It can be understood that the 10th quantity is less than or equal to the 8th quantity, meaning that the virtual speaker of the 8th quantity includes the virtual speaker of the 10th quantity, and the 10th quantity is less than or equal to the 9th quantity, meaning that the virtual speaker of the 9th quantity includes the virtual speaker of the 10th quantity. Furthermore, the 10th quantity is an integer greater than or equal to 1.
[0203] S6109: The encoder 113 obtains the first virtual speaker and the first quantity vote value based on the first vote value of the eighth quantity virtual speaker, the second vote value of the ninth quantity virtual speaker, and the third vote value of the tenth quantity.
[0204] The virtual speaker of the first quantity includes the virtual speaker of the eighth quantity and the virtual speaker of the ninth quantity. The virtual speaker of the fifth quantity includes the virtual speaker of the first quantity. The first quantity is less than or equal to the fifth quantity.
[0205] For example, assuming that the fifth quantity of virtual speakers includes the first, second, third, fourth, and fifth virtual speakers, then the eighth quantity of virtual speakers includes the first and second virtual speakers, the ninth quantity of virtual speakers includes the first and third virtual speakers, the first quantity of virtual speakers includes the first, second, and third virtual speakers, and the first quantity is less than the fifth quantity.
[0206] As another example, if we assume that the fifth quantity of virtual speakers includes the first, second, third, fourth, and fifth virtual speakers, then the eighth quantity of virtual speakers includes the first, second, and third virtual speakers, the ninth quantity of virtual speakers includes the first, fourth, and fifth virtual speakers, the first quantity of virtual speakers includes the first, second, third, fourth, and fifth virtual speakers, and the first quantity is equal to the fifth quantity.
[0207] In some embodiments, if virtual speakers having the same number exist for the eighth quantity virtual speaker and the ninth quantity virtual speaker, the first quantity virtual speaker includes the tenth quantity virtual speaker.
[0208] In one case, the number of virtual speakers for the eighth quantity is exactly the same as the number of virtual speakers for the ninth quantity. The eighth quantity is equal to the ninth quantity, the tenth quantity is equal to the eighth quantity, and the tenth quantity is equal to the ninth quantity. Therefore, the number of virtual speakers for the first quantity is equal to the number of virtual speakers for the tenth quantity, and the vote value for the first quantity is equal to the third vote value for the tenth quantity.
[0209] In another case, the virtual speaker of the eighth quantity is not exactly the same as the virtual speaker of the ninth quantity. For example, the virtual speaker of the eighth quantity includes the virtual speaker of the ninth quantity, and the virtual speaker of the eighth quantity further includes a virtual speaker whose number is different from the number of the virtual speaker of the ninth quantity. The eighth quantity is greater than the ninth quantity, the tenth quantity is less than the eighth quantity, and the tenth quantity is equal to the ninth quantity. The voting value of the first quantity includes the third voting value of the tenth quantity and the first voting value of a virtual speaker whose number is different from the number of the virtual speaker of the ninth quantity.
[0210] As another example, the virtual speaker for the ninth quantity includes the virtual speaker for the eighth quantity, and the virtual speaker for the ninth quantity further includes a virtual speaker whose number is different from the number of the virtual speaker for the eighth quantity. The eighth quantity is less than the ninth quantity, the tenth quantity is equal to the eighth quantity, and the tenth quantity is less than the ninth quantity. The vote value for the first quantity includes the third vote value for the tenth quantity and the second vote value for the virtual speaker whose number is different from the number of the virtual speaker for the eighth quantity.
[0211] As another example, the virtual speaker of the eighth quantity includes the virtual speaker of the tenth quantity, the virtual speaker of the eighth quantity further includes a virtual speaker whose number is different from the number of the virtual speaker of the ninth quantity, the virtual speaker of the ninth quantity includes the virtual speaker of the tenth quantity, the virtual speaker of the ninth quantity further includes a virtual speaker whose number is different from the number of the virtual speaker of the eighth quantity. The tenth quantity is less than the eighth quantity, and the tenth quantity is less than the ninth quantity. The vote value of the first quantity includes the third vote value of the tenth quantity, the first vote value of the virtual speaker whose number is different from the number of the virtual speaker of the ninth quantity, and the second vote value of the virtual speaker whose number is different from the number of the virtual speaker of the eighth quantity.
[0212] In some other embodiments, if there are no virtual speakers with the same number for the eighth quantity virtual speaker and the ninth quantity virtual speaker, the tenth quantity is equal to 0, and the first quantity virtual speaker does not include the tenth quantity virtual speaker. After performing S6106 and S6107, the encoder 113 may perform S6110 directly.
[0213] S6110: The encoder 113 obtains the first quantity virtual speaker and the first quantity vote value based on the first vote value of the eighth quantity virtual speaker and the second vote value of the ninth quantity virtual speaker.
[0214] The virtual speaker of the eighth quantity is completely different from the virtual speaker of the ninth quantity. For example, the virtual speaker of the eighth quantity does not include the virtual speaker of the ninth quantity, and the virtual speaker of the ninth quantity does not include the virtual speaker of the eighth quantity. The virtual speaker of the first quantity includes the virtual speaker of the eighth quantity and the virtual speaker of the ninth quantity, and the vote value of the first quantity includes the first vote value of the virtual speaker of the eighth quantity and the second vote value of the virtual speaker of the ninth quantity.
[0215] In this way, for each coefficient of the current frame, the encoder selects a voting value that is larger from the voting values of the fifth quantity virtual speaker included in the candidate virtual speaker set, and uses the voting value that is larger to determine the virtual speaker of the first quantity and the voting value of the first quantity, thereby reducing the computational complexity of searching for a virtual speaker by the encoder, while ensuring the accuracy of the representative virtual speaker of the current frame, which is the representative virtual speaker selected by the encoder.
[0216] The following describes the method for calculating the vote value, referring to the formula. First, the encoder 113 calculates the vote value P of the l-th virtual speaker for the j-th representative coefficient in the i-th round, based on the correlation between the j-th representative coefficient of the HOA signal and the coefficient of the l-th virtual speaker. jil Step 1 is performed to determine the j-th representative coefficient, which may be any coefficient in the representative coefficient of the third quantity, where l = 1, 2, ..., and Q, where the range of values for l is from 1 to Q, where Q represents the quantity of virtual speakers in the candidate virtual speaker set, where j = 1, 2, ..., and L, where L represents the quantity of the representative coefficient, and where i = 1, 2, ..., and I, where I represents the voting round quantity. The voting value P of the l-th virtual speaker. jil This satisfies equation (6). P jil =log(E jil ) or P jil =E jil Ejil =B ji (θ,φ)·B l (θ,φ) Equation (6) However, θ represents the horizontal angle, φ represents the pitch angle, and B ji (θ, φ) represents the j-th representative coefficient of the HOA signal, and B l (θ, φ) represents the coefficient of the l-th virtual speaker.
[0217] Next, the encoder 113 calculates the voting value P of the Q virtual speakers. jil Based on this, we perform step 2 to obtain a virtual speaker corresponding to the j-th representative coefficient in the i-th round.
[0218] For example, the criterion for selecting a virtual speaker corresponding to the j-th representative coefficient in the i-th round is to select the virtual speaker with the largest absolute value of the votes from the Q virtual speakers for the j-th representative coefficient in the i-th round, where the number of virtual speakers corresponding to the j-th representative coefficient in the i-th round is g ji It is written as l=g ji in the case of,
number
[0219] If i is less than the voting round quantity I, i.e., if the voting round quantity I has been completed cyclically, the encoder 113 performs step 3 to use the remaining virtual speakers in the candidate virtual speaker set as the encoded HOA signal required to calculate the voting value of the virtual speakers for the j-th representative coefficient in the next round. The coefficients of the remaining virtual speakers in the candidate virtual speaker set satisfy equation (7). B j (θ,φ)=B j (θ,φ)-w·B gj,i(θ,φ)·E jig Formula (7) However, E jig represents the voting value of the l-th virtual speaker corresponding to the j-th representative coefficient in the i-th round, and B on the right side of the equation gj,i (θ,φ) represents the coefficient of the HOA signal to be encoded for the jth representative coefficient in the i-th round, and B on the left side of the equation j (θ, φ) represents the coefficient of the HOA signal to be encoded for the j-th representative coefficient in the (i+1)-th round, w is a weight, and a predetermined value can satisfy 0 ≤ w ≤ 1, and furthermore, the weight can satisfy equation (8). w=norm(B gj,i (θ, φ) Equation (8) However, "norm" is an operation that solves a two-norm problem.
[0220] The encoder 113 performs step 4, that is, the encoder 113 calculates the voting value of the virtual speaker corresponding to the j-th representative coefficient in each round.
[0221]
number
[0222] Repeat steps 1 through 3 until the result is calculated.
[0223] Encoder 113 displays the virtual speaker's voting values corresponding to all representative coefficients in each round.
[0224]
number
[0225] Repeat steps 1 through 4 until the result is calculated.
[0226] Finally, encoder 113 outputs the virtual speaker number g corresponding to each representative frequency in each round. j,i And the voting values corresponding to the virtual speaker
[0227]
number
[0228] Based on this, the final vote value for each virtual speaker for the current frame is calculated. For example, encoder 113 accumulates the vote values of virtual speakers with the same number to obtain the final vote value for the virtual speaker for the current frame. The final vote value VOTEg for the virtual speaker for the current frame satisfies equation (9). VOTE g =ΣP jig or VOTE g =VOTE g +P jig Formula (9)
[0229] To increase orientation continuity between consecutive frames and overcome the problem of significantly varying results in selecting a virtual speaker for consecutive frames, encoder 113 adjusts the initial vote value of a virtual speaker in the candidate virtual speaker set for the current frame based on the final vote value of a representative virtual speaker for a previous frame to obtain the final vote value of the virtual speaker for the current frame. Figure 8 is a schematic flowchart of another method for selecting a virtual speaker according to one embodiment of the present application. The method procedure in Figure 8 illustrates the specific calculation process included in S620 in Figure 6.
[0230] S6201: The encoder 113 obtains the final vote value of the 7th quantity in the current frame, which corresponds to the virtual speaker of the 7th quantity, and the current frame, based on the initial vote value of the 1st quantity in the current frame and the final vote value of the 6th quantity in the previous frame.
[0231] The encoder 113 may determine a first quantity of virtual speaker and a first quantity of voting value based on the current frame of the three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, using the method described in S610, and then use the first quantity of voting value as the initial voting value for the current frame corresponding to the first quantity of virtual speaker.
[0232] Each virtual speaker has a one-to-one correspondence with the initial vote value of the current frame; that is, one virtual speaker corresponds to one initial vote value of the current frame. For example, the virtual speaker for the first quantity includes the first virtual speaker, the initial vote value of the first quantity in the current frame includes the initial vote value of the first virtual speaker for the current frame, and the first virtual speaker corresponds to the initial vote value of the first virtual speaker for the current frame. The initial vote value of the first virtual speaker for the current frame represents the priority of using the first virtual speaker when the current frame is encoded.
[0233] The virtual speaker for the sixth quantity, which is included in the representative virtual speaker set for the previous frame, corresponds one-to-one with the final vote value of the sixth quantity for the previous frame. The virtual speaker for the sixth quantity may be the representative virtual speaker for the previous frame, which was used by the encoder 113 to encode the previous frame of the three-dimensional audio signal.
[0234] Specifically, encoder 113 updates the initial vote value of the first quantity in the current frame based on the final vote value of the sixth quantity in the previous frame. Specifically, encoder 113 calculates the sum of the final vote value of the previous frame and the initial vote value of the current frame corresponding to the virtual speakers with the same number in the first quantity virtual speaker and the sixth quantity virtual speaker, to obtain the final vote value of the seventh quantity in the current frame for the seventh quantity virtual speaker, which corresponds to the current frame.
[0235] S6202: The encoder 113 selects a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the seventh quantity, based on the final vote value of the seventh quantity for the current frame.
[0236] The encoder 113 selects a representative virtual speaker for the second quantity for the current frame from the virtual speakers for the seventh quantity based on the final vote value of the seventh quantity for the current frame, and the final vote value for the current frame corresponding to the representative virtual speaker for the second quantity for the current frame is greater than a preset threshold.
[0237] The encoder 113 may, as an alternative, select a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the seventh quantity based on the final vote value of the seventh quantity for the current frame. For example, the final vote value of the second quantity for the current frame is determined from the final vote values of the seventh quantity for the current frame in descending order, and a virtual speaker within the virtual speakers of the seventh quantity that is associated with the final vote value of the second quantity for the current frame is used as the representative virtual speaker for the second quantity for the current frame.
[0238] Optionally, if, in the seventh quantity of virtual speakers, virtual speakers with different numbers have the same voting value, and the voting value of the virtual speakers with different numbers is greater than a preset threshold, the encoder 113 may use the virtual speaker with different numbers as the representative virtual speaker for the current frame.
[0239] It should be noted that the second quantity is less than the seventh quantity. The virtual speaker of the seventh quantity includes a representative virtual speaker of the second quantity for the current frame. The second quantity may be predetermined, or it may be determined based on the number of sound sources in the sound field of the current frame.
[0240] Furthermore, if encoder 113 decides to reuse a representative virtual speaker for a previous frame to encode the next frame before encoding the next frame of the current frame, encoder 113 may encode the next frame of the current frame by using a representative virtual speaker for a second quantity for the previous frame and using a representative virtual speaker for a second quantity for the previous frame.
[0241] In the process of searching for virtual speakers, the positions of actual sound sources unnecessarily overlap with the positions of virtual speakers, so virtual speakers may not be able to form a one-to-one correspondence with actual sound sources. Furthermore, in complex real-world scenarios, a set with a limited number of virtual speakers may not be able to represent all sound sources in the sound field. In this case, the virtual speakers found in different frames may change frequently, and this change clearly affects the listener's auditory perception, resulting in noticeable discontinuities and noise in the three-dimensional audio signal obtained after decoding and reconstruction. According to the method for selecting a virtual speaker provided in this embodiment of the present application, a representative virtual speaker for a previous frame is inherited, specifically, for virtual speakers having the same number, the initial vote value for the current frame is adjusted by using the final vote value for the previous frame. As a result, the encoder is more likely to select a representative virtual speaker for a previous frame, thereby reducing frequent changes in virtual speakers across different frames, increasing signal orientation continuity between frames, improving the audio stability of the reconstructed three-dimensional audio signal, and ensuring the sound quality of the reconstructed three-dimensional audio signal. Furthermore, the parameters are adjusted to ensure that the final vote value for a previous frame is not inherited over a long period of time, preventing the algorithm from being unable to adapt to scenarios where the sound field changes, such as a sound source movement scenario.
[0242] Furthermore, this embodiment of the present application further provides a method for selecting a virtual speaker. The encoder may first determine whether a representative virtual speaker set for a previous frame can be reused to encode the current frame. If the encoder can reuse a representative virtual speaker set for a previous frame to encode the current frame, the encoder does not perform a virtual speaker search process, which effectively reduces the computational complexity of searching for a virtual speaker by the encoder, thereby reducing the computational complexity of performing compressed coding to a three-dimensional audio signal and reducing the computational load on the encoder. If the encoder cannot reuse a representative virtual speaker set for a previous frame to encode the current frame, the encoder selects a representative coefficient, uses the representative coefficient for the current frame to vote for each virtual speaker in the candidate virtual speaker set, and selects a representative virtual speaker for the current frame based on the vote value, thereby reducing the computational complexity of performing compressed coding to a three-dimensional audio signal and reducing the computational load on the encoder. Figure 9 is a schematic flowchart of a method for selecting a virtual speaker according to one embodiment of the present application. Before encoder 113 obtains the coefficient of the fourth quantity of the current frame of the three-dimensional audio signal, and the frequency domain feature value of the coefficient of the fourth quantity, i.e., before S610, the method includes the following steps, as shown in Figure 9.
[0243] S640: The encoder 113 obtains a first correlation between the current frame of the three-dimensional audio signal and a representative virtual speaker set for the previous frame.
[0244] The representative virtual speaker set for a previous frame includes a sixth quantity of virtual speakers, the virtual speakers included in the sixth quantity of virtual speakers are representative virtual speakers for the previous frame used to encode the previous frame of the three-dimensional audio signal. The first correlation represents the priority of reusing the representative virtual speaker set for a previous frame when the current frame is encoded. Priority may be replaced with tendency, specifically, the first correlation is used to determine whether the representative virtual speaker set for a previous frame should be reused when the current frame is encoded. A larger first correlation for the representative virtual speaker set for a previous frame indicates a higher tendency for the representative virtual speaker set for a previous frame, and it can be understood that the encoder 113 is more likely to select the representative virtual speaker for a previous frame to encode the current frame.
[0245] S650: The encoder 113 determines whether the first correlation satisfies the reuse conditions.
[0246] If the first correlation does not satisfy the reuse condition, it indicates that encoder 113 is more likely to search for a virtual speaker and encode the current frame based on a representative virtual speaker for the current frame, and perform S610, specifically, encoder 113 obtains the coefficient of the fourth quantity of the current frame of the three-dimensional audio signal, and the frequency domain feature value of the coefficient of the fourth quantity.
[0247] Optionally, after selecting a third number of representative coefficients from a fourth number of coefficients based on frequency domain feature values of the fourth number of coefficients, the encoder 113 may use the largest representative coefficient among the third number of representative coefficients as the coefficient of the current frame that is used for obtaining a first correlation. In this case, the encoder 113 obtains the first correlation between the largest representative coefficient among the third number of representative coefficients of the current frame and the representative virtual speaker set for a previous frame. If the first correlation does not satisfy a reuse condition, S620 is performed. Specifically, the encoder 113 selects a second number of representative virtual speakers for the current frame from a first number of virtual speakers based on a first number of voting values.
[0248] If the first correlation satisfies the reuse condition, this indicates that the encoder 113 is more likely to select the representative virtual speaker for the previous frame to encode the current frame, and the encoder 113 performs S660 and S670.
[0249] S660: The encoder 113 generates a virtual speaker signal based on the representative virtual speaker set for the previous frame and the current frame.
[0250] S670: The encoder 113 encodes the virtual speaker signal to obtain a bitstream.
[0251] According to the method for selecting a virtual speaker provided in this embodiment of the present application, whether to search for a virtual speaker is determined by using the correlation between the representative coefficient of the current frame and the representative virtual speaker for the previous frame, which effectively reduces the complexity on the encoder side while ensuring the accuracy of selecting the representative virtual speaker for the current frame.
[0252] It can be understood that, in order to implement the functions of the foregoing embodiments, the encoder includes corresponding hardware structures and / or software modules for performing those functions. Those skilled in the art should readily recognize that the units and method steps in the examples described with reference to the embodiments disclosed in the present application can be implemented in the present application in the form of hardware, or a combination of hardware and computer software. Whether the function is performed by hardware or hardware driven by computer software depends on the specific application scenario and design constraints of the technical solution.
[0253] With reference to FIG. 1 to FIG. 9, the foregoing content describes in detail the three-dimensional audio signal coding method provided in this embodiment. With reference to FIG. 10 and FIG. 11, the following describes the three-dimensional audio signal encoding apparatus and encoder provided in embodiments.
[0254] FIG. 10 is a schematic diagram of a possible structure of a three-dimensional audio signal encoding apparatus according to an embodiment. The three-dimensional audio signal encoding apparatus may be configured to implement the function of encoding a three-dimensional audio signal in the foregoing method embodiments, and therefore can also implement the beneficial effects of the foregoing method embodiments. In this embodiment, the three-dimensional audio signal encoding apparatus may be the encoder 113 shown in FIG. 1, or the encoder 300 shown in FIG. 3, or may be a module (such as a chip) applied to a terminal device or a server.
[0255] As shown in FIG. 10, the three-dimensional audio signal encoding apparatus 1000 includes a communication module 1010, a coefficient selection module 1020, a virtual speaker selection module 1030, an encoding module 1040, and a storage module 1050. The three-dimensional audio signal encoding apparatus 1000 is configured to implement the function of the encoder 113 in the method embodiments shown in FIG. 6 to FIG. 9.
[0256] The communication module 1010 is configured to acquire the current frame of a three-dimensional audio signal. Optionally, the communication module 1010 may, as an alternative, receive the current frame of a three-dimensional audio signal acquired by another device, or acquire the current frame of a three-dimensional audio signal from the storage module 1050. The current frame of the three-dimensional audio signal is an HOA signal, where the frequency-domain feature values of the coefficients are determined based on a two-dimensional vector, which contains the HOA coefficients of the HOA signal.
[0257] The virtual speaker selection module 1030 is configured to determine a first quantity virtual speaker and a first quantity voting value based on the current frame of a three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, wherein a virtual speaker corresponds one-to-one with a voting value, the first quantity virtual speaker includes the first virtual speaker, the first quantity voting value includes the voting value of the first virtual speaker, the first virtual speaker corresponds to the voting value of the first virtual speaker, the voting value of the first virtual speaker represents the priority of using the first virtual speaker when the current frame is encoded, the candidate virtual speaker set includes a fifth quantity virtual speaker, the fifth quantity virtual speaker includes the first quantity virtual speaker, the voting round quantity is an integer greater than or equal to 1, and the voting round quantity is less than or equal to the fifth quantity.
[0258] The virtual speaker selection module 1030 is further configured to select a representative virtual speaker for a second quantity for the current frame from among the virtual speakers of a first quantity, based on the voting value of a first quantity, wherein the second quantity is less than the first quantity.
[0259] The voting round quantity is determined based on at least one of the following: the number of directional sources in the current frame of the three-dimensional audio signal, the coding rate, and the coding complexity. The second quantity is either predetermined or determined based on the current frame.
[0260] If the three-dimensional audio signal encoding device 1000 is configured to implement the functions of the encoder 113 in the method embodiment shown in Figures 6 to 9, the virtual speaker selection module 1030 is configured to implement the relevant functions in S610 and S620.
[0261] For example, when selecting a representative virtual speaker for a second quantity for the current frame from the virtual speakers of a first quantity based on the voting value of a first quantity, the virtual speaker selection module 1030 is specifically configured to select a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity and a preset threshold.
[0262] As another example, when selecting a representative virtual speaker for a second quantity for the current frame from the virtual speakers for a first quantity based on the voting value of a first quantity, the virtual speaker selection module 1030 is specifically configured to determine the voting value of the second quantity from the voting values of the first quantity in descending order of the voting values of the first quantity, and to use as the representative virtual speaker for the second quantity for the current frame the virtual speaker for the second quantity that is a virtual speaker for the second quantity in the virtual speakers for the first quantity and is associated with the voting value of the second quantity.
[0263] If the three-dimensional audio signal encoding device 1000 is optionally configured to implement the functions of encoder 113 in the method embodiment shown in Figure 9, the virtual speaker selection module 1030 is configured to implement the relevant functions in S640 and S670. Specifically, the virtual speaker selection module 1030 is configured to obtain a first correlation between a representative virtual speaker set for the current frame and a previous frame, and if the first correlation does not satisfy the reuse condition, to obtain a coefficient of a fourth quantity for the current frame of the three-dimensional audio signal, and a frequency domain feature value of the coefficient of the fourth quantity. The representative virtual speaker set for a previous frame includes a sixth quantity virtual speaker, the virtual speaker included in the sixth quantity virtual speaker is a representative virtual speaker for a previous frame used to encode the previous frame of the three-dimensional audio signal, and the first correlation represents the priority for reusing the sixth quantity virtual speaker when the current frame is encoded.
[0264] When the three-dimensional audio signal encoding device 1000 is configured to implement the functions of the encoder 113 in the method embodiment shown in Figure 8, the virtual speaker selection module 1030 is configured to implement the relevant functions in S620. Specifically, when selecting a representative virtual speaker of a second quantity for the current frame from a first quantity of virtual speakers based on the voting value of a first quantity, the virtual speaker selection module 1030 is particularly configured to obtain the final voting value of a seventh quantity for the current frame corresponding to a seventh quantity virtual speaker, and the current frame, based on the voting value of the first quantity and the final voting value of the sixth quantity for the previous frame of a sixth quantity virtual speaker included in the representative virtual speaker set for the previous frame, which corresponds to the previous frame of the three-dimensional audio signal, and select a representative virtual speaker of a second quantity for the current frame from a seventh quantity virtual speaker based on the final voting value of the seventh quantity for the current frame, provided that the second quantity is less than the seventh quantity. The seventh quantity virtual speaker includes the first quantity virtual speaker, the seventh quantity virtual speaker includes the sixth quantity virtual speaker, and the virtual speaker included in the sixth quantity virtual speaker is a representative virtual speaker for the previous frame, used to encode the previous frame of the three-dimensional audio signal.
[0265] When the three-dimensional audio signal encoding device 1000 is configured to implement the functions of the encoder 113 in the method embodiment shown in Figures 7A and 7B, the coefficient selection module 1020 is configured to implement the relevant functions in S6101. Specifically, when obtaining a representative coefficient of a third quantity in the current frame, the coefficient selection module 1020 is configured to obtain the coefficient of a fourth quantity in the current frame and the frequency domain feature value of the coefficient of the fourth quantity, and to select a representative coefficient of the third quantity from the coefficient of the fourth quantity based on the frequency domain feature value of the coefficient of the fourth quantity, provided that the third quantity is less than the fourth quantity.
[0266] The encoding module 1140 is configured to encode the current frame and obtain a bitstream based on a second quantity of representative virtual speakers for the current frame.
[0267] If the three-dimensional audio signal encoding device 1000 is configured to implement the functions of the encoder 113 in the method embodiment shown in Figures 6 to 9, the encoding module 1140 is configured to implement the relevant functions in S630. For example, the encoding module 1140 is specifically configured to generate a virtual speaker signal based on a second quantity of representative virtual speakers for the current frame and the current frame, encode the virtual speaker signal, and obtain a bitstream.
[0268] The memory module 1050 is configured to store coefficients related to the three-dimensional audio signal, candidate virtual speaker sets, a representative virtual speaker set for a previous frame, selected coefficients and virtual speakers, etc. As a result, the encoding module 1040 encodes the current frame, obtains a bitstream, and transmits the bitstream to the decoder.
[0269] It should be understood that the three-dimensional audio signal coding device 1000 in this embodiment of the present application may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. If the three-dimensional audio signal coding method shown in Figures 6 to 9 is implemented using software, the three-dimensional audio signal coding device 1000 and its modules may be software modules instead.
[0270] For a more detailed description of the communication module 1010, coefficient selection module 1020, virtual speaker selection module 1030, encoding module 1040, and storage module 1050, please refer directly to the relevant descriptions in the method embodiments shown in Figures 6 to 9. Further details are not described again here.
[0271] Figure 11 is a schematic diagram of the structure of an encoder 1100 according to one embodiment. As shown in Figure 11, the encoder 1100 includes a processor 1110, a bus 1120, a memory 1130, and a communication interface 1140.
[0272] In the present embodiment, it should be understood that the processor 1110 may be a central processing unit (CPU), or the processor 1110 may be another general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or another programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or the like. A general-purpose processor may be a microprocessor, or any conventional processor, or the like.
[0273] Alternatively, the processor may be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits configured to control program execution of the solution in the present application.
[0274] The communication interface 1140 is configured to implement communication between the encoder 1100 and an external device or component. In the present embodiment, the communication interface 1140 is configured to receive a three-dimensional audio signal.
[0275] The bus 1120 may include channels configured to transmit information between the aforementioned components (e.g., the processor 1110 and the memory 1130). In addition to a data bus, the bus 1120 may further include a power bus, a control bus, a status signal bus, and the like. However, for the sake of clear description, various types of buses are depicted as the bus 1120 in the drawings.
[0276] For example, the encoder 1100 may include multiple processors. The processors may be multi-core (multi-CPU) processors. A processor as used herein may be one or more devices, circuits, and / or computing units configured to process data (e.g., computer program instructions). The processor 1110 may retrieve coefficients associated with a three-dimensional audio signal, a candidate virtual speaker set, a representative virtual speaker set for a previous frame, and selected coefficients and virtual speakers stored in the memory 1130.
[0277] It should be noted that in Figure 11, only an example is used in which the encoder 1100 includes one processor 1110 and one memory 1130. In this specification, the processor 1110 and the memory 1130 each represent a type of component or device. In a particular embodiment, the number of each type of component or device may be determined based on service requirements.
[0278] The memory 1130 may correspond to a storage medium, such as a mechanical hard disk or a magnetic disk such as a solid-state disk, configured to store information such as coefficients related to the three-dimensional audio signal, a candidate virtual speaker set, a representative virtual speaker set for a previous frame, and a selected coefficient and virtual speaker, as in the method embodiment described above.
[0279] The encoder 1100 may be a general-purpose device or a dedicated device. For example, the encoder 1100 may be an x86-based server or an ARM-based server, or another dedicated server such as a policy control and charging (PCC) server. The type of encoder 1100 is not limited to this embodiment of the present application.
[0280] It should be understood that the encoder 1100 according to this embodiment may correspond to the three-dimensional audio signal coding device 1100 in the embodiment and to the corresponding body configured to perform any of the methods shown in Figures 6 to 9. Furthermore, the aforementioned and other operations and / or functions of the modules within the three-dimensional audio signal coding device 1100 are used, respectively, to implement the corresponding steps of the methods shown in Figures 6 to 9. For brevity, further details are not described here.
[0281] The method steps in the embodiments may be implemented by hardware or by a processor that executes software instructions. The software instructions may include corresponding software modules. The software modules may be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, as a result the processor can read information from and write information to the storage medium. Of course, the storage medium may be a component of the processor. The processor and storage medium may reside in an ASIC. Furthermore, the ASIC may reside in a network device or terminal device. Of course, the processor and storage medium may exist as discrete components in a network device or terminal device.
[0282] All or part of the embodiments described above may be implemented using software, hardware, firmware, or any combination thereof. When software is used for implementation, the embodiments may be implemented all or partly in the form of a computer program product. A computer program product includes one or more computer programs or instructions. When a computer program or instruction is loaded into a computer and executed, all or part of the procedures or functions according to the embodiments of this application are performed. The computer may be a general-purpose computer, a dedicated computer, a computer network, a network device, a user device, or another programmable device. The computer programs or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, a computer program or instruction may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired or wireless means. The computer-readable storage medium may be any available medium accessible by a computer or data storage device, such as a server or data center, which may be an integrated medium of one or more available media. The usable media may be magnetic media, such as floppy disks, hard disks, or magnetic tapes; optical media, such as digital video discs (DVDs); or semiconductor media, such as solid-state drives (SSDs).
[0283] The foregoing description is merely a specific implementation of this application and is not intended to limit the scope of protection of this application. Any equivalent modifications or substitutions that are readily conceivable by a person skilled in the art within the technical scope disclosed in this application should fall within the scope of protection of this application. Accordingly, the scope of protection of this application should be subject to the scope of protection of the claims.
Claims
1. A three-dimensional audio signal encoding method, A step of determining a first quantity of virtual speakers and a first quantity of voting values based on the current frame of a three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, The voting round quantity is determined based on at least one of the number of directional sound sources in the current frame of the three-dimensional audio signal, the coding rate at which the current frame is encoded, and the coding complexity at which the current frame is encoded. The virtual speaker corresponds one-to-one with the vote value, the first quantity of virtual speakers includes the first virtual speaker, the vote value of the first virtual speaker represents the priority for selecting the first virtual speaker when the current frame is encoded, the candidate virtual speaker set includes a fifth quantity of virtual speakers, the fifth quantity of virtual speakers includes the first quantity of virtual speakers, the first quantity is less than or equal to the fifth quantity, the voting round quantity is an integer of 1 or more, and the voting round quantity is less than or equal to the fifth quantity, step, A step of selecting a representative virtual speaker for a second quantity for the current frame from among the virtual speakers of the first quantity, based on the voting value of the first quantity, wherein the second quantity is less than the first quantity. The steps of obtaining the bitstream by encoding the current frame based on the second quantity of representative virtual speakers for the current frame to obtain a compressed version of the current frame, and incorporating the compressed version of the current frame into a bitstream which is a compressed version of the three-dimensional audio signal. A three-dimensional audio signal encoding method, including the above.
2. The method according to claim 1, wherein the second quantity is predetermined or determined based on the current frame.
3. The step of selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity is: Steps to select a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity, based on the voting value of the first quantity and a preset threshold. The method according to claim 1, including the method described in claim 1.
4. The step of selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity is: A step of determining the voting value of a second quantity from the voting value of the first quantity, wherein the virtual speaker of the second quantity within the virtual speaker of the first quantity, which corresponds to the voting value of the second quantity, is a representative virtual speaker of the second quantity for the current frame. The method according to claim 1, including the method described in claim 1.
5. If the first quantity is equal to the fifth quantity, the step of determining the virtual speaker of the first quantity and the voting value of the first quantity based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity is: A step of obtaining a representative coefficient of a third quantity of the current frame, wherein the representative coefficient of the third quantity includes a first representative coefficient and a second representative coefficient. A step of obtaining a first vote value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first vote value for the fifth quantity includes the first vote value of the first virtual speaker. A step of obtaining a second voting value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value for the first virtual speaker. A step of obtaining the respective vote values for the virtual speakers of the fifth quantity based on the first vote value of the fifth quantity and the second vote value of the fifth quantity, wherein the vote value of the first virtual speaker is obtained based on the first vote value of the first virtual speaker and the second vote value of the first virtual speaker. The method according to claim 1, including the method described in claim 1.
6. If the first quantity is less than or equal to the fifth quantity, the step of determining the virtual speaker of the first quantity and the voting value of the first quantity based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity is: A step of obtaining a representative coefficient of a third quantity of the current frame, wherein the representative coefficient of the third quantity includes a first representative coefficient and a second representative coefficient. A step of obtaining a first vote value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first vote value for the fifth quantity includes the first vote value of the first virtual speaker. A step of obtaining a second voting value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value for the first virtual speaker. A step of selecting an eighth quantity of virtual speakers from the fifth quantity of virtual speakers based on a first voting value of the fifth quantity, wherein the eighth quantity is less than the fifth quantity. A step of selecting a virtual speaker of a ninth quantity from virtual speakers of the fifth quantity based on a second voting value of the fifth quantity, wherein the ninth quantity is less than the fifth quantity. A step of obtaining a third vote value for a 10th quantity of a 10th quantity virtual speaker based on a first vote value for an 8th quantity virtual speaker and a second vote value for a 9th quantity virtual speaker, wherein the 8th quantity virtual speaker includes the 10th quantity virtual speaker, the 9th quantity virtual speaker includes the 10th quantity virtual speaker, the 10th quantity virtual speaker includes a second virtual speaker, the third vote value for the second virtual speaker is obtained based on the first vote value for the second virtual speaker and the second vote value for the second virtual speaker, the 10th quantity is less than or equal to the 8th quantity, the 10th quantity is less than or equal to the 9th quantity, and the 10th quantity is an integer of 1 or more. A step of obtaining a first quantity virtual speaker and a first quantity vote value based on the first vote value of the eighth quantity virtual speaker, the second vote value of the ninth quantity virtual speaker, and the third vote value of the tenth quantity, wherein the first quantity virtual speaker includes the eighth quantity virtual speaker and the ninth quantity virtual speaker. The method according to claim 1, including the method described in claim 1.
7. If the first quantity is less than or equal to the fifth quantity, the step of determining the virtual speaker of the first quantity and the voting value of the first quantity based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity is: A step of obtaining a representative coefficient of a third quantity of the current frame, wherein the representative coefficient of the third quantity includes a first representative coefficient and a second representative coefficient. A step of obtaining a first vote value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first vote value for the fifth quantity includes the first vote value of the first virtual speaker. A step of obtaining a second voting value for a fifth quantity of a fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value for the first virtual speaker. A step of selecting an eighth quantity of virtual speakers from the fifth quantity of virtual speakers based on a first voting value of the fifth quantity, wherein the eighth quantity is less than the fifth quantity. A step of selecting a virtual speaker of a ninth quantity from virtual speakers of the fifth quantity based on a second voting value of the fifth quantity, wherein the ninth quantity is less than the fifth quantity, and there is no common part between the virtual speaker of the eighth quantity and the virtual speaker of the ninth quantity. A step of obtaining the first quantity virtual speaker and the first quantity vote value based on the first vote value of the eighth quantity virtual speaker and the second vote value of the ninth quantity virtual speaker, wherein the first quantity virtual speaker includes the eighth quantity virtual speaker and the ninth quantity virtual speaker. The method according to claim 1, including the method described in claim 1.
8. The step of obtaining a first vote value for a fifth quantity of virtual speakers of the fifth quantity, which is obtained by performing a voting round of the voting round quantity using the first representative coefficient, A step of determining a first voting value for the fifth quantity based on the coefficient of the virtual speaker of the fifth quantity and the first representative coefficient. The method according to claim 5, including the method described in claim 5.
9. The step of obtaining a representative coefficient of the third quantity of the current frame is: The steps include obtaining the coefficient of the fourth quantity of the current frame and the frequency domain feature value of the coefficient of the fourth quantity, A step of selecting a representative coefficient of the third quantity from the coefficients of the fourth quantity based on the frequency domain characteristic value of the coefficient of the fourth quantity, wherein the third quantity is less than the fourth quantity. The method according to claim 5, including the method described in claim 5.
10. Before the step of selecting a representative coefficient of the third quantity from the coefficients of the fourth quantity based on the frequency domain characteristic value of the coefficient of the fourth quantity, the method: A step of obtaining a first correlation between the current frame and a representative virtual speaker set for a previous frame, wherein the representative virtual speaker set for the previous frame includes a sixth quantity of virtual speakers, the virtual speakers included in the sixth quantity of virtual speakers are representative virtual speakers for the previous frame used to encode the previous frame of the three-dimensional audio signal, and the first correlation is used to determine whether to reuse the representative virtual speaker set for the previous frame when the current frame is encoded. If the first correlation does not satisfy the reuse conditions, the steps include obtaining the coefficient of the fourth quantity of the current frame of the three-dimensional audio signal, and the frequency domain feature value of the coefficient of the fourth quantity. The method according to claim 9, further comprising:
11. The step of selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity is: Steps to obtain a virtual speaker for a seventh quantity and the final vote value for the seventh quantity of the current frame, corresponding to the current frame, based on the vote value for the first quantity and the final vote value for the sixth quantity of a previous frame, wherein the virtual speaker for the seventh quantity includes the virtual speaker for the first quantity, the virtual speaker for the seventh quantity includes the virtual speaker for the sixth quantity, the virtual speaker for the sixth quantity included in the representative virtual speaker set for the previous frame corresponds one-to-one with the final vote value for the sixth quantity of the previous frame, and the virtual speaker for the sixth quantity is a virtual speaker used when the previous frame of the three-dimensional audio signal is encoded. A step of selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the seventh quantity, based on the final voting value of the seventh quantity for the current frame, wherein the second quantity is less than the seventh quantity. The method according to claim 1, including the method described in claim 1.
12. The method according to claim 1, wherein the current frame of the three-dimensional audio signal is a higher-order ambisonics HOA signal, and the frequency domain feature values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
13. A three-dimensional audio signal encoding device, A virtual speaker selection module configured to determine a first quantity of virtual speakers and a first quantity of voting values based on the current frame of a three-dimensional audio signal, a candidate virtual speaker set, and a voting round quantity, wherein the voting round quantity is determined based on at least one of the number of directional sound sources in the current frame of the three-dimensional audio signal, the coding rate at which the current frame is encoded, and the coding complexity at which the current frame is encoded; the virtual speakers correspond one-to-one with the voting values; the first quantity of virtual speakers includes a first virtual speaker, the voting value of the first virtual speaker represents the priority for selecting the first virtual speaker when the current frame is encoded; the candidate virtual speaker set includes a fifth quantity of virtual speakers, the fifth quantity of virtual speakers includes a first quantity of virtual speakers, the first quantity is less than or equal to the fifth quantity; the voting round quantity is an integer of 1 or more, and the voting round quantity is less than or equal to the fifth quantity. The virtual speaker selection module is further configured to select a representative virtual speaker for the current frame from the first quantity of virtual speakers based on the voting value of the first quantity, wherein the second quantity is less than the first quantity. An encoding module configured to encode the current frame based on a second number of representative virtual speakers for the current frame to obtain a compressed version of the current frame, and to obtain the bitstream by incorporating the compressed version of the current frame into a bitstream which is a compressed version of the three-dimensional audio signal. A three-dimensional audio signal encoding device equipped with the following features.
14. The apparatus according to claim 13, wherein the second quantity is set in advance, or the second quantity is determined based on the current frame.
15. When selecting a representative virtual speaker for the second quantity from the first quantity virtual speakers for the current frame based on the voting value of the first quantity, the virtual speaker selection module: Based on the voting value of the first quantity and a preset threshold, a representative virtual speaker for the second quantity is selected from the virtual speakers of the first quantity for the current frame. The apparatus according to claim 13, configured as follows.
16. When selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity, the virtual speaker selection module: Based on the voting value of the first quantity, the voting value of the second quantity is determined from the voting value of the first quantity, and the virtual speaker of the second quantity within the virtual speaker of the first quantity, which corresponds to the voting value of the second quantity, is used as the representative virtual speaker of the second quantity for the current frame. The apparatus according to claim 13, configured as follows.
17. When the first quantity is equal to the fifth quantity, and the virtual speaker of the first quantity and the voting value of the first quantity are determined based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity, the virtual speaker selection module, The method involves obtaining a representative coefficient of the third quantity in the current frame, wherein the representative coefficient of the third quantity includes the first representative coefficient and the second representative coefficient. The first voting value of the fifth quantity of the fifth virtual speaker is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first voting value of the fifth quantity includes the first voting value of the first virtual speaker. The method involves obtaining a second voting value for the fifth quantity of the fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value of the first virtual speaker. The voting values of each virtual speaker of the fifth quantity are obtained based on the first voting value of the fifth quantity and the second voting value of the fifth quantity, wherein the voting value of the first virtual speaker is obtained based on the first voting value of the first virtual speaker and the second voting value of the first virtual speaker. The apparatus according to claim 13, configured to perform the following:
18. When the first quantity is less than or equal to the fifth quantity, and the virtual speaker of the first quantity and the voting value of the first quantity are determined based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity, the virtual speaker selection module, The method involves obtaining a representative coefficient of the third quantity in the current frame, wherein the representative coefficient of the third quantity includes the first representative coefficient and the second representative coefficient. The first voting value of the fifth quantity of the fifth virtual speaker is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first voting value of the fifth quantity includes the first voting value of the first virtual speaker. The method involves obtaining a second voting value for the fifth quantity of the fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value of the first virtual speaker. Based on the first vote value of the fifth quantity, a virtual speaker of the eighth quantity is selected from the virtual speakers of the fifth quantity, wherein the eighth quantity is less than the fifth quantity. Based on the second voting value of the fifth quantity, a virtual speaker of the ninth quantity is selected from the virtual speakers of the fifth quantity, wherein the ninth quantity is less than the fifth quantity. The method involves obtaining a third vote value for the 10th quantity of a virtual speaker of the 10th quantity based on the first vote value of the virtual speaker of the 8th quantity and the second vote value of the virtual speaker of the 9th quantity, wherein the virtual speaker of the 8th quantity includes the virtual speaker of the 10th quantity, the virtual speaker of the 9th quantity includes the virtual speaker of the 10th quantity, the virtual speaker of the 10th quantity includes the second virtual speaker, the third vote value of the second virtual speaker is obtained based on the first vote value of the second virtual speaker and the second vote value of the second virtual speaker, the 10th quantity is less than or equal to the 8th quantity, the 10th quantity is less than or equal to the 9th quantity, and the 10th quantity is an integer of 1 or more. Obtaining a virtual speaker for the first quantity and a voting value for the first quantity based on the first voting value for the eighth quantity, the second voting value for the ninth quantity, and the third voting value for the tenth quantity, wherein the virtual speaker for the first quantity includes the virtual speaker for the eighth quantity and the virtual speaker for the ninth quantity. The apparatus according to claim 13, configured to perform the following:
19. If the first quantity is less than or equal to the fifth quantity, and the virtual speaker of the first quantity and the voting value of the first quantity are determined based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set, and the voting round quantity, the virtual speaker selection module, The method involves obtaining a representative coefficient of the third quantity in the current frame, wherein the representative coefficient of the third quantity includes the first representative coefficient and the second representative coefficient. The first voting value of the fifth quantity of the fifth virtual speaker is obtained by performing a voting round of the voting round quantity using the first representative coefficient, wherein the first voting value of the fifth quantity includes the first voting value of the first virtual speaker. The method involves obtaining a second voting value for the fifth quantity of the fifth virtual speaker, which is obtained by performing a voting round of the voting round quantity using the second representative coefficient, wherein the second voting value for the fifth quantity includes the second voting value of the first virtual speaker. Based on the first vote value of the fifth quantity, a virtual speaker of the eighth quantity is selected from the virtual speakers of the fifth quantity, wherein the eighth quantity is less than the fifth quantity. Based on the second voting value of the fifth quantity, a virtual speaker for the ninth quantity is selected from the virtual speakers of the fifth quantity, wherein the ninth quantity is less than the fifth quantity, and there is no common portion between the virtual speaker of the eighth quantity and the virtual speaker of the ninth quantity. The method involves obtaining the first quantity virtual speaker and the first quantity vote value based on the first vote value of the eighth quantity virtual speaker and the second vote value of the ninth quantity virtual speaker, wherein the first quantity virtual speaker includes the eighth quantity virtual speaker and the ninth quantity virtual speaker. The apparatus according to claim 13, configured to perform the following:
20. When obtaining the first voting value of the fifth quantity of the fifth quantity of the virtual speaker, which is obtained by performing a voting round of the voting round quantity using the first representative coefficient, the virtual speaker selection module, Based on the coefficient of the virtual speaker of the fifth quantity and the first representative coefficient, the first voting value of the fifth quantity is determined. The apparatus according to claim 17, configured as follows.
21. The apparatus further comprises a coefficient selection module, and when obtaining a representative coefficient of the third quantity of the current frame, the coefficient selection module, Obtain the coefficient of the fourth quantity in the current frame, and the frequency domain feature value of the coefficient of the fourth quantity, Based on the frequency domain characteristic value of the coefficient of the fourth quantity, a representative coefficient of the third quantity is selected from the coefficient of the fourth quantity, wherein the third quantity is less than the fourth quantity. The apparatus according to claim 17, configured to perform the following:
22. The virtual speaker selection module is, The first correlation is obtained between the current frame and a representative virtual speaker set for a previous frame, wherein the representative virtual speaker set for the previous frame includes a sixth quantity of virtual speakers, the virtual speakers included in the sixth quantity of virtual speakers are representative virtual speakers for the previous frame used to encode the previous frame of the three-dimensional audio signal, and the first correlation is used to determine whether to reuse the representative virtual speaker set for the previous frame when the current frame is encoded. If the first correlation does not satisfy the reuse conditions, the coefficient of the fourth quantity of the current frame of the three-dimensional audio signal and the frequency domain feature value of the coefficient of the fourth quantity are obtained. The apparatus according to claim 21, further configured to perform the following:
23. When selecting a representative virtual speaker for the second quantity for the current frame from the virtual speakers of the first quantity based on the voting value of the first quantity, the virtual speaker selection module: Obtaining a virtual speaker for a seventh quantity and the final vote value for the seventh quantity of the current frame, corresponding to the current frame, based on the vote value for the first quantity and the final vote value for the sixth quantity of the previous frame, wherein the virtual speaker for the seventh quantity includes the virtual speaker for the first quantity, the virtual speaker for the seventh quantity includes the virtual speaker for the sixth quantity, the virtual speaker for the sixth quantity included in the representative virtual speaker set for the previous frame corresponds one-to-one with the final vote value for the sixth quantity of the previous frame, and the virtual speaker for the sixth quantity is a virtual speaker used when the previous frame of the three-dimensional audio signal is encoded. Based on the final voting value of the seventh quantity for the current frame, a representative virtual speaker for the second quantity for the current frame is selected from the virtual speakers of the seventh quantity, wherein the second quantity is less than the seventh quantity. The apparatus according to claim 13, configured to perform the following:
24. The apparatus according to claim 13, wherein the current frame of the three-dimensional audio signal is a higher-order ambisonics HOA signal, and the frequency domain feature values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
25. An encoder comprising at least one processor and a memory, wherein the memory is configured to store a computer program, and as a result, when the computer program is executed by the at least one processor, the three-dimensional audio signal coding method according to claim 1 is implemented.
26. A system comprising an encoder according to claim 25 and a decoder, wherein the encoder is configured to perform the operation steps of the method according to claim 1, and the decoder is configured to decode a bitstream generated by the encoder.
27. A computer program wherein, when the computer program is executed, the three-dimensional audio signal encoding method described in claim 1 is implemented.
28. A computer-readable storage medium comprising computer software instructions, wherein when the computer software instructions are executed on an encoder, the encoder is enabled to perform the three-dimensional audio signal encoding method described in claim 1.
Citation Information
Patent Citations
Information processing device, voice processing method and voice processing program
JP2014207568A
Audio signal processor and method for processing encoded multi-channel audio signals
JP2014520473A
Method and Apparatus for Decoding Compressed Hoa Representation and Method and Apparatus for Encoding Compressed Hoa Representation
JP2017523453A
Efficient rendering of virtual soundfields
US20190379992A1