Three-dimensional audio signal encoding method, apparatus, and encoder
By judging the correlation and coefficient characteristics of three-dimensional audio signal frames, the selection of virtual speakers is optimized, the computational complexity of three-dimensional audio signal encoding is reduced, the encoding efficiency and audio stability are improved, and the sound quality is ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-17
- Publication Date
- 2026-04-07
AI Technical Summary
The high computational complexity of existing 3D audio signal encoding results in an excessive computational burden on the encoder, affecting encoding efficiency and sound quality stability.
By determining the correlation between the current frame and the previous frame of the 3D audio signal, it is decided whether to reuse the representative set of virtual speakers from the previous frame for encoding, or to select representative coefficients for voting to choose virtual speakers, thereby reducing computational complexity and improving encoding efficiency.
It effectively reduces the computational complexity of 3D audio signal compression coding, alleviates the encoder burden, improves the continuity between frames and the stability of reconstructed audio, and ensures sound quality.
Smart Images

Figure CN115376528B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia, and in particular to a three-dimensional audio signal encoding method, apparatus and encoder. Background Technology
[0002] With the rapid development of high-performance computers and signal processing technology, listeners have placed increasingly higher demands on voice and audio experiences, and immersive audio can meet these needs. For example, 3D audio technology has been widely used in wireless communication (such as 4G / 5G), virtual reality / augmented reality, and media audio. 3D audio technology is an audio technology that acquires, processes, transmits, renders, and plays back real-world sounds and 3D sound field information, giving sound a strong sense of space, immersion, and presence, providing listeners with an extraordinary auditory experience that makes them feel as if they are actually there.
[0003] Typically, acquisition devices (e.g., microphones) collect large amounts of data to record 3D sound field information and transmit the 3D audio signals to playback devices (e.g., speakers, headphones) for playback. Due to the large data volume of 3D sound field information, a significant amount of storage space is required, and the bandwidth demand for transmitting the 3D audio signals is also high. To address these issues, the 3D audio signals can be compressed, and the compressed data can be stored or transmitted. Currently, the encoder first iterates through the set of candidate virtual speakers and uses the selected virtual speakers to compress the 3D audio signal. Therefore, the computational complexity of the encoder for compressing the 3D audio signal is high. How to reduce the computational complexity of compressing the 3D audio signal is a problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides a method, apparatus, and encoder for encoding three-dimensional audio signals, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals.
[0005] Firstly, this application provides a three-dimensional audio signal encoding method, which can be executed by an encoder, specifically including the following steps: After the encoder obtains the first correlation between the current frame of the three-dimensional audio signal and the representative virtual speaker set of the previous frame, it determines whether the first correlation satisfies the multiplexing condition. If the first correlation satisfies the multiplexing condition, the current frame is encoded according to the representative virtual speaker set of the previous frame to obtain a bitstream. Wherein, the virtual speakers in the representative virtual speaker set of the previous frame are the virtual speakers used to encode the previous frame of the three-dimensional audio signal, and the first correlation is used to determine whether the representative virtual speaker set of the previous frame is reused when encoding the current frame.
[0006] In this way, the encoder can first determine whether the representative virtual speaker set of the previous frame can be reused to encode the current frame. If the encoder reuses the representative virtual speaker set of the previous frame to encode the current frame, it avoids performing the virtual speaker search process again, effectively reducing the computational complexity of the encoder searching for virtual speakers. Therefore, it reduces the computational complexity of compressing and encoding the 3D audio signal and alleviates the computational burden on the encoder. In addition, it can also reduce the frequent jumps of virtual speakers between frames, enhance the continuity of the orientation between frames, improve the stability of the sound image of the reconstructed 3D audio signal, and ensure the sound quality of the reconstructed 3D audio signal.
[0007] If the encoder cannot reuse the representative virtual speaker set of the previous frame to encode the current frame, the encoder selects representative coefficients and uses the representative coefficients of the current frame to vote on each virtual speaker in the candidate virtual speaker set. Based on the voting value, the representative virtual speaker of the current frame is selected, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden on the encoder.
[0008] In one possible implementation, after obtaining the first correlation between the current frame of the 3D audio signal and the representative virtual loudspeakers of the previous frame, the method further includes: the encoder obtaining a second correlation between the current frame and a set of candidate virtual loudspeakers, the second correlation being used to determine whether to use the set of candidate virtual loudspeakers when encoding the current frame, wherein the set of representative virtual loudspeakers of the previous frame is a proper subset of the set of candidate virtual loudspeakers; the reuse condition includes: the first correlation being greater than the second correlation, indicating that the encoder is more inclined to reuse the set of representative virtual loudspeakers of the previous frame to encode the current frame relative to the set of candidate virtual loudspeakers.
[0009] Optionally, obtaining the first correlation between the current frame of the three-dimensional audio signal and the set of representative virtual speakers of previous frames includes: the encoder obtaining the correlation between the current frame and the representative virtual speakers of each previous frame in the set of representative virtual speakers of previous frames; and taking the maximum correlation between the correlation between each representative virtual speaker of previous frames and the current frame as the first correlation.
[0010] For example, the representative set of virtual speakers in the preceding frame includes a first virtual speaker. Obtaining a first correlation between the current frame of the three-dimensional audio signal and the representative set of virtual speakers in the preceding frame includes: the encoder determining the correlation between the current frame and the first virtual speaker based on the coefficients of the current frame and the coefficients of the first virtual speaker.
[0011] Optionally, obtaining the second correlation between the current frame and the candidate virtual speaker set includes: obtaining the correlation between the current frame and each candidate virtual speaker in the candidate virtual speaker set; and taking the maximum correlation between each candidate virtual speaker and the current frame as the second correlation.
[0012] Thus, the encoder selects the typical maximum correlation from multiple correlations and uses the maximum correlation to determine whether the representative virtual speaker set of the previous frame can be reused to encode the current frame. Under the premise of ensuring accurate judgment, the computational complexity of compressing and encoding three-dimensional audio signals is reduced and the computational burden of the encoder is alleviated.
[0013] In another possible implementation, after obtaining the first correlation between the current frame of the 3D audio signal and the representative virtual loudspeakers of the previous frame, the method further includes: obtaining the third correlation between the current frame and the first subset of the candidate virtual loudspeaker set, the third correlation being used to determine whether the first subset of the candidate virtual loudspeaker set is used when encoding the current frame, the first subset being a proper subset of the candidate virtual loudspeaker set; the reuse condition includes: the first correlation being greater than the third correlation, indicating that relative to the first subset of the candidate virtual loudspeaker set, the encoder is more inclined to reuse the representative virtual loudspeaker set of the previous frame to encode the current frame.
[0014] In another possible implementation, after obtaining the first correlation between the current frame of the 3D audio signal and the representative virtual speakers of the previous frame, the method further includes: the encoder obtaining a fourth correlation between the current frame and a second subset of the candidate virtual speaker set, the fourth correlation being used to determine whether the second subset of the candidate virtual speaker set is used when encoding the current frame, the second subset being a proper subset of the candidate virtual speaker set; if the first correlation is less than or equal to the fourth correlation, obtaining a fifth correlation between the current frame and a third subset of the candidate virtual speaker set, the fifth correlation being used to determine whether the third subset of the candidate virtual speaker set is used when encoding the current frame, the third subset being a proper subset of the candidate virtual speaker set, the virtual speakers included in the second subset being different from or partially different from the virtual speakers included in the third subset; the reuse condition includes: the first correlation being greater than the fifth correlation, indicating that relative to the third subset of the candidate virtual speaker set, the encoder is more inclined to reuse the representative virtual speaker set of the previous frame to encode the current frame. In this way, the encoder makes more thorough multi-level judgments on different subsets of the candidate virtual speaker set, ensuring the accuracy of reusing the representative virtual speaker set of the previous frame when encoding the current frame.
[0015] In another possible implementation, if the first relevance does not meet the reuse condition, the method further includes: after the encoder obtains a fourth number of coefficients of the current frame of the 3D audio signal and the frequency domain feature values of the fourth number of coefficients, it selects a third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients; then, it selects a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on the third number of representative coefficients; and finally, it encodes the current frame based on the second number of representative virtual speakers for the current frame to obtain a bitstream. The fourth number of coefficients includes the third number of representative coefficients; the third number being less than the fourth number indicates that the third number of representative coefficients is a subset of the fourth number of coefficients. The current frame of the 3D audio signal is a higher-order ambisonics (HOA) signal; the frequency domain feature values of the coefficients are determined based on the coefficients of the HOA signal.
[0016] Thus, since the encoder selects a portion of the coefficients from all coefficients in the current frame as representative coefficients, and uses a smaller number of representative coefficients to replace all coefficients in the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder searching for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden on the encoder.
[0017] In addition, the encoder encodes the current frame based on the representative virtual speakers of the second number of current frames to obtain the bitstream, including: the encoder generates a virtual speaker signal based on the representative virtual speakers of the second number of current frames and the current frame; and encodes the virtual speaker signal to obtain the bitstream.
[0018] Since the frequency domain eigenvalues of the coefficients of the current frame characterize the sound field characteristics of the three-dimensional audio signal, the encoder selects representative coefficients of the representative sound field components of the current frame based on the frequency domain eigenvalues of the coefficients of the current frame. The representative virtual loudspeaker of the current frame selected from the candidate virtual loudspeaker set using the representative coefficients can fully characterize the sound field characteristics of the three-dimensional audio signal, thereby further improving the accuracy of the encoder in generating virtual loudspeaker signals when compressing the three-dimensional audio signal to be encoded using the representative virtual loudspeaker of the current frame. This is to improve the compression rate of the three-dimensional audio signal and reduce the bandwidth occupied by the encoder's transmission bitstream.
[0019] In another possible implementation, selecting a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on a third number of representative coefficients includes: the encoder determining a first number of virtual speakers and a first number of voting values based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds; selecting a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on the first number of voting values; the second number being less than the first number indicates that the second number of representative virtual speakers for the current frame are a subset of the virtual speakers in the candidate virtual speaker set. Understandably, there is a one-to-one correspondence between virtual speakers and voting values. For example, the first number of virtual speakers includes the first virtual speaker, and the first number of voting values includes the voting value of the first virtual speaker; the first virtual speaker corresponds to the voting value of the first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of using the first virtual speaker when encoding the current frame. The candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number, and the number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number.
[0020] Currently, in the virtual speaker search process, the encoder uses the correlation calculation results between the 3D audio signal to be encoded and the virtual speaker as the selection metric. Furthermore, if the encoder transmits a virtual speaker for each coefficient, efficient data compression cannot be achieved, placing a heavy computational burden on the encoder. The virtual speaker selection method provided in this application involves the encoder using a smaller number of representative coefficients to vote on each virtual speaker in the candidate virtual speaker set, replacing all coefficients of the current frame, and selecting the representative virtual speaker for the current frame based on the voting values. Furthermore, the encoder uses the representative virtual speaker of the current frame to compress and encode the 3D audio signal to be encoded, effectively improving the compression ratio of the 3D audio signal and reducing the computational complexity of the encoder's virtual speaker search, thereby reducing the computational complexity of the 3D audio signal compression and encoding and alleviating the encoder's computational burden.
[0021] The second quantity represents the number of representative virtual speakers for the current frame selected by the encoder. A larger second quantity indicates a larger number of representative virtual speakers for the current frame, resulting in more sound field information in the 3D audio signal; a smaller second quantity indicates a smaller number of representative virtual speakers for the current frame, resulting in less sound field information in the 3D audio signal. Therefore, the number of representative virtual speakers for the current frame selected by the encoder can be controlled by setting the second quantity. For example, the second quantity can be preset, or it can be determined based on the current frame. For instance, the value of the second quantity can be 1, 2, 4, or 8.
[0022] In another possible implementation, selecting a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on a first number of voting values includes: the encoder obtaining a seventh number of final voting values for the current frame corresponding to a seventh number of virtual speakers based on the first number of voting values and a sixth number of final voting values for previous frames; selecting a second number of representative virtual speakers for the current frame from the seventh number of virtual speakers based on the seventh number of final voting values for the current frame; the second number being less than the seventh number indicates that the representative virtual speakers for the second number of current frames are a subset of the seventh number of virtual speakers. The seventh number of virtual speakers includes the first number of virtual speakers and also includes a sixth number of virtual speakers, where the virtual speakers included in the sixth number of virtual speakers are representative virtual speakers for the previous frames used to encode the 3D audio signal. The sixth number of virtual speakers included in the set of representative virtual speakers for the previous frames correspond one-to-one with the final voting values of the sixth number of previous frames.
[0023] During the virtual speaker search process, the locations of real sound sources and virtual speakers may not coincide, leading to a one-to-one correspondence between virtual speakers and real sound sources. Furthermore, in complex real-world scenarios, a limited set of virtual speakers may not be able to represent all sound sources in the sound field. In such cases, frequent jumps in the virtual speakers found between frames can significantly impact the listener's auditory experience, resulting in noticeable discontinuities and noise in the decoded and reconstructed 3D audio signal. The virtual speaker selection method provided in this application inherits the representative virtual speaker from previous frames. Specifically, for virtual speakers with the same number, the initial voting value of the current frame is adjusted using the final voting value of the previous frame. This makes the encoder more inclined to select the representative virtual speaker from the previous frame, thereby reducing frequent jumps in virtual speakers between frames, enhancing the continuity of signal orientation between frames, improving the stability of the sound image in the reconstructed 3D audio signal, and ensuring the sound quality of the reconstructed 3D audio signal.
[0024] Optionally, the method further includes: the encoder can also acquire the current frame of the three-dimensional audio signal so as to compress and encode the current frame of the three-dimensional audio signal to obtain a bitstream, and transmit the bitstream to the decoding end.
[0025] Secondly, this application provides a three-dimensional audio signal encoding apparatus, the apparatus comprising modules for performing the three-dimensional audio signal encoding method of the first aspect or any possible design of the first aspect. For example, the three-dimensional audio signal encoding apparatus includes a virtual speaker selection module and an encoding module. The virtual speaker selection module is used to obtain a first correlation between the current frame of the three-dimensional audio signal and a representative set of virtual speakers of a previous frame, wherein the virtual speakers in the representative set of virtual speakers of the previous frame are the virtual speakers used to encode the previous frame of the three-dimensional audio signal, and the first correlation is used to determine whether to reuse the representative set of virtual speakers of the previous frame when encoding the current frame; the encoding module is used to encode the current frame according to the representative set of virtual speakers of the previous frame if the first correlation satisfies the reuse condition, to obtain a bitstream.
[0026] Thirdly, this application provides an encoder including at least one processor and a memory, wherein the memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, it performs the operation steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation of the first aspect.
[0027] Fourthly, this application provides a system comprising an encoder as described in the third aspect and a decoder, wherein the encoder is configured to perform operational steps of the three-dimensional audio signal encoding method in the first aspect or any possible implementation thereof, and the decoder is configured to decode the bitstream generated by the encoder.
[0028] Fifthly, this application provides a computer-readable storage medium, comprising: computer software instructions; when the computer software instructions are executed in an encoder, causing the encoder to perform the operation steps of the method as described in the first aspect or any possible implementation thereof.
[0029] Sixthly, this application provides a computer program product that, when run on an encoder, causes the encoder to perform the operation steps of the method as described in the first aspect or any possible implementation thereof.
[0030] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the structure of an audio encoding and decoding system provided in an embodiment of this application;
[0032] Figure 2 This is a schematic diagram of an audio encoding / decoding system provided in an embodiment of this application;
[0033] Figure 3 This is a schematic diagram of the structure of an encoder provided in an embodiment of this application;
[0034] Figure 4 A flowchart illustrating a three-dimensional audio signal encoding and decoding method provided in an embodiment of this application;
[0035] Figure 5 A flowchart illustrating a method for selecting a virtual speaker, provided as an embodiment of this application;
[0036] Figure 6 A flowchart illustrating a three-dimensional audio signal encoding method provided in an embodiment of this application;
[0037] Figure 7 A flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application;
[0038] Figure 8 A flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application;
[0039] Figure 9 A flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application;
[0040] Figure 10 A schematic diagram of the structure of an encoding device provided in this application;
[0041] Figure 11 This is a schematic diagram of the structure of an encoder provided in this application. Detailed Implementation
[0042] To ensure clarity and brevity in the description of the following embodiments, a brief introduction to the relevant technologies is given first.
[0043] Sound is a continuous wave produced by the vibration of an object. The object that produces vibrations and emits sound waves is called the sound source. As sound waves propagate through a medium (such as air, solids, or liquids), the auditory organs of humans or animals can perceive the sound.
[0044] Sound waves are characterized by pitch, intensity, and timbre. Pitch indicates the highness or lowness of a sound. Intensity indicates the loudness or volume of a sound. The unit of intensity is the decibel (dB). Timbre is also known as tone color.
[0045] The frequency of a sound wave determines its pitch. The higher the frequency, the higher the pitch. The number of times an object vibrates per second is called its frequency, and the unit of frequency is hertz (Hz). The human ear can distinguish sounds with frequencies between 20 Hz and 20,000 Hz.
[0046] The amplitude of a sound wave determines its intensity. The greater the amplitude, the greater the intensity. The closer to the sound source, the greater the intensity.
[0047] The waveform of a sound wave determines its timbre. Sound wave waveforms include square waves, sawtooth waves, sine waves, and pulse waves, among others.
[0048] Based on the characteristics of sound waves, sound can be divided into regular sound and irregular sound. Irregular sound refers to sound emitted by the irregular vibration of a sound source. Irregular sound is, for example, noise that affects people's work, study, and rest. Regular sound refers to sound emitted by the regular vibration of a sound source. Regular sound includes speech and musical tones. When sound is represented electronically, regular sound is an analog signal that varies continuously in the time and frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.
[0049] Because human hearing has the ability to distinguish the location of sound sources in space, when a listener hears a sound in space, in addition to being able to perceive the pitch, intensity, and timbre of the sound, they can also perceive the location of the sound.
[0050] As people pay increasing attention to and demand higher quality in their auditory experience, three-dimensional audio technology has emerged to enhance the depth, presence, and spatial feel of sound. This allows listeners to not only perceive sounds from front, back, left, and right sources, but also to feel surrounded by the spatial sound field created by these sources, and to experience the sound spreading outwards, creating an immersive audio experience as if the listener were in a cinema or concert hall.
[0051] Three-dimensional audio technology refers to the concept of the space outside the human ear as a system, where the signal received at the eardrum is a three-dimensional audio signal output after the sound emitted from the sound source has been filtered by this external system. For example, the system outside the human ear can be defined as the system impulse response h(n), any sound source can be defined as x(n), and the signal received at the eardrum is the convolution result of x(n) and h(n). The three-dimensional audio signal described in this application's embodiments may refer to a higher-order ambisonics (HOA) signal. Three-dimensional audio can also be called three-dimensional sound effects, spatial audio, three-dimensional sound field reconstruction, virtual 3D audio, or binaural audio, etc.
[0052] As is well known, sound waves propagate in an ideal medium with a wave number of k = w / c and an angular frequency of w = 2πf, where f is the sound wave frequency and c is the speed of sound. The sound pressure p satisfies formula (1), ▽ 2 For the Laplace operator.
[0053] ▽ 2 p+k2 p = 0 Formula (1)
[0054] Assuming the spatial system outside the human ear is a sphere, with the listener at the center, the sound from outside the sphere has a projection on the sphere's surface. Filtering out sounds from outside the sphere, and assuming the sound sources are distributed on this sphere, the sound field generated by the sound sources on the sphere is used to fit the sound field generated by the original sound source. That is, three-dimensional audio technology is a method of fitting a sound field. Specifically, in spherical coordinates, equation (1) is solved. In the passive spherical region, the solution to equation (1) is as follows: equation (2).
[0055]
[0056] Where r represents the radius of the sphere, and θ represents the horizontal angle. denoted by pitch angle, k by wave number, s by amplitude of ideal plane wave, and m by order number of three-dimensional audio signal (or HOA signal). Let represent the spherical Bessel function, also known as the radial basis function, where the first 'j' represents the imaginary unit. It does not change with the angle. Represents θ, spherical harmonic function of direction, The spherical harmonic function represents the direction of the sound source. The coefficients of the three-dimensional audio signal satisfy formula (3).
[0057]
[0058] Substituting formula (3) into formula (2), formula (2) can be transformed into formula (4).
[0059]
[0060] in, The coefficients of the three-dimensional audio signal of order N are used to approximate the sound field. A sound field refers to the region in a medium where sound waves exist. N is an integer greater than or equal to 1. For example, the value of N ranges from 2 to 6. The coefficients of the three-dimensional audio signal described in the embodiments of this application may refer to HOA coefficients or ambisonic coefficients.
[0061] A three-dimensional audio signal is an information carrier that carries the spatial location information of the sound source in the sound field, describing the sound field of the listener in space. Equation (4) shows that the sound field can be expanded on a sphere according to the spherical harmonic function, that is, the sound field can be decomposed into the superposition of multiple plane waves. Therefore, the sound field described by the three-dimensional audio signal can be expressed by the superposition of multiple plane waves, and the sound field can be reconstructed through the coefficients of the three-dimensional audio signal.
[0062] Compared to a 5.1 channel audio signal or a 7.1 channel audio signal, an Nth-order HOA signal has (N+1) 2 With multiple channels, the HOA signal contains a large amount of data describing the spatial information of the sound field. If the acquisition device (e.g., a microphone) transmits this 3D audio signal to the playback device (e.g., a speaker), it consumes a significant amount of bandwidth. Currently, encoders can use spatial squeezed surround audio coding (S3AC) or directional audio coding (DirAC) to compress and encode the 3D audio signal to obtain a bitstream, which is then transmitted to the playback device. The playback device decodes the bitstream, reconstructs the 3D audio signal, and plays the reconstructed 3D audio signal. This reduces the amount of data transmitted to the playback device and the bandwidth usage. However, the computational complexity of compressing and encoding 3D audio signals is high, consuming excessive computational resources. Therefore, how to reduce the computational complexity of compressing and encoding 3D audio signals is a problem that urgently needs to be solved.
[0063] This application provides an audio encoding and decoding technology, particularly a three-dimensional audio encoding and decoding technology for three-dimensional audio signals. Specifically, it provides an encoding and decoding technology that uses fewer channels to represent three-dimensional audio signals, thereby improving traditional audio encoding and decoding systems. Audio encoding (or commonly referred to as encoding) includes two parts: audio encoding and audio decoding. Audio encoding is performed on the source side and typically includes processing (e.g., compressing) the raw audio to reduce the amount of data required to represent the raw audio, thereby enabling more efficient storage and / or transmission. Audio decoding is performed on the destination side and typically includes inverse processing relative to the encoder to reconstruct the original audio. The encoding and decoding parts are also collectively referred to as encoding and decoding. The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0064] Figure 1 This is a schematic diagram of an audio encoding / decoding system provided in an embodiment of this application. The audio encoding / decoding system 100 includes a source device 110 and a destination device 120. The source device 110 is used to compress and encode a three-dimensional audio signal to obtain a bitstream, and transmit the bitstream to the destination device 120. The destination device 120 decodes the bitstream, reconstructs the three-dimensional audio signal, and plays the reconstructed three-dimensional audio signal.
[0065] Specifically, the source device 110 includes an audio acquisition unit 111, a preprocessor 112, an encoder 113, and a communication interface 114.
[0066] Audio acquirer 111 is used to acquire raw audio. Audio acquirer 111 can be any type of audio acquisition device for capturing real-world sounds, and / or any type of audio generation device. Audio acquirer 111 is, for example, a computer audio processor for generating computer audio. Audio acquirer 111 can also be any type of memory or storage device for storing audio. Audio includes real-world sounds, virtual scene sounds (e.g., VR or augmented reality (AR) sounds), and / or any combination thereof.
[0067] The preprocessor 112 receives the raw audio acquired by the audio acquirer 111 and preprocesses the raw audio to obtain a three-dimensional audio signal. For example, the preprocessing performed by the preprocessor 112 includes channel conversion, audio format conversion, or noise reduction.
[0068] Encoder 113 receives the three-dimensional audio signal generated by preprocessor 112 and compresses and encodes the three-dimensional audio signal to obtain a bitstream. For example, encoder 113 may include spatial encoder 1131 and core encoder 1132. Spatial encoder 1131 selects (or searches for) virtual speakers from a set of candidate virtual speakers based on the three-dimensional audio signal, and generates a virtual speaker signal based on the three-dimensional audio signal and the virtual speakers. The virtual speaker signal can also be called a playback signal. Core encoder 1132 encodes the virtual speaker signal to obtain a bitstream.
[0069] The communication interface 114 is used to receive the code stream generated by the encoder 113 and send the code stream to the destination device 120 through the communication channel 130 so that the destination device 120 can reconstruct the three-dimensional audio signal based on the code stream.
[0070] The target device 120 includes a player 121, a post-processor 122, a decoder 123, and a communication interface 124.
[0071] Communication interface 124 is used to receive the bitstream sent by communication interface 114 and transmit the bitstream to decoder 123 so that decoder 123 can reconstruct the three-dimensional audio signal based on the bitstream.
[0072] Communication interfaces 114 and 124 can be used to send or receive relevant data of the original audio through a direct communication link between the source device 110 and the destination device 120, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.
[0073] Both communication interface 114 and communication interface 124 can be configured as follows: Figure 1The arrow pointing from the source device 110 to the corresponding communication channel 130 of the destination device 120 indicates a one-way or two-way communication interface, which can be used to send and receive messages, establish connections, acknowledge and exchange any other information related to the communication link and / or data transmission such as encoded bitstream transmission, etc.
[0074] Decoder 123 is used to decode the bitstream and reconstruct the three-dimensional audio signal. For example, decoder 123 includes a core decoder 1231 and a spatial decoder 1232. Core decoder 1231 decodes the bitstream to obtain a virtual speaker signal. Spatial decoder 1232 reconstructs the three-dimensional audio signal based on a set of candidate virtual speakers and the virtual speaker signal, obtaining the reconstructed three-dimensional audio signal.
[0075] The post-processor 122 receives the reconstructed 3D audio signal generated by the decoder 123 and performs post-processing on the reconstructed 3D audio signal. For example, the post-processing performed by the post-processor 122 includes audio rendering, loudness normalization, user interaction, audio format conversion, or noise reduction.
[0076] Player 121 is used to play the reconstructed sound based on the reconstructed 3D audio signal.
[0077] It should be noted that the audio acquirer 111 and encoder 113 can be integrated into a single physical device or located on different physical devices; there is no limitation on this. For example, such as... Figure 1 The source device 110 shown includes an audio acquirer 111 and an encoder 113, indicating that the audio acquirer 111 and encoder 113 are integrated into a single physical device. Therefore, the source device 110 can also be referred to as a capture device. The source device 110 can be, for example, a media gateway in a wireless access network, a media gateway in a core network, a transcoding device, a media resource server, an AR device, a VR device, a microphone, or other audio capture devices. If the source device 110 does not include the audio acquirer 111, it means that the audio acquirer 111 and encoder 113 are two different physical devices, and the source device 110 can acquire raw audio from other devices (such as audio capture devices or audio storage devices).
[0078] Furthermore, the player 121 and decoder 123 can be integrated into a single physical device or located on different physical devices; there is no limitation on this. For example, such as... Figure 1The destination device 120 shown includes a player 121 and a decoder 123, indicating that the player 121 and decoder 123 are integrated into a single physical device. Therefore, the destination device 120 can also be referred to as a playback device. The destination device 120 has the function of decoding and playing back the reconstructed audio. The destination device 120 can be, for example, a speaker, headphones, or other audio playback device. If the destination device 120 does not include the player 121, it means that the player 121 and decoder 123 are two different physical devices. After decoding and reconstructing the three-dimensional audio signal from the bitstream, the destination device 120 transmits the reconstructed three-dimensional audio signal to other playback devices (such as speakers or headphones) for playback.
[0079] also, Figure 1 It is shown that the source device 110 and the destination device 120 can be integrated into one physical device or set on different physical devices, without limitation.
[0080] For example, such as Figure 2 As shown in (a), source device 110 can be a microphone in a recording studio, and destination device 120 can be a speaker. Source device 110 can acquire the original audio of various musical instruments, transmit the original audio to an encoding / decoding device, perform encoding / decoding processing on the original audio to obtain a reconstructed three-dimensional audio signal, and then play back the reconstructed three-dimensional audio signal by destination device 120. As another example, source device 110 can be a microphone in a terminal device, and destination device 120 can be headphones. Source device 110 can acquire external sounds or audio synthesized by the terminal device.
[0081] For example, such as Figure 2 As shown in (b), if the source device 110 and the destination device 120 are integrated into a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, or an extended reality (XR) device, then the VR / AR / MR / XR device has the functions of capturing raw audio, playing back audio, and encoding / decoding. The source device 110 can capture the sound emitted by the user and the sound emitted by virtual objects in the virtual environment in which the user is located.
[0082] In these embodiments, the source device 110 or its corresponding functions and the destination device 120 or its corresponding functions may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof. As described, Figure 1 The presence and division of different units or functions in the source device 110 and / or destination device 120 shown may vary depending on the actual device and application, which is obvious to those skilled in the art.
[0083] The structure of the audio codec system described above is only illustrative. In some possible implementations, the audio codec system may also include other devices, such as end-side devices or cloud-side devices. After the source device 110 acquires the raw audio, it preprocesses the raw audio to obtain a three-dimensional audio signal; and then transmits the three-dimensional audio to the end-side device or cloud-side device, which performs the encoding and decoding function on the three-dimensional audio signal.
[0084] The audio signal encoding and decoding method provided in this application is mainly applied at the encoding end. Combined with... Figure 3 The structure of the encoder is described in detail. For example... Figure 3 As shown, the encoder 300 includes a virtual speaker configuration unit 310, a virtual speaker set generation unit 320, an encoding analysis unit 330, a virtual speaker selection unit 340, a virtual speaker signal generation unit 350, and an encoding unit 360.
[0085] The virtual speaker configuration unit 310 generates virtual speaker configuration parameters based on encoder configuration information to obtain multiple virtual speakers. Encoder configuration information includes, but is not limited to: the order of the three-dimensional audio signal (or commonly referred to as the HOA order), the encoding bit rate, user-defined information, etc. Virtual speaker configuration parameters include, but are not limited to: the number of virtual speakers, the order of the virtual speakers, the position coordinates of the virtual speakers, etc. The number of virtual speakers can be, for example, 2048, 1669, 1343, 1024, 530, 512, 256, 128, or 64. The order of the virtual speakers can be any from 2nd to 6th order. The position coordinates of the virtual speakers include the horizontal angle and the pitch angle.
[0086] The virtual speaker configuration parameters output by the virtual speaker configuration unit 310 are used as input to the virtual speaker set generation unit 320.
[0087] The virtual speaker set generation unit 320 is used to generate a candidate virtual speaker set based on virtual speaker configuration parameters. The candidate virtual speaker set includes multiple virtual speakers. Specifically, the virtual speaker set generation unit 320 determines the multiple virtual speakers included in the candidate virtual speaker set based on the number of virtual speakers, and determines the coefficients of the virtual speakers based on the position information (e.g., coordinates) and order of the virtual speakers. For example, the method for determining the coordinates of the virtual speakers includes, but is not limited to: generating multiple virtual speakers according to an equidistant rule, or generating multiple non-uniformly distributed virtual speakers based on the principle of auditory perception; then, the coordinates of the virtual speakers are generated based on the number of virtual speakers.
[0088] Based on the above principle of generating three-dimensional audio signals, the coefficients of a virtual speaker can also be generated. The coefficients of the virtual speaker can be calculated using θ in formula (3). s and Set the position coordinates of the virtual speaker respectively. This represents the coefficients of an Nth-order virtual loudspeaker. The coefficients of a virtual loudspeaker can also be called ambisonic coefficients.
[0089] The encoding analysis unit 330 is used to perform encoding analysis on three-dimensional audio signals, such as analyzing the sound field distribution characteristics of three-dimensional audio signals, that is, the number of sound sources, the directionality of sound sources, and the dispersion of sound sources.
[0090] The coefficients of multiple virtual speakers included in the candidate virtual speaker set output by the virtual speaker set generation unit 320 are used as input to the virtual speaker selection unit 340.
[0091] The sound field distribution characteristics of the three-dimensional audio signal output by the encoding analysis unit 330 are used as the input of the virtual speaker selection unit 340.
[0092] The virtual speaker selection unit 340 is used to determine a representative virtual speaker that matches the three-dimensional audio signal based on the three-dimensional audio signal to be encoded, the sound field distribution characteristics of the three-dimensional audio signal, and the coefficients of multiple virtual speakers.
[0093] Not limited to this, the encoder 300 in this embodiment may also exclude the encoding analysis unit 330, that is, the encoder 300 may not analyze the input signal, and the virtual speaker selection unit 340 may use a default configuration to determine the representative virtual speaker. For example, the virtual speaker selection unit 340 may determine the representative virtual speaker that matches the three-dimensional audio signal only based on the coefficients of the three-dimensional audio signal and multiple virtual speakers.
[0094] The encoder 300 can use a three-dimensional audio signal acquired from the acquisition device or a three-dimensional audio signal synthesized using artificial audio objects as its input. Furthermore, the three-dimensional audio signal input to the encoder 300 can be either a time-domain three-dimensional audio signal or a frequency-domain three-dimensional audio signal, without limitation.
[0095] The virtual speaker selection unit 340 outputs the position information and coefficients representing the virtual speaker as inputs to the virtual speaker signal generation unit 350 and the encoding unit 360.
[0096] The virtual speaker signal generation unit 350 generates a virtual speaker signal based on a three-dimensional audio signal and attribute information representing a virtual speaker. The attribute information representing the virtual speaker includes at least one of position information representing the virtual speaker, coefficients representing the virtual speaker, and coefficients of the three-dimensional audio signal. If the attribute information is position information representing the virtual speaker, the coefficients representing the virtual speaker are determined based on the position information; if the attribute information includes coefficients of the three-dimensional audio signal, the coefficients representing the virtual speaker are obtained based on the coefficients of the three-dimensional audio signal. Specifically, the virtual speaker signal generation unit 350 calculates the virtual speaker signal based on the coefficients of the three-dimensional audio signal and the coefficients representing the virtual speaker.
[0097] For example, assume matrix A represents the coefficients of the virtual loudspeaker, and matrix X represents the HOA coefficients of the HOA signal. Matrix X is the inverse of matrix A. The theoretical optimal solution w is obtained using the least squares method, where w represents the virtual loudspeaker signal. The virtual loudspeaker signal satisfies formula (5).
[0098] w = A -1 Formula X (5)
[0099] Among them, A -1 Let represent the inverse matrix of matrix A. Matrix A has a size of (M×C), where C represents the number of virtual loudspeakers, M represents the number of channels in the Nth-order HOA signal, and 'a' represents the coefficients of the virtual loudspeakers. Matrix X has a size of (M×L), where L represents the number of coefficients in the HOA signal, and 'x' represents the coefficients of the HOA signal. The coefficients representing virtual loudspeakers can refer to either the HOA coefficients or the ambisonic coefficients of the virtual loudspeakers. For example,
[0100] The virtual speaker signal output by the virtual speaker signal generation unit 350 is used as the input of the encoding unit 360.
[0101] The encoding unit 360 is used to perform core encoding processing on the virtual speaker signal to obtain a bitstream. Core encoding processing includes, but is not limited to: transformation, quantization, psychoacoustic modeling, noise shaping, bandwidth expansion, downmixing, arithmetic coding, and bitstream generation.
[0102] It is worth noting that the spatial encoder 1131 may include a virtual loudspeaker configuration unit 310, a virtual loudspeaker set generation unit 320, an encoding analysis unit 330, a virtual loudspeaker selection unit 340, and a virtual loudspeaker signal generation unit 350. That is, the virtual loudspeaker configuration unit 310, the virtual loudspeaker set generation unit 320, the encoding analysis unit 330, the virtual loudspeaker selection unit 340, and the virtual loudspeaker signal generation unit 350 implement the functions of the spatial encoder 1131. The core encoder 1132 may include an encoding unit 360. That is, the encoding unit 360 implements the functions of the core encoder 1132.
[0103] Figure 3 The encoder shown can generate one virtual speaker signal or multiple virtual speaker signals. Multiple virtual speaker signals can be generated by... Figure 3 The encoder shown can be obtained by multiple executions, or it can be obtained by... Figure 3 The encoder shown is obtained by executing it once.
[0104] Next, the encoding and decoding process of three-dimensional audio signals will be explained with reference to the accompanying drawings. Figure 4 This is a flowchart illustrating a three-dimensional audio signal encoding and decoding method provided in an embodiment of this application. Figure 1 The process of encoding and decoding three-dimensional audio signals, using the source device 110 and the destination device 120 as an example, will be explained below. Figure 4 As shown, the method includes the following steps.
[0105] S410, source device 110 acquires the current frame of the three-dimensional audio signal.
[0106] As described in the above embodiments, if the source device 110 carries an audio acquisition device 111, the source device 110 can acquire raw audio through the audio acquisition device 111. Optionally, the source device 110 can also receive raw audio acquired by other devices; or acquire raw audio from the memory or other memory in the source device 110. The raw audio may include at least one of real-world sounds acquired in real time, audio stored in the device, and audio synthesized from multiple audio sources. This embodiment does not limit the method of acquiring raw audio or the type of raw audio.
[0107] After acquiring the original audio, the source device 110 generates a three-dimensional audio signal based on three-dimensional audio technology and the original audio, so as to provide the listener with an "immersive" sound effect when playing back the original audio. The specific method for generating the three-dimensional audio signal can be found in the description of the preprocessor 112 in the above embodiments and in the description of existing technologies.
[0108] Furthermore, an audio signal is a continuous analog signal. In audio signal processing, the audio signal can first be sampled to generate a digital signal of frame sequences. A frame can include multiple sample points. A frame can also refer to the sample points obtained from sampling. A frame can also include subframes obtained by dividing a frame. For example, if a frame has a length of L sample points and is divided into N subframes, then each subframe corresponds to L / N sample points. Audio encoding and decoding typically refer to processing audio frame sequences containing multiple sample points.
[0109] An audio frame may include the current frame or a previous frame. In various embodiments of this application, the current frame or previous frame can refer to a frame or a subframe. The current frame refers to the frame undergoing encoding / decoding processing at the current moment. A previous frame refers to a frame that underwent encoding / decoding processing at a time prior to the current moment. A previous frame can be a frame from the moment before the current moment or from several previous moments. In embodiments of this application, the current frame of a 3D audio signal refers to a frame of 3D audio signal undergoing encoding / decoding processing at the current moment. A previous frame refers to a frame of 3D audio signal that underwent encoding / decoding processing at a time prior to the current moment. The current frame of a 3D audio signal can refer to the current frame of the 3D audio signal to be encoded. The current frame of a 3D audio signal can be simply referred to as the current frame. The previous frame of a 3D audio signal can be simply referred to as the previous frame.
[0110] S420 and source device 110 determine the candidate set of virtual speakers.
[0111] In one scenario, the source device 110 has a pre-configured set of candidate virtual speakers in its memory. The source device 110 can read the set of candidate virtual speakers from the memory. The set of candidate virtual speakers includes multiple virtual speakers. Virtual speakers represent speakers that are virtually present in the spatial sound field. The virtual speakers are used to calculate virtual speaker signals based on the three-dimensional audio signal so that the destination device 120 can play back the reconstructed three-dimensional audio signal.
[0112] In another scenario, the source device 110 has virtual speaker configuration parameters pre-configured in its memory. The source device 110 generates a candidate set of virtual speakers based on the virtual speaker configuration parameters. Optionally, the source device 110 generates the candidate set of virtual speakers in real time based on its own computing resources (e.g., processor) capabilities and the characteristics of the current frame (e.g., channel and data volume).
[0113] The specific method for generating a candidate set of virtual speakers can be found in the prior art, as well as in the description of the virtual speaker configuration unit 310 and the virtual speaker set generation unit 320 in the above embodiments.
[0114] S430 and source device 110 select the representative virtual speaker of the current frame from the candidate virtual speaker set based on the current frame of the three-dimensional audio signal.
[0115] The source device 110 votes on the virtual speakers based on the coefficients of the current frame and the coefficients of the virtual speakers, and selects the representative virtual speaker for the current frame from the candidate virtual speaker set based on the voting values of the virtual speakers. A limited number of representative virtual speakers for the current frame are searched from the candidate virtual speaker set as the best matching virtual speaker for the current frame to be encoded, thereby achieving the purpose of data compression of the three-dimensional audio signal to be encoded.
[0116] Figure 5 This is a flowchart illustrating a method for selecting a virtual speaker, as provided in an embodiment of this application. Figure 5 The method described is... Figure 4 This section describes the specific operational procedures included in the S430. (Hereinafter provided is a summary of the details.) Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Specifically, the function of the virtual speaker selection unit 340 is implemented. For example... Figure 5 As shown, the method includes the following steps.
[0117] S510, encoder 113 obtains the representative coefficients of the current frame.
[0118] Representative coefficients can refer to frequency domain representative coefficients or time domain representative coefficients. Frequency domain representative coefficients can also be called frequency domain representative frequency points or spectrum representative coefficients. Time domain representative coefficients can also be called time domain representative sampling points. For specific methods to obtain the representative coefficients of the current frame, please refer to the following: Figure 8 Explanation of S650 and S660.
[0119] S520, encoder 113 selects the representative virtual speaker for the current frame from the candidate virtual speaker set based on the voting values of the virtual speakers in the candidate virtual speaker set according to the representative coefficient of the current frame. Execute S440 to S460.
[0120] Encoder 113 votes on the virtual speakers in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker, and selects (searches) the representative virtual speaker of the current frame from the candidate virtual speaker set based on the final voting value of the virtual speaker in the current frame. The specific method for selecting the representative virtual speaker of the current frame can be found below. Figure 6 , Figure 8 and Figure 9 Explanation of S670.
[0121] It should be noted that the encoder first traverses the virtual speakers included in the candidate virtual speaker set, and then compresses the current frame using the representative virtual speaker of the current frame selected from the candidate virtual speaker set. However, if the results of the virtual speakers selected in consecutive frames differ significantly, it will lead to unstable sound image of the reconstructed 3D audio signal, reducing the sound quality of the reconstructed 3D audio signal. In the embodiments of this application, the encoder 113 can update the initial voting value of the virtual speakers in the current frame included in the candidate virtual speaker set based on the final voting value of the representative virtual speaker of the previous frame, to obtain the final voting value of the virtual speaker in the current frame, and then select the representative virtual speaker of the current frame from the candidate virtual speaker set based on the final voting value of the virtual speaker in the current frame. Thus, by referring to the representative virtual speaker of the previous frame to select the representative virtual speaker of the current frame, the encoder tends to select the same virtual speaker as the representative virtual speaker of the previous frame when selecting the representative virtual speaker of the current frame, increasing the directional continuity between consecutive frames and overcoming the problem of large differences in the results of the virtual speakers selected in consecutive frames. Therefore, the embodiments of this application may also include S530.
[0122] S530, encoder 113 adjusts the initial voting value of the virtual speaker in the current frame of the candidate virtual speaker set according to the final voting value of the representative virtual speaker in the previous frame, and obtains the final voting value of the virtual speaker in the current frame.
[0123] Encoder 113 votes on the virtual speakers in the candidate virtual speaker set based on the representative coefficient of the current frame and the coefficient of the virtual speaker, obtaining the initial voting value of the virtual speakers in the current frame. Then, it adjusts the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame based on the final voting value of the representative virtual speaker in the previous frame, thus obtaining the final voting value of the virtual speakers in the current frame. The representative virtual speaker in the previous frame is the virtual speaker used by encoder 113 when encoding the previous frame. The specific method for adjusting the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame can be found below. Figure 9 The explanation of S6702a to S6702b.
[0124] In some embodiments, if the current frame is the first frame in the original audio, encoder 113 executes steps S510 to S520. If the current frame is any frame above the second frame in the original audio, encoder 113 may first determine whether to reuse the representative virtual speaker of the previous frame to encode the current frame or determine whether to perform a virtual speaker search, to ensure the continuity of orientation between consecutive frames and reduce encoding complexity. Embodiments of this application may also include step S540.
[0125] S540 and encoder 113 determine whether to perform a virtual speaker search based on the representative virtual speaker in the previous frame and the current frame.
[0126] If encoder 113 determines to perform a virtual speaker search, it executes steps S510 to S530. Optionally, encoder 113 may first execute step S510, which involves encoder 113 obtaining the representative coefficients of the current frame. Encoder 113 then determines whether to perform a virtual speaker search based on the representative coefficients of the current frame and the coefficients of the representative virtual speakers in previous frames. If encoder 113 determines to perform a virtual speaker search, it then executes steps S520 to S530.
[0127] If encoder 113 determines that it will not perform a virtual speaker search, execute S550.
[0128] S550, encoder 113 determines the representative virtual speaker of the previous frame to encode the current frame.
[0129] Encoder 113 multiplexes the virtual speaker of the previous frame and the current frame to generate a virtual speaker signal, encodes the virtual speaker signal to obtain a bit stream, and sends the bit stream to the destination device 120, i.e., executes S450 and S460.
[0130] For specific methods on determining whether to perform a virtual speaker search, please refer to the following. Figure 6 The explanation of S610 to S640.
[0131] S440 and source device 110 generate a virtual speaker signal based on the current frame of the three-dimensional audio signal and the representative virtual speaker of the current frame.
[0132] The source device 110 generates a virtual speaker signal based on the coefficients of the current frame and the coefficients representing the virtual speaker in the current frame. Specific methods for generating the virtual speaker signal can be found in existing technologies and the description of the virtual speaker signal generation unit 350 in the above embodiments.
[0133] S450 and source device 110 encode the virtual speaker signal to obtain a bitstream.
[0134] The source device 110 can perform encoding operations such as transformation or quantization on the virtual speaker signal to generate a bitstream, thereby achieving the purpose of data compression of the three-dimensional audio signal to be encoded. Specific methods for generating the bitstream can be found in existing technologies and the description of the encoding unit 360 in the above embodiments.
[0135] S460, source device 110 sends a bitstream to destination device 120.
[0136] Source device 110 can send the original audio bitstream to destination device 120 after encoding the entire original audio. Alternatively, source device 110 can encode the three-dimensional audio signal in real time, frame by frame, and send the bitstream of each frame after encoding. Specific methods for sending the bitstream can be found in existing technologies and the descriptions of communication interfaces 114 and 124 in the above embodiments.
[0137] S470, the destination device 120 decodes the bitstream sent by the source device 110, reconstructs the three-dimensional audio signal, and obtains the reconstructed three-dimensional audio signal.
[0138] After receiving the bitstream, the target device 120 decodes it to obtain the virtual speaker signal. Then, based on the candidate virtual speaker set and the virtual speaker signal, it reconstructs the three-dimensional audio signal to obtain the reconstructed three-dimensional audio signal. The target device 120 plays back the reconstructed three-dimensional audio signal. Alternatively, the target device 120 transmits the reconstructed three-dimensional audio signal to other playback devices, which then play it, making the "immersive" sound effect in places like cinemas, concert halls, or virtual scenes even more realistic.
[0139] Currently, during the virtual speaker search process, the encoder uses the correlation calculation results between the 3D audio signal to be encoded and the virtual speaker as the selection criterion. Furthermore, if the encoder transmits a virtual speaker for each coefficient, data compression cannot be achieved, and it places a heavy computational burden on the encoder. The encoder can first determine whether it can reuse the representative virtual speaker set from previous frames to encode the current frame. If the encoder reuses the representative virtual speaker set from previous frames, it avoids performing the virtual speaker search process again, effectively reducing the computational complexity of the encoder's virtual speaker search, thus reducing the computational complexity of compressing and encoding the 3D audio signal and alleviating the encoder's computational burden. If the encoder cannot reuse the representative virtual speaker set from previous frames to encode the current frame, it then selects representative coefficients and uses these coefficients to vote on each virtual speaker in the candidate virtual speaker set, selecting the representative virtual speaker for the current frame based on the voting value. This achieves the goal of reducing the computational complexity of compressing and encoding the 3D audio signal and alleviating the encoder's computational burden.
[0140] Next, the process of selecting a virtual speaker will be explained in detail with reference to the accompanying diagram. Figure 6 This is a flowchart illustrating a three-dimensional audio signal encoding method provided in an embodiment of this application. Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Figure 6The method described is... Figure 5 The specific operational procedures included in S540 are described below. For example... Figure 6 As shown, the method includes the following steps.
[0141] S610, encoder 113 acquires the first correlation between the current frame of the three-dimensional audio signal and the representative set of virtual speakers in the previous frame.
[0142] The virtual speakers in the representative virtual speaker set of the previous frame are the virtual speakers used to encode the three-dimensional audio signal in the previous frame. The first relevance is used to determine whether to reuse the representative virtual speaker set of the previous frame when encoding the current frame. Understandably, the higher the first relevance of the representative virtual speaker set of the previous frame, the higher the bias of the representative virtual speaker set of the previous frame, and the more likely the encoder 113 is to select the representative virtual speakers of the previous frame for encoding the current frame.
[0143] In some embodiments, encoder 113 may obtain the correlation between the current frame and the representative virtual speakers of each previous frame in the set of representative virtual speakers of previous frames; sort the correlation between the representative virtual speakers of each previous frame and the current frame, and take the maximum correlation between the representative virtual speakers of each previous frame and the current frame as the first correlation.
[0144] For any representative virtual speaker in the set of representative virtual speakers in the previous frame, encoder 113 can determine the correlation between the current frame and the representative virtual speaker in the previous frame based on the coefficients of the current frame and the coefficients of the representative virtual speaker in the previous frame. Assuming that the set of representative virtual speakers in the previous frame includes a first virtual speaker, encoder 113 can determine the correlation between the current frame and the first virtual speaker based on the coefficients of the current frame and the coefficients of the first virtual speaker.
[0145] The correlation between the current frame and the virtual speaker satisfies the following formula (6).
[0146]
[0147] in, Represents the coefficients of the current frame. The coefficients of the representative virtual loudspeakers in the previous frame are represented by l = 1, 2, ..., Q, where Q represents the number of representative virtual loudspeakers in the previous frame's set of representative virtual loudspeakers.
[0148] The coefficients of the current frame can be determined by the ratio of the coefficient values to the number of coefficients in the current frame. The coefficients of the current frame satisfy formula (7).
[0149] or
[0150] Where j = 1, 2, ..., L, indicating that the value of j ranges from 1 to L, L represents the number of coefficients in the current frame, and x represents the coefficients in the current frame.
[0151] Optionally, encoder 113 may also select a third number of representative coefficients according to the method described in S650 and S660 below, and use the largest representative coefficient among the third number of representative coefficients as the coefficient of the current frame for obtaining the first relevance.
[0152] S620 and encoder 113 determine whether the first correlation meets the reuse condition.
[0153] The multiplexing condition is based on the encoder 113 encoding and multiplexing the current frame of the three-dimensional audio signal before using the virtual speaker.
[0154] If the first relevance satisfies the reuse condition, it means that encoder 113 is more inclined to select the representative virtual speaker of the previous frame to encode the current frame, and encoder 113 executes S630 and S640.
[0155] If the first relevance does not meet the reuse condition, it means that the encoder 113 is more inclined to perform virtual speaker search, and encode the current frame according to the representative virtual speaker of the current frame. The encoder 113 executes S650 to S680.
[0156] Optionally, encoder 113 may select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients, and then use the largest representative coefficient among the third number of representative coefficients as the coefficient of the current frame for obtaining the first correlation. Then encoder 113 obtains the first correlation between the largest representative coefficient among the third number of representative coefficients of the current frame and the representative virtual speaker set of the previous frame. If the first correlation does not meet the reuse condition, execute S660, that is, encoder 113 selects a third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients.
[0157] S630, encoder 113 generates a virtual speaker signal based on the representative set of virtual speakers in the previous frame and the current frame.
[0158] Encoder 113 generates a virtual speaker signal based on the coefficients of the current frame and the coefficients representing the virtual speaker in the previous frame. Specific methods for generating the virtual speaker signal can be found in existing technologies and the description of the virtual speaker signal generation unit 350 in the above embodiments.
[0159] The S640 and encoder 113 encode the virtual speaker signal to obtain a bitstream.
[0160] Encoder 113 can perform encoding operations such as transformation or quantization on the virtual speaker signal to generate a bitstream, which is then sent to the destination device 120. This achieves the purpose of data compression of the three-dimensional audio signal to be encoded. Specific methods for generating the bitstream can be found in existing technologies and the description of encoding unit 360 in the above embodiments.
[0161] This application provides two possible implementation methods for encoder 113 to determine whether the first relevance meets the reuse condition. The two methods are described in detail below.
[0162] In a first possible implementation, encoder 113 compares a first relevance with a relevance threshold. If the first relevance is greater than the relevance threshold, encoder 113 encodes the current frame using the representative virtual speakers of previous frames included in the set of representative virtual speakers for previous frames, generating a bitstream, i.e., executing steps S630 and S640. If the first relevance is less than or equal to the relevance threshold, encoder 113 selects a representative virtual speaker for the current frame from the set of candidate virtual speakers, i.e., executing steps S650 to S680. The multiplexing condition includes: the first relevance is greater than the relevance threshold. The relevance threshold can be pre-configured.
[0163] In the second possible implementation, encoder 113 can also obtain the correlation between the current frame and the virtual speakers contained in the candidate virtual speaker set, and determine whether to reuse the representative virtual speaker set of the previous frame to encode the current frame based on the first correlation and the correlation between the virtual speakers contained in the candidate virtual speaker set.
[0164] Figure 7 This is a flowchart illustrating a method for determining whether to perform a virtual speaker search, provided in an embodiment of this application. Figure 7 The method described is... Figure 6 The specific operation process included in S620 is described below. After the encoder 113 acquires the first correlation between the current frame of the three-dimensional audio signal and the previous frame representing the virtual speaker, i.e., S650, the encoder 113 can also execute S6201 and S6202, or S6203 and S6204, or S6205 to S6208.
[0165] S6201, encoder 113 obtains the second correlation between the current frame and the candidate virtual speaker set.
[0166] The second relevance is used to characterize the priority of using the candidate virtual speaker set when encoding the current frame. Understandably, the larger the second relevance of the candidate virtual speaker set, the higher the priority or the stronger the bias of the candidate virtual speaker set, and the more likely the encoder 113 is to select the candidate virtual speaker set to encode the current frame.
[0167] The set of representative virtual speakers in the previous frame is a proper subset of the set of candidate virtual speakers, meaning that the set of candidate virtual speakers includes the set of representative virtual speakers in the previous frame, and all representative virtual speakers in the previous frame that are included in the set of representative virtual speakers in the previous frame belong to the set of candidate virtual speakers.
[0168] In some embodiments, encoder 113 may obtain the correlation between the current frame and each candidate virtual speaker in the candidate virtual speaker set; sort the correlation between each candidate virtual speaker and the current frame, and take the maximum correlation between each candidate virtual speaker and the current frame as the second correlation.
[0169] For any candidate virtual speaker in the candidate virtual speaker set, encoder 113 can determine the correlation between the current frame and the candidate virtual speaker based on the coefficients of the current frame and the coefficients of the candidate virtual speaker. The correlation between the current frame and the candidate virtual speaker satisfies formula (6). It should be noted that... Q can also represent the coefficient of a candidate virtual speaker, and Q can also represent the number of candidate virtual speakers in the candidate virtual speaker set.
[0170] S6202, encoder 113 determines whether the first relevance is greater than the second relevance.
[0171] If the first relevance is greater than the second relevance, encoder 113 executes S630 and S640.
[0172] If the first relevance is less than or equal to the second relevance, encoder 113 executes S650 to S680.
[0173] The reuse criteria include: the first relevance is greater than the second relevance.
[0174] In another scenario, encoder 113 can also obtain the correlation between the current frame and the virtual speakers contained in a subset of the candidate virtual speaker set, and determine whether to reuse the representative virtual speaker set of the previous frame to encode the current frame based on the first correlation and the correlation between the virtual speakers contained in the subset of the candidate virtual speaker set. Execute S6203 and S6204.
[0175] S6203, encoder 113 obtains the third relevance between the current frame and the first subset of the candidate virtual speaker set.
[0176] The third relevance is used to characterize the priority of using the first subset of the candidate virtual loudspeaker set when encoding the current frame. Understandably, the larger the third relevance of the first subset of the candidate virtual loudspeaker set, the higher the priority or the stronger the bias of the first subset of the candidate virtual loudspeaker set, and the more likely the encoder 113 is to select the first subset of the candidate virtual loudspeaker set to encode the current frame.
[0177] The first subset is a proper subset of the candidate virtual speaker set, meaning that the candidate virtual speaker set includes the first subset, and all candidate virtual speakers contained in the first subset belong to the candidate virtual speaker set.
[0178] In some embodiments, encoder 113 may obtain the correlation between the current frame and each candidate virtual speaker in a first subset of the candidate virtual speaker set; sort the correlation between each candidate virtual speaker and the current frame, and take the maximum correlation between each candidate virtual speaker and the current frame as the third correlation.
[0179] For any candidate virtual speaker in the first subset of the candidate virtual speaker set, encoder 113 can determine the correlation between the current frame and the candidate virtual speaker based on the coefficients of the current frame and the coefficients of the candidate virtual speaker. The correlation between the current frame and the candidate virtual speaker satisfies formula (6). It should be noted that... Q can also represent the coefficient of the candidate virtual loudspeakers in the first subset, and Q can also represent the number of candidate virtual loudspeakers in the first subset of the candidate virtual loudspeaker set.
[0180] S6204, encoder 113 determines whether the first relevance is greater than the third relevance.
[0181] If the first relevance is greater than the third relevance, encoder 113 executes S630 and S640.
[0182] If the first relevance is less than or equal to the third relevance, encoder 113 executes S650 to S680.
[0183] The reuse criteria include: the first relevance is greater than the third relevance.
[0184] In another scenario, encoder 113 can also obtain the correlation between the current frame and the virtual speakers contained in multiple subsets of the candidate virtual speaker set, and perform multiple rounds of judgment based on the first correlation and the correlation between the virtual speakers contained in multiple subsets of the candidate virtual speaker set to determine whether to reuse the representative virtual speaker set of the previous frame to encode the current frame. Execute S6205 to S6208.
[0185] S6205, encoder 113 obtains the fourth relevance between the current frame and the second subset of the candidate virtual speaker set.
[0186] The fourth relevance is used to characterize the priority of using the second subset of the candidate virtual loudspeaker set when encoding the current frame. Understandably, the larger the fourth relevance of the second subset of the candidate virtual loudspeaker set, the higher the priority or the stronger the bias of the second subset of the candidate virtual loudspeaker set, and the more likely the encoder 113 is to select the second subset of the candidate virtual loudspeaker set to encode the current frame.
[0187] The second subset is a proper subset of the candidate virtual speaker set, meaning that the candidate virtual speaker set includes the second subset, and all candidate virtual speakers contained in the second subset belong to the candidate virtual speaker set.
[0188] For details on the method by which encoder 113 obtains the fourth relevance between the current frame and the second subset of the candidate virtual speaker set, please refer to the explanation in S6203 above.
[0189] S6206, encoder 113 determines whether the first relevance is greater than the fourth relevance.
[0190] If the first relevance is greater than the fourth relevance, encoder 113 executes steps S630 and S640. The reuse condition includes: the first relevance is greater than the fourth relevance.
[0191] If the first relevance is less than the fourth relevance, encoder 113 executes S650 to S680.
[0192] If the first relevance equals the fourth relevance, encoder 113 executes steps S6207 to S6208. Understandably, encoder 113 can also continue to select other subsets from the candidate virtual speaker set and determine whether the first relevance of the other subsets satisfies the reuse condition.
[0193] S6207, encoder 113 obtains the fifth relevance between the current frame and the third subset of the candidate virtual speaker set.
[0194] The fifth relevance is used to characterize the priority of using the third subset of the candidate virtual loudspeaker set when encoding the current frame. Understandably, the larger the fifth relevance of the third subset of the candidate virtual loudspeaker set, the higher the priority or the stronger the bias of the third subset of the candidate virtual loudspeaker set, and the more likely the encoder 113 is to select the third subset of the candidate virtual loudspeaker set to encode the current frame.
[0195] The third subset is a proper subset of the candidate virtual speaker set, meaning that the candidate virtual speaker set includes the third subset, and all candidate virtual speakers contained in the third subset belong to the candidate virtual speaker set.
[0196] For the specific method of encoder 113 to obtain the fifth relevance between the current frame and the third subset of the candidate virtual speaker set, please refer to the description in S6203 above.
[0197] The virtual speakers included in the second subset are different from or partially different from those included in the third subset. For example, the second subset includes a first virtual speaker and a second virtual speaker, while the third subset includes a third virtual speaker and a fourth virtual speaker. Or, the second subset includes a first virtual speaker and a second virtual speaker, while the third subset includes a first virtual speaker and a fourth virtual speaker.
[0198] S6208, encoder 113 determines whether the first relevance is greater than the fifth relevance.
[0199] If the first relevance is greater than the fifth relevance, encoder 113 executes steps S630 and S640. The reuse condition includes: the first relevance is greater than the fifth relevance.
[0200] If the first relevance is less than the fifth relevance, encoder 113 executes S650 to S680.
[0201] If the first relevance equals the fifth relevance, encoder 113 executes steps S6207 to S6208. Understandably, encoder 113 can also continue to select other subsets from the candidate virtual speaker set and determine whether the first relevance of the other subsets satisfies the reuse condition.
[0202] In some embodiments, if the first relevance is equal to the fifth relevance, encoder 113 can use the second largest relevance between the representative virtual speakers of the previous frame and the current frame as the first relevance, and obtain the sixth relevance between the current frame and the fourth subset of the candidate virtual speaker set. If the first relevance is greater than the sixth relevance, encoder 113 executes S630 and S640. The reuse condition includes: the first relevance is greater than the sixth relevance. If the first relevance is less than the sixth relevance, encoder 113 executes S650 to S680. If the first relevance is equal to the sixth relevance, encoder 113 can continue to select other subsets from the candidate virtual speaker set and determine whether the first relevance of the other subsets meets the reuse condition.
[0203] It should be noted that the embodiments of this application do not limit the number of rounds in which the representative virtual speaker of the previous frame encodes the current frame. Furthermore, the number of correlation values used in each round of judgment is also not limited.
[0204] Furthermore, the subset selected by encoder 113 from the candidate virtual speaker set can be pre-set. Alternatively, encoder 113 can uniformly sample the candidate virtual speaker set to obtain a subset. For example, encoder 113 can select 1 / 10 of the virtual speakers in the candidate virtual speaker set as a subset of the candidate virtual speaker set. The number of virtual speakers included in the subset of the candidate virtual speaker set selected in each round is not limited. For example, the subset in round i+1 may contain more virtual speakers than the subset in round i. Or, the virtual speakers included in the subset in round i+1 may be K virtual speakers in the vicinity of the virtual speakers in the subset in round i. For example, if the subset in round i contains 64 virtual speakers and K = 32, the subset in round i+1 may contain a portion of the 64 × 32 virtual speakers.
[0205] The method for selecting virtual speakers provided in this application uses the correlation between the representative frequency coefficient of the current frame and the representative virtual speaker of the previous frame to determine whether to perform a virtual speaker search. While ensuring the accuracy of the selection of the correlation of the representative virtual speaker of the current frame, it effectively reduces the complexity of the encoding end.
[0206] Typically, a configuration has 2048 virtual speakers. During the virtual speaker search process, the encoder needs to perform 2048 voting operations for each coefficient in the current frame. The method for determining whether to perform a virtual speaker search provided in this application can skip more than 50% of the virtual speaker search steps, improving the encoder's encoding speed. For example, the encoder pre-calculates a grid of 64 virtual speakers approximately uniformly distributed on a sphere, called the coarse scan grid. For each virtual speaker on the coarse scan grid, a coarse scan is performed to find candidate virtual speakers on the coarse scan grid. Then, a second round of fine scanning is performed on the candidate virtual speakers to obtain the final best-matching virtual speaker. After acceleration using this algorithm, the original 2048 scans are reduced to 64 + 64 = 128, resulting in an algorithm speedup of 2048 ÷ 128 = 16 times.
[0207] The following details the process where, if the first relevance does not meet the reuse condition, encoder 113 continues to search for virtual loudspeakers, obtains the representative virtual loudspeaker for the current frame, and performs encoding based on the representative virtual loudspeaker of the current frame. After S620, encoder 113 can also execute S650 to S680. This application embodiment provides a method for selecting virtual loudspeakers. The encoder uses the representative coefficient of the current frame to vote on each virtual loudspeaker in the candidate virtual loudspeaker set, and selects the representative virtual loudspeaker of the current frame based on the voting value, thereby reducing the computational complexity of virtual loudspeaker search and alleviating the computational burden on the encoder.
[0208] S650, encoder 113 acquires the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients.
[0209] Assuming the 3D audio signal is a HOA signal, encoder 113 can sample the current frame of the HOA signal to obtain L·(N+1). 2 The number of sampling points yields the fourth set of coefficients. N represents the order of the HOA signal. For example, assuming the duration of the current frame of the HOA signal is 20 milliseconds, encoder 113 samples the current frame at a frequency of 48 kHz, obtaining 960·(N+1) coefficients in the time domain. 2 Each sampling point can also be called a time-domain coefficient.
[0210] The frequency domain coefficients of the current frame of the 3D audio signal can be obtained by time-frequency transformation based on the time domain coefficients of the current frame of the 3D audio signal. The method of time-domain to frequency-domain transformation is not limited. For example, a modified discrete cosine transform (MDCT) can be used to obtain 960·(N+1) in the frequency domain. 2 Frequency domain coefficients. Frequency domain coefficients can also be called spectral coefficients or frequency points.
[0211] The frequency domain eigenvalues of the sampling points satisfy p(j) = norm(x(j)), where j = 1, 2, ..., L, L represents the number of sampling times, x represents the frequency domain coefficients of the current frame of the three-dimensional audio signal, such as MDCT coefficients, norm is used to calculate the L2 norm, and x(j) represents the (N+1)-th sampling time. 2 Frequency domain coefficients of each sampling point.
[0212] S660 and encoder 113 select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients.
[0213] Encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least one sub-band. Specifically, dividing the spectral range indicated by the fourth number of coefficients into one sub-band means that the spectral range of this sub-band is equal to the spectral range indicated by the fourth number of coefficients, which is equivalent to encoder 113 not dividing the spectral range indicated by the fourth number of coefficients.
[0214] If encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two frequency band sub-bands, in one case, encoder 113 divides the spectral range indicated by the fourth number of coefficients into at least two sub-bands, each of the at least two sub-bands containing the same number of coefficients.
[0215] In another scenario, encoder 113 unequally divides the spectral range indicated by the fourth number of coefficients, such that at least two sub-bands contain different numbers of coefficients, or each of the at least two sub-bands contains a different number of coefficients. For example, encoder 113 can unequally divide the spectral range indicated by the fourth number of coefficients based on a low-frequency range, a mid-frequency range, and a high-frequency range, such that each spectral range includes at least one sub-band. Each sub-band in the at least one low-frequency range contains the same number of coefficients. Each sub-band in the at least one mid-frequency range contains the same number of coefficients. Each sub-band in the at least one high-frequency range contains the same number of coefficients. Sub-bands in the three spectral ranges (low-frequency, mid-frequency, and high-frequency) may contain different numbers of coefficients.
[0216] Furthermore, encoder 113 selects representative coefficients from at least one sub-band included in the spectral range indicated by the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients, thus obtaining a third number of representative coefficients. The third number is less than the fourth number, and the fourth number of coefficients includes the third number of representative coefficients.
[0217] For example, encoder 113 selects Z representative coefficients from each sub-band according to the descending order of the frequency domain characteristic values of the coefficients in at least one sub-band included in the spectral range indicated by the fourth number of coefficients, and combines the Z representative coefficients in at least one sub-band to obtain a third number of representative coefficients, where Z is a positive integer.
[0218] For example, when at least one sub-band includes at least two sub-bands, encoder 113 determines the weight of each sub-band based on the frequency domain feature values of the first candidate coefficients within each of the at least two sub-bands; and adjusts the frequency domain feature values of the second candidate coefficients within each sub-band according to their respective weights, obtaining the adjusted frequency domain feature values of the second candidate coefficients within each sub-band, where the first and second candidate coefficients are partial coefficients within the sub-band. Encoder 113 determines a third number of representative coefficients based on the adjusted frequency domain feature values of the second candidate coefficients within the at least two sub-bands, and the frequency domain feature values of the coefficients within the at least two sub-bands excluding the second candidate coefficients.
[0219] Since the encoder selects a portion of the coefficients from all coefficients in the current frame as representative coefficients, and uses a smaller number of representative coefficients to replace all coefficients in the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder searching for virtual speakers is effectively reduced, thereby reducing the computational complexity of compressing and encoding three-dimensional audio signals and alleviating the computational burden on the encoder.
[0220] S670, encoder 113 selects a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on a third number of representative coefficients.
[0221] Encoder 113 performs correlation calculations between the third number of representative coefficients of the current frame of the three-dimensional audio signal and the coefficients of each virtual speaker in the candidate virtual speaker set, and selects the second number of representative virtual speakers of the current frame.
[0222] Because the encoder selects a subset of coefficients from all coefficients in the current frame as representative coefficients, and uses a smaller number of representative coefficients to replace all coefficients in the current frame to select representative virtual speakers from the candidate virtual speaker set, the computational complexity of the encoder searching for virtual speakers is effectively reduced. This, in turn, reduces the computational complexity of compressing and encoding 3D audio signals and alleviates the computational burden on the encoder. For example, a frame of N-order HOA signal has 960·(N+1) 2 In this embodiment, the top 10% of the coefficients can be selected to participate in the virtual speaker search. At this time, the coding complexity is reduced by 90% compared to the coding complexity of using all coefficients in the virtual speaker search.
[0223] S680 and encoder 113 encode the current frame according to the second number of representative virtual speakers of the current frame to obtain the bit stream.
[0224] Encoder 113 generates a virtual speaker signal based on the representative virtual speakers of the second number of current frames and the current frame, and encodes the virtual speaker signal to obtain a bitstream. The specific method for generating the bitstream can be found in existing technology and in the descriptions of encoding units 360 and S450 in the above embodiments.
[0225] After generating the bitstream, encoder 113 sends the bitstream to destination device 120 so that destination device 120 can decode the bitstream sent by source device 110, reconstruct the three-dimensional audio signal, and obtain the reconstructed three-dimensional audio signal.
[0226] Since the frequency domain eigenvalues of the coefficients of the current frame characterize the sound field characteristics of the three-dimensional audio signal, the encoder selects representative coefficients of the representative sound field components of the current frame based on the frequency domain eigenvalues of the coefficients of the current frame. The representative virtual loudspeaker of the current frame selected from the candidate virtual loudspeaker set using the representative coefficients can fully characterize the sound field characteristics of the three-dimensional audio signal, thereby further improving the accuracy of the encoder in generating virtual loudspeaker signals when compressing the three-dimensional audio signal to be encoded using the representative virtual loudspeaker of the current frame. This is to improve the compression rate of the three-dimensional audio signal and reduce the bandwidth occupied by the encoder's transmission bitstream.
[0227] Figure 8This is a flowchart illustrating another three-dimensional audio signal encoding method provided in an embodiment of this application. Figure 1 The process of selecting a virtual speaker is illustrated using the encoder 113 in the source device 110 as an example. Figure 8 The method described is... Figure 6 The specific operational procedures included in S670 are explained. For example... Figure 8 As shown, the method includes the following steps.
[0228] S6701, encoder 113 determines the first number of virtual speakers and the first number of voting values based on the third number of representative coefficients of the current frame, the candidate virtual speaker set and the number of voting rounds.
[0229] The voting rounds limit the number of times a virtual speaker is voted on. The voting rounds are integers greater than or equal to 1, and are less than or equal to the number of virtual speakers included in the candidate virtual speaker set, and less than or equal to the number of virtual speaker signals transmitted by the encoder. For example, the candidate virtual speaker set may include a fifth number of virtual speakers, which includes a first number of virtual speakers, where the first number is less than or equal to the fifth number, and the voting rounds are integers greater than or equal to 1, and less than or equal to the fifth number. A virtual speaker signal also refers to the transmission channel representing the virtual speaker in the current frame corresponding to the current frame. Typically, the number of virtual speaker signals is less than or equal to the number of virtual speakers.
[0230] In one possible implementation, the number of voting rounds can be pre-configured or determined based on the encoder's computing power. For example, the number of voting rounds can be determined based on the encoder's encoding rate and / or the encoding application scenario.
[0231] In another possible implementation, the number of voting rounds is determined based on the number of directional sound sources in the current frame. For example, when there are 2 directional sound sources in the sound field, the number of voting rounds is set to 2.
[0232] This application provides three possible implementations for determining a first number of virtual speakers and a first number of voting values. The three methods are described in detail below.
[0233] In the first possible implementation, the number of voting rounds is equal to 1. After the encoder 113 samples multiple representative coefficients, it obtains the voting value of each representative coefficient of the current frame for all virtual speakers in the candidate virtual speaker set, and accumulates the voting values of virtual speakers with the same number to obtain a first number of virtual speakers and a first number of voting values. Understandably, the candidate virtual speaker set includes a first number of virtual speakers. The first number is equal to the number of virtual speakers included in the candidate virtual speaker set. Assuming the candidate virtual speaker set includes a fifth number of virtual speakers, then the first number is equal to the fifth number. The first number of voting values includes the voting values of all virtual speakers in the candidate virtual speaker set. The encoder 113 can use the first number of voting values as the final voting value of the first number of virtual speakers for the current frame, and execute S6702, that is, the encoder 113 selects a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on the first number of voting values.
[0234] In this system, each virtual speaker corresponds one-to-one with a voting value. For example, a first number of virtual speakers includes a first virtual speaker, and a first number of voting values includes the voting value of the first virtual speaker; the first virtual speaker corresponds to the voting value of the first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of using the first virtual speaker when encoding the current frame. Priority can also be described as preference; that is, the voting value of the first virtual speaker is used to characterize the preference for using the first virtual speaker when encoding the current frame. Understandably, the larger the voting value of the first virtual speaker, the higher the priority or preference of the first virtual speaker. Compared to virtual speakers in the candidate virtual speaker set with a smaller voting value than the first virtual speaker, the encoder 113 is more inclined to select the first virtual speaker to encode the current frame.
[0235] In the second possible implementation, the difference from the first possible implementation is that after the encoder 113 obtains the voting values of each representative coefficient for all virtual speakers in the candidate virtual speaker set in the current frame, it selects a portion of the voting values from each representative coefficient for all virtual speakers in the candidate virtual speaker set, and accumulates the voting values of the virtual speakers with the same number corresponding to the portion of the voting values, to obtain a first number of virtual speakers and a first number of voting values. Understandably, the candidate virtual speaker set includes the first number of virtual speakers. The first number is less than or equal to the number of virtual speakers included in the candidate virtual speaker set. The first number of voting values includes the voting values of a portion of the virtual speakers included in the candidate virtual speaker set, or the first number of voting values includes the voting values of all the virtual speakers included in the candidate virtual speaker set.
[0236] In the third possible implementation, the difference from the second possible implementation is that the number of voting rounds is an integer greater than or equal to 2. For each representative coefficient of the current frame, the encoder 113 performs at least two rounds of voting on all virtual speakers in the candidate virtual speaker set, selecting the virtual speaker with the largest vote value in each round. After performing at least two rounds of voting on all virtual speakers for each representative coefficient of the current frame, the vote values of virtual speakers with the same number are accumulated to obtain a first number of virtual speakers and a first number of vote values.
[0237] S6702, encoder 113 selects a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on the first number of voting values.
[0238] The encoder 113 selects a second number of representative virtual speakers for the current frame from the first number of virtual speakers based on a first number of voting values, and the voting values of the second number of representative virtual speakers for the current frame are greater than a preset threshold.
[0239] The encoder 113 can also select representative virtual speakers for the second number of current frames from the first number of virtual speakers based on the first number of voting values. For example, the second number of voting values can be determined from the first number of voting values in descending order, and the virtual speakers corresponding to the second number of voting values from the first number of virtual speakers can be used as representative virtual speakers for the second number of current frames.
[0240] Optionally, if the voting values of virtual speakers with different numbers among the first number of virtual speakers are the same, and the voting values of the different virtual speakers are greater than a preset threshold, then the encoder 113 can use all the virtual speakers with different numbers as the representative virtual speakers of the current frame.
[0241] It should be noted that the second number is less than the first number. The first number of virtual speakers includes the second number of representative virtual speakers for the current frame. The second number can be preset, or it can be determined based on the number of sound sources in the sound field of the current frame. For example, the second number can be directly equal to the number of sound sources in the sound field of the current frame, or it can be obtained by processing the number of sound sources in the sound field of the current frame according to a preset algorithm, and using the processed number as the second number. The preset algorithm can be designed as needed. For example, the preset algorithm can be: second number = number of sound sources in the sound field of the current frame + 1, or second number = number of sound sources in the sound field of the current frame - 1, etc.
[0242] The encoder uses a small number of representative coefficients instead of all coefficients in the current frame to vote on each virtual speaker in the candidate virtual speaker set, and selects the representative virtual speaker for the current frame based on the voting values. Then, the encoder uses the representative virtual speaker of the current frame to compress and encode the 3D audio signal to be encoded. This not only effectively improves the compression ratio of the 3D audio signal but also reduces the computational complexity of the encoder searching for virtual speakers, thereby reducing the computational complexity of the 3D audio signal compression and encoding and alleviating the computational burden on the encoder.
[0243] To increase the continuity of orientation between consecutive frames and overcome the problem of large differences in the results of virtual speakers selected in consecutive frames, encoder 113 adjusts the initial voting value of the virtual speakers in the candidate virtual speaker set for the current frame based on the final voting value of the representative virtual speaker in the previous frame, thus obtaining the final voting value of the virtual speaker for the current frame. Figure 9 The diagram shown is a flowchart illustrating another method for selecting a virtual speaker provided in an embodiment of this application. Wherein, Figure 9 The method described is... Figure 8 The specific operational procedures included in S6702 are explained.
[0244] S6702a, encoder 113 obtains the seventh number of final votes of the current frame corresponding to the seventh number of virtual speakers based on the first number of initial votes of the current frame and the sixth number of final votes of the previous frame.
[0245] The encoder 113 can determine a first number of virtual speakers and a first number of voting values based on the current frame of the three-dimensional audio signal, the candidate virtual speaker set and the number of voting rounds according to the method described in S6701 above, and then use the first number of voting values as the initial voting values of the first number of virtual speakers for the current frame.
[0246] Each virtual speaker corresponds one-to-one with the initial voting value of the current frame; that is, one virtual speaker corresponds to one initial voting value of the current frame. For example, a first number of virtual speakers includes a first virtual speaker, and a first number of initial voting values of the current frame include the initial voting value of the first virtual speaker. The first virtual speaker corresponds to the initial voting value of the first virtual speaker in the current frame. The initial voting value of the first virtual speaker in the current frame is used to characterize the priority of using the first virtual speaker when encoding the current frame.
[0247] The sixth set of virtual loudspeakers in the preceding frames corresponds one-to-one with the final voting values of the sixth set of preceding frames. The sixth set of virtual loudspeakers can be the representative virtual loudspeakers of the preceding frames used by encoder 113 to encode the preceding frames of the three-dimensional audio signal.
[0248] Specifically, encoder 113 updates the initial voting values of the first number of current frames based on the final voting values of the sixth number of prior frames. That is, encoder 113 calculates the sum of the initial voting values of the current frames and the final voting values of the prior frames for the first number of virtual speakers and the virtual speakers with the same number among the sixth number of virtual speakers, and obtains the final voting values of the seventh number of current frames corresponding to the seventh number of virtual speakers. The seventh number of virtual speakers includes the first number of virtual speakers, and the seventh number of virtual speakers includes the sixth number of virtual speakers.
[0249] S6702b, encoder 113 selects a representative virtual speaker for the second number of current frames from the seventh number of virtual speakers based on the final voting value of the seventh number of current frames.
[0250] The encoder 113 selects a representative virtual speaker for a second number of current frames from the seventh number of virtual speakers based on the final voting value of the seventh number of current frames, and the final voting value of the representative virtual speaker for the second number of current frames is greater than a preset threshold.
[0251] The encoder 113 can also select representative virtual speakers for the second number of current frames from the seventh number of virtual speakers based on the final voting values of the seventh number of current frames. For example, the final voting values of the second number of current frames can be determined from the final voting values of the seventh number of current frames in descending order, and the virtual speakers associated with the final voting values of the second number of current frames from the seventh number of virtual speakers can be used as representative virtual speakers for the second number of current frames.
[0252] Optionally, if the voting values of virtual speakers with different numbers among the seventh number of virtual speakers are the same, and the voting values of the virtual speakers with different numbers are greater than a preset threshold, then the encoder 113 can use the virtual speakers with different numbers as the representative virtual speakers of the current frame.
[0253] It should be noted that the second number is less than the seventh number. The seventh number of virtual speakers includes the second number of representative virtual speakers for the current frame. The second number can be preset, or it can be determined based on the number of sound sources in the sound field of the current frame.
[0254] In addition, before encoding the next frame of the current frame, if the encoder 113 determines to reuse the representative virtual speakers of the previous frame to encode the next frame, the encoder 113 can use the second number of representative virtual speakers of the current frame as the second number of representative virtual speakers of the previous frame, and use the second number of representative virtual speakers of the previous frame to encode the next frame of the current frame.
[0255] During the virtual speaker search process, the locations of real sound sources and virtual speakers may not coincide, leading to a one-to-one correspondence between virtual speakers and real sound sources. Furthermore, in complex real-world scenarios, a limited set of virtual speakers may not be sufficient to represent all sound sources in the sound field. In such cases, frequent jumps in the virtual speakers found between frames can significantly impact the listener's auditory experience, resulting in noticeable discontinuities and noise in the decoded and reconstructed 3D audio signal. The virtual speaker selection method provided in this application inherits the representative virtual speaker from previous frames. Specifically, for virtual speakers with the same number, the initial voting value of the current frame is adjusted using the final voting value of the previous frame. This makes the encoder more inclined to select the representative virtual speaker from the previous frame, thereby reducing frequent jumps in virtual speakers between frames, enhancing the continuity of signal orientation between frames, improving the stability of the sound image in the reconstructed 3D audio signal, and ensuring the sound quality of the reconstructed 3D audio signal. Additionally, adjusting parameters ensures that the final voting value of the previous frame is not inherited for too long, preventing the algorithm from being unable to adapt to scenarios involving sound source movement and other changes in the sound field.
[0256] It is understood that, in order to achieve the functions in the above embodiments, the encoder includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0257] The above text combines Figures 1 to 9 This document describes in detail the three-dimensional audio signal encoding method provided according to this embodiment. The following will combine... Figure 10 and Figure 11 This describes the three-dimensional audio signal encoding apparatus and encoder provided according to this embodiment.
[0258] Figure 10 This is a schematic diagram of a possible three-dimensional audio signal encoding device provided in this embodiment. These three-dimensional audio signal encoding devices can be used to implement the function of encoding three-dimensional audio signals in the above method embodiments, and therefore can also achieve the beneficial effects of the above method embodiments. In this embodiment, the three-dimensional audio signal encoding device can be as follows: Figure 1 The encoder 113 shown, or as... Figure 3 The encoder 300 shown can also be a module (such as a chip) applied to terminal devices or servers.
[0259] like Figure 10As shown, the three-dimensional audio signal encoding device 1000 includes a communication module 1010, a coefficient selection module 1020, a virtual speaker selection module 1030, an encoding module 1040, and a storage module 1050. The three-dimensional audio signal encoding device 1000 is used to implement the above-mentioned... Figures 6 to 9 The function of encoder 113 in the method embodiment shown.
[0260] The communication module 1010 is used to acquire the current frame of the three-dimensional audio signal. Optionally, the communication module 1010 can also receive the current frame of the three-dimensional audio signal acquired by other devices; or acquire the current frame of the three-dimensional audio signal from the storage module 1050. The current frame of the three-dimensional audio signal is a HOA signal; the frequency domain characteristic values of the coefficients are determined based on the coefficients of the HOA signal.
[0261] The virtual speaker selection module 1030 is used to obtain the first correlation between the current frame of the three-dimensional audio signal and the representative virtual speaker set of the previous frame. The virtual speakers in the representative virtual speaker set of the previous frame are the virtual speakers used to encode the previous frame of the three-dimensional audio signal. The first correlation is used to determine whether to reuse the representative virtual speaker set of the previous frame when encoding the current frame.
[0262] When the three-dimensional audio signal encoding device 1000 is used to implement Figures 6 to 9 In the method embodiment shown, when the encoder 113 functions, the virtual speaker selection module 1030 is used to implement the related functions of S610 to S630 and S670.
[0263] For example, the virtual speaker selection module 1030 obtains the second correlation between the current frame and the candidate virtual speaker set. The second correlation is used to determine whether to use the candidate virtual speaker set when encoding the current frame. The representative virtual speaker set of the previous frame is a proper subset of the candidate virtual speaker set. The reuse conditions include: the first correlation is greater than the second correlation.
[0264] For example, the virtual speaker selection module 1030 obtains the third correlation between the current frame and the first subset of the candidate virtual speaker set. The third correlation is used to determine whether to use the first subset of the candidate virtual speaker set when encoding the current frame. The first subset is a proper subset of the candidate virtual speaker set. The reuse conditions include: the first correlation is greater than the third correlation.
[0265] For example, the virtual speaker selection module 1030 obtains the fourth relevance between the current frame and the second subset of the candidate virtual speaker set. The fourth relevance is used to determine whether the second subset of the candidate virtual speaker set is used when encoding the current frame. The second subset is a proper subset of the candidate virtual speaker set. If the first relevance is less than or equal to the fourth relevance, the module obtains the fifth relevance between the current frame and the third subset of the candidate virtual speaker set. The fifth relevance is used to determine whether the third subset of the candidate virtual speaker set is used when encoding the current frame. The third subset is a proper subset of the candidate virtual speaker set. The virtual speakers included in the second subset are all different from or partially different from the virtual speakers included in the third subset. The reuse condition includes: the first relevance is greater than the fifth relevance.
[0266] When the three-dimensional audio signal encoding device 1000 is used to implement Figure 6 In the method embodiment shown, when the encoder 113 functions, the virtual speaker selection module 1030 is used to implement the related functions of S670. Specifically, the virtual speaker selection module 1030 is specifically used for: when the virtual speaker selection module selects a second number of representative virtual speakers for the current frame from the candidate virtual speaker set according to a third number of representative coefficients, it is specifically used for: determining a first number of virtual speakers and a first number of voting values according to the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds, where each virtual speaker corresponds one-to-one with a voting value; the first number of virtual speakers includes a first virtual speaker, and the voting value of the first virtual speaker is used to characterize the priority of using the first virtual speaker when encoding the current frame; the candidate virtual speaker set includes a fifth number of virtual speakers, and the fifth number of virtual speakers includes the first number of virtual speakers, where the first number is less than or equal to the fifth number; the number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number; and selecting a second number of representative virtual speakers for the current frame from the first number of virtual speakers according to the first number of voting values, where the second number is less than the first number.
[0267] When the three-dimensional audio signal encoding device 1000 is used to implement Figure 9In the method embodiment shown, when the encoder 113 functions, the virtual speaker selection module 1030 is used to implement the related functions of S6701 and S6702. Specifically, the virtual speaker selection module 1030 obtains the seventh number of final voting values of the current frame corresponding to the seventh number of virtual speakers based on the first number of voting values and the sixth number of final voting values of the previous frames. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The virtual speakers included in the sixth number of virtual speakers are representative virtual speakers of the previous frames used to encode the three-dimensional audio signal. Based on the seventh number of final voting values of the current frames, a second number of representative virtual speakers of the current frames are selected from the seventh number of virtual speakers. The second number is less than the seventh number.
[0268] When the three-dimensional audio signal encoding device 1000 is used to implement Figure 6 In the method embodiment shown, when the encoder 113 functions, the coefficient selection module 1020 is used to implement the related functions of S650 and S660. Specifically, when the coefficient selection module 1020 obtains the third number of representative coefficients of the current frame, it is specifically used to: obtain the fourth number of coefficients of the current frame and the frequency domain feature values of the fourth number of coefficients; and select the third number of representative coefficients from the fourth number of coefficients based on the frequency domain feature values of the fourth number of coefficients, wherein the third number is less than the fourth number.
[0269] The encoding module 1040 is used to encode the current frame according to the representative virtual speaker set of the previous frame to obtain the bit stream if the first relevance meets the reuse condition.
[0270] When the three-dimensional audio signal encoding device 1000 is used to implement Figures 6 to 9 In the method embodiment shown, when encoder 113 functions, encoding module 1040 is used to implement the related functions of S630. Specifically, encoding module 1040 is used to generate a virtual speaker signal based on the representative virtual speaker set of the previous frame and the current frame; and to encode the virtual speaker signal to obtain a bitstream.
[0271] The storage module 1050 is used to store coefficients related to the three-dimensional audio signal, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and selected coefficients and virtual speakers, so that the encoding module 1040 can encode the current frame to obtain a bitstream and transmit the bitstream to the decoder.
[0272] It should be understood that the three-dimensional audio signal encoding device 1000 of this application embodiment can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can also be implemented by software. Figures 6 to 9 In the three-dimensional audio signal encoding method shown, the three-dimensional audio signal encoding device 1000 and its various modules can also be software modules.
[0273] For a more detailed description of the aforementioned communication module 1010, coefficient selection module 1020, virtual speaker selection module 1030, encoding module 1040, and storage module 1050, please refer to [the relevant documentation / reference]. Figures 6 to 9 The relevant descriptions in the method embodiments shown are directly obtained and will not be repeated here.
[0274] Figure 11 This is a schematic diagram of the structure of an encoder 1100 provided in this embodiment. Figure 11 As shown, the encoder 1100 includes a processor 1110, a bus 1120, a memory 1130, and a communication interface 1140.
[0275] It should be understood that in this embodiment, the processor 1110 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0276] The processor may also be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, or one or more integrated circuits used to control the execution of the program in this application.
[0277] The communication interface 1140 is used to enable communication between the encoder 1100 and external devices or components. In this embodiment, the communication interface 1140 is used to receive three-dimensional audio signals.
[0278] Bus 1120 may include a pathway for transmitting information between the aforementioned components (such as processor 1110 and memory 1130). In addition to a data bus, bus 1120 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 1120 in the figure.
[0279] As an example, encoder 1100 may include multiple processors. A processor may be a multi-CPU processor. Here, "processor" can refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions). Processor 1110 may access coefficients associated with the three-dimensional audio signal stored in memory 1130, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and selected coefficients and virtual speakers, etc.
[0280] It is worth noting that, Figure 11 Taking encoder 1100 as an example, which includes one processor 1110 and one memory 1130, the processor 1110 and the memory 1130 are used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to business needs.
[0281] The memory 1130 may correspond to the storage medium used in the above method embodiment for storing coefficients related to the three-dimensional audio signal, a set of candidate virtual speakers, a set of representative virtual speakers in the previous frame, and information such as selected coefficients and virtual speakers, for example, a disk, such as a mechanical hard disk or a solid-state drive.
[0282] The encoder 1100 described above can be a general-purpose device or a special-purpose device. For example, the encoder 1100 can be an x86 or ARM-based server, or other special-purpose servers, such as a policy control and charging (PCC) server. This application does not limit the type of encoder 1100.
[0283] It should be understood that the encoder 1100 according to this embodiment can correspond to the three-dimensional audio signal encoding device 1100 in this embodiment, and can correspond to the execution according to Figures 6 to 9 The corresponding subject in any of the methods, and the above and other operations and / or functions of each module in the three-dimensional audio signal encoding device 1100 are respectively for implementing Figures 6 to 9 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.
[0284] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a network device or terminal device. Of course, the processor and storage medium can also exist as discrete components in the network device or terminal device.
[0285] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0286] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A three-dimensional audio signal encoding method, characterized in that, include: Obtain the correlation between the current frame of the 3D audio signal and the representative virtual speakers of each previous frame in the set of representative virtual speakers in the previous frame; The maximum correlation between the representative virtual speakers of each previous frame and the current frame is taken as the first correlation. The virtual speakers in the set of representative virtual speakers of the previous frames are the virtual speakers used to encode the previous frames of the three-dimensional audio signal. The first correlation is used to determine whether the set of representative virtual speakers of the previous frames is reused when encoding the current frame. If the first relevance satisfies the reuse condition, the current frame is encoded according to the representative virtual speaker set of the previous frame to obtain the bitstream.
2. The method according to claim 1, characterized in that, After acquiring the first correlation between the current frame of the three-dimensional audio signal and the previous frame representing the virtual speaker, the method further includes: Obtain the second correlation between the current frame and the candidate virtual speaker set. The second correlation is used to determine whether the candidate virtual speaker set is used when encoding the current frame. The representative virtual speaker set of the previous frame is a proper subset of the candidate virtual speaker set. The reuse conditions include: the first relevance is greater than the second relevance.
3. The method according to claim 2, characterized in that, The step of obtaining the second relevance between the current frame and the candidate virtual speaker set includes: Obtain the correlation between the current frame and each candidate virtual speaker in the candidate virtual speaker set; The maximum correlation between each candidate virtual speaker and the current frame is taken as the second correlation.
4. The method according to claim 1, characterized in that, After acquiring the first correlation between the current frame of the three-dimensional audio signal and the previous frame representing the virtual speaker, the method further includes: Obtain the third correlation between the current frame and the first subset of the candidate virtual speaker set. The third correlation is used to determine whether the first subset of the candidate virtual speaker set is used when encoding the current frame. The first subset is a proper subset of the candidate virtual speaker set. The reuse condition includes: the first relevance is greater than the third relevance.
5. The method according to claim 1, characterized in that, After acquiring the first correlation between the current frame of the three-dimensional audio signal and the previous frame representing the virtual speaker, the method further includes: Obtain the fourth correlation between the current frame and the second subset of the candidate virtual speaker set. The fourth correlation is used to determine whether the second subset of the candidate virtual speaker set is used when encoding the current frame. The second subset is a proper subset of the candidate virtual speaker set. If the first relevance is equal to the fourth relevance, a fifth relevance is obtained between the current frame and the third subset of the candidate virtual speaker set. The fifth relevance is used to determine whether the third subset of the candidate virtual speaker set is used when encoding the current frame. The third subset is a proper subset of the candidate virtual speaker set, and the virtual speakers included in the second subset are all different or partially different from the virtual speakers included in the third subset. The reuse condition includes: the first relevance is greater than the fifth relevance. If the first correlation is less than the fourth correlation, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients; Based on the frequency domain characteristic values of the fourth number of coefficients, a third number of representative coefficients are selected from the fourth number of coefficients, wherein the third number is less than the fourth number; Based on the third number of representative coefficients, select a second number of representative virtual speakers for the current frame from the candidate virtual speaker set; The current frame is encoded using the representative virtual speakers of the second number of current frames to obtain a bitstream.
6. The method according to any one of claims 1-5, characterized in that, The set of representative virtual speakers in the preceding frames includes a first virtual speaker, and the first correlation between the current frame of the three-dimensional audio signal and the set of representative virtual speakers in the preceding frames includes: The correlation between the current frame and the first virtual speaker is determined based on the coefficients of the current frame and the coefficients of the first virtual speaker.
7. The method according to any one of claims 1-5, characterized in that, If the first relevance does not meet the reuse condition, the method further includes: Obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients; Based on the frequency domain characteristic values of the fourth number of coefficients, a third number of representative coefficients are selected from the fourth number of coefficients, wherein the third number is less than the fourth number; Based on the third number of representative coefficients, select a second number of representative virtual speakers for the current frame from the candidate virtual speaker set; The current frame is encoded using the representative virtual speakers of the second number of current frames to obtain a bitstream.
8. The method according to claim 7, characterized in that, The step of selecting a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on the third number of representative coefficients includes: Based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds, a first number of virtual speakers and a first number of voting values are determined. The virtual speakers and the voting values correspond one-to-one. The first number of virtual speakers includes a first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of the first virtual speaker. The candidate virtual speaker set includes a fifth number of virtual speakers. The fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number. Based on the first number of voting values, a second number of representative virtual speakers for the current frame are selected from the first number of virtual speakers, where the second number is less than the first number.
9. The method according to claim 8, characterized in that, The step of selecting representative virtual speakers for the second number of current frames from the first number of virtual speakers based on the first number of voting values includes: Based on the first number of voting values and the sixth number of previous frame final voting values, a seventh number of virtual speakers are obtained, corresponding to the seventh number of current frame final voting values. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The sixth number of virtual speakers included in the representative virtual speaker set of the previous frame correspond one-to-one with the sixth number of previous frame final voting values. The sixth number of virtual speakers are used to encode the previous frames of the three-dimensional audio signal. Based on the final voting values of the seventh number of current frames, a representative virtual speaker for the second number of current frames is selected from the seventh number of virtual speakers, where the second number is less than the seventh number.
10. The method according to any one of claims 1-5, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
11. A three-dimensional audio signal encoding device, characterized in that, include: The virtual speaker selection module is used to obtain the correlation between the current frame of the three-dimensional audio signal and the representative virtual speakers of each previous frame in the set of representative virtual speakers of the previous frame; The maximum correlation between the representative virtual speakers of each previous frame and the current frame is taken as the first correlation. The virtual speakers in the set of representative virtual speakers of the previous frames are the virtual speakers used to encode the previous frames of the three-dimensional audio signal. The first correlation is used to determine whether the set of representative virtual speakers of the previous frames is reused when encoding the current frame. An encoding module is used to encode the current frame according to the set of representative virtual speakers of the previous frame to obtain a bitstream if the first relevance satisfies the reuse condition.
12. The apparatus according to claim 11, characterized in that, The virtual speaker selection module is also used for: Obtain the second correlation between the current frame and the candidate virtual speaker set. The second correlation is used to determine whether the candidate virtual speaker set is used when encoding the current frame. The representative virtual speaker set of the previous frame is a proper subset of the candidate virtual speaker set. The reuse conditions include: the first relevance is greater than the second relevance.
13. The apparatus according to claim 12, characterized in that, When the virtual speaker selection module obtains the second relevance between the current frame and the candidate virtual speaker set, it is specifically used for: Obtain the correlation between the current frame and each candidate virtual speaker in the candidate virtual speaker set; The maximum correlation between each candidate virtual speaker and the current frame is taken as the second correlation.
14. The apparatus according to claim 11, characterized in that, The virtual speaker selection module is also used for: Obtain the third correlation between the current frame and the first subset of the candidate virtual speaker set. The third correlation is used to determine whether the first subset of the candidate virtual speaker set is used when encoding the current frame. The first subset is a proper subset of the candidate virtual speaker set. The reuse condition includes: the first relevance is greater than the third relevance.
15. The apparatus according to claim 11, characterized in that, The virtual speaker selection module is also used for: Obtain the fourth correlation between the current frame and the second subset of the candidate virtual speaker set. The fourth correlation is used to determine whether the second subset of the candidate virtual speaker set is used when encoding the current frame. The second subset is a proper subset of the candidate virtual speaker set. If the first relevance is equal to the fourth relevance, a fifth relevance is obtained between the current frame and the third subset of the candidate virtual speaker set. The fifth relevance is used to determine whether the third subset of the candidate virtual speaker set is used when encoding the current frame. The third subset is a proper subset of the candidate virtual speaker set, and the virtual speakers included in the second subset are all different or partially different from the virtual speakers included in the third subset. The reuse condition includes: the first relevance is greater than the fifth relevance. If the first correlation is less than the fourth correlation, obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients; Based on the frequency domain characteristic values of the fourth number of coefficients, a third number of representative coefficients are selected from the fourth number of coefficients, wherein the third number is less than the fourth number; Based on the third number of representative coefficients, select a second number of representative virtual speakers for the current frame from the candidate virtual speaker set; The current frame is encoded using the representative virtual speakers of the second number of current frames to obtain a bitstream.
16. The apparatus according to any one of claims 11-15, characterized in that, The representative virtual speaker set of the preceding frame includes a first virtual speaker. When the virtual speaker selection module obtains the first correlation between the current frame of the three-dimensional audio signal and the representative virtual speaker set of the preceding frame, it is specifically used for: The correlation between the current frame and the first virtual speaker is determined based on the coefficients of the current frame and the coefficients of the first virtual speaker.
17. The apparatus according to any one of claims 11-15, characterized in that, If the first relevance does not meet the reuse condition, the device further includes a coefficient selection module; The coefficient selection module is used to obtain the fourth number of coefficients of the current frame of the three-dimensional audio signal, and the frequency domain feature values of the fourth number of coefficients. The coefficient selection module is further configured to select a third number of representative coefficients from the fourth number of coefficients based on the frequency domain characteristic values of the fourth number of coefficients, wherein the third number is less than the fourth number; The virtual speaker selection module is further configured to select a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on the third number of representative coefficients. The encoding module is further configured to encode the current frame according to the representative virtual speakers of the second number of current frames to obtain a bitstream.
18. The apparatus according to claim 17, characterized in that, When the virtual speaker selection module selects a second number of representative virtual speakers for the current frame from the candidate virtual speaker set based on the third number of representative coefficients, it is specifically used for: Based on the third number of representative coefficients of the current frame, the candidate virtual speaker set, and the number of voting rounds, a first number of virtual speakers and a first number of voting values are determined. The virtual speakers and the voting values correspond one-to-one. The first number of virtual speakers includes a first virtual speaker. The voting value of the first virtual speaker is used to characterize the priority of the first virtual speaker. The candidate virtual speaker set includes a fifth number of virtual speakers. The fifth number of virtual speakers includes the first number of virtual speakers. The first number is less than or equal to the fifth number. The number of voting rounds is an integer greater than or equal to 1, and the number of voting rounds is less than or equal to the fifth number. Based on the first number of voting values, a second number of representative virtual speakers for the current frame are selected from the first number of virtual speakers, where the second number is less than the first number.
19. The apparatus according to claim 18, characterized in that, When the virtual speaker selection module selects the representative virtual speakers of the second number of current frames from the first number of virtual speakers based on the first number of voting values, it is specifically used for: Based on the first number of voting values and the sixth number of previous frame final voting values, a seventh number of virtual speakers are obtained, corresponding to the seventh number of current frame final voting values. The seventh number of virtual speakers includes the first number of virtual speakers and the sixth number of virtual speakers. The sixth number of virtual speakers included in the representative virtual speaker set of the previous frame correspond one-to-one with the sixth number of previous frame final voting values. The sixth number of virtual speakers are used to encode the previous frames of the three-dimensional audio signal. Based on the final voting values of the seventh number of current frames, a representative virtual speaker for the second number of current frames is selected from the seventh number of virtual speakers, where the second number is less than the seventh number.
20. The apparatus according to any one of claims 11-15, characterized in that, The current frame of the three-dimensional audio signal is a high-order stereo reverberation (HOA) signal; the frequency domain characteristic values of the coefficients of the current frame are determined based on the coefficients of the HOA signal.
21. An encoder, characterized in that, The encoder includes at least one processor and a memory, wherein the memory is used to store a computer program such that when the computer program is executed by the at least one processor, it implements the three-dimensional audio signal encoding method as described in any one of claims 1-10.
22. A codec system, characterized in that, The encoding / decoding system includes an encoder as described in claim 21 and a decoder, wherein the encoder is used to perform the operation steps of the method according to any one of claims 1-10, and the decoder is used to decode the bitstream generated by the encoder.
23. A computer-readable storage medium, characterized in that, Includes computer software instructions; when the computer software instructions are executed in the encoder, the encoder causes the encoder to perform the three-dimensional audio signal encoding method as described in any one of claims 1-10.
24. A computer-readable storage medium, characterized in that, The bitstream obtained by the three-dimensional audio signal encoding method as described in any one of claims 1-10.
Citation Information
Patent Citations
3d immersive spatial audio systems and methods
CN106537942A
Spatial sound reproduction using multichannel loudspeaker systems
CN111869241A