Virtual speaker determination method and related device

By determining target virtual speakers through frame-based attribute information and interpolation, the method addresses spatial jumps in HOA signals, improving the immersive quality of 3D audio decoding.

JP2026502251APending Publication Date: 2026-01-21HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025538517
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-11-22
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

The positions of virtual speakers in adjacent frames of Higher Order Ambisonics (HOA) signals can differ, causing a spatial jump during decoding, which disrupts the immersive auditory experience in 3D audio technology.

Method used

A method to determine target virtual speakers by considering attribute information of both current and previous frames, using interpolation to ensure smooth transitions between frames, thereby maintaining spatial continuity.

Benefits of technology

The method ensures seamless spatial transitions between frames, enhancing the immersive experience in 3D audio by minimizing spatial jumps during decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502251000001_ABST
    Figure 2026502251000001_ABST
Patent Text Reader

Abstract

This application relates to the field of 3D audio encoding and decoding technology and discloses a virtual speaker determination method and related apparatus. The method includes steps of acquiring attribute information of N first virtual speakers, acquiring attribute information of N second virtual speakers, and determining M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers. The target virtual speakers are configured to process a target group of HOA signals, the second virtual speakers are configured to process a reference group of HOA signals, and the first virtual speakers are virtual speakers that match the target group of the HOA signals. The target virtual speaker is determined based on the attribute information of the second virtual speaker and the attribute information of the first virtual speaker, so that it can be ensured that the attribute information of the target virtual speaker is not significantly different from the attribute information of the second virtual speaker, thereby solving the problem that two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to Chinese Patent Application No. 202211717964.9, entitled "Virtual Speaker Determination Method and Related Apparatus," filed on December 29, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of three-dimensional audio encoding and decoding technology, and in particular to a virtual speaker determination method and related apparatus. [Background technology]

[0003] 3D audio technology is an audio technology for acquiring, processing, transmitting, rendering, and reproducing real-world sound events and 3D sound field information through computers, signal processing, etc. 3D audio technology creates a strong sense of space, envelopment, and immersion to provide people with an auditory experience that "makes them feel as if they are actually there." Currently, the mainstream 3D audio technology is higher order ambisonics (HOA) audio technology. HOA technology is attracting attention due to its speaker layout independence in the recording, encoding, and playback phases, and its ability to play HOA format data rotatably, providing a high degree of freedom in the playback of HOA signals.

[0004] In the process of encoding and decoding the HOA signal, based on the HOA coefficients of the current frame of the HOA signal, a virtual speaker that matches the HOA coefficients of the current frame of the HOA signal is selected from the virtual speaker set of the three-dimensional sound field, and the matched virtual speaker is used as the target virtual speaker. In this way, the current frame of the HOA signal is converted into a virtual speaker signal using the target virtual speaker to reduce the number of channels of the HOA signal, thereby improving the efficiency of encoding and decoding the HOA signal.

[0005] However, the positions of the target virtual speakers corresponding to two adjacent frames of the HOA signal in the 3D sound field may be different, i.e., the elevation and azimuth angles of the virtual speakers corresponding to two adjacent frames of the HOA signal may differ. As a result, the two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump. Therefore, how to adjust the virtual speakers corresponding to two adjacent frames of the HOA signal is currently an urgent problem that needs to be solved. Summary of the Invention

[0006] This application provides a virtual speaker determination method and related device to solve the problem in the related art that two adjacent frames of an HOA signal obtained by decoding sound like a spatial jump. The technical solution is as follows: [Means for solving the problem]

[0007] According to a first aspect, there is provided a virtual speaker determination method. The virtual speaker determination method may be applied to an encoder-side device or a decoder-side device. The method includes: a step of acquiring attribute information of N first virtual speakers, the N first virtual speakers being virtual speakers in a virtual speaker set that match HOA coefficients of a target group of the HOA signal, the target group of the HOA signal including at least one frame of the HOA signal, where N is an integer greater than or equal to 1; a step of acquiring attribute information of N second virtual speakers, the N second virtual speakers being virtual speakers in the virtual speaker set that are configured to process a reference group of the HOA signal, where the reference group of the HOA signal is at least one group of HOA signals preceding the target group of the HOA signal; and a step of determining M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, the M target virtual speakers being configured to process the target group of the HOA signal, where M is an integer greater than 1 and M is greater than N; Includes:

[0008] The target virtual speaker is configured to process a target group of HOA signals, and the second virtual speaker is configured to process a reference group of HOA signals. The first virtual speaker is a virtual speaker with which the target group of HOA signals matches. After the first virtual speaker is determined, the target virtual speaker is determined based on the attribute information of the second virtual speaker and the attribute information of the first virtual speaker to ensure that the attribute information of the target virtual speaker is not significantly different from the attribute information of the second virtual speaker. This solves the problem that two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump.

[0009] For example, at least one frame of the HOA signal that currently needs to be encoded and decoded is used as the target group of the HOA signal, where the target group of the HOA signal includes one frame of the HOA signal, or the target group of the HOA signal includes P frames of the HOA signal, where P is an integer greater than 1.

[0010] The virtual speaker set includes a plurality of virtual speakers, each of which has a corresponding HOA coefficient. N first virtual speakers that match the HOA coefficients of the at least one frame of the HOA signal are selected from the virtual speaker set based on the HOA coefficients of the at least one frame of the HOA signal. Then, attribute information of the N first virtual speakers is obtained based on identifiers of the N first virtual speakers from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers.

[0011] For example, the reference group of HOA signals is one group of HOA signals before the target group of HOA signals. Alternatively, the reference group of HOA signals is multiple groups of HOA signals before the target group of HOA signals. Depending on the case, the method for obtaining the attribute information of the N second virtual speakers differs. The following describes the two cases separately.

[0012] In the first case, the reference group of HOA signals is one group of HOA signals preceding the target group of HOA signals, in which case the N virtual speakers configured to process this group of HOA signals are directly used as the N second virtual speakers, and the attribute information of the N second virtual speakers is obtained based on the identifiers of the N second virtual speakers from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers.

[0013] In the second case, the reference group of HOA signals is multiple groups of HOA signals that precede the target group of HOA signals.

[0014] Each group of HOA signals in the multiple groups of HOA signals corresponds to N virtual speakers, and the N virtual speakers corresponding to each group of HOA signals correspond one-to-one to the N virtual speakers corresponding to another group of HOA signals. In this case, virtual speakers that correspond to the multiple groups of HOA signals are used as groups of virtual speakers to obtain N groups of virtual speakers. Any group of virtual speakers in the N groups of virtual speakers includes a virtual speaker corresponding to each group of HOA signals in the multiple groups of HOA signals. Then, based on the identifiers of the multiple virtual speakers, attribute information of the multiple virtual speakers included in any group of virtual speakers in the N groups of virtual speakers is obtained from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers to obtain one group of attribute information. In this way, for each group of virtual speakers in the N groups of virtual speakers, one group of attribute information can be determined according to the above steps to obtain N groups of attribute information. Finally, averaging is performed on the same group of attribute information in the N groups of attribute information to obtain N pieces of attribute information, and the N pieces of attribute information are determined as attribute information of the N second virtual speakers to obtain attribute information of the N second virtual speakers.

[0015] When the attribute information of the virtual speakers includes the elevation angle and the azimuth angle, the M target virtual speakers are determined according to the following steps (1) to (3).

[0016] (1) Based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers, distances between the corresponding first virtual speakers and the corresponding second virtual speakers are determined to obtain N distances.

[0017] (2) Determine M groups of elevation and azimuth angles based on the N distances.

[0018] Based on the above description, the target group of the HOA signal may include one frame of the HOA signal, or the target group of the HOA signal may include P frames of the HOA signal. Depending on the case, the method for determining the M groups of elevation angles and azimuth angles based on the N distances may be different. The following describes the two cases separately.

[0019] In the first case, the target group of the HOA signal includes one frame of the HOA signal, and the current frame of the HOA signal includes H subframes, where H is an integer greater than 1. To obtain H groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the H subframes included in the current frame of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*H=M groups of elevation angles and azimuth angles are obtained.

[0020] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the H subframes are determined by the following operation, i.e., when the target distance is greater than a first distance threshold, determining the elevation angle and azimuth angle corresponding to each of the H subframes based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0021] For example, an implementation process for determining elevation angles and azimuth angles corresponding to H subframes based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a first subframe among the H subframes; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a last subframe among the H subframes; and determining, for an i-th subframe among the H subframes, the elevation angle and azimuth angle corresponding to the i-th subframe by an interpolation process based on the elevation angle and azimuth angle corresponding to the (i-1)-th subframe among the H subframes and the elevation angle and azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1.

[0022] That is, the elevation angle and azimuth angle corresponding to the first subframe of the H subframes are the elevation angle and azimuth angle of the target second virtual speaker of the reference group of the HOA signal, and the elevation angle and azimuth angle corresponding to the last subframe of the H subframes are the elevation angle and azimuth angle of the target first virtual speaker of the current frame of the HOA signal. The elevation angle and azimuth angle corresponding to any subframe other than the first subframe and the last subframe of the H subframes needs to be obtained by interpolation based on the elevation angle and azimuth angle of the previous subframe closest to that subframe and the elevation angle and azimuth angle corresponding to the last subframe. In this way, when the target group of the HOA signal includes one frame of the HOA signal, interpolation is performed among the H subframes included in the current frame of the HOA signal to achieve a smooth transition between the first and second virtual speakers corresponding to the target distance.

[0023] For the i-th subframe among the H subframes, the starting point of the interpolation process for the i-th subframe is the elevation angle and azimuth angle corresponding to the (i-1)-th subframe, and the ending point of the interpolation process is the elevation angle and azimuth angle corresponding to the last subframe. In other words, for any subframe among the H subframes other than the first and last subframes, the starting point of the subframe interpolation process is constantly updated in real time. In this way, the elevation angle and azimuth angle corresponding to each of the H subframes can be more accurately determined.

[0024] It should be noted that in practical applications, the target distance may be less than or equal to the first distance threshold. In other words, the position of the target first virtual speaker in the current frame of the HOA signal is not significantly different from the position of the target second virtual speaker in the reference group of the HOA signal. Optionally, the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to the H subframes, respectively. In other words, the elevation angle corresponding to each frame in the H subframes is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each subframe is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0025] Optionally, the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to the first K subframes of the H subframes, and the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to the remaining subframes of the H subframes, where K is an integer greater than or equal to 1 and K is less than H.

[0026] The first distance threshold is preset, for example, 0.5, and may be adjusted based on different requirements.

[0027] In the second case, the target group of the HOA signal includes P frames of the HOA signal. To obtain P groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the P frames of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*P=M groups of elevation angles and azimuth angles are obtained.

[0028] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal are determined by the following operation, i.e., when the target distance is greater than the second distance threshold, determining the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0029] For example, an implementation process for determining elevation angles and azimuth angles corresponding to P frames of an HOA signal based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the first frame of the HOA signal among the P frames of the HOA signal; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the last frame of the HOA signal among the P frames of the HOA signal; and determining, for the jth frame of the HOA signal among the P frames of the HOA signal, the elevation angle and azimuth angle corresponding to the jth frame of the HOA signal by an interpolation process based on the elevation angle and azimuth angle corresponding to the (j-1)th frame of the HOA signal among the P frames of the HOA signal and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal, where j is greater than 0 and less than P-1.

[0030] That is, the elevation angle and azimuth angle corresponding to the first frame of the HOA signal among the P frames of the HOA signal are the elevation angle and azimuth angle of the target second virtual speaker of the reference group of the HOA signal, and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal among the P frames of the HOA signal are the elevation angle and azimuth angle of the first virtual speaker in the target group of the HOA signal. The elevation angle and azimuth angle corresponding to any frame of the HOA signal other than the first frame of the HOA signal and the last frame of the HOA signal among the P frames of the HOA signal must be obtained by interpolation based on the elevation angle and azimuth angle of the previous frame of the HOA signal that is closest to the current frame of the HOA signal and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal. In this way, when the target group of the HOA signal includes P frames of the HOA signal, interpolation is performed between the P frames of the HOA signal to achieve a smooth transition between the first virtual speaker and the second virtual speaker corresponding to the target distance.

[0031] The starting point of the interpolation process for the jth frame of the HOA signal among the P frames of the HOA signal is the elevation angle and azimuth angle corresponding to the (j-1)th frame of the HOA signal, and the ending point of the interpolation process is the elevation angle and azimuth angle corresponding to the last frame of the HOA signal. In other words, for any frame of the HOA signal among the P frames of the HOA signal other than the first frame and the last frame of the HOA signal, the starting point of the interpolation process for the HOA signal frame is always updated in real time. In this way, the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal can be more accurately determined.

[0032] It should be noted that in practical applications, the target distance may be less than or equal to the second distance threshold. In other words, the position of the first virtual speaker of interest in the target group of the HOA signal is not significantly different from the position of the second virtual speaker of interest in the reference group of the HOA signal. Optionally, the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to P frames of the HOA signal, respectively. In other words, the elevation angle corresponding to each frame of the HOA signal in the P frames of the HOA signal is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each frame of the HOA signal is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0033] Optionally, the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to the first L frames of the P frames of the HOA signal, and the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance are determined as the elevation angle and azimuth angle corresponding to the remaining frames of the HOA signal of the P frames of the HOA signal, where L is an integer greater than or equal to 1 and L is less than P.

[0034] The second distance threshold is preset, and may be equal to or unequal to the first distance threshold, and may be adjusted based on different requirements.

[0035] (3) Virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles are determined as M target virtual speakers.

[0036] After the M groups of elevation angles and azimuth angles are determined based on the N distances according to the above step (2), the virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles are determined as M target virtual speakers, so that the M target virtual speakers will process the target groups of HOA signals thereafter.

[0037] Based on the above description, in practical applications, the attribute information of the virtual speaker may further include other content, such as the HOA coefficient of the virtual speaker. If the attribute information of the virtual speaker includes the HOA coefficient, the HOA coefficient of the virtual speaker needs to be first converted into the elevation angle and azimuth angle of the virtual speaker according to the related algorithm, and then the M target virtual speakers are determined according to the above steps (1) to (3).

[0038] Optionally, in the case of the encoder-side device, after determining the M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, the encoder-side device further needs to encode the attribute information of the M target virtual speakers into the bitstream. In this way, after receiving the bitstream, the decoder-side device can analyze the bitstream to obtain the attribute information of the M target virtual speakers and reconstruct the target group of the HOA signal based on the attribute information of the M target virtual speakers. Alternatively, the encoder-side device directly encodes indexes of the determination method of the M target virtual speakers into the bitstream, so that the decoder-side device analyzes the bitstream to obtain the indexes of the determination method of the M target virtual speakers, and then determines the M target virtual speakers in real time based on the indexes.

[0039] According to a second aspect, there is provided a virtual speaker determination device. The virtual speaker determination device has a function of performing the operations of the virtual speaker determination method in the first aspect. The virtual speaker determination device includes at least one module. The at least one module is configured to perform the virtual speaker determination method provided in the first aspect.

[0040] According to a third aspect, there is provided a computer device including a processor and a memory configured to store a computer program for executing the virtual speaker determination method provided in the first aspect, and the processor configured to execute the computer program stored in the memory to implement the virtual speaker determination method according to the first aspect.

[0041] Optionally, the computing device may further include a communication bus configured to establish a connection between the processor and the memory.

[0042] According to a fourth aspect, there is provided a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the steps of the virtual speaker determination method according to the first aspect.

[0043] According to a fifth aspect, there is provided a computer program product including instructions that, when executed on a computer, enable the computer to perform the steps of the virtual speaker determination method according to the first aspect. In other words, there is provided a computer program that, when executed on a computer, enables the computer to perform the steps of the virtual speaker determination method according to the first aspect.

[0044] The technical effects obtained by the second to fifth aspects are the same as those obtained by the corresponding technical means in the first aspect, and the details will not be repeated here. [Brief explanation of the drawings]

[0045] [Figure 1] 1 is a diagram of an implementation environment according to one embodiment of the present application; [Figure 2] 1 is a diagram of an implementation environment of a terminal scenario according to an embodiment of the present application; [Figure 3]FIG. 1 is a diagram of an implementation environment for a radio and television scenario, according to one embodiment of the present application. [Figure 4] FIG. 1 is a diagram of an implementation environment for a virtual reality streaming scenario, according to one embodiment of the present application. [Figure 5] 1 is a flowchart of a virtual speaker determination method according to an embodiment of the present application; [Figure 6] 10 is a flowchart of another virtual speaker determination method according to an embodiment of the present application. [Figure 7] 1 is a diagram of the structure of a virtual speaker determination device according to an embodiment of the present application; [Figure 8] FIG. 1 is a diagram of the structure of a computing device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0046] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following further describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0047] Before the virtual speaker determination method provided in the embodiment of the present application is described in detail, the implementation environment in the embodiment of the present application will be described first.

[0048] In the process of encoding and decoding an HOA signal, the encoder-side device selects a virtual speaker that matches the HOA coefficients of the current frame of the HOA signal from a virtual speaker set based on the HOA coefficients of the current frame of the HOA signal, uses the selected virtual speaker as a target virtual speaker, and further encodes attribute information of the target virtual speaker into a bitstream. The encoder-side device also further encodes low-order components of the current frame of the HOA signal into a bitstream. After receiving the bitstream, the decoder-side device analyzes the bitstream to obtain attribute information of the target virtual speaker and the low-order components of the current frame of the HOA signal. The decoder-side device then reconstructs the current frame of the HOA signal based on the HOA coefficients of the target virtual speaker and the low-order components of the current frame of the HOA signal. However, in practical applications, the positions of the target virtual speaker corresponding to two adjacent frames of the HOA signal in a three-dimensional sound field may be significantly different. As a result, two adjacent frames of the HOA signal reconstructed by the decoder-side device may sound like a spatial jump. Therefore, one embodiment of the present application provides a virtual speaker determination method. According to the method provided in this embodiment of the present application, the target virtual speaker corresponding to two adjacent frames of the HOA signal can smoothly transition between the two frames of the HOA signal, thereby solving the problem that the two reconstructed adjacent frames of the HOA signal sound like a spatial jump.

[0049] 1 is a diagram of an implementation environment according to an embodiment of the present application. The implementation environment includes a source device 10, a destination device 20, a link 30, and a storage device 40. The source device 10 is configured to encode attribute information of a target virtual speaker and low-order components of an HOA signal. Therefore, the source device 10 may also be referred to as an encoder-side device. The destination device 20 is configured to analyze a bitstream to obtain attribute information of a target virtual speaker and low-order components of an HOA signal. Therefore, the destination device 20 may also be referred to as a decoder-side device.

[0050] Link 30 may receive the bitstream generated by source device 10 and transmit the bitstream to destination device 20. Storage device 40 may receive the bitstream generated by source device 10 and store the bitstream. In this case, destination device 20 may obtain the bitstream directly from storage device 40. Alternatively, storage device 40 may correspond to a file server or another intermediate storage device that can store the bitstream generated by source device 10. In this case, destination device 20 may transmit the bitstream in a streaming format or download the bitstream stored in storage device 40.

[0051] The source device 10 and the destination device 20 each include one or more processors and memory coupled to the one or more processors. The memory may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store necessary program code in the form of instructions or computer-accessible data structures. The source device 10 and the destination device 20 each include a desktop computer, a mobile computing device, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a handheld telephone set such as a so-called "smartphone," a television set, a camera, a display device, a digital media player, a video game console, or an in-vehicle computer.

[0052] The link 30 includes one or more media or devices capable of transmitting a bit stream from the source device 10 to the destination device 20. In a possible implementation, the link 30 includes one or more communication media that can enable the source device 10 to directly transmit the bit stream to the destination device 20 in real time. In this embodiment of the present application, the source device 10 modulates the bit stream according to a communication standard, such as a wireless communication protocol, and transmits the bit stream to the destination device 20. The one or more communication media include wireless communication media and / or wired communication media. For example, the one or more communication media include a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media can be part of a packet-based network. The packet-based network can be a local area network, a wide area network, a global network (e.g., the Internet), etc. The one or more communication media can include a router, a switch, a base station, or another device that facilitates communication from the source device 10 to the destination device 20. This is not specifically limited in this embodiment of the present application.

[0053] In a possible implementation, storage device 40 is configured to store the received bitstream sent by source device 10, and destination device 20 can retrieve the bitstream directly from storage device 40. In this case, storage device 40 may comprise any one of a plurality of distributed or locally accessed data storage media, such as a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium configured to store a bitstream.

[0054] In a possible embodiment, the storage device 40 corresponds to a file server or another intermediate storage device capable of storing the bitstream generated by the source device 10, and the destination device 20 may transmit the bitstream stored on the storage device 40 in a streaming format or download it. The file server is any type of server capable of storing and transmitting the bitstream to the destination device 20. In a possible embodiment, the file server includes a network server, a file transfer protocol (FTP) server, a network attached storage (NAS) device, a local disk drive, etc. The destination device 20 may obtain the bitstream through any standard data connection (including an Internet connection). Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL) or cable modem), or a combination of a wireless channel and a wired connection suitable for obtaining a bitstream stored on a file server. The transmission of the bitstream from the storage device 40 may be a streaming transmission, a download transmission, or a combination thereof.

[0055] The implementation environment shown in Fig. 1 is only a possible implementation. In addition, the technology in the embodiment of the present application is not only applicable to the source device 10 that can encode the HOA signal and the destination device 20 that decodes the bitstream in Fig. 1, but also applicable to other devices that can encode the HOA signal and decode the bitstream. This is not specifically limited in the embodiment of the present application.

[0056] 1, source device 10 includes data source 120, encoder 100, and output interface 140. In some embodiments, output interface 140 includes a modulator / demodulator (modem) and / or a transmitter. A transmitter may also be referred to as a transmitter. Data source 120 includes an HOA signal capture device, an archive containing previously captured HOA signals, a feed-in interface for receiving HOA signals from an HOA signal content provider, and / or a computer graphics system for generating HOA signals, or a combination of these sources of HOA signals.

[0057] Data source 120 is configured to transmit an HOA signal to encoder 100, which is configured to encode the received HOA signal transmitted from data source 120 to obtain a bitstream. The encoder transmits the bitstream to an output interface. In some embodiments, source device 10 transmits the bitstream directly to destination device 20 through output interface 140. In another embodiment, the bitstream may alternatively be stored in storage device 40, so that destination device 20 subsequently retrieves the bitstream for decoding and / or display.

[0058] In the implementation illustrated in FIG. 1 , destination device 20 includes an input interface 240, a decoder 200, and a display device 220. In some embodiments, input interface 240 includes a receiver and / or a modem. Input interface 240 may receive a bitstream over link 30 and / or from storage device 40 and then transmit the bitstream to decoder 200. Decoder 200 may decode the received bitstream to obtain a reconstructed HOA signal. The decoder transmits the reconstructed HOA signal to display device 220. Display device 220 may be integrated with destination device 20 or located external to destination device 20. Generally, display device 220 displays the reconstructed HOA signal. Display device 220 may be any one of several types of display devices. For example, display device 220 may be a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0059] 1, in some aspects, encoder 100 and decoder 200 may be integrated into an audio encoder and decoder, respectively, and may include appropriate multiplexer-demultiplexer (MUX-DEMUX) units or other hardware and software for encoding both audio and video in the same data stream or separate data streams. In some embodiments, if applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol or another protocol, such as the User Datagram Protocol (UDP).

[0060] The encoder 100 and the decoder 200 may each be one of the following circuits: one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. Where the techniques in the embodiments of the present application are partially implemented in software, a device may store instructions for the software in an appropriate non-volatile computer-readable storage medium and execute the instructions in hardware via one or more processors to implement the techniques in the embodiments of the present application. Any one of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered one or more processors. Each of the encoder 100 and the decoder 200 may be included in one or more encoders or decoders. Any of the encoders or decoders may be integrated as part of an encoder / decoder combination (codec) in a corresponding device.

[0061] In embodiments of the present application, the encoder 100 may generally be referred to as "signaling" or "transmitting" some information to another device, such as the decoder 200. The terms "signaling" or "transmitting" may generally refer to the transmission of syntax elements and / or other data used to decode the bitstream. Such transmission may occur in real time or near real time. Alternatively, such communication may occur after a period of time, such as when the syntax elements in the encoded bitstream are stored to a computer-readable storage medium during encoding. The decoding device can then retrieve the syntax elements at any time after they are stored to the medium.

[0062] The virtual speaker determination method provided in this embodiment of the present application can be applied to several scenarios, some of which will be described individually below.

[0063] 2 is a diagram of an implementation environment in which the virtual speaker determination method according to an embodiment of the present application is applied to a terminal scenario. The implementation environment includes a first terminal 101 and a second terminal 201. The first terminal 101 establishes a communication connection with the second terminal 201. The communication connection may be a wireless network connection or a wired network connection, which is not limited in the embodiment of the present application.

[0064] The first terminal 101 may be a sending device or a receiving device. Similarly, the second terminal 201 may be a receiving device or a sending device. If the first terminal 101 is a sending device, the second terminal 201 is a receiving device. If the first terminal 101 is a receiving device, the second terminal 201 is a sending device.

[0065] In the following, an example will be described in which the first terminal 101 is a sending device and the second terminal 201 is a receiving device.

[0066] The first terminal 101 may be the source device 10 in the implementation environment shown in Fig. 1. The second terminal 201 may be the destination device 20 in the implementation environment shown in Fig. 1. The first terminal 101 and the second terminal 201 each include an audio capture module, an audio playback module, an encoder, a decoder, a channel encoding module, and a channel decoding module.

[0067] The audio capture module of the first terminal 101 captures the HOA signal and sends it to the encoder. The encoder determines a target virtual speaker by using the virtual speaker determination method provided in this embodiment of the present application. Furthermore, the attribute information of the target virtual speaker and the low-order components of the current frame of the HOA signal are encoded, which may be referred to as source encoding. Then, to transmit the HOA signal over a channel, the channel encoding module needs to further perform channel encoding, and then the bitstream obtained by encoding is transmitted over a digital channel via a wireless or wired network communication device.

[0068] The second terminal 201 receives the bitstream transmitted on the digital channel through a wireless or wired network communication device. The channel decoding module performs channel decoding on the bitstream, and then the decoder reconstructs the current frame of the HOA signal based on the HOA coefficients of the target virtual speaker and the low-order components of the current frame of the HOA signal, and then plays the HOA signal through the audio playback module.

[0069] The first terminal 101 and the second terminal 201 may be any electronic product capable of performing human-computer interaction with a user via one or more of the following: a keyboard, a touchpad, a touchscreen, a remote control, a voice interaction device, a handwriting device, etc., such as a personal computer (PC), a mobile phone, a smartphone, a personal digital assistant (PDA), a wearable device, a pocket personal computer (PPC), a tablet computer, a smart automobile head unit, a smart TV, or a smart speaker.

[0070] Those skilled in the art should understand that the above terminals are merely examples, and other existing or future possible terminals to which the embodiments of the present application are applicable should fall within the protection scope of the embodiments of the present application and are hereby incorporated by reference.

[0071] 3 is a diagram of an implementation environment in which the virtual speaker determination method according to an embodiment of the present application is applied to radio and television scenarios. Radio and television scenarios include live streaming scenarios and post-production scenarios. In the live streaming scenario, the implementation environment includes a live program 3D audio generation module, a 3D audio encoding module, a set-top box, and a speaker group. The set-top box includes a 3D audio decoding module. In the post-production scenario, the implementation environment includes a post-program 3D audio generation module, a 3D audio encoding module, a network receiver, a mobile terminal, a headset, etc.

[0072] In a live streaming scenario, the live program 3D audio generation module generates a 3D audio signal. The 3D audio signal includes an HOA signal. The 3D audio signal is encoded using an existing encoding method to obtain a bitstream. The bitstream is transmitted to a user via a radio and television network and decoded by a 3D audio decoder in a set-top box using an existing decoding method to reconstruct the 3D audio signal. A group of speakers plays the reconstructed 3D audio signal. Alternatively, the bitstream is transmitted to the user via the Internet and decoded by a 3D audio decoder in a network receiver using an existing decoding method to reconstruct the 3D audio signal. A group of speakers plays the reconstructed 3D audio signal. Alternatively, the bitstream is transmitted to the user via the Internet and decoded by a 3D audio decoder in a mobile device using an existing decoding method to reconstruct the 3D audio signal. A headset plays the reconstructed 3D audio signal.

[0073] In a post-production scenario, the post-program 3D audio generation module generates a 3D audio signal. The 3D audio signal is encoded using an existing encoding method to obtain a bitstream. The bitstream is transmitted to a user via a radio and television network and decoded by a 3D audio decoder in a set-top box using an existing decoding method to reconstruct the 3D audio signal. A group of speakers plays the reconstructed 3D audio signal. Alternatively, the bitstream is transmitted to a user via the Internet and decoded by a 3D audio decoder in a network receiver using an existing decoding method to reconstruct the 3D audio signal. A group of speakers plays the reconstructed 3D audio signal. Alternatively, the bitstream is transmitted to a user via the Internet and decoded by a 3D audio decoder in a mobile device using an existing decoding method to reconstruct the 3D audio signal. A headset plays the reconstructed 3D audio signal.

[0074] 4 is a diagram of an implementation environment in which the virtual speaker determination method is applied to a virtual reality streaming scenario according to an embodiment of the present application. The implementation environment includes an encoder side and a decoder side. The encoder side includes a capture module, a pre-processing module, an encoding module, a packetization module, and a transmission module. The decoder side includes a de-packetization module, a decoding module, a rendering module, and a headset.

[0075] The capture module captures the HOA signal. The pre-processing module then performs pre-processing operations. The pre-processing operations typically include removing low-frequency components from the signal by setting 20 Hz or 50 Hz as the boundary point, and extracting orientation information from the signal. The encoding module then performs encoding by using an existing encoding method. After encoding, the packetization module performs packetization. The transmission module then transmits the packetized signal to the decoder side.

[0076] The de-packetization module on the decoder side first performs de-packetization. Then, the decoding module performs decoding by using an existing decoding method. Then, the rendering module performs binaural rendering on the decoded signal. The rendered signal is mapped to the listener's headset. The headset may be a standalone headset or a headset of a virtual reality-based glasses device.

[0077] It should be noted that the system architecture and service scenarios described in the embodiments of the present application are intended to more clearly describe the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions provided in the embodiments of the present application. With the evolution of system architecture and the emergence of new service scenarios, those skilled in the art may recognize that the technical solutions provided in the embodiments of the present application can also be applied to similar technical problems.

[0078] The following describes in detail the virtual speaker determination method provided in the embodiments of the present application. Referring to the implementation environment shown in Figure 1, it should be noted that the virtual speaker determination method may be performed by the encoder 100 of the source device 10 or the decoder 200 of the destination device 20.

[0079] 5 is a flowchart of a virtual speaker determination method according to an embodiment of the present application. The method is applied to an encoder-side device. Please refer to FIG. 5. The method includes the following steps:

[0080] Step 501: Obtain attribute information of N first virtual speakers that are in a virtual speaker set and match HOA coefficients of a target group of HOA signals, where the target group of HOA signals includes at least one frame of the HOA signal, and N is an integer equal to or greater than 1.

[0081] In some embodiments, at least one frame of the HOA signal that currently needs to be encoded is used as the target group of the HOA signal, where the target group of the HOA signal includes one frame of the HOA signal, or the target group of the HOA signal includes P frames of the HOA signal, where P is an integer greater than 1.

[0082] The virtual speaker set includes a plurality of virtual speakers, each of which has a corresponding HOA coefficient. The encoder-side device selects N first virtual speakers from the virtual speaker set based on the HOA coefficients of the at least one frame of the HOA signal and the HOA coefficients of each virtual speaker. Then, the encoder-side device acquires attribute information of the N first virtual speakers from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers based on identifiers of the N first virtual speakers.

[0083] When the target group of HOA signals includes an HOA signal of one frame, the encoder-side device separately performs an inner product operation between the HOA coefficient of the current frame of the HOA signal and the HOA coefficient of each virtual speaker to obtain multiple operation results. Any operation result among the multiple operation results is a projection component of the current frame of the HOA signal onto the corresponding virtual speaker. Next, the encoder-side device sorts the multiple operation results in descending order of the projection component, and uses the virtual speakers corresponding to the first N operation results among the sorted results as the N first virtual speakers.

[0084] When the target group of the HOA signal includes P frames of the HOA signal, the encoder-side device sequentially performs an inner product operation between the HOA coefficient of each frame of the HOA signal and the HOA coefficient of each virtual speaker for each frame of the HOA signal among the P frames of the HOA signal to obtain multiple operation results. Any operation result among the multiple operation results is a projection component of a specific frame of the HOA signal among the P frames of the HOA signal for a corresponding virtual speaker. Next, the encoder-side device sorts the multiple operation results in descending order of projection component, and uses the virtual speakers corresponding to the first N operation results among the sorted results as N first virtual speakers.

[0085] It should be noted that when the target group of the HOA signal includes P frames of the HOA signal, the N first virtual speakers may not include the first virtual speaker with which the current frame of the HOA signal matches for a particular frame of the HOA signal among the P frames of the HOA signal. In other words, when the P frames of the HOA signal match a total of N first virtual speakers, the number of first virtual speakers with which all of the P frames of the HOA signal match is not equal.

[0086] Certainly, in a practical application, the encoder-side device may further select the N first virtual speakers from the virtual speaker set according to another method, which is not limited in the embodiments of the present application.

[0087] The virtual speaker identifier uniquely identifies the virtual speaker and may be the type, number, name, etc. of the virtual speaker, or a combination of these pieces of information. The attribute information of the virtual speaker includes an elevation angle and an azimuth angle. Indeed, in a practical application, the attribute information of the virtual speaker may further include other content, such as the HOA coefficient of the virtual speaker and the index of the virtual speaker. This is not limited to the embodiments of the present application.

[0088] Optionally, before selecting the N first virtual speakers that match the HOA coefficients of the at least one frame of the HOA signal from the virtual speaker set based on the HOA coefficients of the at least one frame of the HOA signal and the HOA coefficients of each virtual speaker, the encoder-side device needs to separately perform a time-frequency transform on the at least one frame of the HOA signal. In other words, the at least one frame of the time-domain HOA signal is transformed into a frequency-domain HOA signal to obtain frequency-domain coefficients of the at least one frame of the HOA signal, and then the frequency-domain coefficients of the at least one frame of the HOA signal are determined as the HOA coefficients of the at least one frame of the HOA signal.

[0089] Generally, the number of channels of an HOA signal is related to the order of the HOA signal. For example, if one frame of an HOA signal is a Z-th order signal, the number of channels of the current frame of the HOA signal is (Z+1). 2 The encoder-side device selects N first virtual speakers from the virtual speaker set according to the above steps, so that the decoder-side device subsequently selects a virtual speaker set with a channel number of (Z+1) based on the HOA coefficients of the N first virtual speakers. 2 The current frame of the HOA signal is converted into a virtual speaker signal with N channels.

[0090] Step 502: Obtain attribute information of N second virtual speakers in the virtual speaker set, each configured to perform an encoding process on a reference group of HOA signals, where the reference group of HOA signals is at least one group of HOA signals preceding the target group of HOA signals.

[0091] In practical application, for the encoder-side device, the N second virtual speakers are configured to perform the encoding process on the reference group of HOA signals.

[0092] In some embodiments, the reference group of HOA signals is one group of HOA signals preceding the target group of HOA signals. Alternatively, the reference group of HOA signals is multiple groups of HOA signals preceding the target group of HOA signals. Depending on the case, the method by which the encoder-side device obtains the attribute information of the N second virtual speakers differs. The following describes the two cases separately.

[0093] In the first case, the reference group of HOA signals is one group of HOA signals preceding the target group of HOA signals, in which case the encoder-side device directly uses the N virtual speakers configured to perform the encoding process on this group of HOA signals as the N second virtual speakers, and obtains the attribute information of the N second virtual speakers based on the identifiers of the N second virtual speakers from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers.

[0094] The N second virtual speakers configured to perform the encoding process on this group of HOA signals have a one-to-one correspondence with the N first virtual speakers whose target group of HOA signals matches, i.e., for any group of HOA signals, N virtual speakers need to be selected from the virtual speaker set according to the method provided in this embodiment of the present application to obtain N virtual speakers configured to perform the encoding process on this group of HOA signals.

[0095] In the second case, the reference group of HOA signals is multiple groups of HOA signals that precede the target group of HOA signals.

[0096] Each group of HOA signals in the multiple groups of HOA signals corresponds to N virtual speakers, and the N virtual speakers corresponding to each group of HOA signals correspond one-to-one to the N virtual speakers corresponding to another group of HOA signals. In this case, the encoder-side device obtains N groups of virtual speakers by using virtual speakers that correspond to the multiple groups of HOA signals as groups of virtual speakers. Any group of virtual speakers in the N groups of virtual speakers includes a virtual speaker that corresponds to each group of HOA signals in the multiple groups of HOA signals. Next, based on the identifiers of the multiple virtual speakers, attribute information of multiple virtual speakers included in any group of virtual speakers in the N groups of virtual speakers is obtained from the correspondence between the stored identifiers and the stored attribute information of the virtual speakers, thereby obtaining one group of attribute information. In this way, for each group of virtual speakers in the N groups of virtual speakers, one group of attribute information can be determined according to the above steps, thereby obtaining N groups of attribute information. Finally, averaging is performed on the same group of attribute information in the N groups of attribute information to obtain N pieces of attribute information, and the N pieces of attribute information are determined as attribute information of the N second virtual speakers to obtain attribute information of the N second virtual speakers.

[0097] For example, the reference group of HOA signals is three groups of HOA signals before the target group of HOA signals, and each group of HOA signals in the three groups of HOA signals corresponds to four virtual speakers, i.e., N is 4. The four virtual speakers corresponding to the first group of HOA signals are a1, b1, c1, and d1, the four virtual speakers corresponding to the second group of HOA signals are a2, b2, c2, and d2, and the four virtual speakers corresponding to the third group of HOA signals are a3, b3, c3, and d3. In this case, the encoder-side device uses the virtual speakers corresponding to the three groups of HOA signals as the groups of virtual speakers, and the four groups of virtual speakers obtained are [a1, a2, a3], [b1, b2, b3], [c1, c2, c3], and [d1, d2, d3]. Next, for each group of virtual speakers among the four groups of virtual speakers, the encoder-side device averages the attribute information of the three virtual speakers in the same group to obtain four pieces of attribute information, and determines the four pieces of attribute information as the attribute information of the four second virtual speakers.

[0098] Step 503: Determine M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers. The M target virtual speakers are configured to perform encoding processing on a target group of HOA signals, where M is an integer greater than 1, and M is greater than N.

[0099] When the attribute information of the virtual speakers includes an elevation angle and an azimuth angle, the encoder-side device determines M target virtual speakers according to the following steps (1) to (3).

[0100] (1) Based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers, distances between the corresponding first virtual speakers and the corresponding second virtual speakers are determined to obtain N distances.

[0101] Based on the above description, the N first virtual speakers have a one-to-one correspondence with the N second virtual speakers. For any first virtual speaker among the N first virtual speakers, the method for determining the distance between the first virtual speaker and the corresponding second virtual speaker is the same. Therefore, a first virtual speaker is selected from the N first virtual speakers as a target first virtual speaker. Hereinafter, determining the distance between the target first virtual speaker and the target second virtual speaker will be described using the target first virtual speaker as an example. There is a correspondence relationship between the target second virtual speaker and the target first virtual speaker.

[0102] For example, the encoder-side device determines the distance between the target first virtual speaker and the target second virtual speaker according to the following equation (1). d1=arccos[cosβ 11 cosβ 12 cos(∂ 11 -∂ 12 )+sinβ 11 sinβ 12 ]

[0103] In the above formula (1), d1 represents the distance between the target first virtual speaker and the target second virtual speaker, and β 11 denotes the azimuth angle of the first virtual speaker of interest, and β 12 denotes the azimuth angle of the second virtual loudspeaker of interest, and ∂ 11 denotes the elevation angle of the first virtual loudspeaker of interest, and ∂ 12 denotes the elevation angle of the second virtual loudspeaker of interest.

[0104] That is, for any one of the N first virtual speakers, a second virtual speaker corresponding to the first virtual speaker is selected from the N second virtual speakers, and the second virtual speaker and the first virtual speaker correspond to the same channel. Then, based on the elevation angle and azimuth angle of the first virtual speaker and the elevation angle and azimuth angle of the second virtual speaker, the distance between the first virtual speaker and the second virtual speaker is determined according to the above formula (1) to obtain one distance. In this way, for each first virtual speaker in the N first virtual speakers, a second virtual speaker corresponding to the first virtual speaker can be determined according to the above steps, and the distance between the first virtual speaker and the corresponding second virtual speaker is determined to obtain N distances.

[0105] (2) Determine M groups of elevation and azimuth angles based on the N distances.

[0106] Based on the above description, the target group of the HOA signal includes one frame of the HOA signal, or the target group of the HOA signal includes P frames of the HOA signal. In some cases, the encoder-side device determines M groups of elevation angles and azimuth angles based on N distances in different ways. The following describes the two cases separately.

[0107] In the first case, the target group of the HOA signal includes one frame of the HOA signal, and the current frame of the HOA signal includes H subframes, where H is an integer greater than 1. To obtain H groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the H subframes included in the current frame of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*H=M groups of elevation angles and azimuth angles are obtained.

[0108] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the H subframes are determined by the following operation, i.e., when the target distance is greater than a first distance threshold, determining the elevation angle and azimuth angle corresponding to each of the H subframes based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0109] Based on the above description, the N distances are the distances between the corresponding first and second virtual speakers. If the target distance is greater than the first distance threshold, it indicates a large difference between the position of the target first virtual speaker in the current frame of the HOA signal and the position of the target second virtual speaker in the reference group of the HOA signal. As a result, the current frame of the HOA signal and the reference group of the HOA signal subsequently obtained by decoding will sound like a spatial jump. Therefore, the encoder-side device needs to determine the elevation angles and azimuth angles corresponding to the H subframes based on the elevation angles and azimuth angles of the first and second virtual speakers corresponding to the target distances, so that a smooth transition is made between the first and second virtual speakers corresponding to the target distances.

[0110] For example, an implementation process for determining elevation angles and azimuth angles corresponding to H subframes based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a first subframe among the H subframes; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a last subframe among the H subframes; and determining, for an i-th subframe among the H subframes, the elevation angle and azimuth angle corresponding to the i-th subframe by an interpolation process based on the elevation angle and azimuth angle corresponding to the (i-1)-th subframe among the H subframes and the elevation angle and azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1.

[0111] Note that i is the number of any subframe other than the first and last subframes among the H subframes. If the first subframe among the H subframes is numbered from 0, then i is greater than 0 and less than H-1. If the first subframe among the H subframes is numbered from 1, then i is greater than 1 and less than H. In other words, the elevation angle and azimuth angle corresponding to any subframe other than the first and last subframes among the H subframes are determined by an interpolation process.

[0112] That is, the elevation angle and azimuth angle corresponding to the first subframe of the H subframes are the elevation angle and azimuth angle of the target second virtual speaker of the reference group of the HOA signal, and the elevation angle and azimuth angle corresponding to the last subframe of the H subframes are the elevation angle and azimuth angle of the target first virtual speaker of the current frame of the HOA signal. The elevation angle and azimuth angle corresponding to any subframe other than the first subframe and the last subframe of the H subframes needs to be obtained by interpolation based on the elevation angle and azimuth angle of the previous subframe closest to that subframe and the elevation angle and azimuth angle corresponding to the last subframe. In this way, when the target group of the HOA signal includes one frame of the HOA signal, interpolation is performed among the H subframes included in the current frame of the HOA signal to achieve a smooth transition between the first and second virtual speakers corresponding to the target distance.

[0113] For the i-th subframe among the H subframes, the starting point of the interpolation process for the i-th subframe is the elevation angle and azimuth angle corresponding to the (i-1)-th subframe, and the ending point of the interpolation process is the elevation angle and azimuth angle corresponding to the last subframe. In other words, for any subframe among the H subframes other than the first and last subframes, the starting point of the subframe interpolation process is constantly updated in real time. In this way, the elevation angle and azimuth angle corresponding to each of the H subframes can be more accurately determined.

[0114] For example, the encoder-side device determines the elevation angle and azimuth angle corresponding to the i-th subframe according to the following equation (2).

number

[0115] In the above equation (2), ∂ i denotes the elevation angle corresponding to the i-th subframe, and ∂ i-1denotes the elevation angle corresponding to the (i-1)th subframe, and ∂ H denotes the elevation angle corresponding to the last subframe, and β i denotes the azimuth angle corresponding to the i-th subframe, and β i-1 denotes the azimuth angle corresponding to the (i-1)th subframe, and β H indicates the azimuth angle corresponding to the last subframe.

[0116] It should be noted that in the above equation (2), the elevation angle and azimuth angle corresponding to the i-th subframe are determined based on the elevation angle and azimuth angle corresponding to the (i-1)-th subframe and the elevation angle and azimuth angle corresponding to the last subframe by using a linear interpolation method. Indeed, in practical applications, the encoder-side device can further determine the elevation angle and azimuth angle corresponding to the i-th subframe by using a nonlinear interpolation method, such as Lagrange interpolation, which is not limited in the embodiments of the present application.

[0117] For example, the current frame of the HOA signal contains four subframes, and the elevation angle of the first virtual speaker corresponding to the target distance is ∂ 11 , azimuth angle is β 11 and the elevation angle of the second virtual speaker corresponding to the target distance is ∂ 12 , azimuth angle is β 12 If the target distance threshold is greater than the first distance threshold, the elevation angle corresponding to the first subframe is ∂ 12 and the azimuth angle is β 12 The elevation angle corresponding to the fourth subframe is ∂ 11 and the azimuth angle is β 11 The elevation angle ∂2 and azimuth angle β2 correspond to the second subframe, and the elevation angle ∂2 and azimuth angle β2 correspond to the elevation angle ∂ 12 and azimuth angle β 12 and the elevation angle ∂ corresponding to the fourth subframe 11 and azimuth angle β 11The elevation angle ∂3 and azimuth angle β3 corresponding to the third subframe are obtained by interpolation based on the following equation: The elevation angle ∂3 and azimuth angle β3 correspond to the second subframe and the elevation angle ∂2 and azimuth angle β2 correspond to the fourth subframe. 11 and azimuth angle β 11 is obtained by interpolation based on

[0118] It should be noted that in practical applications, the target distance may be less than or equal to the first distance threshold. In other words, the position of the target first virtual speaker in the current frame of the HOA signal is not significantly different from the position of the target second virtual speaker in the reference group of the HOA signal. In some embodiments, the encoder-side device determines the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the H subframes, respectively. In other words, the elevation angle corresponding to each frame in the H subframes is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each subframe is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0119] In some other embodiments, the encoder-side device determines the elevation and azimuth angles of the second virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the first K subframes of the H subframes, and determines the elevation and azimuth angles of the first virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the remaining subframes of the H subframes, where K is an integer greater than or equal to 1 and K is less than H.

[0120] For example, the current frame of the HOA signal contains four subframes, and the elevation angle of the first virtual speaker corresponding to the target distance is ∂ 11 , azimuth angle is β 11 and the elevation angle of the second virtual speaker corresponding to the target distance is ∂ 12 , azimuth angle is β 12 When the target distance threshold is equal to or less than the first distance threshold, the elevation angle corresponding to each of the four subframes is ∂ 11 and the azimuth angle is β11 Alternatively, the elevation angle corresponding to the first of the four subframes is ∂ 12 , and the azimuth angle is β 12 That is, K is 1, and the elevation angle corresponding to each of the remaining three subframes is ∂ 11 , and the azimuth angle is β 11 is.

[0121] The first distance threshold is preset, for example, 0.5, and may be adjusted based on different requirements.

[0122] In the second case, the target group of the HOA signal includes P frames of the HOA signal. To obtain P groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the P frames of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*P=M groups of elevation angles and azimuth angles are obtained.

[0123] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal are determined by the following operation, i.e., when the target distance is greater than the second distance threshold, determining the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0124] Based on the above description, the N distances are the distances between the first virtual speaker and the second virtual speaker that have a corresponding relationship. If the target distance is greater than the second distance threshold, it indicates that there is a large difference between the position of the target first virtual speaker of the target group of the HOA signal and the position of the target second virtual speaker of the reference group of the HOA signal. As a result, the target group of the HOA signal and the reference group of the HOA signal subsequently obtained by decoding will sound like they have a spatial jump. Therefore, the encoder-side device needs to determine the elevation angles and azimuth angles corresponding to each of the P frames of the HOA signal based on the elevation angles and azimuth angles of the first virtual speaker and the second virtual speaker corresponding to the target distance, so that a smooth transition is made between the first virtual speaker and the second virtual speaker corresponding to the target distance.

[0125] For example, an implementation process for determining elevation angles and azimuth angles corresponding to P frames of an HOA signal based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the first frame of the HOA signal among the P frames of the HOA signal; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the last frame of the HOA signal among the P frames of the HOA signal; and determining, for the jth frame of the HOA signal among the P frames of the HOA signal, the elevation angle and azimuth angle corresponding to the jth frame of the HOA signal by an interpolation process based on the elevation angle and azimuth angle corresponding to the (j-1)th frame of the HOA signal among the P frames of the HOA signal and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal, where j is greater than 0 and less than P-1.

[0126] Note that j is the number of any frame of the HOA signal other than the HOA signal of the first frame and the HOA signal of the last frame among the P frames of the HOA signal. If the HOA signal of the first frame among the P frames of the HOA signal is numbered from 0, j is greater than 0 and less than P-1. If the HOA signal of the first frame among the P frames of the HOA signal is numbered from 1, j is greater than 1 and less than P. In other words, the elevation angle and azimuth angle corresponding to any frame of the HOA signal other than the HOA signal of the first frame and the HOA signal of the last frame among the P frames of the HOA signal are determined by an interpolation process.

[0127] That is, the elevation angle and azimuth angle corresponding to the first frame of the HOA signal among the P frames of the HOA signal are the elevation angle and azimuth angle of the target second virtual speaker of the reference group of the HOA signal, and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal among the P frames of the HOA signal are the elevation angle and azimuth angle of the first virtual speaker in the target group of the HOA signal. The elevation angle and azimuth angle corresponding to any frame of the HOA signal other than the first frame of the HOA signal and the last frame of the HOA signal among the P frames of the HOA signal must be obtained by interpolation based on the elevation angle and azimuth angle of the previous frame of the HOA signal that is closest to the current frame of the HOA signal and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal. In this way, when the target group of the HOA signal includes P frames of the HOA signal, interpolation is performed between the P frames of the HOA signal to achieve a smooth transition between the first virtual speaker and the second virtual speaker corresponding to the target distance.

[0128] The starting point of the interpolation process for the jth frame of the HOA signal among the P frames of the HOA signal is the elevation angle and azimuth angle corresponding to the (j-1)th frame of the HOA signal, and the ending point of the interpolation process is the elevation angle and azimuth angle corresponding to the last frame of the HOA signal. In other words, for any frame of the HOA signal among the P frames of the HOA signal other than the first frame and the last frame of the HOA signal, the starting point of the interpolation process for the HOA signal frame is always updated in real time. In this way, the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal can be more accurately determined.

[0129] It should be noted that in practical applications, the target distance may be less than or equal to the second distance threshold. In other words, the position of the first virtual speaker of interest in the target group of the HOA signal is not significantly different from the position of the second virtual speaker of interest in the reference group of the HOA signal. In some embodiments, the encoder-side device determines the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to P frames of the HOA signal, respectively. In other words, the elevation angle corresponding to each frame of the HOA signal in the P frames of the HOA signal is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each frame of the HOA signal is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0130] In some other embodiments, the encoder-side device determines the elevation and azimuth angles of the second virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the first L frames of the P frames of the HOA signal, and determines the elevation and azimuth angles of the first virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the remaining frames of the HOA signal of the P frames of the HOA signal, where L is an integer greater than or equal to 1 and L is less than P.

[0131] The second distance threshold is preset, and may be equal to or unequal to the first distance threshold, and may be adjusted based on different requirements.

[0132] (3) Virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles are determined as M target virtual speakers.

[0133] After determining the M groups of elevation angles and azimuth angles based on the N distances according to the above step (2), the encoder-side device determines the virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles as M target virtual speakers, and as a result, the M target virtual speakers will thereafter perform encoding processing on the target group of HOA signals.

[0134] Based on the above description, in practical applications, the attribute information of the virtual speakers may further include other content, such as the HOA coefficients of the virtual speakers. If the attribute information of the virtual speakers includes the HOA coefficients, the encoder-side device first needs to convert the HOA coefficients of the virtual speakers into the elevation angles and azimuth angles of the virtual speakers according to the relevant algorithm, and then determine the M target virtual speakers according to the above steps (1) to (3).

[0135] Optionally, the encoder-side device may further determine M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, and then encode the attribute information of the M target virtual speakers into the bitstream. In this way, after receiving the bitstream, the decoder-side device may analyze the bitstream to obtain the attribute information of the M target virtual speakers, and reconstruct the target group of the HOA signal based on the attribute information of the M target virtual speakers. Alternatively, the encoder-side device may directly encode indexes of the determination method of the M target virtual speakers into the bitstream, so that the decoder-side device may analyze the bitstream to obtain the indexes of the determination method of the M target virtual speakers, and then determine the M target virtual speakers in real time based on the indexes.

[0136] In this embodiment of the present application, the target virtual speaker is configured to process a target group of HOA signals, and the second virtual speaker is configured to process a reference group of HOA signals. The first virtual speaker is a virtual speaker that matches the target group of the HOA signals. Therefore, after the first virtual speaker is determined, the target virtual speaker is determined based on the attribute information of the second virtual speaker and the attribute information of the first virtual speaker to ensure that the attribute information of the target virtual speaker is not significantly different from the attribute information of the second virtual speaker, thereby solving the problem that two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump.

[0137] 6 is a flowchart of another virtual speaker determination method according to an embodiment of the present application. The method is applied to a decoder-side device. Please refer to FIG. 6. The method includes the following steps:

[0138] Step 601: Obtain attribute information of N first virtual speakers that are in a virtual speaker set and match HOA coefficients of a target group of HOA signals, where the target group of HOA signals includes at least one frame of the HOA signal, and N is an integer equal to or greater than 1.

[0139] In some embodiments, at least one frame of the HOA signal that currently needs to be decoded is used as the target group of the HOA signal, where the target group of the HOA signal includes one frame of the HOA signal, or the target group of the HOA signal includes P frames of the HOA signal, where P is an integer greater than 1.

[0140] The process in which the decoder-side device acquires the attribute information of the N first virtual speakers is similar to the process in which the encoder-side device acquires the attribute information of the N first virtual speakers in the above step 501. Therefore, for details, please refer to the relevant content of the above step 501. The details will not be repeated here.

[0141] Optionally, the encoder-side device may further encode the attribute information of the N first virtual speakers into a bitstream after obtaining the attribute information of the N first virtual speakers according to the aforementioned step 501. In this way, after receiving the bitstream, the decoder-side device may directly analyze the bitstream to obtain the attribute information of the N first virtual speakers.

[0142] The attribute information of the virtual speaker includes an elevation angle and an azimuth angle. In practical applications, the attribute information of the virtual speaker may further include other content, such as the HOA coefficient of the virtual speaker and the index of the virtual speaker. This is not limited to the embodiments of the present application.

[0143] Step 602: Obtain attribute information of N second virtual speakers in the virtual speaker set, each configured to perform a decoding process on a reference group of HOA signals, where the reference group of HOA signals is at least one group of HOA signals preceding the target group of HOA signals.

[0144] In the case of the decoder-side device, the N second virtual speakers are configured to perform decoding processing on the reference group of HOA signals. The process by which the decoder-side device acquires attribute information of the N second virtual speakers is similar to the process by which the encoder-side device acquires attribute information of the N second virtual speakers in the above step 502. Therefore, for details, please refer to the relevant content of the above step 502. The details will not be repeated here.

[0145] Optionally, the encoder-side device may further encode the attribute information of the N second virtual speakers into a bitstream after obtaining the attribute information of the N second virtual speakers according to the aforementioned step 502. In this way, after receiving the bitstream, the decoder-side device may directly analyze the bitstream to obtain the attribute information of the N second virtual speakers.

[0146] Step 603: Determine M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers. The M target virtual speakers are configured to perform decoding processing on a target group of HOA signals, where M is an integer greater than 1, and M is greater than N.

[0147] In some embodiments, the encoder-side device determines M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, and then further encodes an index of a method for determining the M target virtual speakers into the bitstream. Thus, after receiving the bitstream, the decoder-side device can analyze the bitstream to obtain the index of the method for determining the M target virtual speakers, and then determine the M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers in the determination method indicated by the index.

[0148] In some other embodiments, when the attribute information of the virtual speakers includes an elevation angle and an azimuth angle, the decoder-side device determines M target virtual speakers according to the following steps (1) to (3).

[0149] (1) Based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers, distances between the corresponding first virtual speakers and the corresponding second virtual speakers are determined to obtain N distances.

[0150] The process in which the decoder-side device determines N distances based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers is similar to the process in which the encoder-side device determines N distances based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers in the aforementioned step 503. Therefore, for details, please refer to the relevant content of the aforementioned step 503. The details will not be repeated here.

[0151] (2) Determine M groups of elevation and azimuth angles based on the N distances.

[0152] Based on the above description, the target group of the HOA signal includes one frame of the HOA signal, or the target group of the HOA signal includes P frames of the HOA signal. In different cases, the decoder-side device determines M groups of elevation angles and azimuth angles based on N distances in different ways. The following describes the two cases separately.

[0153] In the first case, the target group of the HOA signal includes one frame of the HOA signal, and the current frame of the HOA signal includes H subframes, where H is an integer greater than 1. To obtain H groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the H subframes included in the current frame of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*H=M groups of elevation angles and azimuth angles are obtained.

[0154] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the H subframes are determined by the following operation, i.e., when the target distance is greater than a first distance threshold, determining the elevation angle and azimuth angle corresponding to each of the H subframes based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0155] For example, an implementation process for determining elevation angles and azimuth angles corresponding to H subframes based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a first subframe among the H subframes; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a last subframe among the H subframes; and determining, for an i-th subframe among the H subframes, the elevation angle and azimuth angle corresponding to the i-th subframe by an interpolation process based on the elevation angle and azimuth angle corresponding to the (i-1)-th subframe among the H subframes and the elevation angle and azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1.

[0156] It should be noted that in practical applications, the target distance may be less than or equal to the first distance threshold. In other words, the position of the target first virtual speaker in the current frame of the HOA signal is not significantly different from the position of the target second virtual speaker in the reference group of the HOA signal. In some embodiments, the decoder-side device determines the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the H subframes, respectively. In other words, the elevation angle corresponding to each frame in the H subframes is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each subframe is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0157] In some other embodiments, the decoder-side device determines the elevation and azimuth angles of the second virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the first K subframes of the H subframes, and determines the elevation and azimuth angles of the first virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the remaining subframes of the H subframes, where K is an integer greater than or equal to 1 and K is less than H.

[0158] The first distance threshold is preset, for example, 0.5, and may be adjusted based on different requirements.

[0159] In the second case, the target group of the HOA signal includes P frames of the HOA signal. To obtain P groups of elevation angles and azimuth angles for each distance in the N distances, the elevation angles and azimuth angles corresponding to the P frames of the HOA signal are determined based on the distance until each distance in the N distances is traversed, and N*P=M groups of elevation angles and azimuth angles are obtained.

[0160] One distance among the N distances is used as the target distance, and the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal are determined by the following operation, i.e., when the target distance is greater than the second distance threshold, determining the elevation angle and azimuth angle corresponding to each of the P frames of the HOA signal based on the elevation angle and azimuth angle of the first virtual speaker and the second virtual speaker corresponding to the target distance, until each distance in the N distances is traversed.

[0161] For example, an implementation process for determining elevation angles and azimuth angles corresponding to P frames of an HOA signal based on elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to a target distance includes the steps of: determining the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the first frame of the HOA signal among the P frames of the HOA signal; determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the last frame of the HOA signal among the P frames of the HOA signal; and determining, for the jth frame of the HOA signal among the P frames of the HOA signal, the elevation angle and azimuth angle corresponding to the jth frame of the HOA signal by an interpolation process based on the elevation angle and azimuth angle corresponding to the (j-1)th frame of the HOA signal among the P frames of the HOA signal and the elevation angle and azimuth angle corresponding to the last frame of the HOA signal, where j is greater than 0 and less than P-1.

[0162] It should be noted that in practical applications, the target distance may be less than or equal to the second distance threshold. In other words, the position of the first virtual speaker of interest in the target group of the HOA signal is not significantly different from the position of the second virtual speaker of interest in the reference group of the HOA signal. In some embodiments, the decoder-side device determines the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to P frames of the HOA signal, respectively. In other words, the elevation angle corresponding to each frame of the HOA signal in the P frames of the HOA signal is equal to the elevation angle of the first virtual speaker corresponding to the target distance, and the azimuth angle corresponding to each frame of the HOA signal is equal to the azimuth angle of the first virtual speaker corresponding to the target distance.

[0163] In some other embodiments, the decoder-side device determines the elevation and azimuth angles of the second virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the first L frames of the P frames of the HOA signal, and determines the elevation and azimuth angles of the first virtual speaker corresponding to the target distance as the elevation and azimuth angles corresponding to the remaining frames of the HOA signal of the P frames of the HOA signal, where L is an integer greater than or equal to 1 and L is less than P.

[0164] The second distance threshold is preset, and may be equal to or unequal to the first distance threshold, and may be adjusted based on different requirements.

[0165] (3) Virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles are determined as M target virtual speakers.

[0166] After determining the M groups of elevation angles and azimuth angles based on the N distances according to the above step (2), the decoder-side device determines the virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles as M target virtual speakers, and as a result, the M target virtual speakers will thereafter perform decoding processing on the target group of HOA signals.

[0167] Based on the above description, in practical applications, the attribute information of the virtual speakers may further include other content, such as the HOA coefficients of the virtual speakers. If the attribute information of the virtual speakers includes the HOA coefficients, the decoder-side device first needs to convert the HOA coefficients of the virtual speakers into the elevation angles and azimuth angles of the virtual speakers according to the relevant algorithm, and then determine the M target virtual speakers according to the above steps (1) to (3).

[0168] It should be noted that the above content is described using an example in which the decoder-side device determines M target virtual speakers in real time. In actual application, the encoder-side device determines the M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, and then further encodes the attribute information of the M target virtual speakers into a bitstream. In this way, after receiving the bitstream, the decoder-side device can directly analyze the bitstream to obtain the attribute information of the M target virtual speakers, and reconstruct the target group of the HOA signal based on the attribute information of the M target virtual speakers without determining the M target virtual speakers.

[0169] In this embodiment of the present application, the target virtual speaker is configured to process a target group of HOA signals, and the second virtual speaker is configured to process a reference group of HOA signals. The first virtual speaker is a virtual speaker that matches the target group of the HOA signals. Therefore, after the first virtual speaker is determined, the target virtual speaker is determined based on the attribute information of the second virtual speaker and the attribute information of the first virtual speaker to ensure that the attribute information of the target virtual speaker is not significantly different from the attribute information of the second virtual speaker, thereby solving the problem that two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump.

[0170] 7 is a diagram of the structure of a virtual speaker determination apparatus according to an embodiment of the present application. The virtual speaker determination apparatus may be implemented as part or all of a computer device by using software, hardware, or a combination thereof. The computer device may be the encoder-side device or decoder-side device mentioned above. See FIG. 7. The apparatus includes a first acquisition module 701, a second acquisition module 702, and a determination module 703.

[0171] The first acquisition module 701 is configured to acquire attribute information of N first virtual speakers, the N first virtual speakers being in a virtual speaker set and matching the HOA coefficients of a target group of the HOA signal, the target group of the HOA signal including at least one frame of the HOA signal, and N being an integer greater than or equal to 1. For detailed implementation processes, please refer to the corresponding content of the preceding embodiments. Details will not be repeated here.

[0172] The second acquisition module 702 is configured to acquire attribute information of N second virtual speakers, the N second virtual speakers being in a virtual speaker set and configured to process a reference group of HOA signals, the reference group of HOA signals being at least one group of HOA signals preceding the target group of HOA signals. For detailed implementation processes, please refer to the corresponding content of the preceding embodiments. Details will not be repeated here.

[0173] The determination module 703 determines M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, and the M target virtual speakers are configured to process a target group of HOA signals, where M is an integer greater than 1 and M is greater than N. For detailed implementation processes, please refer to the corresponding content of the foregoing embodiments. Details will not be repeated here.

[0174] Optionally, the attribute information includes elevation angles and azimuth angles, and the N first virtual speakers are in one-to-one correspondence with the N second virtual speakers.

[0175] The decision module 703 a first determination unit configured to determine distances between the first virtual speakers and the second virtual speakers having a correspondence relationship based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers to obtain N distances; a second determining unit configured to determine a group of M elevation angles and azimuth angles based on the N distances; a third determination unit configured to determine virtual speakers in the virtual speaker set corresponding to the M groups of elevation angles and azimuth angles as M target virtual speakers; Includes:

[0176] Optionally, the target group of the HOA signal includes one frame of the HOA signal, and one frame of the HOA signal includes H subframes, where H is an integer greater than 1, and M is a product of H and N.

[0177] The second decision unit is Using one of the N distances as the target distance, the elevation angles and azimuth angles corresponding to the H subframes are calculated by the following operations until each distance in the N distances is traversed: The method is particularly configured to determine, when the target distance is greater than a first distance threshold, by determining the elevation angles and azimuth angles corresponding to the H subframes, respectively, based on the elevation angles and azimuth angles of the first virtual speaker and the second virtual speaker corresponding to the target distance.

[0178] Optionally, the second decision unit is further configured to: determining an elevation angle and an azimuth angle of a second virtual loudspeaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first subframe of the H subframes; determining an elevation angle and an azimuth angle of a first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the last subframe of the H subframes; For an i-th subframe among the H subframes, determine the elevation angle and azimuth angle corresponding to the i-th subframe by an interpolation process based on the elevation angle and azimuth angle corresponding to the (i-1)-th subframe among the H subframes and the elevation angle and azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1; It is specifically configured so that

[0179] Optionally, the second decision unit is further configured to: When the target distance is equal to or less than a first distance threshold, determining the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the H subframes, respectively; or When the target distance is equal to or less than a first distance threshold, determine the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the first K subframes of the H subframes, and determine the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the remaining subframes of the H subframes, where K is an integer greater than or equal to 1 and less than H; It is more particularly configured as follows.

[0180] Optionally, the target group of the HOA signal includes P frames of the HOA signal, where P is an integer greater than 1, and M is the product of P and N.

[0181] The second decision unit is Using one distance among the N distances as the target distance, the elevation angles and azimuth angles corresponding to the P frames of the HOA signal are calculated by the following operation until each distance in the N distances is traversed: determining elevation angles and azimuth angles corresponding to the P frames of the HOA signal based on the elevation angles and azimuth angles of the first virtual speaker and the second virtual speaker corresponding to the target distance when the target distance is greater than a second distance threshold; It is specifically configured so that

[0182] Optionally, the second decision unit is further configured to: determining an elevation angle and an azimuth angle of a second virtual loudspeaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first frame of the HOA signal among the P frames of the HOA signal; determining an elevation angle and an azimuth angle of a first virtual loudspeaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a last frame of the HOA signal among the P frames of the HOA signal; For a j-th frame of the HOA signal among the P frames of the HOA signal, determine an elevation angle and an azimuth angle corresponding to the j-th frame of the HOA signal by an interpolation process based on the elevation angle and the azimuth angle corresponding to the (j-1)-th frame of the HOA signal among the P frames of the HOA signal and the elevation angle and the azimuth angle corresponding to the last frame of the HOA signal, where j is greater than 0 and less than P-1; It is specifically configured so that

[0183] Optionally, the second decision unit is further configured to: If the target distance is less than or equal to a second distance threshold, determine the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the P frames of the HOA signal, respectively; or When the target distance is less than or equal to a second distance threshold, determine the elevation angle and azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the first L frames of the HOA signal among the P frames of the HOA signal, and determine the elevation angle and azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to the remaining frames of the HOA signal among the P frames of the HOA signal, where L is an integer greater than or equal to 1 and L is less than P; It is more particularly configured as follows.

[0184] Optionally, the apparatus is applied in an encoder side device.

[0185] This device is a first encoding module configured to encode the attribute information of the M target virtual speakers into a bitstream; or a second encoding module configured to encode indices of the determination method of the M target virtual speakers into a bitstream; Further includes:

[0186] In this embodiment of the present application, the target virtual speaker is configured to process a target group of HOA signals, and the second virtual speaker is configured to process a reference group of HOA signals. The first virtual speaker is a virtual speaker that matches the target group of the HOA signals. Therefore, after the first virtual speaker is determined, the target virtual speaker is determined based on the attribute information of the second virtual speaker and the attribute information of the first virtual speaker to ensure that the attribute information of the target virtual speaker is not significantly different from the attribute information of the second virtual speaker, thereby solving the problem that two adjacent frames of the HOA signal obtained by decoding sound like they have a spatial jump.

[0187] It should be noted that when the virtual speaker determination device provided in the above embodiment determines a virtual speaker, the division of the above functional modules is merely used as an example for explanation. In actual application, the above functions may be assigned to and completed by different functional modules, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above functions. In addition, the virtual speaker determination device provided in the above embodiment relates to the same concept as the virtual speaker determination method embodiment. For specific implementation processes, please refer to the method embodiment. Details will not be repeated here.

[0188] 8 is a diagram of the structure of a computer device according to one embodiment of the present application. The computer device includes at least one processor 801, a communication bus 802, a memory 803, and at least one communication interface 804.

[0189] The processor 801 may be a general-purpose central processing unit (CPU), a network processor (NP), or a microprocessor, or may be one or more integrated circuits configured to implement the solutions of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0190] A communication bus 802 is used to transmit information between the aforementioned components. The communication bus 802 can be categorized as an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used to represent a bus in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0191] The memory 803 may be read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), an optical disk (including a compact disc read-only memory (CD-ROM), a compact disk, a laser disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be configured to carry or store expected program code in the form of instructions or data structures and that can be accessed by a computer. However, the memory 803 is not limited thereto. The memory 803 may exist independently or be connected to the processor 801 via the communication bus 802. Alternatively, the memory 803 and the processor 801 may be integrated together.

[0192] The communication interface 804 is configured to communicate with another device or a communication network by using any transceiver-type device. The communication interface 804 may include a wired communication interface or a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.

[0193] In a particular implementation, in one embodiment, processor 801 may include one or more CPUs, such as CPU 0 and CPU 1 shown in FIG.

[0194] In one specific implementation, a computing device may include multiple processors, such as processor 801 and processor 805 shown in FIG. 8. Each of the processors may be a single-core processor or a multi-core processor. A processor herein may be one or more devices, circuits, and / or processing cores configured to process data (e.g., computer program instructions).

[0195] In certain implementations, in an embodiment, the computer device may further include an output device and an input device. The output device may communicate with the processor 801 and display information in multiple ways. For example, the output device may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device may communicate with the processor 801 and receive input from a user in multiple ways. For example, the input device may be a mouse, a keyboard, a touchscreen device, a sensor device, etc.

[0196] In some embodiments, the memory 803 may be configured to store program code 810 for implementing the solution of the present application, and the processor 801 may execute the program code 810 stored in the memory 803. The program code 810 may include one or more software modules. By using the processor 801 and the program code 810 in the memory 803, the computing device may implement the virtual speaker determination method provided in the embodiments of FIGS. 5 and 6.

[0197] All or part of the above-described embodiments may be implemented by software, hardware, firmware, or any combination thereof. When software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the procedures or functions according to the embodiments of the present application are generated, in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio waves, or microwave) transmission. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device, such as a server or data center, incorporating one or more available media. The usable medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), a semiconductor medium (e.g., a solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium referred to in the embodiments of the present application may be a non-volatile storage medium, or that is, a non-transitory storage medium.

[0198] That is, an embodiment of the present application further provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to perform the steps of the virtual speaker determination method described above.

[0199] An embodiment of the present application further provides a computer program product including instructions, which, when executed on a computer, enable the computer to perform the steps of the virtual speaker determination method described above. In other words, a computer program is provided, which, when executed on a computer, enables the computer to perform the steps of the virtual speaker determination method.

[0200] It should be understood that "plurality" in this specification means two or more. In the description of the embodiments of the present application, " / " means "or" unless otherwise specified. For example, A / B may refer to A or B. In this specification, "and / or" simply describes an association relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may refer to the following three cases: when only A exists, when both A and B exist, and when only B exists. In addition, to clearly describe the technical solutions in the embodiments of the present application, terms such as "first" and "second" are used in the embodiments of the present application to distinguish between identical or similar items that basically provide the same function or purpose. Those skilled in the art will understand that terms such as "first" and "second" do not limit the number or execution order, and terms such as "first" and "second" do not indicate a clear distinction.

[0201] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals in the embodiments of the present application are used with authorization by the user or full authorization by all parties, and the capture, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, attribute information of virtual speakers in the embodiments of the present application is obtained when sufficient authorization is obtained.

[0202] The foregoing description is only an embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, or improvement made without departing from the spirit and principle of this application shall fall within the protection scope of this application. [Explanation of symbols]

[0203] 10 Source Device 20 destination device 30 Links 40 Storage device 100 Encoder 101 First Terminal 120 Data Sources 140 Output Interface 200 decoder 201 Second Terminal 220 Display device 240 input interface 701 First Acquisition Module 702 Second Acquisition Module 703 Decision Module 801 processor 802 communication bus 803 memory 804 communication interface 805 processor 810 Program Code

Claims

1. 1. A virtual speaker determination method, the method comprising: acquiring attribute information of N first virtual speakers, wherein the N first virtual speakers are virtual speakers in a virtual speaker set and match HOA coefficients of a target group of a higher-order Ambisonics HOA signal, the target group of HOA signals including at least one frame of the HOA signal, and N is an integer greater than or equal to 1; acquiring attribute information of N second virtual speakers, the N second virtual speakers being virtual speakers in a virtual speaker set and configured to process a reference group of HOA signals, the reference group of HOA signals being at least one group of HOA signals preceding the target group of HOA signals; determining M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, wherein the M target virtual speakers are configured to process the target group of HOA signals, where M is an integer greater than 1 and M is greater than N; A method comprising:

2. the attribute information includes an elevation angle and an azimuth angle, and the N first virtual speakers correspond one-to-one to the N second virtual speakers; determining the M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, determining distances between the corresponding first virtual speakers and the corresponding second virtual speakers based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers to obtain N distances; determining M groups of elevation and azimuth angles based on the N distances; determining virtual speakers in the virtual speaker set that correspond to the M groups of elevation angles and azimuth angles as the M target virtual speakers; 2. The method of claim 1, comprising:

3. The target group of HOA signals includes one frame of HOA signals, and the one frame of HOA signals includes H subframes, where H is an integer greater than 1, and M is a product of H and N; determining the M groups of elevation and azimuth angles based on the N distances, Using one of the N distances as a target distance, and performing the following operations on the elevation angles and azimuth angles corresponding to the H subframes respectively until each distance in the N distances is traversed: determining, when the target distance is greater than a first distance threshold, the elevation angles and the azimuth angles corresponding to the H subframes, respectively, based on the elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to the target distance; 3. The method of claim 2, comprising:

4. determining the elevation angles and the azimuth angles corresponding to the H subframes, respectively, based on the elevation angles and the azimuth angles of the first virtual speaker and the second virtual speaker corresponding to the target distance, determining the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first subframe of the H subframes; determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the last subframe of the H subframes; for an ith subframe among the H subframes, determining an elevation angle and an azimuth angle corresponding to the ith subframe by an interpolation process based on an elevation angle and an azimuth angle corresponding to an (i-1)th subframe among the H subframes and the elevation angle and the azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1; 4. The method of claim 3, comprising:

5. The method comprises: When the target distance is equal to or less than the first distance threshold, determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the H subframes, respectively; or determining, when the target distance is equal to or less than the first distance threshold, the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to first K subframes of the H subframes, and determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the remaining subframes of the H subframes, where K is an integer equal to or greater than 1 and is less than H; 4. The method of claim 3, further comprising:

6. the target group of HOA signals includes P frames of HOA signals, where P is an integer greater than 1, and M is the product of P and N; determining the M groups of elevation and azimuth angles based on the N distances, Using one of the N distances as a target distance, and calculating the elevation and azimuth angles corresponding to each of the P frames of the HOA signal by the following operation until each distance in the N distances is traversed: determining, when the target distance is greater than a second distance threshold, the elevation angles and the azimuth angles corresponding to the P frames of the HOA signal based on the elevation angles and the azimuth angles of a first virtual speaker and a second virtual speaker corresponding to the target distance; 3. The method of claim 2, comprising:

7. determining the elevation angles and the azimuth angles corresponding to the P frames of the HOA signal based on the elevation angles and the azimuth angles of the first virtual speaker and the second virtual speaker corresponding to the target distance, determining the elevation angle and the azimuth angle of the second virtual loudspeaker corresponding to the target distance as the elevation angle and azimuth angle corresponding to a first frame of an HOA signal among the P frames of an HOA signal; determining the elevation angle and the azimuth angle of the first virtual loudspeaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a last frame of an HOA signal among the P frames of an HOA signal; for a j-th frame of an HOA signal among the P frames of an HOA signal, determining an elevation angle and an azimuth angle corresponding to the j-th frame of an HOA signal by an interpolation process based on an elevation angle and an azimuth angle corresponding to a (j-1)-th frame of an HOA signal among the P frames of an HOA signal and the elevation angle and the azimuth angle corresponding to the last frame of an HOA signal, where j is greater than 0 and less than P-1; 7. The method of claim 6, comprising:

8. The method comprises: If the target distance is equal to or less than the second distance threshold, determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the P frames of the HOA signal, respectively; or if the target distance is less than or equal to the second distance threshold, determining the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first L frames of an HOA signal among the P frames of an HOA signal, and determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a remaining frame of an HOA signal among the P frames of an HOA signal, where L is an integer greater than or equal to 1 and L is less than P.

7. The method of claim 6, further comprising:

9. the method is applied to an encoder side device, After determining the M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, the method further comprises: Encoding the attribute information of the M target virtual speakers into a bitstream; or encoding an index of a method for determining the M target virtual speakers into the bitstream; 9. The method of claim 1, further comprising:

10. 1. A virtual speaker determination device, comprising: a first acquisition module configured to acquire attribute information of N first virtual speakers, the N first virtual speakers being virtual speakers in a virtual speaker set and matching HOA coefficients of a target group of HOA signals, the target group of HOA signals including at least one frame of the HOA signals, and N being an integer greater than or equal to 1; a second acquisition module configured to acquire attribute information of N second virtual speakers, the N second virtual speakers being virtual speakers in the virtual speaker set and configured to process a reference group of HOA signals, the reference group of HOA signals being at least one group of HOA signals preceding the target group of HOA signals; a determination module configured to determine M target virtual speakers based on the attribute information of the N first virtual speakers and the attribute information of the N second virtual speakers, the M target virtual speakers being configured to process the target group of HOA signals, where M is an integer greater than 1 and M is greater than N; An apparatus comprising:

11. the attribute information includes an elevation angle and an azimuth angle, and the N first virtual speakers correspond one-to-one to the N second virtual speakers; the decision module: a first determination unit configured to determine distances between the first virtual speakers and the second virtual speakers having a correspondence relationship based on the elevation angles and azimuth angles of the N first virtual speakers and the elevation angles and azimuth angles of the N second virtual speakers to obtain N distances; a second determining unit configured to determine M groups of elevation angles and azimuth angles based on the N distances; a third determination unit configured to determine virtual speakers in the virtual speaker set and corresponding to the M groups of elevation angles and azimuth angles as the M target virtual speakers; The apparatus of claim 10, comprising:

12. The target group of HOA signals includes one frame of HOA signals, and the one frame of HOA signals includes H subframes, where H is an integer greater than 1, and M is a product of H and N; The second determination unit: Using one of the N distances as a target distance, the elevation angles and azimuth angles corresponding to the H subframes are calculated by the following operation until each distance in the N distances is traversed: determining the elevation angles and the azimuth angles corresponding to the H subframes, respectively, based on the elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to the target distance when the target distance is greater than a first distance threshold; Specifically configured to:

12. The apparatus of claim 11.

13. The second determination unit: determining the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first subframe of the H subframes; determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the last subframe of the H subframes; For an i-th subframe among the H subframes, determining an elevation angle and an azimuth angle corresponding to the i-th subframe by an interpolation process based on an elevation angle and an azimuth angle corresponding to an (i-1)-th subframe among the H subframes and the elevation angle and the azimuth angle corresponding to the last subframe, where i is greater than 0 and less than H-1; 13. The device according to claim 12, specifically adapted to:

14. The second determination unit: When the target distance is equal to or less than the first distance threshold, determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the H subframes, respectively; or when the target distance is equal to or less than the first distance threshold, determine the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to first K subframes of the H subframes, and determine the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the remaining subframes of the H subframes, where K is an integer greater than or equal to 1 and less than H; 13. The apparatus of claim 12, further specifically configured to:

15. the target group of HOA signals includes P frames of HOA signals, where P is an integer greater than 1, and M is the product of P and N; The second determination unit: Using one of the N distances as a target distance, the elevation angles and azimuth angles corresponding to the P frames of the HOA signal are calculated by the following operation until each distance in the N distances is traversed: determining the elevation angles and the azimuth angles corresponding to the P frames of the HOA signal based on the elevation angles and azimuth angles of a first virtual speaker and a second virtual speaker corresponding to the target distance when the target distance is greater than a second distance threshold; Specifically configured to:

12. The apparatus of claim 11.

16. The second determination unit: determining the elevation angle and the azimuth angle of the second virtual loudspeaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first frame of an HOA signal among the P frames of an HOA signal; determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a last frame of an HOA signal among the P frames of an HOA signal; For a j-th frame of an HOA signal among the P frames of an HOA signal, determine an elevation angle and an azimuth angle corresponding to the j-th frame of the HOA signal by an interpolation process based on an elevation angle and an azimuth angle corresponding to a (j-1)-th frame of an HOA signal among the P frames of an HOA signal and the elevation angle and the azimuth angle corresponding to the last frame of the HOA signal, where j is greater than 0 and less than P-1; 16. The device according to claim 15, specifically adapted to:

17. The second determination unit: When the target distance is equal to or less than the second distance threshold, determining the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to the P frames of the HOA signal, respectively; or when the target distance is less than or equal to the second distance threshold, determine the elevation angle and the azimuth angle of the second virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a first L frames of the HOA signal among the P frames of the HOA signal, and determine the elevation angle and the azimuth angle of the first virtual speaker corresponding to the target distance as the elevation angle and the azimuth angle corresponding to a remaining frame of the HOA signal among the P frames of the HOA signal, where L is an integer greater than or equal to 1 and L is less than P; 16. The device of claim 15, further specifically configured to:

18. The apparatus is applied to an encoder side device, The device, a first encoding module configured to encode the attribute information of the M target virtual speakers into a bitstream; or a second encoding module configured to encode an index of a method for determining the M target virtual speakers into the bitstream; 18. Apparatus according to any one of claims 10 to 17.

19. 10. A computing device comprising a memory and a processor, the memory configured to store a computer program, and the processor configured to execute the computer program stored in the memory to perform the steps of the method of any one of claims 1 to 9.

20. 10. A computer-readable storage medium having stored thereon instructions that, when executed on a computer, enable the computer to perform the steps of the method of any one of claims 1 to 9.

21. 10. A computer program comprising instructions that, when executed on a computer, enable the computer to carry out the method of any one of claims 1 to 9.