Spatial audio rendering method and related device

By converting multi-channel and multi-object audio sources into FOA/HOA format audio signals and rendering them using rendering data expressions in the FOA/HOA domain, the high computational complexity and large storage requirements of existing technologies are solved, achieving a more natural and continuous spatial audio experience.

WO2026051478A1PCT designated stage Publication Date: 2026-03-12HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing technologies have high computational complexity and large storage requirements when rendering spatial audio from multi-channel and multi-object sources, resulting in unnatural sound and poor continuity, especially when the head is turned.

Method used

The input audio source is converted into an audio signal in FOA/HOA format using FOA/HOA format. The signal is then rendered using rendering data expressions in the FOA/HOA domain, which reduces computational complexity, preserves spatial information, and reduces storage requirements. Head motion information is used to process the rotation of the spatial audio signal.

Benefits of technology

While reducing computational complexity, it preserves the spatial information of the original input sound source, improves the naturalness and continuity of the listening experience in head motion tracking mode, and reduces memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025100277_12032026_PF_FP_ABST
    Figure CN2025100277_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio processing, and discloses a spatial audio rendering method and a related device. The method comprises: acquiring a spatial audio signal, wherein the spatial audio signal is an FOA audio signal or an HOA audio signal; acquiring spatial rendering data, wherein the spatial rendering data is FOA rendering data or HOA rendering data; and rendering the spatial audio signal on the basis of the spatial rendering data to obtain a binaural spatial audio output signal. In the present application, an input audio source is converted into an FOA / HOA format, so that all spatial information of the input audio source can be retained, and the number of channels and / or the number of sound objects of the spatial input audio source can also be converted into the number of channels of an audio signal in the FOA / HOA format. In this way, even if the number of channels and / or the number of sound objects of the spatial input audio source increases, the number of channels of the audio signal in the FOA / HOA format of the same order remains unchanged, thereby reducing the computational complexity of signal processing.
Need to check novelty before this filing date? Find Prior Art

Description

Spatial audio rendering method and related apparatus

[0001] The present application claims priority to the Chinese Patent Application No. 202411244874.1, filed on September 5, 2024, and entitled "Spatial audio rendering method and related apparatus", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of audio processing, in particular to a spatial audio rendering method and related apparatus. BACKGROUND

[0003] Spatial audio technology is widely used in products such as headphones, which simulates the characteristic that sound is ubiquitous in the real world, bringing users a real immersive three-dimensional audio experience, and becomes a key demand in audio-visual entertainment scenarios. Currently, the input sound sources of spatial audio mainly include multi-channel sound sources and / or multi-object sound sources, the multi-channel sound sources include sound sources of multiple channels, and the multi-object sound sources include sound sources of multiple sound emitting objects. How to render the multi-channel sound sources and / or multi-object sound sources to obtain a binaural spatial audio output signal has become one of the current concerns. SUMMARY

[0004] The present application provides a spatial audio rendering method and related apparatus, which can preserve all spatial information of the original input sound source while reducing the computational complexity of signal processing. The technical solution is as follows:

[0005] In a first aspect, a spatial audio rendering method is provided, the method comprising: obtaining a spatial audio signal, the spatial audio signal being obtained by converting a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including sound sources of multiple channels, the multi-object sound source including sound sources of multiple sound emitting objects, the spatial audio signal being a first order ambisonics (FOA) audio signal or a higher order ambisonics (HOA) audio signal; obtaining spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal; and rendering the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal.

[0006] The different formats of input sound sources are uniformly converted into FOA / HOA formats through stereo sound transformation, which not only can preserve all the spatial information of the original input sound sources, but also can convert the number of channels and / or the number of sound emitting objects of the spatial input sound sources into the number of channels of the FOA / HOA format audio signals, which is only related to the order of the audio signals and is irrelevant to the number of channels and / or the number of sound emitting objects of the spatial input sound sources. In this way, even if the number of channels and / or the number of sound emitting objects of the spatial input sound sources increases, the number of channels of the FOA / HOA format audio signals of the same order will remain unchanged, thereby reducing the computational complexity of subsequent signal processing.

[0007] When the order of the spatial audio signal is equal to 1, the FOA rendering data can be determined based on the order of the spatial audio signal through the expression of the spatial rendering data in the FOA / HOA domain stored in the database. When the order of the spatial audio signal is greater than 1, the HOA rendering data can be determined based on the order of the spatial audio signal through the expression of the spatial rendering data in the FOA / HOA domain stored in the database. That is, the expression of the spatial rendering data in the FOA / HOA domain is stored in the database, and the HOA rendering data can be determined through the expression. In this way, the database does not need to store discrete spatial rendering data of each direction, greatly reducing the occupation of memory space.

[0008] After obtaining the spatial audio signal, the spatial audio signal is rendered through the spatial rendering data once, which increases the processing freedom without increasing the rendering computational complexity.

[0009] In a possible implementation, the rendering process of the spatial audio signal can be deployed on an audio processing terminal or a binaural playback terminal, that is, the spatial audio rendering method provided in the present application can be applied to an audio processing terminal or a binaural playback terminal. If the spatial audio rendering method provided in the present application is applied to an audio processing terminal, the audio processing terminal can obtain a spatial input sound source, perform stereo sound transformation on the spatial input sound source, and obtain a spatial audio signal. If the spatial audio rendering method provided in the present application is applied to a binaural playback terminal, the binaural playback terminal can receive a spatial audio signal sent by an audio processing terminal, and the spatial audio signal is obtained by performing stereo sound transformation on a spatial input sound source by the audio processing terminal.

[0010] The process of stereo sound transformation of the audio processing terminal on the spatial input sound source is introduced below.

[0011] In a possible implementation, the stereo sound transformation is performed on the spatial input sound source to obtain the target direct sound signal, and the spatial audio signal is determined based on the target direct sound signal.

[0012] The input sound source in different formats is converted into the FOA / HOA format through the stereo sound transformation, which not only preserves all the spatial information of the original input sound source, but also converts the number of sound channels and / or the number of sound objects of the spatial input sound source into the number of channels of the FOA / HOA format audio signal, which is only related to the order of the audio signal and is irrelevant to the number of sound channels and / or the number of sound objects of the spatial input sound source. In this way, even if the number of sound channels and / or the number of sound objects of the spatial input sound source increases, the number of channels of the FOA / HOA format audio signal of the same order remains unchanged, thereby reducing the computational complexity of subsequent signal processing.

[0013] In a possible implementation, the stereo sound transformation is performed on the spatial input sound source to obtain the target direct sound signal based on the order of the stereo sound transformation corresponding to the spatial input sound source, and the spatial audio signal is determined based on the target direct sound signal.

[0014] For each group of input sound sources in the at least one group of input sound sources, if the order of the stereo sound transformation corresponding to the group of input sound sources is equal to 1, the FOA transformation is performed on the group of input sound sources to obtain the corresponding direct sound signal, and the direct sound signal is a FOA signal. If the order of the stereo sound transformation corresponding to the group of input sound sources is greater than 1, the HOA transformation is performed on the group of input sound sources to obtain the corresponding direct sound signal, and the direct sound signal is a HOA signal.

[0015] It should be noted that the order of the stereo sound transformation refers to the order of the stereo sound transformation and the order of the direct sound signal obtained after the stereo sound transformation. The orders of the stereo sound transformation corresponding to the at least one group of input sound sources can be the same or different.

[0016] In a possible implementation, the spatial input sound sources are grouped to obtain at least one group of input sound sources, including: if the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are taken as one group of input sound sources, and the multi-object sound sources are taken as one group of input sound sources; or, the spatial input sound sources are grouped according to a sound source feature to obtain the at least one group of input sound sources, where the sound source feature is used to represent a characteristic of a sound source.

[0017] In a case where the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are taken as one group of input sound sources, and the multi-object sound sources are taken as one group of input sound sources, which is simple in operation and high in grouping efficiency. According to the sound source feature, the spatial input sound sources can be grouped into the at least one group of input sound sources without considering whether the multi-channel sound sources and the multi-object sound sources exist simultaneously, so that subsequent processing can be performed on each group of input sound sources respectively, the degree of freedom of processing is increased, and the flexibility of processing of the input sound sources is high.

[0018] In a possible implementation, the sound source feature includes at least one of a sound source direction, a sound source distance, and sound source content, where the sound source content is used to indicate a type of a sound emitting object in a corresponding sound source.

[0019] For example, the type of the sound emitting object can be a human, an animal, an environment, a musical instrument, or the like. Of course, the type of the sound emitting object can also be an adult, a child, a male, a female, a type of animal, a type of musical instrument, or the like.

[0020] In a possible implementation, before the at least one group of input sound sources is respectively subjected to stereo sound transformation based on a stereo sound transformation order corresponding to each group of input sound sources, the method further includes: displaying a first parameter configuration interface; and obtaining the stereo sound transformation order corresponding to each group of input sound sources from the first parameter configuration interface.

[0021] That is, the stereo sound transformation order corresponding to each group of input sound sources is set in the first parameter configuration interface.

[0022] As an example, the first parameter configuration interface includes a default transformation order corresponding to each group of input sound sources. For each group of input sound sources, if an adjustment operation of the default transformation order corresponding to the group of input sound sources is detected, an adjusted transformation order is determined as the stereo sound transformation order corresponding to the group of input sound sources, and if the adjustment operation of the default transformation order corresponding to the group of input sound sources is not detected, the default transformation order is determined as the stereo sound transformation order corresponding to the group of input sound sources.

[0023] That is, the user can adjust the default transform order corresponding to the group of input sound sources, or can not adjust the default transform order corresponding to the group of input sound sources. If the user adjusts the default transform order corresponding to the group of input sound sources, the adjusted transform order is determined as the stereo sound transform order corresponding to the group of input sound sources. If the user does not adjust the default transform order corresponding to the group of input sound sources, the default transform order corresponding to the group of input sound sources is directly determined as the stereo sound transform order corresponding to the group of input sound sources.

[0024] The determination of the stereo sound transform order corresponding to each of the at least one group of input sound sources through the first parameter configuration interface can support customization of the stereo sound transform order, thereby improving the flexibility of autonomous selection of the stereo sound transform order. Moreover, the display of the default transform order corresponding to each of the at least one group of input sound sources in the first parameter configuration interface can provide a basis for adjustment of the transform order for the user, so that the user can determine the stereo sound transform order corresponding to each of the at least one group of input sound sources through simple adjustment operations, thereby improving the determination efficiency of the stereo sound transform order.

[0025] In a possible implementation, before superimposing the at least one direct sound signal, the method further includes: performing audio effect processing on the at least one direct sound signal based on the direct sound audio effect processing parameter corresponding to each of the at least one group of input sound sources.

[0026] In this way, after superimposing the at least one direct sound signal after audio effect processing, the realism of the target direct sound signal can be improved.

[0027] In a possible implementation, the method further includes: displaying a second parameter configuration interface; and obtaining the direct sound audio effect processing parameter corresponding to each of the at least one group of input sound sources from the second parameter configuration interface.

[0028] That is, the direct sound audio effect processing parameter corresponding to each of the at least one group of input sound sources is set in the second parameter configuration interface.

[0029] As an example, the second parameter configuration interface includes the default audio effect processing parameter corresponding to each of the at least one group of input sound sources. For each group of input sound sources, if an adjustment operation of the default audio effect processing parameter corresponding to the group of input sound sources is detected, the adjusted audio effect processing parameter is determined as the direct sound audio effect processing parameter corresponding to the group of input sound sources, and if the adjustment operation of the default audio effect processing parameter corresponding to the group of input sound sources is not detected, the default audio effect processing parameter is determined as the direct sound audio effect processing parameter corresponding to the group of input sound sources.

[0030] That is, the user can adjust the default sound effect processing parameters corresponding to the group of input sound sources, or can not adjust the default sound effect processing parameters corresponding to the group of input sound sources. If the user adjusts the default sound effect processing parameters corresponding to the group of input sound sources, the adjusted sound effect processing parameters are determined as the direct sound effect processing parameters corresponding to the group of input sound sources. If the user does not adjust the default sound effect processing parameters corresponding to the group of input sound sources, the default sound effect processing parameters corresponding to the group of input sound sources are directly determined as the direct sound effect processing parameters corresponding to the group of input sound sources.

[0031] By determining the direct sound effect processing parameters corresponding to at least one group of input sound sources through the second parameter configuration interface, customization of the sound effect processing parameters can be supported, thereby improving the flexibility of autonomous selection of the sound effect processing parameters. Moreover, the default sound effect processing parameters corresponding to at least one group of input sound sources are displayed in the second parameter configuration interface, and the user can be provided with a basis for adjusting the sound effect processing parameters through the default sound effect processing parameters, so that the user can determine the direct sound effect processing parameters corresponding to at least one group of input sound sources through simple adjustment operations, thereby improving the determination efficiency of the direct sound effect processing parameters.

[0032] In a possible implementation, the method further includes: performing reverberation processing on the spatial input sound source based on reverberation data corresponding to the spatial input sound source to obtain a target reverberation sound signal, the reverberation data being FOA reverberation data or HOA reverberation data; and determining the spatial audio signal based on the target direct sound signal, including: superimposing the target direct sound signal and the target reverberation sound signal to obtain the spatial audio signal.

[0033] In this way, the target direct sound signal and the target reverberation sound signal can be superimposed to obtain the spatial audio signal, so that the spatial audio signal can bring a more realistic and immersive spatial audio experience.

[0034] In a possible implementation, the spatial input sound source is divided into at least one group of input sound sources; and the reverberation processing on the spatial input sound source based on the reverberation data corresponding to the spatial input sound source to obtain a target reverberation sound signal includes: performing reverberation processing on the at least one group of input sound sources based on reverberation data corresponding to the at least one group of input sound sources, respectively, to obtain at least one reverberation sound signal; and superimposing the at least one reverberation sound signal to obtain the target reverberation sound signal.

[0035] The reverberation data is FOA reverberation data or HOA reverberation data. For each group of input sound sources, if the reverberation data corresponding to the group of input sound sources is FOA reverberation data, the group of input sound sources is subjected to reverberation processing based on the reverberation data corresponding to the group of input sound sources, and a corresponding reverberation sound signal is an FOA signal. If the reverberation data corresponding to the group of input sound sources is HOA reverberation data, the group of input sound sources is subjected to reverberation processing based on the reverberation data corresponding to the group of input sound sources, and a corresponding reverberation sound signal is an HOA signal.

[0036] It should be noted that the reverberation data refers to data for converting input sound sources into reverberation sound signals, and the reverberation data corresponding to at least one group of input sound sources can be the same or different.

[0037] In a possible implementation, before the at least one group of input sound sources is subjected to reverberation processing based on the reverberation data corresponding to the at least one group of input sound sources respectively, the method further includes: obtaining the reverberation data corresponding to the at least one group of input sound sources respectively based on the reverberation parameters corresponding to the at least one group of input sound sources respectively.

[0038] In a possible implementation, the reverberation parameters include a reverberation sound rendering order and a room type, and the room type is used to indicate a type of a room currently simulated by the reverberation sound. The reverberation data corresponding to the at least one group of input sound sources respectively is obtained based on the reverberation sound rendering order and the room type corresponding to the at least one group of input sound sources respectively.

[0039] The reverberation sound rendering order refers to an order of the reverberation data and an order of the reverberation sound signal obtained after the reverberation processing. The reverberation sound rendering orders corresponding to the at least one group of input sound sources respectively can be the same or different. In addition, for each group of input sound sources, a reverberation sound rendering order corresponding to the group of input sound sources is less than or equal to a binaural response conversion order corresponding to the group of input sound sources. That is, an order of a reverberation sound signal corresponding to the group of input sound sources is less than or equal to an order of a direct sound signal corresponding to the group of input sound sources.

[0040] In a possible implementation, the reverberation parameters further include a spatial sense parameter and / or a room size, and the spatial sense parameter is used to describe spatial information in the reverberation sound. After the reverberation data corresponding to the at least one group of input sound sources respectively is obtained based on the reverberation sound rendering order and the room type corresponding to the at least one group of input sound sources respectively, the method further includes: adjusting the reverberation data corresponding to the at least one group of input sound sources respectively based on the spatial sense parameter and / or the room size corresponding to the at least one group of input sound sources respectively.

[0041] If the reverberation parameter comprises a spaciousness parameter, for each of the at least one group of input sound sources, a direct-mix ratio corresponding to the group of input sound sources is obtained from a correspondence between stored spaciousness parameters and direct-mix ratios based on the spaciousness parameter corresponding to the group of input sound sources, and the reverberation data corresponding to the group of input sound sources is adjusted according to the direct-mix ratio corresponding to the group of input sound sources. For example, the reverberation data corresponding to the group of input sound sources is multiplied by the inverse of the direct-mix ratio corresponding to the group of input sound sources to obtain adjusted reverberation data of the group of input sound sources.

[0042] The spaciousness parameter is used to describe the spatial information in the reverberation sound, that is, the spaciousness parameter is used to describe the spatial component in the reverberation sound, so that the audio signal sounds with spaciousness. In other words, the spaciousness parameter refers to the listening feeling of the indoor sound field currently simulated by the reverberation sound. For example, if the spaciousness parameter is smaller, it means that the listening feeling and spaciousness of the indoor sound field currently simulated by the reverberation sound are weaker, and if the spaciousness parameter is larger, it means that the listening feeling and spaciousness of the indoor sound field currently simulated by the reverberation sound are stronger. The spaciousness parameter has a correspondence with the direct-mix ratio, different spaciousness parameters correspond to different direct-mix ratios, and the energy ratio of the direct sound signal and the reverberation signal in the audio signal can be adjusted by the direct-mix ratio, so as to adjust the listening feeling of the indoor sound field, so that the listening feeling of the indoor sound field is more realistic.

[0043] It should be noted that adjusting the reverberation data by the direct-mix ratio, and then adjusting the spaciousness of the audio signal is an example, that is, the direct-mix ratio is one of the parameters for adjusting the spaciousness. In actual application, the spaciousness of the audio signal can also be adjusted by other ways, that is, the parameters for adjusting the spaciousness can also include other parameters in addition to the direct-mix ratio, and the adjustment of the spaciousness is realized by adjusting the reverberation data by other parameters, which is not limited by the embodiments of the present application.

[0044] In a possible implementation, the at least one group of input sound sources is divided by a sound source feature; and the obtaining of the reverberation data corresponding to the at least one group of input sound sources based on the reverberation sound rendering order and the room type corresponding to the at least one group of input sound sources comprises: obtaining the reverberation data corresponding to the at least one group of input sound sources based on the reverberation sound rendering order, the room type and the sound source feature corresponding to the at least one group of input sound sources.

[0045] In a possible implementation, before the rendering of the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal, the method further comprises: obtaining head movement information, the head movement information indicating a posture of a head wearing a binaural playback terminal; and performing rotation processing on the spatial audio signal based on the head movement information. That is, when the binaural playback terminal is in a head movement tracking mode, the spatial audio signal can also be processed based on the posture of the head.

[0046] As an example, the head movement information comprises Euler angles, which comprise a nutation angle, a precession angle and a spin angle. A rotation matrix is determined by the nutation angle, the precession angle and the spin angle. The rotation matrix is multiplied with the spatial audio signal to achieve the rotation processing of the spatial audio signal.

[0047] In a second aspect, a spatial audio rendering method is provided, which is applied to an audio processing terminal. The method comprises: obtaining a spatial input sound source, the spatial input sound source comprising a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source comprising sound sources of multiple channels, and the multi-object sound source comprising sound sources of multiple sound emitting objects; performing a binaural response conversion on the spatial input sound source to obtain a spatial audio signal, the spatial audio signal being a first-order binaural response (FOA) audio signal or a higher-order binaural response (HOA) audio signal; and sending the spatial audio signal to a binaural playback terminal, so that the binaural playback terminal renders the spatial audio signal based on spatial rendering data to obtain a binaural spatial audio output signal; wherein the spatial rendering data is FOA rendering data or HOA rendering data, and the order of the spatial rendering data is the same as the order of the spatial audio signal.

[0048] In a third aspect, a spatial audio rendering apparatus is provided, which has the function of implementing the method behavior of the first aspect. The spatial audio rendering apparatus comprises at least one module for implementing the method provided in the first aspect.

[0049] In a fourth aspect, a spatial audio rendering apparatus is provided, which has the function of implementing the method behavior of the second aspect. The spatial audio rendering apparatus comprises at least one module for implementing the method provided in the second aspect.

[0050] In a fifth aspect, an electronic device is provided, which comprises a processor and a memory, the memory being configured to store a computer program for executing the method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the method of the first aspect.

[0051] Optionally, the electronic device can further comprise a communication bus for establishing a connection between the processor and the memory.

[0052] In a sixth aspect, an audio processing terminal is provided, which comprises a processor and a memory, the memory being configured to store a computer program for executing the method provided in the second aspect. The processor is configured to execute the computer program stored in the memory to implement the method of the second aspect.

[0053] Optionally, the electronic device further comprises a communication bus for establishing a connection between the processor and the memory.

[0054] In a seventh aspect, a computer-readable storage medium is provided, and the storage medium stores instructions therein, when the instructions are run on a computer, the computer is caused to perform the method in the first aspect or the second aspect.

[0055] In an eighth aspect, a computer program product is provided, and the computer program product comprises instructions, when the instructions are run on a computer, the computer is caused to perform the method in the first aspect or the second aspect.

[0056] The technical effects obtained by the second aspect to the eighth aspect are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0057] FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0058] FIG. 2 is a schematic diagram of another implementation environment provided by an embodiment of the present application;

[0059] FIG. 3 is a schematic diagram of another implementation environment provided by an embodiment of the present application;

[0060] FIG. 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0061] FIG. 5 is a flowchart of a spatial audio rendering method provided by an embodiment of the present application;

[0062] FIG. 6 is a schematic diagram of a first parameter configuration interface provided by an embodiment of the present application;

[0063] FIG. 7 is a schematic diagram of another first parameter configuration interface provided by an embodiment of the present application;

[0064] FIG. 8 is a schematic diagram of another first parameter configuration interface provided by an embodiment of the present application;

[0065] FIG. 9 is a schematic diagram of another first parameter configuration interface provided by an embodiment of the present application;

[0066] FIG. 10 is a schematic diagram of a second parameter configuration interface provided by an embodiment of the present application;

[0067] FIG. 11 is a schematic diagram of a parameter configuration interface provided by an embodiment of the present application;

[0068] FIG. 12 is a schematic diagram of a third parameter configuration interface provided by an embodiment of the present application;

[0069] FIG. 13 is a schematic diagram of another parameter configuration interface according to an embodiment of the present application;

[0070] FIG. 14 is a block diagram of a spatial audio rendering method according to an embodiment of the present application;

[0071] FIG. 15 is a block diagram of another spatial audio rendering method according to an embodiment of the present application;

[0072] FIG. 16 is a schematic diagram of a distance of a spatial input sound source according to an embodiment of the present application;

[0073] FIG. 17 is a schematic diagram of a structure of a spatial audio rendering device according to an embodiment of the present application;

[0074] FIG. 18 is a schematic diagram of a structure of another spatial audio rendering device according to an embodiment of the present application. DETAILED DESCRIPTION

[0075] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0076] Before the spatial audio rendering method provided by the embodiments of the present application is explained in detail, the terms, application scenarios and implementation environments related to the embodiments of the present application are introduced.

[0077] First, the terms related to the embodiments of the present application are introduced.

[0078] 1. Multi-channel sound source

[0079] The multi-channel sound source includes a sound source of multiple channels. Usually, the positions of the multiple channels are according to a standard layout. For example, 5.1, 7.1, 5.1.4, 7.1.4, 9.1.6, etc. are all multi-channel sound sources, and the positions of each channel in these multi-channel sound sources are all according to a standard layout.

[0080] 2. Multi-object sound source

[0081] The multi-object sound source includes a sound source of multiple sound emitting objects. The metadata of each sound emitting object is used to describe sound source related information of the sound emitting object, such as position information, sound source size, etc. The sound emitting object can be a person, an animal, an environment, a musical instrument, etc. The person can be an adult, a child, a male, a female, etc. The animal can be a tiger, an elephant, a cat, etc. The musical instrument can be a keyboard, a saxophone, a guitar, etc. The sound source size of the sound emitting object can be the sound emitting area of the sound emitting object.

[0082] 3. First order ambisonics (FOA) and higher order ambisonics (HOA)

[0083] Ambisonics technology is a spatial audio method capable of presenting the entire sound field. The audio signal obtained by the ambisonics technology can be referred to as a spatial audio signal, and the spatial audio signal includes a plurality of channel signals, and the number of channels of the spatial audio signal is related to the order of the spatial audio signal. That is, the number of channels of the spatial audio signal is (X+1) 2 , X is the order of the spatial audio signal.

[0084] Generally, the spatial audio signal includes FOA audio signal and HOA audio signal. The FOA audio signal is a first-order signal, and the number of channels of the FOA audio signal is (1+1) 2 =4. The HOA audio signal is a high-order signal, and assuming that the HOA audio signal is an A-order signal, the number of channels of the HOA audio signal is (A+1) 2 .

[0085] 4, direct sound signal, reverberation sound signal

[0086] The sound wave signal directly radiated by the sound source to the indoor space without any reflection is referred to as a direct sound signal. Reverberation is an acoustic phenomenon in a room. After the sound is emitted by the sound source, it propagates in the air, and during the propagation process, it produces reflection, absorption, diffusion, transmission, interference and diffraction on the interface of the room, forming a complex indoor sound field, and people have a reverberation feeling. Under the condition that the indoor sound field reaches stability, the sound continues due to multiple reflections or scattering after the sound source stops emitting sound, which is called reverberation. Therefore, the superposition of the sound wave signal emitted by the sound source after one or more reflections is referred to as a reverberation sound signal.

[0087] 5, equalizer (equalizer, EQ)

[0088] The EQ is an electronic device for adjusting audio signals in various frequency ranges such as high frequency, medium frequency and low frequency. By adjusting the audio signals in different frequency ranges, the quality of the sound is improved.

[0089] 6, direct-to-reverberation ratio

[0090] The ratio between the direct sound signal and the reverberation sound signal in the audio signal. It can be referred to as an energy ratio, or an intensity ratio.

[0091] Secondly, the application scenarios related to the embodiments of the present application are introduced.

[0092] The spatial audio technology is applied to the binaural playback terminal such as earphones, and through simulating the characteristic that sound is ubiquitous in the real world, a real immersive three-dimensional audio experience is brought to the user, and in the audio-visual entertainment scene, it becomes a key demand.

[0093] At present, mainstream formats of spatial input sound sources (i.e., input sound sources of spatial audio) include multi-channel sound sources and multi-object sound sources. A spatial audio rendering method provided by the related art includes: performing binaural rendering on each channel in a multi-channel sound source to obtain an audio output signal after binaural rendering of each channel; and performing binaural rendering on each sound emitting object in a multi-object sound source to obtain an audio output signal after binaural rendering of each sound emitting object. Then, the audio output signal after left ear rendering of each channel and the audio output signal after left ear rendering of each sound emitting object are superimposed to obtain a left ear spatial audio output signal, and the audio output signal after right ear rendering of each channel and the audio output signal after right ear rendering of each sound emitting object are superimposed to obtain a right ear spatial audio output signal.

[0094] Before performing binaural rendering on the input sound source of each channel in the multi-channel sound source, a corresponding binaural head related impulse response (HRIR) / binaural room impulse response (BRIR) of each channel needs to be obtained, and then the input sound source of each channel is binaurally rendered through the corresponding binaural HRIR / BRIR of each channel. For example, for any one channel, the corresponding binaural HRIR / BRIR of the channel is obtained from the correspondence between the orientation and the binaural HRIR / BRIR stored in the database through the orientation of the channel, and then the input sound source of the channel is convolved with the binaural HRIR / BRIR of the channel to obtain an audio output signal after binaural rendering of the channel. In some embodiments, if the binaural playback terminal is in a head tracking mode, the relative orientation between the orientation of the channel and the head pose is determined, and the corresponding binaural HRIR / BRIR of the channel is obtained from the correspondence between the orientation and the binaural HRIR / BRIR stored in the database through the relative orientation.

[0095] Similarly, before binaural rendering is performed on the input sound source of each sound emitting object in the multi-object sound source, the binaural HRIR / BRIR corresponding to each sound emitting object also needs to be obtained, and then the input sound source of each sound emitting object is binaurally rendered through the binaural HRIR / BRIR corresponding to each sound emitting object. For example, for any sound emitting object, the binaural HRIR / BRIR corresponding to the sound emitting object is obtained from the correspondence between the orientation and the binaural HRIR / BRIR stored in the database through the orientation of the sound emitting object, and then the input sound source of the sound emitting object is convolved with the binaural HRIR / BRIR of the sound emitting object to obtain the audio output signal after binaural rendering of the sound emitting object. In some embodiments, if the binaural playback terminal is in the head motion tracking mode, the relative orientation between the orientation of the sound emitting object and the head pose is determined, and the binaural HRIR / BRIR corresponding to the sound emitting object is obtained from the correspondence between the orientation and the binaural HRIR / BRIR stored in the database through the relative orientation. The orientation of each sound emitting object is determined by the position information in the metadata of the sound emitting object.

[0096] Therefore, this spatial audio rendering scheme needs to store discrete binaural HRIR / BRIRs that can cover the entire three-dimensional space in the database to provide binaural HRIR / BRIRs at any orientation, which occupies a large amount of memory space, and the discrete binaural HRIR / BRIRs differ at different orientations, resulting in a non-smooth audio output signal during head rotation, and thus causing unnatural and discontinuous listening in the head motion tracking mode. In addition, in the scenario where the number of channels / sound emitting objects of the input sound source is large, the above rendering method of processing each channel / sound emitting object increases the computational complexity, so the related technology also provides two schemes to reduce the number of channels. One is to use downmixing to reduce the number of channels, for example, in a scenario where there are multi-channel sound sources and multi-object sound sources, a vector-based amplitude panning (VBAP) method is used to map the multi-object sound source to the multi-channel sound source to reduce the total processing channel number. The other is to combine channels with similar orientations into one channel for spatial audio rendering, and to combine sound emitting objects with similar orientations into one sound emitting object for spatial audio rendering. Whether it is downmix processing or channel combination processing, the spatial information of the original input sound source is lost.

[0097] Ambisonics is a sound format covering the whole globe, FOA is a first-order representation of Ambisonics, the order of FOA corresponds to the number of channels of FOA, HOA is a high-order representation of Ambisonics, and the order of HOA also corresponds to the number of channels of HOA. Therefore, the embodiment of the present application provides a spatial audio rendering method, which converts spatial input sound sources into FOA / HOA format, thereby representing input sound sources in different formats, and also preserves the spatial information of the input sound sources, and at the same time, the order is used to control the number of channels of the FOA / HOA format audio signal (which is no longer related to the number of sound channels / number of sound emitting objects of the input sound source), thereby realizing the premise of reducing the calculation complexity under the premise of preserving all the spatial information of the original input sound source in the scene of a large number of sound channels / sound emitting objects. In addition, the rendering of the FOA / HOA format audio signal does not need to store discrete binaural HRIR / BRIR of each direction in the database, but stores binaural HRIR / BRIR in the FOA / HOA domain. The binaural HRIR / BRIR in the FOA / HOA domain is a distributed expression of binaural HRIR / BRIR in three-dimensional space, and through the expression, the binaural HRIR / BRIR of any direction in three-dimensional space can be determined. That is, the database does not need to store discrete binaural HRIR / BRIR of each direction, but store the expression of binaural HRIR / BRIR in the FOA / HOA domain, which greatly reduces the occupation of memory space. Moreover, the binaural HRIR / BRIR in the FOA / HOA domain is a continuous expression in three-dimensional space, and after determining the binaural HRIR / BRIR through the expression, the audio output signal rendered in the process of head rotation is relatively smooth, thereby improving the naturalness and continuity of the listening experience in the head tracking mode.

[0098] Finally, the implementation environment related to the embodiments of the present application is introduced.

[0099] Please refer to FIG. 1, which is a schematic diagram of an implementation environment according to an embodiment of the present application. The implementation environment includes an audio processing terminal 101 and a binaural playback terminal 102. The audio processing terminal 101 and the binaural playback terminal 102 are in communication connection, which can be wired connection or wireless connection, and the present application does not limit this.

[0100] In some embodiments, referring to FIG. 2, the audio processing terminal 101 is configured to obtain a spatial audio signal, the spatial audio signal being converted from a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the spatial audio signal being a FOA audio signal or a HOA audio signal; obtain spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal; render the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal. Then, the binaural spatial audio output signal is sent to the binaural playback terminal 102 through short-distance transmission. The binaural playback terminal 102 is configured to play back the binaural spatial audio output signal.

[0101] In some other embodiments, referring to FIG. 3, the audio processing terminal 101 is configured to obtain a spatial audio signal, the spatial audio signal being converted from a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the spatial audio signal being a FOA audio signal or a HOA audio signal. Then, the spatial audio signal is sent to the binaural playback terminal 102 through short-distance transmission. The binaural playback terminal 102 is configured to obtain spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal; render the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal, and play back the binaural spatial audio output signal.

[0102] As can be seen from FIGS. 2 and 3, the rendering process of the spatial audio signal can be deployed on the audio processing terminal 101 or the binaural playback terminal 102.

[0103] It should be noted that the audio processing terminal 101 can be any electronic device capable of processing spatial audio, such as a personal computer (PC), a mobile phone, a smart phone, a personal digital assistant (PDA), a wearable device, a pocket PC (PPC), a tablet computer, a smart car, a smart television, a smart speaker, etc. The binaural playback terminal 102 can be any electronic device capable of playing a binaural audio output signal, such as a headset, glasses, etc.

[0104] Those skilled in the art should understand that the audio processing terminal 101 and the binaural playback terminal 102 described above are only examples, and other existing or future audio processing terminals or binaural playback terminals, such as those applicable to the embodiments of the present application, should also be included in the protection scope of the embodiments of the present application, and are hereby incorporated by reference.

[0105] Referring to FIG. 4, FIG. 4 is a structural schematic diagram of an electronic device according to an embodiment of the present application. The electronic device can be the audio processing terminal or the binaural playback terminal shown in the above. The electronic device includes a processor 401, a communication bus 402, a memory 403, and at least one communication interface 404.

[0106] The processor 401 can be a general central processing unit (CPU), a network processing unit (NP), a microprocessor, or can be one or more integrated circuits used to implement the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0107] The communication bus 402 is used to transmit information between the above components. The communication bus 402 can be divided into an address bus, a data bus, a control bus, and the like. For the sake of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or only one type of bus.

[0108] The memory 403 can be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disk (including a compact disc read-only memory (CD-ROM), a compact disc, a laser disk, a digital versatile disc, a Blu-ray disc, and the like), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory 403 can exist independently and be connected to the processor 401 through the communication bus 402. The memory 403 can also be integrated with the processor 401.

[0109] The communication interface 404 uses any transceiver-like mechanism for communicating with other devices or communication networks. The communication interface 404 includes a wired communication interface and can also include a wireless communication interface. The wired communication interface can be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface can be, for example, a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof.

[0110] In some embodiments, the memory 403 is configured to store program code 410 for implementing the solutions of the present application, and the processor 401 can execute the program code 410 stored in the memory 403. The program code 410 can include one or more software modules, and the electronic device can implement the spatial audio rendering method provided in the embodiments below by means of the processor 401 and the program code 410 in the memory 403.

[0111] It should be noted that the application scenarios and implementation environments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It can be known by those skilled in the art that, with the appearance of new application scenarios and the evolution of implementation environments, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0112] Next, the spatial audio rendering method provided by the embodiments of the present application will be explained in detail.

[0113] FIG. 5 is a flowchart of a spatial audio rendering method provided by an embodiment of the present application. The method can be applied to an audio processing terminal or a binaural playback terminal. Please refer to FIG. 5, the method includes the following steps.

[0114] Step 501: Obtain a spatial audio signal, the spatial audio signal is obtained by converting a spatial input sound source, the spatial input sound source includes a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source includes a sound source of multiple channels, the multi-object sound source includes a sound source of multiple sound emitting objects, and the spatial audio signal is a FOA audio signal or a HOA audio signal.

[0115] Based on the above description, the rendering process of the spatial audio signal can be deployed on an audio processing terminal or a binaural playback terminal, i.e., the spatial audio rendering method provided in the embodiments of the present application can be applied to an audio processing terminal or a binaural playback terminal. If the spatial audio rendering method provided in the embodiments of the present application is applied to an audio processing terminal, the audio processing terminal can obtain a spatial input sound source, perform stereo ambisonic transformation on the spatial input sound source, and obtain a spatial audio signal. If the spatial audio rendering method provided in the embodiments of the present application is applied to a binaural playback terminal, the binaural playback terminal can receive a spatial audio signal sent by an audio processing terminal, and the spatial audio signal is obtained by performing stereo ambisonic transformation on a spatial input sound source by the audio processing terminal.

[0116] The process of performing stereo ambisonic transformation on the spatial input sound source by the audio processing terminal is described below.

[0117] In some embodiments, the process of performing stereo ambisonic transformation on the spatial input sound source by the audio processing terminal includes the following steps (1)-(2):

[0118] (1) performing stereo ambisonic transformation on the spatial input sound source based on the order of stereo ambisonic transformation corresponding to the spatial input sound source, to obtain a target direct sound signal.

[0119] The process of performing stereo ambisonic transformation on the spatial input sound source based on the order of stereo ambisonic transformation corresponding to the spatial input sound source includes the following steps A1-A3:

[0120] Step A1, grouping the spatial input sound source to obtain at least one group of input sound sources.

[0121] In some embodiments, if the spatial input sound source includes a multi-channel sound source and a multi-object sound source, the multi-channel sound source is taken as a group of input sound sources, and the multi-object sound source is taken as a group of input sound sources; or, the spatial input sound source is grouped according to a sound source feature to obtain at least one group of input sound sources, and the sound source feature is used to represent the characteristics of the sound source.

[0122] The sound source feature includes at least one of the following: sound source direction, sound source distance, sound source content; wherein the sound source content is used to indicate the type of the sound object in the corresponding sound source. For example, the type of the sound object can be a person, an animal, an environment, a musical instrument, etc. Of course, the type of the sound object can also be an adult, a child, a male, a female, a type of animal, a type of musical instrument, etc.

[0123] The grouping of the spatial input sound sources according to the sound source features can be performed in various manners. In some embodiments, the input sound sources of the same sound source feature are grouped, the input sound sources of the sound source features within the same feature range are grouped, or the input sound sources with a difference between the sound source features less than a difference threshold are grouped, the feature range and the difference threshold being determined based on the sound source features of the spatial input sound sources.

[0124] For example, assuming that the sound source features include sound source directions, the input sound sources of the same sound source direction are grouped, or the input sound sources with a direction difference between the sound source directions less than a difference threshold are grouped, i.e., the input sound sources with close sound source directions are grouped. Assuming that the sound source features include sound source distances, the input sound sources of the sound source distances within the same distance range are grouped, the distance range being obtained by dividing the sound source distances of the spatial input sound sources. Assuming that the sound source features include sound source contents, the input sound sources of the same type of sound production object indicated by the sound source contents are grouped.

[0125] The above-mentioned grouping manners are examples, and in actual applications, the above-mentioned grouping manners can be combined, or other grouping manners can be used, which are not limited in the embodiments of the present application.

[0126] In the case where the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are grouped as a group of input sound sources, and the multi-object sound sources are grouped as a group of input sound sources, which is relatively simple and efficient. The grouping according to the sound source features does not need to consider whether the multi-channel sound sources and the multi-object sound sources exist simultaneously, and the spatial input sound sources can be divided into at least one group of input sound sources, so that subsequent processing can be performed on each group of input sound sources, increasing the degree of freedom of processing and improving the flexibility of processing of the input sound sources.

[0127] The above-mentioned two grouping manners (i.e., grouping according to the multi-channel sound sources and the multi-object sound sources, and grouping according to the sound source features) are examples, and in actual applications, other grouping manners can be used. For example, if the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are grouped according to the sound source features to obtain at least one group of input sound sources, and the multi-object sound sources are grouped according to the sound source features to obtain at least one group of input sound sources.

[0128] In step A2, the at least one group of input sound sources is respectively subjected to stereo sound transformation based on the respective stereo sound transformation orders of the at least one group of input sound sources, to obtain at least one direct sound signal.

[0129] For each of the at least one group of input sound sources, if the stereo sound transform order corresponding to the group of input sound sources is equal to 1, the FOA transform is performed on the group of input sound sources to obtain the direct sound signal corresponding to the group of input sound sources, and the direct sound signal is an FOA signal. If the stereo sound transform order corresponding to the group of input sound sources is greater than 1, the HOA transform is performed on the group of input sound sources to obtain the direct sound signal corresponding to the group of input sound sources, and the direct sound signal is an HOA signal.

[0130] It should be noted that the stereo sound transform order refers to the order of stereo sound transform, and also refers to the order of the direct sound signal obtained after the stereo sound transform. The stereo sound transform orders corresponding to the at least one group of input sound sources can be the same or different, and the embodiments of the present application do not limit this.

[0131] The stereo sound transform can uniformly convert input sound sources of different formats into FOA / HOA formats, which not only can retain all spatial information of the original input sound sources, but also can convert the number of channels of the spatial input sound sources and / or the number of sound emitting objects into the number of channels of the FOA / HOA format audio signal, which is only related to the order of the audio signal and is irrelevant to the number of channels of the spatial input sound sources and / or the number of sound emitting objects. In this way, even if the number of channels of the spatial input sound sources and / or the number of sound emitting objects increases, the number of channels of the FOA / HOA format audio signal of the same order will remain unchanged, thereby reducing the computational complexity of subsequent signal processing.

[0132] In some embodiments, before the stereo sound transform is performed on the at least one group of input sound sources, a first parameter configuration interface can also be displayed to obtain the stereo sound transform order corresponding to each group of input sound sources from the first parameter configuration interface. That is, the stereo sound transform order corresponding to each group of input sound sources is set in the first parameter configuration interface.

[0133] As an example, the first parameter configuration interface includes the default transform order corresponding to each group of input sound sources. For each group of input sound sources, if an adjustment operation of the default transform order corresponding to the group of input sound sources is detected, the adjusted transform order is determined as the stereo sound transform order corresponding to the group of input sound sources, and if the adjustment operation of the default transform order corresponding to the group of input sound sources is not detected, the default transform order is determined as the stereo sound transform order corresponding to the group of input sound sources.

[0134] That is, the user can adjust the default transform order corresponding to the group of input sound sources, or can not adjust the default transform order corresponding to the group of input sound sources. If the user adjusts the default transform order corresponding to the group of input sound sources, the adjusted transform order is determined as the stereo sound transform order corresponding to the group of input sound sources. If the user does not adjust the default transform order corresponding to the group of input sound sources, the default transform order corresponding to the group of input sound sources is directly determined as the stereo sound transform order corresponding to the group of input sound sources.

[0135] Wherein, when the user adjusts the default transform order corresponding to the group of input sound sources, the adjusted transform order can be directly input in the display position of the default transform order in the first parameter configuration interface, thereby covering the default transform order. Alternatively, the first parameter configuration interface can also provide a plurality of candidate transform orders, and the user can select a candidate transform order from the plurality of candidate transform orders as the adjusted transform order, thereby covering the default transform order. Of course, in actual application, the default transform order can also be adjusted by other ways, which are not limited by the embodiments of the present application.

[0136] The default transform order corresponding to each group of input sound sources can be determined by the audio processing terminal based on the characteristics of each group of input sound sources. For example, the audio processing terminal stores a correspondence between sound source characteristics and default transform orders. For each group of input sound sources in the at least one group of input sound sources, the default transform order corresponding to the group of input sound sources is obtained from the correspondence between sound source characteristics and default transform orders based on the characteristics of the group of input sound sources.

[0137] The above is to display the default transform order corresponding to each group of input sound sources in the first parameter configuration interface, and then to determine the stereo sound transform order corresponding to each group of input sound sources by adjusting or not adjusting the default transform order by the user. In other embodiments, the first parameter configuration interface can also not display the default transform order corresponding to each group of input sound sources, and the user can directly input the stereo sound transform order corresponding to each group of input sound sources in the first parameter configuration interface, or select the stereo sound transform order corresponding to each group of input sound sources from a plurality of candidate transform orders provided by the first parameter configuration interface.

[0138] Determining the stereo sound transform order corresponding to each group of input sound sources through the first parameter configuration interface can support customization of the stereo sound transform order, thereby improving the flexibility of autonomous selection of the stereo sound transform order. Moreover, displaying the default transform order corresponding to each group of input sound sources in the first parameter configuration interface can provide the user with a basis for adjusting the transform order through the default transform order, so that the user can determine the stereo sound transform order corresponding to each group of input sound sources through simple adjustment operation, thereby improving the determination efficiency of the stereo sound transform order.

[0139] For example, spatial input audio sources include multi-channel audio sources and multi-object audio sources. The first parameter configuration interface displayed on the audio processing terminal is shown in Figures 6 and 7, where the multi-channel audio sources are treated as one group of input audio sources and the multi-object audio sources as another. In Figures 6 and 7, the first parameter configuration interface includes "Multi-channel" and "Multi-object" controls. Users can set the stereo transformation order of the multi-channel and multi-object audio sources by clicking these controls. Assuming the user currently clicks the "Multi-channel" control, as shown in Figure 6, the first parameter configuration interface displays an order configuration area, and the default transformation order 2 is displayed in the order configuration area. If the user adjusts this default transformation order 2 to 3, then the stereo transformation order of the multi-channel audio source is determined to be 3. Alternatively, as shown in Figure 7, the order configuration area in the first parameter configuration interface does not display the default transformation order, and the order configuration area provides multiple candidate transformation orders. Users can select one candidate transformation order from these multiple candidate transformation orders as the stereo transformation order of the multi-channel sound source. Suppose the user selects 3 as the stereo transformation order of the multi-channel sound source.

[0140] For example, spatial input sound sources are divided into three groups according to their characteristics. The first parameter configuration interface displayed on the audio processing terminal is shown in Figures 8 and 9. In Figures 8 and 9, the first parameter configuration interface includes controls for "Group 1," "Group 2," and "Group 3." Users can set the stereo transformation order of these three groups of input sound sources by clicking these three controls. Assuming the user clicks the "Group 1" control, as shown in Figure 8, the first parameter configuration interface displays an order configuration area, and the order configuration area displays a default transformation order of 2. If the user adjusts the default transformation order 2 to 3, then the stereo transformation order of the first group of input sound sources is determined to be 3. Alternatively, as shown in Figure 9, the order configuration area in the first parameter configuration interface does not display a default transformation order, and the order configuration area provides multiple candidate transformation orders. The user can select one candidate transformation order from these multiple candidate transformation orders as the stereo transformation order of the first group of input sound sources. Assuming the user selects 3 as the stereo transformation order of the first group of input sound sources.

[0141] It should be noted that the above is to set the stereo sound transform order corresponding to the at least one group of input sound sources through the first parameter configuration interface. In other embodiments, the corresponding default transform order can also be directly obtained from the stored correspondence between the sound source characteristics and the default transform order according to the sound source characteristics of the at least one group of input sound sources, and the default transform order corresponding to the at least one group of input sound sources is determined as the corresponding stereo sound transform order. That is, the default transform order can be directly used as the final stereo sound transform order, without being displayed to the user through the first parameter configuration interface and adjusted by the user. Of course, the stereo sound transform order corresponding to the at least one group of input sound sources can also be determined through other ways, and the embodiments of the present application do not limit this.

[0142] Since the default transform order has a correspondence with the sound source characteristics, that is, the default transform order is determined based on the sound source characteristics, the effect of stereo sound transform on the input sound source through the default transform order is usually good, so directly using the default transform order as the stereo sound transform order without user adjustment can avoid user adjustment operation and improve the efficiency of stereo sound transform.

[0143] Step A3, superimposing the at least one direct sound signal to obtain a target direct sound signal.

[0144] Based on the above description, the stereo sound transform order corresponding to the at least one group of input sound sources can be equal to 1 or greater than 1. When the stereo sound transform order is equal to 1, the obtained direct sound signal is a FOA signal, and when the stereo sound transform order is greater than 1, the obtained direct sound signal is a HOA signal. That is, the at least one direct sound signal can include a FOA signal or a HOA signal, and the order of the HOA signal can also be the same or different. Based on this, when superimposing the at least one direct sound signal, signals of the same channel can be superimposed to obtain a target direct sound signal.

[0145] For example, the spatial input sound sources are divided into three groups of input sound sources, and three direct sound signals are obtained by performing stereo sound conversion on the three groups of input sound sources, which are FOA signal, 2-order HOA signal and 3-order HOA signal respectively. The FOA signal includes 4 channels of signals, the 2-order HOA signal includes 9 channels of signals, and the 3-order HOA signal includes 16 channels of signals. The signal of the first channel of the FOA signal, the signal of the first channel of the 2-order HOA signal and the signal of the first channel of the 3-order HOA signal are superimposed to obtain the signal of the first channel of the target direct sound signal. The signal of the second channel of the FOA signal, the signal of the second channel of the 2-order HOA signal and the signal of the second channel of the 3-order HOA signal are superimposed to obtain the signal of the second channel of the target direct sound signal. The signal of the third channel of the FOA signal, the signal of the third channel of the 2-order HOA signal and the signal of the third channel of the 3-order HOA signal are superimposed to obtain the signal of the third channel of the target direct sound signal. The signal of the fourth channel of the FOA signal, the signal of the fourth channel of the 2-order HOA signal and the signal of the fourth channel of the 3-order HOA signal are superimposed to obtain the signal of the fourth channel of the target direct sound signal. The signal of the fifth channel of the 2-order HOA signal and the signal of the fifth channel of the 3-order HOA signal are superimposed to obtain the signal of the fifth channel of the target direct sound signal. The signal of the sixth channel of the 2-order HOA signal and the signal of the sixth channel of the 3-order HOA signal are superimposed to obtain the signal of the sixth channel of the target direct sound signal. The signal of the seventh channel of the 2-order HOA signal and the signal of the seventh channel of the 3-order HOA signal are superimposed to obtain the signal of the seventh channel of the target direct sound signal. The signal of the eighth channel of the 2-order HOA signal and the signal of the eighth channel of the 3-order HOA signal are superimposed to obtain the signal of the eighth channel of the target direct sound signal. The signal of the ninth channel of the 2-order HOA signal and the signal of the ninth channel of the 3-order HOA signal are superimposed to obtain the signal of the ninth channel of the target direct sound signal. The signals of the tenth to sixteenth channels of the 3-order HOA signal are the signals of the tenth to sixteenth channels of the target direct sound signal.

[0146] In some embodiments, before superimposing the at least one direct sound signal, the at least one direct sound signal can also be subjected to sound effect processing based on the direct sound effect processing parameters corresponding to the at least one group of input sound sources respectively. In this way, after superimposing the at least one direct sound signal subjected to sound effect processing, the realism of the target direct sound signal can be improved.

[0147] In some embodiments, the direct sound effect processing parameter corresponding to each group of input sound sources comprises a gain and / or an EQ parameter. In this way, for each group of input sound sources, if the direct sound effect processing parameter corresponding to the group of input sound sources comprises a gain, the direct sound signal corresponding to the group of input sound sources is gain-adjusted based on the gain. If the direct sound effect processing parameter corresponding to the group of input sound sources comprises an EQ parameter, the direct sound signal corresponding to the group of input sound sources is EQ-adjusted based on the EQ parameter.

[0148] As an example, the direct sound signal corresponding to the group of input sound sources is multiplied by the gain corresponding to the group of input sound sources to obtain a gain-adjusted direct sound signal. The EQ parameter comprises a plurality of frequency point EQ parameters, and the direct sound signal corresponding to the group of input sound sources is filtered by the plurality of frequency point EQ parameters to obtain an EQ-adjusted direct sound signal.

[0149] It should be noted that when the direct sound effect processing parameter of the group of input sound sources comprises a gain and an EQ parameter, the direct sound signal corresponding to the group of input sound sources can be first gain-adjusted by the gain, and then the gain-adjusted direct sound signal is EQ-adjusted by the EQ parameter. The direct sound signal corresponding to the group of input sound sources can also be first EQ-adjusted by the EQ parameter, and then the EQ-adjusted direct sound signal is gain-adjusted by the gain. The order of the two processing operations is not limited by the embodiments of the present application.

[0150] In some embodiments, the direct sound effect processing parameter corresponding to each group of input sound sources can be the same or different, and the embodiments of the present application also do not limit this.

[0151] In some embodiments, before the at least one direct sound signal is processed, a second parameter configuration interface can also be displayed to obtain the direct sound effect processing parameter corresponding to each group of input sound sources from the second parameter configuration interface. That is, the direct sound effect processing parameter corresponding to each group of input sound sources is set in the second parameter configuration interface.

[0152] As an example, the second parameter configuration interface comprises default sound effect processing parameters corresponding to each group of input sound sources. For each group of input sound sources, if an adjustment operation of the default sound effect processing parameter corresponding to the group of input sound sources is detected, the adjusted sound effect processing parameter is determined as the direct sound effect processing parameter corresponding to the group of input sound sources, and if the adjustment operation of the default sound effect processing parameter corresponding to the group of input sound sources is not detected, the default sound effect processing parameter is determined as the direct sound effect processing parameter corresponding to the group of input sound sources.

[0153] That is, the user can adjust the default sound effect processing parameter corresponding to the group of input sound sources, or can not adjust the default sound effect processing parameter corresponding to the group of input sound sources. If the user adjusts the default sound effect processing parameter corresponding to the group of input sound sources, the adjusted sound effect processing parameter is determined as the direct sound effect processing parameter corresponding to the group of input sound sources. If the user does not adjust the default sound effect processing parameter corresponding to the group of input sound sources, the default sound effect processing parameter corresponding to the group of input sound sources is directly determined as the direct sound effect processing parameter corresponding to the group of input sound sources.

[0154] When the user adjusts the default sound effect processing parameter corresponding to the group of input sound sources, the slider on the sliding bar where the default sound effect processing parameter is located in the second parameter configuration interface can be directly slid to obtain the adjusted sound effect processing parameter, thereby covering the default sound effect processing parameter. Alternatively, the second parameter configuration interface can also provide a plurality of candidate sound effect processing parameters, and the user can select a candidate sound effect processing parameter from the plurality of candidate sound effect processing parameters as the adjusted sound effect processing parameter, thereby covering the default sound effect processing parameter. Of course, in actual application, the default sound effect processing parameter can also be adjusted by other ways, and the embodiments of the present application do not limit this.

[0155] The default sound effect processing parameter corresponding to each group of input sound sources can be determined by the audio processing terminal based on the characteristics of each group of input sound sources. For example, the audio processing terminal includes a correspondence between sound source characteristics and default sound effect processing parameters. For each group of input sound sources in at least one group of input sound sources, the default sound effect processing parameter corresponding to the group of input sound sources is obtained from the correspondence between sound source characteristics and default sound effect processing parameters based on the characteristics of the group of input sound sources.

[0156] The above is that the default sound effect processing parameters corresponding to at least one group of input sound sources are displayed in the second parameter configuration interface, and then the default sound effect processing parameters are adjusted or not adjusted by the user to determine the direct sound effect processing parameters corresponding to at least one group of input sound sources. In some other embodiments, the second parameter configuration interface can also not display the default sound effect processing parameters corresponding to at least one group of input sound sources, and the user directly inputs the direct sound effect processing parameter corresponding to each group of input sound sources in the second parameter configuration interface, or selects the direct sound effect processing parameter corresponding to each group of input sound sources from a plurality of candidate sound effect processing parameters provided by the user in the second parameter configuration interface.

[0157] The second parameter configuration interface is used to determine the direct sound effect processing parameters corresponding to the at least one group of input sound sources, which can support customization of the effect processing parameters, and further improve the flexibility of autonomous selection of the effect processing parameters. Moreover, the default effect processing parameters corresponding to the at least one group of input sound sources are displayed in the second parameter configuration interface, and the default effect processing parameters can provide a basis for the user to adjust the effect processing parameters, so that the user can determine the direct sound effect processing parameters corresponding to the at least one group of input sound sources through simple adjustment operations, and further improve the determination efficiency of the direct sound effect processing parameters.

[0158] For example, the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are taken as a group of input sound sources, and the multi-object sound sources are taken as a group of input sound sources. The second parameter configuration interface displayed by the audio processing terminal is shown in FIG. 10. In FIG. 10, the second parameter configuration interface includes controls of “multi-channel” and “multi-object”, and the user sets the direct sound effect processing parameters of the multi-channel sound sources and the multi-object sound sources by clicking the two controls. Assuming that the user currently clicks the “multi-channel” control, the second parameter configuration interface displays an effect parameter configuration area, and the effect parameter configuration area displays a sliding bar of gain and sliding bars of EQ parameters of multiple frequency points, as shown in FIG. 10. The sliders (black dots on the sliding bars in FIG. 10) on each sliding bar are in the default position, indicating the default gain of the multi-channel sound sources and the default EQ parameters of the multiple frequency points. The user adjusts the position of the slider on each sliding bar to adjust the default effect processing parameters, and further determines the adjusted effect processing parameters as the direct sound effect processing parameters of the multi-channel sound sources.

[0159] It should be noted that the second parameter configuration interface and the first parameter configuration interface can be located in the same interface or different interfaces, and the embodiments of the present application do not limit this. For example, the first parameter configuration interface and the second parameter configuration interface are located in the same interface, and the interface is shown in FIG. 11.

[0160] The above is to set the direct sound effect processing parameters corresponding to the at least one group of input sound sources through the second parameter configuration interface. In other embodiments, the corresponding default effect processing parameters can also be directly obtained from the stored correspondence between the sound source characteristics and the default effect processing parameters according to the sound source characteristics of the at least one group of input sound sources, and the default effect processing parameters corresponding to the at least one group of input sound sources are determined as the respective corresponding direct sound effect processing parameters. That is, the default effect processing parameters can be directly taken as the final direct sound effect processing parameters, without being displayed to the user through the second parameter configuration interface and adjusted by the user. Of course, the direct sound effect processing parameters corresponding to the at least one group of input sound sources can also be determined through other ways, and the embodiments of the present application do not limit this.

[0161] Since the default sound effect processing parameter has a corresponding relationship with the sound source feature, that is, the default sound effect processing parameter is determined based on the sound source feature, the effect of sound effect processing on the input sound source through the default sound effect processing parameter is usually good, and therefore directly using the default sound effect processing parameter as the direct sound effect processing parameter, without user adjustment, can avoid user adjustment operation and improve the efficiency of sound effect processing.

[0162] It should be noted that the above is to group the spatial input sound source, perform stereo sound transformation on each group to obtain the corresponding direct sound signal of each group, and then superimpose the corresponding direct sound signal of each group to obtain the target direct sound signal. Of course, the spatial input sound source can also be directly subjected to stereo sound transformation to obtain the target direct sound signal without grouping, and the embodiments of the present application do not limit this. The processing operation involved in directly performing stereo sound transformation on the spatial input sound source is similar to the processing operation involved in performing stereo sound transformation on the at least one group of input sound sources, and the detailed implementation process can be referred to the above description, and the embodiments of the present application will not be repeated here.

[0163] (2) Based on the target direct sound signal, determine the spatial audio signal.

[0164] In some embodiments, the target direct sound signal can be directly determined as the spatial audio signal. At this time, the spatial audio signal includes the direct sound signal and does not include the reverberation sound signal.

[0165] In other embodiments, the reverberation sound signal can also be added to the spatial audio signal. That is, based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is subjected to reverberation processing to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data. In this way, the target direct sound signal and the target reverberation sound signal can be superimposed to obtain the spatial audio signal, so that the spatial audio signal can bring a more realistic and more immersive spatial audio experience.

[0166] The process of the audio processing terminal performing reverberation processing on the spatial input sound source is described below.

[0167] In some embodiments, the spatial input sound source is divided into at least one group of input sound sources; at this time, the process of performing reverberation processing on the spatial input sound source based on the reverberation data corresponding to the spatial input sound source includes the following steps B1-B2:

[0168] Step B1, based on the reverberation data corresponding to the at least one group of input sound sources, respectively performing reverberation processing on the at least one group of input sound sources to obtain at least one reverberation sound signal.

[0169] Based on the above description, the reverberation data is FOA reverberation data or HOA reverberation data, and for each group of input sound sources in the at least one group of input sound sources, if the reverberation data corresponding to the group of input sound sources is FOA reverberation data, the group of input sound sources is subjected to reverberation processing based on the reverberation data corresponding to the group of input sound sources, and a corresponding reverberation sound signal is obtained as an FOA signal. If the reverberation data corresponding to the group of input sound sources is HOA reverberation data, the group of input sound sources is subjected to reverberation processing based on the reverberation data corresponding to the group of input sound sources, and a corresponding reverberation sound signal is obtained as an HOA signal.

[0170] It should be noted that the reverberation data refers to data for converting the input sound source into a reverberation sound signal, and the reverberation data corresponding to the at least one group of input sound sources can be the same or different, which is not limited by the embodiments of the present application.

[0171] In some embodiments, before the at least one group of input sound sources is subjected to reverberation processing respectively, the reverberation data corresponding to the at least one group of input sound sources can be obtained based on the reverberation parameters corresponding to the at least one group of input sound sources respectively.

[0172] As an example, the reverberation parameters include a reverberation sound rendering order and a room type, and the room type is used to indicate the type of the room currently simulated by the reverberation sound. At this time, based on the reverberation parameters corresponding to the at least one group of input sound sources respectively, the reverberation data corresponding to the at least one group of input sound sources is obtained, including: based on the reverberation sound rendering order and the room type corresponding to the at least one group of input sound sources respectively, the reverberation data corresponding to the at least one group of input sound sources is obtained.

[0173] The reverberation sound rendering order refers to the order of the reverberation data, and also refers to the order of the reverberation sound signal obtained after the reverberation processing. The reverberation sound rendering order corresponding to the at least one group of input sound sources can be the same or different, which is not limited by the embodiments of the present application. In addition, for each group of input sound sources in the at least one group of input sound sources, the reverberation sound rendering order corresponding to the group of input sound sources is less than or equal to the stereo sound conversion order corresponding to the group of input sound sources. That is, the order of the reverberation sound signal corresponding to the group of input sound sources is less than or equal to the order of the direct sound signal corresponding to the group of input sound sources.

[0174] The room type can be a cinema, a concert hall, a church, etc., which is not limited by the embodiments of the present application.

[0175] For each group of input sound sources in the at least one group of input sound sources, the reverberation data corresponding to the group of input sound sources is obtained from the stored corresponding relationship between the room type, the reverberation sound rendering order and the reverberation data, based on the room type and the reverberation sound rendering order corresponding to the group of input sound sources. The corresponding relationship between the room type, the reverberation sound rendering order and the reverberation data is determined in advance, and the determination method of the corresponding relationship is not limited by the embodiments of the present application.

[0176] The above-mentioned manner of obtaining the reverberation data corresponding to each group of input sound sources can be applied to any division manner of the spatial input sound sources. That is, no matter how the spatial input sound sources are divided, the reverberation data corresponding to each group of input sound sources can be obtained by the above-mentioned manner. In some other embodiments, if at least one group of input sound sources is divided based on the sound source feature, and based on the reverberation sound rendering order and the room type corresponding to each group of input sound sources, the reverberation data corresponding to each group of input sound sources is obtained based on the reverberation sound rendering order, the room type and the sound source feature corresponding to each group of input sound sources.

[0177] As an example, for each group of input sound sources in the at least one group of input sound sources, the reverberation data corresponding to the group of input sound sources is obtained from the stored correspondence between the room type, the reverberation sound rendering order, the sound source feature and the reverberation data, based on the room type, the reverberation sound rendering order and the sound source feature corresponding to the group of input sound sources. The correspondence between the room type, the reverberation sound rendering order, the sound source feature and the reverberation data is determined in advance, and the determination manner of the correspondence is not limited in the embodiments of the present application.

[0178] In some other embodiments, the reverberation parameters further include a spatial parameter and / or a room size, the spatial parameter is used to describe the spatial information in the reverberation sound, and the room size indicates the size of the room currently simulated by the reverberation sound. Thus, after obtaining the reverberation data corresponding to each group of input sound sources, the reverberation data corresponding to each group of input sound sources can be adjusted based on the spatial parameter and / or the room size corresponding to each group of input sound sources.

[0179] If the reverberation parameters include the spatial parameter, for each group of input sound sources in the at least one group of input sound sources, the direct-to-reverberation ratio corresponding to the group of input sound sources is obtained from the stored correspondence between the spatial parameter and the direct-to-reverberation ratio, based on the spatial parameter corresponding to the group of input sound sources, and the reverberation data corresponding to the group of input sound sources is adjusted according to the direct-to-reverberation ratio corresponding to the group of input sound sources. For example, the reverberation data corresponding to the group of input sound sources is multiplied by the inverse of the direct-to-reverberation ratio corresponding to the group of input sound sources, to obtain the adjusted reverberation data corresponding to the group of input sound sources.

[0180] Since the spatial parameter is used to describe the spatial information in the reverberation sound, i.e., the spatial parameter is used to describe the spatial component in the reverberation sound, so that the audio signal sounds with spatiality. In other words, the spatial parameter refers to the listening feeling of the indoor sound field currently simulated by the reverberation sound. For example, if the spatial parameter is smaller, it means that the listening feeling and spatiality of the indoor sound field currently simulated by the reverberation sound are weaker, and if the spatial parameter is larger, it means that the listening feeling and spatiality of the indoor sound field currently simulated by the reverberation sound are stronger. The spatial parameter has a corresponding relationship with the direct-mix ratio, different spatial parameters correspond to different direct-mix ratios, and the energy ratio of the direct sound signal and the reverberation sound signal in the audio signal can be adjusted through the direct-mix ratio, so as to adjust the listening feeling of the indoor sound field, so that the listening feeling of the indoor sound field is more realistic.

[0181] It should be noted that adjusting the reverberation data through the direct-mix ratio, and then adjusting the spatiality of the audio signal is an example, i.e., the direct-mix ratio is one of the parameters for adjusting the spatiality. In actual application, the spatiality of the audio signal can also be adjusted in other ways, i.e., the parameters for adjusting the spatiality can also include other parameters in addition to the direct-mix ratio, and the adjustment of the spatiality is realized by adjusting the reverberation data through other parameters, which is not limited by the embodiments of the present application.

[0182] If the reverberation parameter further includes the room size, for each group of input sound sources in the at least one group of input sound sources, based on the room size corresponding to the group of input sound sources, the adjustment parameters are obtained from the corresponding relationship between the stored room size and the adjustment parameters, here the adjustment parameters include EQ parameters, delay parameters, etc., and the reverberation data corresponding to the group of input sound sources is adjusted according to the adjustment parameters corresponding to the group of input sound sources. For example, the EQ parameters include the EQ parameters of a plurality of frequency points, the reverberation data corresponding to the group of input sound sources is filtered through the EQ parameters of the plurality of frequency points, the delay parameters include the delay parameters of each reflection signal in the reverberation data, and each reflection signal in the reverberation data corresponding to the group of input sound sources is adjusted in delay, so as to obtain the adjusted reverberation data of the group of input sound sources.

[0183] It should be noted that when the reverberation parameter of the group of input sound sources includes the spatial parameter and the room size, the spatiality of the reverberation data corresponding to the group of input sound sources can be adjusted through the spatial parameter first, and then the room size of the spatially adjusted reverberation data is adjusted through the adjustment parameter corresponding to the room size. The room size of the reverberation data corresponding to the group of input sound sources can also be adjusted through the adjustment parameter corresponding to the room size first, and then the spatiality of the room size adjusted reverberation data is adjusted through the direct-mix ratio corresponding to the spatial parameter. The order of the two processing operations is not limited by the embodiments of the present application.

[0184] Wherein, the reverberation parameters corresponding to each group of input sound sources can be the same or different, and embodiments of the present application do not limit this.

[0185] In some embodiments, based on the reverberation parameters corresponding to each group of input sound sources, before obtaining the reverberation data corresponding to each group of input sound sources, a third parameter configuration interface can be displayed, and the reverberation parameters corresponding to each group of input sound sources are obtained from the third parameter configuration interface. That is, the reverberation parameters corresponding to each group of input sound sources are set in the third parameter configuration interface.

[0186] As an example, the third parameter configuration interface includes the default reverberation parameters corresponding to each group of input sound sources. For each group of input sound sources, if an adjustment operation of the default reverberation parameter corresponding to the group of input sound sources is detected, the adjusted reverberation parameter is determined as the reverberation parameter corresponding to the group of input sound sources, and if the adjustment operation of the default reverberation parameter corresponding to the group of input sound sources is not detected, the default reverberation parameter is determined as the reverberation parameter corresponding to the group of input sound sources.

[0187] That is, the user can adjust the default reverberation parameter corresponding to the group of input sound sources, or can not adjust the default reverberation parameter corresponding to the group of input sound sources. If the user adjusts the default reverberation parameter corresponding to the group of input sound sources, the adjusted reverberation parameter is determined as the reverberation parameter corresponding to the group of input sound sources. If the user does not adjust the default reverberation parameter corresponding to the group of input sound sources, the default reverberation parameter corresponding to the group of input sound sources is directly determined as the reverberation parameter corresponding to the group of input sound sources.

[0188] Wherein, when the user adjusts the default reverberation parameter corresponding to the group of input sound sources, the slider on the slider bar where the default reverberation parameter is located in the third parameter configuration interface can be directly slid to obtain the adjusted reverberation parameter, thereby covering the default reverberation parameter. Alternatively, the third parameter configuration interface can also provide a plurality of candidate reverberation parameters, and the user can select a candidate reverberation parameter from the plurality of candidate reverberation parameters as the adjusted reverberation parameter, thereby covering the default reverberation parameter. Of course, in actual application, the default reverberation parameter can also be adjusted by other ways, and embodiments of the present application do not limit this.

[0189] The default reverberation parameter corresponding to each group of input sound sources can be determined by the audio processing terminal based on the characteristics of each group of input sound sources. For example, the audio processing terminal includes a correspondence between sound source characteristics and default reverberation parameters. For each group of input sound sources in the at least one group of input sound sources, the default reverberation parameter corresponding to the group of input sound sources is obtained from the correspondence between sound source characteristics and default reverberation parameters based on the characteristics of the group of input sound sources.

[0190] The above is to display the default reverberation parameters corresponding to the at least one group of input sound sources in the third parameter configuration interface, and then determine the reverberation parameters corresponding to the at least one group of input sound sources by adjusting or not adjusting the default reverberation parameters by the user. In other embodiments, the third parameter configuration interface can also not display the default reverberation parameters corresponding to the at least one group of input sound sources, and directly input the reverberation parameters corresponding to each group of input sound sources by the user in the third parameter configuration interface, or select the reverberation parameters corresponding to each group of input sound sources from the multiple candidate reverberation parameters provided by the user in the third parameter configuration interface.

[0191] The determination of the reverberation parameters corresponding to the at least one group of input sound sources through the third parameter configuration interface can support the customization of the reverberation parameters, thereby improving the flexibility of the autonomous selection of the reverberation parameters. Moreover, the display of the default reverberation parameters corresponding to the at least one group of input sound sources in the third parameter configuration interface can provide the user with a basis for adjusting the reverberation parameters through the default reverberation parameters, so that the user can determine the reverberation parameters corresponding to the at least one group of input sound sources through a simple adjustment operation, thereby improving the determination efficiency of the reverberation parameters.

[0192] For example, the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are taken as a group of input sound sources, and the multi-object sound sources are taken as a group of input sound sources. The third parameter configuration interface displayed by the audio processing terminal is as shown in FIG. 12. In FIG. 12, the third parameter configuration interface includes the controls of “multi-channel” and “multi-object”, and the user sets the reverberation parameters of the multi-channel sound sources and the multi-object sound sources by clicking the two controls. Assuming that the user currently clicks the “multi-channel” control, as shown in FIG. 12, the third parameter configuration interface displays a reverberation parameter configuration area, and the reverberation parameter configuration area displays the slide bar of the spatial sense parameter, the slide bar of the room size, and the default room type. The slider on each slide bar is in the default position, indicating the default spatial sense parameter and the default room size of the multi-channel sound sources. The user adjusts the default spatial sense parameter by sliding the position of the slider on the slide bar of the spatial sense parameter, adjusts the default room size by sliding the position of the slider on the slide bar of the room size, and can also adjust the default room type, thereby obtaining the reverberation parameters of the multi-channel sound sources.

[0193] It should be noted that FIG. 12 is illustratively explained by taking the reverberation parameters including the spatial sense parameter, the room size, and the room type as an example. The reverberation sound rendering order included in the reverberation parameters can also be set in a similar manner. Of course, in other embodiments, the reverberation sound rendering order can also be the same as the stereo sound transformation order, so the stereo sound transformation order set in the first parameter configuration interface can be directly taken as the reverberation sound rendering order.

[0194] The above is setting the reverberation parameters corresponding to the at least one group of input sound sources through the third parameter configuration interface. In other embodiments, the corresponding default reverberation parameters can also be directly obtained from the stored correspondence between the sound source characteristics and the default reverberation parameters according to the sound source characteristics of the at least one group of input sound sources, and the default reverberation parameters corresponding to the at least one group of input sound sources are determined as the respective corresponding reverberation parameters. That is, the default reverberation parameters can be directly used as the final reverberation parameters without being displayed to the user through the third parameter configuration interface and adjusted by the user. Of course, the reverberation parameters corresponding to the at least one group of input sound sources can also be determined in other ways, which are not limited in the embodiments of the present application.

[0195] Since the default reverberation parameters have a correspondence with the sound source characteristics, that is, the default reverberation parameters are determined based on the sound source characteristics, the effect of reverberation processing on the input sound source through the default reverberation parameters is usually good, so directly using the default reverberation parameters as the final reverberation parameters without user adjustment can avoid user adjustment operations and improve the efficiency of reverberation processing.

[0196] The third parameter configuration interface, the second parameter configuration interface, and the first parameter configuration interface can be located in the same interface or different interfaces, which are not limited in the embodiments of the present application. For example, the first parameter configuration interface, the second parameter configuration interface, and the third parameter configuration interface are located in the same interface, which is shown in FIG. 13.

[0197] Step B2, superimposing the at least one reverberation sound signal to obtain a target reverberation sound signal.

[0198] Based on the above description, the reverberation data corresponding to the at least one group of input sound sources can be FOA reverberation data or HOA reverberation data. When the reverberation data is FOA reverberation data, the obtained reverberation sound signal is a FOA signal, and when the reverberation data is HOA reverberation data, the obtained reverberation sound signal is a HOA signal. That is, the at least one reverberation sound signal can include a FOA signal or a HOA signal, and the orders of the HOA signals can be the same or different. Based on this, when superimposing the at least one reverberation sound signal, signals of the same channel can be superimposed to obtain a target reverberation sound signal.

[0199] The process of superimposing the at least one reverberation sound signal is similar to the above-mentioned process of superimposing the at least one direct sound signal, and the details are described above, which will not be repeated here.

[0200] In some embodiments, before the superposition of the at least one reverberation sound signal, the at least one reverberation sound signal can also be subjected to sound effect processing based on the sound effect processing parameters corresponding to the at least one group of input sound sources respectively. In this way, the at least one reverberation sound signal subjected to sound effect processing is superposed to obtain the target reverberation sound signal.

[0201] The sound effect processing manner of the at least one reverberation sound signal is similar to the sound effect processing manner of the at least one direct sound signal, and details are described above and will not be repeated here.

[0202] Based on the above description, when the reverberation parameter also includes a spatial sense parameter, the reverberation data corresponding to the at least one group of input sound sources can also be adjusted based on the spatial sense parameter corresponding to the at least one group of input sound sources respectively, so that the reverberation data of each channel needs to be adjusted. In some other embodiments, before the at least one group of input sound sources is subjected to reverberation processing, the reverberation data corresponding to the at least one group of input sound sources can not be adjusted, but after the target direct sound signal and the target reverberation sound signal are superposed through the direct-to-reverberation ratio corresponding to the spatial sense parameter, only one spatial sense parameter needs to be set, and the number of channels of the target direct sound signal and the target reverberation sound signal is relatively small, which can reduce the computational complexity.

[0203] It should be noted that the above is to group the spatial input sound sources, process each group of input sound sources to obtain the corresponding reverberation sound signal, and then superpose each group of corresponding reverberation sound signals to obtain the target reverberation sound signal. Of course, the spatial input sound sources can also not be grouped, but directly subjected to reverberation processing to obtain the target reverberation sound signal, and the embodiments of the present application do not limit this. The processing operations involved in directly processing the spatial input sound sources are similar to the processing operations involved in processing the at least one group of input sound sources, and the detailed implementation process can be referred to the above description, and the embodiments of the present application will not be repeated here.

[0204] Step 502: Obtain spatial rendering data, which is FOA rendering data or HOA rendering data, and the order of the spatial rendering data is the same as the order of the spatial audio signal.

[0205] When the order of the spatial audio signal is equal to 1, the FOA rendering data can be determined based on the order of the spatial audio signal and the expression of the spatial rendering data of the FOA / HOA domain stored in the database. When the order of the spatial audio signal is greater than 1, the HOA rendering data can be determined based on the order of the spatial audio signal and the expression of the spatial rendering data of the FOA / HOA domain stored in the database. That is, the expression of the spatial rendering data of the FOA / HOA domain stored in the database can be used to determine the HOA rendering data. In this way, the database does not need to store discrete spatial rendering data of each direction, which greatly reduces the occupation of memory space.

[0206] The spatial rendering data can be binaural HRIR / BRIR, and can also be other data, which is not limited in the embodiments of the application.

[0207] In addition, the spatial rendering data can include left ear spatial rendering data and right ear spatial rendering data, or the left ear spatial rendering data and the right ear spatial rendering data can be combined together by using the symmetry of the left ear and the right ear.

[0208] Step 503: rendering the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal.

[0209] Based on the above description, the spatial rendering data can include left ear spatial rendering data and right ear spatial rendering data, or the left ear spatial rendering data and the right ear spatial rendering data can be combined together. For ease of description, the left ear spatial rendering data and the right ear spatial rendering data are combined together and referred to as binaural spatial rendering data. In different cases, the way of rendering the spatial audio signal to obtain a binaural spatial audio output signal is different, which will be introduced next.

[0210] In the first way, the spatial rendering data includes left ear spatial rendering data and right ear spatial rendering data. At this time, the spatial audio signal and the data corresponding to the channel in the left ear spatial rendering data are convolved to obtain the convolution result of each channel, and the convolution results of each channel are superimposed to obtain the left ear spatial audio output signal. The spatial audio signal and the data corresponding to the channel in the right ear spatial rendering data are convolved to obtain the convolution result of each channel, and the convolution results of each channel are superimposed to obtain the right ear spatial audio output signal.

[0211] In the second mode, the spatial rendering data includes binaural spatial rendering data. At this time, the spatial audio signal and the data of the corresponding channel in the binaural spatial rendering data are convoluted to obtain the convolution result of each channel. The convolution results of the channels of the even symmetry mode are superimposed to obtain a first superimposed result, and the convolution results of the channels of the odd symmetry mode are superimposed to obtain a second superimposed result. The first superimposed result and the second superimposed result are superimposed again and divided by 2 to obtain the spatial audio output signal for the left ear. The first superimposed result and the second superimposed result are subtracted again and divided by 2 to obtain the spatial audio output signal for the right ear.

[0212] After obtaining the spatial audio signal through the above steps, the spatial audio signal can be rendered by the spatial rendering data once, which increases the processing freedom without increasing the rendering calculation complexity.

[0213] In some embodiments, before rendering the spatial audio signal, head motion information indicating the attitude of the head wearing the binaural playback terminal can also be obtained, and the spatial audio signal is processed based on the head motion information. That is, when the binaural playback terminal is in the head motion tracking mode, the spatial audio signal can also be processed based on the attitude of the head.

[0214] Based on the above description, the spatial audio rendering method provided by the embodiments of the present application can be applied to an audio processing terminal or a binaural playback terminal. If the method is applied to an audio processing terminal, the audio processing terminal can receive the head motion information sent by the binaural playback terminal, and if the method is applied to a binaural playback terminal, the binaural playback terminal can directly detect the head motion information.

[0215] It should be noted that the head motion information can be measured by an inertial measurement unit (IMU) included in the binaural playback terminal. Of course, it can also be measured by other means by the binaural playback terminal, and the embodiments of the present application do not limit this.

[0216] As an example, the head motion information includes Euler angles, which include a nutation angle, a precession angle, and a rotation angle. A rotation matrix is determined through the nutation angle, the precession angle, and the rotation angle, and the rotation matrix is multiplied by the spatial audio signal to realize the rotation processing of the spatial audio signal.

[0217] In some embodiments, after the spatial audio signal is rendered to obtain the binaural spatial audio output signal, the binaural spatial audio output signal can also be processed for sound effects. For example, the binaural spatial audio output signal is processed for EQ adjustment, dynamic range control (DRC), etc.

[0218] If the method is applied to an audio processing terminal, the audio processing terminal can send the binaural spatial audio signal to a binaural playback terminal after obtaining the binaural spatial audio output signal, so that the binaural playback terminal plays the binaural spatial audio signal. If the method is applied to a binaural playback terminal, the binaural playback terminal plays the binaural spatial audio output signal after obtaining the binaural spatial audio output signal.

[0219] In summary, the embodiments of the present application convert input sound sources of different formats into FOA / HOA format through stereo sound transformation, which not only preserves all spatial information of the original input sound sources, but also converts the number of channels of the spatial input sound sources and / or the number of sound emitting objects into the number of channels of the FOA / HOA format audio signal, which is only related to the order of the audio signal and is irrelevant to the number of channels of the spatial input sound sources and / or the number of sound emitting objects. In this way, even if the number of channels of the spatial input sound sources and / or the number of sound emitting objects increases, the number of channels of the FOA / HOA format audio signal of the same order remains unchanged, thereby reducing the computational complexity of subsequent signal processing. In addition, the rendering of the FOA / HOA format audio signal does not need to store discrete spatial rendering data of each direction in the database, but stores spatial rendering data of the FOA / HOA domain, which is a distributed expression of discrete spatial rendering data in three-dimensional space, and through which the spatial rendering data of any direction in three-dimensional space can be determined. That is, the database does not need to store discrete spatial rendering data of each direction, but stores the expression of the spatial rendering data of the FOA / HOA domain, which greatly reduces the occupation of memory space. Moreover, the spatial rendering data of the FOA / HOA domain is a continuous expression in three-dimensional space, and after the spatial rendering data is determined through this expression, the audio output signal rendered in the process of head rotation is relatively smooth, thereby improving the naturalness and continuity of the listening experience in the head tracking mode. Furthermore, the above parameter configuration interface supports customization of parameters, improves the flexibility of autonomous selection of parameters, and thus realizes customized stereo sound transformation effect and spatial rendering effect.

[0220] Next, the spatial audio rendering method provided by the embodiments of the present application will be introduced schematically through FIG. 14 and FIG. 15.

[0221] Please refer to FIG. 14, the spatial input sound source includes multi-channel sound source and multi-object sound source. The spatial audio rendering method mainly includes two parts of obtaining spatial audio signal and rendering spatial audio signal. Among them, FIG. 14 takes HOA as an example for illustration.

[0222] The obtaining step of the spatial audio signal is as follows:

[0223] 1、The multi-channel sound source is divided into a group, the multi-object sound source is divided into a group, HOA transformation is carried out based on the HOA order corresponding to the multi-channel sound source, HOA transformation is carried out based on the HOA order corresponding to the multi-object sound source, and the two HOA direct sound signals corresponding to the HOA transformation are output respectively.

[0224] 2、The two HOA direct sound signals pass through the respective sound effect processing modules, and the respective HOA direct sound signals are gain-adjusted and EQ-adjusted according to the sound effect processing parameters set by the user.

[0225] 3、The two HOA direct sound signals after sound effect processing are superimposed by the adder, and then the direct sound effect processing module is used for direct sound effect processing to form the final HOA direct sound signal. If the HOA orders of the two groups are different, the signals of the channels corresponding to the same order are superimposed, for example, 2-order (9-channel) HOA signals and 3-order (16-channel) HOA signals are added, and the first 9 channels are added.

[0226] 4、The two groups of input sound sources enter the reverberation processing module at the same time, the reverberation processing module performs reverberation processing on the two groups of input sound sources based on the HOA reverberation data corresponding to the two groups of input sound sources, obtains two HOA reverberation sound signals, and superimposes the two HOA reverberation sound signals. The HOA reverberation data corresponding to the two groups of input sound sources is obtained by reading the corresponding HOA reverberation data from the preset reverberation data according to the room type and the reverberation rendering order, and then adjusting the HOA reverberation data according to the spatial sense parameter and the room size.

[0227] 5、The superimposed HOA reverberation sound signal is subjected to reverberation sound effect processing by the reverberation sound effect processing module to form the final HOA reverberation sound signal.

[0228] 6、The HOA direct sound signal and the HOA reverberation sound signal are superimposed again to form the final HOA audio signal, i.e. the spatial audio signal. Then, output to the spatial audio signal rendering module.

[0229] The rendering steps of the spatial audio signal are as follows:

[0230] 1、The HOA rotation & rendering module performs two processing processes, first, the received HOA audio signal (i.e. the spatial audio signal) is subjected to rotation processing according to the head movement information (the head movement information is obtained from the IMU on the binaural playback terminal, generally using Euler angles to represent), and then the rotated HOA audio signal is subjected to rendering to obtain the binaural spatial audio output signal. The rendering of the HOA audio signal needs to be rendered by HOA rendering data, and the order of the HOA rendering data is the same as that of the spatial audio signal.

[0231] 2. The spatial audio output signal is post-processed (e.g. EQ, DRC, etc.) by the output sound effect processing module, and then the final binaural spatial audio signal is output.

[0232] Please refer to FIG. 15, the spatial input sound source includes multi-channel sound source and multi-object sound source. The spatial audio rendering method mainly includes two parts of spatial audio signal acquisition and spatial audio signal rendering, wherein FIG. 14 takes HOA as an example for illustration.

[0233] The spatial audio signal acquisition step is as follows:

[0234] 1. Different types of input sound sources enter the sound source grouping module, which groups the input sound sources according to their characteristics, including the direction of the sound source, the distance of the sound source, the content of the sound source, etc. to obtain N groups of input sound sources.

[0235] 2. Based on the respective HOA order corresponding to the N groups of input sound sources, the N groups of input sound sources are respectively subjected to HOA transformation to obtain N groups of HOA direct sound signals.

[0236] 3. The N groups of HOA direct sound signals respectively pass through the respective sound effect processing modules, and the gain adjustment and EQ adjustment are performed on the respective corresponding HOA direct sound signals according to the user-set sound effect processing parameters.

[0237] 4. After the N groups of HOA direct sound signals subjected to sound effect processing are superimposed by the adder, the direct sound effect processing module is subjected to direct sound effect processing to form the final HOA direct sound signal. If the HOA orders of the N groups of input sound sources are not the same, the signals corresponding to the channels of the same order are superimposed accordingly.

[0238] 5. The N groups of input sound sources simultaneously enter the reverberation grouping selection module, which reads the corresponding HOA reverberation data from the pre-stored reverberation data according to the sound source characteristics, room type and reverberation rendering order of the above grouping, and then adjusts the HOA reverberation data according to the spatial sense parameters and room size to obtain the HOA reverberation data corresponding to the N groups of input sound sources.

[0239] 6. Based on the HOA reverberation data corresponding to the N groups of input sound sources, the N groups of input sound sources are subjected to reverberation processing to obtain N groups of HOA reverberation sound signals, which are superimposed by the adder.

[0240] 7. The superimposed HOA reverberation sound signal is subjected to sound effect processing by the reverberation sound effect processing module to form the final HOA reverberation sound signal.

[0241] 8. The HOA direct sound signal and the HOA reverberation sound signal are superimposed again to form the final HOA audio signal, i.e. the spatial audio signal. Then, it is output to the spatial audio rendering module.

[0242] The spatial audio rendering module performs the rendering operation of the spatial audio signal similarly as in FIG. 14, and details are not repeated here.

[0243] Next, the spatial audio rendering method provided by the embodiment of the present application is illustrated by FIG. 16.

[0244] Referring to FIG. 16, the spatial input sound sources include multi-channel sound sources and multi-object sound sources. The multi-channel sound sources include five channels (represented by small black circles in FIG. 16), which are located at a near distance. The multi-object sound sources include two object sound sources (represented by triangles in FIG. 16) located at a middle distance, and two object sound sources (represented by hexagonal stars in FIG. 16) located at a far distance.

[0245] The obtaining of the spatial audio signal is as follows:

[0246] 1. Grouping the spatial input sound sources according to distance characteristics, i.e., the five multi-channel sound sources at a near distance form a group, the two object sound sources at a middle distance form a group, and the two object sound sources at a far distance form a group, for a total of three groups.

[0247] 2. According to the HOA order and sound effect processing parameters configured in the parameter configuration interface, sequentially performing HOA transformation and sound effect processing on the three groups of sound sources to obtain HOA direct sound signals of the three groups of sound sources. The HOA direct sound signals of the three groups of sound sources are superimposed to obtain a total HOA direct sound signal.

[0248] 3. Obtaining the room type and the reverberation rendering order configured in the parameter configuration interface, for example, the user selects a concert hall, and the order is 2 orders, then selecting 2-order HOA concert hall reverberation data of the near, middle, and far three groups from the concert hall reverberation data, and applying the 2-order HOA concert hall reverberation data of the three groups to the subsequent reverberation processing of the three groups of sound sources.

[0249] 4. Adjusting the three kinds of reverberation data selected above according to the spatial sense parameters and room sizes configured in the parameter configuration interface for the three groups of sound sources, for example, if the spatial sense of the sound sources at a near distance is to be reduced, the reverberation energy of the sound sources at a near distance can be reduced to achieve this purpose.

[0250] 5. Obtaining three groups of reverberation data through the third and fourth steps, and sending the three groups of reverberation data together with the corresponding three groups of sound sources to a reverberation processing module for reverberation processing to obtain HOA reverberation sound signals corresponding to the three groups of sound sources.

[0251] 6. The three groups of HOA reverberation sound signals are superimposed and subjected to reverberation sound effect processing by a reverberation sound effect processing module to form a total HOA reverberation sound signal. The total HOA reverberation sound signal is superimposed with the total HOA direct sound signal obtained in the second step and subjected to subsequent rendering.

[0252] FIG. 17 is a structural schematic diagram of a spatial audio rendering apparatus provided in an embodiment of the present application, which can be realized by software, hardware or a combination of both as part or all of an electronic device, which can be the audio processing terminal or the binaural playback terminal shown in FIG. 1. Referring to FIG. 17, the apparatus includes an audio signal obtaining module 1701, a rendering data obtaining module 1702 and an audio signal rendering module 1703.

[0253] The audio signal obtaining module 1701 is configured to obtain a spatial audio signal, the spatial audio signal being obtained by converting a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including a sound source of multiple channels, the multi-object sound source including a sound source of multiple sound emitting objects, the spatial audio signal being a first order ambisonic (FOA) audio signal or a higher order ambisonic (HOA) audio signal;

[0254] The rendering data obtaining module 1702 is configured to obtain spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal;

[0255] The audio signal rendering module 1703 is configured to render the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal.

[0256] In a possible implementation, the apparatus is included in the audio processing terminal;

[0257] The audio signal obtaining module 1701 is specifically configured to:

[0258] obtain a spatial input sound source;

[0259] perform ambisonic conversion on the spatial input sound source to obtain the spatial audio signal.

[0260] In a possible implementation, the audio signal obtaining module 1701 is specifically configured to:

[0261] perform ambisonic conversion on the spatial input sound source based on an ambisonic conversion order corresponding to the spatial input sound source to obtain a target direct sound signal;

[0262] determine the spatial audio signal based on the target direct sound signal.

[0263] In a possible implementation, the audio signal obtaining module 1701 is specifically configured to:

[0264] group the spatial input sound source to obtain at least one group of input sound sources;

[0265] perform stereo sound transformation on the at least one group of input sound sources respectively based on the stereo sound transformation order corresponding to each group of input sound sources, to obtain at least one direct sound signal;

[0266] superimpose the at least one direct sound signal to obtain a target direct sound signal.

[0267] In a possible implementation, the audio signal obtaining module 1701 is specifically configured to:

[0268] If the spatial input sound sources include multi-channel sound sources and multi-object sound sources, the multi-channel sound sources are taken as one group of input sound sources, and the multi-object sound sources are taken as one group of input sound sources; or,

[0269] group the spatial input sound sources according to sound source features to obtain at least one group of input sound sources, the sound source features being used to represent characteristics of the sound sources.

[0270] In a possible implementation, the sound source features include at least one of the following: sound source direction, sound source distance, and sound source content; the sound source content is used to indicate a type of a sound emitting object in the corresponding sound source.

[0271] In a possible implementation, the apparatus further includes:

[0272] an interface display module configured to display a first parameter configuration interface;

[0273] a transformation order obtaining module configured to obtain, from the first parameter configuration interface, the stereo sound transformation order corresponding to each group of input sound sources.

[0274] In a possible implementation, the apparatus further includes:

[0275] a direct sound effect processing module configured to perform sound effect processing on the at least one direct sound signal based on the direct sound effect processing parameter corresponding to each group of input sound sources.

[0276] In a possible implementation, the apparatus further includes:

[0277] an interface display module configured to display a second parameter configuration interface;

[0278] a sound effect processing parameter obtaining module configured to obtain, from the second parameter configuration interface, the direct sound effect processing parameter corresponding to each group of input sound sources.

[0279] In a possible implementation, the apparatus further includes:

[0280] a reverberation processing module configured to perform reverberation processing on the spatial input sound sources based on reverberation data corresponding to the spatial input sound sources to obtain a target reverberation sound signal, the reverberation data being FOA reverberation data or HOA reverberation data.

[0281] The audio signal obtaining module is specifically configured to:

[0282] superimpose the target direct sound signal and the target reverberation sound signal to obtain a spatial audio signal.

[0283] In a possible implementation manner, the spatial input sound sources are divided into at least one group of input sound sources;

[0284] The reverberation processing module is specifically configured to:

[0285] perform reverberation processing on the at least one group of input sound sources respectively based on the reverberation data corresponding to the at least one group of input sound sources respectively, to obtain at least one reverberation sound signal;

[0286] superimpose the at least one reverberation sound signal to obtain a target reverberation sound signal.

[0287] In a possible implementation manner, the apparatus further includes:

[0288] The reverberation data obtaining module is configured to obtain reverberation data corresponding to the at least one group of input sound sources respectively based on reverberation parameters corresponding to the at least one group of input sound sources respectively.

[0289] In a possible implementation manner, the reverberation parameters include a reverberation sound rendering order and a room type, and the room type is used to indicate a type of a room currently simulated by the reverberation sound.

[0290] The reverberation data obtaining module is specifically configured to:

[0291] obtain the reverberation data corresponding to the at least one group of input sound sources respectively based on the reverberation sound rendering order and the room type corresponding to the at least one group of input sound sources respectively.

[0292] In a possible implementation manner, the reverberation parameters further include a spatial sense parameter and / or a room size, and the spatial sense parameter is used to describe spatial information in the reverberation sound.

[0293] The apparatus further includes:

[0294] The reverberation data adjusting module is configured to adjust the reverberation data corresponding to the at least one group of input sound sources respectively based on a spatial sense parameter and / or a room size corresponding to the at least one group of input sound sources respectively.

[0295] In a possible implementation manner, the at least one group of input sound sources is obtained by dividing the input sound sources based on sound source features.

[0296] The reverberation data obtaining module is specifically configured to:

[0297] The reverberation data corresponding to the at least one group of input sound sources is obtained based on the reverberation rendering order, the room type and the sound source characteristics corresponding to the at least one group of input sound sources respectively.

[0298] In a possible implementation, the apparatus is included in a binaural playback terminal.

[0299] The audio signal obtaining module is specifically configured to:

[0300] The spatial audio signal is received from a binaural playback terminal.

[0301] In a possible implementation, the apparatus further includes:

[0302] The head movement information obtaining module is configured to obtain head movement information, the head movement information indicating a posture of a head wearing the binaural playback terminal.

[0303] The rotation processing module is configured to perform rotation processing on the spatial audio signal based on the head movement information.

[0304] The embodiments of the present application convert different formats of input sound sources into FOA / HOA formats through stereo sound transformation, which not only preserves all the spatial information of the original input sound sources, but also converts the number of sound channels and / or the number of sound objects of the spatial input sound sources into the number of channels of the FOA / HOA format audio signal, which is only related to the order of the audio signal and is irrelevant to the number of sound channels and / or the number of sound objects of the spatial input sound sources. In this way, even if the number of sound channels and / or the number of sound objects of the spatial input sound sources increases, the number of channels of the FOA / HOA format audio signal of the same order remains unchanged, thereby reducing the computational complexity of subsequent signal processing. In addition, the rendering of the FOA / HOA format audio signal does not need to store discrete spatial rendering data of each direction in the database, but stores spatial rendering data of the FOA / HOA domain. The spatial rendering data of the FOA / HOA domain is a distribution expression of the discrete spatial rendering data in three-dimensional space, and through the expression, the spatial rendering data of any direction in three-dimensional space can be determined. That is, the database does not need to store discrete spatial rendering data of each direction, but stores the expression of the spatial rendering data of the FOA / HOA domain, which greatly reduces the occupation of the memory space. Moreover, the spatial rendering data of the FOA / HOA domain is a continuous expression in three-dimensional space, and after the spatial rendering data is determined through the expression, the audio output signal rendered in the process of head rotation is relatively smooth, thereby improving the naturalness and continuity of the listening experience in the head tracking mode. Furthermore, the above parameter configuration interface supports customization of parameters, improves the flexibility of autonomous selection of parameters, and thus realizes customized stereo sound transformation effect and spatial rendering effect.

[0305] It should be noted that the spatial audio rendering apparatus provided in the above embodiments is only used for example to illustrate the division of the above functional modules when performing spatial audio rendering. In actual applications, the above functions can be distributed to different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the spatial audio rendering apparatus and the spatial audio rendering method provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.

[0306] FIG. 18 is a structural schematic diagram of a spatial audio rendering apparatus provided in an embodiment of the present application, which can be realized by software, hardware or a combination of both to become part or all of an audio processing terminal. Referring to FIG. 18, the apparatus includes an input sound source acquisition module 1801, a binaural conversion module 1802 and an audio signal sending module 1803.

[0307] The input sound source acquisition module 1801 is configured to acquire a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including a sound source of multiple channels, and the multi-object sound source including a sound source of multiple sound emitting objects.

[0308] The binaural conversion module 1802 is configured to perform binaural conversion on the spatial input sound source to obtain a spatial audio signal, the spatial audio signal being a first-order binaural sound FOA audio signal or a high-order binaural sound HOA audio signal.

[0309] The audio signal sending module 1803 is configured to send the spatial audio signal to a binaural playback terminal, so that the binaural playback terminal performs rendering on the spatial audio signal based on spatial rendering data to obtain a binaural spatial audio output signal, wherein the spatial rendering data is FOA rendering data or HOA rendering data, and the order of the spatial rendering data is the same as the order of the spatial audio signal.

[0310] The embodiments of the present application convert input sound sources of different formats into FOA / HOA format through stereo sound conversion, which not only preserves all spatial information of the original input sound sources, but also converts the number of sound channels and / or the number of sound objects of the spatial input sound sources into the number of channels of the FOA / HOA format audio signal, which is only related to the order of the audio signal and is irrelevant to the number of sound channels and / or the number of sound objects of the spatial input sound sources. In this way, even if the number of sound channels and / or the number of sound objects of the spatial input sound sources increases, the number of channels of the FOA / HOA format audio signal of the same order remains unchanged, thereby reducing the computational complexity of subsequent signal processing. In addition, the rendering of the FOA / HOA format audio signal does not need to store discrete spatial rendering data of each direction in the database, but stores spatial rendering data of the FOA / HOA domain, which is a distributed expression of discrete spatial rendering data in three-dimensional space, and through which the spatial rendering data of any direction in three-dimensional space can be determined. That is, the database does not need to store discrete spatial rendering data of each direction, but stores the expression of the spatial rendering data of the FOA / HOA domain, which greatly reduces the occupation of memory space. Moreover, the spatial rendering data of the FOA / HOA domain is a continuous expression of three-dimensional space, and after the spatial rendering data is determined through this expression, the audio output signal rendered in the process of head rotation is relatively smooth, thereby improving the naturalness and continuity of the listening experience in the head tracking mode. Furthermore, the above parameter configuration interface supports customization of parameters, improves the flexibility of autonomous selection of parameters, and thus realizes customized stereo sound conversion effect and spatial rendering effect.

[0311] It should be noted that the spatial audio rendering device provided by the above embodiments is only exemplified by the division of the above functional modules during spatial audio rendering. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the spatial audio rendering device and the spatial audio rendering method provided by the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0312] The embodiments of the present application also provide a computer readable storage medium, wherein instructions are stored in the storage medium, and when the instructions run on a computer, the computer executes the above method.

[0313] The embodiments of the present application also provide a computer program product containing instructions, and when the instructions run on a computer, the computer executes the above method.

[0314] In the above embodiments, all or part can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example: coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example: infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example: floppy disk, hard disk, magnetic tape), optical media (for example: digital versatile disc (DVD)) or semiconductor media (for example: solid state disk (SSD)) and the like. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.

[0315] It should be understood that "multiple" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" herein only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, "first", "second" and the like are used to distinguish the same items or similar items with basically the same function and role in the embodiments of the present application. Those skilled in the art can understand that "first", "second" and the like do not limit the quantity and execution order, and "first", "second" and the like do not necessarily mean different.

[0316] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0317] The above describes the embodiments provided by the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A spatial audio rendering method, characterized by, The method comprises: obtaining a spatial audio signal, the spatial audio signal being converted from a spatial input sound source, the spatial input sound source comprising a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source comprising a sound source of multiple channels, and the multi-object sound source comprising a sound source of multiple sound emitting objects, the spatial audio signal being a first order ambisonic (FOA) audio signal or a higher order ambisonic (HOA) audio signal; obtaining spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal; based on the spatial rendering data, rendering the spatial audio signal to obtain a binaural spatial audio output signal.

2. The method of claim 1, wherein, The method is applied to an audio processing terminal. The method comprises: obtaining the spatial input sound source; performing ambisonic conversion on the spatial input sound source to obtain the spatial audio signal.

3. The method of claim 2, wherein, The method comprises: based on the ambisonic conversion order corresponding to the spatial input sound source, performing ambisonic conversion on the spatial input sound source to obtain a target direct sound signal; based on the target direct sound signal, determining the spatial audio signal.

4. The method of claim 3, wherein, The method comprises: grouping the spatial input sound source to obtain at least one group of input sound sources; based on the ambisonic conversion order corresponding to each of the at least one group of input sound sources, performing ambisonic conversion on each of the at least one group of input sound sources to obtain at least one direct sound signal; superimposing the at least one direct sound signal to obtain the target direct sound signal.

5. The method of claim 4, wherein, The method comprises: if the spatial input sound source comprises a multi-channel sound source and a multi-object sound source, taking the multi-channel sound source as one group of input sound sources and taking the multi-object sound source as one group of input sound sources; or grouping the spatial input sound source according to a sound source feature to obtain the at least one group of input sound sources, the sound source feature being used to represent the characteristics of the sound source.

6. The method of claim 5, wherein, The sound source feature comprises at least one of the following: a sound source direction, a sound source distance, and a sound source content, wherein the sound source content is used to indicate the type of sound emitting object in the corresponding sound source.

7. The method according to any one of claims 4 to 6, wherein, Before the method performs ambisonic conversion on each of the at least one group of input sound sources based on the ambisonic conversion order corresponding to each of the at least one group of input sound sources, the method further comprises: displaying a first parameter configuration interface; obtaining the ambisonic conversion order corresponding to each of the at least one group of input sound sources from the first parameter configuration interface.

8. The method according to any one of claims 4 to 7, wherein, Before the method superimposes the at least one direct sound signal, the method further comprises: based on the direct sound effect processing parameters corresponding to each of the at least one group of input sound sources, performing sound effect processing on the at least one direct sound signal.

9. The method of claim 8, wherein, The method further comprises: displaying a second parameter configuration interface; Obtaining, from the second parameter configuration interface, direct sound effect processing parameters corresponding to the at least one group of input sound sources respectively.

10. The method of any one of claims 3-9, wherein, The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; 11. The method of claim 10, wherein, The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; 12. The method of claim 11, wherein, The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; 13. The method of claim 12, wherein, The reverberation parameters include a reverberation sound rendering order and a room type, and the room type is used to indicate a type of a room currently simulated by the reverberation sound; The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; 14. The method of claim 13, wherein, The method further includes: Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; The method further includes:

15. The method of claim 13 or 14, wherein, Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; The at least one group of input sound sources is divided by a sound source feature; The method further includes:

16. The method of claim 1, wherein, Based on the reverberation data corresponding to the spatial input sound source, the spatial input sound source is reverberated to obtain a target reverberation sound signal, and the reverberation data is FOA reverberation data or HOA reverberation data; The method is applied to a binaural playback terminal; The method further includes:

17. The method of any one of claims 1-16, wherein, Receiving the spatial audio signal sent by an audio processing terminal. The method further includes: The method further includes: Head movement information is acquired, the head movement information indicating a posture of a head wearing a binaural playback terminal; Based on the head movement information, the spatial audio signal is processed by rotation.

18. A spatial audio rendering method, characterized by, The method is applied to an audio processing terminal, and includes: A spatial input sound source is acquired, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including a sound source of multiple channels, and the multi-object sound source including a sound source of multiple sound emitting objects; The spatial input sound source is processed by stereo sound conversion to obtain a spatial audio signal, the spatial audio signal being a first-order stereo sound (FOA) audio signal or a high-order stereo sound (HOA) audio signal; The spatial audio signal is sent to a binaural playback terminal, so that the binaural playback terminal renders the spatial audio signal based on spatial rendering data to obtain a binaural spatial audio output signal; wherein the spatial rendering data is FOA rendering data or HOA rendering data, and the order of the spatial rendering data is the same as the order of the spatial audio signal.

19. A spatial audio rendering apparatus, characterized by The device includes: An audio signal acquisition module is configured to acquire a spatial audio signal, the spatial audio signal being obtained by conversion of a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including a sound source of multiple channels, and the multi-object sound source including a sound source of multiple sound emitting objects, the spatial audio signal being a first-order stereo sound (FOA) audio signal or a high-order stereo sound (HOA) audio signal; A rendering data acquisition module is configured to acquire spatial rendering data, the spatial rendering data being FOA rendering data or HOA rendering data, and the order of the spatial rendering data being the same as the order of the spatial audio signal; An audio signal rendering module is configured to render the spatial audio signal based on the spatial rendering data to obtain a binaural spatial audio output signal.

20. The apparatus of claim 19, wherein, The device is included in an audio processing terminal; The audio signal acquisition module is specifically configured to: Acquire the spatial input sound source; Process the spatial input sound source by stereo sound conversion to obtain the spatial audio signal.

21. The apparatus of claim 20, wherein, The audio signal acquisition module is specifically configured to: Process the spatial input sound source by stereo sound conversion based on a stereo sound conversion order corresponding to the spatial input sound source to obtain a target direct sound signal; Determine the spatial audio signal based on the target direct sound signal.

22. The apparatus of claim 21, wherein, The audio signal acquisition module is specifically configured to: Group the spatial input sound source to obtain at least one group of input sound sources; Process the at least one group of input sound sources by stereo sound conversion based on a stereo sound conversion order corresponding to each of the at least one group of input sound sources to obtain at least one direct sound signal; Superimpose the at least one direct sound signal to obtain the target direct sound signal.

23. The apparatus of claim 22, wherein, The audio signal acquisition module is specifically configured to: If the spatial input sound source includes a multi-channel sound source and a multi-object sound source, the multi-channel sound source is taken as a group of input sound sources, and the multi-object sound source is taken as a group of input sound sources; Or, The spatial input sound sources are grouped according to source characteristics to obtain the at least one group of input sound sources, and the source characteristics are used to represent characteristics of the sound sources.

24. The apparatus of claim 23, wherein, The source characteristics include at least one of a source direction, a source distance, and a source content, and the source content is used to indicate a type of a sound object in the corresponding sound source.

25. The apparatus of any one of claims 22-24, wherein, The device further includes: an interface display module configured to display a first parameter configuration interface; a transform order acquisition module configured to acquire, from the first parameter configuration interface, a stereo sound transform order corresponding to each of the at least one group of input sound sources.

26. The apparatus of any one of claims 22-25, wherein, The device further includes: a direct sound effect processing module configured to perform sound effect processing on the at least one direct sound signal based on direct sound effect processing parameters corresponding to each of the at least one group of input sound sources.

27. The apparatus of claim 26, wherein, The device further includes: an interface display module configured to display a second parameter configuration interface; a sound effect processing parameter acquisition module configured to acquire, from the second parameter configuration interface, direct sound effect processing parameters corresponding to each of the at least one group of input sound sources.

28. The apparatus of any one of claims 21-27, wherein, The device further includes: a reverberation processing module configured to perform reverberation processing on the spatial input sound sources based on reverberation data corresponding to the spatial input sound sources to obtain a target reverberation sound signal, the reverberation data being FOA reverberation data or HOA reverberation data; The audio signal acquisition module is specifically configured to: superimpose the target direct sound signal and the target reverberation sound signal to obtain the spatial audio signal.

29. The apparatus of claim 28, wherein, The spatial input sound sources are divided into at least one group of input sound sources; The reverberation processing module is specifically configured to: perform reverberation processing on each of the at least one group of input sound sources based on reverberation data corresponding to each of the at least one group of input sound sources to obtain at least one reverberation sound signal; and superimpose the at least one reverberation sound signal to obtain the target reverberation sound signal.

30. The apparatus of claim 29, wherein, The device further includes: a reverberation data acquisition module configured to acquire, based on reverberation parameters corresponding to each of the at least one group of input sound sources, reverberation data corresponding to each of the at least one group of input sound sources.

31. The apparatus of claim 30, wherein, The reverberation parameters include a reverberation sound rendering order and a room type, and the room type is used to indicate a type of a room currently simulated by the reverberation sound; The reverberation data acquisition module is specifically configured to: acquire, based on the reverberation sound rendering order and the room type corresponding to each of the at least one group of input sound sources, the reverberation data corresponding to each of the at least one group of input sound sources.

32. The apparatus of claim 31, wherein, The reverberation parameters further include a spatial sense parameter and / or a room size, and the spatial sense parameter is used to describe spatial information in the reverberation sound; The device further includes: a reverberation data adjustment module configured to adjust, based on the spatial sense parameter and / or the room size corresponding to each of the at least one group of input sound sources, the reverberation data corresponding to each of the at least one group of input sound sources.

33. The apparatus of claim 31 or 32, wherein, The at least one group of input sound sources is obtained by grouping according to source characteristics; The reverberation data acquisition module is specifically configured to: acquire, based on the reverberation sound rendering order, the room type, and the source characteristics corresponding to each of the at least one group of input sound sources, the reverberation data corresponding to each of the at least one group of input sound sources.

34. The apparatus of claim 19, wherein, The device is included in a binaural playback terminal; The audio signal acquisition module is specifically configured to: Receive the spatial audio signal sent by the audio processing terminal.

35. The apparatus of any one of claims 19-34, wherein, The device further comprises: A head movement information acquisition module is configured to acquire head movement information, the head movement information indicating a posture of a head wearing a binaural playback terminal. A rotation processing module is configured to perform rotation processing on the spatial audio signal based on the head movement information.

36. A spatial audio rendering apparatus, characterized by The device is included in an audio processing terminal, and the device comprises: An input sound source acquisition module is configured to acquire a spatial input sound source, the spatial input sound source including a multi-channel sound source and / or a multi-object sound source, the multi-channel sound source including a sound source of multiple channels, and the multi-object sound source including a sound source of multiple sound emitting objects. A stereo sound transformation module is configured to perform stereo sound transformation on the spatial input sound source to obtain a spatial audio signal, the spatial audio signal being a first-order stereo sound (FOA) audio signal or a high-order stereo sound (HOA) audio signal. An audio signal sending module is configured to send the spatial audio signal to a binaural playback terminal, so that the binaural playback terminal performs rendering on the spatial audio signal based on spatial rendering data to obtain a binaural spatial audio output signal, wherein the spatial rendering data is FOA rendering data or HOA rendering data, and the order of the spatial rendering data is the same as the order of the spatial audio signal.

37. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory storing a computer program, and the processor is configured to execute the computer program to execute the method of any one of claims 1-17.

38. An audio processing terminal, characterized by The audio processing terminal comprises a memory and a processor, the memory storing a computer program, and the processor is configured to execute the computer program to execute the method of claim 18.

39. A computer-readable storage medium, characterized in that, The storage medium stores instructions, when the instructions run on the computer, causing the computer to execute the method of any one of claims 1-18.

40. A computer program product comprising instructions, wherein: When the instructions run on the computer, the computer is caused to execute the method of any one of claims 1-18.

Citation Information

Patent Citations

  • Obtaining symmetry information for higher order ambisonic audio renderers

    CN106465029A

  • Audio signal processing method and apparatus using ambisonics signal

    CN111034225A

  • Binauralization of rotated higher order ambisonics

    US20140355766A1

  • Method and apparatus for enhancing directivity of a 1st order ambisonics signal

    US20160057556A1

  • Apparatus, Method or Computer Program for Synthesizing a Spatially Extended Sound Source Using Variance or Covariance Data

    US20240284132A1