Audio processing method and apparatus

The Hierarchical Sound Field Representation (HSR) addresses compatibility and efficiency issues in spatial audio split rendering by defining IR formats, ensuring low-latency and adaptable processing on resource-constrained devices.

WO2025232856A1PCT designated stage Publication Date: 2025-11-13DOUYIN VISION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/093565
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-05-08
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing spatial audio processing technologies face challenges in ensuring compatibility and efficient transmission of intermediate representations (IRs) in split rendering scenarios, leading to issues with data latency, bandwidth consumption, and hardware/software modularization in resource-constrained devices.

Method used

The implementation of a Hierarchical Sound Field Representation (HSR) as an IR format that defines supported degrees of freedom, metadata, and data structure, enabling versatile and compatible rendering across different stages of the split rendering pipeline.

Benefits of technology

HSR ensures low-latency, low-bandwidth IR transmission and adapts to varying hardware capabilities, enhancing the effectiveness of spatial audio processing on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025093565_13112025_PF_FP_ABST
    Figure CN2025093565_13112025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of audio processing technology, and in particular to an audio processing method, an audio processing apparatus, a computer-readable storage medium and a computer program product. The audio processing method includes: receiving IR information of audio data transmitted by a first renderer; determining, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; and processing, based on the corresponding rendering algorithm, the IR information by using a second renderer.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO PROCESSING METHOD AND APPARATUSCROSS-REFERENCE TO RELATED APPLICATIONSThis application claims priority to PCT Patent Application No. PCT / CN2024 / 092222, and filed on May 10, 2024. The entire disclosure of the prior application is hereby incorporated by reference in its entirety.TECHNICAL FIELDThis disclosure relates to the field of audio processing technology, in particular to an audio processing method, an audio processing apparatus, an electronic device, a computer-readable storage medium and a computer program product.BACKGROUNDThe IVAS codec is an extension of the 3GPP Enhanced Voice Services (EVS) codec. It provides full and bit exact EVS codec functionality for mono speech / audio signal input. It further provides:1. Encoding and decoding of stereo and immersive audio formats such as multi-channel audio, scene-based audio (Ambisonics) , metadata-assisted spatial audio (MASA) , object-based audio (ISM) , and their combination;2. VAD / DTX / CNG for rate efficient stereo and immersive conversational voice transmissions;3. Error concealment mechanisms to combat the effects of transmission errors and lost packets. Jitter buffer management is also provided;4. The IVAS codec operates on 20-ms audio frames. In addition, rendering is possible with 5-ms granularity;5. Support for bit rate switching upon command;6. Stereo and immersive audio coding at the following discrete bit rates [kbps] : 13.2, 16.4, 24.4, 32, 48, 64, 80, 128, 160, 192, 256, 384, and 512, with supported bit rate ranges listed in table 1.Table 1 Ranges of supported bitrates for stereo and immersive coding of the IVAS codec13.2 kbps-128 kbps for 1 ISM, 16.4 kbps-256 kbps for 2 ISMs, 24.4 kbps-384 kbps for 3 ISMs, 24.4 kbps-512 kbps for 4 ISMs.The codec for Immersive Voice and Audio Services is part of a framework including of an encoder, decoder, and renderer.SUMMARYAccording to some embodiments of the present disclosure, there is provided an audio processing method, including: receiving IR (Intermediate Representation) information of audio data transmitted by a first renderer; determining, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; and processing, based on the corresponding rendering algorithm, the IR information by using a second renderer.In some embodiments, the IR information includes supported DoF (degrees of freedom) of a listener for the audio data, and the determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information includes: identifying the sound field structure based on the supported DoF.In some embodiments, the identifying the sound field structure based on the supported DoF includes: identifying the sound field structure based on metadata of the IR information, the metadata including at least one of:a sample rate of the audio data, a type of the IR information, a data channel offset of the IR information in a multi-channel data stream, or sound field transform information of the IR information.In some embodiments, the type of the IR information is configured to indicate a channel number and / or a channel order of the IR information.In some embodiments, the second renderer supports a data structure of the IR information, and the determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information includes: obtaining the corresponding rendering algorithm from an algorithm database.In some embodiments, the second renderer does not support a data structure of the IR information, and the determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information includes: determining a Downmix operator and / or Mute operator as the corresponding rendering algorithm.In some embodiments, the audio processing method further includes: sending an exception message to a local logger and / or to the first renderer, the exception message indicating that the second renderer does not support the data structure of the IR information.In some embodiments, the processing, based on the corresponding rendering algorithm, the IR information by using the second renderer includes: processing the IR information based on listener pose information for the audio data.In some embodiments, the listener pose information includes listener position information and / or orientation information.According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, including: a receiving unit configured to receive IR information of audio data transmitted by a first renderer; a determining unit configured to determine, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; and a processing unit configured to process based on the corresponding rendering algorithm, the IR information by using a second renderer.In some embodiments, the IR information includes supported DoF (degrees of freedom) of a listener for the audio data, and the determining unit identifies the sound field structure based on the supported DoF.In some embodiments, the determining unit identifies the sound field structure based on metadata of the IR information, the metadata including at least one of: a sample rate of the audio data, a type of the IR information, a data channel offset of the IR information in a multi-channel data stream, or sound field transform information of the IR information.In some embodiments, the type of the IR information is configured to indicate a channel number and / or a channel order of the IR information.In some embodiments, the second renderer supports a data structure of the IR information, and the determining unit obtains the corresponding rendering algorithm from an algorithm database.In some embodiments, the second renderer does not support a data structure of the IR information, and the determining unit determines a Downmix operator and / or Mute operator as the corresponding rendering algorithm.In some embodiments, the processing unit sends an exception message to a local logger and / or to the first renderer, the exception message indicating that the second renderer does not support the data structure of the IR information.In some embodiments, the processing unit processes the IR information based on listener pose information for the audio data.In some embodiments, the listener pose information includes listener position information and / or orientation information.According to some embodiments of the present disclosure, there is provided an audio processing method, including: loading an appropriate rendering algorithm; and processing an input HSR using the appropriate rendering algorithm.In some embodiments, the loading the appropriate rendering algorithm includes: identifying a structure of a sound field to load the appropriate rendering algorithm.In some embodiments, the s processing an input HSR using the appropriate rendering algorithm including: processing an input HSR using a current listener pose.According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, including: loading module for loading an appropriate rendering algorithm; and processing module for processing an input HSR using the appropriate rendering algorithm.In some embodiments, the loading module identifies a structure of a sound field to load the appropriate rendering algorithm.In some embodiments, the processing module processes an input HSR using a current listener pose.According to still other embodiments of the present disclosure, there is provided an electronic device, including: a memory; a processor coupled to the memory, the processor configured to, based on instructions stored in the memory, carry out the audio processing method according to any one of the above embodiments.According to still other embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements the audio processing method according to any one of the above embodiments.According to still other embodiments of the present disclosure, there is provided a computer program product, including: instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of the above embodiments.BRIEF DESCRIPTION OF THE DRAWINGSThe accompanying drawings, which are incorporated in and constitute a portion of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.The present disclosure will be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:FIG. 1 shows schematic diagrams of overview of the audio processing functions of the receive side of the codec;FIG. 2 shows schematic diagrams of overview of TD binaural renderer;FIG. 3 shows schematic diagrams of structure of an HSR Node;FIG. 4 shows schematic diagrams of Spatial Audio Split Rendering Pipeline using HSR as IR;FIG. 5A shows schematic diagrams of a 0dof HSR that contains single-perspective binaural audio;FIG. 5B shows schematic diagrams of a spatial audio split rendering pipeline using a single-perspective binaural HSR;FIG. 6A shows schematic diagrams of a 3dof HSR that contains a multichannel signal locked to listeners position;FIG. 6B shows schematic diagrams of a spatial audio split rendering pipeline using this single-position Multichannel HSR;FIG. 7A shows schematic diagrams of a 3dof HSR that contains a 3-th order Ambisonics signal rotated to a given world orientation;FIG. 7B shows schematic diagrams of a spatial audio split rendering pipeline using this single-position Ambisonic HSR;FIG. 8A shows schematic diagrams of a 3dof combined HSR;FIG. 8B shows schematic diagrams of a spatial audio split rendering pipeline using this multi-perspective binaural HSR;FIG. 9A shows schematic diagrams of a 6dof combined HSR is as follows. This HSR is formed by 4 single-perspective Ambisonic HSRs;FIG. 9B shows schematic diagrams of a spatial audio split rendering pipeline using this multi-position Ambisonics HSR;FIG. 10A shows schematic diagrams of a 6dof combined HSR;FIG. 10B shows schematic diagrams of a spatial audio split rendering pipeline using this mixed HSR;FIG. 11 shows a block diagram of the electronic device according to other embodiments of the present disclosure;FIG. 12 shows a block diagram of the electronic device according to further embodiments of the present disclosure;FIG. 13a shows a flow diagram of the audio processing method according to some embodiments of the present disclosure; andFIG. 13b shows a block diagram of the audio processing apparatus according to some embodiments of the present disclosure.DETAILED DESCRIPTIONVarious exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Notice that, unless otherwise specified, the relative arrangement, numerical expressions and numerical values of the components and steps set forth in these examples do not limit the scope of the disclosure.At the same time, it should be understood that, for ease of description, the dimensions of the plurality of parts shown in the drawings are not drawn to actual proportions.The following description of at least one exemplary embodiment is in fact merely illustrative and is in no way intended as a limitation to the disclosure, its application or use.Techniques, methods, and apparatus known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, these techniques, methods, and apparatuses should be considered as part of the specification.Of all the examples shown and discussed herein, any specific value should be construed as merely illustrative and not as a limitation. Thus, other examples of exemplary embodiments may have different values.Notice that, similar reference numerals and letters are denoted by the like in the accompanying drawings, and therefore, once an item is defined in a drawing, there is no need for further discussion in the accompanying drawings.The terms used in this disclosure are described as following:Application: an application is a service enabler deployed by service providers, manufacturers or users. Individual applications will often be enablers for a wide range of services;Receiver side: in end to end media system, 3 steps should be implemented, capturing and encoding, network transmission, decoding and rendering, the 3rd step is usually called receiver side;Loudspeaker reproduction: Use one or more speakers for audio playback. Common loudspeaker setup includes stereo, 5.1, 7.1.4, 22.2, which specifies the number of loudspeakers as well as azimuth and elevation angles of each loudspeaker;Headphone reproduction: Use headphone for audio playback;Room acoustics: Room acoustics is the simulation of how sound behaves in an enclosed space, influencing sound quality through early reflections and late reflections;Head tracking: Dynamic tracking of the user's head posture when playing back audio using headphones;Sound field: Sound field refers to the distribution and characteristics of sound waves in a given space or environment. It includes factors such as sound intensity, directionality, reflections, and reverberation within the area of interest;Motion-to-Sound Latency: The amount of time it takes between user's motion takes place and the output signal of spatial audio rendering system is changed.

[0023] The abbreviations used in this disclosure are described as following:3GPP: Third Generation Partnership Project;AR: Augmented Reality;DoF: Degrees of Freedom;FOA: First-Order Ambisonics;HMD: Head-Mounted Display;HOA: Higher Order Ambisonics;HRIR: Head-Related Impulse Response;HRTF: Head-Related Transfer Function;HSR: Hierarchical Sound-field Representation;IR: Intermediate Representation: a sound field representation being transmitted between different devices in ISAR framework;ISAR: Immersive Audio for Split Rendering Senarios;ISM: Independent Streams with Metadata;ITD: Inter-Channel Time Delay;IVAS: Immersive Voice and Audio Services;LFE: Low-Frequency Effects;MASA: Metadata-Assisted Spatial Audio;MC: Multi Channel;McMASA: Multi-Channel MASA;MCT: Multi-channel Coding Tool;OSBA: Objects with SBA;OTT: Over-The-Top;PLC: Packet loss concealment;QoE: Quality of Experience;QoS: Quality of Service;SBA: Scene-Based Audio;UE: User Equipment;VBAP: Vector Base Amplitude Panning;VBR: Variable Bit Rate;XR: eXtended Reality;XY + M A: 2dof ambisonic format, which is formed by FOA minus vertical (Z axis) vibration modality.An overview of the audio processing functions of the receive side of the codec is shown in Fig. 1, with rendering features highlighted.In Fig. 1, the interfaces are marked consistently with [5] using the following numbers:3: Encoded audio frames (50 frames / s) , number of bits depending on IVAS codec mode;4: Encoded Silence Insertion Descriptor (SID) frames;5: RTP Payload packets;6: Lost Frame Indicator (BFI) ;7: Renderer config data;8: Head-tracker pose information and scene orientation control data;9: Audio output channels (16-bit linear PCM, sampled at 8 (only EVS) , 16, 32, or 48 kHz) ;10: Metadata associated with output audio.For Multi-Channel (MC) Operation, coding of multi-channel inputs is available for the channel layouts 5.1, 7.1, 5.1+2, 5.1+4, and 7.1+4. The coding technique is selected from a set of coding modes based on the available bitrate and specified channel layout. The general principle in technique selection is to aim for best possible quality given the allowed bitrate. For all techniques, LFE channel coding is also offered either separately or within the technique. The multi-channel operation supports output to mono, stereo, multi-channel (at the same or any other layout with up to 16 speakers) , Ambisonics (at up to order 3) , and binaural.For Scene-based Audio (Ambisonics) Operation, Coding of ambisonics signals is supported for 1st-to 3rd-order inputs throughout the full bitrate range. The decoder's output can be SBA (of order 1, 2, or 3) , mono, stereo, binaural or multi-channel. This flexibility of input order, bitrate and output format combinations is in part achieved by the combination of covariance-based and directional analysis at different frequencies.For the lowest 8 bands, covariance analysis with a 20 ms stride is performed at the encoder and corresponding reconstruction is performed at the decoder. For the highest 4 bands, an estimation of the parameters of a psychoacoustic model is implemented with a time resolution of 5 ms.At the encoder, the covariance-analysis metadata for the higher bands are estimated from these model parameters and combined with the directly calculated metadata for the lower bands. Based on these metadata, a downmix to 1 to 4 channels (dependent on bitrate) is obtained. The downmix channels are then coded with the appropriate core coder.At the decoder, the downmix channels plus the metadata are received. The latter include the transmitted covariance-analysis metadata and model parameters for the lower and higher bands, respectively. These metadata are used to reconstruct the HOA signal and render to the requested output format. In this, the psychoacoustic model parameters for the 8 lower frequency bands are estimated from the reconstructed audio channels. The model-based reconstruction allows for the output SBA order on the decoder side to be higher than the input order on the encoder side.Low-latency operation (less than or equal to 38 ms end-to-end) is achieved by using a very-low-latency 1-ms MDFT-based filterbank at the encoder and a 5-ms CLDFB-based filterbank at the decoder, which additionally enables the type of signal modifications that are necessary for rendering to a broad variety of output configuration.For Metadata-assisted Spatial Audio (MASA) Operation, the IVAS codec supports coding of parametric spatial audio format called metadata-assisted spatial audio (MASA) . This format is specifically optimized for the direct immersive audio capture from smartphones and other form factors that can be unsuitable for dedicated spherical microphone arrays.The MASA format is based on 1-2 audio channels and associated metadata that is provided for each audio frame. The MASA spatial metadata describes the spatial audio characteristics of the captured immersive audio using several spatial parameters including spatial direction information, directional and non-directional energy ratios, and two types of coherence information. The spatial metadata is provided in each frame according to a time-frequency resolution of 4 subframes and 24 frequency bands. The MASA descriptive metadata provides additional information relating to the creation and understanding of the MASA audio signal.The coding in this operation is based on compression of the metadata exploiting detected redundancies and prioritization of selected parameters at each bitrate. The 1 or 2 audio transport channels are coded using the SCE and CPE coding capabilities. A joint bitrate allocation between these two coding blocks is based on a metadata analysis and simplification processing.The MASA format can be flexibly rendered for binaural or loudspeaker reproduction, including mono and stereo playback. Rendering to Ambisonics is also supported. Furthermore, decoded MASA format bitstream can be directly output from the decoder without rendering as a fully compliant MASA format output for further processing.Rendering is the process of generating digital audio output from the decoded digital audio signal. Rendering is used when output format is different than input format. In case output format is the same as input format, the decoded audio channels are simply passed through to the output channels. Binaural rendering is a special case, where binaural output channels are prepared for headphone reproduction. This process includes head-tracking and scene orientation control, head-related transfer function processing, and room acoustic synthesis. IVAS rendering is integrated with IVAS decoder but can also be operated standalone as external rendering while bypassing the internal renderer. The external renderer can be applied e.g., in the case of rendering outputs originating from multiple sources, such as decoders or audio streams.For rendering for loudspeaker reproduction, Digital audio decoded by IVAS decoder can be rendered for loudspeaker reproduction. The process of rendering depends on the decoded audio format. In case of multi-channel formats, the decoded format can match the loudspeaker configuration or output channels can be generated by application of multi-channel conversion gains from conversion tables. For SBA, MASA and ISM formats the spatial audio needs to be mapped to the loudspeaker positions of the loudspeaker setup (e.g., 5.1 or 7.1.4) . Depending on the decoded format amplitude panning is employed, with either the vector-base amplitude panning (VBAP) scheme (using triangles) , or an improved edge-fading amplitude panning (EFAP) scheme (using polygons) . Specifically for SBA audio, the AllRAD loudspeaker decoding scheme is used with EFAP.Rendering for binaural headphone reproduction.Time Domain binaural renderer. The time domain (TD) renderer operates on signals in time domain. In the IVAS decoder it is used for binaural rendering of discrete ISM, where each audio signal is encoded and decoded with a dedicated SCE module. This covers all ISM bit rates, except 3-4 objects for bit rates 24.4 kbps and 32 kbps. Further it is used in the decoder for binaural rendering of 5.1 and 7.1 signals when head-tracking is enabled. In the external renderer it is used for all ISM configurations and all multichannel loudspeaker configurations, both with and without head-tracking enabled. An overview of the TD binaural renderer is found in Fig. 2 below. An HRIR model accepts the object position metadata along with the head-tracking data and generates an HRIR filter pair. The ITD may be modeled as a part of the HRIR, or it may be modeled as a separate parameter. In case an ITD parameter is output, the ITD synthesis is performed in the ITD synthesis stage. The time aligned signals are then convolved with the HRIR filter pair to form a binauralized signal.Parametric binauralizer and stereo renderer.The parametric binauralizer and stereo renderer operates on the following IVAS formats and operations: MASA, OMASA, multi-channel (in McMASA mode) , SBA, OSBA, and ISM, i.e., the input to the encoder has been audio signals (and potentially spatial metadata) in one of these formats, and it is now being rendered to binaural or stereo output. The IVAS format being processed (i.e., whether operating on MASA, OMASA, multi-channel (in McMASA mode) , SBA, OSBA, or ISM format) is obtained.Room acoustics rendering.IVAS rendering supports synthesis of room acoustics for realistic immersive effect. The room acoustics can be synthesized using room impulse response convolution or late reverb, optionally combined with early reflections. The room impulse response data (BRIRs) , late reverb and early reflection synthesis are driven by the set of parameters that are discussed in detail in section Rendering control.Rendering control.Head tracking.In addition to supporting head rotation processing, the IVAS renderer supports a number of modes for listener orientation tracking. The listener orientation tracking refers to a set of methods used to provide or estimate listener’s frontal orientation (being listener’s head neutral position or torso position) . Such a listener’s frontal orientation is further referred to as reference orientation. The following orientation tracking modes are supported:External reference orientation;External reference vector orientation;External reference leveled vector orientation;Adaptive long-term average reference orientation.Room acoustics parameters.The late reverb is driven by the set of parameters including of:RT60 –indicating the time that it takes for the reflections to drop 60 dB in energy level;DSR –diffuse to source signal energy ratio;Pre-delay –delay at which the computation of DSR values was made. Can be interpreted as the threshold between early reflections and late reverberation phase;Spatialized, rotation-responsive, first-order early reflections can be added when using multichannel input (any configuration accepted) . The early reflection rendering is determined by several parameters that drive a shoebox model using the image-source method. The set of parameters consists of:3D rectangular virtual room dimensions;Broadband energy absorption coefficient per wall;Listener origin coordinates within room (optional) ;Low-complexity mode (optional) –favours efficient early reflection rendering over spatial accuracy.Room acoustics parameters are provided to the renderer as metadata. Two metadata formats are supported in the IVAS decoder / renderer implementation: binary renderer config metadata format, and text renderer config metadata format . Regardless of the metadata format, the general metadata processing is shared. Both metadata formats support multiple acoustic environment datasets, allowing for selecting between such acoustic environments.Splitted Rendering.Spatial audio rendering is computationally intensive, especially with realtime room acoustics rendering at high fedality and low latency. At the same time, rendered spatial audio is often consumed on light-weight mobile end-devices, such as mobile phones, XR devices, or other wearable devices. It's a difficult task to hit the balance between rendering quality, end-to-end latency, and power consumption, especially when all the computation is done totally on resource-constrained devices.To solve this issue, some might distribute spatial audio processing load on two or more hardware devices. This technology is called split rendering of spatial audio. A common strategy of split rendering of spatial audio is to render computationally-intense, but less latency-demanding tasks on cloud or edge devices; and at the same time, to render computationally-light, but latency-demanding tasks on light-weight resource-constrained end devices. Recent spatial audio standards like 3GPP ISAR have employed this idea. A representation of intermediate sound field is transmitted between stages of such split rendering pipeline, which can be called Intermediate Representation, i.e. IR.Different IRs have different levels of vulnerability to data transmission latency between stages of split rendering. For example:Binaural IR is prone to data transmission latency because such latency adds up to motion-to-sound latency when listeners translating or rotating their heads;3DOF IR has moderate vulnerability to transmission latency because such latency adds up to motion-to-sound latency only when listeners translating their heads;6DOF IR is highly robust to transmission latency because data transmission latency doesn't affect motion-to-sound latency. it only results in delays in the sound content itself.Different IRs also have different data sizes. Binaural IRs have the least amount of data size, followed by 3dof and 6dof IR, respectively. Higher data size usually requires higher bit rate or transmission bandwidth, which eventually levels up transmission cost for IR between stages of split rendering.A critical part of designing a spatial audio split rendering pipeline is to hit the balance between latency robustness and transmission cost. A suitable IR is the key to this problem. Our Hierarchical Sound Field Representation can be used as IR in various spatial audio split rendering scenarios. To form a complete technical solution, a split rendering pipeline that uses HSR as IR needs to be defined as well.Technical problems to be solved.A suitable IR is the key to a successful spatial audio split rendering system. However, there hasn't been any format specified for IR, resulting in two issues:It is impossible to ensure that IR can be received and processed by downstream rendering stages;Rendering stages cannot inform the outside world of the format of its IR output;To solve this problem, an IR needs the following features:It supports connectionless networking, which could ensure low-transmission latency and low-bandwidth consumption;It has versatile data structure, so that adopters of this standard can easily hit proper balance between transmission latency immunity and low bit rate based on their product definition;It has low data overhead when used to represent conventional sound field representation like multichannel or Ambisonics signals;Good forward and backward compatibility, which help modularize the hardware and software in multiple stages of spatial audio split rendering.Our Hierarchical Sound Field Representation can be used as IR in various spatial audio split rendering scenarios.To form a complete technical solution, a split rendering pipeline that uses HSR as IR needs to be defined as well.Hierachial Soundfield Representation.A novel sound-field representation format is proposed: Hierarchical Sound-field Representation (HSR) . An HSR has the following features.An HSR clearly defines its supported degree of freedom (dof) , which specifies the dof of listener movements can be used to query sound field information without using data from other HSRs or other sources.An HSR of given dof can be hierarchically constructed with one or more lower-dof HSRs. Each HSR can have multiple children HSRs, but can only have one parent HSR.Each HSR carries its metadata along with its data stream. Metadata can be transmitted once, or being transmitted every data packet.Metadata.Metadata of an HSR includes each HSR's property that can be used by downstream spatial audio split rendering stages to render the end result or IR. An HSR node contains at least one of the following metadata items:Sample Rate: Sample rate of audio data, for example 16kHz, 32kHz, 44.1kHz, 48kHz, etc;Type: specify channel number and channel order of this HSR, for example binaural, 5.1, 7.1.4, FOA, 7-th order ambisonics, combined HSR, customized, etc;Data Channel Offset (optional) : Channel offset for this HSR in the multi-channel data stream. Default to 0;Transform (optional) : sound field transform, which includes following fields:IsRelative (optional) : is this transform relative to parent HSR's transform. isRelative = false indicates that this HSR transform is relative to world coordinate, whatever its parent HSR transform is;Position (optional) : Positional information of this HSR;Rotation (optional) : rotational information of this HSR;If this field is void, its default value is:IsRelative: True;Position: [0, 0, 0] ;Rotation: azimuth = 0 degrees, elevation = 0 degrees, roll = 0 degrees;Expansion metadata (optional) : allow extra metadata to be transmitted, usually contains implementation-specific fields. For example, implementations can store extra sound field properties for HSRs with Type equal to "customized" ;Children (optional) : Define children of this HSR node as shown in Fig. 3.Metadata can be transmitted each N HSR frame, N >= 1; or it can be transmitted only when it's updated; Metadata can be transmitted in whole, or only transmitting changed metadata fields.Multi-channel Data Stream.The data stream contains raw data of the HSR, which is a channel-wise concatenation of all HSRs in the transmitted HSR Hierarchy. Multi-channel audio data can be arranged in interleaved or planar manner.FIG. 13a shows a flow diagram of the audio processing method according to some embodiments of the present disclosure.As shown in FIG. 13a, in step 110, IR information of audio data transmitted by a first renderer is received. For example, An HSR is uses as the IR information which is clearly defines its supported DoF, specifying the DoF of listener movements.In step 120, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information is determined.In some embodiments, the IR information includes supported DoF (Binaural, 3DOF or 6DOF) of a listener for the audio data, and the sound field structure is identified based on the supported DoF. For example, the sound field structure is identified based on metadata of the IR information, and the metadata including at least one of: a sample rate of the audio data, a type of the IR information, a data channel offset of the IR information in a multi-channel data stream, or sound field transform information of the IR information.For example, the type of the IR information is configured to indicate a channel number and / or a channel order of the IR information.In some embodiments, the second renderer supports a data structure of the IR information, and the corresponding rendering algorithm is obtained from an algorithm database.In some embodiments, the second renderer does not support a data structure of the IR information, and a Downmix operator and / or Mute operator is determined as the corresponding rendering algorithm. For example, an exception message is sentto a local logger and / or to the first renderer, and the exception message indicating that the second renderer does not support the data structure of the IR information.In step 130, based on the corresponding rendering algorithm, the IR information is processed by using a second renderer.In some embodiments, the IR information is processed based on listener pose information for the audio data. For example, the listener pose information includes listener position information and / or orientation information.In the above embodiments, the downstream renderer obtains the sound field structure by analyzing the IR information transmitted by the upstream renderer, and selects a rendering algorithm appropriate for the sound field structure to process the audio data, enabling the rendering algorithm to adapt to the current scenario, thereby improving the effectiveness of the audio processing.The audio processing method of the present disclosure is illustrated exemplarily in the following embodiments.Split Rendering Pipeline using HSR.A brief split rendering pipeline is defined in Fig. 4, which illustrates how to build a versatile spatial audio split rendering system using HSR as IR. For one stage in spatial audio split rendering pipeline, our rendering pipeline can be defined as the following steps.Rendering algorithm selection: The metadata included in input HSR stream and output HSR format helps the downstream renderer to identify the structure of sound field.The downstream renderer loads the appropriate rendering algorithm from algorithm database.If the downstream renderer implementation doesn't support this HSR structure, a default Downmix  / Mute operator can be loaded as a last resort, and an optional exception message can be sent to local logger as well as to upstream renderer.Renderer: Processes the input HSR from upstream using selected rendering algorithm and current listener pose.Listener Pose (Optional) : Current listener position and orientation is sent to local as well as upstream renderers.Exceptions and Logs (Optional) : Each renderer can output logs or exceptions to upstream renderers, in which their compatibility to their input HSR is reported. The upstream renderers can then adjust their output HSR format accordingly.In some embodiments, Rendering Pipeline with 0dof HSR IR.Single-perspective Binaural.Some embodiments of a 0dof HSR that contains single-perspective binaural audio can be as follows:Table 2Here some embodiments of a spatial audio split rendering pipeline using a single-perspective binaural HSR as shown in Fig. 5B, in which:Renderer 0 renders original data to single-perspective binaural signal;Renderer 1 passes through the binaural signal to output.In some embodiments, Rendering Pipeline with 3dof HSR IR.Single-position Multichannel.Some embodiments of a 3dof HSR that contains a multichannel signal locked to listeners position:Table 3Here are some embodiments of a spatial audio split rendering pipeline using this single-position Multichannel HSR as shown in Fig. 6B, in which:Renderer 0 renders original data to single-perspective multichannel signal;Renderer 1 binauralizes multichannel HSR to binaural output signal.Single-position Ambisonics.Some embodiments of a 3dof HSR that contains a 3-th order Ambisonics signal rotated to a given world orientation:Table 4Here are some embodiments of a spatial audio split rendering pipeline using this single-position Ambisonic HSR as shown in Fig. 7B, in which:Renderer 0 renders original data to single-perspective Ambisonics signal;Renderer 1 binauralizes Ambisonics HSR to binaural output signal.Multi-perspective Binaural.Some embodiments of a 3dof combined HSR can be as follows. This HSR is formed by 3 0dof binaural HSR oriented at 3 different angles, and the entire combined HSR is rotated pi  / 3 on azimuth axis, relative to world coordinate:Table 5Here are some embodiments of a spatial audio split rendering pipeline using this multi-perspective binaural HSR as shown in Fig. 8B, in which:Renderer 0 renders original data to multi-perspective binaural signal;Renderer 1 selects the binaural signal with closest orientation to real listener orientation.In some embodiments, Rendering Pipeline with 6dof HSR IR.Multi-positional Ambisonics.Some embodiments of a 6dof combined HSR is as follows. This HSR is formed by 4 single-perspective Ambisonic HSRs at different positions:Table 6Here are embodiments of a spatial audio split rendering pipeline using this multi-position Ambisonics HSR as shown in Fig. 9B, in which:Renderer 0 renders original data to multi-position Ambisonics HSR, which contains 4 Ambisonics HSR;Renderer 1 renders multi-position Ambisonics HSR to one Ambisonics signal at listener's current position;Renderer 2 renders Ambisonic HSR to output binaural signal.Mixed soundfield.Some embodiments of a 6dof combined HSR is as follows. This is a rather complex HSR formed by 3 3-dof HSRs, 2 of which are combined HSRs formed by lower-dof HSRs:Table 7Here are some embodiments of a spatial audio split rendering pipeline using this mixed HSR as shown in Fig. 10B, in which:Renderer 0 renders original data to mixed HSR, which contains the mixed HSR Renderer 1 renders this 6dof mixed HSR to a 3dof mixed HSRRenderer 2 renders the 3dof mixed HSR to output binaural signal.According to some embodiments of the present disclosure, there is provided an audio processing method, including: loading an appropriate rendering algorithm; and processing an input HSR using the appropriate rendering algorithm.In some embodiments, the loading the appropriate rendering algorithm includes: identifying a structure of a sound field to load the appropriate rendering algorithm.In some embodiments, the s processing an input HSR using the appropriate rendering algorithm including: processing an input HSR using a current listener pose.According to some other embodiments of the present disclosure, there is provided an audio processing apparatus, including: loading module for loading an appropriate rendering algorithm; and processing module for processing an input HSR using the appropriate rendering algorithm.In some embodiments, the loading module identifies a structure of a sound field to load the appropriate rendering algorithm.In some embodiments, the processing module processes an input HSR using a current listener pose.According to still other embodiments of the present disclosure, there is provided an electronic device, including: a memory; a processor coupled to the memory, the processor configured to, based on instructions stored in the memory, carry out the audio processing method according to any one of the above embodiments.According to still other embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, implements the audio processing method according to any one of the above embodiments.According to still other embodiments of the present disclosure, there is provided a computer program product, including: instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of the above embodiments.FIG. 11 shows a block diagram of the electronic device according to some embodiments of the present disclosure.As shown in FIG. 11, the electronic device 6 includes: a memory 61 and a processor 62 coupled to the memory 61, the processor 62 configured to, based on instructions stored in the memory 61, carry out the audio processing method according to any one of the embodiments of the present disclosure.Wherein, the memory 61 may include, for example, system memory, a fixed non-transitory storage medium, or the like. The system memory stores, for example, an operating system, applications, a boot loader, a database, and other programs.FIG. 12 shows a block diagram of the electronic device according to some embodiments of the present disclosure.As shown in FIG. 12, the electronic device 7 for generating the fitness regimen information of this embodiment includes: a memory 710 and a processor 720 coupled to the memory 710, the processor 720 configured to, based on instructions stored in the memory 710, carry out the audio processing method according to any one of the embodiments of the present disclosure.The memory 710 may include, for example, system memory, a fixed non-transitory storage medium, or the like. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.The electronic device 7 for generating the fitness regimen information may further include an input-output interface 730, a network interface 740, a storage interface 750, and the like. These interfaces 730, 740, 750, the memory 710 and the processor 720 may be connected through a bus 760, for example. Wherein, the input-output interface 730 provides a connection interface for input-output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, a loudspeaker, etc. The network interface 740 provides a connection interface for various networked devices. The storage interface 750 provides a connection interface for external storage devices such as an SD card and a USB flash disk.FIG. 13b shows a block diagram of the audio processing apparatus according to some embodiments of the present disclosure.As shown in FIG. 13b, an audio processing apparatus 13, including: a receiving unit 131 configured to receive IR information of audio data transmitted by a first renderer; a determining unit 132 configured to determine, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; and a processing unit 133 configured to process based on the corresponding rendering algorithm, the IR information by using a second renderer.In some embodiments, the IR information includes supported DoF (degrees of freedom) of a listener for the audio data, and the determining unit 132 identifies the sound field structure based on the supported DoF.In some embodiments, the determining unit 132 identifies the sound field structure based on metadata of the IR information, the metadata including at least one of: a sample rate of the audio data, a type of the IR information, a data channel offset of the IR information in a multi-channel data stream, or sound field transform information of the IR information.In some embodiments, the type of the IR information is configured to indicate a channel number and / or a channel order of the IR information.In some embodiments, the second renderer supports a data structure of the IR information, and the determining unit 132 obtains the corresponding rendering algorithm from an algorithm database.In some embodiments, the second renderer does not support a data structure of the IR information, and the determining unit 132 determines a Downmix operator and / or Mute operator as the corresponding rendering algorithm.In some embodiments, the processing unit 133 sends an exception message to a local logger and / or to the first renderer, the exception message indicating that the second renderer does not support the data structure of the IR information.In some embodiments, the processing unit 133 processes the IR information based on listener pose information for the audio data.In some embodiments, the listener pose information includes listener position information and / or orientation information.In the above embodiments, the downstream renderer obtains the sound field structure by analyzing the IR information transmitted by the upstream renderer, and selects a rendering algorithm appropriate for the sound field structure to process the audio data, enabling the rendering algorithm to adapt to the current scenario, thereby improving the effectiveness of the audio processing.Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, embodiments of the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment containing both hardware and software elements. Moreover, the present disclosure may take the form of a computer program product embodied on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical memory, etc. ) having computer-usable program code embodied therein.Heretofore, the method, apparatus device, computer-readable storage medium and computer program product according to the present disclosure have been described in detail. In order to avoid obscuring the concepts of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can understand how to implement the technical solutions disclosed herein.The method and system of the present disclosure may be implemented in many ways. For example, the method and system of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above sequence of steps of the method is merely for the purpose of illustration, and the steps of the method of the present disclosure are not limited to the above-described specific order unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the method according to the present disclosure. Thus, the present disclosure also covers a recording medium storing programs for executing the method according to the present disclosure.Although some specific embodiments of the present disclosure have been described in detail by way of example, those skilled in the art should understand that the above examples are only for the purpose of illustration and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure.The scope of the disclosure is defined by the following claims.

Claims

1.An audio processing method, comprising:receiving Intermediate Representation (IR) information of audio data transmitted by a first renderer;determining, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; andprocessing, based on the corresponding rendering algorithm, the IR information by using a second renderer.2.The audio processing method according to claim 1, wherein the IR information comprises supported degrees of freedom (DoF) of a listener for the audio data, andthe determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information comprises:identifying the sound field structure based on the supported DoF.3.The audio processing method according to claim 2, wherein the identifying the sound field structure based on the supported DoF comprises:identifying the sound field structure based on metadata of the IR information, the metadata comprising at least one of: a sample rate of the audio data, a type of the IR information, a data channel offset of the IR information in a multi-channel data stream, or sound field transform information of the IR information.4.The audio processing method according to claim 3, wherein the type of the IR information is configured to indicate a channel number and / or a channel order of the IR information.5.The audio processing method according to any one of claims 1 to 4, wherein the second renderer supports a data structure of the IR information, andthe determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information comprises:obtaining the corresponding rendering algorithm from an algorithm database.6.The audio processing method according to any one of claims 1 to 4, wherein the second renderer does not support a data structure of the IR information, andthe determining, based on the sound field structure of the IR information, the corresponding rendering algorithm for the IR information comprises:determining a Downmix operator and / or Mute operator as the corresponding rendering algorithm.7.The audio processing method according to claim 6, further comprising:sending an exception message to a local logger and / or to the first renderer, the exception message indicating that the second renderer does not support the data structure of the IR information.8.The audio processing method according to any one of claims 1 to 4, wherein the processing, based on the corresponding rendering algorithm, the IR information by using the second renderer comprises:processing the IR information based on listener pose information for the audio data.9.The audio processing method according to claim 8, wherein the listener pose information comprises listener position information and / or orientation information.10.An audio processing apparatus, comprising:a receiving unit configured to receive Intermediate Representation (IR) information of audio data transmitted by a first renderer;a determining unit configured to determine, based on a sound field structure of the IR information, a corresponding rendering algorithm for the IR information; anda processing unit configured to process based on the corresponding rendering algorithm, the IR information by using a second renderer.11.An electronic device, comprising:a processor;a memory for storing processor executable instructions;wherein the processor is used to read the executable instructions from the memory and execute the instructions to implement a audio processing method of any one of claims 1 to 9.12.A computer readable storage medium storing thereon a computer program that, when executed by a processor, causes the processor to implement an audio processing method of any one of claims 1 to 9.13.A computer program product, comprising:instructions that, when executed by a processor, cause the processor to implement an audio processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Rendering item processing method and device of audio renderer, equipment and storage medium

    CN115038029A

  • Audio signal processing device and audio signal processing system

    US20200053461A1

  • Deferred audio rendering

    US20200178016A1

  • Three-dimensional audio playing method and playing apparatus

    US20200374646A1

  • Method and apparatus for providing audio content in immersive reality

    US20210211826A1