Audio renderer for modeling auditory position lag and operation method thereof

The audio renderer addresses the challenge of modeling auditory positional delay by calculating sound propagation delays and adjusting object positions, resulting in more realistic audio experiences in virtual reality environments.

WO2026155406A1PCT designated stage Publication Date: 2026-07-23ELECTRONICS & TELECOMM RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ELECTRONICS & TELECOMM RES INST
Filing Date
2025-12-18
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing audio rendering technologies struggle to accurately model auditory positional delay for object-based audio services, particularly in complex virtual reality environments, leading to unrealistic sound experiences due to the complexity of incorporating acoustic spatial information.

Method used

An audio renderer that determines the distance between an audio object and a listener, calculates the time delay for sound propagation, adjusts the audio object's position based on this delay, and models auditory positional delay to render sound sources more realistically, considering factors like distance gain, medium absorption, and Doppler effect.

Benefits of technology

Enhances the realism of audio experiences in virtual reality by accurately modeling auditory positional delay, ensuring synchronized audio-visual feedback for moving objects, thereby improving the overall immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025022145_23072026_PF_FP_ABST
    Figure KR2025022145_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an audio renderer for modeling an auditory position lag and an operation method thereof. The disclosed operation method of the audio renderer comprises the operations of: determining the distance between a position of an audio object and a listener; on the basis of the distance, determining a time delay required for a sound of the audio object to cover the distance; determining, on a trajectory of the audio object, a previous position of the audio object corresponding to the time delay; modifying a position of the audio object to the determined previous position; and modeling an auditory position lag of the audio object on the basis of the modified position of the audio object.
Need to check novelty before this filing date? Find Prior Art

Description

Audio renderer modeling auditory positional delay and method of operation thereof

[0001] The following disclosure relates to an audio renderer for modeling auditory positional delay and a method of operation thereof.

[0002] Audio services have evolved from mono and stereo services through 5.1 and 7.1 channels to multichannel services such as 9.1, 11.1, 10.2, 13.1, 15.1, and 22.2 channels. Unlike existing channel-based audio services, object-based audio service technology, which treats a single sound source as an object, is being developed. Object-based audio services can store, transmit, and play object audio signals and object audio-related information (e.g., location of object audio, size of object audio).

[0003] Information required for rendering audio signals includes the relative angle and distance between the audio object and the listener; however, there are cases where audio signals are rendered by additionally utilizing acoustic spatial information. This is because acoustic spatial information enables the better realization of acoustic transmission characteristics depending on the space. Using acoustic spatial information to finely realize acoustic transmission characteristics and render object-based audio signals can require very complex computations. To simplify the realization of acoustic transmission characteristics depending on the space, a method has been proposed to render object-based audio signals by dividing them into direct sound, early reflections, and late reverberations.

[0004] The aforementioned background technology was possessed or acquired during the process of deriving the present disclosure and cannot be considered as prior art disclosed to the general public prior to the filing of the present disclosure.

[0005] Various embodiments can model the auditory positional delay of an audio object based on the distance between the location of the audio object and the listener.

[0006] Various embodiments determine the auditory location of an audio object based on the distance between the visual location of the audio object and the listener, and can render the sound source of the audio object based on the determined auditory location.

[0007] Various embodiments can determine whether to apply an auditory position delay of an audio object to an audio object depending on the situation.

[0008] Other objects and advantages of the present invention may be understood from the following description and will become more clearly apparent from the embodiments of the present invention. Furthermore, it will be readily apparent that the objects and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims.

[0009] A method of operation of an audio renderer according to one embodiment includes: determining the distance between the position of an audio object and a listener; determining, based on the distance, the time delay required for the sound of the audio object to cover the distance; determining the previous position of the audio object corresponding to the time delay on the trajectory of the audio object; modifying the position of the audio object to the determined previous position; and modeling the auditory position lag of the audio object based on the modified position of the audio object.

[0010] The operation of determining the previous position of the audio object may determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame based on the time delay, and determine the position of the audio object stored at the frame index as the previous position.

[0011] The sound source of the above audio object may be an object-based audio signal.

[0012] The operation of modeling the above auditory position delay can determine whether to apply the auditory position delay to the audio object according to a predetermined flag for the audio object.

[0013] The operation of modeling the above auditory position delay can render the sound source of the audio object at the location of the modified audio object according to the above auditory position delay.

[0014] The operation of modeling the above auditory position delay determines the distance gain and medium absorption gain according to the position of the modified audio object and the Doppler effect according to the change in distance, and can render the sound source of the audio object based on the distance gain, the medium absorption gain and the Doppler effect.

[0015] The operation of modeling the above auditory position delay can skip modeling the auditory position delay of the audio object that is separated from the listener by a predetermined distance or more.

[0016] A method of operation of an audio renderer according to one embodiment may include an operation of determining the distance between the visual position of an audio object and a listener, an operation of determining, based on the distance, the time delay required for the sound of the audio object to cover the distance, an operation of determining the auditory position of the audio object corresponding to the time delay on the trajectory of the audio object, and an operation of rendering the sound source of the audio object based on the auditory position.

[0017] The operation of determining the auditory location of the above audio object can determine the frame index of the location corresponding to the time delay among the locations of the above audio object stored per frame based on the above time delay, and determine the location of the above audio object stored in the frame index as the auditory location.

[0018] The sound source of the above audio object may be an object-based audio signal.

[0019] The operation of rendering the sound source of the above audio object can determine whether to render the sound source of the above audio object according to a flag predetermined for the above audio object.

[0020] The operation of rendering the sound source of the above audio object determines the distance gain and medium absorption gain according to the auditory position and the Doppler effect according to the change in distance, and can render the sound source of the above audio object based on the distance gain, the medium absorption gain and the Doppler effect.

[0021] The operation of rendering the sound source of the above audio object may skip rendering the sound source of the above audio object when it is separated from the listener by a predetermined distance or more.

[0022] An audio renderer according to one embodiment includes a processor and a memory for storing instructions, and when the instructions are executed by the processor, the audio renderer determines the distance between the location of an audio object and a listener, determines the time delay required for the sound of the audio object to cover the distance based on the distance, determines the previous location of the audio object corresponding to the time delay on the trajectory of the audio object, modifies the location of the audio object to the determined previous location, and models the auditory positional delay of the audio object based on the modified location of the audio object, and the sound source of the audio object may be an object-based audio signal.

[0023] When the above instructions are executed by the processor, the audio renderer may determine, based on the time delay, the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame, and determine the position of the audio object stored at the frame index as the previous position.

[0024] When the above instructions are executed by the processor, the audio renderer may determine whether to apply the auditory position delay to the audio object according to a predetermined flag for the audio object.

[0025] When the above instructions are executed by the processor, the audio renderer may render the sound source of the audio object at the location of the modified audio object according to the auditory position delay.

[0026] When the above instructions are executed by the processor, the audio renderer may determine the distance gain and medium absorption gain according to the position of the modified audio object and the Doppler effect according to the change in distance, and render the sound source of the audio object based on the distance gain, the medium absorption gain and the Doppler effect.

[0027] When the above instructions are executed by the processor, the audio renderer may be able to skip modeling the auditory positional delay of the audio object that is spaced more than a predetermined distance from the listener.

[0028] Various embodiments can render the sound source of an audio object more realistically and enhance the sense of realism experienced by the listener by modeling auditory positional delay based on the distance between the audio object and the listener and the speed of the audio object.

[0029] Various embodiments can more realistically model audio objects moving at high speed in virtual reality (VR) environments.

[0030] FIG. 1 is a block diagram showing an overview of the components of an audio renderer according to one embodiment.

[0031] Figure 2 is a diagram illustrating the encoder structure of an audio renderer.

[0032] Figure 3 is a diagram illustrating the renderer stages of the renderer pipeline of an audio renderer.

[0033] FIG. 4 is a diagram illustrating an audio renderer according to one embodiment.

[0034] FIGS. 5 and 6 are flowcharts illustrating the operation method of an audio renderer according to one embodiment.

[0035] FIGS. 7 and FIGS. 8 are drawings for illustrating an operation for modeling an auditory position delay according to one embodiment.

[0036] FIG. 9 is a diagram illustrating code for rendering a render item according to one embodiment.

[0037] FIGS. 10 and FIGS. 11 are drawings for illustrating code that models auditory position delay according to one embodiment.

[0038] FIG. 12 is a diagram illustrating code for rendering a sound source according to one embodiment.

[0039] FIGS. 13 to 15 are drawings for explaining syntax and flags according to one embodiment.

[0040] FIG. 16 is a diagram illustrating metadata for a render item according to one embodiment.

[0041] FIG. 17 is a diagram illustrating a data structure for a render item according to one embodiment.

[0042] FIG. 18 is a diagram for explaining parameters according to one embodiment.

[0043] FIG. 19 is a block diagram showing an audio renderer according to one embodiment.

[0044] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, actual implementations are not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or substitutions included in the technical concept described by the embodiments.

[0045] In this document, each of the following phrases may include any one of the items listed together in the corresponding phrase, or any combination of A, B, and C, or all possible combinations thereof. Terms such as "A or B," "at least one of A and B," "at least one of A, B, and C," "at least one of A, B, or C," and "a combination of one or more of A, B, and C" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may also be named the first component.

[0046] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or coupled with that other component, or that there may be other components in between.

[0047] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0048] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0049] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are given the same reference numeral regardless of the drawing number, and redundant descriptions thereof will be omitted.

[0050]

[0051] FIG. 1 is a block diagram showing an overview of the components of an audio renderer according to one embodiment.

[0052] According to one embodiment, an audio renderer (10) may represent a device that renders a sound source in virtual reality or augmented reality (AR) and provides it to a listener. For example, the audio renderer (10) may allow a listener to experience virtual sound in a virtual reality (VR) or augmented reality (AR) simulation through immersive audio playback utilizing the listener's 6 degrees of freedom (6DoF) movement in an audio scene. Audio effects and phenomena known in real sound, such as localization, distance attenuation, reflections, reverberation, occlusion, diffraction, and Doppler effect, may be modeled by the audio renderer (10), which is controlled by additional input of metadata transmitted as a bitstream and interactive listener position data. Here, 6DoF may represent spatial navigation (x, y, z) and user head direction (yaw, pitch, roll). In this specification, for convenience of explanation, the user of the audio renderer (10) may be a listener.

[0053] While VR presentations provide the user with the feeling that they are actually present in a virtual world, AR can enrich the real world by naturally perceiving virtual elements as part of the real world. The user can interact with virtual scenes or virtual elements, and the audio renderer (10) can respond to the interaction to generate sounds that match the user experience in the real world and feel realistic. The audio renderer (10) can render a real-time interactive audio presentation while allowing the user 6DoF movement.

[0054] The audio renderer (10) can support real-time audibility of complex 6DoF audio scenes in which a user can directly interact with entities within the scene. To this end, the software architecture of the audio renderer (10) can be divided into multiple workflows and components. In one embodiment, the audio renderer (10) can support rendering of VR and AR scenes. For VR and AR scenes, rendering metadata and audio scene information can be obtained from the bitstream.

[0055] In one embodiment, the audio renderer (10) can perform a control workflow and a rendering workflow. For example, the audio renderer (10) may include a control unit and a rendering unit. In the control workflow, the control unit may include a clock (101), a scene controller (103), and a stream manager (107). In the rendering workflow, the rendering unit may include a renderer pipeline (110), a spatializer, and a limiter. The audio renderer (10) can render audio objects through the control workflow and the rendering workflow. Additionally, the audio renderer (10) can interface with external systems and components through the control unit.

[0056] The control workflow is the entry point of the audio renderer (10) and can be responsible for interfacing with external systems and components. The main functions of the control workflow are embedded in the scene controller component and can coordinate the state of all entities in the 6DoF scene and implement the interactive interface of the audio renderer (10). The scene controller (103) supports external updates to modifiable attributes of scene objects and can complete the information of the bitstream by receiving scene space information (LSI; listener space information). Additionally, the scene controller (103) can track time-dependent or position-dependent attributes of scene objects (e.g., interpolated positions or listener proximity conditions).

[0057] The scene state used by the scene controller (103) may reflect the current state of scene objects, including audio elements, transforms / anches, and geometry. Other components of the audio renderer (10) may reflect changes in the scene state. All objects of the entire scene are created before rendering begins, and the metadata of the objects may be updated to a state that reflects the desired scene configuration at the start of playback. In one embodiment, the scene controller (103) may process changes in all scene information (103_3), internal or external. The inputs of the scene controller (103) may be information received from an external interface of the renderer (e.g., LSI and listener location and dynamic update information (103_1)) and information transmitted by the bitstream (105) (e.g., scene update information). The scene controller (103) may include a scene information module. The scene information module can update the current state of all metadata (e.g., acoustic elements, physical objects) related to the 6DoF rendering of the scene. The scene information module can output the current scene information (103_3) to the renderer pipeline (110).

[0058] The stream manager (107) may provide an integrated interface that allows components of the audio renderer (10) to access audio streams associated with audio elements of the scene state, as well as basic audio playback variables such as audio sample frequency and audio frame length. For example, the stream manager (107) may provide an interface for inputting an acoustic signal (e.g., audio input (100)) for acoustic elements of the scene information module. The audio stream may be input to the audio renderer (10) as pulse code modulation (PCM) floating-point samples. The source of the audio stream may be, for example, a decoded audio stream or locally captured audio. The audio input (100) may be a pre-encoded or decoded sound source signal, or a local sound source or a remote sound source. The stream manager (107) may output the acoustic signal to the renderer pipeline (110). The renderer pipeline (110) can render an audio signal received from the stream manager (107) using the current scene information (103_3). The renderer pipeline (110) may include renderer stages for processing rendering parameters and signals of an audio signal to be rendered (e.g., a render item (RI)).

[0059] The clock (101) may provide an interface for components of the audio renderer (10) to obtain the current scene time in seconds. The clock input (101_1) may be, for example, a synchronization signal of other subsystems or an internal clock of the audio renderer (10). The clock input to the scene controller (103) may not be related to audio synchronization. The clock (101) may output current time information of the scene to the scene controller (103).

[0060] The rendering workflow can generate PCM floating-point audio output signals. The rendering workflow is separate from the control workflow, and the rendering workflow can access the scene state (for conveying changes to the 6DoF scene) and the stream manager (for providing input audio streams) for communication between the two workflows.

[0061] The renderer pipeline (110) can audibly render the input audio stream provided by the stream manager (107) based on the current scene state. Rendering is composed of sequential pipelines, and each stage of the renderer pipeline (110) implements independent perceptual effects and can utilize the processing of the previous and subsequent stages. Each stage of the renderer pipeline (110) can be instantiated during the initialization process of the audio renderer (10) and can be processed according to a predetermined order.

[0062] The spatializer is located after the renderer pipeline (110) and can audibly convert the output of the stages of the renderer pipeline (110) into a single output audio stream suitable for a desired playback method (e.g., binaural or adaptive loudspeaker rendering).

[0063] A limiter can provide clipping protection to the audibly multi-channel output signal.

[0064] According to one embodiment, when an audio renderer (10) renders an audio object and a sound source of an audio object, it may model an auditory positional delay of an audio object by considering the distance between the position of the audio object and the listener. For example, the audio renderer (10) may model an auditory positional delay by determining the time delay required for the sound of the audio object to cover the distance based on the distance between the position of the audio object and the listener, and by modifying the position of the audio object based on the determined time delay. In one embodiment, the audio renderer (10) may determine the auditory position based on the visual position and time delay of the audio object, and render the sound source of the audio object based on the determined auditory position. The process of the audio renderer (10) modeling the auditory positional delay will be explained in detail below through FIGS. 4 to 19.

[0065]

[0066] Figure 2 is a diagram illustrating the encoder structure of an audio renderer.

[0067] Referring to FIG. 2, the audio renderer may include an encoder (200). The encoder (200) may include an EIF (encoder input format) parser module (210), a scene metadata module (230), and a bitstream generation module (250).

[0068] The EIF parser module (210) can take directional information of the EIF and / or SOFA (spatial oriented format for audio) formats, which are common input formats of the encoder (200), as input. The EIF parser module (210) can analyze the information of the EIF and / or SOFA formats to extract elements that constitute scene information of the content (e.g., geometric structure information of the space, sound source information (e.g., location, shape, and directionality of the sound source), acoustic characteristic information of the material and space, and update information (e.g., motion information)).

[0069] In one embodiment, the metadata of the EIF may include data for an audio renderer to render an auditory position delay. For example, the metadata of the EIF may include at least one of the position of an audio object, a past position, a movement path, and a flag determining whether to apply an auditory position delay.

[0070] The scene metadata module (230) may include a sound source metadata generation module, a multi HOA (higher order ambisonics) metadata generation module, a reverberation parameterization module, a low-complexity early reflections parameterization module, a portal generation module, a sound source / object mobility analysis module, a mesh merge module, a diffraction path analysis module, and an initial reflective surface and array analysis module.

[0071] The bitstream generation module (250) can receive the metadata and directional information of the SOFA file generated by each module of the encoder (200), and generate a bitstream by quantizing and multiplexing.

[0072]

[0073] Figure 3 is a diagram illustrating the renderer stages of the renderer pipeline of an audio renderer.

[0074] Referring to FIG. 3, renderer steps for an audio renderer to render audio objects are illustrated as examples. Each renderer step may be executed in a predetermined order. For example, each renderer step may be executed in the order illustrated in FIG. 3, but the embodiment is not limited thereto. In each renderer step, render items (e.g., audio objects) may be optionally disabled or enabled. Each renderer step may process the rendering of the enabled render items. Below, each renderer step will be described.

[0075] The effect activator stage (301) may be a stage that manages the activation and deactivation of render items related to sound effect playback. Scene objects may be activated and deactivated within the scene state during runtime.

[0076] The room assigning stage (303) may be a stage for applying metadata of acoustic environment information for the room into which a listener has entered to each render item when the listener enters a room containing acoustic environment information. In this specification, for convenience of explanation, the room assigning stage (303) may also be referred to as the acoustic environment assigning stage. The room assigning stage (303) may update the metadata of each render item and the listener in each update stage to reflect the current scene configuration based on the acoustic environments defined in the scene.

[0077] The granular synthesis stage (305) is a method of rendering procedural audio and may be a stage in which sound evolves in a controlled manner using real-time input. Grain synthesis may be based on the original recording divided into small pieces, for example, grains. The granular synthesis stage (305) operates by connecting the grains at the time of rendering, and the grains to be used may be controlled through real-time user input, changes in virtual scene state, or predefined trajectories.

[0078] The reverberation stage (307) may be a stage for rendering diffuse late reverberation of each acoustic environment. For example, the reverberation stage (307) may be a stage for generating reverberation according to acoustic environment information of the current space (e.g., a room containing acoustic environment information). The reverberation stage (307) may be a stage for receiving reverberation parameters from a bitstream (bitstream (105) in FIG. 1), attenuating a feedback delay network (FDN) reverberator, and initializing delay parameters.

[0079] The portal stage (309) may be a stage that manages the activation and deactivation of two source types associated with the portal. Here, the portal may be an abstract concept that models the transmission of sound from one space to another through a geometrically defined open part. First, in the portal stage (309), the audio renderer may set up a reverberation extension source so that reverberation is heard through an acoustic opening (e.g., a portal) outside the acoustic environment, and manage the mixing of signals played from said reverberation extension source. Second, in the portal stage (309), the audio renderer may manage coupling sources that render materials and render sources on the opposite side of the portal, and may simulate vibrations occurring in a door or window, etc., to become an extension source itself.

[0080] The early reflections stage (311) may be a stage for calculating specular reflections on a reflective surface using transmitted geometric data. In the early reflections stage (311), the image source model may be used to verify the visibility of a potential propagation path from the sound source to the listener.

[0081] The airflow simulation stage (313) may be a stage that simulates the sound perceived by the listener as air passes through the listener's ear. The sound heard by the listener may vary depending on the speed of the airflow and the direction of the listener.

[0082] The SESS (spatially extended sound sources) detection step (315) may be an auxiliary step for rendering the SESS.

[0083] The occlusion stage (317) may be a stage that provides occlusion information for a direct path (e.g., line of sight) from the sound source to the listener. If the path is obscured by an acoustically opaque or partially transparent object, geometry / mesh information that appears along the line of sight may be updated in a dedicated data structure.

[0084] The heterogeneous extent rendering stage (319) may be a stage for rendering spatially heterogeneous audio elements. Spatially heterogeneous audio elements may be audio elements having source signals with an extended size and two or more audio channels. Audio elements may include object sources with two or more source channels and HOA sources with an extended size. Rendering may appropriately represent audio elements at listening positions within and around an extent that includes both width and height information using the provided extended size information.

[0085] The diffraction stage (321) may be a stage for generating information necessary to generate a diffracted sound source transmitted to a listener from a sound source blocked by an obstacle. In the diffraction stage (321), preprocessed geometric data of a bitstream including edge, path, and voxel data may be used. For a fixed sound source, a pre-calculated diffraction path may be used to generate information. For a moving sound source, a diffraction path calculated from potential diffraction edges may be used to generate information.

[0086] The directivity stage (323) may be a stage for auditoryizing the directional characteristics of an audio element. The directivity stage (323) may include directional data coding and directional data rendering.

[0087] The distance stage (325) may be a stage for rendering independent perceptual effects related to the transmission of sound in the air, such as propagation delay, distance gain, and medium absorption. The distance stage (325) may calculate the current distance between each render item (e.g., audio object) and the listener, and may interpolate the distance between calls to the update routine based on a constant velocity model. Propagation delay may be applied to the signal associated with the render item to generate physically accurate delay and Doppler effects using a variable delay line that includes subsample interpolation. To mitigate jitter in head-tracking listener position and render item position updates, smoothing may be applied to the distance used for propagation delay rendering when updating the model velocity. The conversion from distance to propagation delay may be calculated at the speed of sound given by local configuration parameters. In one embodiment, the audio renderer can model the auditory positional delay of the audio object based on the distance between the position of the audio object and the listener in the distance step (325).

[0088] In one embodiment, the distance between the listener location and the render item may be calculated as a Euclidean distance when a location update is provided. This distance may correspond to an instantaneous propagation delay reproduced using an interpolated variable delay line. A continuous change in propagation delay can produce an essentially physically accurate Doppler effect. The Doppler pitch shift may be a function of the relative velocity between the sound source and the observer, for example, the derivative of the distance.

[0089] The directional focus stage (327) may be a stage that attenuates disruptive sounds outside the spatial area of ​​interest to improve accessibility. The focus may be radially symmetric with one main lobe area.

[0090] The metadata culling stage (329) may be a stage that saves computations that may occur in subsequent stages by disabling render items that become inaudible due to very low gain or EQ (equalizer) (e.g., strong distance attenuation or occlusion). Additionally, reflection render items that are perceived as part of a parent base render item due to precedence effects may also be disabled and culled.

[0091] The consolidation stage (331) may be a stage for combining render items with similar localization attributes to reduce the total number of render items or computational complexity of the renderer pipeline. Render items whose difference in perceptual localization attributes is below a given threshold may be identified through a psychoacoustic model. A group of render items within the threshold may be selected through a computationally efficient clustering algorithm based on the psychoacoustic model. Render items within each group may be consolidated into a common representative render item. For temporal stability, render item assignment is optimized to avoid unnecessarily frequent reassignment, and crossfades may be applied when render item assignment changes.

[0092] The equalizer stage (333) may be a stage of applying frequency-dependent gain to all relevant audio signals after frequency-dependent attenuation has accumulated for acoustic effects (e.g., shielding, diffraction, reflection, directivity, medium attenuation) from previous stages.

[0093] The LC early reflections stage (low-complexity early reflections stage) (335) may be a stage for applying early reflections to sound. In an indoor acoustic environment, the impulse response may include direct sound, early reflections (ER), and late reverberation. In the LC early reflections stage (335), as the listener and / or sound source moves in the environment, the direct sound and all early reflections may dynamically change their individual directions and distances from the listener. In the LC early reflections stage (335), a single common reflection pattern may be applied to all primary sound sources within the scene.

[0094] The fade stage (337) may be a stage that applies fade-in and fade-out ramps to the audio signal, respectively, before the render items are activated or deactivated. For example, the fade stage (337) may be a stage that reduces discontinuous distortion that may occur when the activation status of a render item changes or when a listener suddenly moves through a fade-in-out process.

[0095] A single point higher order ambisonics stage (339) may be a stage that renders a single HOA source binaurally according to the listener's position and orientation relative to the sound source location. For example, the single HOA stage (339) may be a stage that renders background sound by a single HOA source. The single HOA stage (339) may be a stage that converts a signal in an equivalent spatial domain (ESD) format input from a bitstream into an HOA and converts it into a binaural signal through a MagLS (magnitude least squares) decoder. The single HOA stage (339) may be a stage that converts the input audio into an HOA and spatially combines and transforms the signal through HOA decoding.

[0096] The homogeneous extent rendering stage (341) may be a stage for synthesizing SESS for headphone playback of an object source in which a predetermined flag (e.g., objectSourceHasExtent) is set to "1". Here, the predetermined flag may represent a flag indicating whether the object source is spatially extended.

[0097] The panner stage (343) may be a stage for panning the sound source to a virtual loudspeaker (LS) setting. The panner stage (343) may be a stage for implementing vector-base-amplitude-panning (VBAP) through additional control functions such as a configurable spatial spread.

[0098] A multi-point higher-order ambisonics stage (345) may be a stage that provides a 6DoF listening environment to a listener by rendering audio scenes containing one or more sets of multi-channel signals represented as HOA sources. For example, the multi-point HOA stage (345) may be a stage that renders HOA sound sources in 6DoF with respect to the listener's location using information from a spatial metadata frame.

[0099] Hereinafter, with reference to FIGS. 4 to 19, an audio renderer for modeling auditory positional delay according to one embodiment and a method of operation thereof will be described. According to one embodiment, the audio renderer (1900) of FIG. 19 can perform an audio rendering method.

[0100]

[0101] FIG. 4 is a diagram illustrating an audio renderer according to one embodiment.

[0102] Referring to FIG. 4, the audio renderer (410) can determine the auditory position (440) of the audio object based on the distance between the visual position (430) of the audio object and the listener (420). Alternatively, the audio renderer (410) can modify the position of the audio object based on the distance between the position of the audio object and the listener (420).

[0103] When audio objects and video objects need to be synchronized in a VR scene, more precise audio playback may be required. For example, in the case of a 6DoF VR application, it may be required to process the relative position of the audio object according to the listener's position and orientation to reflect physical phenomena. Since the speed of light is faster than the speed of sound, a perceptible difference may occur between the visual position (430) of an audio object moving fast in the real world (e.g., airplane, jet, rocket, vehicle) and the position of the audio object's sound source. In one embodiment, an audio renderer (410) may determine the auditory position (440) of the audio object based on the visual position (430) of the audio object in VR or AR, and render the sound source of the audio object at the auditory position (440) to reflect the difference and provide more realistic sound source playback to the listener (420). However, in this specification, the audio object may include not only fast-moving objects but also objects moving in VR or AR regardless of speed.

[0104] In FIG. 4, the auditory location (440) of the audio object may be delayed relative to the visual location (430) of the audio object. At the moment the listener (420) hears the sound source of the audio object at the auditory location (440), the audio object can be seen at the visual location (430). Accordingly, the audio renderer (410) may determine the auditory location (440) of the audio object to be any of the past locations of the audio object that have moved. The delay of the auditory location (440) may correspond to the time required for the sound to propagate by the distance (sDist) between the auditory location (440) and the listener (420). According to one embodiment, the audio renderer (410) may determine the auditory location (440) of the audio object using the distance (dist) between the visual location (430) of the audio object and the listener (420).

[0105] The audio renderer (410) can determine the visual position (430) of the audio object from the movement trajectory of the audio object. For example, the audio renderer (410) can determine the visual position (430) of the audio object through a function for obtaining the movement trajectory of the audio object. For instance, the audio renderer (410) can determine the visual position (430) of the audio object at the current time through the movement trajectory of the audio object as shown in Equation 1 below.

[0106]

[0107] Here, vp is the visual location of the audio object (430), trajectory is the movement trajectory of the audio object, getLocation is a function that determines the location at that time in the movement trajectory, and ct may represent the current time.

[0108] The audio renderer (410) can determine the distance between the audio object and the listener (420) based on the visual position (430) of the audio object. The audio renderer (410) can obtain the position of the listener from the head tracking information of the rendering system. For example, the audio renderer (410) can determine the absolute value of the position difference between the visual position (430) of the audio object and the position of the listener (420) as the distance between the visual position (430) of the audio object and the listener (420), as shown in Equation 2 below.

[0109]

[0110] The audio renderer (410) can determine the time delay required for the sound of the audio object to cover the distance based on the distance between the visual location (430) of the audio object and the listener (420). For example, the audio renderer (410) can determine the time delay using the speed of sound movement as shown in Equation 3 below.

[0111]

[0112] Here, sdly may represent a time delay, and SpeedOfSound may represent the speed of sound. For example, the speed of sound may be 343 m / s, but may vary depending on the embodiment.

[0113] The audio renderer (410) can determine the auditory location (440) of an audio object based on a time delay. The audio renderer (410) can determine the auditory location (440) of an audio object based on the visual location (430) of an audio object and a time delay. For example, the audio renderer (410) can determine the auditory location (440) of an audio object by subtracting a time delay from the current time as shown in Equation 4 below.

[0114]

[0115] Here, ap can represent the auditory location (440) of the audio object.

[0116] The audio renderer (410) can implement sound delay by rendering the sound source at an auditory location (440) instead of the visual location (430) of the audio object in the specializer. Additionally, when rendering the sound source at the auditory location (440), the audio renderer (410) can determine at least one of the distance gain, medium absorption, and Doppler effect of the sound source according to the auditory location, and render the sound source of the audio object by reflecting the determined distance gain, medium absorption, and Doppler effect. For example, the audio renderer (410) can determine the distance between the auditory location (440) of the audio object and the listener (420) as shown in Equation 5 below, and determine the distance gain, medium absorption, and Doppler effect of the sound source based on the determined distance between the auditory location (440) and the listener (420).

[0117]

[0118] Here, sDist may represent the distance between the auditory location (440) of the audio object and the listener (420).

[0119] For example, the positional delay according to the velocity of each audio object may be as shown in Table 1 below.

[0120] Speed ​​Distance sDist Time Delay sdly Localization Distance Airplane 250 [m / sec] 5,000 [m] 15 [sec] 3,750 [m] Jet 400 [m / sec] 100 ~ 5,000 [m] 0.3 ~ 15 [sec] 120 ~ 6,000 [m] Rocket 1,000 [m / sec] 5,000 [m] 15 [sec] 15,000 [m]

[0121] According to one embodiment, the audio renderer (410) can determine the auditory location (440) of the audio object again by using the distance sDist between the determined auditory location (440) of the audio object and the listener (420) and the time delay sdly, in order to determine the auditory location (440) of the audio object more accurately.

[0122]

[0123] FIGS. 5 and 6 are flowcharts illustrating the operation method of an audio renderer according to one embodiment.

[0124] Referring to FIG. 5, in the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. Operations (510) to (550) may be performed by at least one component of the audio renderer (e.g., a processor, etc.).

[0125] In operation (510), the audio renderer can determine the distance between the location of the audio object and the listener. The sound source of the audio object can be classified as an object-based audio signal, a channel-based audio signal, or a scene-based audio signal. In one embodiment, the sound source of the audio object may be an object-based audio signal.

[0126] In operation (520), the audio renderer can determine the time delay required for the sound of the audio object to cover the distance based on the distance.

[0127] In operation (530), the audio renderer can determine the previous position of the audio object corresponding to the time delay on the trajectory of the audio object. Based on the time delay, the audio renderer can determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame, and determine the position of the audio object stored at the frame index as the previous position.

[0128] In operation (540), the audio renderer can modify the position of the audio object to a determined previous position.

[0129] In operation (550), the audio renderer can model the auditory positional delay of the audio object based on the position of the modified audio object. The audio renderer can determine whether to apply the auditory positional delay to the audio object based on a flag predetermined for the audio object. The audio renderer can render the sound source of the audio object at the position of the modified audio object based on the auditory positional delay. The audio renderer determines the distance gain and medium absorption gain and the Doppler effect based on the change in distance based on the position of the modified audio object, and can render the sound source of the audio object based on the distance gain, medium absorption gain, and Doppler effect. The audio renderer can skip modeling the auditory positional delay of the audio object if it is separated from the listener by more than a predetermined distance.

[0130]

[0131] Referring to FIG. 6, in the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. Operations (610) to (640) may be performed by at least one component of the audio renderer (e.g., a processor, etc.).

[0132] In operation (610), the audio renderer can determine the distance between the visual location of the audio object and the listener. The sound source of the audio object may be an object-based audio signal.

[0133] In operation (620), the audio renderer can determine the time delay required for the sound of the audio object to cover the distance based on the distance.

[0134] In operation (630), the audio renderer can determine the auditory location of the audio object corresponding to the time delay on the trajectory of the audio object. Based on the time delay, the audio renderer can determine the frame index of the location corresponding to the time delay among the locations of the audio object stored per frame, and determine the location of the audio object stored at the frame index as the auditory location.

[0135] In operation (640), the audio renderer can render the sound source of an audio object based on an auditory position. The audio renderer can determine whether to render the sound source of an audio object based on a predetermined flag for the audio object. The audio renderer determines the distance gain and medium absorption gain based on the auditory position and the Doppler effect based on the change in distance, and can render the sound source of an audio object based on the distance gain, medium absorption gain, and Doppler effect. The audio renderer can skip rendering the sound source of an audio object that is separated from the listener by more than a predetermined distance.

[0136]

[0137] FIGS. 7 and FIGS. 8 are drawings for illustrating an operation for modeling an auditory position delay according to one embodiment.

[0138] Referring to FIG. 7, the audio renderer can model an auditory position delay for a listener (710) by modifying the position (730) of an audio object to a position (740) based on the movement trajectory (720) of an audio object. In this specification, for convenience of explanation, the movement trajectory (720) of an audio object may also be referred to as the trajectory of an audio object.

[0139] According to one embodiment, the audio renderer can determine the previous position (740) of the audio object corresponding to the time delay on the movement trajectory (720) of the audio object. In one embodiment, the audio renderer can determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame based on the time delay, and determine the position of the audio object stored at the corresponding frame index as the previous position (740). The audio renderer can determine the previous position (740) as the auditory position of the audio object. Through this, the audio renderer can render the sound source more realistically even for audio objects moving non-linearly, and enhance the sense of realism experienced by the listener (710).

[0140] For example, the audio renderer can store the positions of the audio object as it moves, frame by frame. The audio renderer can determine a frame index (idx) corresponding to the time delay among the stored positions of the audio object based on a time delay based on the distance (dist) between the listener (710) and the audio object. The audio renderer can modify the position of the audio object to the position (740) of the audio object at the determined frame index and model the auditory position delay of the audio object based on the modified position (740). For instance, the audio renderer can render the sound source of the audio object at the corresponding position (740).

[0141] According to one embodiment, an audio renderer may determine whether to apply auditory positional delay to an audio object based on a predetermined flag for the audio object. Here, the predetermined flag may indicate whether auditory positional delay, which is caused by modeling the propagation delay of sound delivered to a listener, is applied to the corresponding audio object. The predetermined flag may be included in the syntax representing the corresponding audio object. For example, the audio renderer may apply auditory positional delay to the audio object if the value of the predetermined flag "objectSourcePositionLagEnabled" included in the bitstream is a first value (e.g., "TRUE") (e.g., "objectSourcePositionLagEnabled == TRUE"). Additionally, the audio renderer may not apply auditory positional delay to the audio object if the value of the predetermined flag "objectSourcePositionLagEnabled" included in the bitstream is a second value (e.g., "FALSE") (e.g., "objectSourcePositionLagEnabled == FALSE"). For example, an audio renderer may not apply auditory positional delay to audio objects by skipping the modeling of auditory positional delay of audio objects.

[0142] According to one embodiment, the audio renderer may modify the location itemLocation of a render item (e.g., audio object) from location (730) to location (740) when a predetermined flag is a first value. For example, the audio renderer may modify the location (730) of the audio object based on the past location of the audio object included in the time delay and movement trajectory (720). For instance, the audio renderer may modify the location itemLocation of the render item from location (730) to location (740) through a predetermined function (e.g., "itemLocation = PosLag(traj, dist);"). In this specification, for convenience of explanation, the predetermined function for modifying the location (730) of the render item may be referred to as an auditory location delay function. According to one embodiment, the audio renderer may not modify the location (730) of the render item when a predetermined flag is a second value. For example, the audio renderer may not modify the position (730) of the render item by not applying a predetermined function.

[0143] An auditory position delay function (e.g., "PosLag()") can provide a past position on the movement traj of an audio object corresponding to the time delay T required for the sound to travel a distance dist. The audio renderer can determine the past position of the audio object corresponding to the time delay through the auditory position delay function and modify the position (730) of the audio object to the position (740) corresponding to the determined past position.

[0144] Referring to FIG. 8, pseudo-code (800) of an auditory position delay function for modifying the position of a render item is illustrated as an example. The calculation of the render item position "traj.pos(idx)" can be performed as defined in the pseudo-code (800) illustrated in FIG. 8. However, the pseudo-code (800) illustrated in FIG. 8 is for illustrative purposes only and is not limited thereto.

[0145]

[0146] FIG. 9 is a diagram illustrating code for rendering a render item according to one embodiment.

[0147] Referring to FIG. 9, code (900) representing a render item implemented in software for rendering an audio object is illustrated as an example. However, the code (900) illustrated in FIG. 9 is an example for illustrative purposes only, and embodiments are not limited thereto, and render items may be implemented in various ways.

[0148] According to one embodiment, code (900) representing a render item may include a flag (e.g., "useDelayedPosition") controlling whether to apply an auditory position and a variable (e.g., "actualDelayedPosition") representing the auditory position of the render item. An audio renderer may model an auditory position delay of the render item based on the flag controlling whether to apply an auditory position and the variable representing the auditory position of the render item included in the code (900). The flag controlling whether to apply an auditory position may help reduce the complexity of the audio renderer by allowing VR content creators to selectively apply sound delay effects to necessary audio objects. For example, depending on the flag, the distance between the visual position of the audio object and the listener position, or the distance between the auditory position and the listener position, may be selectively used.

[0149]

[0150] FIGS. 10 and FIGS. 11 are drawings for illustrating code that models auditory position delay according to one embodiment.

[0151] Referring to FIG. 10, codes (1010, 1020, 1030, 1040, 1050, 1060) regarding the distance between a render item and a listener are illustrated as examples. However, the codes (1010, 1020, 1030, 1040, 1050, 1060) illustrated in FIG. 10 are exemplary for illustrative purposes only, and embodiments are not limited thereto and can be implemented in various ways.

[0152] In code (1010), the audio renderer can determine the current position of the audio object. Additionally, the audio renderer can determine the current distance between the audio object and the listener.

[0153] In code (1020), the audio renderer can determine the time delay required for the sound of the audio object's source to propagate. Additionally, the audio renderer can determine the position of the delayed audio object based on the time delay. The audio renderer can modify the position of the audio object based on the time delay.

[0154] In code (1030), the audio renderer can determine the distance between the modified position of the audio object and the listener.

[0155] In code (1040), the audio renderer can apply a Doppler effect to the audio object at a modified position of the audio object. The audio renderer can determine a cursor position for the distance of the audio object delayed for the Doppler effect.

[0156] In code (1050), the audio renderer can determine the attenuation gain for the audio object at the modified location of the audio object. The audio renderer can apply the determined attenuation gain to the sound source of the audio object.

[0157] In code (1060), the audio renderer can determine the medium absorption gain for the audio object at a modified location of the audio object. The audio renderer can determine the medium absorption gain for each frequency band. The audio renderer can apply the determined medium absorption gain to the sound source of the audio object.

[0158] Referring to FIG. 11, the code (1100) regarding the distance between a render item and a listener may include code (1110) for calculating values ​​for a delayed sound source.

[0159]

[0160] FIG. 12 is a diagram illustrating code for rendering a sound source according to one embodiment.

[0161] Referring to FIG. 12, code (1210) for rendering an audio object at a modified position of the audio object is illustrated as an example. The audio renderer determines whether to apply an auditory position delay to the audio object according to a predetermined flag for the audio object, and can render the audio object and the sound source at the modified position. According to the predetermined flag, the listener can hear the sound of the sound source from the auditory position (modified position) instead of the visual position of the audio object.

[0162]

[0163] FIGS. 13 to 15 are drawings for illustrating syntax and flags according to one embodiment. The syntaxes shown in FIGS. 13 to 15 may be different from each other. The flags shown in FIGS. 13 to 15 are exemplary for illustrative purposes and the embodiment is not limited thereto, and the syntax may include other flags or some flags may be omitted.

[0164] Referring to FIG. 13, the syntax (1300) for an audio object may include a flag (1310) (e.g., "objectSourcePositionLagEnabled") that determines whether to apply auditory positional delay to the audio object. The syntax (1300) may relate to an object source. The flag (1310) may indicate whether auditory positional delay, which is caused by modeling the sound propagation delay transmitted to the listener, is applied to the corresponding audio object. For example, the flag (1310) may indicate that auditory positional delay is not applied to an audio object moving quickly from a distance. In one embodiment, the audio renderer may skip modeling auditory positional delay for audio objects located more than a predetermined distance from the listener. The flag (1310) may not be applied to audio objects associated with a 3D extended sound source.

[0165] Referring to FIG. 14, the syntax (1400) for an audio object may include a flag (1410) (e.g., "hoaSourcePositionLagEnabled") that determines whether to apply an auditory position delay to the audio object. The syntax (1400) may be for an HOA source.

[0166] Referring to FIG. 15, the syntax (1500) for an audio object may include a flag (1510) (e.g., "channelSourcePositionLagEnabled") that determines whether to apply an auditory position delay to the audio object. The syntax (1500) may be for a channel source.

[0167]

[0168] FIG. 16 is a diagram illustrating metadata for a render item according to one embodiment.

[0169] Referring to FIG. 16, metadata fields (1600) for a render item are illustrated as an example. The metadata fields (1600) may include a field for the current position of the audio object and a field (1610) for the modified position (e.g., "actualDelayPosition"). The field (1610) may indicate the delayed sound position (including direction) relative to the visual object of the render item in global coordinates. Additionally, the metadata fields (1600) may further include a field for compensation for distance to a listener to synchronize multiple render items at different locations in terms of propagation delay and distance attenuation.

[0170]

[0171] FIG. 17 is a diagram illustrating a data structure for a render item according to one embodiment.

[0172] Referring to FIG. 17, data (1700) representing a render item is illustrated as an example. The data (1700) may include a flag (e.g., "PositionLagEnabled") (1710) representing an auditory positional delay. An audio renderer may render the delay of the sound source of the corresponding audio object based on the auditory position rather than the visual position of the sound source in the specializer according to the flag (1710).

[0173]

[0174] FIG. 18 is a diagram for explaining parameters according to one embodiment.

[0175] Referring to FIG. 18, parameters (1800) for an audio object are illustrated as an example. The parameters (1800) may include a parameter (1810) (e.g., "PositionLagEnabled") for whether to skip modeling auditory positional delay.

[0176]

[0177] FIG. 19 is a block diagram showing an audio renderer according to one embodiment.

[0178] Referring to FIG. 19, the audio renderer (1900) may include a processor (1910). The processor (1910) may include at least one processor. Additionally, the audio renderer (1900) may further include a memory (1920).

[0179] The memory (1920) can store instructions (e.g., programs) executable by the processor (1910). For example, the instructions may include instructions for executing the operation of the processor (1910) and / or the operation of each component of the processor (1910).

[0180] The processor (1910) is a device that executes instructions or programs or controls the audio renderer (1900), and may include various processors such as a CPU (Central Processing Unit) or a GPU (Graphic Processing Unit). The processor (1910) can determine the distance between the location of the audio object and the listener. Based on the distance, the processor (1910) can determine the time delay required for the sound of the audio object to cover the distance. The processor (1910) can determine the previous location of the audio object corresponding to the time delay on the trajectory of the audio object. The processor (1910) can modify the location of the audio object to the determined previous location. Based on the modified location of the audio object, the processor (1910) can model the auditory positional delay of the audio object.

[0181] The processor (1910) can determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame based on the time delay, and determine the position of the audio object stored at the frame index as the previous position. The processor (1910) can determine whether to apply an auditory position delay to the audio object according to a predetermined flag for the audio object. The processor (1910) can render the sound source of the audio object at the modified position of the audio object according to the auditory position delay. The processor (1910) determines the distance gain and medium absorption gain according to the position of the modified audio object and the Doppler effect according to the change in distance, and can render the sound source of the audio object based on the distance gain, medium absorption gain, and Doppler effect. The processor (1910) can skip modeling the auditory position delay of the audio object that is separated from the listener by more than a predetermined distance.

[0182] The processor (1910) can determine the distance between the visual location of the audio object and the listener. Based on the distance, the processor (1910) can determine the time delay required for the sound of the audio object to cover the distance. The processor (1910) can determine the auditory location of the audio object corresponding to the time delay on the trajectory of the audio object. Based on the auditory location, the processor (1910) can render the sound source of the audio object.

[0183] The processor (1910) can determine the frame index of the location corresponding to the time delay among the locations of audio objects stored per frame based on the time delay, and determine the location of the audio object stored at the frame index as an auditory location. The processor (1910) can determine whether to render the sound source of the audio object according to a flag predetermined for the audio object, which is an operation of rendering the sound source of the audio object. The processor (1910) determines the distance gain and medium absorption gain according to the auditory location and the Doppler effect according to the change in distance, and can render the sound source of the audio object based on the distance gain, medium absorption gain, and Doppler effect. The processor (1910) can skip rendering the sound source of the audio object if it is separated from the listener by more than a predetermined distance.

[0184] In addition, regarding the audio renderer (1900), the above-described operation can be processed.

[0185]

[0186] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. In addition, other processing configurations, such as parallel processors, are also possible.

[0187] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on computer-readable recording media.

[0188] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program instructions, data files, data structures, etc., either individually or in combination, and the program instructions recorded on the medium may be those specifically designed and configured for the embodiment or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0189] The hardware device described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0190] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based thereon. For example, appropriate results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0191] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

1. Regarding the operation method of the audio renderer, An operation to determine the distance between the location of an audio object and a listener; An operation to determine the time delay required for the sound of the audio object to cover the distance based on the distance; An operation to determine the previous position of the audio object corresponding to the time delay on the trajectory of the audio object; An operation to modify the position of the above audio object to the determined previous position; and An operation to model the auditory position lag of the audio object based on the position of the modified audio object. including How the audio renderer operates.

2. In Paragraph 1, The operation of determining the previous position of the above audio object is Based on the above time delay, determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame, and Determining the position of the audio object stored in the frame index as the previous position How the audio renderer operates.

3. In Paragraph 1, The sound source of the above audio object object-based audio signals, How the audio renderer operates.

4. In Paragraph 1, The operation modeling the above auditory position delay is Determining whether to apply the auditory position delay to the audio object according to a predetermined flag for the audio object, How the audio renderer operates.

5. In Paragraph 1, The operation modeling the above auditory position delay is Rendering the sound source of the audio object at the location of the modified audio object according to the above auditory position delay, How the audio renderer operates.

6. In Paragraph 1, The operation modeling the above auditory position delay is Determining the distance gain and medium absorption gain according to the position of the above modified audio object and the Doppler effect according to the change in distance, Rendering the sound source of the audio object based on the distance gain, the medium absorption gain, and the Doppler effect, How the audio renderer operates.

7. In Paragraph 1, The operation modeling the above auditory position delay is Skip modeling the auditory position delay of the audio object separated from the listener by a predetermined distance or more, How the audio renderer operates.

8. A computer-readable recording medium storing a computer program that executes the method of any one of paragraphs 1 through 7.

9. In the method of operation of the audio renderer, An action that determines the distance between the visual position of an audio object and the listener; An operation to determine the time delay required for the sound of the audio object to cover the distance based on the distance; An operation to determine the auditory position of the audio object corresponding to the time delay on the trajectory of the audio object; Operation of rendering the sound source of the audio object based on the above auditory location including How the audio renderer operates.

10. In Paragraph 9, The operation of determining the auditory location of the above audio object is Based on the above time delay, determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame, and Determining the location of the audio object stored in the above frame index as the above auditory location, How the audio renderer operates.

11. In Paragraph 9, The sound source of the above audio object object-based audio signals, How the audio renderer operates.

12. In Paragraph 9, The operation of rendering the sound source of the above audio object is Determining whether to render the sound source of the audio object according to a predetermined flag for the audio object, How the audio renderer operates.

13. In Paragraph 9, The operation of rendering the sound source of the above audio object is Determining the distance gain and medium absorption gain according to the above auditory position and the Doppler effect according to the change in the above distance, Rendering the sound source of the audio object based on the distance gain, the medium absorption gain, and the Doppler effect, How the audio renderer operates.

14. In Paragraph 9, The operation of rendering the sound source of the above audio object is Skip rendering of the sound source of the audio object separated from the aforementioned listener by a predetermined distance or more, How the audio renderer operates.

15. In audio renderers, processor; and Memory that stores instructions Includes, When the above instructions are executed by the processor, the audio renderer, Determine the location of the audio object and the distance between it and the listener, Based on the above distance, determine the time delay required for the sound of the audio object to cover the above distance, and Determining the previous position of the audio object corresponding to the time delay on the trajectory of the audio object, and Modify the position of the above audio object to the previously determined position, and Based on the location of the modified audio object above, the auditory positional delay of the audio object is modeled, and The sound source of the above audio object is an object-based audio signal, Audio renderer.

16. In Paragraph 15, When the above instructions are executed by the processor, the audio renderer, Based on the above time delay, determine the frame index of the position corresponding to the time delay among the positions of the audio object stored per frame, and Determining the position of the audio object stored in the frame index as the previous position, Audio renderer.

17. In Paragraph 15, When the above instructions are executed by the processor, the audio renderer, Determining whether to apply the auditory position delay to the audio object according to a predetermined flag for the audio object, Audio renderer.

18. In Paragraph 15, When the above instructions are executed by the processor, the audio renderer, Rendering the sound source of the audio object at the location of the modified audio object according to the above auditory position delay, Audio renderer.

19. In Paragraph 15, When the above instructions are executed by the processor, the audio renderer, Determining the distance gain and medium absorption gain according to the position of the above modified audio object and the Doppler effect according to the change in distance, Rendering the sound source of the audio object based on the distance gain, the medium absorption gain, and the Doppler effect, Audio renderer.

20. In Paragraph 15, When the above instructions are executed by the processor, the audio renderer, skipping the modeling of the auditory position delay of the audio object separated from the listener by a predetermined distance or more, Audio renderer.