Apparatus and method for rendering a sound scene using pipeline stages - Patents.com

The pipelined rendering architecture with reconfigurable audio data processors and control layers addresses the challenge of rendering complex sound scenes in virtual or augmented reality by enabling dynamic reconfiguration and synchronization, resulting in high-quality, real-time auditoryization.

JP7675089B2Active Publication Date: 2025-05-12FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022555053
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-13
Filing Date
2021-03-12
Publication Date
2025-05-12
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

Existing audio rendering technologies struggle to efficiently handle complex sound scenes with many sources in virtual or augmented reality applications, particularly when frequent changes occur, leading to audible artifacts and limitations in real-time processing.

Method used

A pipelined rendering architecture with multiple stages, each equipped with a reconfigurable audio data processor and a control layer, allows for dynamic reconfiguration and synchronization across stages, enabling high-quality, real-time auditoryization of complex sound scenes with changing elements.

Benefits of technology

This approach enables high-quality, real-time rendering of complex audio scenes with dynamically changing elements, such as moving sources and listeners, ensuring a perceptually compelling soundscape without audible artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675089000001
    Figure 0007675089000001
  • Figure 0007675089000002
    Figure 0007675089000002
  • Figure 0007675089000003
    Figure 0007675089000003
Patent Text Reader

Abstract

An apparatus for rendering a sound scene (50), comprising: a first pipeline stage (200) having a first control layer (201) and a reconfigurable first audio data processor (202), the first pipeline stage (200) being configured to operate according to a first configuration of the reconfigurable first audio data processor (202); a second pipeline stage (300) located after the first pipeline stage (200) in a pipeline flow, the second pipeline stage (300) having a second control layer (301) and a reconfigurable second audio data processor (302), the second pipeline stage (300) being configured to operate according to a first configuration of the reconfigurable second audio data processor (302); and a control circuit for controlling the first control layer (201) and the second control layer (301) in response to the sound scene (50). a central controller (100) for reconfiguring the reconfigurable first audio data processor (202) into the second configuration for the reconfigurable first audio data processor (202), wherein the first control layer (201) prepares a second configuration for the reconfigurable first audio data processor (202) during or after operation of the reconfigurable first audio data processor (202) in the first configuration of the reconfigurable first audio data processor (202), or the second control layer (301) prepares a second configuration for the reconfigurable second audio data processor (302) during or after operation of the reconfigurable second audio data processor (302) in the first configuration of the reconfigurable second audio data processor (302), wherein the central controller (100) controls, at a particular moment,The switch control (110) is configured to control the first control layer (201) or the second control layer (301).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to audio processing, in particular audio signal processing of sound scenes occurring, for example, in virtual reality or augmented reality applications. [Background technology]

[0002] Geometric acoustics is applied to auralization, i.e. real-time and offline audio rendering of auditory scenes and environments. This includes virtual reality (VR) and augmented reality (AR) systems, such as the MPEG-I 6-DoF audio renderer. To render complex audio scenes with six degrees of freedom (DoF), the field of geometric acoustics is applied, where the propagation of sound data is modeled using methods known from optics, such as ray tracing. In particular, reflections on walls are modeled based on models derived from optics, where the angle of incidence of a ray reflected on a wall results in an angle of reflection equal to the angle of incidence.

[0003] Real-time auralization systems, such as audio renderers in virtual reality (VR) or augmented reality (AR) systems, typically render early reflections based on geometric data of the reflecting environment. Then, geometric acoustic methods, such as image source methods combined with ray tracing, are used to find valid propagation paths of the reflected sound. These methods are effective when the reflecting plane is large compared to the wavelength of the incident sound. The distance of the reflecting point on the surface to the boundary of the reflecting plane must also be large compared to the wavelength of the incident sound.

[0004] Virtual reality (VR) or augmented reality (AR) sounds are rendered to a listener (user). The input to this process is the (typically anechoic) audio signal of the sound source. A number of signal processing techniques are then applied to these input signals to simulate and incorporate relevant acoustic effects such as sound transmission through walls / windows / doors, ambient diffraction and occlusion by solid or permeable structures, sound propagation over longer distances, reflections in semi-open and closed environments, Doppler shift of a moving source / listener, etc. The output of the audio rendering is an audio signal that when delivered to the listener via headphones or loudspeakers creates a realistic three-dimensional acoustic impression of the presented VR / AR scene.

[0005] Rendering is performed in a listener-centric manner, and the system must react instantaneously to user movements and interactions without significant delay. Processing of audio signals must therefore be done in real time. User input manifests itself in changes to the signal processing (e.g., different filters). These changes should be incorporated into the rendering without audible artifacts.

[0006] Most audio renderers used a predefined, fixed signal processing structure (see, for example, [1] for a block diagram applied to multiple channels) with a fixed computational time budget for each individual audio source (e.g., 16x object sources, 2x 3rd order Ambisonics). These solutions allow for the rendering of dynamic scenes by updating position-dependent filter and reverb parameters, but do not allow for dynamically adding / removing sources during runtime.

[0007] Furthermore, a fixed signal processing architecture can be rather ineffective when rendering complex scenes, since a large number of sources must be processed in the same way. Newer rendering concepts facilitate the concepts of clustering and level of detail (LOD), where sources are combined and rendered with different signal processing, depending on the perception. Source clustering (see [2]) can enable renderers to handle complex scenes containing hundreds of objects. In such settings, the cluster budget is still fixed, which can result in audible artifacts of extensive clustering in complex scenes. Summary of the Invention [Problem to be solved by the invention]

[0008] It is an object of the present invention to provide an improved concept for rendering an audio scene. [Means for solving the problem]

[0009] This object is achieved by an apparatus for rendering a sound scene according to claim 1, or a method for rendering a sound scene according to claim 21, or a computer program according to claim 22.

[0010] The invention is based on the discovery that a pipelined rendering architecture is useful for rendering complex sound scenes with many sound sources in an environment where frequent changes of the sound scene may occur. The pipelined rendering architecture comprises a first pipeline stage comprising a first control layer and a reconfigurable first audio data processor. Further, a second pipeline stage is provided, which is located after the first pipeline stage in terms of the pipeline flow. This second pipeline stage also comprises a second control layer and a reconfigurable second audio data processor. Both the first and second pipeline stages are configured to operate according to a specific configuration of the reconfigurable first audio data processor at a specific time during processing. To control the pipeline architecture, a central controller is provided for controlling the first control layer and the second control layer. The control is performed in response to the sound scene, i.e. in response to the original sound scene or to changes in the sound scene.

[0011] In order to achieve a synchronous operation of the device between all pipeline stages, when a reconfiguration task of the first or second reconfigurable audio data processor is required, the central controller controls the control layer of the pipeline stage such that during or after the operation of the reconfigurable audio data processor in the first configuration, the first control layer or the second control layer prepares another configuration, such as a second configuration of the first or second reconfigurable audio data processor. Thus, a new configuration for the reconfigurable first or second audio data processor is prepared while the reconfigurable audio data processor belonging to this pipeline stage is still operating according to a different configuration, or is configured in a different configuration if a processing task with a previous configuration has already been performed. In order to ensure that both pipeline stages operate synchronously in order to obtain a so-called "atomic operation" or "atomic update", the central controller controls the first and second control layers using a switch control to reconfigure the reconfigurable first audio data processor or the reconfigurable second audio data processor into a second different configuration at a specific moment in time. Even if only a single pipeline stage is reconfigured, embodiments of the present invention nevertheless ensure that switch control at a particular time instance causes the correct audio sample data to be processed in the audio workflow via provisioning of audio stream input or output buffers included in the corresponding rendering list.

[0012] Preferably, the apparatus for rendering a sound scene has more pipeline stages than the first and second pipeline stages, but in systems that already have the first and second pipeline stages and no additional pipeline stages, synchronized switching of pipeline stages in response to switch controls is necessary to obtain an improved, high-quality audio rendering operation that is at the same time very flexible.

[0013] In particular, in complex virtual reality scenes where a user can move in three directions and further where the user can move his or her head in three additional directions, i.e., a six-degree-of-freedom (6-DoF) scenario, frequent and abrupt changes of filters in the rendering pipeline, e.g., to switch from one head-related transfer function to another in case of a listener's head movement or a listener walking around, may be required.

[0014] Another problematic situation for high-quality flexible rendering is that the number of rendered sources constantly changes as the listener moves around the virtual or augmented reality scene. This can occur, for example, due to the fact that certain image sources become visible at a certain position of the user, or due to the fact that additional diffraction effects must be taken into account. Furthermore, another procedure is that in certain situations, clustering of many different closely spaced sources is possible, but when the user approaches these sources, the clustering is no longer feasible, since they are so close that each source needs to be rendered at its separate position. Such audio scenes are therefore problematic in that changes in filters or the number of rendered sources, or generally changes in parameters, are always required. On the other hand, to ensure that real-time rendering in complex audio environments is achievable, it is useful to distribute the different operations for rendering to different pipeline stages, so that efficient and fast rendering is possible.

[0015] A further example of a completely variable parameter is that the frequency-dependent distance attenuation and the propagation delay change with the distance between the user and the sound source as soon as the user approaches the source or image source. Similarly, the frequency-dependent characteristics of a reflective surface may change depending on the configuration between the user and the reflective object. Furthermore, depending on whether the user is closer to the diffractive object or further away from it or at a different angle, the frequency-dependent diffraction characteristics also change. Therefore, if all these tasks are distributed to different pipeline stages, a continuous modification of these pipeline stages must be possible and must be performed synchronously. All this is achieved by a central controller that controls the control layer of the pipeline stages to prepare for a new configuration during or after the operation of the corresponding configurable audio data processor in a previous configuration. In response to the switch control of all stages in the pipeline brought about by the control update via the switch control, a reconfiguration is performed at a specific moment that is identical or at least very similar between the pipeline stages in the device for rendering a sound scene.

[0016] The present invention is advantageous because it enables high-quality real-time auralization of auditory scenes with dynamically changing elements, e.g., moving sources and listeners, and thus contributes to achieving perceptually compelling soundscapes, a key element for an immersive experience of a virtual scene. Embodiments of the present invention apply separate and concurrent workflows, threads or processes that are very well suited to the context of rendering dynamic auditory scenes.

[0017] 1. Interaction Workflow: Handling the changes in the virtual scene that occur at any given point in time (e.g. user movements, user interactions, scene animations, etc.). 2. Control workflow: A snapshot of the current state of the virtual scene results in signal processing and updating of its parameters. 3. Processing Workflow: Performing real-time signal processing, i.e. taking a frame of input samples and computing the corresponding frame of output samples.

[0018] The execution of a control workflow, similar to a frame loop in visual computing, varies in execution time depending on the required computations that trigger changes. Preferred embodiments of the present invention are advantageous in that such variations in the execution of the control workflow do not have any adverse effect on the processing workflows that run simultaneously in the background. Because real-time audio is processed in blocks, the allowable computation time of the processing workflows is typically limited to a few milliseconds.

[0019] The processing workflows, which run simultaneously in the background, are processed by the first reconfigurable audio data processor and the second reconfigurable audio data processor, and the control workflows are initiated by the central controller and then implemented at the pipeline stage level by the control layers of the pipeline stages in parallel with the background operations of the processing workflows. The interaction workflows are implemented at the pipelined rendering apparatus level by the central controller's interface to external devices such as a head tracker or similar device, or are controlled by audio scenes with moving sources or geometry that represent changes in the sound scene as well as changes in the user's orientation or position, i.e. generally changes in the user's position.

[0020] The invention is advantageous in that it allows multiple objects in a scene to be coherently modified and sampled synchronously by a centralized switch control procedure, which further allows so-called atomic updates of multiple elements that must be supported by the control and processing workflows in order not to interrupt the audio processing due to changes at the highest level, i.e. the interaction workflow, or at an intermediate level, i.e. the control workflow.

[0021] A preferred embodiment of the invention relates to an apparatus for rendering sound scenes implementing a modular audio rendering pipeline, where the necessary steps for the auralization of a virtual auditory scene are divided into several stages, each independently responsible for a specific perceptual effect. The individual division into at least two, or preferably even more, individual pipeline stages is application dependent and is preferably defined by the creator of the rendering system, as will be shown later.

[0022] The present invention provides a general structure for a rendering pipeline that facilitates parallel processing and dynamic reconfiguration of signal processing parameters depending on the current state of a virtual scene. In the process, embodiments of the present invention:

[0023] a) each stage can dynamically change their DSP processing (e.g. number of channels, updated filter coefficients) without producing audible artifacts, and any updates to the rendering pipeline are handled synchronously and atomically as necessary based on recent changes in the scene; b) Scene changes (e.g., listener movements) can be received at any time and do not affect the real-time performance of the system, especially the DSP processing.

[0024] c) Individual stages can benefit from the capabilities of other stages in the pipeline (e.g. unified directional rendering for primary and image sources or opacity clustering to reduce complexity). Ensure that: Preferred embodiments of the present invention are described below with reference to the accompanying drawings. [Brief description of the drawings]

[0025] [Figure 1] FIG. 1 is an input / output diagram of the rendering stage. [Diagram 2] FIG. 13 illustrates state transitions of rendering items. [Diagram 3] FIG. 1 illustrates an overview of a rendering pipeline. [Figure 4] FIG. 1 illustrates an example structure of a virtual reality auralization pipeline. [Diagram 5] FIG. 1 shows a preferred implementation of an apparatus for rendering a sound scene. [Figure 6] FIG. 1 illustrates an example implementation for modifying metadata for an existing rendering item. [Figure 7] FIG. 13 illustrates another example for reducing rendering items, e.g., by clustering. [Figure 8] FIG. 13 illustrates another exemplary implementation for adding new rendering items for early reflections, etc. [Figure 9] 1 is a flow chart illustrating the control flow from a high level event that is an audio scene (change), to a low level fade-in or fade-out of old and new items, or a cross-fade of filters or parameters. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026] Fig. 5 shows an apparatus for rendering a sound scene or audio scene received by the central controller 100. The apparatus comprises a first pipeline stage 200 having a first control layer 201 and a reconfigurable first audio data processor 202. Furthermore, the apparatus comprises a second pipeline stage 300 located after the first pipeline stage 200 with respect to the pipeline flow. The second pipeline stage 300 can be located immediately after the first pipeline stage 200 or with one or more pipeline stages between the pipeline stage 300 and the pipeline stage 200. The second pipeline stage 300 comprises a second control layer 301 and a reconfigurable second audio data processor 302. Furthermore, an optional nth pipeline stage 400 is shown, comprising an nth control layer 401 and a reconfigurable nth audio data processor 402. In the exemplary embodiment of Fig. 5, the result of the pipeline stage 400 is the already rendered audio scene, i.e. the result of all processing of the audio scene or audio scene changes that have arrived at the central controller 100. The central controller 100 is configured to control the first control layer 201 and the second control layer 301 in response to a sound scene.

[0027] Responsive to a sound scene means representing the complete sound scene processed by the central controller 100 in response to an input of the entire scene at a particular initialization or start moment, or in response to a sound scene change, together with the preceding scene present before the sound scene changes again. In particular, the central controller 100 controls the first and second control layers, and, if available, any other control layers, such as the nth control layer 401, such that a new or second configuration of the first reconfigurable audio data processor, the second reconfigurable audio data processor and / or the nth reconfigurable audio data processor is prepared while the corresponding reconfigurable audio data processor operates in the background according to the previous or first configuration. In this background mode, it is not determined whether the reconfigurable audio data processor is still operating, i.e. receiving input samples and calculating output samples. Instead, it may also be a situation in which a particular pipeline stage has already completed its task. Thus, the preparation of the new configuration takes place during or after the operation of the corresponding reconfigurable audio data processor in the previous configuration.

[0028] To ensure that atomic updates of the individual pipeline stages 200, 300, 400 are possible, the central controller outputs switch controls 110 to reconfigure the individual reconfigurable first or second audio data processors at a particular moment. Depending on a particular application or sound scene change, only a single pipeline stage may be reconfigured at a particular moment, or both of two pipeline stages, such as pipeline stages 200, 300, are reconfigured at a particular moment, or it is also possible to provide switch controls to reconfigure at a particular moment for all pipeline stages of the entire device for rendering a sound scene, or only for a subgroup having more than two pipeline stages but less than all pipeline stages. For this purpose, the central controller 100 has, in addition to the processing workflow connections connecting the pipeline stages in series, a control line to each control layer of the corresponding pipeline stage. Furthermore, a control workflow connection, which will be described later, may also be provided via a first structure for the central switch control 110. However, in a preferred embodiment, the control workflow is also executed via a serial connection between the pipeline stages, whereby a central connection between each control layer of the individual pipeline stages and the central controller 100 is reserved only for the switch control 110 in order to obtain atomic updates and thus accurate, high quality audio rendering even in complex environments.

[0029] In the following sections, we describe a typical audio rendering pipeline that consists of independent rendering stages, each with a separate, synchronous control and processing workflow (Figure 1). A high-level controller ensures that all stages in the pipeline can be updated together atomically.

[0030] Every rendering stage has a control and a processing section with separate inputs and outputs corresponding to the control and processing workflow, respectively. In the pipeline, the output of one rendering stage is the input of the subsequent rendering stage, but the common interface ensures that rendering stages can be rearranged and replaced depending on the application.

[0031] This common interface is described as a flat list of render items that are provided to the rendering stage of the control workflow. A render item combines processing instructions (i.e. metadata such as position, orientation, equalization, etc.) with an audio stream buffer (single-channel or multi-channel). The mapping of buffers to render items is arbitrary, and multiple render items can reference the same buffer.

[0032] Every rendering stage ensures that the subsequent stage can read the correct audio samples from the audio stream buffers corresponding to the connected rendering items at the speed of the processing workflow. To achieve this, every rendering stage creates a processing diagram from information in the rendering items that describes the required DSP steps and their input and output buffers. Additional data may be required to build the processing diagram (e.g., geometry in the scene or a personalized HRIR set) and is provided by the controller. After the control updates are propagated throughout the pipeline, the processing diagrams are queued for synchronization and passed to the processing workflow simultaneously for all rendering stages. The exchange of processing diagrams is triggered without interfering with the real-time audio block rate, but individual stages must ensure that the exchange does not result in audible artifacts. If the rendering stage acts only on metadata, the DSP workflow can be no action.

[0033] The controller maintains a list of rendering items that correspond to real audio sources in the virtual scene. In the control workflow, the controller initiates a new control update by passing a new list of rendering items to the first rendering stage and atomically accumulating all metadata changes resulting from user interactions and other modifications of the virtual scene. Control updates are triggered at a fixed rate that may depend on available computational resources, but only after the previous update has finished. The rendering stage creates a new list of output rendering items from the input list. In the process, it can modify existing metadata (e.g., add equalization characteristics), add new rendering items, and deactivate or remove existing rendering items. The rendering items follow a defined life cycle (Figure 2) that is communicated via state indicators (e.g., "activate", "deactivate", "is active", "is inactive") on each rendering item. This allows subsequent rendering stages to update their DSP diagrams according to newly created or retired rendering items. Artifact-free fading in and fading out of rendering items upon state changes is handled by the controller.

[0034] In real-time applications, the processing workflow is triggered by callbacks from the audio hardware: when a new block of samples is required, the controller fills the rendering item buffer it holds with input samples (e.g., from disk or from an incoming audio stream). The controller then sequentially triggers the processing units of the rendering stages that act on the audio stream buffer according to their current processing diagram.

[0035] The rendering pipeline may include one or more spatialization stages (Figure 3) similar to the rendering stage, but the output of those processing stages is a mixed representation of the entire virtual auditory scene as described by the final list of rendering items, which can be directly played back in the specified playback method (e.g., binaural over headphones or multi-channel loudspeaker setup). However, spatialization may be followed by additional rendering stages (e.g., to limit the dynamic range of the output signal). Advantages of the proposed solution

[0036] Compared to the state of the art, our audio rendering pipeline is able to handle highly dynamic scenes with the flexibility to adapt the processing to different hardware or user requirements. In this section, some advances over established methods are listed. New audio elements can be added and removed from the virtual scene at runtime. Similarly, rendering stages can dynamically adjust the level of detail of their rendering based on available computational resources and perceptual requirements.

[0037] Depending on the application, rendering stages can be reordered, or new rendering stages can be inserted anywhere in the pipeline (e.g., the clustering or visualization stage), without having to modify other parts of the software. The implementation of individual rendering stages can be changed without having to modify other rendering stages. Multiple stereoscopic renderings can share a common processing pipeline, so e.g. Multi-user VR setups, headphones and loudspeaker rendering can be done in parallel with minimal computational effort.

[0038] Virtual scene changes (e.g. caused by a high-speed head-tracking device) are accumulated with a dynamically adjustable controlled rate, reducing the computational effort for e.g. filter switching. At the same time, scene updates that explicitly require atomicity (e.g. translation of an audio source) are guaranteed to be performed simultaneously across all rendering stages. Control and processing speed can be adjusted separately based on user and (audio playback) hardware requirements. Working Example

[0039] A practical example of a rendering pipeline for creating a virtual acoustic environment for a VR application may include the following rendering stages in the given order (see also FIG. 4):

[0040] 1. Transmission: Reducing complex scenes with multiple adjacent subspaces by downmixing the signal and reverb of parts far from the listener into a single rendering item (possibly with spatial extent).

[0041] Processing section: Downmixing the signal into a combined audio stream buffer, and processing the audio samples using established techniques to create late reverberation

[0042] 2. Extent: Rendering the perceived effect of a spatially extended sound source by creating multiple spatially separated rendering items. Processing section: Distribution of the input audio signal into several buffers for new rendering items (possibly with additional processing such as decorrelation)

[0043] 3. Early Reflections: Incorporate perceptually relevant geometric reflections into surfaces by creating representative rendering items with corresponding equalization and position metadata. Processing section: Distribution of the input audio signal to several buffers for new rendering items.

[0044] 4. Clustering: Combine multiple rendering items with perceptually indistinguishable locations into a single rendering item to reduce the computational complexity of subsequent stages. Processing section: Downmixing the signal into a combined audio stream buffer 5. Diffraction: Adds the perceptual effect of propagation path obstruction and diffraction due to geometry. 6. Propagation: Rendering the perceptual effects on the propagation path (e.g. direction-dependent radiation characteristics, medium absorption, propagation delay, etc.). Processing section: filtering, fractional delay lines, etc. 7. Binaural Spatialization: Render the remaining rendering items into a listener-centered binaural sound output. Processing section: HRIR filtering, downmix, etc.

[0045] 1 to 4 are now restated. Fig. 1 shows a first pipeline stage 200, also called a "render stage", comprising, for example, a control layer 201, shown as "controller" in Fig. 1, and a reconfigurable first audio data processor 202, shown as "DSP" (digital signal processor). However, the pipeline stage or rendering stage 200 of Fig. 1 can also be considered as the second pipeline stage 300 of Fig. 1 or the n-th pipeline stage 400 of Fig. 5.

[0046] The pipeline stage 200 receives an input rendering list 500 as input via an input interface and outputs an output rendering list 600 via an output interface. In the case of connection immediately after the second pipeline stage 300 in Figure 5, the input rendering list of the second pipeline stage 300 becomes the output rendering list 600 of the first pipeline stage 200 since the pipeline stages are connected in series for pipeline flow.

[0047] Each rendering list 500 contains a selection of rendering items indicated by a column of the input rendering list 500 or the output rendering list 600. Each rendering item comprises a rendering item identifier 501, rendering item metadata 502, indicated as "x" in Fig. 1, and one or more audio stream buffers depending on the number of audio objects or individual audio streams belonging to the rendering item. The audio stream buffers are indicated with "O" and are preferably implemented by memory references to real physical buffers in a word memory part of the device for rendering sound scenes, which can for example be managed by a central controller or by any other memory management method. Alternatively, the rendering list can contain audio stream buffers representing physical memory parts, but it is preferable to implement the audio stream buffers 503 as said references to a specific physical memory.

[0048] Similarly, the output rendering list 600 also has one column for each rendering item, the corresponding rendering item being identified by a rendering item identification 601, a corresponding metadata 602, and an audio stream buffer 603. The metadata 502 or 602 for a rendering item may include the location of the source, the type of source, the equalizer associated with the particular source, or in general, the frequency selection behavior associated with the particular source. Thus, the pipeline stage 200 receives the input rendering list 500 as input and generates the output rendering list 600 as output. Within the DSP 202, the audio sample values ​​identified by the corresponding audio stream buffer are processed as required by the corresponding configuration of the reconfigurable audio data processor 202, for example as illustrated by a particular processing diagram generated by the control layer 201 for the digital signal processor 202. The input rendering list 500 may, for example, include three rendering items and the output rendering list 600 may, for example, include four rendering items, i.e., more rendering items than the input, so that the pipeline stage 202 may, for example, perform an upmix. Another implementation may be, for example, that a first rendering item with four audio signals is downmixed to a rendering item with a single channel. The second rendering item may be left unchanged by the processing, i.e., may be, for example, only copied from input to output, and the third rendering item may be, for example, left unchanged by the rendering stage. For example, only the last output rendering item in the output rendering list 600 may be generated by the DSP by combining the second and third rendering items of the input rendering list 500 into a single output audio stream for the corresponding audio stream buffer of the fourth rendering item of the output rendering list.

[0049] Fig. 2 shows a state diagram for defining the "live" of a rendering item. The corresponding states of the state diagram are preferably also stored in the rendering item metadata 502 or in the rendering item identification field. At the start node 510, two different activation methods can be performed. One method is a normal activation to reach the activation state 511. The other method is an immediate activation procedure for already reaching the active state 512. The difference between both procedures is that a fade-in procedure is performed from the activation state 511 to the active state 512.

[0050] If a rendering item is active, it is processed and can be either immediately deactivated or normally deactivated. In the latter case, a deactivation state 514 is obtained and a fade-out procedure is performed to go from the deactivation state 514 to the inactive state 513. In the case of immediate deactivation, a direct transition from state 512 to state 513 is performed. The inactive state can return to immediate reactivation or enter a reactivation command to reach the activation state 511, or if neither reactivation nor immediate reactivation control is obtained, control can proceed to the located output node 515.

[0051] FIG. 3 shows an overview of a rendering pipeline, where an audio scene is shown in block 50 and the individual control flows are also shown. The central switch control flow is shown at 110. The control workflow 130 is shown entering the first stage 200 from the controller 100 and from there via a corresponding serial control workflow line 120. Thus, FIG. 3 shows an implementation in which the control workflow is also fed to the starting stage of the pipeline and propagated from there successively to the final stage. Similarly, the processing workflow 120 starts from the controller 120 via the reconfigurable audio data processors of the individual pipeline stages and enters the final stage, of which FIG. 3 shows two final stages, namely one loudspeaker output stage or spatializer 1 stage 400a or a headphone spatializer output stage 400b.

[0052] FIG. 4 shows an exemplary virtual reality rendering pipeline with an audio scene representation 50, a controller 100, and a transmission pipeline stage 200 as a first pipeline stage. The second pipeline stage 300 is implemented as an extent rendering stage. The third pipeline stage 400 is implemented as an early reflection pipeline stage. The fourth pipeline stage is implemented as a clustering pipeline stage 551. The fifth pipeline stage is implemented as a diffraction pipeline stage 552. The sixth pipeline stage is implemented as a propagation pipeline stage 553, and finally the seventh pipeline stage 554 is implemented as a binaural stereoscopic to finally obtain a headphone signal for headphones worn by a listener navigating within the virtual reality or augmented reality audio scene.

[0053] Subsequently, Figures 6, 7, and 8 are shown and described to provide specific examples of how pipeline stages can be configured and how pipeline stages can be reconfigured. FIG. 6 illustrates the procedure for modifying metadata for an existing rendering item. scenario

[0054] The two object audio sources are represented as two Rendering Items (RIs). The Directivity Stage is responsible for directional filtering of the sound source signals. The Propagation Stage is responsible for rendering the propagation delay based on the distance to the listener. The Binaural Spatializer is responsible for binauralization and downmixing of the scene to a binaural stereo signal.

[0055] A control step requires changes in the DSP processing of the individual stages because the RI position changes with respect to the previous control step. The acoustic scene should be updated synchronously, so that, for example, the perceived effect of changing distance is synchronized with the perceived effect of a change in the angle of incidence relative to the listener. Implementation

[0056] The Render List is propagated through the complete pipeline at each control step. During the control step, the DSP processing parameters remain constant for all stages until the last Stage / Spatializer processes the new Render List. After that, all Stages change their DSP parameters synchronously at the start of the next DSP step.

[0057] It is the responsibility of each Stage to update the DSP processing parameters without noticeable artifacts (e.g. output crossfades for FIR filter updates, linear interpolation for delay lines).

[0058] RI can contain fields for metadata pooling. This way, for example, a Directivity stage does not need to filter the signal itself, but can update an EQ field in the RI metadata. Subsequent EQ stages then apply the combined EQ fields of all preceding stages to the signal. Key benefits - The guaranteed atomicity of a scene changes (both between stages and between RIs) -Larger DSP reconstructions are performed synchronously when all Stage / Spatializers are ready, without blocking audio processing

[0059] -With clearly defined responsibilities, other stages in the pipeline are independent of the algorithm used for a particular task (e.g., clustering method or even availability). -Metadata pooling allows many stages (Directivity, Occlusion, etc.) to operate only on control steps.

[0060] In particular, the input rendering list is the same as the output rendering list 500 of the example of Figure 6. In particular, the rendering list has a first rendering item 511 and a second rendering item 512, each of which has a single audio stream buffer.

[0061] In the first rendering or pipeline stage 200, which in this example is a directional stage, a first FIR filter 211 is applied to the first rendering item and another directional or FIR filter 212 is applied to the second rendering item 512. Furthermore, within the second rendering stage or second pipeline stage 33, which in this embodiment is a propagation stage, a first interpolation delay line 311 is applied to the first rendering item 511 and another second interpolation delay line 312 is applied to the second rendering item 512.

[0062] Also, in the third pipeline stage 400 connected following the second pipeline stage 300, the first stereo FIR filter 411 for the first rendering item 511 is used, and the second FIR filter 412 or the second rendering item 512 is used. In the binaural spatializer, a downmix of the two filter output data is performed in the adder 413 to obtain a binaural output signal. This generates two object signals represented by the rendering items 511, 512, a binaural signal at the output of the adder 413 (not shown in FIG. 6). Thus, as explained, all elements 211, 212, 311, 312, 411, 412 are changed in response to switch control at the same specific moment under the control of the control layers 201, 301, 401. In FIG. 6, a situation is shown in which the number of objects represented in the rendering list 500 remains the same, but the metadata for the objects changes due to the different positions of the objects. Alternatively, the metadata of the objects, and in particular the object's position, remains the same, but taking into account the movement of the listener, the relationship between the listener and the corresponding (fixed) object changes, resulting in the FIR filters 211, 212 changing, the delay lines 311, 312 changing and the FIR filters 411, 412 changing, which are implemented, for example, as head-related transfer function filters that change with each change in the source or object position or the listener position, as measured, for example, by a head tracker. FIG. 7 shows a further example related to rendering item reduction (by clustering). scenario

[0063] In a complex auditory scene, the Render List may contain many RIs that are perceptually close, i.e., the difference between their positions cannot be distinguished by the listener. To reduce the computational load of subsequent stages, the Clustering Stage can replace multiple individual RIs with a single representative RI.

[0064] At some control step, the scene composition may change such that clustering is no longer perceptually feasible, in which case the Clustering Stage becomes inactive and passes the Render List on unchanged. Implementation

[0065] When some incoming RIs are clustered, the original RIs are deactivated in the outgoing Render List. The reduction is opaque to subsequent Stages, and the Clustering Stage must ensure that valid samples are provided to the buffer associated with the representative RI as soon as the new outgoing Render List becomes active.

[0066] When a cluster becomes infeasible, the new outgoing Render List of the Clustering stage contains the original, unclustered RIs, and subsequent stages have to process them separately, starting from the next DSP parameter change (e.g. by adding a new FIR filter, delay line, etc. to their DSP diagram).

[0067] Key benefits -RI opacity reduction reduces the computational load of subsequent stages without explicit reconfiguration Due to the atomicity of DSP parameter changes, the Stage can handle a variable number of incoming and outgoing RIs without artifacts. In the example of FIG. 7, the input rendering list 500 includes three rendering items 521, 522, 523, and the output renderer 600 includes two rendering items 623, 624.

[0068] The first rendering item 521 comes from the output of the FIR filter 221. The second rendering item 522 is generated by the output of the FIR filter 222 of the directional stage and the third rendering item 523 is obtained at the output of the FIR filter 223 of the first pipeline stage 200, which is the directional stage. It should be noted that when it is outlined that a rendering item is at the output of a filter, this refers to the audio samples of the audio stream buffer of the corresponding rendering item.

[0069] 7, rendering item 523 is not affected by the clustering state 300 and becomes output rendering item 623. However, rendering item 521 and rendering item 522 are downmixed to a downmix rendering item 324 that emerges in the renderer 600 as output rendering item 624. The downmix in the clustering stage 300 is indicated by location 321 for the first rendering item 521 and location 322 for the second rendering item 522.

[0070] Again, the third pipeline stage in FIG. 7 is binaural spatialisation 400, where rendering item 624 is processed by a first stereo FIR filter 424 and rendering item 623 is processed by a stereo filter FIR filter 423, and the outputs of both filters are summed in summer 413 to give the binaural output. FIG. 8 shows another example showing the addition of a new rendering item (for early reflections).

[0071] scenario In geometric room acoustics, it can be beneficial to model reflected sound as an image source (i.e. two point sources with the same signal and their positions are mirrored on the reflecting surfaces). If the configuration between the listener, source, and reflecting surfaces in the scene is suitable for reflections, the Early Reflections Stage adds a new RI to its outgoing Render List representing the image source.

[0072] The audibility of an image source typically changes rapidly as the listener moves. The Early Reflections Stage can activate and deactivate RIs at each control step, and subsequent Stages should adjust their DSP processing accordingly. Implementation

[0073] The Early Reflections Stage ensures that the associated audio buffer contains the same samples as the original RI, so stages after the Early Reflections Stage can process the reflected RI normally. In this way, perceptual effects such as propagation delays can be handled for the original RI and reflections, etc., without explicit reconstruction. To increase efficiency when the activity status of the RI changes frequently, the Stage can retain necessary DSP artifacts (such as FIR filter instances) for reuse.

[0074] A Stage can treat render items with certain characteristics differently. For example, a render item created by a Reverb Stage (represented by item 532 in Figure 8) may not be processed by the Early Reflections Stage, but only by the Spatializer. In this way, a render item can provide the functionality of a downmix bus. Similarly, a Stage can treat a render item produced by an Early Reflections Stage with lower quality DSP algorithms, since they are usually not acoustically noticeable. Key benefits -Different Render Items can be treated differently based on their characteristics. -Stages that create new Render Items can benefit from the processing of subsequent Stages without explicit reconfiguration.

[0075] The rendering list 500 includes a first rendering item 531 and a second rendering item 532. Each has a single audio stream buffer that can, for example, carry a mono or stereo signal.

[0076] The first pipeline stage 200 is for example a reverb stage with a generated render item 531. The rendering list 500 further comprises a render item 532. In the previous deflection stage 300, the render item 531, in particular its audio samples, is represented by an input 331 for a copy operation. The input 331 of the copy operation is copied to an output audio stream buffer 331, which corresponds to the audio stream buffer of the render item 631 of the output rendering list 600. Also, another copied audio object 333 corresponds to the render item 633. Furthermore, as mentioned above, the render item 532 of the input rendering list 500 is simply copied or fed to the render item 632 of the output rendering list.

[0077] Then, in the third pipeline stage, i.e. in the above example, in binaural spatialization, the stereo FIR filter 431 is applied to the first rendering item 631, the stereo FIR filter 433 is applied to the second rendering item 633 and the third stereo FIR filter 432 is applied to the third rendering item 632. The contributions of all three filters are then correspondingly added, i.e. added per channel by the adder 413, the output of which is the left signal on the one hand and the right signal on the other hand for headphones or generally binaural playback.

[0078] FIG. 9 shows an overview of the individual control procedures, from the high-level control by the audio scene interface of the central controller to the low-level control performed by the control layers of the pipeline stages.

[0079] At certain points in time, which may be irregular and dependent on the listener's behavior, for example as determined by a head tracker, the central controller receives an audio scene or a change in the audio scene, as indicated by step 91. In step 92, the central controller determines a rendering list for each pipeline stage under the control of the central controller. In particular, the control updates sent from the central controller to the individual pipeline stages are triggered at a regular rate, i.e. at a certain update rate or update frequency.

[0080] The central controller sends the individual rendering lists to the respective pipeline stage control layers, as shown in step 93. This can be done centrally, for example via a switch control infrastructure, but preferably this is done sequentially through the first pipeline stage and from there to the next pipeline stage, as shown by control workflow line 130 in Figure 3. In a further step 94, each control layer builds its corresponding processing diagram for a new configuration for the corresponding reconfigurable audio data processor, as shown in step 94. The old configuration is also denoted to be the "first configuration" and the new configuration is denoted to be the "second configuration".

[0081] In step 95, the control layer receives a switch control from the central controller and reconfigures its associated reconfigurable audio data processor to the new configuration. This control layer switch control reception in step 95 can be in response to the reception of a ready message for all pipeline stages by the central controller, or in response to the sending of a corresponding switch control instruction from the central controller after a certain period of time relative to the update trigger, as was done in step 93. Then, in step 96, the control layer of the corresponding pipeline stage takes care of fading out items that are not present in the new configuration, or of fading in new items that were not present in the old configuration. In the case of the same object in the old and new configurations, and in the case of metadata changes regarding distances to sources or new HRTF filters, etc. due to listener head movements, etc., the cross-fading of filters or cross-fading of filtered data to come smoothly from one distance, for example, to the other, is also controlled by the control layer in step 96.

[0082] The actual processing in the new configuration is initiated by a callback from the audio hardware. In other words, the processing workflow is therefore triggered after the reconfiguration to the new configuration in the preferred embodiment. When a new block of samples is requested, the central controller fills the audio stream buffers of the rendering items it holds with input samples from the disk or from the incoming audio stream. The controller then sequentially triggers the processing units of the rendering stage, i.e. the reconfigurable audio data processors, which act on the audio stream buffers according to their current configuration, i.e. according to their current processing diagram. Thus, the central controller fills the audio stream buffers of the first pipeline stage in the device for rendering a sound scene. However, there are also situations in which the input buffers of the other pipeline stages are filled from the central controller. This situation may arise, for example, if there was no spatially extended sound source in the previous situation of the audio scene. Thus, in this previous situation, the stage 300 of FIG. 4 did not exist. However, the listener may then move to a particular location in the virtual audio scene where a spatially extended sound source is visible, or the listener is so close to this sound source that it must be rendered as a spatially extended sound source.The central controller 100 then supplies a new rendering list, typically via the transmission stage 200, to the extension rendering stage 300, in order to introduce this spatially extended sound source at this point via block 300.

[0083] References [1] Wenzel, E. M., Miller, J. D., and Abel, J. S. "Sound Lab: A real-time, software-based system for the study of spatial hearing." Audio Engineering Society Convention 108. Audio Engineering Society, 2000.

[0084] [2] Tsingos, N., Gallo, E., and Drettakis, G "Perceptual audio rendering of complex virtual environments." ACM Transactions on Graphics (TOG) 23.3 (2004): 249-258.

Claims

1. An apparatus for rendering a sound scene (50), comprising: a first pipeline stage (200) comprising a first control layer (201) and a reconfigurable first audio data processor (202), the reconfigurable first audio data processor (202) configured to operate according to a first configuration of the reconfigurable first audio data processor (202); a second pipeline stage (300) located after the first pipeline stage (200) with respect to a pipeline flow, the second pipeline stage (300) comprising a second control layer (301) and a reconfigurable second audio data processor (302), the reconfigurable second audio data processor (302) configured to operate according to a first configuration of the reconfigurable second audio data processor (302); a central controller (100) for controlling the first control layer (201) and the second control layer (301) in response to the sound scene (50), the first control layer (201) preparing a second configuration of the reconfigurable first audio data processor (202) during or after operation of the reconfigurable first audio data processor (202) in the first configuration of the reconfigurable first audio data processor (202) or the second control layer (301) preparing a second configuration of the reconfigurable second audio data processor (302) during or after operation of the reconfigurable second audio data processor (302) in the first configuration of the reconfigurable second audio data processor (302); Equipped with the central controller (100) is configured to control, at a particular moment, the first control layer (201) or the second control layer (301) using a switch control (110) to reconfigure the first reconfigurable audio data processor (202) into the second configuration for the first reconfigurable audio data processor (202) or to reconfigure the second reconfigurable audio data processor (302) into the second configuration for the second reconfigurable audio data processor (302), An apparatus for rendering a sound scene (50).

2. The central controller (100) controlling the first control layer (201) to prepare the second configuration of the reconfigurable first audio data processor (202) during the operation of the reconfigurable first audio data processor (202) in the first configuration of the reconfigurable first audio data processor (202); during the operation of the reconfigurable second audio data processor (302) in the first configuration of the reconfigurable second audio data processor (302), controlling the second control layer (301) to prepare the second configuration of the reconfigurable second audio data processor (302); using said switch control (110) to control said first control layer (201) and said second control layer (301) to reconfigure said first reconfigurable audio data processor (202) into said second configuration for said first reconfigurable audio data processor (202) and to reconfigure said second reconfigurable audio data processor (302) into said second configuration for said second reconfigurable audio data processor (302) at said particular moment in time. The apparatus of claim 1 configured to:

3. the first pipeline stage (200) or the second pipeline stage (300) comprises an input interface configured to receive an input rendering list (500), the input rendering list (500) comprising an input list of rendering items (501), metadata (502) for each rendering item in the input list of the rendering items (501), and an audio stream buffer (503) for each rendering item in the input list of the rendering items (501); at least the first pipeline stage (200) comprises an output interface configured to output an output rendering list (600), the output rendering list (600) including an output list of rendering items (601), metadata (602) for each rendering item in the output list of the rendering items (601), and an audio stream buffer (603) for each rendering item in the output list of the rendering items (601); When the second pipeline stage (300) is connected to the first pipeline stage (200), the output rendering list (600) of the first pipeline stage (200) is the input rendering list (500) of the second pipeline stage (300).

3. Apparatus according to claim 1 or 2.

4. 4. The apparatus of claim 3, wherein the first pipeline stage (200) is configured to write audio samples to an audio stream buffer (603) indicated by the output rendering list (600) such that the second pipeline stage (300) following the first pipeline stage (200) can retrieve the audio samples from the audio stream buffer (603) at a processing workflow speed.

5. the central controller (100) is configured to provide the input rendering list (500) or the output rendering list (600) to the first pipeline stage (200) or the second pipeline stage (300), the first or second configuration of the reconfigurable first audio data processor (202) or the reconfigurable second audio data processor (302) comprising a processing diagram, the first control layer (201) or the second control layer (301) being configured to create the processing diagram for the second configuration from the input rendering list (500) or the output rendering list (600) received from the central controller (100) or a previous pipeline stage, the process diagram includes audio data processor steps and references to input and output buffers of the first reconfigurable audio data processor (202) or the second reconfigurable audio data processor (302); 4. The apparatus of claim 3.

6. 6. The apparatus of claim 5, wherein the central controller (100) is configured to provide to the first pipeline stage (200) or the second pipeline stage (300) additional data required to create the process diagram, the additional data being provided by the central controller (100).

7. the central controller (100) is configured to receive a sound scene change (50) from a previous sound scene to a current sound scene at a sound scene change moment; the central controller (100) is configured to generate, in response to the sound scene change (50) and based on the current sound scene, a first rendering list for the first pipeline stage (200) and a second rendering list for the second pipeline stage (300), and the central controller (100) is configured to transmit, following the sound scene change moment, the first rendering list for the first pipeline stage (200) to the first control layer (201) and the second rendering list for the second pipeline stage (300) to the second control layer (301); 7. Apparatus according to any one of claims 1 to 6.

8. the first control layer (201) is configured to calculate the second configuration of the first reconfigurable audio data processor (202) from the first rendering list following the sound scene change moment; the second control layer (301) is configured to calculate the second configuration of the second reconfigurable audio data processor (302) from the second rendering list; the central controller (100) is configured to trigger the switch control (110) simultaneously for the first and second pipeline stages (200, 300); 8. The apparatus of claim 7.

9. 9. The apparatus of claim 1, wherein the central controller (100) is configured to use the switch control (110) without interfering with audio sample calculation operations performed by the reconfigurable first audio data processor (202) and the reconfigurable second audio data processor (302).

10. the central controller (100) is configured to receive changes to the sound scene (50) at moments of random change (91) dependent on the listener's behavior; the central controller (100) is configured to provide control instructions to the first and second control layers (201, 301) at a constant control rate (93); the reconfigurable first audio data processor (202) and the reconfigurable second audio data processor (302) operate at an audio block rate to calculate output audio samples from input audio samples received from an input buffer of the reconfigurable first audio data processor (202) or the reconfigurable second audio data processor (302), the output audio samples being stored in an output buffer of the reconfigurable first audio data processor (202) or the reconfigurable second audio data processor (302); 10. Apparatus according to any one of claims 1 to 9.

11. the central controller (100) is configured to trigger the switch control (110) at a specific time period after controlling the first and second control layers (201, 202) to prepare the second configuration or in response to a ready signal received from the first and second pipeline stages (200, 300) indicating that the first and second pipeline stages (200, 300) are ready for a change to the second configuration; 11. Apparatus according to any one of claims 1 to 10.

12. the first pipeline stage (200) or the second pipeline stage (300) is configured to create the output rendering list (600) from the input rendering list (500); said creating includes modifying metadata of rendering items in said input rendering list (500) and writing the modified metadata to said output rendering list (600); or calculating output audio data for the rendering item using input audio data read from an input stream buffer of the input rendering list (500), and writing the output audio data to an output stream buffer of the output rendering list (600).

7. Apparatus according to any one of claims 3 to 6.

13. the first control layer (201) or the second control layer (301) is configured to control the first reconfigurable audio data processor (202) or the second reconfigurable audio data processor (302) to fade in new rendering items that are processed after the switch control (110) or to fade out old rendering items that are no longer present after the switch control (110) but were present before the switch control (110).

13. Apparatus according to any one of claims 1 to 12.

14. Each rendering item in the list of rendering items (501, 601) in the input rendering list (500) or the output rendering list (600) of the first pipeline stage (200) or the second pipeline stage (300) includes a state indicator indicating at least one of the following states: rendering is active, rendering is activated, rendering is inactive, rendering is deactivated, 4. The apparatus of claim 3.

15. the central controller (100) is configured to fill an input buffer of rendering items maintained by the central controller (100) with new audio samples in response to a request from the first pipeline stage (200) or the second pipeline stage (300); the central controller (100) is configured to sequentially trigger the first reconfigurable audio data processor (202) and the second reconfigurable audio data processor (302), whereby the first reconfigurable audio data processor (202) and the second reconfigurable audio data processor (302) act on the input buffers of the rendering items according to the first configuration or the second configuration, depending on which configuration is currently active; 15. Apparatus according to any one of claims 1 to 14.

16. The second pipeline stage (300) is a spatialization stage that provides as output a channel representation for headphone playback or a loudspeaker setup.

16. Apparatus according to any one of claims 1 to 15.

17. The first and second pipeline stages (200, 300) include: a transmission stage (200), an extent stage (300), an early reflection stage (400), a clustering stage (551), a diffraction stage (552), a propagation stage (553), and a stereolithography stage (554); 16. The apparatus of claim 1 , further comprising at least one of:

18. the first pipeline stage (200) is a direction stage (200) for one or more rendering items, and the second pipeline stage (300) is a propagation stage (300) for one or more rendering items; the central controller (100) is configured to receive a change in the sound scene (50) indicating that the one or more rendering items have one or more new positions; the central controller (100) is configured to control the first control layer (201) and the second control layer (301) to adapt filter settings for the first reconfigurable audio data processor (202) and the second reconfigurable audio data processor (302) to the one or more new positions; the first control layer (201) or the second control layer (301) is configured to change to the second configuration at the particular moment, and when changing to the second configuration, a cross-fade operation from the first configuration to the second configuration is performed by the first control layer (201) and the second control layer (301); 17. Apparatus according to any one of claims 1 to 16.

19. the first pipeline stage (200) is a directional stage (200) and the second pipeline stage (300) is a clustering stage (300); the central controller (100) is configured to receive a change in the sound scene (50) indicating that clustering of the rendering items should be stopped; the central controller (100) is configured to control the first control layer (201) to deactivate the reconfigurable second audio data processor (302) of the clustering stage and to copy the rendering items of the input rendering list (500) to the rendering items of the output rendering list (600) of the second pipeline stage (300), 7. Apparatus according to any one of claims 3 to 6.

20. 1. A method of rendering a sound scene (50) using an apparatus comprising: a first pipeline stage (200) comprising a first control layer (201) and a reconfigurable first audio data processor (202), the reconfigurable first audio data processor (202) being configured to operate according to a first configuration of the reconfigurable first audio data processor (202); and a second pipeline stage (300) located after the first pipeline stage (200) in a pipeline flow, the second pipeline stage (300) comprising a second control layer (301) and a reconfigurable second audio data processor (302), the reconfigurable second audio data processor (302) being configured to operate according to a first configuration of the reconfigurable second audio data processor (302), the method comprising: controlling the first control layer (201) and the second control layer (301) in response to the sound scene (50), wherein the first control layer (201) prepares a second configuration of the reconfigurable first audio data processor (202) during or after operation of the reconfigurable first audio data processor (202) in the first configuration of the reconfigurable first audio data processor (202) or the second control layer (301) prepares a second configuration of the reconfigurable second audio data processor (302) during or after operation of the reconfigurable second audio data processor (302) in the first configuration of the reconfigurable second audio data processor (302); - controlling, at a particular moment, the first control layer (201) or the second control layer (301) using a switch control (110) to reconfigure the first reconfigurable audio data processor (202) into the second configuration for the first reconfigurable audio data processor (202) or to reconfigure the second reconfigurable audio data processor (302) into the second configuration for the second reconfigurable audio data processor (302); A method comprising:

21. A computer program for performing the method of claim 20 when the computer program is run on a computer or processor.

Citation Information

Patent Citations

  • Sound processing apparatus and sound processing method, and control program

    JP2003061200A

  • Acoustic processing device with low latency

    JP2017530413A

  • Processor with pipelined core and algorithm-matching pipeline compiler using reconfigurable algorithms

    JP2019506695A