Rendering technology

The renderer device addresses the challenge of personalized audio rendering for users with hearing impairments by processing audio scene representations with context-specific rules, enhancing accessibility and reducing latency in VR/AR and metaverse applications.

JP2026515680APending Publication Date: 2026-05-19FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2024-04-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing audio rendering technologies fail to personalize and adapt to the user's specific profile and environment, particularly for users with hearing impairments, leading to increased latency and complexity, and limited accessibility in VR/AR and metaverse applications.

Method used

A renderer device that processes audio scene representations using context-specific rules and parameters, including position, orientation, and environmental data to generate personalized audio signals, adapting to individual hearing impairments and environmental conditions.

Benefits of technology

Enhances accessibility and reduces latency by personalizing audio rendering for users with hearing impairments, improving the user experience in VR/AR and metaverse environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515680000030
    Figure 2026515680000030
  • Figure 2026515680000031
    Figure 2026515680000031
  • Figure 2026515680000032
    Figure 2026515680000032
Patent Text Reader

Abstract

The disclosed renderer device (400) is, A rendering unit (430) configured to process rendering audio scene representations (402, 412) and to receive at least one context-specific rule or parameter (441, 442), wherein the rendering unit (430) is configured to generate a rendered audio signal (422) from audio scene representations (402, 412) conditioned by at least one context-specific rule or parameter (441, 442), The system comprises a contextualization unit (440) configured to receive and / or derive context-specific data (461, 462), and a contextualization unit (440) configured to provide at least one context-specific rule or parameter (441, 442) to a rendering unit (430) based on the context-specific data (461, 462).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates, for example, to audio rendering for vehicles and / or hearing impairment aids. [Background technology]

[0002] The inventors noted that audio renderers generally provide audio rendering that is independent of the user or the environment in which the user consumes audio content. For example, in the case of virtual reality (VR) and augmented reality (AR), even if the user provides some feedback, this feedback is still limited to the options provided by a particular audio scene and does not necessarily fit the user's specific profile. In other words, the feedback provided by the user is the feedback that the creator anticipates (e.g., head movements to enable viewing different viewports or hearing different sounds), but it cannot actually be tailored to a particular user. For example, if the user is hearing impaired, it is desirable to render the audio scene taking the user's disability into account, and it is even better if the content is redefined to suit the user's sensitivities. However, this work should be done during authoring, which increases the burden on the content creator.

[0003] A specific example is explained below. Because social VR and the metaverse promise and aim to connect everyone, audio technology needs to be designed to accommodate users of all ages and a wide range of hearing abilities. However, in most consumer electronics, the typical target user group is the average hearing population, and people with hearing impairments receive little consideration.

[0004] The populations of all industrialized countries are aging rapidly. In Germany, the population aged 67 and over will increase by 22% by 2035. Already, the median age in Germany and Japan is approximately 48. Globally, the group of people over 65 will significantly outnumber the total population.

[0005] According to a study titled "Hearing Loss in Old Age - Characteristics and Location" (https: / / www.aerzteblatt.de / archiv / 48807 / Hoerminderung-im-Alter-Auspraegung-und-Lokalisation), "In old age, hearing loss is statistically more likely, although it is not a natural process. [...] The majority of hearing loss in older adults is due to changes in inner ear hair cells and the degenerative process of the central auditory pathway. Of older adults who clearly indicate a need for hearing aids, only 15% actually use them. The reasons for this shortage are that expectations for hearing aids are too high, tolerance is low, and there are technically unresolved speech processing strategies to compensate for the central nervous system component of age-related hearing loss, which is thought to play a greater role in old age." Although this publication is from 2005, the situation has generally not changed.

[0006] The global hearing aid market is projected to grow from $10.23 billion in 2022 to $17.68 billion by 2029, with the recently FDA-approved over-the-counter hearing aid segment playing a significant role.

[0007] VR / AR (along with MPEG-I as one form) a) To adapt these technologies to this large, often affluent, user group by lowering the threshold for use by people with hearing impairments, b) Making social VR, metaverse, and other future communication concepts available to specific groups. c) Open up specific areas of AR / VR applications for the hearing impaired. This could be training and learning sessions on how to use new hearing devices, or sessions that combine hearing devices with general AR technologies to enrich auditory information, for example, by providing visual transcripts of sounds heard in AR glasses. Alternatively, reverberation cancellation technology could be applied to improve speech intelligibility.

[0008] Cosmic hearing loss results in a) deterioration of hearing (increased auditory threshold), b) volume perception (reduced dynamic range), and c) frequency discrimination (recognition of auditory objects in complex sound scenes) [1].

[0009] A common strategy for adapting audio to a listener with hearing impairment is a post-filter 120, which takes the (rendered) stereo output 112 (output by rendering unit 110) and generates a separate audio signal by applying frequency-dependent amplification (EQ), frequency-dependent dynamic range compression / automatic gain control (DRC / AGC) adjustments, and (in more severe cases) frequency underpassing (see Figure 1). The EQ and volume adjustment parameters are predetermined and stored in the hearing loss profile, for example, by a hearing aid specialist or by a hearing test app on the user's phone. Common techniques for mapping frequency-dependent hearing loss to amplification (primarily to improve speech intelligibility) include CAMEQ, CAMREST, DSL[i / o], FIG6, or NAL-NL1.

[0010] To address severe high-frequency hearing loss, a technique called frequency downsampling (FL) is used. This involves compressing and shifting the spectrum so that frequency components in the range of severe hearing loss (usually high frequencies) appear in the frequency range where hearing loss is less severe.

[0011] Generally, if a wireless link (e.g., Bluetooth) exists between the playback device (where rendering takes place) and the headphones or hearing aid, audio adaptation may be handled by the headphones / hearing aid. This additional processing can introduce undesirable latency, increasing complexity and ultimately reducing the hearable's operating time due to increased battery consumption.

[0012] Devices in the newly established over-the-counter (OTC) hearing aid product category primarily focus on improving speech clarity in noisy environments, rather than selective amplification to address hearing loss.

[0013] In the context of object-based audio, accessibility has been investigated, for example, in [2]–[4].

[0014] The problem with conventional solutions is that audio adaptation can only affect the rendered signal (after rendering output). This is not ideal because it increases latency and complexity and limits the possibility of creating an easily accessible audio stream.

[0015] The possibility of implementing HRTF (Head-Related Transfer Function) is being explored. However, HRTF does not act on decoding or rendering, but rather on audio signals that have already been decoded (or rendered). Therefore, the problems of the prior art persist. [Prior art documents] [Patent Documents]

[0016] [Non-Patent Document 1] https: / / www.aerzteblatt.de / archiv / 48807 / Hoerminderung-im-Alter-Auspraegung-und-Lokalisation [Overview of the project]

Problems to be Solved by the Invention

[0017] From the above, the discovery of a rendering technique that enables personalization and environment-specific adaptation is expected.

Means for Solving the Problems

[0018] According to one aspect, a renderer device is provided. The renderer device includes: a rendering unit configured to process an audio scene representation to be rendered and receive at least one context-specific rule or parameter, and generate a rendered audio signal from the audio scene representation conditioned by the at least one context-specific rule or parameter; and a contextualization unit configured to receive and / or derive context-specific data, and provide at least one context-specific rule or parameter to the rendering unit based on the context-specific data.

[0019] According to one aspect, the rendering unit is configured to process an audio scene representation as including audio elements, and generate a rendered audio signal from the audio elements and the at least one context-specific rule or parameter.

[0020] According to one aspect, the rendering unit is configured to process an audio scene representation including audio elements and metadata, and generate a rendered audio signal from the audio elements, the metadata, and the at least one context-specific rule or parameter.

[0021] According to one aspect, the metadata includes position metadata that provides information regarding at least one of the position, orientation, directivity, source, and width of at least one object to be rendered, at least one object to be rendered is part of an audio scene representation, and the rendering unit is configured to generate a rendered audio signal from an audio element, metadata, and at least one context-specific rule or parameter.

[0022] According to one aspect, the rendering unit is configured to modify the metadata based on at least one context-specific rule or parameter to obtain modified metadata, and the rendering unit is configured to apply spatial audio processing and synthesis to the audio element based on the modified metadata.

[0023] According to one aspect, the rendering unit is configured to apply spatial audio processing and synthesis to the audio element based on the metadata and at least one context-specific rule or parameter.

[0024] According to one aspect, the rendering unit is configured to combine audio elements based on the metadata and at least one context-specific rule or parameter.

[0025] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter that includes a context-specific position rule or parameter that associates a context-specific gain weight with distance, position, gesture, and / or orientation in order to correspondingly apply the context-specific gain weight to the object to be rendered based on the context-specific position rule or parameter.

[0026] In one embodiment, a context-specific position rule or parameter defines the gain weights to be frequency-dependent, thereby allowing the rendering unit to apply a first context-specific gain weight to a first frequency band and a second context-specific gain weight to a second frequency band according to the position rule or parameter.

[0027] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter, including a context-specific position rule or parameter based on a distance threshold, a position threshold, or an orientation threshold, and the rendering unit is configured to compare the distance, position, or orientation of an object to be rendered with the distance threshold, position threshold, or orientation threshold, respectively, thereby refraining from rendering the object if the distance, position, or orientation exceeds the distance threshold, position threshold, or orientation threshold, and rendering the object if the distance, position, or orientation falls below the distance threshold, position threshold, or orientation threshold.

[0028] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter based on a context-specific gain threshold, and the rendering unit is configured to compare the gain of an object to be rendered with the context-specific gain threshold, thereby refraining from rendering the object if the gain is below the context-specific gain threshold, and rendering the object if the gain is above the context-specific gain threshold.

[0029] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter, which includes a position rule or parameter that describes distance-dependent attenuation, position-dependent attenuation, or orientation-dependent attenuation of the gain of the object being rendered, the position rule or parameter includes a context-specific attenuation parameter that applies to the context-specific distance-dependent attenuation, position-dependent attenuation, or orientation-dependent attenuation.

[0030] In one embodiment, the contextualization unit is configured to define position-dependent decay as distance-dependent decay that is inversely proportional to the distance of the rendered object, increased by an exponential amount defined by a context-specific decay parameter.

[0031] In one embodiment, the contextualization unit is configured to define position-dependent decay as distance-dependent decay, which is inversely proportional to the distance of the rendered object and increases or decreases according to context-specific decay parameters.

[0032] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter to provide context-specific information regarding at least one channel-specific gain weight to be applied to the corresponding audio element of the rendered audio signal, thereby enabling the rendering unit to apply the channel-specific gain weight to the corresponding audio element of the rendered audio signal.

[0033] In one embodiment, the channel-specific weights include multiple channel-specific gains, each channel-specific gain being specific to a frequency band, thereby allowing the rendering unit to apply a first channel-specific gain weight to a first frequency band and a second channel-specific gain weight to a second frequency band according to at least one context-specific rule or parameter.

[0034] In one embodiment, the contextualization unit is configured to include a context-specific reverberation level reduction rule or parameter by defining at least one context-specific rule or parameter, thereby enabling the renderer unit to perform a context-specific reduction of the reverberation level based on the context-specific reverberation level reduction rule or parameter.

[0035] According to one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter, which includes a context-specific initial reflection level reduction rule or parameter, thereby enabling the rendering unit to perform a context-specific reduction of the initial reflection level based on the context-specific initial reflection level reduction rule or parameter.

[0036] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter that includes a dynamic range control rule or parameter, thereby enabling the rendering unit to perform dynamic range control based on the dynamic range control rule or parameter.

[0037] In one embodiment, the contextualization unit is configured to derive context-specific rules or parameters from background noise, thereby applying a higher gain to the rendered audio signal when the background noise is higher, and a lower gain to the rendered audio signal when the background noise is lower.

[0038] According to one embodiment, the contextualization unit is configured to define context-specific floor attenuation parameters that are used by the rendering unit to perform floor attenuation according to context-specific floor attenuation parameters.

[0039] In one embodiment, the contextualization unit is configured to define context-specific culling gain parameters used by the rendering unit to perform linear fade-out of objects close to a first source distance culling value for primary reflections, or to perform linear fade-out of objects close to a second source distance culling value for secondary reflections.

[0040] In one embodiment, the contextualization unit is configured to define context-specific culling gain parameters used by the rendering unit to correct cylinder reflections.

[0041] A renderer device in one embodiment is configured to process an audio scene representation to obtain a version of the audio scene representation that includes audio elements, and the renderer device is further configured to generate a rendered audio signal from the audio elements and metadata.

[0042] A renderer device according to one embodiment is configured to process an audio scene representation and obtain a version of the audio scene representation including audio elements and metadata, and the renderer device is further configured to generate a rendered audio signal from the audio elements and metadata.

[0043] One renderer device is configured to process an audio scene representation to obtain a version of the audio scene representation that includes audio elements in a core decoder block, and to process the version of the audio scene representation to generate rendered audio signals from the audio elements in a rendering block.

[0044] In one embodiment, at least one context-specific rule or parameter includes a context-specific rule or parameter for frequency band modification, thereby associating an input frequency band with an output frequency band, and thereby the rendering unit modifies at least one frequency band of the audio scene representation to a different frequency band of the audio element according to the context-specific rule or parameter for frequency band modification.

[0045] According to one embodiment, a context-specific rule or parameter for changing the frequency band reduces the frequency of at least one frequency band.

[0046] According to one embodiment, at least one context-specific rule or parameter includes a context-specific rule or parameter for frequency-dependent gain amplification that associates an input frequency band with context-specific frequency band weights, thereby the rendering unit correspondingly applies context-specific gain weights to the spectral values ​​of at least one bin of at least one frequency band according to the context-specific rule or parameter.

[0047] In one embodiment, the contextualization unit is configured to define at least one rule or parameter as a geometric range.

[0048] In one embodiment, the renderer device may be configured to use at least one context-specific rule or parameter, which includes a proposed amplification curve for each channel, and the proposed amplification curve follows a contextualized profile based on context-specific data.

[0049] In one embodiment, the renderer device is configured to use at least one context-specific rule or parameter as parameterized based on a parameter or measurement, thereby modulating at least one context-specific rule or parameter in accordance with the parameter or measurement.

[0050] In one embodiment, the rendering unit is configured to process an audio scene representation to derive a version of the audio scene representation that includes multiple metadata sets within the metadata, and the rendering unit is configured to emit one of the metadata sets based on at least one context-specific rule or parameter.

[0051] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter based on a contextualization profile containing multiple context-specific data, and the contextualization unit is configured to extract context-specific data related to the audio signal and / or rendering configuration to be rendered from the contextualization profile, and to derive at least one context-specific rule or parameter from the related context-specific data.

[0052] In one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter based on the parameters of the rendering unit, such that at least one context-specific rule or parameter conforms the parameters of the rendering unit to the contextualization profile.

[0053] A rendering unit in one embodiment, wherein the contextualization unit is configured to derive a degradation model from a contextualization profile and / or context-specific data, the degradation model indicating a specific degradation in the ability of a particular human user or another audio receiving entity to obtain a rendered audio signal, and the contextualization unit is further configured to define at least one context-specific rule or parameter to compensate for the specific degradation based on the degradation model.

[0054] According to one embodiment, the renderer device may be configured to perform simplified rendering when it is determined that the user has a hearing impairment and / or reduced cognitive and / or physical sensitivity.

[0055] According to one embodiment, the contextualization unit is configured to define at least one context-specific rule or parameter based on the contextualization settings received in the audio scene representation.

[0056] In one embodiment, the contextualization unit is configured to select at least one context-specific rule or parameter from a plurality of contextualization settings received in an audio scene representation, and to select the most appropriate context-specific rule or parameter from the plurality of received contextualization settings based on feedback such as user-specific physical and / or cognitive hearing impairment information from the user, from pre-configuration, or from manual selection.

[0057] According to one embodiment, the contextualization unit is configured to access user-specific physical and / or cognitive hearing impairment information, which provides information about user-specific physical and / or cognitive hearing impairment, and to define at least one context-specific rule or parameter as a user-specific rule or parameter based on the user-specific physical and / or cognitive hearing impairment information.

[0058] According to one embodiment, user-specific physical and / or cognitive auditory deterioration information is or includes information regarding the user's cochlear deterioration.

[0059] In one embodiment, the contextualization unit is configured to generate at least one context-specific rule, which is a user-specific rule or parameter for compensating for user-specific physical and / or cognitive hearing impairment.

[0060] In one embodiment, the renderer device may be configured to perform a configuration session to acquire user-specific physical and / or cognitive hearing deterioration information through multiple acquisitions, thereby deriving a user-specific physical and / or cognitive hearing deterioration model, and to derive at least one context-specific rule or parameter by applying parameters instructed to compensate for the user-specific physical and / or cognitive hearing deterioration model.

[0061] In one embodiment, the renderer device may be configured to perform an upload session for uploading user-specific physical and / or cognitive hearing impairment information, derive a user-specific physical and / or cognitive hearing impairment model, and derive at least one context-specific rule or parameter by applying parameters that compensate for the user-specific physical and / or cognitive hearing impairment model.

[0062] According to one embodiment, at least one context-specific rule or parameter associates different potential characteristics of an audio scene representation with different parameters applied to the audio scene representation.

[0063] In one embodiment, the contextualization unit includes simplification rules or parameters that instruct the rendering unit to reduce the number of audio elements to be rendered, thereby reducing the number of objects to be rendered.

[0064] According to one embodiment, the rendering unit is A first default mode in which the rendering unit operates using at least one first selectable rule or parameter, wherein the first selectable rule or parameter is either a default rule or parameter, or a first context-specific rule or parameter of at least one context-specific rule or parameter, The rendering unit is configured to perform a selection from a second contextualization mode, in which it operates using at least one second selectable rule or parameter instead of a first selectable rule or parameter, the second selectable rule or parameter being a context-specific rule or parameter of at least one context-specific rule or parameter different from the first selectable rule or parameter.

[0065] According to one embodiment, the selection is controlled by manual selection. In one embodiment, the selection is controlled by the presence or absence of a second selectable rule or parameter.

[0066] In one embodiment, the selection is controlled via measurements of biological and / or physiological parameters, thereby selecting a first mode when the measurements of biological and / or physiological parameters match a predetermined standard model representing the physical and / or cognitive hearing of a standard user, and selecting a second contextualized mode when the measurements of biological and / or physiological parameters do not match the predetermined standard model, thereby indicating a deterioration in the user's physical and / or cognitive hearing.

[0067] In one embodiment, selection is controlled through measurements of biological and / or physiological parameters, including EEG measurements.

[0068] According to one embodiment, selection is controlled through measurements of biological and / or physiological parameters, including heart rate measurements.

[0069] According to one embodiment, selection is controlled through measurements of biological and / or physiological parameters, including electrocutaneous responses.

[0070] In one embodiment, the selection is controlled via a feedback signal.

[0071] The renderer device includes a first selectable rule or parameter which is a standard default rule or parameter independent of context-specific data, and a second selectable rule or parameter which is a selectable rule or parameter for at least one contextualization.

[0072] In one embodiment, the first selectable rule or parameter includes a first context-specific rule or parameter of at least one context-specific rule or parameter, and the second selectable rule or parameter includes a second context-specific rule or parameter of at least one context-specific rule or parameter.

[0073] According to one embodiment, the selection is controlled through measurements of biological and / or physiological parameters, including pupillary measurements that measure changes in pupil size.

[0074] According to one embodiment, context-specific data is or includes user-specific personalized data for a particular non-human user unit or non-human user layer.

[0075] According to one embodiment, context-specific data is user-specific and provides context-specific data obtained from feedback signals.

[0076] According to one embodiment, the renderer device may be configured to transmit rendered audio signals to an audio consuming device.

[0077] According to one embodiment, the renderer device may be configured to wirelessly connect to an audio consuming device.

[0078] In one embodiment, the renderer device may be configured to receive feedback signals from an audio consuming device indicating the user's movement, position, and / or orientation, thereby the rendering unit provides rendered audio signals based on the feedback signals.

[0079] According to one embodiment, the rendered audio signal is part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation.

[0080] According to one embodiment, the rendered audio signal is part of an audio scene in the metaverse environment, and the rendered audio signal is defined based on the user's position and / or orientation in the metaverse.

[0081] According to one embodiment, the rendered audio signal is part of the audio scene in the video game environment, and the rendered audio signal is defined based on its position and / or orientation in the video game.

[0082] According to one embodiment, the contextualization unit is configured to define and / or modify context-specific data based on manual input.

[0083] According to one embodiment, the renderer device is The rendering unit may be configured to receive a feedback signal indicating the user's location so that it provides a rendered audio signal based on the feedback signal. The rendering unit is configured to select rendered audio signals from the audio scene representation based on motion, position, and / or orientation detected from the user's location information. The contextualization unit is configured to define at least one context-specific rule or parameter based on context-specific data, independently of the feedback signal.

[0084] According to one embodiment, the renderer device may be configured to receive at least one context-specific rule or parameter at a refresh rate lower than that of the feedback signal.

[0085] According to one embodiment, the renderer device may be further configured to provide a video consuming device with visual metadata specific to the audio scene representation.

[0086] In one embodiment, the contextualization unit is configured to provide at least one visualization command, instructing the rendering unit to provide visual metadata to a video consuming device.

[0087] According to one embodiment, the renderer device may be configured to be installed in a vehicle, and the renderer device is configured to receive a first contextualized feedback signal that provides location measurements relating to the vehicle and a second feedback signal that provides user-provided location measurements relating to the user. The contextualization unit is configured to receive a first contextualization feedback signal as context-specific data, and to derive context-specific rules or parameters based on the first contextualization feedback signal, thereby, Render the audio scene representation using context-specific rules or parameters, and / or render the audio scene representation based on a second feedback signal.

[0088] According to one embodiment, the renderer device is The system can be configured to select between rendering the audio scene representation using context-specific rules or parameters, and rendering the audio scene representation based on a second feedback signal.

[0089] According to one embodiment, the renderer device may be configured to render the audio scene in a virtual environment linked to the vehicle when rendering an audio scene representation using context-specific rules or parameters. When rendering an audio scene representation based on a second feedback signal, the system may be configured to render the audio scene in relation to the user's position.

[0090] In one embodiment, the renderer device may be configured to receive a second feedback signal, which includes a gyroscope and / or accelerometer measurement (value) of position feedback, and may be further configured to receive contextual feedback, which includes a gyroscope and / or accelerometer measurement (value) of contextual feedback, and the renderer device is configured to subtract the gyroscope and / or accelerometer measurement (value) of contextual feedback from the gyroscope and / or accelerometer measurement (value) of position feedback, thereby rendering an audio scene representation using the result of the subtraction.

[0091] In one embodiment, an audio scene representation includes a first audio scene representation which is mixed with a second audio scene representation, the first audio signal representation which is rendered according to vehicle position data, and the second audio signal representation which is rendered independently of vehicle position data, and the renderer device is configured to mix the first audio scene representation with the second audio scene representation to obtain a mixed version of the first and second audio scene representations using a mixing weight defined based on vehicle position data according to at least one context-specific rule or parameter.

[0092] In one embodiment, the first audio signal representation is rendered according to position data, and is rendered according to the relative position of the vehicle and the external position in such a way that the relative position conditions the mixed weights.

[0093] According to one embodiment, the first audio signal representation is rendered according to the distance from the external position of the vehicle, thereby increasing the mixed weight of the second audio signal representation when the distance decreases and decreasing the mixed weight of the second audio signal representation when the distance increases.

[0094] In one embodiment, the second audio signal representation is rendered according to the user's position data in such a way that the mixed weights follow the user's position data.

[0095] In one embodiment, the renderer device may be configured to receive contextualized input as context-specific data and user input input to the rendering unit, wherein the number of times contextualized input is received is lower than the number of times user input is received.

[0096] In one embodiment, the renderer device may be configured to receive contextualized input as context-specific data and user input that is input to the rendering unit, wherein the refresh frequency of context-specific rules is lower than the input frequency of user input.

[0097] According to one embodiment, the audio element includes an audio object.

[0098] According to one embodiment, the audio element includes an audio channel.

[0099] According to one embodiment, the audio element includes an ambisonic signal or an ambisonic coefficient.

[0100] According to one embodiment, the rendering unit is configured to provide the rendered audio signal to the audible unit.

[0101] According to one embodiment, the rendering unit provides the rendered audio signal to the speaker.

[0102] According to one embodiment, the renderer device may be configured to perform a first decompression operation by receiving a compressed version of the audio scene representation and converting the audio scene representation to a version that includes audio elements.

[0103] According to one embodiment, context-specific data is user-specific personal data of a particular human user, or includes such data.

[0104] In one embodiment, the system may be configured to provide a mute function according to at least one context-specific rule or parameter, the mute function being associated with displayed output and / or user position data such that the mute function forces the audio object of the audio signal representation to mute if the audio object is not displayed and / or is not within the display range or viewport.

[0105] According to one embodiment, a system for providing video and audio scenes is provided, the system comprising a video renderer for decoding and rendering video scenes, and a renderer device according to any of the above embodiments.

[0106] According to one embodiment, a system for providing audio content is provided, the system comprising a renderer device according to one embodiment and a background noise sensor, wherein the contextualization unit is configured to derive context-specific rules or parameters from background noise, thereby applying a higher gain to the rendered audio signal when the background noise is higher and a lower gain to the rendered audio signal when the background noise is lower.

[0107] According to one embodiment, the system may be installed in a vehicle.

[0108] According to one aspect, An audio rendering method is provided, which includes processing an audio scene representation that generates a rendered audio signal from an audio scene representation conditioned by at least one context-specific rule or parameter. This method involves generating at least one context-specific rule or parameter based on context-specific data.

[0109] A non-temporary storage unit that stores instructions, and when an instruction is executed by the processor, the processor performs the method described above. [Brief explanation of the drawing]

[0110] [Figure 1] This figure shows an example using conventional technology. [Figure 2] This diagram shows the technology for changing the frequency band. [Figure 3] This figure shows an example of frequency-dependent amplification. [Figure 4a] This figure shows an embodiment of the present invention. [Figure 4b] This figure shows an embodiment of the present invention. [Figure 4c] This figure shows an embodiment of the present invention. [Figure 5a] This is a diagram showing the technology according to the present invention. [Figure 5b] This is a diagram showing the technology according to the present invention. [Figure 6] This is a diagram showing the technology according to the present invention. [Figure 7a] This is a diagram showing the technology according to the present invention. [Figure 7b] This is a diagram showing the technology according to the present invention. [Figure 7c] This is a diagram showing the technology according to the present invention. [Figure 8] This is a diagram showing the technology according to the present invention. [Figure 9a] This figure shows an embodiment of the present invention. [Figure 9b] This figure shows an embodiment of the present invention. [Figure 10] This is a diagram showing the technology according to the present invention. [Figure 11a] This is a diagram showing the technology according to the present invention. [Figure 11b] This is a diagram showing the technology according to the present invention. [Figure 12a] This is a diagram showing the technology according to the present invention. [Figure 12b] This is a diagram showing the technology according to the present invention. [Modes for carrying out the invention]

[0111] Figure 4a shows a first general example of a renderer device 400. The renderer device 400 may include a rendering unit 430. The renderer device may include a contextualization unit 440. The rendering unit can receive an audio scene representation 402 or 412 to be rendered. The rendering unit 430 can process the audio scene representations 402, 412 to generate a rendered audio signal 422 from the audio scene representations 402, 412. The contextualization unit 440 (for example, a personalization unit, or may include a personalization unit) can receive and / or derive context-specific data 461 (for example, context-specific feedback). Based on the context-specific data 461, the contextualization unit 440 can generate and / or derive at least one context-specific rule and / or parameter 441, 442 for the rendering unit 430. The rendered audio signal 422 may be, for example, an uncompressed audio signal. The rendered audio signal 422 may include specific audio channels for each specific speaker (other options are also possible). The rendered audio signal may be transmitted to each speaker, for example, via a wireless (e.g., Bluetooth) connection and / or a wired connection. The audio scene representations 402, 412 may be compressed versions of the rendered audio signal 422, or versions of the rendered audio signal 422 being rendered. The rendered audio signal 422 may be, for example, a bitstream, or another representation with respect to audio elements (e.g., audio objects, downmixed audio channels, ambisonic elements, etc., also referred to below as "objects").

[0112] Figure 4b shows a more detailed view of the renderer device 400 of Figure 4a (however, in some cases, the devices in Figures 4a and 4b may be considered as two different embodiments). The renderer device 400 may be part of a system 400a for providing media scenes, such as video and audio scenes 472, 496 (e.g., virtual reality scenes or augmented reality scenes). System 400a also includes, together with the renderer device 400, a video decoder and renderer 495 for decoding and rendering video scenes (e.g., provided in a compressed version 402aa).

[0113] The renderer device 400 may be independent of system 400a and the video decoder and renderer 495. The renderer device 400 may receive a bitstream 402 as input. The renderer device 400 may output a rendered audio signal 422 (bitstream 402 may be part of a media bitstream, including video bitstream 402aa if present). The renderer device 400 may comprise a rendering unit 430. The rendering unit 430 may include a core decoder 410 to which bitstream 402 may be input. The core decoder 410 may output audio elements and optionally metadata (e.g., indicating at least one of the position, orientation, directivity, source, and width of at least one object to be rendered). The audio elements may be, for example, channels (e.g., downmix channels, or more generally, downmix representations of the represented audio scene). In addition, or alternatively, the audio elements may be objects (e.g., considering the position, directivity, etc., of the audio source). In addition, or alternatively, the audio element may be, or include, an ambisonic element (e.g., a compressed ambisonic representation of the audio scene to be rendered). In particular, the audio element may include, for example, parameters (e.g., in the form of metadata) which, when applied to the audio element (e.g., objects, ambisonic elements, and / or channels), provide a rendered signal 422 of the encoded audio scene representation 402, 412. In some examples, the audio element 412 may have a lower compression ratio than the bitstream 402. However, the audio element 412 does not have to be in a form required for rendering to a speaker.

[0114] The renderer device 400 may include a renderer 420 (for example, part of a rendering unit 430). The renderer 420 may receive audio elements 412 (channels, objects, ambisonics) as input and can generate a rendered audio signal 422. The rendered audio signal 422 may be provided, for example, to a speaker (and / or headphones 470) in a wired and / or wireless format.

[0115] The rendered audio signal 422 may, for example, be in the time domain. The bitstream 422 may be in a compressed format such as the frequency domain (e.g., modified discrete cosine transform MDCT, modified discrete sine transform MDST, etc.) or the ambisonic domain. The audio elements of the audio scene representations 402, 412 may, according to certain examples, be in the time domain or the frequency or ambisonic domain, and may be converted (e.g., to the time domain) by the renderer 420 (or, in the examples, more generally, the rendering unit 430). Thus, the core decoder 410 and / or the renderer 420 (and more generally, the rendering unit 430) can, in some examples, perform the conversion from a compressed audio format (e.g., the frequency domain) to a time-domain audio format.

[0116] The core decoder 410 and / or renderer 420 (and more generally, rendering unit 430) can perform audibility (e.g., binauralization) in some examples, but in other examples, audibility (e.g., binauralization) may be performed by an external unit outside the renderer device 400 or downstream of the rendering unit 430.

[0117] The rendered signal 422 may be supplied, for example, to a speaker and / or headphones 470. Block 470 may, additionally or alternatively, be an audible device. The sound 472 produced by the speaker and / or headphones 470 may be provided to a human user 449 or another receiving entity. Here, it is generally assumed that the user 449 is human, but it may be replaced by a living organism (e.g., an animal) or a receiving entity (e.g., an automated system, e.g., a robotic system, e.g., a non-human or non-human layer).

[0118] The contextualization unit 440 (which may also be a personalization unit) can receive context-specific data 461, for example, from a user 449 (or from an environment such as a vehicle, as will be shown later). The contextualization unit 440 can provide context-specific rules or parameters 442 and / or 441 to the renderer 420 (or more generally to the rendering unit 430). In particular, it is shown here that context-specific rules and / or parameters 441 may be provided to the core decoder 410. In particular, it is shown that rendering parameters 442 may be provided to the renderer 420 (or more generally to the rendering unit 430). In any case, the parameters or rules 441 and 442 are treated almost uniformly here, without much distinction. Basically, the distinction between units 410 and 420 (and the relevant distinction between parameters or rules 441 and 442) should be understood here as merely an example and not limiting.

[0119] The contextualization unit 440 may include an accessibility interface 460. The accessibility interface 460 can provide context-specific data to the data preprocessor 450 (for example, with respect to a contextualization profile or a context-specific profile 462).

[0120] The data preprocessor 450 can provide rules and / or parameters 441, 442. Contextualization rules and / or parameters 441, 442 consider a specific context (e.g., personalization) and thereby condition the rendering of the audio scene representation 402, 412, resulting in the production of a rendered audio signal 422 by continuing to consider the context (e.g., personalization). Thus, it will be understood that the rendering conforms to the contextualization rules and / or parameters 441, 442. Many possibilities for embodying contextualization, as well as methods for deriving the rules and / or parameters 441, 442, will be shown. It should be noted that context-specific rules and / or parameters 441, 442 can be based, for example, on input 461 from a user 449 (or other receiving entity), or on preferences or other data (e.g., received from a storage device). As can be seen from input 461 (or more generally context-specific data), the accessibility interface 460 can extract (or more generally derive) context-specific data (e.g., contextualized profile 462), such as personalized data (e.g., personalization profile). In some examples, the contextualized profile 462 is independent of a particular type of core decoder 410 and / or renderer 420 (or more generally rendering unit 430). It is the data preprocessor 450 that receives the contextualized profile 462 (or more generally context-specific data 462), re-transforms that information into context-specific parameters 441, 442, and provides them to the rendering unit 430 (e.g., the accessibility interface 460).

[0121] The renderer device 400 may include or be connected to a feedback unit 490 (for example, a unit that includes at least one of a head tracker, an eye tracker, and a position sensor that provides positional information such as gestures, position, angle, direction, and motion; in some examples, the feedback unit 490 may include an accelerometer and / or a gyroscope; in some examples, the feedback unit may include a visual sensor such as an image acquisition sensor, a video acquisition sensor, or an audio acquisition sensor) in order to provide feedback information 492 from feedback 491 from a user 449 (or other receiving entity). In some examples, input 461 and feedback input 491 may be the same (thus line 461 is replaced by line 461b), but in other examples they may be different and perform different tasks, and feedback 491 may allow defining the audio signal to be rendered at each point in time according to user feedback (or the entity's response to rendered audio provided to the entity) (e.g., from the position of the head, from the distance to a virtual object, from the currently displayed viewport, etc.), but without changing the personalized data (or, more generally, context-specific data) and keeping it the same, and without changing the personalization rules and / or parameters 441, 442 (or context-specific data rules or parameters), whereas input (or, more generally, context-specific data) 461 may be intended as personalized data (or, more generally, context-specific data) that changes the rendering of the scene. In some cases, personalized data (or more generally, context-specific data) 461 may be acquired during a configuration session, thereby deriving and storing a general context-specific profile 462 (e.g., by a contextualization unit, e.g., by an accessibility interface 460). Thus, the contextualization profile 462 may be used whenever a context exists.Contextualization may also be personalization, and input 461 (context-specific data) may be acquired during a configuration session to derive a personalization profile 462, thereby, whenever the renderer device 400 is used for that particular human user (or another type of user) 449, context-specific rules and / or parameters(s) targeting the personalization profile or individual-specific profile (or more generally, contextualization profile or context-specific profile) 462 are used (for example, a specific rule 442 for defining gain according to a particular location may be defined based on profile 462). In contrast, feedback 491, 492 does not necessarily involve the definition of either context-specific (e.g., personalization) profile, but simply provides real-time information about, for example, the location of user 449, so that a particular rendered audio signal 422, 472 is modulated according to user location feedback 491, 492, for example, by modulating a specific gain according to the user's location data (but the rules for defining gain remain unchanged). The user position feedback 491, 492 may be understood as feedback within the scene, and the input 461 (context-specific data) may be understood as modifying the rendering of the represented audio scene. In the example, feedback 491 may be provided as a feedback signal, while context-specific data 461 may be considered a feedforward signal, which, once acquired, is retained across multiple sessions of rendering. In the example, the feedback signals 491, 492 may change multiple times (e.g., multiple times per second), while context-specific data 461 may remain unchanged. Figure 4b does not show the unit providing context-specific data 461, which may be provided by the same feedback unit 490, or, in other examples, by a different unit (e.g., operated by a clinician).Figure 4b shows arrow 461b indicating that the feedback unit 490 can optionally provide context-specific data 461.

[0122] Generally speaking, both input 461 (contextualized input or context-specific input, context-specific data) and feedback 491 may be obtained from at least one sensor, including one of a unit that includes at least one of a head tracker, eye tracker, or position sensor that provides positional information such as gestures, position, angle, orientation, and motion. In some examples, the at least one sensor may include an accelerometer and / or gyroscope, while in some examples, the feedback unit may include a visual sensor, such as an image acquisition sensor, a video acquisition sensor, or an audio acquisition sensor. The at least one sensor may be, for example, in an immersive device (in which case the at least one sensor may include, for example, at least one gyroscope and / or at least one accelerometer) and / or in at least one non-immersive unit (e.g., an audio, image, and / or video acquisition unit). It should be noted that input 491 and feedback 461 are not necessarily acquired by the same unit and / or the same type of sensor. In some cases (e.g., when input 491 is acquired by a clinician), input 491 (context-specific data) may be acquired by at least one first sensor, and feedback 461 may be acquired by a second sensor (which may be the same or a different type as the first sensor). In some cases (e.g., see below for vehicle 900 in Figures 9a and 9b), input 491 and feedback 461 may be acquired simultaneously by different units (4619 and 470) (but in some cases they may be the same unit), while in some other cases (e.g., when a clinician acquires contextualized input 461), feedback 461 is acquired after input 491 (which may be acquired in the configuration session described above) is acquired (e.g., during the operation session).

[0123] It should be noted that blocks 470 and 490 may, for example, be part of a media content consuming device 475 that may be applied to a user. The media content consuming device 475 (e.g., an immersive device) may be a device for consuming virtual reality content, augmented reality content, etc. In some cases, context-specific data 461 may be obtained by the same media content consuming device 475. Alternatively, the context-specific data 461 may be obtained by a device different from the media content consuming device 475, and possibly from a device different from the renderer device 400. It should be noted that the media content consuming device 475 is just an example, and in some cases, inputs 461 and / or 491 may be obtained from other different units (e.g., audio, image, and / or video acquisition units, etc.).

[0124] The system 400a for providing video and audio scenes may also include a video renderer 495 that can provide the user with a rendered scene 496 obtained from a video bitstream 402a. In some examples, the video bitstream and audio bitstream 402 may be obtained from the same source. The video renderer 495 may be conditioned by visualization metadata 482. The visualization metadata 482 may be received from a first-person view metadata output unit 480. The first-person view metadata output unit 480 may be input via an input 424 from a contextualization unit 440 (e.g., from a data preprocessor 450). A visualization command may be provided (e.g., by the contextualization unit) that instructs the rendering unit to provide the visualization metadata 482 to a video consuming device (e.g., the video renderer 495).

[0125] Generally speaking, it will be understood that the rendered audio signal may be rendered according to certain context-specific (e.g., personalization) rules and / or parameters. Figures 6, 7a, and 7b show examples of context-specific (e.g., personalization) rules or parameters. Figure 6 shows user 449 (e.g., virtually) placed in environment 150 (e.g., a virtual environment). Environment 150 contains two objects, for example, virtual objects (sound sources S1, 152 and sound sources S2, 152) placed in different positions. Each position corresponds to, for example, the orientation of user 449 in virtual environment 150. 0° corresponds to orientation 803. 90° corresponds to orientation 802. 180° corresponds to orientation 801. The two objects (sound sources S1, 152 and S2, 152) are placed close to positions 801 and 803, respectively. The gain profile changes along the different orientations that user 449 can take. The gain profile for providing gain to the rendered audio signal 422(472) may be conditioned by context-specific rules and / or parameters 411, 442. For example, according to the gain profile, when the user looks towards position 803, the user experiences a higher gain from sensor S2 and a lower gain from sound source S1. On the other hand, when user 449 is oriented towards orientation 801, the user experiences a higher gain from sound source S1 and a lower gain from sound source S2. Apart from this general rule (orientation-based attenuation rule), the gain profile may have different attenuation rules that can be defined by context-specific rules and / or parameters 442. For example, according to a particular personalization (contextualization), the gain profile may be modified based on a particular user 449 (e.g., based on a particular personalization or contextualization profile 462). For example, a user 449 with good hearing may have a gain profile that allows for different hearing than a user 449 with hearing impairment.For example, certain audio sources may be deactivated according to certain personalization profiles 462 and personalization rules 441, 442 resulting from those personalization profiles. For instance, if sound source S2 is less relevant than sound source S1, and personalization profile 462 indicates that user 449 has a hearing impairment, the less relevant sound source S2 may be deactivated. In some other cases, for example, if personalization profile 462 indicates that user 449 requires replacement of certain auditory bands, the gain profile may be modified to provide auditory bands that are actually properly acquired by the impaired user 449. However, more generally, the gain profile (or more generally, the orientation-change profile) changes according to a particular personalization 442 (or even more generally, according to a particular contextualization).

[0126] Another example is provided by Figures 7a and 7b, where user 449 is in the first position (x' u , y' u This indicates that the user is moving within the virtual space 150 from a position (with a virtual distance d'1 from sound source 1, 151-1, and a distance d'2 from sound source 2, 152-2). In Figure 7b, the user is at position (x' u , y' u ) from position (x'' u , y'' u The audio source is moved to a location where it has a distance d''1 from sound source 1, 152-1 and a distance d''2 from sound source 2, 152-2. Even in this case, a decay law (distance-based decay law) can be defined, and in particular, it can be conditioned by personalization rules and / or parameters 442 (which can then be derived from personalization profile 442). It is shown that different rules may be applied to define decay depending on the distance from the virtual audio source. Even in this case, for example, in the case of a disabled user, a specific personalized rule may be selected, and for example, some secondary (less relevant) audio sources may be deactivated.

[0127] Next, we will present examples that can be based on the examples in Figures 6 to 7b, or examples that can be based on different examples.

[0128] Figure 4c illustrates how context-specific rules and / or parameters 441, 442 may be instantiated by modifying the parameters (or more generally, metadata) of the bitstream 402 (or more generally, the rendered audio scene representation 402 or 412). Figure 4c shows that the rendered audio scene representation 402 or 412 may include elements (e.g., channels) 402a and / or 412a (which may be part of the bitstream 402 or audio element 412) as well as metadata 402b and / or 412b. The metadata 402b or 412b may include, for example, parameters used in normal compression operations (e.g., each LPC, linear predictive coding, parameters, whitening parameters, etc.) and / or, for example, a mixture matrix. A rendering unit 430 (e.g., 410 and / or 420) may include a spatial audio processing and synthesis unit 425, which can use modified metadata 402c or 412c modified from metadata 402b, 412b obtained from audio scene representations 402, 412. A metadata modification unit 427 may exist that modifies metadata 402b or 412b and provides modified metadata 402c or 412c based on context-specific rules or parameters 441 or 442 (indicated here as 441b and 442b, respectively). For example, if metadata 402b, 412b includes compression parameters (e.g., LPC parameters, whitening parameters, etc.), then context-specific rules and / or parameters 441b or 442b may modify those parameters (e.g., according to specific rules and / or according to at least one weight defined, for example, by the contextualization unit 440, for example, by the data preprocessor 450).If metadata 402b, 412b is defined, for example, with respect to a matrix (e.g., a confusion matrix and / or a covariance matrix and / or correlation matrix), the metadata modification unit 427 may modify the confusion matrix according to context-specific parameters or rules 441b or 442b (e.g., for a user 449 with hearing impairment, some channels may be deactivated or some channels may be attenuated).

[0129] The rendering unit 430 can combine audio elements 412 based on metadata (e.g., included in or modified from 412b and / or 412) and context-specific rules or parameters.

[0130] In the example, audio scene representations 402, 412 may include multiple metadata sets in their metadata. The rendering unit 430 may emit one of the metadata sets based on at least one context-specific rule or parameter. For example, the rendered audio signal 422 (472) may be conditioned by the remaining metadata sets.

[0131] In some cases, a default rule (a first rule) may exist, which is either standard or, more generally, not context-conditional, or context-conditional. An example is shown in Figure 5a. Switch 440a is shown to switch between a first default rule and a second context-specific rule provided by the contextualization unit 440. Here, switch 440a may be controlled by either feedback (e.g., from 492 such as 461b, or from unit 440) or another type of input. Thus, the context-specific rules 441, 442 applied to the rendering unit 430 may be applied differently depending on the different inputs.

[0132] Figure 8 shows an example of use in clinical applications (e.g., for users with hearing impairments). Here, a contextualized profile (e.g., a personalized profile) 462 may be provided. The contextualized profile may be provided by clinical staff after evaluating feedback or other inputs from the user 449 (e.g., 461, 491, 492, 461b). The contextualized profile (e.g., a personalized profile) 462 may be independent of a particular renderer device 400. In this case, the contextualized profile 462 may be incorporated into a contextualization unit 440 (e.g., a personalized unit), thereby allowing the contextualization unit 440 to derive context-specific rules or parameters 441, 442 (e.g., personalization-specific rules or parameters). For example, a degradation model 445' can be derived by a degradation model definer 445. The degradation model 445' may represent the user's cochlear degradation (or, more generally, physical and / or cognitive hearing degradation) or other degradation. The degradation model 445' can be provided to a "context-specific rule or parameter generator" 446. The context-specific rule or parameter generator 446 can define context-specific rules and / or parameters 441, 442 in such a way that the degradation suffered by user 449 is compensated for. For example, if user 449 has hearing loss in one ear, information about the ear with hearing loss may be part of the degradation model 445', so that the context-specific rule and / or parameter generator 446 can modify the speaker gain applied to the specific affected ear, while the other speakers may have different gains (for example, the gain for the ear with poor hearing may be greater than the gain for the ear with normal hearing). For example, the same may apply if one ear does not perceive certain frequency bands, in which case certain frequency bands may be modified for only one ear.

[0133] The example in Figure 8 shows that, for example, the contextualized (e.g., personalized) profile 462 can be replaced by another technique that also incorporates the degradation model 445', which is provided directly to the contextualization unit 440, thereby incorporating the context-specific rule or parameter generator 446 into the degradation model. In Figure 8, the degradation model definer 445 corresponds to the accessibility interface 460, and the context-specific rule or parameter generator 446 corresponds to the data preprocessor 450. However, this correspondence does not need to be maintained, and in some examples, blocks 450 and 460 are completely replaced by blocks 445 and / or 446.

[0134] In some examples, the contextualized profile 462 may be defined directly by the renderer device 400 itself by analyzing the user's behavioral response (or other types of feedback or other inputs 461, 492, 461b) to several pilot signals 422, 472 provided by the renderer device 400, for example. By analyzing the behavior or feedback from the user 449, the renderer 400 (in particular the degradation model definer 445) may generate its own degradation model 445', and / or the context-specific rule or parameter generator 446 may define context-specific (e.g., personalization-specific) rules or parameters 441, 442.

[0135] As shown in the figures, referring to Figures 7a and 7b, the contextualization unit 440 can define at least one context-specific rule(s) and / or parameters(s) 441, 442, including at least one context-specific position rule(s) and / or parameters, based on a distance threshold (e.g., between user 449 and object 152-1 or 152-2 to be rendered). The rendering unit 430 may be configured to compare distance (e.g., d'1 and / or d''1 and / or d'2 and / or d''2) to the distance threshold. For example, rules 441, 442 may include refraining from rendering an object if the distance or position or orientation exceeds the distance threshold, and rendering an object if the distance or position or orientation falls below the distance threshold or position threshold or orientation threshold. For example, in Figure 7a, object 152-2 may not be rendered because distance d'2 is greater than a given distance threshold, while in Figure 7b, d'2 may not be rendered because d''2 is greater than a given distance threshold. This may be defined, for example, by at least one specific rule and / or parameter 441, 442, for example, for a user 449 whose contextualized profile 462 indicates hearing impairment, at least one rule and / or parameter 441, 442 may stipulate that at least one object 152-2 to be rendered (e.g., a secondary and less relevant object) will not be rendered beyond a predetermined distance threshold. Essentially, the relevance of each object is associated with the object's priority, and for example, if user 449 is hearing impaired, the priority may be compared to a priority threshold (see Figure 5a, e.g., through the comparison in 440a). Therefore, for a hearing-impaired user 449, objects with a priority (relevance) lower than the priority threshold are excluded from rendering, resulting in a smaller number of objects being rendered (e.g., within the audio element 412).

[0136] As shown in Figure 6, the contextualization unit 440 can define at least one context-specific rule and / or parameter 441, 442, which includes at least one context-specific position rule and / or parameter based on a specific orientation (e.g., the user's orientation), for example by comparing orientation thresholds with orientation. The rendering unit 430 may be configured to compare angular positions (e.g., 801, 802, 803, etc.) with angular (orientation) thresholds. For example, rules 441, 442 may include refraining from rendering an object if the angular position (orientation) exceeds a predetermined angular (orientation) threshold, and rendering the object if the distance, position, or orientation falls below a predetermined angular (orientation) threshold. For example, in Figure 7a, object s1 may not be rendered when user 449 is facing position 803, but may be rendered when user 449 turns their head toward position 801. This may be defined by at least one rule and / or parameter 441, 442.

[0137] More generally, location information (e.g., gestures) relating to the user may be taken into consideration, and one or more location (e.g., distance and / or angle) measurements may be compared to one or more location (e.g., distance and / or angle) thresholds, so that a particular object may be rendered or not rendered based on the result of the comparison with one or more location (e.g., distance and / or angle) thresholds.

[0138] Instead of comparing position measurements (values) (distance, angle, gesture measurement, etc.) to a threshold(s), additionally or alternatively, at least one rule and / or parameter 441, 442 may suggest comparing gain (e.g., gain of at least one channel, and / or gain of at least one object, and / or gain of at least one ambisonic component) to a gain threshold, thereby making it possible to refrain from rendering certain elements (e.g., objects, channels, or ambisonic elements) based on the results of comparison with one or more position (e.g., distance and / or angle) thresholds. This can be achieved, for example, based on a degradation model 445', thereby allowing a hearing-impaired user to have a simplified rendering.

[0139] More generally, the relevance of each element 412 (e.g., channel, ambisonic component, or object) may be associated with the priority (relevance) of the audio element, which may be compared to a priority threshold, for example, if the user 449 is hearing impaired (see Figure 5a, e.g., through the comparison in 440a). Thus, for example, in the case of a hearing-impaired user 449, audio elements with a priority (relevance) lower than the priority threshold are excluded from rendering, thereby reducing the number of elemental objects (e.g., within audio element 412). This enables simplified rendering for impaired users. The decision to initiate simplified rendering may be based, for example, on the recognition of a specific user 449, or manual selection, or pre-selection (e.g., based on pre-configuration), for example, on feedback indicating the user 449's physical and / or cognitive hearing impairment (see also below). In some examples, priority (e.g., per element such as object, ambisonic component, and / or channel) may be read as metadata from, for example, bitstream 402 (or more generally, within audio scene representations 402, 412).

[0140] In addition, or alternatively, the contextualization unit 440 may define at least one context-specific rule and / or parameter 441, 442, which includes at least one position rule and / or parameter that presents distance-dependent attenuation (as in Figures 7a and 7b), orientation-dependent attenuation (as in Figure 6), gesture-dependent attenuation, or more generally, position-dependent attenuation of the gain of the element being rendered (or another characteristic of the rendered audio signals 422, 472). At least one position rule and / or parameter 441, 442 may include a context-specific attenuation parameter that applies to context-specific distance-dependent attenuation or position-dependent attenuation or orientation-dependent attenuation. Position-dependent attenuation may be defined as distance-dependent attenuation that is inversely proportional to the distance (e.g., d'1, d''1, d'2, d''2 in Figures 7a and 7b) of the object being rendered (e.g., source 152-1 or 152-2 in Figures 7a and 7b), increasing according to the context-specific attenuation parameter. More specifically, position-dependent attenuation is defined as a context-specific attenuation parameter (hereinafter, a dist It can be defined as distance-dependent decay inversely proportional to the distance of the rendered object (e.g., source 152-1 or 152-2 in Figures 7a and 7b), increased by an exponential amount defined by (also shown as). To enable faster distance-dependent decay, the formula is This can be expanded to TIFF2026515680000001.tif1319 (where r is the distance from the audio source, e.g., one of d'1, d''1, d'2, d''2 in Figures 7a and 7b). In the case of TIFF2026515680000002.tif719, the value The larger the value of TIFF2026515680000003.tif710, the greater the gain attenuation with respect to a given distance. In the case of TIFF2026515680000004.tif719, the value The smaller TIFF2026515680000005.tif710 is, the less pronounced the gain attenuation with respect to a given distance. This example is shown in Figure 5b, in particular as an example of Figure 5a. Figure 7c is The example TIFF2026515680000006.tif725 is shown. TIFF2026515680000007.tif722 or TIFF2026515680000008.tif722 or It follows TIFF2026515680000009.tif722, or specific context-specific data 461 or 462. Note that the default values ​​(for example, for default rules, see below for the first default mode) are: TIFF2026515680000010.tif722 is also acceptable.

[0141] At least one rule and / or parameter for providing context-specific information may define at least one channel-specific gain weight to be applied to the corresponding element (e.g., channel, object, ambisonic element), thereby allowing the rendering unit 430 to apply the channel-specific gain weight to the corresponding element of the rendered audio signal 422.

[0142] Channel-specific weights can include multiple object-channel-specific gains (e.g., channel-specific gains), each of which is specific to a particular frequency band. Therefore, the rendering unit 430 applies the first channel-specific gain weights to the first frequency band and the second channel-specific gain weights to the second frequency band according to context-specific rules and / or parameters. This can, for example, follow a degradation model 445', thereby attenuating some frequency bands (which are difficult for the user 449 to hear) and simplifying the rendering result.

[0143] At least one context-specific rule and / or parameter may include a context-specific reverberation level reduction rule and / or parameter, thereby causing the rendering unit 430 to perform a context-specific reduction of the reverberation level based on at least one context-specific reverberation level reduction rule and / or parameter.

[0144] Context-specific rules and / or parameters 441, 442 may be defined to include at least one context-specific initial reflection level reduction rule and / or parameter, thereby causing the rendering unit 430 to perform a context-specific reduction of the initial reflection level based on at least one context-specific initial reflection level reduction rule and / or parameter.

[0145] At least one context-specific rule and / or parameter 441, 442 may include a context-specific dynamic range control rule and / or parameter, thereby causing the rendering unit 430 to perform dynamic range control based on the dynamic range control rule or parameter.

[0146] At least one context-specific rule and / or parameter 441, 442 may include a context-specific floor attenuation parameter used by the rendering unit 430 to perform floor attenuation according to that context-specific floor attenuation parameter.

[0147] At least one context-specific rule and / or parameter 441, 442 may include a context-specific culling gain parameter used by the rendering unit 430 to correct cylinder reflections.

[0148] At least one context-specific rule and / or parameter 441, 442 may include a geometric range of the acoustic space, thereby defining, for example, a wider acoustic space or a more restricted acoustic space by at least one context-specific rule and / or parameter 441, 442.

[0149] At least one context-specific rule and / or parameter 441, 442 can associate different potential characteristics of the audio scene representation with different parameters applied to the audio scene representation, thereby rendering the audio signals 422, 472 accordingly.

[0150] At least one context-specific rule and / or parameter 441, 442 may include text-specific rules or parameters for frequency band modification. At least one context-specific rule and / or parameter for frequency band modification may associate an input frequency band (e.g., for audio scene representations 402, 412) with an output frequency band (e.g., for rendered signals 422, 472). As shown in Figure 3, the rendering unit 430 may change at least one frequency band of the audio scene representations 402, 412 to a different frequency band of the audio element according to at least one context-specific rule and / or parameter 441, 442 for frequency band modification. This may be based, for example, on hearing degradation (e.g., as shown in degradation model 445'), so that frequency bands in which the sensitivity of the hearing-impaired user 449 is degraded are moved to frequency bands in which the sensitivity of the hearing-impaired user 449 is not degraded (or is less degraded). Figure 3 shows the audio scene representation 402 or 412 (in particular, channels 402a or 412a of the audio scene representation 402 or 412) in the frequency domain version (e.g., frequency on the horizontal axis, value of each bin on the vertical axis). Each bin can be moved from one frequency to another according to the rules and / or parameters 441, 442. For example, each band after the move may be kept identical (or at least based on the input band). Following the degradation of general hearing impairment, the frequency band is moved from the input band (of representations 402, 412) to the output band (of rendered signals 422, 472) which has lower frequencies than the input band, but this may vary depending on the specific degradation.For example, if the degradation model 445' indicates that the user's sensitivity is degraded in a specific frequency band (degradation band), the contextualization unit 440 can define at least one rule and / or parameters 441, 442 to shift the band in a way that avoids the degradation band or reduces its use (for example, the degradation model in Figure 3 indicates that user 449's sensitivity is reduced between 4000Hz and 8000Hz, but acceptable below 4000Hz, which is why the band is shifted from between 4000Hz and 8000Hz to below 4000Hz). Thus, the user's hearing impairment is compensated.

[0151] In the example in Figure 2, the gain (for each element, such as the channel, object, or ambisonic component) changes depending on the frequency and pressure level.

[0152] Generally speaking, in the example, the contextualization unit 440 can access user-specific physical and / or cognitive hearing deterioration information (e.g., deterioration model 445') that provides information about user-specific physical and / or cognitive hearing deterioration. Thus, the contextualization unit 440 can define at least one context-specific rule and / or parameter as a user-specific rule and / or parameter based on the user-specific physical and / or cognitive hearing deterioration information. An upload session may be provided to upload user-specific physical and / or cognitive hearing deterioration information in order to derive at least one context-specific rule and / or parameter by deriving a user-specific physical and / or cognitive hearing deterioration model and applying parameters that compensate for the user-specific physical and / or cognitive hearing deterioration model.

[0153] In the example, an audio scene representation (e.g., bitstream 402) may be encoded within it (e.g., metadata 402b, 412b) and may contain multiple contextualization settings. These contextualization settings may be provided to a contextualization unit 440, which can define context-specific rules and / or parameters 441, 442 according to the contextualization settings. For example, there may be an instruction (e.g., encoded in a specific field of bitstream 402) to select one pre-stored rule and / or parameter from a group of pre-stored rules and / or parameters, and the context-specific rules and / or parameters 441, 442 are selected according to the contextualization settings. However, it is also possible that a limited number of parameters and / or rules are provided in the metadata 402b, 412b, and the contextualization unit 440 selects a limited number of context-specific parameters and / or rules based on context-specific data 461.

[0154] Here, we will further elaborate on the example illustrated in Figure 5a. The rendering unit 430 is - A first default mode that operates using a first selectable rule and / or parameter 441, 442 (the first selectable rule and / or parameter being either the default rule and / or parameter, or the first context-specific rule and / or parameter), - A selection can be made (for example, via switch 440a) between a second contextualized mode that operates using at least one second selectable rule and / or parameter instead of the first selectable rule and / or parameter (the second selectable rule and / or parameter being a context-specific rule and / or parameter different from the first selectable rule and / or parameter).

[0155] (The example in Figure 5a is an example of the first mode in which the default rule is independent of the context, but this can be changed to an example of the first mode in which the default rule is also a context-specific rule, whereas the second rule is always a context-specific rule.)

[0156] The selection between the first mode and the second mode (for example, via switch 440a) may be performed manually or based on a preset.

[0157] Alternatively (for example, the example in Figure 5a is also the example in Figure 8), the selection between a first (default) mode (having a first rule that is either contextualized or context-independent) and a second (contextualized) mode (e.g., via switch 440a) may also be performed on a context basis. For example, feedback (or other inputs) 461, 461b, or 491 (492) may be considered.

[0158] For example, the selection (e.g., via switch 440a) may be controlled by measurements of biological and / or physiological parameters (e.g., feedback or part of other inputs 461, 461b, or 491), thereby, - A first mode in which the measured values ​​of biological and / or physiological parameters match a predetermined standard model representing the physical and / or cognitive hearing of a standard user, - Select a second contextualization mode if the measured values ​​of biological and / or physiological parameters do not match a given standard model, thereby indicating a deterioration in the user's physical and / or cognitive hearing.

[0159] (In the example in Figure 7c, the first mode is This means using TIFF2026515680000011.tif719, and the second contextualization mode follows specific context-specific rules. TIFF2026515680000012.tif721 or (This may mean using TIFF2026515680000013.tif722).

[0160] Measurements of biological and / or physiological parameters may include pupillary measurements that measure changes in pupil size. Measurements of biological and / or physiological parameters may include EEG measurements. Measurements of biological and / or physiological parameters may include heart rate measurements. Measurements of biological and / or physiological parameters may include electrocutaneous responses.

[0161] Essentially, the physical condition of user 449 may be determined by measurements of biological and / or physiological parameters, thereby determining whether user 449 has a physical hearing or attention impairment, which may result in a simplified rendering of the audio signal; otherwise, if it is determined that there is no physical hearing or attention impairment, a full rendering (or less simplified rendering) may be performed.

[0162] Generally speaking, the contextualization unit 440 may receive contextualization input as context-specific data (461). Furthermore, user input (491, 492) may be input to the rendering unit (430). The frequency of contextualization input may be lower than the frequency of user input (491, 492). Generally speaking, the refresh frequency (e.g., refresh rate) of context-specific rules may be lower than the input frequency (491, 492) of user input (e.g., the refresh period may be longer than the input period). Therefore, context-specific rules generally have more inertia and slower correction than user input.

[0163] Generally speaking, the rendered audio signal (422) may be transmitted to an audio consuming device (475). The renderer device 400 may be configured to wirelessly connect to the audio consuming device (449). The renderer device may receive a feedback signal (491) from the audio consuming device (475) indicating the user's movement, or more generally, their position (e.g., orientation, gestures, etc.), so that the rendering unit (430) provides the rendered audio signal (422) based on the feedback signal (492). The rendered audio signal (422) may be part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation. The rendered audio signal (422) may be part of an audio scene in a metaverse environment, and the rendered audio signal is defined based on the user's position and / or orientation in the metaverse. The rendered audio signal (422) may be part of an audio scene in a video game environment, and the rendered audio signal is defined based on the position and / or orientation in the video game. The contextualization unit (440) can define and / or modify context-specific data based on manual input.

[0164] In some of the above examples, when the goal is to address at least users with hearing impairments, rules 141 and 142 apply. 1) To compensate for specific hearing impairments, and / or 2) It can be defined as simplifying the sound scene (for example, fewer objects, fewer channels, fewer ambisonic elements, etc.).

[0165] In another example, context-specific data 461 may be a measurement of background noise. Here, context-specific rules 441, 442 can be applied to rendered audio signals 422, 472 if the background noise 461 is high, and to rendered audio signals 422, 472 if the background noise 461 is low. In this case, background noise 461 may be obtained via input 461b from a sensor 490, which may be part of an audio consumption device 475, for example.

[0166] It should be noted that at least one rule and / or parameter 441, 442 may vary according to a particular pressure parameter. For example, rule 441 may require different parameters (e.g., gain) for different pressure parameters. In the example in Figure 6 (or the examples in Figures 7a and 7b, respectively), the gain may be modified not only based on, for example, angle (or distance from the audio source, respectively), but also based on a particular pressure parameter, thereby modulating the gain according to the pressure parameter. Thus, at least one rule and / or parameter 441, 442 may be parameterized with respect to a parameter such as pressure level (other parameters may be selected). The pressure parameter may be read, for example, from bitstream 402 (or more generally from audio scene representations 402, 412), for example, as metadata 402b. More generally, at least one rule 441, 442 may depend on multiple values ​​(e.g., pressure, frequency, and at least two of positional data such as position, distance, angle, orientation, motion, velocity, acceleration, gesture, etc.).

[0167] A more in-depth example of Figure 6 is shown in Figures 11a and 11b (alternative versions thereof are shown in Figures 12a and 12b). Here, the region of interest 800 is defined by the aperture angle 802 (e.g., twice the angle) (e.g., according to at least one context-specific rule and / or parameter 441, 442). The region of interest may also be defined by the position of the user 449 (when acquired by a head tracker or eye tracker, which may be sensor 470 in Figure 9a, or by a user-mounted sensor, such as sensor 470 in Figure 9b). According to at least one context-specific rule and / or parameter 441, 442, the region of interest 800 can be configured such that audio objects of audio scene representations 402, 412 outside the region of interest 800 are not rendered, or are rendered with a lower gain (e.g., attenuated) according to gradient attenuation (e.g., angles outside the region of interest 800 but close to it are attenuated less than angles outside the region of interest 800 but further away from it). At least one context-specific rule and / or parameter can define the amplitude of the region of interest (e.g., aperture angle 802). At least one context-specific rule or parameter can define the gain attenuation for audio objects outside the region of interest so that it attenuates according to their particular location. The angle considered may be, for example, the angle between the user's position and the object being rendered. (Figures 12a and 12b also include stopband attenuation at specific decibel values).

[0168] Figures 9a and 9b illustrate an example in which the above example is applied to an environment 900 (e.g., a manned environment), which is described primarily as a vehicle 900 (e.g., a manned vehicle such as a car). The difference between Figures 9a and 9b is that in Figure 9a, the position sensor 490 is external to the user (e.g., a visual or audio acquisition unit), whereas in Figure 9b, the position sensor 490 is coupled to the user 449 (e.g., an immersive device), which may be an accelerometer or gyroscope, for example. In Figure 9b, the position sensor 490 is shown as being separate from the speaker 470, although the speaker 470 may be integrated into a single unit (e.g., an immersive unit). System 900 includes specific examples of System 400, also shown as 409. Here, a rendering unit 430 and a contextualization unit 440 are shown, which may also be the units in Figures 4a and / or 4b, and therefore they may inherit any of the features outlined above and below. The first feedback 492 (the feedback 492 in Figure 4b, or corresponding to feedback of other positions or gestures of user 449 in the vehicle 900) may be user position feedback. Thus, element 490 may be a position sensor (e.g., a sensor that acquires position measurements such as position and / or orientation and / or gesture measurements, in particular as in Figure 9a, as well as a sensor that acquires acceleration, curves, etc., in particular as in Figure 9b). The position sensor 490 can acquire measurements of position data from user 449 as it is acquired, for example, from an image (or other type of feedback) acquired from a visual signal 491 as in Figure 9a, or from acceleration as in Figure 9b. Thus, the first feedback 492 provided from the position sensor 490 (Figure 9a or Figure 9b) to the rendering unit 430 may be able to condition the rendering of the audio scene representation 402 or 412 as described above by applying contextualization rules and / or parameters 441, 442 that are not defined based on the first feedback 492.Here, the contextualization unit 440 can receive context feedback 461 (context-specific data) that is different from the first feedbacks 491, 492. The context feedback 461 can be, for example, vehicle position feedback independent of the position feedback 461 obtained from the user 449, such as vehicle position feedback. The vehicle position feedback 461 can be provided, for example, from a vehicle position sensor 4619 (including, for example, a global positioning system, GPS), and the GPS can be applied to the vehicle 900 to register the position (e.g., geographical position) of the vehicle 900. Thus, the vehicle position sensor 4619 can provide position information (e.g., position, orientation, etc.) 461 of the vehicle 900. The vehicle position feedback 461 can be provided to the contextualization unit 440, for example, with a reduced occurrence rate (reduced refresh rate, or reduced refresh frequency, or reduced refresh period) with respect to the position feedback 492 of the user 449 (e.g., if the position feedback 492 is provided n times per second, the vehicle position feedback 461 is provided m times per second, where m < n, for example, m << n, for example, m < n / 10). The refresh frequency of at least one context-specific rule and / or parameter (441, 442) can be lower than the input frequency of the user input (491, 492). Thus, at least one context-specific rule and / or parameter 441, 442 can change at a lower frequency than the user input. Thus, the vehicle context-specific rules and / or parameters 441, 442 can condition the rendering of the audio scene representations 402, 412 to be rendered towards the user as a rendered audio signal 472 (e.g., provided to the speaker 470). In the example, it is possible to perform a selection of selectively activating or deactivating the contextualization unit 440, whereby, when deactivated, the rendering of the audio scene representations 402, 412 by the rendering unit 430 will not be conditioned by the vehicle position feedback.This selection may, in some examples, be made manually or by other types of selection (e.g., by pre-configuration). If context-specific information 441, 442 is provided to the rendering unit 430, it may be presented that rendering is conditioned by context feedback 461 from the vehicle 900 (e.g., from the vehicle's position and / or orientation), while if the contextualization unit 440 is deactivated, rendering may follow only (or at least primarily) the user's position feedback 492. In the example provided, if the contextualization unit 440 is deactivated, the position of the virtual audio source follows the movement of the user's 449 head, while if the contextualization unit 440 is activated, the position of the virtual audio source follows the position feedback 461 from the vehicle 900. Thus, it is possible to choose between two different behaviors in which a particular feedback is dominant over another feedback (e.g., selectively dominant) (for example, the dominant feedback may be the specific feedback that is rendered or the specific feedback that is fully rendered).

[0169] In the example of Figure 9a or Figure 9b, the audio scene representation (e.g., bitstream 402, e.g., version 412 thereof) may include a first audio scene representation that is mixed with a second audio scene representation (for example, the first audio scene representation may include sounds related to the external environment encountered by the vehicle as it moves, and the second audio signal representation may include music or other media content that can be consumed by user 449). In one example, the first audio signal representation may be rendered according to the vehicle's position data, and the second audio signal representation may be rendered independently of the vehicle's position data. The renderer device 400 (409) can mix the first audio scene representation with the second audio scene representation (for example, in the rendering unit 430, for example, in the renderer 420) and obtain a rendered audio signal (422, 472) as a mixed version of the first and second audio scene representations, using mixing weights defined (for example, by the contextualization unit 440) (for example, in spatial synthesis) according to at least one context-specific rule or parameter (441, 442) based on the position data of the vehicle 900. For example, if a sound related to an external location (for example, a sound that draws the user's attention to the presence of a specific commercial area, such as a rest area or a gas station) is rendered, the rendered sound (encoded in the first audio scene representation) is positioned in the direction of the commercial area's location, giving the user an impression of the commercial area's location. Therefore, the mixing weights used in the mixing are assigned to speakers located in the direction of the commercial area. Therefore, the first audio signal representation may be rendered in such a way that the relative position conditions the mixed weights, according to position data (e.g., context-specific data acquired as input 461) corresponding to the relative position of the vehicle and the external position of the commercial area (e.g., higher gain is assigned to the rendered channel associated with the speaker facing the commercial area).The first audio signal representation may also be rendered according to the relative distance and / or relative orientation of the vehicle from the commercial area (or, more generally, the external location). For example, in the case of distance relative to the second scene representation, the mixed weight of the first audio signal representation is increased. Thus, the first audio signal representation can be rendered according to the distance of the vehicle from the external location, increasing the mixed weight of the first audio signal representation relative to the second audio signal representation when the distance decreases, and / or decreasing the mixed weight of the first audio signal representation relative to the second audio signal representation when the distance increases. Similarly, additionally or alternatively, the angle between the vehicle 900 and the external location (commercial area) can be considered. The second audio signal representation will be rendered according to the user's (449) location data in such a way that the mixed weight follows the user's location data 491.

[0170] In some examples (particularly in the example of Figure 9b), it is possible to have embodiments in which user inputs 491, 492 are “refined” from the motion of the vehicle (e.g., acceleration) acquired as feedback 461 by the vehicle position sensor 4619 (this can occur in particular when the position sensor 490 in Figure 9b is or includes an accelerometer or gyroscope). More specifically, the measurement 492 obtained from feedback 491 may be the result of subtracting the measurement 461 from the vehicle position sensor 4619, thereby having a “refinement” effect. An example is shown in Figure 10, which shows feedback 492 being subtracted from input (context-specific data) 461 in a subtractor 492a, thereby providing the subtracted feedback 492' to the renderer unit 430. The subtracted feedback 492' (with the acceleration and / or motion of the vehicle 900 “refined”) can then be used by the renderer unit 430 to render an audio signal 422. This may be used in particular when both the input (context-specific data) 461 and the feedback 492 are provided, for example, by an accelerometer and / or gyroscope (similar to Figure 9b), and the accelerometer or gyroscope may be used as a sensor 4619 for measuring the movement of the vehicle 900, and another accelerometer or gyroscope may be used to measure the position (e.g., gesture) measurement 492 of the user 449 (this is not appropriately shown in Figure 9a where sensor 470 is shown as an image or sound acquisition sensor, however, since position measurement 491 such as gestures can also be impaired in principle by the acceleration and movement of the vehicle 900, the inventors understand that the technique of Figure 10 can also be applied in the case of Figure 9b, where sensor 470 is a visual and / or audio sensor and the feedback 461 includes at least one of the user's position, angle, orientation, gesture, etc.). Therefore, it should be noted that at least one context-specific rule or parameter 441, 442 may include (or at least define) the subtractor 492a.

[0171] Considering the above, it will be understood that context-specific rules and / or parameters(s)441, 442 (regardless of whether the context refers to a specific environment or a specific user449) can increase the degree of personalization (or contextualization) of the rendered audio scene without changing the authoring.

[0172] In this example, the speaker / headphones 470 may be, for example, an audible device or part thereof. The user 449 may be human or non-human, for example, a digital assistant.

[0173] In the example, despite the fact that the output (e.g., 470) and / or input (e.g., 490) can be physically isolated from the renderer device 400 (e.g., the connections 422, 492 and / or 461 may be wireless), the advantage of reduced latency is achieved because rendering (e.g., at the renderer unit 430) is performed within the same hardware device (e.g., within the same digital board, or even within the same integrated circuit) as contextuation (e.g., at the contextuation unit 440), thereby reducing latency.

[0174] Regardless of whether the connection (e.g., 422, 492, and / or 461) is wireless or not, the rendered audio signal 422 may be simplified, for example, with respect to the audio scene representations 402 and / or 412, so that only the signals to be rendered are processed and nothing else is processed according to at least one rule and / or parameter 441, 442, thus reducing undesirable latency. For example, in the case of audio signal simplification (e.g., rendering of lower-priority audio elements is avoided if their priority falls below a priority threshold), transmission of unused audio elements is avoided. This is in contrast to the example in Figure 1, where the wireless connection between rendering and EQ, dynamic compression, and FL is compromised by all rendered signals. Thus, in this example, the transmission of unnecessary data is reduced, especially when there is wireless transmission of the rendered audio signal 422.

[0175] Discussion Here, we will describe a specific example of a rendering device 400, which is an accessible immersive device.

[0176] In the proposal, the audio renderer 400 may include an accessibility interface 460 that reads, processes, and / or generates the user's hearing impairment profile and / or other user preferences (e.g., preferences to limit the complexity of the scene), or another example of the input 461. Next, a data preprocessor module (unit) 450 maps the data (hearing impairment profile) 462 from the accessibility interface 460 to internal rendering parameters. These parameters are then parsed for the core decoder 410 and / or rendering module (renderer) 420 (or more generally by the rendering unit 430), after which the audio scene is rendered according to the user's needs. In some applications, sensory information 491 such as head posture, pupil size, or gaze data may be provided and included when creating these rendering instructions.

[0177] Some applications can support the visualization of audio scenes, for example, to visually highlight audible objects in a video representation of a VR scene. This is indicated by a processing box called PoV (First-Person View) Description Output. For this purpose, the renderer device 400 can provide a First-Person View Metadata Output Interface 480 that provides metadata 482 such as the location of the currently active source or the transmitted text string of a TTS-generated audio object. Furthermore, the application may supply or derive an optimized and / or simpler representation of the audio scene, either general or suitable for a hearing impairment profile that focuses on relevant scene elements (e.g., speech, sounds close to the listener) and excludes distracting elements (background noise, ambient noise, early reflections, reverb, diffuse sound energy, etc.).

[0178] This example primarily aims to provide the individual use of 3D and immersive audio rendering, such as an MPEG-I renderer, particularly for people with hearing impairments (user 449). The general idea is to enable rendering commands so that the renderer 400 can generate spatial audio scenes 422, 472 in a way that is more enjoyable / understandable (i.e., accessible) for listeners with specific impairments 449.

[0179] The following subsections describe several possible methods for improving accessibility during rendering.

[0180] -EQ (Equalization) The HRTF (Head-Related Transfer Function) used for binauralization can be filtered based on the proposed amplification curve of the auditory profile in offline processing (e.g., from measurement 461, further processed with, e.g., degradation model 445' in Figure 8) (e.g., as one of the context-specific rules and / or parameters 441, 442). This avoids additional latency and runtime complexity. To compensate for different hearing loss in each ear, the EQ may need to be processed individually for each ear (e.g., for each channel of the rendered signal 422).

[0181] -DRC / AGC (Dynamic Range Control / Automatic Gain Control) Based on auditory profile data (e.g., 441, 442), the dynamic range of the output signal can be corrected using the DRC function within the renderer pipeline.

[0182] -FL (Frequency reduction) In the frequency domain of the renderer device 400, one or more mapping rules and / or parameters 441, 442 can determine how the signal energy of a frequency bin affected by (e.g., severe) hearing impairment is allocated to other frequency bins, in which case the spectrum is compressed (e.g., pre-inverse MDCT or MDST) because auditory sensitivity remains.

[0183] Simplification of sound scenes Limit the number of sound elements that are rendered. Render only the essential elements. Example: ISO / IEC 23008-3 MPEG-H Audio. The MPEG-H Decoder(410) may have a decoding parameter dynamic_object_priority that defines the priority of audio objects (e.g., the audio object of audio element 412). If the priority is lower than a certain threshold (e.g., 7, which may be the priority value assigned to a particular object), the audio object(412) may be discarded from rendering and decoding. If an object(412) is discarded (e.g., as a result of a particular context-specific rule, for example, based on a selection in 440a in Figure 5a or Figure 5b), the object(412) with the lowest priority is discarded first according to rules 441, 442.

[0184] Using this function, the contextualization unit 440 can signal the renderer unit 430 (e.g., the core decoder 410) not to decode certain audio elements (412, e.g., objects with lower priority) according to defined priorities.

[0185] Faster gain attenuation of distant sound sources For example, the gain attenuation of a point source is generally calculated as a function of distance, given by g = 1 / r, where r is the distance between the source and the listener, and g is the attenuation gain (see, for example, ISO / IEC 23090-4:202X MPEG-I part 4 Immersive Audio, WD2, Clause 6.6.12.4 -- Distance attenuation due to geometrical spreading).

[0186] To enable faster distance-dependent decay, the formula is: It can be expanded to TIFF2026515680000014.tif721. In the case of TIFF2026515680000015.tif719, the value The larger the TIFF2026515680000016.tif710, the greater the gain attenuation for a given distance. In the case of TIFF2026515680000017.tif719, the value The smaller the TIFF2026515680000018.tif710, the less significant the gain attenuation for a given distance.

[0187] (In particular, this is illustrated by FIG. 5b as a specific case of FIG. 5a) · Limitation of spatial complexity o Acoustic flashlight effect: Mute all sounds outside the current field of view · Alternatively, if supported by the HMD, it can also be made dynamic based on the eye-tracking data o Reduce the contribution from reverberation and early reflections Example: In many auditory virtual environments and artificial reverberations, the gain (or mixing) of early reflections and late reverberations with respect to the direct sound can be adjusted. For this specific example, see, for example, ISO / IEC 23090-4:202X MPEG-I part 4 Immersive Audio, WD2, Clause 6.6.4.3.7 - RI Gain.

[0188] Here, the following gain parameters that may be modified to reduce the contribution of early reflection acoustic energy are defined.

[0189] · Adjustment gain g tuning : Attenuate the sound level of the image source.

[0190] · Floor attenuation g floor : Further reduce the sound level of the floor reflection · Culling gain g culling implements a linear fade-out of the image source close to the primary reflection source distance culling value earlySourceCullingDistanceOrder1 or the secondary reflection earlySourceCullingDistanceOrder2 · Cylinder gain g cylinder: Gain value to correct for calculation reflection · Rendering to mono To our knowledge, there are currently no 3D audio renderers that explicitly feature accessibility as a key feature. Detectability can be achieved through specific user interfaces and associated signal processing behaviors.

[0191] Here, we offer some discussion regarding the examples in Figures 11a and 11b. To enhance clarity and attentional capacity in a virtual sound scene, the audio renderer can attenuate sounds outside the field of view (e.g., behind the user) and / or amplify sounds in front of the user. Of course, this is a dynamic effect as a function of position input data, since user movement affects the user's position and orientation in the scene.

[0192] This amplification and attenuation can be achieved by creating a beam function that can be parameterized from a user interface.

[0193] For example, a classic microphone beam pattern can be created using the following equation. TIFF2026515680000019.tif740TIFF2026515680000020.tif73 are the directions of the audio source that are rendered with respect to a specific field direction (for example, in the examples in Figures 11a and 11b). TIFF2026515680000021.tif73 refers to the beamforming weights.

[0194] a and b are in the range of 0 to 1. For example, the following beam pattern directivity can be achieved.

[0195] In all directions (i.e., no effect): a=1, b=0 TIFF2026515680000022.tif715 Cardioid: a=0.5, b=0.5, TIFF2026515680000023.tif715 Hypercardioid: a=0.3, b=0.7, TIFF2026515680000024.tif715 parameters TIFF2026515680000025.tif713 sharpens the beam pattern.

[0196] An alternative parameterization, such as the one proposed here, independent of classical microphone beam patterns, can be defined, for example, by specifying the aperture angle in the region of interest (i.e., between 0° and 180°) and a parameter defining the gradient of energy attenuation outside the region of interest (e.g., -6dB attenuation for every 10 degrees). Perhaps an additional parameter would define the desired attenuation of sound immediately behind the user (i.e., -16dB at 180°). Using these parameters, omnidirectional beam patterns can be created.

[0197] Alternatively, the direction the user is looking can be estimated using gaze information provided by an eye-tracking device (for example, in Figure 9a), and this direction can be used as the primary direction of interest (instead of the forward direction). The beam pattern is then guided so that the main lobe aligns with this direction.

[0198] In the examples of Figures 11a and 11b, (for example, the equation for beamforming weights) The beam pattern (using TIFF2026515680000026.tif742 or an alternative formula) may be an example of a context-specific rule and may be modified according to specific context-specific data (e.g., personalized data) 461 (or input 461b or feedback 491).

[0199] In the example, according to a first context-specific rule (implied by a first context-specific data 461 or input 461b or feedback 491) (for example, in the default mode for a user without hearing impairment), the beam pattern attenuates direct sound only and not reflections; and according to a second context-specific rule (implied by a second context-specific data 461 or input 461b or feedback 491) (for example, in a contextualized mode for a user with hearing impairment), the beam pattern attenuates not only direct sound but also early reflections (and / or late reflections).

[0200] Alternatively (however, in some cases, as in the examples in Figures 11a and 11b or Figure 6), according to a first context-specific rule (implied by a first context-specific data 461 or input 461b or feedback 491) (e.g., for a non-hearing user, e.g., in default mode), the beam pattern attenuates direct sound and early reflections but not late reflections; whereas, according to a second context-specific rule (implied by a second context-specific data 461 or input 461b or feedback 491) (e.g., for a hearing-impaired user, e.g., in contextualized mode), the beam pattern attenuates not only direct sound and early reflections but also late reflections.

[0201] For example (for instance, in the examples of Figures 11a and 11b or Figure 6), according to a first context-specific rule (implied by a first context-specific data 461 or input 461b or feedback 491) (for example, for a user without hearing impairment, e.g., in default mode), the beam pattern attenuates direct sound and reflected(or reflected) at the same rate; and according to a second context-specific rule (implied by a second context-specific data 461 or input 461b or feedback 491) (for example, for a user with hearing impairment, e.g., in contextualized mode), the beam pattern attenuates reflected(or reflected) at a greater rate than the beam pattern attenuates direct sound.

[0202] In the example, context-specific rules can be a composition of rules. For example, (for example, expression The beam pattern (which has TIFF2026515680000027.tif742) is attenuated with distance (for example). It can be combined with TIFF2026515680000028.tif716). For example, the resulting formula is: This could be TIFF2026515680000029.tif1377. When defining multiple rules (for example, regarding multiple modes), it may be possible to modify both rules that are combined with each other using one of the techniques described above and below.

[0203] Attenuation due to beam patterns can have different effects on different components of the sound field. For example, - The beam pattern can only affect direct sound, but it may or may not attenuate early reflections, and it may or may not attenuate late reflections (i.e., late reverb). - All initial reflections of a particular sound source can be attenuated by the same amount as directional sound attenuation. - Using the attenuation of the direct sound component of an audio element due to the beam pattern and direction of the sound source (e.g., a specific first attenuation rate), all associated early reflections of that audio element in the virtual space can be attenuated (e.g., by the same first attenuation rate). - The beam pattern can attenuate sound differently depending on the distance to the user.

[0204] In some examples, an audio scene representation may include metadata indicating that at least one element (e.g., at least one object) is not subject to contextualization (e.g., cannot be personalized), thereby preventing the rendering unit from performing contextualization (e.g., personalization) (e.g., the rendering unit may be suppressed from applying certain contextualization rules (e.g., personalization rules)). In other examples, this possibility is not anticipated.

[0205] Here, we offer some discussion regarding Figure 7c. For audio objects without range, i.e., point source audio objects, the distance attenuation curve generated by the model can be the classic 1 / r point source distance attenuation curve, where r represents the distance from the source to the listener.

[0206] To improve accessibility (for example, for users with hearing impairments), an adjustment parameter α is introduced into the calculation of distance-based gain attenuation, and the distance r is r α It can be changed to this. The intended effect is to allow the user to increase or decrease the depth of the sound scene (consisting of audio objects) as needed. For example, if α > 1, the depth of the scene expands, and objects further away become harder to hear and disappear. Conversely, if α < 1, the depth of the scene shrinks, and objects further away become easier to hear. The default value, for example, α=1 (like the first 7c), shows the effect of the exponent α on three different values: 0.7, 1.0, and 1.1.

[0207] The distance index α may be provided as part of a context-specific rule.

[0208] We will provide some considerations with reference to Figures 11a, 11b, 12a, and 12b.

[0209] Intended as a feature to improve accessibility, directional focus means attenuating distracting sounds from directions outside the spatial domain of interest. The focus may be radially symmetric with, for example, one "main lobe" region. Its attenuation behavior can be configured using, for example, three parameters provided by the contextualization unit 430. The default direction is towards the user's front view, but it can be reoriented in other directions, for example, to enable control via other services or modalities (e.g., eye trackers, handheld controllers, etc.).

[0210] -Signed audio elements associated with the listener may not be processed by directional focus.

[0211] Final example Depending on specific implementation requirements, the examples of this disclosure may be implemented in hardware or software. Implementation may be carried out using digital storage media, such as floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memory, hard disks, or any other magnetic or optical memory that stores electronically readable control signals that can cooperate with or be controlled by a programmable computer system to perform the respective methods. This is why digital storage media can be computer-readable.

[0212] Accordingly, some examples provided in the pre-configured disclosure include a data carrier containing electronically readable control signals that can work with a programmable computer system to perform any of the methods described herein.

[0213] Generally, the examples of the present disclosure may be implemented as computer program products having program code, which is effective for performing any of the present methods when the computer program product is executed on a computer.

[0214] The program code may also be stored in a machine-readable carrier, for example.

[0215] Other examples include a computer program for performing any of the methods described herein, the computer program being stored in a machine-readable carrier.

[0216] In other words, an example of the method of the present invention is a computer program having program code for performing any of the methods described herein when the computer program is executed on a computer.

[0217] Therefore, a further example of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program for performing any of the methods described herein is recorded.

[0218] Therefore, a further example of the methods of the present invention is a data stream or sequence of signals representing a computer program for performing any of the methods described herein. The data stream or sequence of signals may be configured to be transmitted, for example, over a data communication link, for example, over the Internet.

[0219] Further examples include processing means configured or adapted to perform any of the methods described herein, such as a computer or a programmable logic device.

[0220] Further examples include computers on which computer programs for performing any of the methods described herein are installed.

[0221] Further examples include devices or systems configured to transmit a computer program to a receiver for performing at least one of the methods described herein. The transmission may be, for example, electronic or optical. The receiver may be, for example, a computer, a mobile device, a memory device, or a similar device. The device or system may include, for example, a file server for transmitting the computer program to the receiver.

[0222] In some examples, programmable logic devices (e.g., field-programmable gate arrays, FPGAs) can be used to perform some or all of the functions of the methods described herein. In some examples, a field-programmable gate array can work with a microprocessor to perform any of the methods described herein. In general, the methods are performed by any hardware device in some examples. Such hardware device may be any universally applicable hardware, such as a computer processor (CPU), or it may be method-specific hardware, such as an ASIC.

[0223] While the present invention has been described in relation to several advantageous embodiments, there are modifications, substitutions, and equivalents that fall within the scope of the invention. It should also be noted that there are many alternative methods for implementing the methods and compositions of the present invention. Accordingly, the following appended claims are intended to be interpreted as including all such modifications, substitutions, and equivalents that fall within the true spirit and scope of the invention.

[0224] References [1]BC Moore, “Perceptual consequences of cochlear hearing loss and their implications for the design of hearing aids,” Ear ​​and hearing, vol. 17, no. 2, pp. 133-161, 1996. [2]I. McClenaghan, L. Pardoe, and L. Ward, “The next generation of audio accessibility,” 2022. [3]L. A. Ward, “Improving Broadcast Accessibility for Hard of Hearing Individuals: using object-based audio personalisation and narrative importance.,” 2020, doi: 10.13140 / RG.2.2.31454.46405. [4]J. Paulus, M. Torcoli, C. Uhle, J. Herre, S. Disch, and H. Fuchs, “Source Separation for Enabling Dialogue Enhancement in Object-based Broadcast with MPEG-H,” J. Audio Eng. Soc., vol. 67, no. 7 / 8, pp. 510-521, Aug. 2019, doi: 10.17743 / jaes.2019.0032.

Claims

1. A renderer device (400), A rendering unit (430) configured to process rendering audio scene representations (402, 412) and to receive at least one context-specific rule or parameter (441, 442), wherein the rendering unit (430) is configured to generate a rendered audio signal (422) from the audio scene representations (402, 412) conditioned by the at least one context-specific rule or parameter (441, 442), A contextualization unit (440) configured to receive and / or derive context-specific data (461, 462), and the contextualization unit (440) configured to provide the rendering unit (430) with at least one context-specific rule or parameter (441, 442) based on the context-specific data (461, 462). A renderer device equipped with the following features.

2. The renderer apparatus according to claim 1, wherein the rendering unit (430) is configured to process the audio scene representation (402) as including audio elements, and the rendering unit (430) is configured to generate the rendered audio signal (422) from the audio elements and the at least one context-specific rule or parameter.

3. The renderer apparatus according to any one of claims 1 to 2, wherein the rendering unit (430) is configured to process the audio scene representation (402) including audio elements and metadata, and the rendering unit (430) is configured to generate the rendered audio signal (422) from the audio elements, the metadata and the at least one context-specific rule or parameter.

4. The renderer apparatus according to claim 3, wherein the metadata includes position metadata providing information about at least one of the position, orientation, directivity, source, and width of at least one object to be rendered, the at least one object to be rendered is part of the audio scene representation (402), and the rendering unit (430) is configured to generate the rendered audio signal from the audio elements, the metadata, and at least one context-specific rule or parameter.

5. The renderer apparatus according to any one of claims 3 to 4, wherein the rendering unit is configured to modify the metadata based on at least one context-specific rule or parameter to obtain modified metadata, and the rendering unit is configured to apply spatial audio processing and synthesis to the audio elements based on the modified metadata.

6. The renderer apparatus according to any one of claims 3 to 5, wherein the rendering unit (430) is configured to apply spatial audio processing and synthesis to the audio element (412) based on the metadata and the at least one context-specific rule or parameter (422).

7. The renderer apparatus according to any one of claims 3 to 6, wherein the rendering unit (430) is configured to combine the audio elements (412) based on the metadata and the at least one context-specific rule or parameter (421).

8. The renderer apparatus according to any one of claims 1 to 7, wherein the contextualization unit (440) is configured to define the at least one contextual rule or parameter (441) including a contextual rule or parameter that associates contextual gain weights with distance, position, gesture, and / or orientation, in order to apply contextual gain weights correspondingly to objects to be rendered based on the contextual position rule or parameter.

9. The renderer apparatus according to claim 8, wherein the context-specific position rule or parameter defines the gain weights to be frequency-dependent, so that the rendering unit (430) applies a first context-specific gain weight to a first frequency band and a second context-specific gain weight to a second frequency band according to the position rule or parameter.

10. The renderer apparatus according to any one of claims 1 to 9, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter, including a context-specific position rule or parameter based on a distance threshold, a position threshold, or an orientation threshold, and the rendering unit is configured to compare the distance, position, or orientation of an object to be rendered with the distance threshold, position threshold, or orientation threshold, respectively, and thereby refrain from rendering the object if the distance, position, or orientation exceeds the distance threshold, position threshold, or orientation threshold, and render the object if the distance, position, or orientation falls below the distance threshold, position threshold, or orientation threshold.

11. The renderer apparatus according to any one of claims 1 to 10, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter based on a context-specific gain threshold, and the rendering unit is configured to compare the gain of an object to be rendered with the context-specific gain threshold, thereby refraining from rendering the object if the gain is below the context-specific gain threshold, and rendering the object if the gain is above the context-specific gain threshold.

12. The renderer apparatus according to any one of claims 1 to 11, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter such that it includes a position rule or parameter that describes distance-dependent attenuation, position-dependent attenuation, or orientation-dependent attenuation of the gain of the object to be rendered, the position rule or parameter includes a context-specific attenuation parameter that is applied to the context-specific distance-dependent attenuation, position-dependent attenuation, or orientation-dependent attenuation.

13. The renderer apparatus according to claim 12, wherein the contextualization unit (440) is configured to define the position-dependent attenuation as distance-dependent attenuation inversely proportional to the distance of the rendered object, increased by an exponential amount defined by the context-specific attenuation parameter.

14. The renderer apparatus according to claim 12 or 13, wherein the contextualization unit (440) is configured to define the position-dependent attenuation as a distance-dependent attenuation that is inversely proportional to the distance of the object being rendered, increasing or decreasing according to the context-specific attenuation parameter.

15. The renderer apparatus according to any one of claims 1 to 14, wherein the contextualization unit (440) is configured to define the at least one contextual rule or parameter to provide contextual information relating to the at least one channel-specific gain weight applied to the corresponding audio element of the rendered audio signal (422), thereby enabling the rendering unit to apply the channel-specific gain weight to the corresponding audio element of the rendered audio signal (422).

16. The renderer apparatus according to claim 15, wherein the channel-specific weights include a plurality of channel-specific gains, and each channel-specific gain is specific to each frequency band, so that the rendering unit (430) applies a first channel-specific gain weight to a first frequency band and a second channel-specific gain weight to a second frequency band according to the at least one context-specific rule or parameter.

17. The renderer apparatus according to any one of claims 1 to 16, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter, which includes a context-specific reverberation level reduction rule or parameter, thereby enabling the renderer unit to perform a context-specific reduction of the reverberation level based on the context-specific reverberation level reduction rule or parameter.

18. The renderer apparatus according to any one of claims 1 to 17, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter, which includes a context-specific initial reflection level reduction rule or parameter, so that the rendering unit performs a context-specific reduction of the initial reflection level based on the context-specific initial reflection level reduction rule or parameter.

19. The renderer apparatus according to any one of claims 1 to 18, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter, including a dynamic range control rule or parameter, so that the rendering unit performs dynamic range control based on the dynamic range control rule or parameter.

20. The renderer apparatus according to any one of claims 1 to 19, wherein the contextualization unit (440) is configured to derive context-specific rules or parameters from background noise, thereby applying a higher gain to the rendered audio signal when the background noise is higher and a lower gain to the rendered audio signal when the background noise is lower.

21. The renderer apparatus according to any one of claims 1 to 20, wherein the contextualization unit (440) is configured to define context-specific floor attenuation parameters used by the rendering unit (430) to perform floor attenuation according to the context-specific floor attenuation parameters.

22. The renderer apparatus according to any one of claims 1 to 21, wherein the contextualization unit (440) is configured to define context-specific culling gain parameters used by the rendering unit (430) to perform linear fade-out of objects closer to a first source distance culling value for primary reflections, or to perform linear fade-out of objects closer to a second source distance culling value for secondary reflections.

23. The renderer apparatus according to any one of claims 1 to 22, wherein the contextualization unit (440) is configured to define context-specific culling gain parameters used by the rendering unit (430) to correct cylinder reflections.

24. A renderer device according to any one of claims 1 to 23, configured to process the audio scene representation (402) to obtain a version (412) of the audio scene representation including audio elements, wherein the renderer device (400) is further configured to generate the rendered audio signal (422) from the audio elements and metadata (412).

25. A renderer device according to any one of claims 1 to 24, wherein the renderer device (400) is configured to process the audio scene representation (402) to obtain a version (412) of the audio scene representation including audio elements and metadata, and the renderer device (400) is further configured to generate the rendered audio signal (422) from the audio elements and metadata (412).

26. A renderer device according to claim 24 or 25, configured to process the audio scene representation (402) to obtain a version (412) of the audio scene representation including audio elements in a core decoder block (410), and to process the version (412) of the audio scene representation to generate the rendered audio signal (422) from the audio elements (412) in a rendering block (430).

27. The renderer apparatus according to any one of claims 1 to 26, wherein the at least one context-specific rule or parameter (441, 421) includes a context-specific rule or parameter for frequency band modification, the context-specific rule or parameter associates an input frequency band with an output frequency band, and thereby the rendering unit modifies at least one frequency band of the audio scene representation to a different frequency band of the audio element in accordance with the context-specific rule or parameter for frequency band modification.

28. The renderer device according to claim 27, wherein the context-specific rules or parameters for changing the frequency band reduce the frequency of at least one frequency band.

29. The renderer apparatus according to any one of claims 1 to 28, wherein the at least one context-specific rule or parameter (441, 442) includes a context-specific rule or parameter for frequency-dependent gain amplification that associates an input frequency band with a context-specific frequency band weight, thereby causing the rendering unit to apply the context-specific gain weight correspondingly to the spectral values ​​of at least one bin of at least one frequency band according to the context-specific rule or parameter.

30. The renderer device according to any one of claims 1 to 29, wherein the contextualization unit is configured to define the at least one rule or parameter as a geometric range.

31. A renderer device according to any one of claims 1 to 30, configured to use at least one context-specific rule or parameter (441, 442) including a proposed amplification curve for each channel, wherein the proposed amplification curve follows a contextualized profile based on the context-specific data.

32. The apparatus according to any one of claims 1 to 31, configured to use the at least one context-specific rule or parameter (441, 442) as one that is parameterized by a parameter or measurement, thereby modulating the at least one context-specific rule or parameter (441, 442) according to the parameter or measurement.

33. The renderer apparatus according to any one of claims 1 to 32, wherein the rendering unit is configured to process the audio scene representation in order to derive a version of the audio scene representation as having a plurality of metadata sets in the metadata, and the rendering unit is configured to emit one of the metadata sets based on at least one context-specific rule or parameter.

34. The renderer apparatus according to any one of claims 1 to 33, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter (441, 442) based on a contextualization profile (462) which includes a plurality of context-specific data (461), and the contextualization unit (440) is configured to extract from the contextualization profile (462) the context-specific data and / or rendering configuration related to the audio signal to be rendered, and to derive the at least one context-specific rule or parameter (441, 442) from the related context-specific data (462).

35. The renderer apparatus according to any one of claims 1 to 34, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter based on the parameters of the rendering unit (430) in such a manner that the at least one context-specific rule or parameter conforms the parameters of the rendering unit (430) to the contextualization profile.

36. A rendering unit according to any one of claims 1 to 35, wherein the contextualization unit (440) is configured to derive a degradation model (435') from the contextualization profile (462) and / or the context-specific data (461), the degradation model (435') indicating a specific degradation in the ability of a particular human user (449) or another audio receiving entity to obtain the rendered audio signal (422), and the contextualization unit (440) is further configured to define, based on the degradation model (435'), the at least one context-specific rule or parameter (441, 442) to compensate for the specific degradation.

37. A rendering unit according to any one of claims 1 to 36, configured to perform simplified rendering when the user is determined to have a hearing impairment and / or reduced cognitive and / or physical sensitivity.

38. The renderer device according to any one of claims 1 to 37, wherein the contextualization unit (440) is configured to define the at least one context-specific rule or parameter based on the contextualization settings received in the audio scene representation (402).

39. The renderer device according to any one of claims 1 to 38, wherein the contextualization unit is configured to select at least one context-specific rule or parameter from a plurality of contextualization settings received in the audio scene representation (402), and to select the most appropriate context-specific rule or parameter from the plurality of received contextualization settings based on feedback such as user-specific physical and / or cognitive hearing impairment information from the user, from pre-configuration, or from manual selection.

40. The renderer device according to any one of claims 1 to 37, wherein the contextualization unit (440) is configured to access user-specific physical and / or cognitive hearing deterioration information that provides information about user-specific physical and / or cognitive hearing deterioration, and to define the at least one context-specific rule or parameter as a user-specific rule or parameter based on the user-specific physical and / or cognitive hearing deterioration information.

41. The renderer device according to claim 40, wherein the user-specific physical and / or cognitive auditory deterioration information is or includes information relating to the user's cochlear deterioration.

42. The renderer device according to claim 40 or 41, wherein the contextualization unit (440) is configured to generate the at least one context-specific rule which is a user-specific rule or parameter for compensating for the user-specific physical and / or cognitive hearing impairment.

43. A renderer device according to any one of claims 40 to 42, configured to run a configuration session to acquire user-specific physical and / or cognitive hearing deterioration information through multiple acquisitions, and thus derive a user-specific physical and / or cognitive hearing deterioration model, and to derive the at least one context-specific rule or parameter by applying parameters instructed to compensate for the user-specific physical and / or cognitive hearing deterioration model.

44. A renderer device according to any one of claims 40 to 43, configured to perform an upload session for uploading user-specific physical and / or cognitive hearing deterioration information in order to derive a user-specific physical and / or cognitive hearing deterioration model, and to derive the at least one context-specific rule or parameter by applying parameters that compensate for the user-specific physical and / or cognitive hearing deterioration model.

45. The renderer apparatus according to any one of claims 1 to 44, wherein the at least one context-specific rule or parameter relates different potential characteristics of the audio scene representation to different parameters applied to the audio scene representation.

46. The renderer apparatus according to any one of claims 1 to 45, wherein the contextualization unit (440) includes simplification rules or parameters that command a reduction in the number of audio elements to be rendered, thereby reducing the number of objects to be rendered by the rendering unit.

47. The rendering unit (430) is The rendering unit (430) operates using at least one first selectable rule or parameter, and the first selectable rule or parameter is either a default rule or parameter, or a first context-specific rule or parameter among the at least one context-specific rule or parameter, in a first default mode, The rendering unit (430) operates using at least one second selectable rule or parameter instead of the first selectable rule or parameter, and the second selectable rule or parameter is a context-specific rule or parameter of the at least one context-specific rule or parameter that is different from the first selectable rule or parameter, in a second contextualization mode. A renderer device according to any one of claims 1 to 46, configured to perform a selection from.

48. The renderer apparatus according to claim 47, wherein the selection is controlled by manual selection.

49. The renderer apparatus according to claim 47 or 48, wherein the selection is controlled by the presence or absence of the second selectable rule or parameter.

50. The renderer device according to any one of claims 47 to 49, wherein the selection is controlled via measurements of biological and / or physiological parameters, thereby selecting the first mode when the measurements of the biological and / or physiological parameters match a predetermined standard model representing the physical and / or cognitive hearing of a standard user, and selecting the second contextualized mode when the measurements of the biological and / or physiological parameters do not match the predetermined standard model, thereby indicating a deterioration of the user's physical and / or cognitive hearing.

51. The renderer apparatus according to any one of claims 47 to 50, wherein the selection is controlled via measurements of biological and / or physiological parameters, including EEG measurements.

52. The renderer device according to any one of claims 47 to 51, wherein the selection is controlled via measurements of biological and / or physiological parameters, including heart rate measurements.

53. The renderer apparatus according to any one of claims 47 to 52, wherein the selection is controlled through measurements of biological and / or physiological parameters, including electrocutaneous reactions.

54. The renderer apparatus according to any one of claims 47 to 53, wherein the selection is controlled via a feedback signal (461).

55. The renderer device according to any one of claims 47 to 54, wherein the first selectable rule or parameter includes a standard default rule or parameter independent of the context-specific data, and the second selectable rule or parameter includes the selectable rule or parameter for at least one contextualization.

56. The renderer device according to any one of claims 47 to 55, wherein the first selectable rule or parameter includes a first context-specific rule or parameter from the at least one context-specific rule or parameter, and the second selectable rule or parameter includes a second context-specific rule or parameter from the at least one context-specific rule or parameter.

57. The renderer device according to any one of claims 47 to 56, wherein the selection is controlled via measurements of biological and / or physiological parameters, including pupillary measurements for measuring the change in pupil size.

58. The renderer device according to any one of claims 1 to 57, wherein the context-specific data is or includes user-specific personalized data for a specific non-human user unit or non-human user layer.

59. The renderer device according to any one of claims 1 to 58, wherein the context-specific data is user-specific and provides context-specific data obtained from a feedback signal.

60. The renderer device according to any one of claims 1 to 59, configured to transmit the rendered audio signal (422) to an audio consuming device (475).

61. The renderer device according to claim 60, configured to connect wirelessly to the aforementioned audio consumption device (475).

62. The renderer apparatus according to claim 60 or 61, configured to receive a feedback signal (461) indicating the user's movement, position, and / or orientation from the audio consuming device (475), thereby causing the rendering unit (430) to provide the rendered audio signal (422) based on the feedback signal (492).

63. The renderer device according to any one of claims 1 to 62, wherein the rendered audio signal (422) is part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation.

64. The renderer device according to any one of claims 1 to 63, wherein the rendered audio signal (422) is part of an audio scene in a metaverse environment, and the rendered audio signal is defined based on the user's position and / or orientation in the metaverse.

65. The renderer apparatus according to any one of claims 1 to 64, wherein the rendered audio signal (422) is part of an audio scene in a video game environment, and the rendered audio signal is defined based on its position and / or orientation in the video game.

66. The renderer device according to any one of claims 1 to 65, wherein the contextualization unit (440) is configured to define and / or modify the context-specific data based on manual input.

67. The rendering unit (430) receives a feedback signal (492) indicating the user's location information so that it provides the rendered audio signal (422) based on the feedback signal (492). It is configured in such a way, The rendering unit (420) is configured to select the rendered audio signal (422) from the audio scene representations (402, 412) based on the motion, position, and / or orientation detected from the user's location information (492). The contextualization unit (440) is configured to define the at least one context-specific rule or parameter (441, 442) based on the context-specific data (462), independently of the feedback signal (492). A renderer device according to any one of claims 1 to 66.

68. The renderer device according to claim 67, configured to receive the at least one context-specific rule or parameter at a refresh rate lower than the feedback signal (492).

69. The renderer apparatus according to any one of claims 1 to 68, further configured to provide visual metadata specific to the audio scene representation to a video consuming device (495).

70. The renderer apparatus according to claim 69, wherein the contextualization unit (440) is configured to provide at least one visualization command and instructs the rendering unit (430) to provide the visual metadata (482) to the video consuming device.

71. It is configured to be installed on a vehicle (900) and to receive a first contextualized feedback signal (461, 461b) that provides a position measurement of the vehicle (900) and a second feedback signal (491) that provides a position measurement of the user (449), The contextualization unit (440) is configured to receive the first contextualization feedback signals (461, 461b) as context-specific data, and to derive the context-specific rules or parameters (441, 442) based on the first contextualization feedback signals (461, 461b), thereby, The audio scene representation (402, 412) is rendered using the context-specific rules or parameters (441, 442), and / or Based on the second feedback signal (461), the audio scene representations (402, 412) are rendered. A renderer device according to any one of claims 1 to 70.

72. Rendering the audio scene representation (402, 412) using the context-specific rules or parameters (441, 442), Rendering the audio scene representations (402, 412) based on the second feedback signal (461) and The renderer device according to claim 71, configured to select from among the following.

73. When rendering the audio scene representation (402, 412) using the context-specific rules or parameters, the system is configured to render the audio scene in a virtual environment that interacts with the vehicle. When rendering the audio scene representations (402, 412) based on the second feedback signal (461), the system is configured to render the audio scene in relation to the user's position. The renderer device according to claim 71 or 72.

74. The system is configured to receive the second feedback signal (461) which includes a measurement (value) from a gyroscope and / or accelerometer for position feedback, The contextualized feedback (461, 461b) is further configured to be input as including a gyroscope and / or accelerometer measurement (value) of the contextualized feedback, and / or The system is configured to subtract (492a) the measurement (value) (461) of the contextualized feedback from the measurement (value) (492) of the position feedback from the measurement (value) (492) of the gyroscope and / or accelerometer, thereby using the result (492a) of the subtraction to render the audio scene representation. The renderer device according to claims 71 to 73.

75. The aforementioned audio scene representations (402, 412) include a first audio scene representation that is mixed with a second audio scene representation. The first audio signal representation will be rendered according to the vehicle's position data, and the second audio signal representation will be rendered independently of the vehicle's position data. The renderer device is configured to mix the first audio scene representation with the second audio scene representation in order to obtain the rendered audio signal (422, 472) as a mixed version of the first audio scene representation and the second audio scene representation, using mixed weights defined based on at least the vehicle's position data according to the at least one context-specific rule or parameter (441, 442). A renderer device according to any one of claims 71 to 74.

76. The renderer apparatus according to any one of claims 71 to 75, wherein the audio scene representations (402, 412) include a first audio scene representation that is mixed with a second audio scene representation.

77. The renderer apparatus according to claim 75, wherein the first audio signal representation is rendered according to position data in such a manner that the relative position conditions the mixed weight, and is rendered according to the relative position and external position of the vehicle.

78. The renderer apparatus according to claim 75 or 77, wherein the first audio signal representation is rendered according to the distance of the vehicle from the external position, thereby increasing the mixed weight of the first audio signal representation with respect to the second audio signal representation when the distance decreases, and / or decreasing the mixed weight of the first audio signal representation with respect to the second audio signal representation when the distance increases.

79. The renderer apparatus according to any one of claims 75 to 78, wherein the second audio signal representation is rendered in accordance with the user's (449) position data in such a manner that the mixed weights conform to the user's position data.

80. A renderer device according to any one of claims 1 to 79, configured to receive contextualized input as context-specific data (461) and user inputs (491, 492) input to the rendering unit (430), wherein the number of occurrences of receiving the contextualized input is lower than the number of occurrences of receiving the user inputs (491, 492).

81. A renderer device according to any one of claims 1 to 80, configured to receive contextualized input as context-specific data (461) and user inputs (491, 492) input to the rendering unit (430), wherein the refresh frequency of the context-specific rules is lower than the input frequency of the user inputs (491, 492).

82. The renderer apparatus according to any one of claims 1 to 81, wherein the audio element includes an audio object.

83. The renderer device according to any one of claims 1 to 82, wherein the audio element includes an audio channel.

84. The renderer apparatus according to any one of claims 1 to 83, wherein the audio element includes an ambisonic signal or an ambisonic coefficient.

85. The renderer apparatus according to any one of claims 1 to 84, wherein the rendering unit is configured to provide the rendered audio signal to the audible unit.

86. The renderer apparatus according to any one of claims 1 to 85, wherein the rendering unit provides the rendered audio signal to a speaker.

87. A renderer device according to any one of claims 1 to 86, configured to receive a compressed version of the audio scene representation (402) and perform a first decompression operation by converting the audio scene representation (402) to a version (412) that includes the audio elements.

88. The renderer device according to any one of claims 1 to 87, wherein the context-specific data is or includes user-specific personal data of a specific human user.

89. The system according to any one of claims 1 to 88, configured to define the region of interest in such a way that objects of the audio scene representation outside the region of interest are not rendered or are rendered at a lower gain, according to at least one context-specific rule or parameter.

90. The system according to claim 89, wherein the at least one context-specific rule or parameter defines the amplitude of the region of interest.

91. The system according to claim 89 or 90, wherein the at least one context-specific rule or parameter defines the attenuation of the gain for the audio object outside the region of interest so as to attenuate according to the particular location of the audio object.

92. The system according to claim 91, configured to define the angle based on the position of the object and to define the user's head position based on a user feedback signal (492).

93. A system (400a) for providing video and audio scenes (472, 495), comprising a video renderer (495) for decoding and rendering video scenes, and a renderer device (400) according to any one of claims 1 to 92.

94. A system for providing audio content, comprising the renderer device according to any one of claims 1 to 88 and a background noise sensor, wherein the contextualization unit (440) is configured to derive context-specific rules or parameters from background noise, thereby applying a higher gain to the rendered audio signal when the background noise is higher and a lower gain to the rendered audio signal when the background noise is lower.

95. A system according to any one of claims 89 to 94, which is installed in a vehicle.

96. Audio rendering method, Processing audio scene representations (402, 412) that generate rendered audio signals (422) from audio scene representations (402, 412) conditioned by at least one context-specific rule or parameter (441, 442). Includes, This includes generating the at least one context-specific rule or parameter (441, 442) based on the context-specific data (461, 462), method.

97. A non-temporary storage unit for storing instructions, wherein when an instruction is executed by a processor, the non-temporary storage unit causes the processor to execute the method according to claim 96.