Rendering techniques

By using rendering and contextualization units in the renderer device to process audio scene representations and apply context-specific rules and parameters, the problem of existing technologies being unable to provide personalized adaptations for hearing-impaired users is solved, thus improving the audio rendering effect in virtual reality and augmented reality.

CN121264061APending Publication Date: 2026-01-02FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480035848.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-05
Filing Date
2024-04-04
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing audio rendering technologies cannot be personalized for specific users with hearing impairments, especially in virtual reality and augmented reality, resulting in a poor user experience for users with hearing impairments. Furthermore, existing technologies increase latency and complexity.

Method used

Employing a renderer device, which includes a rendering unit and a contextualization unit, it processes audio scene representations by receiving context-specific rules and parameters to generate personalized audio signals that adapt to the user's hearing impairment and environmental characteristics, including adjustments to parameters such as position, orientation, and gain weight.

Benefits of technology

It enables personalized audio rendering for hearing-impaired users in virtual reality and augmented reality, reducing latency and complexity, and improving the accessibility and user experience of audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

A renderer device (400) comprises a rendering unit (430) configured to process an audio scene representation (402, 412) to be rendered and to receive at least one context-specific rule or parameter (441, 442), the rendering unit (430) configured to generate a rendered audio signal (422) from the audio scene representation (402, 412) adjusted by the at least one context-specific rule or parameter (441, 442), a contextualization unit (440) configured to generate a rendered audio signal (422) from the audio scene representation (402, 412) adjusted by the at least one context-specific rule or parameter (441, 442). And a contextualization unit (440) configured to receive and / or derive context-specific data (461, 462), the contextualization unit (440) configured to provide at least one context-specific rule or parameter (441, 442) to the rendering unit (430) based on the context-specific data (461, 462).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to audio rendering, for example, for vehicles and / or for assistive devices for the hearing impaired. Background Technology

[0002] The inventors noted that audio renderers typically provide audio rendering independent of the user or the environment in which the user consumes the audio content. For example, in virtual reality (VR) and augmented reality (AR) scenarios, even if the user provides some feedback, that feedback is limited by the options available for the specific audio scene and may not be tailored to the user's specific profile. In other words, while the user-provided feedback is anticipated by the author (e.g., head movements to allow different perspectives and sounds), it cannot be truly customized for a specific user. For instance, if the user has hearing impairment, the audio scene should ideally be rendered considering the user's impairment, and even better, the content should be redefined to suit the user's sensitivity. However, this should be done during production, which increases the burden on the author in generating the content.

[0003] Specific examples are discussed below.

[0004] Based on the vision and goal of connecting everyone through social VR and the metaverse, audio technology needs to be designed to meet the needs of users of all ages and with a wide range of hearing abilities. However, for most consumer electronics devices, the common target user group is the normal hearing population, while people with hearing impairments are rarely considered.

[0005] The populations of all industrialized countries are aging rapidly. By 2035, the number of people aged 67 and over in Germany will increase by 22%. Currently, the median age in Germany or Japan is around 48. Globally, the growth of the population aged 65 and over will significantly outpace the growth of the total population.

[0006] According to the study "Hearing loss in the elderly - characteristics and location" (https: / / www.aerzteblatt.de / archiv / 48807 / Hoerminderung-im-Alter-Auspraegung-und-Lokalisation): Hearing loss has a high statistical probability in older adults, although it is not a natural process. ...Most hearing loss in older adults is caused by changes in the inner ear hair cells and degenerative processes in the central auditory pathway. Only 15% of older adults who meet the criteria for hearing aid use actually use them. The undersupply is due to factors including overly high expectations of hearing aids, poor acceptance, and technically unresolved speech processing strategies—strategies that compensate for the more centrally involved components of age-related hearing loss. This publication was made in 2005, but the situation has largely remained unchanged.

[0007] The global hearing aid market is projected to grow from $10.23 billion in 2022 to $17.68 billion in 2029, with the recently FDA-approved over-the-counter hearing aid segment being a significant part of this growth.

[0008] VR / AR (using MPEG-I as a form of representation) can:

[0009] a) By taking into account hearing impairment, this technology can be adapted to this large and often wealthier user group, and also lower the barrier for hearing-impaired individuals to use it.

[0010] b) Utilize social VR, metaverse, and other communication concepts soon to be launched for groups.

[0011] c) Specific areas for opening up AR / VR applications to people with hearing impairments. This could be for training and learning sessions on how to use new auditory devices or combine auditory devices with general AR technologies to enrich auditory information through visual transcripts of speech heard, for example, in AR glasses. Alternatively, dereverberation techniques can be applied to increase speech clarity.

[0012] The consequences of cochlear hearing loss are a) hearing degradation (increased hearing threshold), b) loudness perception degradation (reduced dynamic range) and c) frequency division degradation (recognition of auditory objects in complex sound scenes)[1].

[0013] A common strategy for adapting audio to impaired listeners is a post-filter 120, which uses a (rendered) stereo output 112 (output by the rendering unit 110) and generates a personalized audio signal by applying frequency-dependent amplification (EQ), frequency-dependent dynamic range compression / automatic gain control (DRC / AGC) adjustments, and (in more severe cases) frequency reduction (see [link]). Figure 1 The parameters for EQ and loudness adjustment are predetermined and stored in a hearing loss profile, for example, through a hearing aid specialist or through a hearing test application on the user's phone. Common techniques for mapping frequency-related hearing loss to amplification (primarily to improve speech intelligibility) include CAMEQ, CAMREST, DSL[i / o], FIG6, or NAL-NL1.

[0014] To address severe high-frequency hearing loss, a technique called frequency downsampling (FL) is used. Here, the spectrum is compressed and shifted so that frequency components from the region of severe hearing loss (typically high frequencies) appear in the region of less severe hearing loss.

[0015] Typically, when there is a wireless link (such as Bluetooth) between the playback device (where rendering occurs) and the headphones or hearing aid, audio adaptation can be handled on the headphones / hearing aid. This additional processing can introduce unwanted latency and reduce the ear-wearing device's usability due to its complexity and consequently increased battery consumption.

[0016] The newly established over-the-counter (OTC) hearing aid product category focuses more on improving speech clarity in noisy environments rather than selectively amplifying to address hearing loss.

[0017] In the context of object-based audio, accessibility has been explored, for example, in [2] through [4].

[0018] The problem with previous technical solutions was that audio adaptation could only affect the rendered signal (after the rendered output). This was less than ideal because it increased latency, complexity, and limited the possibility of creating more accessible audio streams.

[0019] The possibility of implementing a head-related transfer function (HRTF) has been investigated. However, HRTF does not operate on decoding or rendering, but rather on an already decoded (or rendered) audio signal. Therefore, the problems of previous techniques remain.

[0020] Based on the above, the goal is to find a rendering technology that allows for personalized or environment-specific adaptation. Summary of the Invention

[0021] According to one aspect, a renderer device is provided, comprising:

[0022] A rendering unit is configured to process an audio scene representation to be rendered and to receive at least one context-specific rule or parameter. The rendering unit is configured to generate a rendered audio signal from the audio scene representation adjusted by at least one context-specific rule or parameter.

[0023] A contextualization unit is configured to receive and / or derive context-specific data, and the contextualization unit is configured to provide at least one context-specific rule or parameter to the rendering unit based on the context-specific data.

[0024] According to one aspect, the rendering unit is configured to process an audio scene representation as including audio elements, and the rendering unit is configured to generate a rendered audio signal from the audio elements and at least one context-specific rule or parameter.

[0025] According to one aspect, the rendering unit is configured to process an audio scene representation including audio elements and metadata, and the rendering unit is configured to generate a rendered audio signal from the audio elements, metadata, and at least one context-specific rule or parameter.

[0026] According to one aspect, the metadata includes location metadata that provides information about at least one of the following: position, orientation, directivity, origin, and width of at least one object to be rendered, the at least one object to be rendered being part of an audio scene representation, wherein the rendering unit is configured to generate a rendered audio signal from audio elements, metadata, and at least one context-specific rule or parameter.

[0027] According to one aspect, the rendering unit is configured to modify metadata based on at least one context-specific rule or parameter to obtain modified metadata, wherein the rendering unit is configured to apply spatial audio processing and synthesis to audio elements based on the modified metadata.

[0028] According to one aspect, the rendering unit is configured to apply spatial audio processing and synthesis to audio elements based on metadata and at least one context-specific rule or parameter.

[0029] According to one aspect, the rendering unit is configured to be based on metadata and at least one context-specific rule or parameter combination of audio elements.

[0030] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as a context-specific location rule or parameter that associates context-specific gain weights with distance, position, pose, and / or orientation, so as to apply context-specific gain weights to the object to be rendered accordingly based on the context-specific location rule or parameter.

[0031] According to one aspect, the context-specific location rules or parameters define the gain weights as frequency-dependent, such that the rendering unit applies a first context-specific gain weight to a first frequency band and a second context-specific gain weight to a second frequency band according to the location rules or parameters.

[0032] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as including a context-specific location rule or parameter based on a distance threshold, a position threshold, or an orientation threshold, wherein the rendering unit is configured to compare the distance or position or orientation of the object to be rendered with the distance threshold, position threshold, or orientation threshold respectively, so as to avoid rendering the object if the distance or position or orientation exceeds the distance threshold, position threshold, or orientation threshold, and to render the object if the distance or position or orientation is below the distance threshold, position threshold, or orientation threshold.

[0033] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter based on a context-specific gain threshold, wherein the rendering unit is configured to compare the gain of the object to be rendered with the context-specific gain threshold so as to avoid rendering the object if the gain is lower than the context-specific gain threshold, and to render the object if the gain is higher than the context-specific gain threshold.

[0034] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as a position rule or parameter that includes a distance-dependent attenuation, a position-dependent attenuation, or an orientation-dependent attenuation that specifies the gain of the object to be rendered, wherein the position rule or parameter includes a context-specific attenuation parameter to be applied to the context-specific distance-dependent attenuation, position-dependent attenuation, or orientation-dependent attenuation.

[0035] According to one aspect, the contextualization unit is configured to define position-dependent decay as distance-dependent decay that is inversely proportional to the distance to the object to be rendered, which increases exponentially by a context-specific decay parameter.

[0036] According to one aspect, the contextualization unit is configured to define position-dependent attenuation as distance-dependent attenuation that is inversely proportional to the distance to the object to be rendered, which increases or decreases according to a context-specific attenuation parameter.

[0037] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as providing context-specific information about at least one channel-specific gain weight of the corresponding audio element to be applied to the rendered audio signal, such that the rendering unit applies the channel-specific gain weight to the corresponding audio element of the rendered audio signal.

[0038] According to one aspect, the channel-specific weights include multiple channel-specific gains, each channel-specific gain being specific to a specific frequency band, such that the rendering unit applies a first channel-specific gain weight to a first frequency band and a second channel-specific gain weight to a second frequency band according to at least one context-specific rule or parameter.

[0039] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as including a context-specific reverberation level reduction rule or parameter, such that the renderer unit performs a context-specific reduction of the reverberation level based on the context-specific reverberation level reduction rule or parameter.

[0040] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as including a context-specific early reflection level reduction rule or parameter, such that the rendering unit performs a context-specific reduction of the early reflection level based on the context-specific early reflection level reduction rule or parameter.

[0041] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter as including a dynamic range control rule or parameter, such that the rendering unit performs dynamic range control based on the dynamic range control rule or parameter.

[0042] According to one aspect, the contextualization unit is configured to derive context-specific rules or parameters from background noise so that a higher gain is applied to the rendered audio signal when the background noise is high, and a lower gain is applied to the rendered audio signal when the background noise is low.

[0043] According to one aspect, the contextualization unit is configured to define context-specific floor damping parameters for use by the rendering unit to perform floor damping based on the context-specific floor damping parameters.

[0044] According to one aspect, the contextualization unit is configured to define a context-specific culling gain parameter as used by the rendering unit to perform a linear fade-out of objects that are closer to a first source distance culling value for first-order reflections or to perform a linear fade-out of objects that are closer to a second source distance culling value for second-order reflections.

[0045] According to one aspect, the contextualization unit is configured to define a context-specific culling gain parameter that is used by the rendering unit to modify cylindrical reflections.

[0046] According to one aspect, the renderer device is configured to process an audio scene representation to obtain an audio scene representation version including audio elements, and the renderer device is further configured to generate a rendered audio signal from the audio elements and metadata.

[0047] According to one aspect of the renderer device, it is configured to process an audio scene representation to obtain an audio scene representation version including audio elements and metadata, and the renderer device is further configured to generate a rendered audio signal from the audio elements and metadata.

[0048] According to one aspect of the renderer device, it is configured to process the audio scene representation in the core decoder block to obtain an audio scene representation version including audio elements, and to process the audio scene representation version in the render block to generate a rendered audio signal from the audio elements.

[0049] According to one aspect, at least one context-specific rule or parameter includes a context-specific rule or parameter for frequency band changing, the context-specific rule or parameter associating an input frequency band with an output frequency band, such that the rendering unit changes at least one frequency band of the audio scene representation to a different frequency band of the audio element according to the context-specific rule or parameter for frequency band changing.

[0050] According to one aspect, a situation-specific rule or parameter used for frequency band changes reduces the frequency of at least one frequency band.

[0051] According to one aspect, at least one context-specific rule or parameter includes a context-specific rule or parameter for frequency-dependent gain amplification, the context-specific rule or parameter associating an input frequency band with a context-specific frequency band weight, such that the rendering unit applies the context-specific gain weight to the spectral values ​​of at least one interval of at least one frequency band in accordance with the context-specific rule or parameter.

[0052] According to one aspect, the contextualization unit is configured to define at least one rule or parameter as a geometric range.

[0053] According to one aspect, the renderer device can be configured to use at least one context-specific rule or parameter including a recommended amplification curve for each channel, the recommended amplification curve following a contextualized profile based on context-specific data.

[0054] According to one aspect, the renderer device can be configured to use at least one context-specific rule or parameter parameterized with respect to a parameter or a measurement value, so as to modulate at least one context-specific rule or parameter according to the parameter or measurement value.

[0055] According to one aspect, the rendering unit is configured to process the audio scene representation to export the audio scene representation version as including multiple metadata sets in the metadata, wherein the rendering unit is configured to release one of the metadata sets based on at least one context-specific rule or parameter.

[0056] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter based on a contextualization profile that includes multiple context-specific data, and the contextualization unit is configured to extract context-specific data related to the audio signal to be rendered and / or the rendering configuration from the contextualization profile, and derive at least one context-specific rule or parameter from the relevant context-specific data.

[0057] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter based on the parameters of the rendering unit, such that at least one context-specific rule or parameter adapts the parameters of the rendering unit to the contextualization profile.

[0058] According to a rendering unit of one aspect, wherein the contextualization unit is configured to derive a degradation model from a contextualization profile and / or context-specific data, the degradation model indicating a specific degradation in the ability of a particular human user or other audio receiving entity to acquire a rendered audio signal, wherein the contextualization unit is further configured to define at least one context-specific rule or parameter based on the degradation model to compensate for the specific degradation.

[0059] According to one aspect, the renderer device can be configured to perform simplified rendering when it is determined that the user has impaired hearing and / or reduced cognitive and / or physical sensitivity.

[0060] According to one aspect, the contextualization unit is configured to define at least one context-specific rule or parameter based on contextualization settings received in the audio scene representation.

[0061] According to one aspect, the contextualization unit is configured to select at least one context-specific rule or parameter from multiple contextualization settings received from an audio scene representation, and to select the most appropriate context-specific rule or parameter from multiple received contextualization settings based on feedback from the user, or from presets, or from manual selection, such as user-specific physical and / or cognitive auditory degeneration information.

[0062] According to one aspect, the contextualization unit is configured to access user-specific physical and / or cognitive-auditory degeneration information that provides information about user-specific physical and / or cognitive-auditory degeneration, and to define at least one context-specific rule or parameter as a user-specific rule or parameter based on the user-specific physical and / or cognitive-auditory degeneration information.

[0063] According to one aspect, user-specific physical and / or cognitive hearing degeneration information includes or comprises information about the user's cochlear degeneration.

[0064] According to one aspect, the contextualization unit is configured to generate at least one context-specific rule, which is a user-specific rule or parameter for compensating for user-specific physical and / or cognitive auditory degeneration.

[0065] According to one aspect, the renderer device can be configured to perform a configuration session to acquire user-specific physical and / or cognitive auditory degeneration information through multiple acquisitions, thereby deriving a user-specific physical and / or cognitive auditory degeneration model, and deriving at least one context-specific rule or parameter by applying parameters designed to compensate for the user-specific physical and / or cognitive auditory degeneration model.

[0066] According to one aspect, the renderer device can be configured to perform an upload session to upload user-specific physical and / or cognitive auditory degeneration information, thereby deriving a user-specific physical and / or cognitive auditory degeneration model, and deriving at least one context-specific rule or parameter by applying parameters that compensate for the user-specific physical and / or cognitive auditory degeneration model.

[0067] According to one aspect, at least one context-specific rule or parameter associates different potential characteristics of the audio scene representation with different parameters to be applied to the audio scene representation.

[0068] According to one aspect, contextualization units include simplification rules or parameters that command a reduction in the number of audio elements to be rendered, thereby reducing the number of objects to be rendered by the rendering unit.

[0069] According to one aspect, the rendering unit is configured to make a selection among the following:

[0070] A first default mode, wherein the rendering unit operates using at least one first selectable rule or parameter, wherein the first selectable rule or parameter is a default rule or parameter or a first context-specific rule or parameter among at least one context-specific rule or parameter; and

[0071] The second contextualized mode, wherein the rendering unit operates using at least one second selectable rule or parameter instead of the first selectable rule or parameter, wherein the second selectable rule or parameter is a context-specific rule or parameter that is different from the first selectable rule or parameter among at least one context-specific rule or parameter.

[0072] Depending on one aspect, the option is to control it through manual selection.

[0073] Depending on one aspect, the choice is controlled by the presence or absence of a second alternative rule or parameter.

[0074] According to one aspect, control is selected through the measurement of biological and / or physiological parameters, such that a first mode is selected when the measurement results of the biological and / or physiological parameters match a predetermined standard model of the user's physical and / or cognitive hearing, indicating a standard, and a second contextualized mode is selected when the measurement results of the biological and / or physiological parameters do not match the predetermined standard model, thereby indicating a deterioration in the user's physical and / or cognitive hearing.

[0075] According to one aspect, control is selected based on measurements of biological and / or physiological parameters, including EEG measurements.

[0076] According to one aspect, control is selected based on measurements of biological and / or physiological parameters, including heart rate measurements.

[0077] Depending on one aspect, control may be selected based on measurements of biological and / or physiological parameters, including skin conductance.

[0078] Based on one aspect, control is chosen to be achieved through feedback signals.

[0079] The renderer device, wherein the first selectable rule or parameter includes standard default rules or parameters independent of context-specific data, and the second selectable rule or parameter includes at least one contextualized selectable rule or parameter.

[0080] According to one aspect, the first optional rule or parameter includes a first context-specific rule or parameter among at least one context-specific rule or parameter, and the second optional rule or parameter includes a second context-specific rule or parameter among at least one context-specific rule or parameter.

[0081] According to one aspect, control is selected based on measurements of biological and / or physiological parameters of pupillary measurement measures, including changes in pupil size.

[0082] According to one aspect, context-specific data is or includes user-specific personalized data for a specific non-human user unit or non-human user layer.

[0083] According to one aspect, context-specific data is user-specific and provides context-specific data obtained from feedback signals.

[0084] According to one aspect, the renderer device can be configured to send rendered audio signals to an audio consumption device.

[0085] According to one aspect, the renderer device can be configured to connect wirelessly to an audio consumption device.

[0086] According to one aspect, the renderer device can be configured to receive feedback signals from an audio consumption device that indicate the user's movement, position, and / or orientation, such that the rendering unit provides a rendered audio signal based on the feedback signals.

[0087] According to one aspect, the rendered audio signal is part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation.

[0088] According to one aspect, the rendered audio signal is part of the audio scene in the metaverse environment, and the rendered audio signal is defined based on the user's position and / or orientation in the metaverse.

[0089] According to one aspect, the rendered audio signal is part of the audio scene in the video game environment, and the rendered audio signal is defined based on the position and / or orientation in the video game.

[0090] According to one aspect, the contextualization unit is configured to be based on manually input definitions and / or changes to context-specific data.

[0091] According to one aspect, the renderer device can be configured to receive:

[0092] Feedback signals indicating user location information enable the rendering unit to provide rendered audio signals based on these feedback signals; and

[0093] The rendering unit is configured to select rendered audio signals from the audio scene representation based on detected movement, position, and / or orientation from the user's location information.

[0094] The contextualization unit is configured to define at least one context-specific rule or parameter based on context-specific data, independent of the feedback signal.

[0095] According to one aspect, the renderer device can be configured to receive at least one context-specific rule or parameter having a refresh period lower than that of the feedback signal.

[0096] According to one aspect, the renderer device can be further configured to provide visual metadata specific to the audio scene representation to the video consumption device.

[0097] According to one aspect, the contextualization unit is configured to provide at least one visualization command, thereby instructing the rendering unit to provide visual metadata to the video consumption device.

[0098] According to one aspect, the renderer device can be configured to be installed in a vehicle, and the renderer device is configured to receive a first contextualized feedback signal providing position measurement results about the vehicle and a second feedback signal providing position measurement results about the user.

[0099] The contextualization unit is input with a first contextualization feedback signal as context-specific data and is configured to derive context-specific rules or parameters based on the first contextualization feedback signal, so as to:

[0100] Use context-specific rules or parameters to render the audio scene representation; and / or render the audio scene representation based on a second feedback signal.

[0101] Depending on one aspect, the renderer device can be configured to select from the following:

[0102] Rendering an audio scene representation using context-specific rules or parameters; and rendering an audio scene representation based on a second feedback signal.

[0103] According to one aspect, the renderer device can be configured to render audio scenes consistent with vehicles in a virtual environment when rendering audio scene representations using context-specific rules or parameters, and

[0104] It is configured to render the audio scene in accordance with the user's position when rendering the audio scene representation based on the second feedback signal.

[0105] According to one aspect, the renderer device may be configured to receive a second feedback signal including position feedback from gyroscopes and / or acceleration measurements, the renderer device is further configured to receive contextual feedback including contextual feedback from gyroscopes and / or acceleration measurements, and the renderer device is configured to subtract the contextual feedback from the position feedback from the contextual feedback gyroscope and / or acceleration measurements in order to render an audio scene representation using the subtraction result.

[0106] According to one aspect, the audio scene representation includes a first audio scene representation to be mixed with a second audio scene representation, wherein the first audio signal representation needs to be rendered based on vehicle location data, and the second audio signal representation needs to be rendered independently of the vehicle location data, wherein the renderer device is configured to mix the first audio scene representation and the second audio scene representation using blending weights to obtain a rendered audio signal as a blended version of the first audio scene representation and the second audio scene representation, the blending weights being defined based at least on the vehicle location data according to at least one context-specific rule or parameter.

[0107] According to one aspect, the first audio signal indicates that it needs to be rendered based on location data, and needs to be rendered based on the relative position of the vehicle and the external location, so that the relative position adjusts the blending weight.

[0108] According to one aspect, the first audio signal representation needs to be rendered based on the distance between the vehicle and the external location, so as to increase the blending weight of the second audio signal representation when the distance decreases, and decrease the blending weight of the second audio signal representation when the distance increases.

[0109] According to one aspect, the second audio signal indicates that it needs to be rendered based on the user's location data, so that the blending weights follow the user's location data.

[0110] According to one aspect, the renderer device can be configured to receive contextual input as context-specific data and user input to be input to the rendering unit, wherein the frequency of receiving contextual input is lower than the frequency of receiving user input.

[0111] According to one aspect, the renderer device can be configured to receive contextualized input as context-specific data and user input to be input to the rendering unit, wherein the refresh rate of the context-specific rules is lower than the input frequency of the user input.

[0112] According to one aspect, audio elements include audio objects.

[0113] According to one aspect, audio elements include audio channels.

[0114] According to one aspect, audio elements include stereo reverberation signals or stereo reverberation coefficients.

[0115] According to one aspect, the rendering unit is configured to provide the rendered audio signal to the audible unit.

[0116] According to one aspect, the rendering unit provides the rendered audio signal to the loudspeaker.

[0117] According to one aspect, the renderer device can be configured to receive a compressed version of the audio scene representation and perform a first decompression operation by converting the audio scene representation into a version that includes audio elements.

[0118] According to one aspect, context-specific data is or includes user-specific personalized data for a particular human user.

[0119] According to one aspect, the system can be configured to provide a mute function based on at least one context-specific rule or parameter, the mute function being associated with location data via display output and / or user, such that the mute function forces the audio object represented by the audio signal to be muted when the audio object is not displayed and / or is not within the display range or visible area.

[0120] According to one aspect, a system for providing video and audio scenes is provided, the system comprising a video renderer for decoding and rendering the video scene and a renderer device as described in any of the preceding aspects.

[0121] According to one aspect, a system for providing audio content is provided, the system including a renderer device according to one aspect and a background noise sensor, wherein a contextualization unit is configured to derive context-specific rules or parameters from the background noise so as to apply a higher gain to the rendered audio signal when the background noise is high, and to apply a lower gain to the rendered audio signal when the background noise is low.

[0122] According to one aspect, the system can be installed in a vehicle.

[0123] According to one aspect, an audio rendering method is provided, which includes:

[0124] Processing audio scene representation to generate a rendered audio signal from an audio scene representation adjusted by at least one context-specific rule or parameter.

[0125] The method involves generating at least one context-specific rule or parameter based on context-specific data.

[0126] A non-temporary storage unit for storing instructions, which, when executed by a processor, cause the processor to perform the aforementioned methods. Attached Figure Description

[0127] Figure 1 An example based on prior technology is shown.

[0128] Figure 2 The technique used to change the frequency band is shown.

[0129] Figure 3 An example of frequency-dependent amplification is shown.

[0130] Figures 4a, 4b and 4c illustrate examples according to the present invention.

[0131] Figures 5a and 5b illustrate the technology according to the present invention.

[0132] Figure 6 Figures 7a, 7b, 7c and Figure 8 The technique according to the present invention is shown.

[0133] Figures 9a and 9b illustrate examples according to the present invention.

[0134] Figure 10 The technique according to the present invention is shown.

[0135] Figures 11a, 11b, 12a and 12b illustrate the technology according to the present invention. Detailed Implementation

[0136] Figure 4a illustrates a first general example of a renderer device 400. The renderer device 400 may include a rendering unit 430. The renderer device may include a contextualization unit 440. The rendering unit may receive an audio scene representation 402 or 412 to be rendered. The rendering unit 430 may process the audio scene representations 402, 412 to generate a rendered audio signal 422 from the audio scene representations 402, 412. The contextualization unit 440 (which may be or include, for example, a personalization unit) may receive and / or derive context-specific data 461 (such as context-specific feedback). The contextualization unit 440 may generate and / or otherwise derive at least one context-specific rule and / or parameter 441, 442 to the rendering unit 430 based on the context-specific data 461. The rendered audio signal 422 may be, for example, an uncompressed audio signal. For each loudspeaker, the rendered audio signal 422 may include a specific audio channel for that particular loudspeaker (other options are possible). For example, the rendered audio signal can be transmitted to each loudspeaker via, for example, wireless (such as Bluetooth) and / or wired connections. Audio scene representations 402 and 412 can be compressed versions of the rendered audio signal 422, or any version of the rendered audio signal 422 to be rendered. For example, the rendered audio signal 422 can be a bitstream or other representations, for example, of audio elements (such as audio objects, hereinafter also referred to as "objects"), downmixed audio channels, stereo reverb elements, etc.

[0137] Figure 4b shows a more detailed view of the renderer device 400 of Figure 4a (but in some cases, the devices of Figures 4a and 4b may be considered as two different embodiments). The renderer device 400 may be part of a system 400a that provides media scenes (such as video and audio scenes 472, 496, such as virtual reality scenes or augmented reality scenes). System 400a may include a video decoder and renderer 495 for decoding and rendering video scenes (such as those provided in compressed version 402aa), as well as the renderer device 400.

[0138] Renderer device 400 may be independent of system 400a and video decoder and renderer 495. Renderer device 400 may be input bitstream 402. Renderer device 400 may output rendered audio signal 422 (bitstream 402 may be part of a media bitstream (if present) that also includes video bitstream 402aa). Renderer device 400 may include rendering unit 430. Rendering unit 430 may include core decoder 410, which may be input bitstream 402. Core decoder 410 may output audio elements and optionally output metadata (such as indicating at least one of the following: position, orientation, directivity, source, and width of at least one object to be rendered). Audio elements may be, for example, channels (such as downmixed channels, or more generally, downmixed representations of the audio scene to be represented). Alternatively or additionally, audio elements may be objects (such as positions considering audio source, directivity, etc.). Alternatively or additionally, audio elements may be or include stereo reverb elements (such as compressed stereo reverb representations of the audio scene to be rendered). Notably, audio elements may include parameters (such as metadata) that, when applied to audio elements (such as objects, stereo reverb elements, and / or channels), generate a rendered signal 422 of encoded audio scene representations 402 and 412. In some instances, audio element 412 may be less compressed than bitstream 402. However, audio element 412 may not possess the necessary form for direct rendering to a loudspeaker.

[0139] Renderer device 400 may include renderer 420 (such as part of rendering unit 430). Renderer 420 may be input audio elements 412 (channels, objects, stereo reverb) and may generate rendered audio signal 422. For example, rendered audio signal 422 may be provided to a loudspeaker (and / or headphones 470) in wired and / or wireless form.

[0140] The rendered audio signal 422 may be in a time-domain format, for example. The bitstream 422 may be in a compressed format, such as the frequency domain (e.g., Improved Discrete Cosine Transform MDCT, Improved Discrete Sine Transform MDST, etc.), the stereo reverberation domain, etc. Depending on the specific example, the audio elements of the audio scene representations 402, 412 may be in the time domain, frequency domain, or stereo reverberation domain, and these audio elements may be converted (e.g., to the time domain) by the renderer 420 (or more generally, in this example, by the rendering unit 430). Therefore, in some instances, the core decoder 410 and / or the renderer 420 (and more generally, the rendering unit 430) may perform the conversion from the compressed audio format (e.g., the frequency domain) to the time-domain audio format.

[0141] In some examples, the core decoder 410 and / or renderer 420 (more generally, rendering unit 430) may perform auditory speech (such as binauralization), but in other examples, auditory speech (such as binauralization) may be performed by an external unit outside of renderer device 400 or downstream of rendering unit 430.

[0142] For example, the rendered signal 422 can be provided to a loudspeaker and / or headphones 470. Alternatively, block 470 can be an auditory device. Sound 472 generated by the loudspeaker and / or headphones 470 can be provided to a human user 449 or other receiving entity. It is generally assumed here that user 449 is human, but the user can be replaced by a living being (such as an animal) or a receiving entity (such as an automated system, such as a robotic system, such as a non-human or non-human layer).

[0143] Contextualization unit 440 (which may be a personalization unit) may receive context-specific data 461, for example, from user 449 (or from an environment such as a vehicle, as will be shown later). Contextualization unit 440 may provide context-specific rules or parameters 442 and / or 441 to renderer 420 (or more generally, to rendering unit 430). Specifically, it is shown here that context-specific rules and / or parameters 441 may be provided to core decoder 410. Specifically, it is shown that rendering parameters 442 may be provided to renderer 420 (or more generally, to rendering unit 430). In any case, parameters or rules 441 and 442 are largely uniformly treated here without much distinction. Essentially, the differences between units 410 and 420 (and the related differences between parameters or rules 441 and 442) should be understood here only as illustrative and not restrictive.

[0144] Contextualization unit 440 may include accessibility interface 460. Accessibility interface 460 may provide context-specific data (such as in the form of a contextualization profile or context-specific profile 462) to data preprocessor 450.

[0145] The data preprocessor 450 can provide rules and / or parameters 441, 442. Contextualizing rules and / or parameters 441, 442 takes into account specific contexts (such as personalization) to adjust the rendering of audio scene representations 402, 412, and thus produces a rendered audio signal 422 by taking into account context (such as personalization). It is understood that rendering is therefore adapted to contextualizing rules and / or parameters 441, 442. The diverse possibilities for contextualization implementation and the methods for deriving rules and / or parameters 441, 442 will be explained later. It should be noted that context-specific rules and / or parameters 441, 442 may be based, for example, on input 461 from user 449 (or other receiving entity), or preferences, or other data (such as data received from a storage device). As described in input 461 (or more generally, context-specific data), accessibility interface 460 can extract (or more generally, derive) context-specific data (such as contextualized profile 462), such as personalized information (such as a personalized profile). In some instances, the contextual profile 462 is therefore independent of the specific type of the core decoder 410 and / or renderer 420 (or more generally, rendering unit 430). The data preprocessor 450 receives the contextual profile 462 (or more generally, context-specific data 462) and re-converts its information into context-specific parameters 441, 442 for use by the rendering unit 430 (e.g., for the accessibility interface 460).

[0146] The renderer device 400 may include or be connected to a feedback unit 490 (such as a unit including at least one of a head tracker and an eye tracker, a position sensor that provides position information such as posture, position, angle, direction, and movement; in some examples, the feedback unit 490 may include an accelerometer and / or a gyroscope, but in some examples, the feedback unit may include a visual sensor, such as an image acquisition sensor, a video acquisition sensor, an audio acquisition sensor, etc.) to provide feedback information 492 based on feedback 491 from the user 449 (or other receiving entity). While in one example, input 461 and feedback input 491 may be the same (and therefore line 461 may be replaced by line 461b), in other examples they may be different and perform different tasks: feedback 491 may allow the definition of the audio signal to be rendered at each moment based on user feedback (such as from head position, distance from virtual objects, currently displayed visible area, etc.) (or entity's response to rendered audio provided to it), without changing the unchanging personalized data (or more generally, context-specific data), and without changing the personalized rules and / or parameters 441, 442 (or context-specific data rules or parameters); while input (or more generally, context-specific data) 461 may serve as personalized data (or more generally, context-specific data) that changes the rendering of the scene. In some examples, personalized data (or more generally, context-specific data) 461 may be obtained during a configuration session, allowing it to be exported (e.g., via a contextualization unit, such as via accessibility interface 460) and stored as a generic contextualization profile 462. Therefore, the contextual profile 462 can be used at any time when a context needs to be presented. Contextualization can be personalized: input 461 (context-specific data) can be obtained during a configuration session to derive a personalized profile 462, such that at any time the renderer device 400 is used for a particular human user (or other type of user) 449, context-specific rules and / or parameters for the personalized profile or individual-specific profile (or more generally, a contextual profile or context-specific profile) 462 are used (e.g., specific rules 442 for defining gain based on a specific location can be defined based on profile 462). In contrast, feedback 491, 492 does not necessarily participate in the definition of any context-specific (e.g., personalized) profile, but only provides, for example, real-time information about the location of the user 449, such that specific rendered audio signals 422, 472 are modulated based on the user's location feedback 491, 492, for example by modulating a specific gain based on the user's location data (but the rules for defining the gain remain unchanged). User location feedback 491 and 492 can be understood as feedback within the scene, while input 461 (context-specific data) can be understood as modifying the rendering of the audio scene to be represented.In the example, feedback 491 can be provided as a feedback signal, while context-specific data 461 can be considered as a feedforward signal, which is maintained for multiple sessions of rendering once acquired. In the example, context-specific data 461 can remain constant, while feedback signals 491, 492 can be considered to change multiple times (e.g., multiple times per second). Although the unit providing context-specific data 461 is not shown in Figure 4b, this data can be provided by the same feedback unit 490, or in other instances, by different units (e.g., operated by a clinician). Arrow 461b in Figure 4b indicates that feedback unit 490 may optionally provide context-specific data 461.

[0147] Generally, both input 461 (contextualized input or context-specific input, context-specific data) and feedback 491 can be acquired from at least one sensor, which includes any of the following: a unit including at least one of a head tracker and an eye tracker; a position sensor providing positional information such as posture, position, angle, orientation, and movement; in some examples, at least one sensor may include an accelerometer and / or a gyroscope, but in some instances, the feedback unit may include a visual sensor, such as an image acquisition sensor, a video acquisition sensor, an audio acquisition sensor, etc. For example, at least one sensor may be in an immersive device (in which case, at least one sensor may include, for example, at least one gyroscope and / or at least one accelerometer) and / or at least one non-immersive unit (such as an audio, image, and / or video acquisition unit, etc.). It should be noted that input 491 and feedback 461 do not always have to be acquired by the same unit and / or by the same type of sensor: in some examples (such as when input 491 is acquired by a clinician), input 491 (context-specific data) may be acquired by at least one first sensor, and feedback 461 may be acquired by a second sensor (which may be of the same or different type as the first sensor). In some cases (such as in the case of vehicle 900 in Figures 9a and 9b, see below), input 491 and feedback 461 are acquired by units (4619 and 470) that may be different from each other (but in some examples they may be the same unit), but the input and feedback are acquired simultaneously. In another example (e.g., in the case where a clinician acquires contextual input 461), feedback 461 is acquired after input 491 (which may have been acquired in a previously configured session) is acquired (e.g., during an operational session).

[0148] It should be noted that blocks 470 and 490 may be part of, for example, a media content consuming device 475, which may be applied to a user. The media content consuming device 475 (such as an immersive device) may be a device for consuming virtual reality content, augmented reality content, etc. In some cases, context-specific data 461 may be acquired by the same media content consuming device 475. Alternatively, context-specific data 461 may be acquired by a device different from the media content consuming device 475, and in some cases, from a device different from the renderer device 400. It should be noted that the media content consuming device 475 is merely an example, and in some cases, input 461 and / or 491 may be acquired from other different units (such as audio, video, and / or video acquisition units, etc.).

[0149] System 400a for providing video and audio scenes may also include a video renderer 495 that can provide a rendered scene 496 taken from video bitstream 402a to a user. In some instances, the video bitstream and audio bitstream 402 may be obtained from the same source. The video renderer 495 may be modulated by metadata 482 for visualization. The metadata 482 for visualization may be received from a point-of-view metadata output unit 480. The point-of-view metadata output unit 480 may be input to input 424 from contextualization unit 440 (e.g., from data preprocessor 450). Visualization commands may be provided (e.g., through contextualization unit) to instruct the rendering unit to provide the metadata 482 for visualization to a video consumption device (e.g., video renderer 495).

[0150] Generally speaking, it can be understood that the audio signal to be rendered can be rendered according to specific (such as personalized) rules and / or parameters in a specific context. Figure 6 Figures 7a and 7b show examples of context-specific (e.g., personalized) rules or parameters. Figure 6A user 449 is presented (e.g., virtually) in an environment 150 (e.g., a virtual environment). Within environment 150, two objects, such as virtual objects (sound source S1 152 and sound source S2 152), are presented at different locations. For example, each location corresponds to the user 449's orientation within the virtual environment 150. 0° corresponds to orientation 803. 90° corresponds to orientation 802. 180° corresponds to orientation 801. The two objects (sound source S1 152; sound source S2 152) are positioned near locations 801 and 803, respectively. The gain profile varies depending on the user 449's possible orientation. The gain profile used to provide gain to the rendered audio signal 422 (472) can be adjusted via context-specific rules and / or parameters 411, 442. For example, according to the gain profile, when the user looks towards location 803, he / she will experience higher gain from sensor S2 and lower gain from sound source S1. On the other hand, when user 449 is directed toward 801, he / she will experience higher gain from sound source S1 and lower gain from sound source S2. In addition to this general rule (the orientation-based attenuation law), gain profiles may have different attenuation laws, which may be defined by context-specific rules and / or parameters 442. For example, depending on a specific personalization (contextualization), the gain profile may vary based on a specific user 449 (e.g., based on a specific personalization or contextual profile 462). For example, a user 449 with good hearing may have a gain profile that allows hearing in a different way than a user 449 with impaired hearing. For example, based on some personalized profiles 462 and personalized rules 441, 442 generated by the personalized profiles, the activation of some audio sources can be deactivated. For example, if sound source S2 is less relevant than sound source S1, the activation of the less relevant sound source S2 can be deactivated if personalized profile 462 indicates that user 449 has impaired hearing. In some other cases, such as when personalization profile 462 instructs user 449 to replace certain auditory bands, the gain profile may be modified to provide the auditory bands that the impaired user 449 can actually receive normally. However, more generally, the gain profile (or more generally, the orientation profile) varies according to a particular personalization (or even more generally, according to a particular contextualization) 442.

[0151] Another example is provided by Figures 7a and 7b, showing the situation in virtual space 150 from the first position (x' u , y' u User 449 (who has a virtual distance d'1 from sound source 1 151-1 and a distance d'2 from sound source 2 152-2) has moved. In Figure 7b, the user has moved from position (x') u , y' u Move to position (x)u , y" u ), the position (x" u , y" u The device now has a distance d"1 from sound source 1 152-1 and a distance d"2 from sound source 2 152-2. Even in this case, the attenuation law (distance-based attenuation law) can be defined, specifically adjusted by personalized rules and / or parameter 442 (derived from personalized profile 442). Different rules can be applied to define the attenuation as a function of distance from the virtual audio source. Even in this case, such as in the case of an impaired user, specific personalized rules can be selected, for example, the activation of some secondary (less relevant) audio sources can be undone.

[0152] Subsequently, it will be provided that can be based on Figure 6 The example in Figure 7b may be based on different examples.

[0153] Figure 4c illustrates an example of how context-specific rules and / or parameters 441, 442 can be instantiated by modifying (or more generally, metadata of) parameters (or more generally, metadata of) the audio scene representation 402 or 412 to be rendered. Figure 4c shows that the audio scene representation 402 or 412 to be rendered may contain elements (e.g., channels) 402a and / or 412a (which may be part of bitstream 402 or audio element 412) and metadata 402b and / or 412b. Metadata 402b or 412b may include, for example, parameters for normal compression operations (e.g., individual LPCs, linear predictive coding, parameters, whitening parameters, etc.), and / or may include, for example, a mixing matrix. Rendering unit 430 (e.g., 410 and / or 420) may include a spatial audio processing and synthesis unit 425, which may use the modified metadata 402c or 412c, which is modified based on metadata 402b, 412b obtained from the audio scene representations 402, 412. There may be a metadata modification unit 427 that modifies metadata 402b or 412b to provide modified metadata 402c or 412c based on context-specific rules or parameters 441 or 442 (indicated here by 441b and 442b, respectively). For example, if metadata 402b, 412b includes compression parameters (such as LPC parameters, whitening parameters, etc.), context-specific rules and / or parameters 441b or 442b may modify these parameters (e.g., according to specific rules and / or according to at least one weight, such as defined by contextualization unit 440, or by data preprocessor 450). If metadata 402b, 412b is defined, for example, with respect to matrices (such as a blending matrix and / or a covariance matrix and / or a correlation matrix, etc.), metadata modification unit 427 may modify the blending matrix according to context-specific parameters or rules 441b or 442b (e.g., in the case of a hearing-impaired user 449, some channels may be deactivated, or some channels may be attenuated, etc.).

[0154] The rendering unit 430 can combine audio elements 412 based on metadata (such as 412b and / or as included in or modified in 412) and context-specific rules or parameters.

[0155] In the example, audio scene representations 402 and 412 may include multiple metadata sets in the metadata. Rendering unit 430 may release one of the metadata sets based on at least one context-specific rule or parameter. For example, the rendered audio signal 422 (472) may be modulated using the remaining metadata sets.

[0156] In some cases, a default rule (first rule) may exist, which is standard or, more generally, not context-sensitive but context-sensitive. An example is provided in Figure 5a. Switch 440a is shown to toggle between a first default rule and a second context-specific rule provided by contextualization unit 440. Here, switch 440a can be controlled by feedback (such as 492, 461b, or from unit 440) or by other types of input. Thus, context-specific rules 441, 442 to be applied to rendering unit 430 can be applied differently depending on the input.

[0157] Figure 8 An example is shown that could be used, for example, in clinical applications (such as for users with hearing impairment). Here, a contextualized profile (such as a personalized profile) 462 may be provided. The contextualized profile may be provided, for example, by a clinical staff member after feedback from user 449 or other inputs (such as 461, 491, 492, 461b) has been evaluated. The contextualized profile (such as a personalized profile) 462 may be independent of a specific renderer device 400. In this case, the contextualized profile 462 may be input to a contextualization unit 440 (which may be, for example, a personalization unit), causing the contextualization unit 440 to derive context-specific rules or parameters 441, 442 (such as personalization-specific rules or parameters). For example, a degradation model 445' may be derived by a degradation model definer 445. The degradation model 445' may indicate cochlear degeneration in the user (or more generally, physical and / or cognitive hearing degeneration), or other degeneration. The degradation model 445' may be provided to a "context-specific rule or parameter generator" 446. Context-specific rule or parameter generator 446 can define context-specific rules and / or parameters 441, 442 to compensate for the degradation suffered by user 449. For example, if user 449 has abnormal hearing in one ear, the information from the abnormal hearing ear can be part of the degradation model 445', causing context-specific rule and / or parameter generator 446 to change the gain of the amplifier to be applied to the specific affected ear, while other amplifiers can have different gains (e.g., the gain for the affected ear can be greater than the gain for the normal hearing ear). This can also be applied to situations where, for example, one ear cannot distinguish certain frequency bands, in which case only some frequency bands can be modified for that ear, and so on.

[0158] Figure 8 The example could be replaced, for instance, by another technique in which the contextualized (e.g., personalized) profile 462 also integrates a degradation model 445' and provides it directly to the contextualization unit 440, so that the context-specific rule or parameter generator 446 already has the degradation model built in. Figure 8The diagram shows that the degradation model definer 445 corresponds to the accessibility interface 460 and the context-specific rule or parameter generator 446 corresponds to the data preprocessor 450. However, this correspondence is not necessarily true, and in some examples, blocks 450 and 460 are completely replaced by blocks 445 and / or 446.

[0159] In some examples, the contextual profile 462 may be defined directly by the renderer device 400 itself, for example, by analyzing the behavioral responses of the user 449 to some cues 422, 472 provided by the renderer device 400 (or other kinds of feedback or other inputs 461, 492, 461b). By analyzing the behavior or feedback from the user 449, the renderer 400 (and specifically, the degradation model definer 445) may generate its own degradation model 445', and / or the context-specific rule or parameter generator 446 may define context-specific (such as personalization-specific) rules or parameters 441, 442.

[0160] As shown in Figures 7a and 7b, the contextualization unit 440 can define at least one context-specific rule and / or parameter 441, 442 based on a distance threshold (e.g., between the user 449 and the object to be rendered 152-1 or 152-2), including at least one context-specific location rule and / or parameter. The rendering unit 430 can be configured to compare distances (e.g., d'1 and / or d"1 and / or d'2 and / or d"2) with the distance threshold. For example, rules 441, 442 may include avoiding rendering the object if the distance, position, or orientation exceeds the distance threshold, and rendering the object if the distance, position, or orientation is below the distance threshold, position threshold, or orientation threshold. For example, in Figure 7a, object 152-2 may not be rendered because the distance d'2 is greater than a predetermined distance threshold, while in Figure 7b, d'2 may not be rendered because d"2 is greater than the predetermined distance threshold. Such rules can be defined, for example, by at least one specific rule and / or parameter 441, 442: for example, for a user 449 whose contextual profile 462 indicates hearing impairment, at least one rule and / or parameter 441, 442 may specify that at least one object 152-2 to be rendered (such as a minor, less relevant object) will not be rendered if the distance exceeds the predetermined distance threshold. Essentially, the relevance of each object can be associated with its priority, and, for example, the priority can be compared with a priority threshold in the case of user 449's hearing impairment (e.g., by comparison at 440a, see Figure 5a). Therefore, in the case of user 449 with hearing impairment, the number of objects to be rendered (e.g., in audio element 412) is reduced because objects with a priority (relevance) below the priority threshold are excluded from rendering.

[0161] like Figure 6As shown, contextualization unit 440 may define at least one context-specific rule and / or parameter 441, 442 based on a specific orientation (such as the user's orientation), for example, by comparing the orientation with an orientation threshold. Rendering unit 430 may be configured to compare angular positions (such as 801, 802, 803, etc.) with an angle (or orientation) threshold. For example, rules 441, 442 may include avoiding rendering an object if the angular position (or orientation) exceeds a predetermined angle (or orientation) threshold, and rendering an object if the distance, position, or orientation is below the predetermined angle (or orientation) threshold. For example, in Figure 7a, object s1 may not be rendered when user 449 is guided to orientation position 803, and may be rendered when user 449 rotates their head towards position 801. These rules and / or parameters may be defined by at least one rule and / or parameter 441, 442.

[0162] More generally, information about the user's position (such as pose) can be considered, and one or more position (such as distance and / or angle) measurements can be compared with one or more position (such as distance and / or angle) thresholds, such that a particular object may or may not be rendered based on the comparison results with one or more position (such as distance and / or angle) thresholds.

[0163] Instead of comparing position measurements (distance, angle, pose, etc.) with thresholds, at least one rule and / or parameter 441, 442 may specify comparing gains (such as the gain of at least one channel, and / or the gain of at least one object, and / or the gain of at least one stereo reverberation component) with gain thresholds to avoid rendering specific elements (such as objects, channels, or stereo reverberation elements) based on comparisons with one or more position (such as distance and / or angle) thresholds. This could, for example, be based on a degradation model 445', allowing for simplified rendering for users with hearing impairments.

[0164] More generally, the relevance of each element 412 (e.g., channel, stereo reverberation component, or object) can be associated with the priority (relevance) of the audio element, and, for example, the priority can be compared with a priority threshold, for example, in the case of hearing impairment of user 449 (e.g., by comparison at 440a, see Figure 5a). Therefore, for example, in the case of hearing-impaired user 449, the number of objects (e.g., in audio element 412) to be rendered can be reduced because audio elements with a priority (relevance) below the priority threshold are excluded from rendering. This allows for simplified rendering for impaired users. The decision to initiate simplified rendering can be based, for example, on the identification of a specific user 449, manual selection, or pre-selection (e.g., based on a preset), for example, on feedback indicating physical and / or cognitive hearing degradation in user 449 (see below). In some examples, priorities (e.g., for each element, such as object, stereo reverberation component, and / or channel) can be read as metadata, for example, from bitstream 402 (or more generally, in audio scene representations 402, 412).

[0165] Alternatively, the contextualization unit 440 may define at least one context-specific rule and / or parameter 441, 442 as including a distance-dependent attenuation (as shown in Figures 7a and 7b) or orientation-dependent attenuation (as shown in Figures 7b) of the gain of the element to be rendered (or another property of the rendered audio signal 422, 472). Figure 6 At least one position rule and / or parameter of (in the middle) or pose-related attenuation or, more generally, position-related attenuation. At least one position rule and / or parameter 441, 442 may include a context-specific attenuation parameter to be applied to context-specific distance-related attenuation, position-related attenuation, or orientation-related attenuation. Position-related attenuation may be defined as a distance-related attenuation inversely proportional to the distance (d'1, d"1, d'2, d"2) to the distance to the object to be rendered (source 152-1 or 152-2 in Figures 7a and 7b), which increases according to the context-specific attenuation parameter. More specifically, position-related attenuation may be defined as a distance-related attenuation inversely proportional to the distance to the object to be rendered (source 152-1 or 152-2 in Figures 7a and 7b), which increases by an exponent defined by the context-specific attenuation parameter (also indicated here hereinafter as a). dist The equation can be extended to allow for faster distance-dependent decay. (Where r can be the distance from the audio source, such as one of d'1, d"1, d'2, and d"2 in Figures 7a and 7b). In In the case of, value The larger the value, the greater the gain attenuation over a given distance. In the case of, value The smaller the value, the less noticeable the gain attenuation for a given distance. This example is particularly illustrated in Figure 5b as in Figure 5a. Figure 7c shows... Examples, where ,or ,or Or, depending on the specific context, specific data 461 or 462. It is worth noting that default values ​​(such as for default rules, for example, in the first default mode, see below) can be... .

[0166] At least one rule and / or parameter used to provide context-specific information may define at least one channel-specific gain weight to be applied to the corresponding element (such as a channel, object, or stereo reverb element), such that the rendering unit 430 applies the channel-specific gain weight to the corresponding element of the rendered audio signal 422.

[0167] Channel-specific weights may include multiple object-channel-specific gains (such as channel-specific gains), each specific to a particular frequency band, such that rendering unit 430 applies a first channel-specific gain weight to a first frequency band and a second channel-specific gain weight to a second frequency band according to context-specific rules and / or parameters. This may follow, for example, a degradation model 445', causing some frequency bands (frequency bands that user 449 cannot hear) to attenuate, and simplifying the rendering result.

[0168] At least one context-specific rule and / or parameter may include a context-specific reverberation level reduction rule and / or parameter, such that the rendering unit 430 performs a context-specific reduction of the reverberation level based on at least one context-specific reverberation level reduction rule and / or parameter.

[0169] Context-specific rules and / or parameters 441, 442 can be defined as including at least one context-specific early reflection level reduction rule and / or parameter, such that rendering unit 430 performs context-specific reduction of early reflection level based on at least one context-specific early reflection level reduction rule and / or parameter.

[0170] At least one context-specific rule and / or parameter 441, 442 may include a context-specific dynamic range control rule and / or parameter, such that the rendering unit 430 performs dynamic range control based on the dynamic range control rule or parameter.

[0171] At least one scenario-specific rule and / or parameter 441, 442 may include a scenario-specific floor damping parameter, which is to be used by the rendering unit 430 to perform floor damping according to the scenario-specific floor damping parameter.

[0172] At least one context-specific rule and / or parameter 441, 442 may include a context-specific culling gain parameter to be used by the rendering unit 430 to modify cylindrical reflections.

[0173] At least one context-specific rule and / or parameter 441, 442 may include the geometric extent of the acoustic space, such that, for example, a wider acoustic space or a more restricted acoustic space is defined by at least one context-specific rule and / or parameter 441, 442.

[0174] At least one context-specific rule and / or parameter 441, 442 can associate different potential characteristics of the audio scene representation with different parameters to be applied to the audio scene representation in order to render the audio signal 422, 472 accordingly.

[0175] At least one context-specific rule and / or parameter 441, 442 may include a text-specific rule or parameter for frequency band changing. At least one context-specific rule and / or parameter for frequency band changing may associate an input frequency band (such as in audio scene representations 402, 412) with an output frequency band (of the rendered signals 422, 472). Figure 3 As shown, rendering unit 430 can change at least one frequency band of audio scene representation 402, 412 to a different frequency band of audio elements according to at least one context-specific rule and / or parameters 441, 442 for frequency band change. This can, for example, be based on hearing degradation (as indicated in degradation model 445'), causing the frequency band where the sensitivity of hearing-impaired user 449 has degraded to be moved to a frequency band where the sensitivity of hearing-impaired user 449 has not degraded (or degraded less). Figure 3 The audio scene representation 402 or 412 (specifically, channel 402a or 412a of audio scene representation 402 or 412) is shown in a frequency domain version (e.g., frequency on the horizontal axis and values ​​of each interval on the vertical axis). Each interval can be shifted from one frequency to a different frequency according to rules and / or parameters 441, 442. For example, each shifted frequency band can remain the same (or at least based on the input frequency band). After common hearing impairment degradation, the frequency band may shift from the input frequency band (represented by 402, 412) to the output frequency band (of rendered signals 422, 472) with a lower frequency than the input frequency band, but this can change according to specific degradation. For example, if degradation model 445' indicates that the user's sensitivity degrades in a specific frequency band (degradation band), contextualization unit 440 can define at least one rule and / or parameter 441, 442 that shifts the frequency band to avoid the degradation band or reduce the use of the degradation band (e.g., Figure 3The degradation model in the model suggests that user 449 has reduced sensitivity between 4000 Hz and 8000 Hz, but acceptable sensitivity below 4000 Hz (which is why the frequency band shifts from between 4000 Hz and 8000 Hz to below 4000 Hz). Therefore, this compensates for the user's hearing impairment.

[0176] exist Figure 2 In the example, the gain (e.g., for each element, such as channel, object, or stereo reverberation component) varies according to frequency and sound pressure level.

[0177] Generally, in the example, contextualization unit 440 may have access to user-specific physical and / or cognitive-auditory degeneration information (as in the degeneration model 445') that provides information about user-specific physical and / or cognitive-auditory degeneration. Therefore, contextualization unit 440 may define at least one context-specific rule and / or parameter based on the user-specific physical and / or cognitive-auditory degeneration information. An upload session may be provided to upload user-specific physical and / or cognitive-auditory degeneration information to derive a user-specific physical and / or cognitive-auditory degeneration model, and at least one context-specific rule and / or parameter may be derived by applying parameters that compensate for the user-specific physical and / or cognitive-auditory degeneration model.

[0178] In the example, the audio scene representation (e.g., bitstream 402) may include multiple contextual settings encoded therein (e.g., in metadata 402b, 412b). Contextual settings are available to contextualization unit 440, which can define context-specific rules and / or parameters 441, 442 based on the contextual settings. For example, there may be an indication (e.g., encoded in a specific field of bitstream 402) to select a pre-stored rule and / or parameter from a plurality of pre-stored rules and / or parameters, and context-specific rules and / or parameters 441, 442 will be selected based on the contextual settings. However, a limited number of parameters and / or rules may be provided in the metadata 402b, 412b, and contextualization unit 440 can select context-specific parameters and / or rules from a limited number of context-specific parameters and / or rules based on context-specific data 461.

[0179] Now, let's further illustrate the example in Figure 5a. The rendering unit 430 can perform selection among the following (e.g., via switch 440a):

[0180] - First default mode, wherein the rendering unit operates using first selectable rules and / or parameters 441, 442 (the first selectable rules and / or parameters are either default rules and / or parameters or first context-specific rules and / or parameters); and

[0181] - Second contextualized mode, wherein the rendering unit operates using at least one second selectable rule and / or parameter instead of the first selectable rule and / or parameter (the second selectable rule and / or parameter is a context-specific rule and / or parameter that is different from the first selectable rule and / or parameter).

[0182] (The example in Figure 5a is an example of a first pattern with a default rule that is independent of the context, but it can be changed to an example of a first pattern that has a default rule that is also a context-specific rule, while the second rule is always a context-specific rule).

[0183] The selection between the first and second modes (e.g., via switch 440a) can be performed manually or based on preset settings.

[0184] Alternatively (as in the example of Figure 5a), it is also... Figure 8 For example, the choice between the first (default) mode (with a first rule, contextual or context-independent) and the second (contextual) mode (as via switch 440a) can also be performed based on context. For example, feedback (or other input) 461, 461b or 491 (492) can be considered.

[0185] For example, selection (e.g. via switch 440a) can be controlled by the measurement results of biological and / or physiological parameters (e.g., feedback or other inputs 461, 461b, or 491) to enable selection:

[0186] - First mode, where the measurement results of biological and / or physiological parameters match a predetermined standard model of the user's physical and / or cognitive hearing, and

[0187] - A second contextualized mode, in cases where measurements of biological and / or physiological parameters do not match a predetermined standard model, thereby indicating a deterioration in the user's physical and / or cognitive hearing.

[0188] (In the example of Figure 7c, the first mode may imply the use of) Furthermore, the second contextualized mode can be used according to specific rules implied by a specific context. or ).

[0189] Measurements of biological and / or physiological parameters may include pupillary measurements of changes in pupil size. Measurements of biological and / or physiological parameters may include EEG measurements. Measurements of biological and / or physiological parameters may include heart rate measurements. Measurements of biological and / or physiological parameters may include conductance of skin responses.

[0190] Basically, the measurement results of biological and / or physiological parameters allow the determination of the user 449's physical state, thereby determining whether the user 449 is in a state of hearing or attention impairment, and thereby performing simplified rendering of the audio signal; otherwise, if it is determined that there is no hearing or attention impairment, full rendering (or less simplified rendering) can be performed.

[0191] Generally, the contextualization unit 440 can receive contextual input as context-specific data (461). Additionally, user input (491, 492) can be input to the rendering unit (430). The timing of contextual input can be lower than the timing of user input (491, 492). Generally, the refresh frequency (e.g., refresh rate) of context-specific rules can be lower than the input frequency (e.g., refresh cycle can be longer than input cycle) of user input (491, 492). Therefore, context-specific rules typically have greater inertia than user input and are modified more slowly.

[0192] Generally, the rendered audio signal (422) can be sent to an audio consuming device (475). The renderer device 400 can be configured to wirelessly connect to the audio consuming device (449). The renderer device can receive feedback signals (491) from the audio consuming device (475) indicating the user's movement or more generally, position (such as orientation, posture, etc.), such that the rendering unit (430) provides the rendered audio signal (422) based on the feedback signals (492). The rendered audio signal (422) can be part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation. The rendered audio signal (422) can be part of an audio scene in a metaverse environment, and the rendered audio signal is defined based on the user's position and / or orientation in the metaverse. The rendered audio signal (422) can be part of an audio scene in a video game environment, and the rendered audio signal is defined based on the position and / or orientation in the video game. The contextualization unit (440) can be defined and / or modified based on manually entered context-specific data.

[0193] In the examples above, at least in the context of dealing with hearing-impaired users, rules 141 and 142 can be defined as:

[0194] 1) Compensating for specific hearing impairments, and / or

[0195] 2) Simplify the acoustic scene (e.g., fewer objects, fewer channels, fewer stereo reverberation elements, etc.)

[0196] According to another example, context-specific data 461 may be a measurement of background noise. Here, context-specific rules 441, 442 may apply higher gain to the rendered audio signals 422, 472 when the background noise 461 is high, and lower gain to the rendered audio signals 422, 472 when the background noise 461 is low. In this case, the background noise 461 can be obtained from a sensor 490 via input 461b, which may be, for example, part of an audio consumption device 475.

[0197] It should be noted that at least one rule and / or parameter 441, 442 can be changed based on a specific pressure parameter. For example, rule 441 may require different parameters (such as gain) for different pressure parameters. Figure 6 In the examples (or in the examples of Figures 7a and 7b respectively), for instance, the gain may vary not only based on the angle (or, respectively, based on the distance from the audio source) but also based on a specific pressure parameter in order to modulate the gain according to the pressure parameter. Therefore, at least one rule and / or parameter 441, 442 may be parameterized with respect to a parameter such as pressure level (other parameters may be selected). The pressure parameter may be read, for example, from bitstream 402 (or more generally, from audio scene representations 402, 412), for example, as metadata 402b. More generally, at least one rule 441, 442 may depend on more than one value (such as pressure, frequency, or at least two of location data such as position, distance, angle, orientation, movement, velocity, acceleration, posture, etc.).

[0198] Figures 11a and 11b show Figure 6A more detailed example (alternative versions of which are Figures 12a and 12b) is provided. Here, the region of interest 800 is defined by an opening angle 802 (e.g., twice the angle) (e.g., according to at least one context-specific rule and / or parameter 441, 442). Notably, the region of interest may be defined by the position of the user 449 (e.g., by a head tracker or eye tracker, which may be sensor 470 in Figure 9a; or by a sensor attached to the user, such as sensor 470 in Figure 9b). According to at least one context-specific rule and / or parameter 441, 442, the region of interest 800 may cause audio objects outside the region of interest 800 of the audio scene representations 402, 412 to be either unrendered or rendered with a lower gain (e.g., attenuation gain) (e.g., causing angles outside but close to the region of interest 800 to cause less attenuation than angles outside but angularly further away from the region of interest 800). At least one context-specific rule and / or parameter can define the amplitude of the region of interest (e.g., aperture angle 802). At least one context-specific rule or parameter can define the gain decay of audio objects outside the region of interest, so that the gain for the audio object is attenuated according to the specific location of the audio object. For example, the angle considered could be the angle between the user's position and the object to be rendered. (Figures 12a and 12b also add stopband attenuation at specific decibel values).

[0199] Specific examples of applying the above examples to an environment 900 (such as a manned environment) are shown in FIGS. 9a and 9b, which are mainly described here as a vehicle 900 (such as a manned vehicle, such as a car). The difference between FIGS. 9a and 9b is that in FIG. 9a, the position sensor 490 is outside the user (such as a vision or audio acquisition unit), while in FIG. 9b, the position sensor 490 is engaged with the user 449 (such as an immersive device) and can be, for example, an accelerometer or a gyroscope. In FIG. 9b, the position sensor 490 is shown as separate from the loudspeaker 470, but the loudspeaker 470 can be integrated in a single unit (such as an immersive unit). The system 900 includes a specific example of the system 400, also indicated by 409. Here, the rendering unit 430 and the contextualization unit 440 are shown, and the rendering unit and the contextualization unit can be the rendering unit and the contextualization unit of FIGS. 4a and / or 4b, and thus, they can inherit any of the features generally described above and below. The first feedback 492 (corresponding to the feedback 492 of FIG. 4b or other position or pose feedback of the user 449 inside the vehicle 900) can be the position feedback of the user. Therefore, the component 490 can be a position sensor (such as a sensor that obtains position measurements, such as measurements of position and / or orientation and / or pose, etc., especially as in FIG. 9a, and also a sensor that obtains acceleration, curves, etc., especially as in FIG. 9b). The position sensor 490 can obtain measurements of position data from the user 449, for example, from the acquired images (or other types of feedback) such as visual, optical signals 491 as in FIG. 9a or acceleration as in FIG. 9b. Therefore, the first feedback 492 provided from the position sensor 490 (of FIG. 9a or FIG. 9b) to the rendering unit 430 can allow the rendering of the audio scene representation 402 or 412 to be adjusted by applying contextualization rules and / or parameters 441, 442 not defined based on the first feedback 492, as described above. Here, the contextualization unit 440 can receive the contextual feedback 461 (context-specific data), which can be different from the first feedbacks 491, 492. For example, the contextual feedback 461 can be, for example, a vehicle position feedback independent of the position feedback 461 obtained from the user 449. For example, the vehicle position feedback 461 can be provided from a vehicle position sensor 4619 (such as including a global positioning system GPS) applied to the vehicle 900 and registering the position of the vehicle 900 (such as a geographical location). Therefore, the vehicle position sensor 4619 can provide the position information (such as position, orientation, etc.) 461 of the vehicle 900. The vehicle position feedback 461 can be provided to the contextualization unit 440, for example, at a reduced temporal incidence (reduced refresh rate, or reduced refresh frequency, or reduced refresh period) relative to the position feedback 492 of the user 449 (such as, if the position feedback 492 is provided n times per second, then the vehicle position feedback 461 is provided m times per second, where m < n, for example, m << n, for example, m < n / 10).The refresh rate of at least one context-specific rule and / or parameter (441, 442) may be lower than the input frequency of the user's input (491, 492). Therefore, at least one context-specific rule and / or parameter 441, 442 may change at a lower frequency than the user's input. Thus, the vehicle context-specific rules and / or parameters 441, 442 may adjust the rendering of the audio scene representation 402, 412 to be displayed to the user (e.g., provided to the loudspeaker 470) as a rendered audio signal 472. In the example, it is possible to perform a selection based on which the contextualization unit 440 can selectively initiate and de-initiate, such that when initiated, the rendering unit 430 does not adjust the rendering result of the audio scene representation 402, 412 through any vehicle position feedback. In some instances, this selection may be performed manually or through other types of selection (e.g., through presets). When context-specific information 441, 442 is provided to rendering unit 430, rendering may be specified to be regulated by contextual feedback 461 from vehicle 900 (e.g., from the vehicle's position and / or orientation), while rendering may follow only (or at least primarily) the user's positional feedback 492 when contextual unit 440 is deactivated. An example may be provided where, when contextual unit 440 is deactivated, the position of the virtual audio source follows the head movement of user 449, while when contextual unit 440 is activated, the position of the virtual audio source may follow positional feedback 461 from vehicle 900. Therefore, it is possible to choose between two different operations where a particular feedback is dominant relative to another (e.g., selectively dominant) (e.g., the dominant feedback may be unique to rendering or unique to full rendering).

[0200] In the examples of Figure 9a or Figure 9b, the audio scene representation (such as bitstream 402, for example in its version 412) may include a first audio scene representation to be mixed with a second audio scene representation (such as the first audio scene representation may include sounds related to the external environment encountered by the vehicle during its movement, and the second audio signal representation may include music or other media content that can be consumed by user 449). In one example, the first audio signal representation may be rendered based on vehicle location data, and the second audio signal representation may be rendered independently of the vehicle location data. Renderer device 400 (409) may use blending weights (such as in spatial composition) to blend the first audio scene representation with the second audio scene representation (such as at rendering unit 430, as in renderer 420) to obtain a rendered audio signal (422, 472) as a blended version of the first audio scene representation and the second audio scene representation, the blending weights being defined at least based on the location data of vehicle 900 according to at least one context-specific rule or parameter (441, 442) (such as by contextualization unit 440). As an example, if the sound to be rendered is related to an external location (such as a sound that attracts the user's attention to the existence of a specific commercial area, such as a rest area or gas station), the sound to be rendered (encoded in the first audio scene representation) will be positioned in the direction of the commercial area's location to give the user an impression of the commercial area's location. Therefore, the mixing weights to be used for mixing are assigned to the loudspeaker in the direction of the commercial area. Thus, the first audio signal representation can be rendered based on location data (such as context-specific data acquired as input 461) according to the relative position of the vehicle and the external location of the commercial area, such that the relative position adjusts the mixing weights (e.g., higher gain is assigned to the rendered channel associated with the loudspeaker in the direction of the commercial area). The first audio signal representation can also be rendered based on the relative distance and / or relative orientation of the vehicle and the commercial area (or more generally, the external location). For example, in the case of distance relative to the second scene representation, the mixing weights for the first audio signal representation are increased. Therefore, the first audio signal representation can be rendered based on the distance between the vehicle and the external location, so that the blending weight for the first audio signal representation is increased relative to the second audio signal representation when the distance decreases, and / or the blending weight for the first audio signal representation is decreased relative to the second audio signal representation when the distance increases. Similarly, additionally or alternatively, the angle between the vehicle 900 and the external location (commercial area) can be considered. The second audio signal representation needs to be rendered based on the user's (449) location data, such that the blending weight follows the user's location data 491.

[0201] In some examples (specifically, in the example of Figure 9b), it is possible to have an embodiment in which the user's inputs 491, 492 are "polished" based on movements (such as acceleration) attributable to the motion of the vehicle, which are acquired by the vehicle's position sensor 4619 as feedback 461 (especially when the position sensor 490 of Figure 9b is or includes an accelerometer or gyroscope). More specifically, the measurement result 492 obtained from the feedback 491 can be subtracted from the measurement result 461 from the vehicle's position sensor 4619, thereby having a "polished" effect. Figure 10 An example is shown where the input (context-specific data) 461 is subtracted from the feedback 492 at subtractor 492a, thereby providing the subtracted feedback 492' to the renderer unit 430. The subtracted feedback 492' (the "polished" acceleration and / or movement of the vehicle 900) can then be used to render the audio signal 422 via the renderer unit 430. This is particularly useful in cases where the input (context-specific data) 461 and feedback 492 are provided, for example, via an accelerometer and / or gyroscope (as shown in Figure 9b): an accelerometer or gyroscope can be used as sensor 4619 to measure the motion of the vehicle 900, and another accelerometer or gyroscope can be used to measure the position (e.g., posture) measurement 492 of the user 449 (this is not correctly shown in Figure 9a, which shows sensor 470 as an image or sound acquisition sensor; however, since in principle position measurements 491 such as posture can also be affected by the acceleration and movement of the vehicle 900, the inventors have understood that...). Figure 10 The technique can also be applied in the case of Figure 9b, where sensor 470 is a visual and / or audio sensor, and feedback 461 includes at least one of the user's position, angle, orientation, and posture. It should be noted that at least one context-specific rule or parameter 441, 442 may therefore include a subtractor 492a (or at least define the operation of the subtractor).

[0202] In light of the foregoing, it is understood that context-specific rules and / or parameters 441 and 442 (whether context refers to a specific environment or a specific user 449) allow for an increase in the personalization (or contextualization) of the audio scene to be rendered without altering the production process.

[0203] In this example, the loudspeaker / headphones 470 may be, for example, part of an auditory aid. The user 449 may be human or non-human, such as a digital assistant.

[0204] In the example, although the output (e.g., 470) and / or input (e.g., 490) may be physically separate from the renderer device 400 (e.g., connections 422, 492, and / or 461 may be wireless), the following advantage is achieved: rendering (e.g., at renderer unit 430) is performed within the same hardware device as contextualization (e.g., on the same digital board or even on the same integrated circuit), thereby reducing latency.

[0205] With or without wireless connectivity (e.g., 422, 492, and / or 461), the rendered audio signal 422 can be simplified, for example, relative to the audio scene representation 402 and / or 412: thus, unwanted latency is reduced because only the signal to be rendered is processed according to at least one rule and / or parameter 441, 442, and nothing more. For example, in the case of simplified audio signals (where rendering of lower-priority audio elements is avoided if their priority is below a priority threshold), the transmission of unused audio elements is avoided. This is related to... Figure 1 In contrast, the example shown here illustrates how the wireless connection between rendering and EQ, dynamic compression, and FL is compromised due to all rendered signals. Therefore, in this example, the transmission of unnecessary data is reduced, especially given the presence of wireless transmission of rendered audio signal 422.

[0206] illustrate

[0207] This section illustrates a specific example, such as a rendering device 400 used as an accessible immersive device.

[0208] In the proposal, the audio renderer 400 may be equipped with an accessibility interface 460 that reads, processes, and / or generates a user's hearing loss profile and / or other user preferences (such as preferences for extreme scene complexity), or another instance of input 461. Subsequently, a data preprocessor module (unit) 450 maps the data (hearing loss profile) 462 from the accessibility interface 460 to internal rendering parameters. These parameters are then parsed to a core decoder 410 and / or a rendering module (renderer) 420 (or more generally, via rendering unit 430), which then renders the audio scene according to the user's requirements. In some applications, sensory information 491, such as head pose, pupil size, or eye gaze data, may be provided and included in the creation of these rendering instructions.

[0209] Some applications can support the visualization of acoustic scenes, for example, to visually highlight audible objects in a video representation of a VR scene. This functionality is indicated by a processing box labeled as a viewpoint (PoV) description output. For this purpose, the renderer device 400 can provide a viewpoint metadata output interface 480, which provides metadata 482 (such as the location of the current active source or the transmitted text string of a speech object generated by TTS). Furthermore, the application can be fed with or derive an optimized and / or simpler representation of the audio scene, which is generic or suited to an auditory loss profile, such as focusing on relevant scene elements (e.g., speech, sounds near the listener, etc.) and omitting distracting scene elements (background noise, atmospheric sounds, early reflections, reverberation, diffuse sound energy, etc.).

[0210] This example primarily aims to provide personalized use of 3D and immersive audio renderers (such as MPEG-I renderers), especially for people with hearing impairments (user 449). The general idea is to enable rendering instructions so that renderer 400 can generate spatial audio scenes 422, 472 in a more enjoyable / understandable (i.e., more accessible) way for specific impaired listeners 449.

[0211] The following subsections describe some possible approaches to increasing accessibility during rendering:

[0212] -Emotional Balance (EQ)

[0213] The head-related transfer function (HRTF) for binauralization can be filtered offline based on the suggested amplification curve of the auditory profile (e.g., as a context-specific rule and / or parameter 441, 442) (e.g., based on measurement result 461, for example, further combined with...). Figure 8 (The degradation model 445' is processed). This avoids additional latency and runtime complexity. EQ may need to be processed individually for each ear (e.g., for each channel of the rendered signal 422) to compensate for the different hearing losses in each ear.

[0214] Dynamic Range Control / Automatic Gain Control (DRC / AGC)

[0215] Based on the data in the auditory profile (such as 441, 442), the dynamic range of the output signal can be modified using the DRC function in the renderer pipeline.

[0216] -Frequency Decrease (FL)

[0217] In the frequency domain of the renderer device 400, one or more mapping rules and / or parameters 441, 442 can determine how the signal energy of the frequency range affected by hearing impairment (such as severe) is distributed to other frequency ranges that maintain hearing sensitivity, such that the spectrum is compressed (as in the previous inverse MDCT or MDST).

[0218] Acoustic scene simplification

[0219] Limit the number of sound elements rendered

[0220] Render only important elements

[0221] Example: ISO / IEC 23008-3 MPEG-H audio. The MPEG-H decoder (410) may have a decoding parameter dynamic_object_priority, which defines the priority of audio objects (such as the audio object of audio element 412). If the priority is lower than a certain threshold (such as 7, which may be the priority value assigned to a particular object), the audio object (412) may be discarded from rendering and decoding. If an object (412) is to be discarded (e.g., due to specific rules for a specific context, such as based on the selection at 440a in Figure 5a or Figure 5b), the object (412) with the lowest priority is discarded first according to rules 441, 442.

[0222] Using this feature, contextualization unit 440 can signal renderer unit 430 (such as core decoder 410) to not decode certain audio elements (412, such as objects with lower priority) based on their defined priorities.

[0223] Faster gain attenuation for distant sound sources

[0224] Example: The gain attenuation of a point sound source that varies with distance is generally calculated using g = 1 / r, where r is the distance between the sound source and the listener and g is the attenuation gain. (See, for example, ISO / IEC 23090-4:202X MPEG-I Part 4 Immersive Audio, WD2, Clause 6.6.12.4 – Distance Attenuation Due to Geometric Spread)

[0225] To allow for faster range-dependent decay, the equations can be extended to... .

[0226] exist In the case of, value The larger the value, the greater the gain attenuation for a given distance.

[0227] exist In the case of, value The smaller the value, the less noticeable the gain attenuation for a given distance.

[0228] (It is worth noting that Figure 5b describes a special case of Figure 5a)

[0229] Constrained space complexity

[0230] Acoustic flashlight effect: Mute all sounds outside the current field of view.

[0231] Or even make it dynamic based on eye-tracking data (if HMD supports it).

[0232] Reduce the effects of reverberation and early reflections

[0233] Example: In many auditory virtual environments and artificial reverberation, the gain (or mixing) of early reflections and late reverberation relative to the direct sound can be adjusted. For this specific example, see, for example, ISO / IEC 23090-4:202X MPEG-I Part 4 Immersive Audio, WD2, Clause 6.6.4.3.7 – RI Gain.

[0234] Here, the following gain parameters are defined that can be modified to reduce the effect of early reflected acoustic energy:

[0235] Tuning gain g 调谐 : Attenuates the sound level of the image source.

[0236] Floor damping g 地板 Further reduce the sound level reflected from the floor.

[0237] Remove gain g 剔除 Linear fade-out of image sources that approximate the source distance culling value earlySourceCullingDistanceOrder1 for first-order reflections or earlySourceCullingDistanceOrder2 for second-order reflections.

[0238] Cylindrical gain g 柱面 : Modify the gain value of cylindrical reflection

[0239] Rendered as mono

[0240] To the best of our knowledge, no 3D audio renderer currently explicitly includes accessibility functionality. Accessibility can be achieved through specific user interfaces and associated signal processing behaviors.

[0241] This section provides some discussion of the examples in Figures 11a and 11b. To enhance intelligibility and focus in virtual sound scenes, audio renderers can attenuate (or even mute) sounds outside the user's field of vision (such as behind the user) and / or amplify sounds in front of the user. The user's movement will, of course, affect the user's position and orientation within the scene; therefore, this effect is dynamic as a function of the positional input data.

[0242] This amplification and attenuation can be achieved by creating a beam function, which can be parameterized through a user interface.

[0243] For example, equations can be used to create classic microphone beam patterns.

[0244]

[0245] Where δ is the direction of the audio source to be rendered relative to a specific field direction (as shown in the examples of Figure 11a and Figure 11b), and Γ refers to the beamforming weight.

[0246] The ranges of a and b are between 0 and 1. For example, the following beam pattern directivity can be achieved:

[0247] Omnidirectional (i.e., no effect) a=1, b=0,

[0248] Heart shape: a = 0.5, b = 0.5

[0249] Supercardioid: a = 0.3, b = 0.7

[0250] parameter Sharpen the beam pattern.

[0251] The alternative parameterization proposed in this paper, independent of classic microphone beamforming, can be defined, for example, by parameters consisting of the opening angle (between 0° and 180°) of the region of interest and the energy attenuation slope outside the region of interest (e.g., -6 dB per 10 degrees). Additional parameters can define the desired attenuation of sound behind the user (e.g., -16 dB at 180°). Using these parameters, beamforming patterns for all directions can be created.

[0252] Alternatively, gaze information provided by an eye tracker (as shown in Figure 9a) can be used to estimate the direction the user is looking in and use that direction (rather than the forward direction) as the primary direction of interest. The beam pattern is then guided to align the main lobe with that direction.

[0253] In the examples shown in Figures 11a and 11b, the beam pattern (e.g., using a formula for beamforming weights) is... (Or alternative formula) can be an example of a context-specific rule and can be changed based on context-specific data (such as personalized data) 461 (or input 461b or feedback 491).

[0254] In the example, according to the first context-specific rule (implied by the first context-specific data 461 or input 461b or feedback 491) (e.g. in the default mode, for example, for a non-hearing impaired user), the beam pattern attenuates only the direct sound without attenuating the reflections, while in the second context-specific rule (implied by the second context-specific data 461, or input 461b or feedback 491) (e.g. in the contextualized mode, for example, for a hearing impaired user), the beam pattern attenuates not only the direct sound but also the early reflections (and / or late reflections).

[0255] Alternatively (but in some cases, as shown in Figures 11a and 11b or) Figure 6 In the example), according to the first context-specific rule (implied by the first context-specific data 461 or input 461b or feedback 491) (e.g. in the default mode, for example for a non-hearing impaired user), the beam pattern attenuates direct sound and early reflections, but does not attenuate late reflections, while in the second context-specific rule (implied by the second context-specific data 461 or input 461b or feedback 491) (e.g. in the contextualized mode, for example for a hearing impaired user), the beam pattern attenuates not only direct sound and early reflections, but also late reflections.

[0256] In the examples (e.g., in Figures 11a and 11b or examples of the figures), according to the first context-specific rule (implied by the first context-specific data 461 or input 461b or feedback 491) (e.g. in the default mode, e.g. for a non-hearing impaired user), the beam pattern attenuates direct sound and reflection by the same percentage, while in the second context-specific rule (implied by the second context-specific data 461 or input 461b or feedback 491) (e.g. in the contextualized mode, e.g. for a hearing impaired user), the beam pattern attenuates reflection by a greater percentage than the beam pattern attenuates direct sound.

[0257] In the example, context-specific rules can be combinations of rules. For example, beam patterns (such as those with formulas) Distance-dependent attenuation (e.g.) (Combination). For example, the resulting formula could be: When multiple rules are defined (e.g., for multiple patterns), two combined rules can be modified simultaneously using any of the techniques discussed above and below.

[0258] The attenuation caused by the beam pattern can have different effects on different components of the sound field. For example:

[0259] - Beam pattern can affect only the direct sound, but can attenuate or attenuate early reflections, and can attenuate or not attenuate later reflections (i.e., later reverberation).

[0260] - All early reflections from a specific sound source can be attenuated, and the attenuation value can be the same as that of a directional sound.

[0261] - The attenuation of the direct sound component of an audio element due to beam pattern and source direction (such as a specific first attenuation percentage) can be used for all associated early reflections of the audio element in the attenuation virtual space (such as the same first attenuation percentage).

[0262] - Beam pattern can attenuate sound based on the distance between the sound and the user.

[0263] In some examples, the context representation may include metadata for at least one element (such as at least one object) indicating that at least one element (such as at least one object) does not undergo contextualization (such as not being personalized), so that the rendering unit does not perform contextualization (such as personalization) (e.g., thus preventing the rendering unit from applying specific contextualization rules (such as personalization rules)). In other examples, this possibility is not foreseen.

[0264] Now, some explanations are provided regarding Figure 7c. For audio objects without extension, i.e., point source audio objects, the distance attenuation curve generated by the model can be the classic 1 / r point source distance attenuation curve, where r represents the distance from the sound source to the listener.

[0265] To increase accessibility (e.g., for hearing-impaired users), a tuning parameter α can be introduced to calculate distance-based gain attenuation, thereby changing the distance r to r α The intended effect is that the depth of the sound scene (composed of audio objects) can be increased or decreased to meet the user's needs. For example, when α > 1, the scene depth expands and distant objects become less audible and disappear. Conversely, when α < 1, the scene depth shrinks and distant objects become more audible. α = 1 (as the default value, for example in the first 7c) depicts the effect of the exponent α for three different values: 0.7, 1.0, and 1.1.

[0266] The distance index α can be provided as part of a context-specific rule.

[0267] Some explanations are provided with reference to Figures 11a, 11b, 12a, and 12b.

[0268] Directional focusing is designed to improve accessibility and is intended to attenuate distracting sounds from directions outside the region of interest. The focus can be radially symmetrical, for example, having a single "main lobe" region. Its attenuation behavior is configurable, for example, with three parameters provided by the contextualization unit 430. The default orientation is directed towards the user's forward gaze direction, but it can also be redirected to other directions, for example, controlled via other services or modalities (e.g., eye trackers, handheld controllers, etc.).

[0269] - Audio elements represented by signals and associated with the listener may not be processed by directional focusing.

[0270] Final Example

[0271] The examples of this disclosure may be implemented in hardware or software, depending on the specific implementation requirements. Implementations may be performed using digital storage media, such as floppy disks, DVDs, Blu-ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs, flash memory, hard disks, or any other magnetic or optical storage media storing electronically readable control signals, which may cooperate with a programmable computer system to perform the corresponding methods. Therefore, the digital storage can be read by a computer.

[0272] Therefore, some examples according to this disclosure include a data carrier containing electronically readable control signals that are capable of cooperating with a programmable computer system to perform any of the methods described herein.

[0273] Generally, the examples of this disclosure can be implemented as a computer program product having program code that, when the computer program product is run on a computer, can effectively perform any of the methods.

[0274] For example, program code can also be stored on a machine-readable medium.

[0275] Other examples include computer programs for performing any of the methods described herein, which are stored on a machine-readable medium.

[0276] In other words, therefore, an example of the method of the present invention is a computer program having program code that, when run on a computer, performs any of the methods described herein.

[0277] Therefore, another example of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which a computer program for performing any of the methods described herein is recorded.

[0278] Therefore, another example of the method of the present invention is a data stream or signal sequence representing a computer program for performing any of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication link (such as via the Internet).

[0279] Another example includes a processing component, such as a computer or programmable logic device, that is configured or adapted to perform any of the methods described herein.

[0280] Another example includes a computer on which a computer program for performing any of the methods described herein is installed.

[0281] Another example includes an apparatus or system configured to transmit a computer program for performing at least one of the methods described herein to a receiver. For example, the transmission may be electronic or optical. For example, the receiver may be a computer, mobile device, storage device, or similar device. For example, the apparatus or system may include a file server for transmitting the computer program to the receiver.

[0282] In some examples, a programmable logic device (such as a field-programmable gate array (FPGA)) may be used to perform some or all of the functionalities of the methods described herein. In some examples, the FPGA may cooperate with a microprocessor to perform any of the methods described herein. Generally, in some examples, these methods are performed by any hardware device. This hardware device may be any general-purpose hardware, such as a computer processor (CPU), or it may be method-specific hardware, such as an ASIC.

[0283] While the invention has been described with reference to several advantageous embodiments, modifications, substitutions, and equivalents that fall within the scope of the invention are possible. It should also be noted that many alternative ways of implementing the methods and compositions of the invention exist. Therefore, the appended claims should be interpreted as covering all modifications, substitutions, and equivalents that conform to the true spirit and scope of the invention.

[0284] References

[0285] [1]BC Moore, “Perceptual consequences of cochlear hearing loss and their implications for the design of hearing aids,” Ear ​​and hearing, vol. 17, no. 2, pp. 133-161, 1996.

[0286] [2]I. McClenaghan, L. Pardoe, and L. Ward, “The next generation ofaudio accessibility, ” 2022.

[0287] [3]L. A. Ward, “Improving Broadcast Accessibility for Hard of HearingIndividuals: using object-based audio personalisation and narrativeimportance.,” 2020, doi: 10.13140 / RG.2.2.31454.46405.

[0288] [4]J. Paulus, M. Torcoli, C. Uhle, J. Herre, S. Disch, and H. Fuchs,“Source Separation for Enabling Dialogue Enhancement in Object-basedBroadcast with MPEG-H,” J. Audio Eng. Soc., vol. 67, no. 7 / 8, pp. 510-521,Aug. 2019, doi: 10.17743 / jaes.2019.0032.

Claims

1. A renderer device (400), comprising: A rendering unit (430) is configured to process an audio scene representation (402, 412) to be rendered and to receive at least one context-specific rule or parameter (441, 442), the rendering unit (430) being configured to generate a rendered audio signal (422) from the audio scene representation (402, 412) adjusted by at least one context-specific rule or parameter (441, 442). A contextualization unit (440) is configured to receive and / or derive context-specific data (461, 462), and the contextualization unit (440) is configured to provide at least one context-specific rule or parameter (441, 442) to the rendering unit (430) based on the context-specific data (461, 462).

2. The renderer device of claim 1, wherein the rendering unit (430) is configured to process an audio scene representation (402) to include audio elements, and the rendering unit (430) is configured to generate a rendered audio signal (422) from the audio elements and at least one context-specific rule or parameter.

3. The renderer device as claimed in any of the preceding claims, wherein the rendering unit (430) is configured to process an audio scene representation (402) including audio elements and metadata, and the rendering unit (430) is configured to generate a rendered audio signal (422) from the audio elements, metadata and at least one context-specific rule or parameter.

4. The renderer device of claim 3, wherein the metadata includes location metadata providing information about at least one of the following: position, orientation, directivity, sound source, and width of at least one object to be rendered, the at least one object to be rendered being part of an audio scene representation (402), wherein the rendering unit (430) is configured to generate a rendered audio signal from audio elements, metadata, and at least one context-specific rule or parameter.

5. The renderer device of any one of claims 3 to 4, wherein the rendering unit is configured to modify metadata based on at least one context-specific rule or parameter to obtain modified metadata, wherein the rendering unit is configured to apply spatial audio processing and synthesis to audio elements based on the modified metadata.

6. The renderer device according to any one of claims 3 to 5, wherein the rendering unit (430) is configured to apply spatial audio processing and synthesis to audio elements (412) based on metadata and at least one context-specific rule or parameter (422).

7. The renderer device of any one of claims 3 to 6, wherein the rendering unit (430) is configured to combine audio elements (412) based on metadata and at least one context-specific rule or parameter (421).

8. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter (441) as a context-specific location rule or parameter that associates context-specific gain weights with distance, position, pose and / or orientation, so as to apply context-specific gain weights to the object to be rendered accordingly based on the context-specific location rule or parameter.

9. The renderer device of claim 8, wherein a context-specific location rule or parameter defines the gain weight as frequency-dependent, such that the rendering unit (430) applies a first context-specific gain weight to a first frequency band and a second context-specific gain weight to a second frequency band according to the location rule or parameter.

10. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as including a context-specific location rule or parameter based on a distance threshold, a position threshold, or an orientation threshold, wherein the rendering unit is configured to compare the distance or position or orientation of an object to be rendered with the distance threshold, the position threshold, or the orientation threshold, respectively, so as to avoid rendering the object if the distance or position or orientation exceeds the distance threshold, the position threshold, or the orientation threshold, and to render the object if the distance or position or orientation is below the distance threshold, the position threshold, or the orientation threshold.

11. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter based on a context-specific gain threshold, wherein the rendering unit is configured to compare the gain of an object to be rendered with the context-specific gain threshold to avoid rendering the object if the gain is below the context-specific gain threshold, and to render the object if the gain is above the context-specific gain threshold.

12. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as a position rule or parameter including a distance-dependent attenuation, a position-dependent attenuation, or an orientation-dependent attenuation that specifies the gain of the object to be rendered, wherein the position rule or parameter includes a context-specific attenuation parameter to be applied to the context-specific distance-dependent attenuation, the position-dependent attenuation, or the orientation-dependent attenuation.

13. The renderer device of claim 12, wherein the contextualization unit (440) is configured to define position-dependent decay as distance-dependent decay that is inversely proportional to the distance of the object to be rendered by an exponential increase defined by a context-specific decay parameter.

14. The renderer device of claim 12 or 13, wherein the contextualization unit (440) is configured to define position-dependent attenuation as distance-dependent attenuation that is inversely proportional to the distance by which the object to be rendered increases or decreases according to a context-specific attenuation parameter.

15. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as context-specific information providing at least one channel-specific gain weight for a corresponding audio element of the rendered audio signal (422) to be applied, such that the rendering unit applies the channel-specific gain weight to the corresponding audio element of the rendered audio signal (422).

16. The renderer device of claim 15, wherein the channel-specific weights include a plurality of channel-specific gains, each channel-specific gain being specific to a frequency band, such that the rendering unit (430) applies a first channel-specific gain weight to a first frequency band and applies a second channel-specific gain weight to a second frequency band according to at least one context-specific rule or parameter.

17. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as including a context-specific reverberation level reduction rule or parameter, such that the renderer unit performs context-specific reduction of the reverberation level based on the context-specific reverberation level reduction rule or parameter.

18. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as including a context-specific early reflection level reduction rule or parameter, such that the renderer unit performs a context-specific reduction of the early reflection level based on the context-specific early reflection level reduction rule or parameter.

19. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter as including a dynamic range control rule or parameter, such that the renderer unit performs dynamic range control based on the dynamic range control rule or parameter.

20. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to derive context-specific rules or parameters from background noise to apply a higher gain to the rendered audio signal when the background noise is high, and to apply a lower gain to the rendered audio signal when the background noise is low.

21. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define a context-specific floor damping parameter for use by the rendering unit (430) to perform floor damping according to the context-specific floor damping parameter.

22. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define a context-specific culling gain parameter as used by the rendering unit (430) to perform a linear fade-out on objects that are close to a first source distance culling value for first-order reflection or on objects that are close to a second source distance culling value for second-order reflection.

23. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define a context-specific culling gain parameter to be used by the rendering unit (430) to modify cylindrical reflections.

24. The renderer device as claimed in any of the preceding claims is configured to process an audio scene representation (402) to obtain an audio scene representation version (412) including audio elements, and the renderer device (400) is further configured to generate a rendered audio signal (422) from the audio elements and metadata (412).

25. The renderer device as claimed in any of the preceding claims is configured to process an audio scene representation (402) to obtain an audio scene representation version (412) including audio elements and metadata, and the renderer device (400) is further configured to generate a rendered audio signal (422) from the audio elements and metadata (412).

26. The renderer device of claim 24 or 25, configured to process an audio scene representation (402) in a core decoder block (410) to obtain an audio scene representation version (412) including audio elements, and to process the audio scene representation version (412) in a render block (430) to generate a rendered audio signal (422) from the audio elements (412).

27. The renderer device as claimed in any of the preceding claims, wherein at least one context-specific rule or parameter (441, 421) includes a context-specific rule or parameter for frequency band changing, the context-specific rule or parameter associating an input frequency band with an output frequency band, such that the rendering unit changes at least one frequency band of the audio scene representation to a different frequency band of the audio element according to the context-specific rule or parameter for frequency band changing.

28. The renderer device of claim 27, wherein the context-specific rules or parameters for frequency band changing reduce the frequency of at least one frequency band.

29. The renderer device as claimed in any of the preceding claims, wherein at least one context-specific rule or parameter (441, 442) includes a context-specific rule or parameter for frequency-dependent gain amplification, the context-specific rule or parameter associating an input frequency band with a context-specific frequency band weight, such that the rendering unit applies the context-specific gain weight accordingly to the spectral values ​​of at least one interval of at least one frequency band based on the context-specific rule or parameter.

30. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit is configured to define at least one rule or parameter as a geometric range.

31. The renderer device as claimed in any of the preceding claims is configured to use at least one context-specific rule or parameter (441, 442) including a suggested amplification curve for each channel, the suggested amplification curve following a contextualized profile based on context-specific data.

32. The device as claimed in any of the preceding claims is configured to use at least one context-specific rule or parameter (441, 442) parameterized with respect to parameters and measurements in order to modulate at least one context-specific rule or parameter (441, 442) according to said parameter or measurement.

33. The renderer device as claimed in any of the preceding claims, wherein the rendering unit is configured to process an audio scene representation to derive an audio scene representation version comprising a plurality of metadata sets in the metadata, wherein the rendering unit is configured to release one of the metadata sets based on at least one context-specific rule or parameter.

34. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter (441, 442) based on a contextualization profile (462) including a plurality of context-specific data (461), the contextualization unit (440) being configured to extract context-specific data related to the audio signal to be rendered and / or the rendering configuration from the contextualization profile (462), and to derive at least one context-specific rule or parameter (441, 442) from the related context-specific data (462).

35. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter based on the parameters of the renderer unit (430), such that the at least one context-specific rule or parameter adapts the parameters of the renderer unit (430) to the contextualization profile.

36. The rendering unit as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to derive a degradation model (435') from a contextualization profile (462) and / or context-specific data (461), the degradation model (435') indicating a specific degradation in the ability of a particular human user (449) or other audio receiving entity to acquire a rendered audio signal (422), wherein the contextualization unit (440) is further configured to define at least one context-specific rule or parameter (441, 442) as compensating for the specific degradation based on the degradation model (435').

37. The rendering unit as described in any of the preceding claims is configured to perform simplified rendering upon determining that the user has impaired hearing and / or reduced cognitive and / or physical sensitivity.

38. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define at least one context-specific rule or parameter based on contextualization settings received in the audio scene representation (402).

39. The renderer device of any one claim, wherein the contextualization unit is configured to select at least one context-specific rule or parameter from a plurality of contextualization settings received from the audio scene representation (402), and to select the most suitable context-specific rule or parameter from the plurality of received contextualization settings based on feedback from the user, or from preset, or from manually selected, such as user-specific physical and / or cognitive hearing degradation information.

40. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to access user-specific physical and / or cognitive-auditory degeneration information that provides information about user-specific physical and / or cognitive-auditory degeneration, and to define at least one context-specific rule or parameter as a user-specific rule or parameter based on the user-specific physical and / or cognitive-auditory degeneration information.

41. The renderer device of claim 40, wherein user-specific physical and / or cognitive hearing degeneration information is or includes information about the user's cochlear degeneration.

42. The renderer device of claim 40 or 41, wherein the contextualization unit (440) is configured to generate at least one context-specific rule, the at least one context-specific rule being a user-specific rule or parameter for compensating for user-specific physical and / or cognitive hearing degeneration.

43. The renderer device of any one of claims 40 to 42, configured to perform a configuration session to acquire user-specific physical and / or cognitive auditory degeneration information through multiple acquisitions in order to thereby derive a user-specific physical and / or cognitive auditory degeneration model, and to derive at least one context-specific rule or parameter by applying parameters for compensating the user-specific physical and / or cognitive auditory degeneration model.

44. The renderer device of any one of claims 40 to 43, configured to perform an upload session to upload user-specific physical and / or cognitive-auditory degeneration information, thereby deriving a user-specific physical and / or cognitive-auditory degeneration model, and deriving at least one context-specific rule or parameter by applying parameters that compensate for the user-specific physical and / or cognitive-auditory degeneration model.

45. The renderer device as claimed in any of the preceding claims, wherein at least one context-specific rule or parameter associates different potential characteristics of the audio scene representation with different parameters to be applied to the audio scene representation.

46. ​​The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) includes a simplified rule or parameter that commands a reduction in the number of audio elements to be rendered, such that the rendering unit reduces the number of objects to be rendered.

47. The renderer device as claimed in any of the preceding claims, wherein the rendering unit (430) is configured to perform a selection among: In the first default mode, the rendering unit (430) operates using at least one first selectable rule or parameter, wherein the first selectable rule or parameter is a default rule or parameter or a first context-specific rule or parameter among at least one context-specific rule or parameter; and In the second contextualized mode, the rendering unit (430) operates using at least one second selectable rule or parameter instead of the first selectable rule or parameter, wherein the second selectable rule or parameter is a context-specific rule or parameter that is different from the first selectable rule or parameter among at least one context-specific rule or parameter.

48. The renderer device of claim 47, wherein the selection is controlled by manual selection.

49. The renderer device of claim 47 or 48, wherein the selection is controlled by the presence or absence of a second selectable rule or parameter.

50. The renderer device of any one of claims 47 to 49, wherein selection is controlled by measurements of biological and / or physiological parameters to select a first mode if the measurements of the biological and / or physiological parameters match a predetermined standard model of a user's physical and / or cognitive hearing indicating a standard, and to select a second contextualized mode if the measurements of the biological and / or physiological parameters do not match the predetermined standard model, thereby indicating deterioration in the user's physical and / or cognitive hearing.

51. The renderer device of any one of claims 47 to 50, wherein selection is controlled by measurements of biological and / or physiological parameters including EEG measurements.

52. The renderer device of any one of claims 47 to 51, wherein the selection is controlled by measurements of biological and / or physiological parameters, including heart rate measurements.

53. The renderer device of any one of claims 47 to 52, wherein the selection is controlled by measurements of biological and / or physiological parameters including electrodermal response.

54. The renderer device as claimed in any one of claims 47 to 53, wherein the selection is controlled by a feedback signal (461).

55. The renderer device of any one of claims 47 to 54, wherein the first selectable rule or parameter includes a standard default rule or parameter independent of context-specific data, and the second selectable rule or parameter includes at least one contextualized selectable rule or parameter.

56. The renderer device of any one of claims 47 to 55, wherein the first selectable rule or parameter includes a first context-specific rule or parameter of at least one context-specific rule or parameter, and the second selectable rule or parameter includes a second context-specific rule or parameter of at least one context-specific rule or parameter.

57. The renderer device of any one of claims 47 to 56, wherein the selection is controlled by the measurement of biological and / or physiological parameters, including a pupillary measurement that measures changes in pupil size.

58. The renderer device as claimed in any of the preceding claims, wherein context-specific data is or includes user-specific personalized data for a particular non-human user unit or non-human user layer.

59. The renderer device as claimed in any of the preceding claims, wherein the context-specific data is user-specific and provides context-specific data obtained from feedback signals.

60. The renderer device as claimed in any of the preceding claims is configured to transmit a rendered audio signal (422) to an audio consumption device (475).

61. The renderer device of claim 60, configured to wirelessly connect to an audio consumption device (475).

62. The renderer device of claim 60 or 61, configured to receive a feedback signal (461) from an audio consumption device (475) indicating the user’s movement, position and / or orientation, such that the rendering unit (430) provides a rendered audio signal (422) based on the feedback signal (492).

63. The renderer device as claimed in any of the preceding claims, wherein the rendered audio signal (422) is part of an audio scene in a virtual reality or augmented reality environment, and the rendered audio signal is defined based on the user's position and / or orientation.

64. The renderer device as claimed in any of the preceding claims, wherein the rendered audio signal (422) is part of an audio scene in a metaverse environment, the rendered audio signal being defined based on the user’s position and / or orientation in the metaverse.

65. The renderer device as claimed in any of the preceding claims, wherein the rendered audio signal (422) is part of an audio scene in a video game environment, the rendered audio signal being defined based on position and / or orientation in the video game.

66. The renderer device as claimed in any of the preceding claims, wherein the contextualization unit (440) is configured to define and / or change context-specific data based on manual input.

67. The renderer device as claimed in any of the preceding claims is configured to receive: A feedback signal (492) indicating the user's location information causes the rendering unit (430) to provide a rendered audio signal (422) based on the feedback signal (492); and The rendering unit (420) is configured to select rendered audio signals (422) from the audio scene representation (402, 412) based on detected movement, position, and / or orientation from the user's position information (492). The contextualization unit (440) is configured to define at least one context-specific rule or parameter (441, 442) based on context-specific data (462) independently of the feedback signal (492).

68. The renderer device of claim 67, configured to receive at least one context-specific rule or parameter having a refresh period lower than that of the feedback signal (492).

69. The renderer device as claimed in any of the preceding claims is further configured to provide visual metadata specific to the audio scene representation to the video consumption device (495).

70. The renderer device of claim 69, wherein the contextualization unit (440) is configured to provide at least one visualization command that instructs the rendering unit (430) to provide visual metadata (482) to the video consumption device.

71. The renderer device as claimed in any of the preceding claims, configured to be installed in a vehicle (900), the renderer device being configured to receive a first contextualized feedback signal (461, 461b) providing position measurement results regarding the vehicle (900) and a second feedback signal (491) providing position measurement results regarding a user (449). The contextualization unit (440) is input with a first contextualization feedback signal (461, 461b) as context-specific data, and is configured to derive context-specific rules or parameters (441, 442) based on the first contextualization feedback signal (461, 461b) so as to: Use context-specific rules or parameters (441, 442) to render the audio scene representation (402, 412); and / or The audio scene representation (402, 412) is rendered based on the second feedback signal (461).

72. The renderer device of claim 71, configured to select among: Use context-specific rules or parameters (441, 442) to render audio scene representations (402, 412); and The audio scene representation (402, 412) is rendered based on the second feedback signal (461).

73. The renderer device as claimed in claim 71 or 72, configured to render an audio scene in a virtual environment consistent with the vehicle when rendering an audio scene representation using context-specific rules or parameters (402, 412), and It is configured to render the audio scene in accordance with the user's position when rendering the audio scene representation (402, 412) based on the second feedback signal (461).

74. The renderer device as claimed in claims 71 to 73, configured to receive a second feedback signal (461) including position feedback from a gyroscope and / or acceleration measurement. The renderer device is further configured to receive contextual feedback (461, 461b), including contextual feedback of gyroscope and / or acceleration measurements, and / or The renderer device is configured to subtract (492a) contextualized feedback gyroscope and / or acceleration measurement results (461) from position feedback gyroscope and / or acceleration measurement results (492) in order to render an audio scene representation using the subtraction result (492a).

75. The renderer device of any one of claims 71 to 74, wherein the audio scene representation (402, 412) includes a first audio scene representation to be mixed with the second audio scene representation. The first audio signal indicates that rendering needs to be based on vehicle location data, while the second audio signal indicates that rendering needs to be independent of vehicle location data. The renderer device is configured to use blending weights to blend a first audio scene representation with a second audio scene representation to obtain a rendered audio signal (422, 472) as a blended version of the first audio scene representation and the second audio scene representation, wherein the blending weights are defined based at least on vehicle position data according to at least one context-specific rule or parameter (441, 442).

76. The renderer device of any one of claims 71 to 75, wherein the audio scene representation (402, 412) includes a first audio scene representation to be mixed with the second audio scene representation.

77. The renderer device of claim 75, wherein the first audio signal indicates that rendering is required based on location data, and that rendering is required based on the relative positions of the vehicle and external locations, such that the relative positions adjust the blending weights.

78. The renderer device of claim 75 or 77, wherein the first audio signal representation needs to be rendered based on the distance between the vehicle and the external location, so as to increase the blending weight of the first audio signal representation relative to the second audio signal representation when the distance decreases, and / or decrease the blending weight of the first audio signal representation relative to the second audio signal representation when the distance increases.

79. The renderer device of any one of claims 75 to 78, wherein the second audio signal indicates that rendering is required based on the user's (449) position data, such that the blending weights follow the user's position data.

80. The renderer device as claimed in any of the preceding claims is configured to receive contextual input as context-specific data (461) and user input (491, 492) to be input to the rendering unit (430), wherein the time occurrence of receiving the contextual input is lower than the time occurrence of the user input (491, 492).

81. The renderer device as claimed in any of the preceding claims is configured to receive contextual input as context-specific data (461) and user input (491, 492) to be input to the rendering unit (430), wherein the refresh frequency of the context-specific rules is lower than the input frequency of the user input (491, 492).

82. The renderer device as claimed in any of the preceding claims, wherein the audio element comprises an audio object.

83. The renderer device as claimed in any of the preceding claims, wherein the audio element includes an audio channel.

84. The renderer device as claimed in any of the preceding claims, wherein the audio element includes a stereo reverb signal or a stereo reverb coefficient.

85. The renderer device as claimed in any of the preceding claims, wherein the rendering unit is configured to provide a rendered audio signal to the auditory unit.

86. The renderer device as claimed in any of the preceding claims, wherein the rendering unit provides the rendered audio signal to the loudspeaker.

87. The renderer device as claimed in any of the preceding claims is configured to receive a compressed version of an audio scene representation (402) and to perform a first decompression operation by converting the audio scene representation (402) into a version (412) including audio elements.

88. The renderer device as claimed in any of the preceding claims, wherein the context-specific data is or includes user-specific personalized data of a particular human user.

89. The system as claimed in any of the preceding claims is configured to define a region of interest according to at least one context-specific rule or parameter, such that objects outside the region of interest in the audio scene representation are not rendered or are rendered with lower gain.

90. The system of claim 89, wherein at least one context-specific rule or parameter defines the magnitude of the region of interest.

91. The system of claim 89 or 90, wherein at least one context-specific rule or parameter defines the attenuation of the gain of an audio object outside the region of interest, the attenuation being based on a specific location of the audio object.

92. The system of claim 91, configured to define an angle based on the location of the object and the user’s head position based on the user’s feedback signal (492).

93. A system (400a) for providing video and audio scenes (472, 495), the system (400a) comprising a video renderer (495) for decoding and rendering the video scenes and a renderer device (400) as claimed in any of the preceding claims.

94. A system for providing audio content, the system comprising a renderer device as claimed in any one of claims 1 to 88 and a background noise sensor, wherein a contextualization unit (440) is configured to derive context-specific rules or parameters from the background noise to apply a higher gain to the rendered audio signal when the background noise is high, and to apply a lower gain to the rendered audio signal when the background noise is low.

95. The system as claimed in any one of claims 89 to 94, wherein it is installed in a vehicle.

96. An audio rendering method, comprising: Process the audio scene representation (402, 412) to generate a rendered audio signal (422) from the audio scene representation (402, 412) adjusted by at least one context-specific rule or parameter (441, 442). The method includes generating at least one context-specific rule or parameter (441, 442) based on context-specific data (461, 462).

97. A non-transitory storage unit for storing instructions that, when executed by a processor, cause the processor to perform the method of claim 96.