Acoustic optimization for augmented reality experience
By generating a reverberation control signal outside the XR system and combining it with user input and system sensing, the limitations of XR systems in controlling and tuning the reverberation of media content are overcome, enabling personalized acoustic optimization and improving sound quality and user experience.
Patent Information
- Application Number
- CN202510985122.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-17
- Filing Date
- 2025-07-17
- Publication Date
- 2026-03-03
AI Technical Summary
Existing XR systems have limitations in controlling and tuning the simulated reverberation of media content, making it difficult to personalize and optimize for different applications and environments, resulting in poor sound quality.
By generating a reverberation control signal outside the media content, and combining it with user input and system sensing, the reverberation control signal is used to tune user-configurable acoustic settings to generate reverberation settings suitable for different applications, and simulated reverberation is synthesized through virtual acoustic simulation.
It optimizes sound quality and speech intelligibility based on media content and environment analysis, provides customized sound effects, and enhances the user's XR experience.
Smart Images

Figure CN121603834A_ABST
Abstract
Description
Background Technology
[0001] Related patent applications
[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 683,564, filed August 15, 2024, which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to extended reality (XR), and more specifically to acoustic optimizations for XR experiences. Other aspects are also described.
[0004] Background Information
[0005] A physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic devices. A physical environment can include physical features, such as physical surfaces or physical objects. For example, a physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with a physical environment through senses such as sight, touch, hearing, taste, and smell. Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic devices. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, and / or virtual reality (VR) content. Summary of the Invention
[0006] Specific embodiments of this disclosure include utilizing reverberation control signals generated outside the media content, such as user input and / or system sensing. The reverberation control signals can be used to control or tune analog reverberation applied to the media content of a given application running on a user-worn XR system on a per-application basis. In some cases, the reverberation control signals can control or tune user-configurable acoustic settings to customize each of multiple applications running on the XR system. The reverberation control signals may be based on an abstraction of one or more reverberation parameters presented to the user. The system can then alter the analog reverberation applied to the media content (e.g., video, music, or telephone calls).
[0007] Therefore, this system provides users with a tunable user interface to adjust how the audio of media content might sound when played by a specific application of the XR system. This allows users to optimize sound quality, speech intelligibility, etc., as needed based on the content of the media (which can be captured by recording) and / or analysis of the environment (room size, doors, windows, and reverberation control signals). In some cases, the system can utilize reverberation control signals to generate customized sounds (e.g., natural soundscapes or background sound enhancement), such as by sensing the environment to determine where windows and doors are located and whether they are open or closed, and then enhancing the sound blocked by closed windows or doors; or virtually generating sounds to make closed windows or doors appear open.
[0008] Some specific implementations may include a method for acoustic optimization, comprising: receiving a reverberation control signal from a user interface (UI) connected to an extended reality (XR) system worn by a user, wherein the UI presents to the user an abstraction of one or more reverberation parameters; generating a reverberation setting including multiple reverberation parameters, wherein the reverberation setting is used to control simulated reverberation of audio content, for example via headphones of the XR system, and wherein the reverberation control signal tunes the reverberation setting for a given application running on the XR system; and rendering the audio content using the reverberation setting for the given application to generate the simulated reverberation in the user's environment.
[0009] Some specific implementations may include a method for acoustic optimization, comprising: determining a video category based on the video portion of media content and determining reverberation characteristics based on the audio portion of the media content; receiving a reverberation control signal outside the media content at an XR system worn by a user; generating a reverberation setting including a plurality of reverberation parameters for controlling, for example, analog reverberation of the audio portion via headphones of the XR system, wherein the reverberation setting is generated based on i) the video category, ii) the reverberation characteristics, and iii) the reverberation control signal; and rendering the audio portion using the reverberation setting during playback of the media content in the user's environment. Other aspects are also described and protection of these other aspects is claimed.
[0010] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed descriptions below and specifically pointed out in the claims section. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description
[0011] The aspects of this disclosure are illustrated by way of example and are not limited to the illustrations in the accompanying drawings, in which similar reference numerals indicate similar elements. It should be noted that references to “a” or “an” aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a given drawing may be used to illustrate more than one aspect of this disclosure, and for a given aspect, not all elements in that drawing may be necessary.
[0012] Figure 1 This is an example of an acoustically optimized system that provides an XR experience.
[0013] Figure 2 This is an example of a user interface connected to an XR system.
[0014] Figure 3 This is an example of acoustic optimization in a user environment.
[0015] Figure 4 This is an example of the process for acoustic optimization for XR experiences.
[0016] Figure 5 This is another example of a process for acoustic optimization for XR experiences. Detailed Implementation
[0017] Using consumer electronics devices or professional recording equipment, individuals and teams can capture two-dimensional (2D) or three-dimensional (3D) video content. The resulting video files can contain audio tracks that can be mono, stereo, or spatially encoded (e.g., based on High-Order Ambient Stereo (HOA) encoding). The video content can then be played back by applications on an XR system as windowed content (such as watching 2D or 3D video on a virtual rectangular TV screen) or immersive content (such as a 360° viewing environment where the user feels immersed in the rendered video).
[0018] Audio rendering accompanying video can take various forms. For example, audio rendering can include: i) direct mono or stereo rendering of audio files to headphones; ii) point-source spatial audio rendering, where the audio is positioned at a target location, such as the center of the screen or field of view (or as a virtual stereo speaker attached to the edge of a virtual screen); or iii) immersive (surround sound) audio rendering, where spatially encoded audio in an audio track is positioned within the 3D environment surrounding the user.
[0019] Spatial audio rendering refers to a process in which a single audio track (e.g., mono) or multiple audio tracks encoded in stereo or spatial audio formats are processed into a binaural (e.g., one channel per ear of the user) audio stream to create the illusion for the user that the sound is coming from somewhere in a 3D environment (e.g., in front of or around the user). At a high level, directional information can be encoded in the head-related transfer function (HRTF) portion of the spatial audio rendering process, and the size and timbre of the space in which the audio is playing can be modeled through the reverberation portion of the rendering process.
[0020] When rendering spatial audio from captured video content, subtle or "dry" (reduced) reverb can be used to prevent spatial reverb from clashing with the reverb in the content or otherwise creating a distracting effect. However, when rendering spatial audio, too subtle or dry reverb can suppress the illusion that the sound is coming from somewhere in the space rather than from the headphones (an externalization effect). Additionally, in some cases, users may want to enhance the spatial reverb of a room to create an aesthetic "blur" for their content.
[0021] Furthermore, while some systems may attempt to achieve realistic reverberation (e.g., faithfully recreating the physical properties of the space in which the system resides), this may not be suitable for all applications running on the system. For example, some applications may expect to optimize sound effects (e.g., media), or the spectrum of human speech (e.g., telephone), or narration (e.g., storytelling or meditation), each of which may include different reverberations. However, conventional XR systems are generally limited in their ability to allow users to customize reverberation in different environments for different applications in this way.
[0022] Specific embodiments of this disclosure address problems such as these by utilizing reverberation control signals generated externally to the media content (such as user input and / or system sensing). The reverberation control signals can be used to control or tune analog reverberation applied to media content running on a given application on a user-worn XR system on a per-application basis. In some cases, the reverberation control signals can control or tune user-configurable acoustic settings to customize each of multiple applications running on the XR system. The reverberation control signals can be based on non-technical abstractions of one or more reverberation parameters presented to the user. The system can then alter the analog reverberation applied to the media content (e.g., video, music, or telephone calls).
[0023] Therefore, this system provides users with a tunable user interface to adjust how the audio of media content might sound when played by a specific application of the XR system. This allows users to optimize sound quality, speech intelligibility, etc., as needed based on the content of the media (which can be captured by recording) and / or analysis of the environment (room size, doors, windows, and reverberation control signals). In some cases, the system can utilize reverberation control signals to generate customized sounds (e.g., natural soundscapes or background sound enhancement), such as by sensing the environment to determine where windows and doors are located and whether they are open or closed, and then enhancing the sound blocked by closed windows or doors; or virtually generating sounds to make closed windows or doors appear open. Other aspects are also described and protections for other aspects are required.
[0024] In some implementations, the XR system may apply reverberation estimation to analyze recorded audio signals to determine the reverberation characteristics of the environment in which the recording took place. The system may also apply video and / or image classification to analyze image or video signals and assign categories to the location where the recording occurred (e.g., indoors vs. outdoors, or on a beach, in a classroom, in a park, in a car, etc.). The system may use sensors such as cameras, microphones, and / or lidar (light detection and ranging) to estimate the geometry and materials of the environment in which the user of the XR system is currently located (e.g., a physical room). The system can then generate simulated reverberation applicable to the environment for the virtual content.
[0025] In some implementations, based on reverberation estimation combined with video and / or image classification, the media content of audio / video files can be analyzed and output to a series of heuristic algorithms to determine the optimal reverberation settings for rendering the media content. This may include simulating reverberation and / or accessing a lookup table at runtime to match an estimated real-world space of the environment. Simulated reverberation may include a set of parameters for dynamically synthesizing reverberation using virtual acoustic simulation. In some cases, reverberation settings may be input into a lookup table to determine the optimal reverberation preset from a library. Thus, the simulated reverberation experienced by the user achieves an optimal balance between externalization, immersion, and fidelity to the original content.
[0026] In some implementations, reverb settings may include a set of parameters for dynamically synthesizing reverb using virtual acoustic simulation. In some cases, reverb settings may be entered into a lookup table to determine the optimal reverb preset from a library. When a reverb setting is selected, a reverb control signal may be utilized, which may include one or more parameters external to the audio / video content. For example, the reverb control signal may include the reverberation characteristics of the actual room or environment in which the user of the system is currently located, sensed or detected by the XR system. In some cases, the reverb control signal may include input from the user, such as playback volume level, immersion level (e.g., MR vs. full VR), apparent virtual size of the video playback window, distance of the video playback window from the user, and / or visual rendering configuration for the media (e.g., portal mode vs. full immersive mode).
[0027] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Where the shape, relative positioning, and other aspects of the described components are not explicitly defined, the scope of the invention is not limited to the components shown, which are for illustrative purposes only. Furthermore, while many details have been set forth, it should be understood that some aspects of this disclosure can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0028] In the case of an XR system, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. As an example, an XR system may detect head movement and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. As another example, an XR system may detect movement of an electronic device (e.g., a mobile phone, tablet, or laptop computer) presenting the XR environment and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), an XR system may adjust the properties of the graphical content in the XR environment in response to a representation of physical movement (e.g., a voice command).
[0029] Many different types of electronic systems enable humans to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display may utilize digital light projection, organic light-emitting diodes (OLEDs), LEDs, micro-LEDs, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some implementations, transparent or translucent displays may be configured to selectively become opaque. Projection-based systems may employ retinal projection techniques that project graphic images onto the human retina. Projection systems may also be configured to project virtual objects onto a physical environment, such as as holograms or onto a physical surface.
[0030] Figure 1 This is an example of an acoustically optimized system 100 for providing an XR experience. System 100 may be part of an XR system worn by a user, including headphones, a microphone, sensors, and a video display. System 100 may include a storage system 102 for storing or buffering media content, including video and audio portions. For example, the media content may be locally stored or streamed audio / video media files. System 100 also includes a video analyzer 104 and an audio analyzer 106. The video analyzer 104 may analyze the video portion of the media content to determine a video classification representing the environment in which the recording took place. For example, the video analyzer 104 may use video estimation algorithms to analyze the recorded video signal (represented by the video portion) and assign a classification to the location where the recording occurred, such as indoors vs. outdoors, on a beach, in a classroom, in a park, in a car, etc. Thus, the video classification may be within or dependent on the video portion of the media content.
[0031] The audio analyzer 106 can also analyze the audio portion of the media content to determine reverberation characteristics based on the audio portion representing the environment in which it was recorded. For example, the audio analyzer 106 can use a reverberation estimation algorithm to analyze the recorded audio signal (represented by the audio portion) to determine one or more reverberation characteristics, such as reverberation time, reverberation intensity, direct transmission level, reverberation transmission level, room size, absorption coefficient, and / or scattering coefficient. Therefore, reverberation characteristics can be within or dependent on the audio portion of the media content.
[0032] System 100 may also receive a reverberation control signal 108 generated outside of the media content (e.g., independent of the media content in storage system 102). In some cases, the reverberation control signal 108 may be generated via a UI connected to system 100. The UI may present an abstraction of one or more reverberation parameters to the user. This abstraction may be based on non-technical abstractions such as size, color, mood, etc. For example, the abstraction may include “humidity” versus “dryness” to control parameters such as RT60 and reverberation intensity, absorption, and scattering; or “brightness” versus “darkness” to control parameters such as high frequencies absorbed / scattered on materials. In some cases, the reverberation control signal 108 may be generated via a remote control operated (physically or virtually) by the user. In some cases, the reverberation control signal 108 may be generated via authoring tools or other programs. In some cases, the reverberation control signal 108 may be generated via sensing (e.g., a camera, microphone, LiDAR). The reverb control signal 108 can tune the reverb settings generated by the reverb setting generator 110 of the system 100 for a given application running on the XR system. The reverb control signal 108 can also customize the simulated reverb for a specific application generated by the reverb synthesizer 114 of the system 100.
[0033] For example, see also: Figure 2 UI 130 can be connected to system 100. UI 130 can be a remote control operated by the user. In some cases, the remote control can be a virtual remote control presented in the user's XR environment (e.g., an application window UI panel). In other cases, the remote control can be a physical remote control held by the user, for example, wirelessly connected to system 100.
[0034] UI 130 allows the reverberation control signal 108 to be generated by the user in a simplified manner, where one or more abstractions may correspond to one or more of a plurality of reverberation parameters presented to the user (as opposed to directly presenting technical parameters). For example, the plurality of reverberation parameters may include technical parameters such as reverberation time, reverberation intensity, direct delivery level, reverberation delivery level, room size, absorption coefficient, and / or scattering coefficient. In some cases, UI 130 may include a button 132 with presets and / or a slider 134 for adjusting the reverberation settings via parameters, such as a spatialized panel similar to a dimmer switch for a light. In some cases, button 132 may include automatic presets for various desired optimizations, such as presets for human speech, music, film, nature, underwater, silence / space, dereverberation, custom presets, enhanced background / exterior, added background noise (environment for sound design), etc. In some cases, presets may optimize sound quality for various conditions, such as sound effects (e.g., media), human speech spectrum (e.g., telephone), or narration (e.g., storytelling or meditation), each of which potentially involves different reverberations. Slider 134 can indicate an abstraction of infinitely adjustable reverb quality to simplify reverb adjustments for the user (e.g., more reverb versus less reverb on a sliding scale). This could include, for example, indications such as “humidity” versus “dryness” (for controlling RT60 and reverb intensity, or absorption and scattering), “brightness” versus “darkness” (for controlling high frequencies of absorption / scattering on materials), “warmth,” etc. UI 130 can also implement size adjustments, direct send, and / or reverb send keys corresponding to one or more of a plurality of reverb parameters.
[0035] Therefore, a single input to UI 130 allows the user to simultaneously adjust and control multiple reverberation parameters to tune the reverberation settings. Adjustments made via one or more abstractions may also include limitations on reverberation parameters (e.g., upper / lower limits) to maintain minimum spatial quality. Speaker button 136 also allows the user to provide manual placement of virtual speakers in the environment (e.g., manual placement of virtual sources).
[0036] Refer again Figure 1In some cases, the reverberation control signal 108 (generated from a source external to the media content itself, such as user input) may also indicate conditions sensed or detected in the environment. For example, the reverberation control signal 108 may indicate detected reverberation characteristics of the physical environment (measured in samples taken in the room via cameras, microphones, and / or lidar), the playback volume or selected immersion level of the XR system, the virtual size of the playback window, the distance between the playback window and the user, and / or whether a window or door is open or closed. The reverberation settings can be tuned via the reverberation control signal 108 based on sensed or detected conditions. In some cases, acoustic parameters of the user's physical (real) surrounding environment (such as RT60 parameters) may be used to provide the reverberation control signal 108, which may be determined based on sound sensed in the room using one or more microphones, or based on the size, materials, and obstacles in the room (e.g., sensed by a camera that may be integrated into the XR system worn by the user, such as headphones or XR headsets).
[0037] Therefore, system 100 can utilize reverberation setting generator 110 to generate reverberation settings with acoustic optimization for XR systems. The reverberation settings may include multiple reverberation parameters (e.g., reverberation time, reverberation intensity, direct transmission level, reverberation transmission level, room size, absorption coefficient, and / or scattering coefficient). In some cases, multiple reverberation parameters may be adjusted based on actual reverberation parameters detected in the user's physical environment. Multiple reverberation parameters can be used on an application-by-application basis (e.g., specific to each application or media content that may be running) to control simulated reverberation (generated by reverberation synthesizer 114) via headphones of system 100. The reverberation settings configured via multiple reverberation parameters can be continuously updated to dynamically synthesize reverberation over time using virtual acoustic simulation and / or lookup tables to determine the optimal reverberation preset from a library for each application.
[0038] These parameters may include the reverberation characteristics of the actual room or environment in which the user of system 100 is currently located, the playback volume level or immersion level of system 100 (e.g., MR vs. full VR), the apparent virtual size of the video playback window, the distance of the video playback window from the user, and / or the visual rendering configuration of the media (e.g., portal mode vs. full immersive mode). Reverberation control signal 108 can tune the reverberation settings for a given application running on system 100. In some implementations, the reverberation settings may be generated based on video classification from video analyzer 104, reverberation characteristics from audio analyzer 106, and / or reverberation control signal 108. In some implementations, reverberation setting generator 110 may include a series of heuristic algorithms to determine the optimal reverberation settings for rendering media content. This may include generating output that allows for the simulation of reverberation or, via reverberation synthesizer 114, access to a lookup table at runtime to fit the real-world space of the environment. In some implementations, the reverberation settings may be saved as presets accessible to the user through UI 130 (e.g., saved to a button in button 132).
[0039] System 100 may also utilize cameras, microphones, and / or lidar to estimate the geometry and materials of the environment (e.g., a physical room) in which the user of system 100 is currently located. System 100 may use reverb synthesizer 114 to generate simulated reverb 116 based on reverb settings to apply to the environment of media content. In some cases, reverb settings including multiple reverb parameters may be used to synthesize simulated reverb based on machine learning or ray tracing. In some cases, reverb settings including multiple reverb parameters may be used to look up reverb presets 118 for simulated reverb from a library.
[0040] System 100 can then utilize spatial audio renderer 120 to render audio content (e.g., the audio portion of media content) using reverb settings tailored to a given application running on an XR system. For example, the system can render audio content via headphones worn by the user's XR system during playback of the media content. Spatial audio renderer 120 can process the audio content into a binaural (e.g., one channel per ear) audio stream to create the illusion for the user that the sound originates somewhere in the 3D environment (e.g., in front of or around the user). The audio content can be rendered using acoustic optimizations to improve the user's XR experience.
[0041] For example, the audio portion of media content can be rendered as windowed or immersive content during playback of the video portion of the media content in an XR system. See also... Figure 3Acoustic optimization can be used to render audio content to improve the user's XR experience in environment 150. The user may be wearing an XR system including system 100 implemented by head-mounted display (HMD) 152. The user may also be accessing UI 130 via a remote control (e.g., an application window UI panel indicated by a dashed line and / or a handheld wireless controller indicated by a solid line). The acoustic optimization provided by system 100 can result in one or more virtual sources (e.g., virtual speakers A through D providing immersive surround sound rendering, one or more of which can be manually placed by the user at a target location in the environment via button 136 of UI 130) appearing to the user to be distributed in environment 150 to achieve the desired simulated reverberation for the application's media content.
[0042] In some implementations, system 100 may also perform pre-spatial processing 122 on the audio content before rendering using reverb settings. For example, pre-spatial processing 122 may perform content-associated de-reverb or up-mixing, including those selected via UI 130.
[0043] Therefore, System 100 can generate simulated reverberation to achieve an optimal balance between externalization, immersion, and fidelity to the original content using user input. System 100 can utilize a reverberation control signal 108 generated outside the media content to control simulated reverberation for a given application running on the XR system worn by the user, based on the user's environment (e.g., physical room). The reverberation control signal 108 provides user-configurable acoustic settings to customize each application running on the XR system using their own reverberation settings. This allows System 100 to alter the reverberation in the user's environment (e.g., the room the user is in) so that the virtual source is played back by the application. Therefore, System 100 can provide the user with tunable user interface parameters to adjust how the audio sounds when played by System 100. This allows the user to optimize sound quality, speech intelligibility, etc., as needed based on the content of the media content (captured by recording) and analysis of the user's environment (room size, doors, windows, and reverberation control signal).
[0044] In some cases, system 100 may generate custom sounds (e.g., natural soundscapes or background sound enhancement), such as by using a scan of the environment to determine where windows and doors are located and whether they are open or closed, and by enhancing the sound blocked by closed windows or doors, or virtually generating sounds to make closed windows or doors appear open.
[0045] As described above, system 100 may include a connection to headphones. Headphones may be over-ear, on-ear, loose-fitting earbuds, and in-ear sealed headphones. In some cases, system 100 may include simulated reverberation for headphones containing real-world sounds from a physical environment, such as ambient speech or background noise. This can provide the user with improved noise cancellation or a transparency mode. For example, in some cases, system 100 may include simulated reverberation containing real-world speech (e.g., speech not generated by the headphones) from another person nearby in the physical environment. This can provide a noise cancellation mode (cancellation of background noise). In another example, in some cases, system 100 may include simulated reverberation to enhance real-world background noise in the physical environment (e.g., noise that would normally be filtered out or reaches the user's ear without modification). This can provide a transparency mode.
[0046] Now refer to a flowchart of an example procedure for performing passive hearing tests. These procedures can utilize computing devices (such as those related to...) Figures 1 to 3 The processes are performed by the systems, hardware, and software described herein. These processes may be performed, for example, by executing machine-readable programs or other computer-executable instructions (such as routines, instructions, programs, or other code). The operation of the processes or other techniques, methods, or algorithms described in connection with the specific embodiments disclosed herein may be implemented directly in hardware, firmware, software executed by hardware, circuitry, or combinations thereof.
[0047] For the sake of simplicity, these processes are depicted and described herein as a series of operations. However, the operations according to this disclosure can be performed in various orders and / or simultaneously. Additionally, other operations not presented and described herein may be used. Furthermore, not all exemplified operations are required to implement the processes according to the disclosed subject matter.
[0048] Figure 4 This is an example of a process 400 for acoustic optimization of an XR experience. At operation 402, the system can receive a reverberation control signal from a UI connected to the XR system at the XR system worn by the user. The XR system may include headphones (e.g., over-ear), AR, VR, or MR video displays, and sensors (e.g., cameras, microphones, and / or LiDAR). The UI may present an abstraction of one or more reverberation parameters to the user. For example, the XR system 100 worn by the user may receive a reverberation control signal 108 external to media content, such as via a UI 130 connected to the system. In some cases, when received via the UI 130, the reverberation control signal may be generated by a remote control held and operated by the user. In some cases, the reverberation control signal may be generated via an authoring tool or other program.
[0049] At operation 404, the system can generate a reverberation setting that includes multiple reverberation parameters. This reverberation setting can be used to control a user's analog reverberation via headphones from the XR system (e.g., applied to an audio signal for audio content such as video, music, or a phone call). For example, reverberation setting generator 110 can generate the reverberation setting based on a reverberation control signal. The multiple reverberation parameters may include reverberation time, reverberation intensity, direct transmission level, reverberation transmission level, room size, absorption coefficient, and / or scattering coefficient associated with the user's environment (e.g., environment 150). The reverberation control signal can tune the reverberation setting for a given application running on the XR system. For example, multiple applications may be running on the system, and each application may have its own reverberation setting tuned for the user based on the reverberation control signal.
[0050] At operation 406, the system can render audio content using reverb settings for a given application to generate simulated reverb in the user's environment. For example, spatial audio renderer 120 can render audio content (e.g., the audio portion of media content, which may also include video portions) using reverb settings for a given application running on an XR system. The system can render the audio content via headphones worn by the user's XR system during playback of the media content. The audio content can be rendered using acoustic optimizations to improve the user's XR experience. For example, acoustic optimizations can cause virtual sources (e.g., virtual speakers A through D) to appear distributed in the user's environment to achieve the desired reverb.
[0051] At operation 408, the system can determine whether there is a detected change in the reverb control signal, such as a change in the same reverb control signal that was previously sent or received as an additional reverb control signal. If the reverb control signal has not changed (“No”), the system can continue rendering the audio content at operation 406. However, if the reverb control signal has changed (“Yes”), the system can return to operation 404 to further tune the reverb settings based on the change.
[0052] Figure 5 This is an example of a process 500 for acoustic optimization of an XR experience. At operation 502, the system determines a video category based on the video portion of the media content and determines reverberation characteristics based on the audio portion of the media content. For example, an XR system 100 worn by a user can utilize a video analyzer 104 and an audio analyzer 106 to determine the video category based on the video portion of the media content and the reverberation characteristics based on the audio portion of the media content, respectively.
[0053] At operation 504, the system can receive reverberation control signals external to the media content at the XR system worn by the user. For example, the XR system may include headphones (e.g., over-ear type), AR, VR, or MR video displays, and sensors (e.g., cameras, microphones, and / or LiDAR). The XR system 100 worn by the user can receive the reverberation control signals 108 via a UI 130 and / or via sensing. For example, in some cases, the reverberation control signals may be generated via a UI connected to the system (such as a remote control operated by the user) or via authoring tools or other programs. In some cases, the reverberation control signals may be generated via sensing (such as the system's camera, microphone, and / or LiDAR).
[0054] At operation 506, the system can generate a reverberation setting that includes multiple reverberation parameters for simulating reverberation of the audio portion of audio content or media content controlled via headphones of the XR system. The reverberation setting is generated based on i) video classification, ii) reverberation characteristics, and iii) a reverberation control signal. For example, reverberation setting generator 110 can generate the reverberation setting based on a reverberation control signal. The multiple reverberation parameters may include reverberation time, reverberation intensity, direct transmission level, reverberation transmission level, room size, absorption coefficient, and / or scattering coefficient associated with the user's environment (such as environment 150). The reverberation setting can be generated based on video classification from video analyzer 104, reverberation characteristics from audio analyzer 106, and reverberation control signal 108 from a UI (e.g., a remote control, authoring tool, or other program). One or more of the video classification, reverberation characteristics, and / or reverberation control signal can be used to tune the reverberation setting for a given application running on the XR system. For example, multiple applications may be running on the system, and each application may have its own reverb settings that are tuned for the user based on one or more of video classification, reverb characteristics, and / or reverb control signals.
[0055] At operation 508, the system can render audio portions using reverb settings during playback of media content in the user's environment. For example, spatial audio renderer 120 can render audio content (e.g., the audio portion of the media content) using reverb settings during playback of media content on an XR system. The system can render audio content via headphones worn by the user's XR system during playback of the media content. The audio content can be rendered using acoustic optimizations to improve the user's XR experience. For example, acoustic optimizations can cause virtual sources (e.g., virtual speakers A through D) to appear distributed in the user's environment to achieve the desired reverb.
[0056] At operation 510, the system can determine whether there is a detected change in one or more of the video classification, reverberation characteristics, and / or reverberation control signals. If there is no change (“No”), the system can continue rendering the audio content at operation 406. However, if there is a change (“Yes”), the system can return to operation 404 to further tune the reverberation settings based on that change.
[0057] As described above, one aspect of the present invention is the collection and use of data available from specific and legitimate sources for acoustic optimization of XR experiences. This disclosure contemplates that, in some instances, the collected data may include personal information data that uniquely identifies or can be used to identify specific individuals. Such personal information data may include demographic data, location-based data, online identifiers, telephone numbers, email addresses, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other personal information.
[0058] This disclosure recognizes that the use of such personal information data in the techniques of this invention can benefit users. For example, personal information data can be used for acoustic optimization of XR experiences. Therefore, the use of such personal information data enables users to have greater control over the content delivered.
[0059] This disclosure anticipates that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information data will comply with established privacy policies and / or privacy practices. Specifically, it is expected that such entities will implement and consistently apply privacy practices generally recognized as meeting or exceeding industry or governmental requirements for protecting user privacy. Such information regarding the use of personal data should be highlighted and easily accessible to users, and should be updated as data collection and / or use change. Users' personal information should be collected only for lawful use. Furthermore, such collection / sharing should only occur after receiving user consent or other lawful grounds provided for in applicable law. Additionally, such entities should consider taking any necessary steps to protect and safeguard the right to access such personal information data and to ensure that other entities with access to personal information data comply with their privacy policies and procedures. Furthermore, such entities may be subject to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices. Moreover, policies and practices should be tailored to the specific types of personal information data collected and / or accessed, and made applicable to applicable laws and standards, including jurisdiction-specific considerations that may allow for the application of higher standards. For example, in the United States, the collection or access to certain health data may be governed by federal and / or state laws (such as the Health Insurance Portability and Accountability Act (HIPAA)); while health data in other countries may be subject to other regulations and policies and should be handled accordingly.
[0060] Regardless of the foregoing, this disclosure also contemplates implementation schemes for users to selectively block the use or access to personal information data. That is, this disclosure contemplates hardware and / or software components to prevent or block access to such personal information data. For example, such as with respect to acoustic optimization of the XR experience, the inventive technology can be configured to allow users to opt-in or opt-out at any time during or after registration for the service to participate in the collection of personal information data.
[0061] Furthermore, the intent of this disclosure is that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Deidentification can be facilitated, where appropriate, by removing identifiers, controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods (such as differentiated privacy).
[0062] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, content can be selected and delivered to the user based on aggregated non-personal information data or an absolute minimum amount of personal information, such as content disposed solely on the user's device or other non-personal information available for content delivery services.
[0063] When utilizing the various aspects of the embodiments, it will become apparent to those skilled in the art that combinations or variations of the above embodiments are possible for acoustic optimization of the XR experience. Although the embodiments have been described in language specific to structural features and / or methodological behavior, it should be understood that the appended claims are not necessarily limited to the specific features or behaviors described. Rather, the specific features and behaviors disclosed should be understood as embodiments of the illustrative claims.
Claims
1. A method for acoustic optimization, the method comprising: At the extended reality (XR) system worn by the user, a reverberation control signal is received from a user interface (UI) connected to the XR system, wherein the UI presents to the user an abstraction of one or more reverberation parameters; Generate a reverb setting that includes multiple reverb parameters, wherein the reverb setting is used to control the analog reverb of audio content, and wherein the reverb control signal tunes the reverb setting for a given application running on the XR system; as well as The audio content is rendered using the reverb settings for the given application to generate the simulated reverb for the user.
2. The method of claim 1, wherein the reverberation control signal is generated via a remote control operated by the user.
3. The method of claim 1, wherein multiple reverb parameters are adjusted simultaneously by a single input from the user to the UI.
4. The method of claim 1, wherein the abstract indication corresponds to one or more of the one or more reverberation parameters of humidity and dryness, brightness and darkness, or warmth.
5. The method of claim 1, wherein the reverberation control signal is generated via a creation tool that customizes the simulated reverberation for each of a plurality of applications.
6. The method of claim 1, wherein the plurality of reverberation parameters includes two or more of reverberation time, reverberation intensity, direct transmission level, reverberation transmission level, room size, absorption coefficient, or scattering coefficient.
7. The method of claim 1, wherein the plurality of reverberation parameters are used to synthesize the simulated reverberation based on machine learning or ray tracing.
8. The method of claim 1, wherein the plurality of reverberation parameters are used to find a reverberation preset for the simulated reverberation.
9. The method of claim 1, wherein the simulated reverberation in noise cancellation mode is combined with the real-world speech of another person in the physical environment.
10. The method of claim 1, wherein the simulated reverberation is included in a transparency mode to enhance real-world background noise in the physical environment.
11. A method for acoustic optimization, the method comprising: The video category is determined based on the video portion of the media content, and the reverberation characteristics are determined based on the audio portion of the media content. The reverberation control signal outside the media content is received at the extended reality (XR) system worn by the user; Generate a reverb setting, the reverb setting including multiple reverb parameters for controlling the analog reverb of the audio portion, wherein the reverb setting is generated based on i) the video classification, ii) the reverb characteristics and iii) the reverb control signal; as well as The audio portion is rendered using the reverb settings during playback of the media content to the user.
12. The method of claim 11, wherein the reverberation control signal is generated via a user interface (UI) that enables the user to tune the simulated reverberation for a given application playing the media content.
13. The method of claim 11, wherein the reverberation control signal is generated via a tool that customizes the simulated reverberation for a given application running on the XR system.
14. The method of claim 11, wherein the reverberation control signal indicates the reverberation characteristics of the environment.
15. The method of claim 11, wherein the reverberation control signal indicates the playback volume or a selected immersion level.
16. The method of claim 11, wherein the reverberation control signal indicates the virtual size of the playback window or the distance between the playback window and the user.
17. The method of claim 11, wherein the reverberation control signal indicates whether a window or door is open or closed.
18. The method of claim 11, wherein the reverberation parameter is adjusted based on a reverberation parameter detected in the environment.
19. The method according to claim 11, further comprising: The reverb settings are saved as presets that the user can access via a UI connected to the XR system.
20. The method of claim 11, wherein the audio portion is rendered as windowed content or immersive content during playback of the video portion by the XR system.