Neuroacoustic modeling for audio environments
By using neuroacoustic modeling technology and generating implicit acoustic representations of the audio environment through audio rendering models, the design and placement challenges of audio capture devices in the audio environment are solved, achieving optimal acoustic performance and audio experience for the audio system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies make it difficult to accurately design and deploy audio capture devices in audio environments to provide the desired acoustic performance, and traditional 3D models lack the functionality of audio attributes, leading to inaccuracies and inefficiencies in audio systems.
Using neuroacoustic modeling techniques, an implicit acoustic representation of the audio environment is generated through an audio rendering model based on impulse response and geometric representation. The model is then trained using neural radiation fields to optimize the layout and configuration of the audio system.
It achieves optimal design and installation of audio systems in audio environments, providing precise acoustic performance and audio experience, reducing noise and improving audio quality.
Smart Images

Figure CN121693776A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 519,959, filed August 16, 2023, entitled “NEURALACOUSTIC MODELING FOR AN AUDIO ENVIRONMENT,” the entire contents of which are hereby incorporated by reference. Technical Field
[0003] Embodiments of this disclosure generally relate to audio processing, and more specifically, to systems configured to provide and / or utilize neuroacoustic modeling for an audio environment. Background Technology
[0004] A three-dimensional (3D) model of a scene (e.g., a room, building, city, or any other type of indoor or outdoor space) can be generated using multiple images captured from different viewpoints within the scene. However, 3D models of scenes typically lack the functionality to include scene-related audio attributes. Summary of the Invention
[0005] Various embodiments of this disclosure relate to apparatuses, systems, methods, and computer-readable media for providing neuroacoustic modeling of audio environments. These features, as well as additional features, functions, and details of the various embodiments, are described below. The claims set forth herein further serve as an overview of this disclosure. Attached Figure Description
[0006] Some embodiments have thus been described in a general sense. Reference will now be made to the accompanying drawings, which are not necessarily drawn to scale, and in which:
[0007] Figure 1 An example audio signal processing system using neural rendering and inference is shown according to one or more embodiments disclosed herein;
[0008] Figure 2 An example neural rendering processing apparatus configured according to one or more embodiments disclosed herein is shown;
[0009] Figure 3 An example neural rendering model generation stream enabled by a neural rendering model generation engine according to one or more embodiments disclosed herein is shown;
[0010] Figure 4 An example neural rendering model inference flow enabled by a neural rendering model inference engine according to one or more embodiments disclosed herein is shown;
[0011] Figure 5AAn example audio environment is shown according to one or more embodiments disclosed herein;
[0012] Figure 5B Another example audio environment is shown according to one or more embodiments disclosed herein;
[0013] Figure 6 Example models are shown according to one or more embodiments disclosed herein;
[0014] Figure 7 An example neural rendering stream enabled by a neural rendering system according to one or more embodiments disclosed herein is shown;
[0015] Figure 8A An example neural rendering stream enabled by a neural rendering system according to one or more embodiments disclosed herein is shown;
[0016] Figure 8B Another example neural rendering flow enabled by a neural rendering system according to one or more embodiments disclosed herein is shown;
[0017] Figure 8C This illustrates yet another example of a neural rendering flow enabled by a neural rendering system according to one or more embodiments disclosed herein;
[0018] Figure 9 A training stream enabled by a neural rendering system is shown according to one or more embodiments disclosed herein;
[0019] Figure 10 An inference stream enabled by a neural rendering system is illustrated according to one or more embodiments disclosed herein;
[0020] Figure 11 Example methods for providing neuroacoustic modeling of an audio environment, according to one or more embodiments disclosed herein, are illustrated; and
[0021] Figure 12 Example methods for providing audio inference using neuroacoustic modeling for an audio environment, according to one or more embodiments disclosed herein, are shown. Detailed Implementation
[0022] Various embodiments of this disclosure will now be described more fully below with reference to the accompanying drawings, which illustrate some, but not all, of the embodiments of this disclosure. In fact, this disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements.
[0023] Overview
[0024] For certain types of audio systems, such as conference room audio systems, multiple audio capture devices can be positioned at different locations within the audio environment. However, it is often difficult to design, install, and / or configure audio capture devices within an audio environment to provide the desired acoustic performance. For example, it may be necessary to add or improve multimedia capabilities to a meeting space to provide real-time streaming of audio and / or video with minimal disruption. Therefore, technical experts may design, install, and / or configure audio / video equipment for real-time streaming by recording images and measurements of the meeting space, enabling the subsequent manual creation of drawings, blueprints, and / or acoustic performance estimates of the meeting space. This process involves manually converting images and measurements into blueprints, which can lead to inaccuracies and / or inefficiencies in the audio system. Furthermore, even with technical experts, accurately estimating the acoustic performance of a meeting space is challenging. Consequently, it is difficult to optimally arrange and / or configure audio capture devices within the meeting space to optimize the audio system's performance.
[0025] As discussed above, a three-dimensional (3D) model of a scene (e.g., a room, building, city, or any other type of indoor or outdoor space) can be generated using multiple images captured from different viewpoints within the scene. However, the synthesis of 3D models is typically based on the visual properties of the scene. Therefore, even using a typical 3D model of a scene, it is difficult to accurately estimate the acoustic characteristics and / or performance of the scene. Furthermore, typical audio capture devices and / or other audio equipment cannot measure or utilize a precise spatial and acoustic understanding of the audio environment. For example, different audio environments require different placements and configurations of audio capture devices and / or other audio equipment to provide the desired listening experience at different locations within the audio environment. However, sound transforms and propagates differently depending on the audio environment.
[0026] The various examples disclosed in this paper provide neuroacoustic modeling for audio environments. Neuroacoustic modeling can provide implicit representations of the physical and acoustic properties of an audio environment. In some examples, neuroacoustic modeling can be based on a set of impulse response samples and a geometric representation of the audio environment. Additionally, neuroacoustic modeling can provide rendering of the ambient acoustic properties, acoustic characteristics, and / or acoustic simulations of an audio environment. In some examples, neuroacoustic modeling can be utilized during acoustic rendering to provide acoustic simulations associated with source-receiver location pairs in the audio environment.
[0027] Exemplary neuroacoustic modeling systems and methods
[0028] Figure 1An audio signal processing system 100 configured to provide neuroacoustic modeling of an audio environment is illustrated according to embodiments of the present disclosure. The audio signal processing system 100 may be, for example, a conferencing system (e.g., a conference audio system, video conferencing system, digital conferencing system, etc.), an audio performance system, an audio recording system, a music performance system, a music recording system, a digital audio workstation, a lecture hall microphone system, a broadcast microphone system, an augmented reality system, a virtual reality system, an online gaming system, or another type of audio system. Additionally, the audio signal processing system 100 may be implemented as an audio signal processing device and / or as software configured to execute on a microphone, smartphone, laptop computer, personal computer, digital conferencing system, wireless conferencing unit, audio workstation device, augmented reality device, virtual reality device, recording device, headset, earphone, speaker, or other device. The audio signal processing system 100 disclosed herein may additionally or alternatively be integrated with other conferencing DSP processing into a virtual DSP processing system (e.g., DSP processing via a virtual processor or virtual machine).
[0029] The audio signal processing system 100 provides neuroacoustic modeling for the audio environment. Neuroacoustic modeling enables the rendering of the ambient acoustic properties and / or acoustic simulations of the audio environment. In some examples, neuroacoustic modeling provides the rendering of the ambient acoustic properties and / or acoustic simulations of the audio environment. In some examples, neuroacoustic modeling provides volumetric audio simulation and rendering. In some examples, neuroacoustic modeling provides accurate 3D audio simulation and rendering. In some examples, neuroacoustic modeling provides ray-traced neuroacoustic modeling for volumetric audio simulation and rendering of the audio environment. Additionally, neuroacoustic modeling can be trained using neural radiation fields and can be enhanced with the acoustic properties and / or material characteristics of the audio environment.
[0030] Neuroacoustic modeling can be provided via an audio rendering model. The audio rendering model can provide a rendering of the acoustic characteristics and / or acoustic properties of an audio environment. In some examples, the audio rendering model can be a neuroacoustic model. In other examples, the audio rendering model can be an audiovisual rendering model that combines the acoustic characteristics and / or acoustic properties of the audio environment to provide a rendering of the visual characteristics and / or acoustic properties of the audio environment. In some examples, the audio rendering model can be a neural radiation field (NeRF) model enhanced with audio coding or another type of 3D visualization model enhanced with audio coding. In some examples, the audio rendering model can be a multilayer perceptron (MLP) model containing one or more nonlinear activation functions. In some examples, the input to the MLP model can be a positional encoding (e.g., positional embedding, etc.). Additionally, the output of the MLP model can be based on the energy reflected and / or absorbed for each frequency band associated with a position in the audio environment corresponding to the positional encoding of the input.
[0031] In some examples, the audio rendering model can be configured to output acoustic properties associated with the audio environment. In some examples, photogrammetry can be used to estimate the relative 3D position (e.g., x, y, z coordinates) and orientation (e.g., pitch, yaw, and roll) of captured images. The captured images, along with the position and orientation data, can then be used to train the weights of the audio rendering model. The weights can encode scene-related spatial information. For example, the weights can be associated with: the corresponding position of a structure or object within the scene, a volumetric density representing the corresponding opacity of the structure or object, color information associated with the structure or object (e.g., RGB color), and / or other scene-related spatial data. It should be understood that color information can be associated with directionality relative to the audio environment. For example, the audio rendering model can model color information (e.g., RGB color) in various directions from a corresponding position in the audio environment, such that the color can change depending on the orientation. In some examples, color information can be used to determine the reflectivity or other properties of a structure or object. Once the audio rendering model is trained, the latent space represented by the trained weights of the audio rendering model can be queried to infer a view of the scene. Therefore, in addition to the visual properties of the audio environment, the audio rendering model can also be extended to include acoustic information related to the audio environment.
[0032] In some examples, the audio rendering model can be trained as a NeRF model or another 3D representation model augmented with audio encoding related to material properties. The augmented encoding can be used to capture the location of the sound source, properties associated with the propagation of sound within the audio environment, acoustic properties associated with specific objects within the audio environment, and / or other acoustic properties associated with the audio environment.
[0033] In some examples, an image of the audio environment can be incorporated to capture acoustic measurements of the audio environment. These acoustic measurements can then be correlated with a visual view associated with the image of the audio environment. In some examples, acoustic measurements and / or other acoustic data can be correlated with point cloud locations, voxel raster locations, mesh locations, or other 3D representative locations of the neural rendering model. In some examples, the neural rendering model can query a trained NeRF model or other 3D representations of the audio environment to infer the geometry of the audio environment relative to the propagation of sound within it. In some examples, a separate model (e.g., decoupled from the neural rendering model) can be trained to infer material properties of structures or objects that can be used to generate acoustic transfer functions, impulse responses, or other audio renderings of the audio environment.
[0034] Using the audio signal processing system 100, audio systems for audio environments can be optimally designed, installed, arranged, and / or configured. For example, the audio signal processing system 100 can provide automatic, enhanced, and / or strengthened configurations of audio capture devices for audio systems. The audio signal processing system 100 can also, alternatively, provide optimal arrangement of audio capture devices in an audio environment to provide optimal acoustic and / or computational performance of the audio system.
[0035] In some examples, the audio signal processing system 100 can be used to infer the location of an audio source and / or an acoustic transfer function that estimates the transformation of audio as it propagates through an audio environment. For example, an acoustic transfer function (e.g., impulse response) can be inferred from a trained neural rendering model associated with the audio signal processing system 100 to enable one or more downstream audio applications for the audio environment. In some examples, the trained neural rendering model associated with the audio signal processing system 100 can be queried to support the design and / or configuration of audio equipment within the audio environment.
[0036] In some examples, to provide an audio rendering model associated with the audio signal processing system 100, multiple audio capture devices can be positioned in the audio environment (e.g., in the case of intelligent and / or expert positioning where no audio capture devices are available in the audio environment) to construct the audio rendering model through a precise spatial and acoustic understanding of the audio environment. The audio rendering model associated with the audio signal processing system 100 can also allow for the adjustment of audio parameters to optimize the audio experience of the audio environment by utilizing a "mind's eye" view of the audio environment via the audio rendering model.
[0037] In addition, the audio signal processing system 100 can utilize neuroacoustic modeling to provide various improvements related to audio processing, such as: optimally controlling and / or configuring audio equipment in an audio environment, optimally arranging audio equipment in an audio environment, automatically tracking sound sources or emitted sounds in an audio environment, reducing noise in an audio environment, and / or improving one or more other audio processes related to the audio system in the audio environment.
[0038] The audio signal processing system 100 can also be adapted to produce an improved audio signal with reduced noise, reverberation, and / or other undesirable audio artifacts. In applications focused on noise reduction, this reduced noise can be stationary and / or non-stationary noise. Additionally, the audio signal processing system 100 can provide improved audio quality for audio signals in an audio environment. The audio environment can be an indoor environment, an outdoor environment, a room, a performance hall, a broadcast environment, a stadium or arena, a virtual environment, or another type of audio environment. In various examples, the audio signal processing system 100 can be configured to remove or suppress noise, reverberation, and / or other undesirable sounds from the audio signal via digital signal processing. The audio signal processing system 100 can alternatively be used for other types of sound enhancement applications, such as, but not limited to, active noise cancellation, adaptive noise cancellation, etc.
[0039] The audio signal processing system 100 includes one or more capture devices 102. The one or more capture devices 102 may be audio capture devices configured to capture audio from one or more sound sources. The one or more capture devices 102 may include one or more sensors configured to capture audio by converting sound into one or more electrical signals. The audio captured by the one or more capture devices 102 may also be converted into audio data 106. The audio data 106 may be digital audio data associated with one or more electrical signals, or alternatively, analog audio data.
[0040] One or more capture devices 102 may additionally or alternatively be corresponding video capture devices configured to capture video and / or images in relation to an audio environment. One or more capture devices 102 may include one or more sensors configured to capture video and / or images by converting light into one or more electrical signals. The video and / or images captured by one or more capture devices 102 may also be converted into video data 108. Video data 108 may be digital video data and / or digital image data associated with one or more electrical signals, or alternatively, analog video data and / or analog image data. It should be understood that video data 108 may be represented as one or more images (e.g., image data) captured by one or more capture devices 102.
[0041] In some examples, one or more capture devices 102 include one or more consumer cameras, smartphone cameras, 3D cameras, and / or other types of cameras. 3D cameras may include, but are not limited to: LiDAR, RADAR, inertial measurement units (IMUs), magnetic field sensors, accelerometers, gyroscopes, or other types of sensors capable of capturing video, images, and / or position. In some examples, position and / or orientation data provided by the 3D camera may be used to augment video data 108.
[0042] In the example, one or more capture devices 102 are one or more microphone arrays. For example, one or more capture devices 102 may correspond to one or more array microphones, one or more beamforming lobes of an array microphone, one or more linear array microphones, one or more ceiling array microphones, one or more desktop array microphones, or another type of array microphone. In alternative examples, one or more capture devices 102 are another type of capture device, such as, but not limited to, one or more condenser microphones, one or more microelectromechanical systems (MEMS) microphones, one or more dynamic microphones, one or more piezoelectric microphones, one or more virtual microphones, one or more network microphones, one or more ribbon microphones, and / or another type of microphone configured to capture audio. It should be understood that in some examples, one or more capture devices 102 may additionally or alternatively include one or more video capture devices, one or more image capture devices, one or more infrared capture devices, one or more sensor devices, and / or one or more other types of capture devices. Additionally, one or more capture devices 102 may be located within a specific audio environment.
[0043] The audio signal processing system 100 also includes a neuroacoustic modeling system 104. The neuroacoustic modeling system 104 can be configured to perform one or more modeling processes relative to audio data 106 and / or video data 108 to provide neuroacoustic modeling data 110. The neuroacoustic modeling system 104 can also be additionally or alternatively configured to perform one or more inference processes relative to a digital exploration request 109 to provide neuroacoustic modeling inference data 111.
[0044] Figure 1The neuroacoustic modeling system 104 depicted includes a model generation engine 112 and / or a model inference engine 113. The neuroacoustic modeling system 104 utilizes the model generation engine 112 to generate, train, and / or retrain an audio rendering model 105 using audio data 106 and / or video data 108. The audio rendering model 105 can provide a rendering of the acoustic characteristics and / or acoustic properties of an audio environment. In some examples, the audio rendering model 105 can be a neuroacoustic model for an audio environment. For example, the audio rendering model 105 can be an implicit representation of the physical and acoustic characteristics of the audio environment. In some examples, the audio rendering model 105 can be optimized based on the impulse response set and the geometric representation of the audio environment. In some examples, the audio rendering model 105 can be a machine learning model, such as a neural network model, a deep learning model, or another type of machine learning model.
[0045] In some examples, the audio rendering model 105 may be an audiovisual rendering model that combines the acoustic characteristics and / or acoustic properties of the audio environment to provide a rendering of the visual characteristics and / or acoustic properties of the audio environment. In some examples, the audio rendering model 105 may be an enhanced neural rendering model that includes a neural rendering volume representation of the audio environment enhanced with audio encoding. For example, the audio rendering model 105 may be a NeRF model (e.g., an enhanced NeRF model) or another type of 3D visualization model enhanced with audio encoding (e.g., an enhanced 3D model).
[0046] In some examples, the audio rendering model 105 may be an MLP model containing one or more non-linear activation functions. In some examples, the input to the MLP model may be a positional encoding (e.g., positional embedding, etc.). Additionally, the output of the MLP model may be based on the energy reflected and / or absorbed for each frequency band associated with a position in the audio environment corresponding to the positional encoding of the input.
[0047] The model generation engine 112 can utilize audio data 106 and / or video data 108 to generate a set of impulse responses associated with audio samples representing the acoustic properties of an audio environment. Additionally, the model generation engine 112 can determine and / or associate the locations of corresponding audio sources and corresponding audio receivers for the corresponding impulse responses in the impulse response set. The impulse response set can be the corresponding acoustic measurements between two audio sample locations in the audio environment. In some examples, the impulse response can be derived from the frequency response associated with the audio sample. In some examples, the model generation engine 112 can utilize audio data 106 and / or video data 108 to generate a set of images. The image set can contain multiple images, each associated with an audio sample representing the acoustic properties of the audio environment. In some examples, the information captured from the audio data 106 and / or video data 108 can include: volume density, light radiation, ambient audio volume density, reflected acoustic transfer function, transmitted acoustic transfer function, and / or other information.
[0048] In some examples, impulse responses from a set of impulse responses can be used to estimate impulse responses across locations of the audio source, and frequency responses can be based on impulse response cascades. In some examples, the impulse response can be estimated by evaluating a volumetric room density function between two locations in the audio environment. In some examples, the volumetric density function can utilize environmental measurements, such as, but not limited to, temperature, humidity, air velocity and direction, pressure, density, fluid flow direction and intensity, viscosity, shear force, elasticity, and / or other environmental measurements at the location. For example, the impulse response can be adjusted based on environmental measurements.
[0049] In some examples, the impulse response in the impulse response set can be calculated based on a digital transformation of the audio frequency response (e.g., a Fourier transform). The audio frequency response can be between any two locations in the audio environment and can be calculated based on the estimated acoustic transfer function (e.g., the estimated reflection and / or transmission acoustic transfer function) and / or the characteristic impedance of the air in the audio environment. In some examples, the audio frequency response can be represented by the following equation (1):
[0050] (1)
[0051] Where Z represents the characteristic acoustic impedance of air. Variables A, B, C, and D can be represented by the following equation (2):
[0052] (2)
[0053] Where S11 corresponds to the reflected audio transfer function at point 1, S12 corresponds to the transmitted audio transfer function between point 1 and point 2, S22 corresponds to the reflected audio transfer function at point 2, and S21 corresponds to the transmitted audio transfer function between point 2 and point 3.
[0054] Based on audio samples, model generation engine 112 can determine a set of camera properties including the relative position of the audio samples and the camera orientation associated with the audio samples. Then, model generation engine 112 can generate an audio rendering model 105 based on an impulse response set, an image set, audio samples, and / or the camera property set. Audio rendering model 105 can be a neural rendering volume representation of an audio environment enhanced with audio coding. For example, using audio rendering model 105, different viewpoints within the audio environment can be enhanced with audio coding, allowing the visual and acoustic properties of the audio environment to be provided via audio rendering model 105. In some examples, the visual and acoustic properties of the audio environment can be represented via the neural radiation field of audio rendering model 105. Therefore, audio rendering model 105 can be an implicit representation of the physical and acoustic characteristics of the audio environment.
[0055] Audio rendering model 105 can map audio characteristics to specific colors or color ranges. Audio characteristics may include, but are not limited to, amplitude, frequency, delay, crest factor, beamforming, and / or other types of audio characteristics. In some examples, audio rendering model 105 can map audio amplitudes to corresponding pixels in an image set. For example, an image can be converted from RGB color format to YUV format, and audio amplitudes can be mapped to the U and / or V components of the YUV format. After audio encoding, the YUV format image can be converted back to RGB color format. It should be understood that audio rendering model 105 can alternatively map audio amplitudes to different types of color space formats.
[0056] In some examples, the neuroacoustic modeling data 110 may include data representing visual attributes and audio characteristics of the audio environment. For example, the neuroacoustic modeling data 110 may include 3D model data associated with visual attributes of the audio environment. The 3D model data may include data associated with the 3D position, viewing orientation, color data, density data, camera pose data, and / or other visual attributes of the audio environment. In some examples, the 3D model data may be NeRF model data associated with one or more neural radiation fields of the audio environment. Additionally, the neuroacoustic modeling data 110 may include audio data associated with audio characteristics of the audio environment. For example, the audio data may include audio encoding (e.g., encoded audio characteristics) associated with the audio environment. In some examples, the audio data may include one or more impulse responses associated with corresponding locations within the audio environment.
[0057] The neuroacoustic modeling system 104 additionally or alternatively utilizes the model inference engine 113 to provide one or more inferences, predictions, and / or insights about an audio environment using the audio rendering model 105. For example, the audio rendering model 105 can be used to provide one or more inferences, predictions, and / or insights related to the visual and / or acoustic properties of the audio environment. The neuroacoustic modeling inference data 111 may contain one or more inferences, predictions, and / or insights about the audio environment. In some examples, the neuroacoustic modeling inference data 111 may contain one or more impulse responses provided by the audio rendering model 105. In some examples, the impulse responses contained in the neuroacoustic modeling inference data 111 may be room impulse responses (RIR) or acoustic impulse responses (AIR) representing sound waves, reflections, reverberation, echoes, and / or other acoustic characteristics of the audio environment. The neuroacoustic modeling inference data 111 may additionally or alternatively contain the locations of one or more candidate audio components associated with the audio environment, one or more tuning settings for one or more audio devices in the audio environment, inferred acoustic properties of one or more surfaces in the audio environment, and / or one or more other inferences associated with the audio environment.
[0058] In some examples, the model inference engine 113 can output one or more candidate audio component locations associated with the audio environment. One or more candidate audio component locations can be generated based on the audio rendering model 105. The one or more candidate audio component locations can each represent a candidate location within the audio environment of an audio component (e.g., an audio capture device, an audio output device, or another type of audio component).
[0059] In some examples, inference, prediction, and / or insights about the audio environment can enable streamlined workflows (e.g., accurate 3D acoustic modeling or visual / acoustic modeling of the audio environment). In some examples, inference, prediction, and / or insights about the audio environment can enable optimized recommendations for placing audio devices and / or equipment within the audio environment.
[0060] In some examples, for a constellation of camera, speaker, and / or microphone devices installed in multiple locations within an audio environment, the constellation of devices can automatically determine the 3D shape of the audio environment, where the devices are located within the audio environment, how the devices are oriented within the audio environment, and / or how the devices should be optimally configured. Therefore, sound within the audio environment can be output and / or captured for listeners and speakers within the audio environment at the desired level and / or quality.
[0061] In some examples, one or more impulse responses provided by audio rendering model 105 can be used to improve the placement of audio devices within an audio environment. The impulse responses provided by audio rendering model 105 can be room impulse response (RIR) or acoustic impulse response (AIR) representing the sound waves, reflections, reverberation, echoes, and / or other acoustic characteristics of source-receiver location pairs in the audio environment. One or more impulse responses provided by audio rendering model 105 can be used to accurately simulate the acoustic characteristics of the audio environment to achieve realistic sound reproduction and / or audio analysis. In some examples, one or more impulse responses provided by audio rendering model 105 can be used to determine locations within the audio environment where audio sources and / or receivers are optimally located and / or oriented such that sound sources and / or reflections in the audio environment are captured in an optimal manner.
[0062] In some examples, inferences, predictions, and / or insights about the audio environment can enable the generation of additional image and / or audio data to model the audio environment for use in simulating and rendering audio (e.g., for scenarios where there are not enough devices in the constellation to capture a complete representation of the audio environment).
[0063] In some examples, audio rendering model 105 can be used in the acoustic rendering process of an audio environment to render the impulse response of source-receiver location pairs in the audio environment. In some examples, the rendered impulse response can be RIR or AIR. The source-receiver location pairs can be associated with a specific audio source and a specific capture device in the audio environment. In some examples, audio rendering model 105 can implement a rendering process that utilizes the geometric representation of the audio environment to render the impulse response (RIR, AIR, etc.) of the source-receiver pair. Therefore, audio rendering model 105 can implement acoustic simulation for one or more source-receiver location pairs within the audio environment. In some examples, the source-receiver location pair associated with the rendered impulse response can be a new source-receiver location pair not included in the training dataset of audio rendering model 105. Therefore, audio rendering model 105 can render the impulse response from a previously unknown source-receiver location pair in the audio environment.
[0064] In some examples, one or more portions of the neuroacoustic modeling inference data 111 may be output via an audio output device and / or a display output device. For example, one or more portions of the neuroacoustic modeling inference data 111 may be output via an audio mixer device, a DSP processing device, a smartphone, tablet computer, laptop computer, personal computer, audio workstation device, wearable device, augmented reality device, virtual reality device, recording device, microphone, headset, earphone, speaker, haptic device, or another type of output device. In some examples, one or more portions of the neuroacoustic modeling inference data 111 may be rendered via a user interface (e.g., an electronic interface) of a user device, such as a smartphone, tablet computer, laptop computer, personal computer, audio workstation device, touch controller device, augmented reality device, virtual reality device, or another type of user device.
[0065] In some examples, acoustic simulation can be used to support one or more downstream systems, such as, but not limited to, systems for determining the placement of audio devices in an audio environment, systems for tuning audio devices in an audio environment, systems for designing and / or installing audio systems in an audio environment, systems for designing architectural structures for an audio environment, virtual reality systems, and / or one or more other downstream systems relative to the neuroacoustic modeling system 104.
[0066] In some examples, inferences, predictions, and / or insights about the audio environment can enhance augmented reality, virtual reality, metaverse applications, and / or video game applications to model 3D environments and render realistic 3D audio for immersive participants in 3D space, so that listeners wearing augmented reality or virtual reality devices in the simulated space will hear acoustically realistic sounds in the simulated space. For example, audio rendering model 105 can be used to render audio in a simulated manner to represent how sound will propagate from the audio source location to the audio receiver in the audio environment. In some examples, impulse responses (RIR, AIR, etc.) provided by audio rendering model 105 can be used to apply audio effects to the audio signal to simulate how the audio will sound if it is transmitted between the audio source location and the audio receiver in the audio environment. In some examples, the simulated audio can be uniquely simulated as binaural audio for the different ears of a listener at the audio receiver location. Therefore, audio rendering model 105 can be used to achieve an immersive audio experience for audio environments associated with augmented reality, virtual reality, metaverse applications, and / or video game applications.
[0067] In some examples, inferences, predictions and / or insights about the audio environment can be encapsulated and / or encoded via an interchangeable format (e.g., a generic scene description (USD) format (e.g., OpenUSD) or another type of interchangeable format) to transform environmental information (e.g., 3D geometry, color, material information, etc.) into specific technological applications (e.g., augmented reality, virtual reality, metaverse, etc.) for immersive collaboration, gaming, etc.
[0068] In some examples, one or more impulse responses provided by the audio rendering model 105 can be used to configure and / or tune one or more audio settings of one or more audio devices located within an audio environment. For example, one or more impulse responses provided by the audio rendering model 105 can be used to configure and / or tune receiver devices (e.g., microphones, etc.) located within an audio environment. Alternatively or additionally, one or more impulse responses provided by the audio rendering model 105 can be used to configure and / or tune sound source devices (e.g., speakers, etc.) located within an audio environment. In some examples, one or more impulse responses provided by the audio rendering model 105 can be used to configure and / or tune digital signal processing of audio captured by the audio devices.
[0069] In some examples, the directional sensitivity of the audio device can be adjusted based on one or more impulse responses provided by the audio rendering model 105 to optimally capture the audio source directly while minimizing unwanted sounds such as reflections and echoes. Alternatively, the equalization of the audio device can be adjusted based on one or more impulse responses provided by the audio rendering model 105 to enhance or attenuate one or more audio frequencies captured by the audio device to adapt to the influence of specific audio characteristics of the audio environment. Alternatively, the equalization of the sound source device can be adjusted based on one or more impulse responses provided by the audio rendering model 105 to modify the frequency response associated with the audio environment. Alternatively, the reverberation of the audio device can be adjusted based on one or more impulse responses provided by the audio rendering model 105 to provide improved dereverberation and / or reduce the effects of audio reflections in the audio environment.
[0070] In some examples, one or more impulse responses provided by the audio rendering model 105 can be used for room or building design associated with an audio environment. For example, room acoustics software applications can utilize one or more impulse responses provided by the audio rendering model 105 to provide room acoustic simulations and / or measurements to enable the construction and / or configuration of a room or building associated with an audio environment. Thus, regions of interest in a room or building can be identified and / or the impulse responses can be used to inform the construction or treatment of the room or building. In some examples, the placement of acoustic panels and / or recommended materials in a room or building can be determined based on one or more impulse responses provided by the audio rendering model 105.
[0071] In some examples, one or more impulse responses provided by the audio rendering model 105 can be used to infer the acoustic properties of one or more surfaces within the audio environment. Alternatively, based on the inferred acoustic properties of one or more surfaces within the audio environment, the material type of one or more surfaces can be inferred. The material type can include specific materials, such as, but not limited to, plaster, wood, metal, etc. In some examples, the inferred acoustic properties of one or more surfaces within the audio environment can be used to query a lookup table or database associated with the mapping between acoustic properties and material types. In some examples, the lookup table or database can be associated with a mapping between audio absorption characteristics (e.g., absorption coefficient, etc.) and material types. Therefore, instead of pre-assigning a specific material type (e.g., by the user) to surfaces or objects within the environment model, one or more impulse responses provided by the audio rendering model 105 can be used to infer the material type of one or more surfaces or objects within the audio environment.
[0072] In some examples, inferences, predictions, and / or insights about the audio environment may be enhanced additionally or alternatively by: modeling the architecture of the room and the effects of altered surfaces; generating datasets of room types and using them as input for training other models; training models to enhance déreverberation in specific rooms and geometries; modeling of directional sound sources (e.g., speakers) and / or sensors; data collection for training neural representations; and / or one or more other technical improvements related to audio processing.
[0073] In some examples, the model inference engine 113 receives a digital exploration request 109. The digital exploration request 109 may be a request to digitally explore the visual and / or audio properties of an audio environment. The digital exploration request 109 may also include an audio environment identifier for the audio environment. The audio environment identifier may be a numeric code, a bit string, an alphanumeric string, or another type of identifier that identifies the audio environment. For example, the audio environment may be referred to as “Meeting Room A,” and the audio environment identifier may be a numeric code, a bit string, an alphanumeric string, or another type of identifier corresponding to the phrase “Meeting Room A.” In some examples, the digital exploration request 109 may additionally or alternatively include information related to the number of audio devices, the location of candidate audio devices, camera location, camera angle, X, Y, Z position, pitch information, roll information, yaw information, viewpoint-related image, spherical coordinate information, azimuth information, elevation information, and / or other information related to the audio environment to facilitate the digital exploration of the visual and / or audio properties of the audio environment.
[0074] In some examples, a digital exploration request 109 is received from a user device, such as a smartphone, tablet, laptop, personal computer, audio workstation, touch controller, augmented reality, virtual reality, or another type of user device. In some examples, the digital exploration request 109 is generated via a user interface on the user device's display. The user interface may be a graphical user interface for designing applications, installing applications, configuring audio devices, audio processing applications, system configuration applications, or another type of application associated with a software platform.
[0075] Based on the audio environment identifier, the model inference engine 113 can determine the audio rendering model (e.g., audio rendering model 105) associated with the audio environment. For example, the model inference engine 113 can associate the audio environment identifier with the audio rendering model associated with the audio environment (e.g., audio rendering model 105). In some examples, the audio rendering model is selected from a set of audio rendering models, wherein the corresponding audio rendering model in the set of audio rendering models is previously generated for a specific audio environment.
[0076] Therefore, compared to traditional modeling techniques, the neuroacoustic modeling system 104 can provide improved modeling and / or inference about the audio environment. Additionally, the accuracy of sound source localization in the audio environment can be improved by employing the neuroacoustic modeling system 104. The neuroacoustic modeling system 104 can be additionally or alternatively adapted to generate improved audio signals with reduced noise, reverberation, and / or other undesirable audio artifacts, even considering stringent audio delay requirements. For example, the neuroacoustic modeling system 104 can optimize the configuration and / or location of audio capture devices in the audio environment to generate improved audio signals with reduced noise, reverberation, and / or other undesirable audio artifacts. Therefore, audio can be provided to the user without undesirable sound reflections. The neuroacoustic modeling system 104 can also improve the operational efficiency of denoising, dereverberation, and / or other audio filtering while also optimizing the acoustic performance of the audio system.
[0077] Compared to conventional audio processing systems, the neuroacoustic modeling system 104 can also employ fewer computational resources. Alternatively, in one or more examples, the neuroacoustic modeling system 104 can be configured to allocate a smaller amount of memory resources for denoising, dereverberation, and / or other audio processing of the audio signal. In yet another example, the neuroacoustic modeling system 104 can be configured to improve the processing speed of modeling operations, denoising operations, dereverberation operations, and / or audio processing operations. These improvements can enable improved audio processing systems and / or hardware / software configurations in audio environments where high-fidelity audio is desired and / or processing efficiency is critical.
[0078] Figure 2 An example neuroacoustic modeling device 202 configured according to one or more embodiments of the present disclosure is shown. The neuroacoustic modeling device 202 can be configured to perform... Figure 1 One or more technologies described herein and / or one or more other technologies described herein.
[0079] The neuroacoustic modeling device 202 may be a computing system communicatively coupled to one or more circuit modules associated with audio processing. The neuroacoustic modeling device 202 may include a processor 204, a memory 206, a model generation circuit system 208, a model inference circuit system 210, an input / output circuit system 212, and / or a communication circuit system 214, or otherwise communicate with the processor 204, memory 206, model generation circuit system 208, model inference circuit system 210, input / output circuit system 212, and / or communication circuit system 214. In some examples, the processor 204 (which may include multiple or coprocessors or any other processing circuitry associated with the processor) may communicate with the memory 206.
[0080] Memory 206 may include a non-transitory memory circuitry and may include one or more volatile and / or non-volatile memories. In some examples, memory 206 may be an electronic storage device (e.g., a computer-readable storage medium) configured to store data retrievable by processor 204. In some examples, the data stored in memory 206 may include audio data, stereo audio signal data, mono audio signal data, radio frequency signal data, video data, image data, training data for an audio rendering model, a set of weights for an audio rendering model, a set of trained audio rendering models, etc., to enable the neuroacoustic modeling device 202 to perform various functions or methods according to embodiments of the present disclosure described herein.
[0081] In some examples, processor 204 can be embodied in a variety of different ways. For example, processor 204 can be embodied as one or more of a variety of hardware processing components, such as a central processing unit (CPU), microprocessor, coprocessor, DSP, field-programmable gate array (FPGA), neural processing unit (NPU), graphics processing unit (GPU), system-on-a-chip (SoC), cloud server processing element, controller, or processing element with or without an accompanying DSP. Processor 204 can also be embodied in a variety of other processing circuitry systems, including integrated circuits such as microcontroller units (MCUs), ASICs (Application-Specific Integrated Circuits), hardware accelerators, cloud computing chips, or application-specific electronic chips. Furthermore, in some examples, processor 204 can include one or more processing cores configured to execute independently. Multi-core processors can implement multiprocessing within a single physical package. Alternatively or additionally, processor 204 can include one or more processors configured via a bus to enable independent execution of instructions, pipeline processing, and / or multi-threaded processing.
[0082] In some examples, processor 204 may be configured to execute instructions, such as computer program code or instructions, stored in memory 206 or otherwise accessible to processor 204. Alternatively or additionally, processor 204 may be configured to perform hard-coded functionality. Thus, whether configured by hardware or software instructions, or by a combination thereof, processor 204 may represent a computational entity (e.g., physically embodied in a circuit system) configured to perform operations according to embodiments of the present disclosure described herein. For example, when processor 204 is embodied as a CPU, DSP, ARM, FPGA, ASIC, or similar device, the processor may be configured as hardware for performing operations of embodiments of the present disclosure. Alternatively, when processor 204 is embodied to execute software or computer program instructions, the instructions may specifically configure processor 204 to perform the algorithms and / or operations described herein when executing the instructions. However, in some examples, processor 204 may be a processor specifically configured to employ means of embodiments of the present disclosure by further configuring the processor using instructions for performing the algorithms and / or operations described herein. In addition, the processor 204 may further include a clock, an arithmetic logic unit (ALU), and logic gates configured to support the operation of the processor 204.
[0083] In one or more examples, the neuroacoustic modeling device 202 may include a model generation circuitry system 208. The model generation circuitry system 208 may be any component embodied in hardware or a combination of hardware and software, configured to perform one or more functions disclosed herein related to the model generation engine 112. In one or more examples, the neuroacoustic modeling device 202 may include a model inference circuitry system 210. The model inference circuitry system 210 may be any component embodied in hardware or a combination of hardware and software, configured to perform one or more functions disclosed herein related to the model inference engine 113.
[0084] In some examples, the neuroacoustic modeling device 202 may include an input / output circuitry system 212 that can communicate with the processor 204 to provide output to a user and, in some examples, receive indications of user input. The input / output circuitry system 212 may include a user interface and may include a display. In some examples, the input / output circuitry system 212 may also include a keyboard, touchscreen, touch area, softkeys, buttons, knobs, or other input / output mechanisms.
[0085] In some examples, the neuroacoustic modeling device 202 may include a communication circuitry system 214. The communication circuitry system 214 may be any component embodied in hardware or a combination of hardware and software, configured to receive data from a network and / or to any other means or module communicating with the neuroacoustic modeling device 202 and / or to the network and / or to any other means or module communicating with the neuroacoustic modeling device 202. In this regard, the communication circuitry system 214 may include, for example, an antenna or one or more other communication means for enabling communication with a wired or wireless communication network. For example, the communication circuitry system 214 may include an antenna, one or more network interface cards, a bus, a switch, a router, a modem, and supporting hardware and / or software, or any other means suitable for enabling communication via a network. Additionally or alternatively, the communication circuitry system 214 may include circuitry for interacting with the antenna to transmit signals via the antenna or for processing the reception of signals received via the antenna.
[0086] Figure 3 This illustration shows one or more embodiments of a device for use by... Figure 1 The model generation engine 112 enables neural rendering model generation stream 300. Neural rendering model generation stream 300 includes audio-enhanced image generation 302, camera property generation 304, and audio rendering model generation 306.
[0087] Audio-enhanced image generation 302 utilizes audio data 106 and / or video data 108 to generate audio sample data 308. Audio sample data 308 may contain a set of impulse responses for each source-listener audio device pair within the audio environment. Audio-enhanced image generation 302 additionally or alternatively utilizes video data 108 to generate image data 310. Based on image data 310, camera property generation 304 generates camera property data 312. In some examples, camera property generation 304 may infer the position of one or more capture devices 102. In some examples, the position of one or more capture devices 102 may correspond to a camera position provided by a photogrammetric process associated with camera property generation 304. Additionally, based on audio sample data 308, image data 310, and / or camera property data 312, audio rendering model generation 306 may generate neuroacoustic modeling data 110. In some examples, camera property data 312 contains sensor information related to the position and / or orientation of one or more capture devices 102 and / or one or more audio sources.
[0088] Figure 4 This illustration shows one or more embodiments of a device for use by... Figure 1The model inference engine 113 enables neural rendering model inference stream 400. Neural rendering model inference stream 400 includes neuroacoustic modeling inference 402.
[0089] Neuroacoustic modeling inference 402 utilizes digital exploration request 109 to generate neuroacoustic modeling inference data 111. In some examples, neuroacoustic modeling inference 402 determines an audio environment identifier 404 contained in digital exploration request 109. Based on audio environment identifier 404, neuroacoustic modeling inference 402 can query a model data storage area 406 containing a set of audio rendering models 105a-n. Model data storage area 406 may be integrated into or communicatively coupled to a cloud platform (e.g., a server system) or a user device, such as a smartphone, tablet computer, laptop computer, personal computer, audio workstation device, touch controller device, augmented reality device, virtual reality device, or another type of user. In some examples, the corresponding audio rendering models in the set of audio rendering models 105a-n are trained via a cloud platform. In some examples, the corresponding audio rendering models in the set of audio rendering models 105a-n are trained via a user device. Therefore, training of the corresponding audio rendering models in the set of audio rendering models 105a-n can be performed on the same or different devices relative to the corresponding audio rendering models in the set of audio rendering models 105a-n stored via the model data storage area 406.
[0090] In some examples, neuroacoustic modeling inference 402 can associate an audio environment identifier 404 with an audio rendering model 105 from the set of audio rendering models 105a-n. The audio rendering model 105 can be used by neuroacoustic modeling inference 402 to generate neuroacoustic modeling inference data 111. In some examples, the neuroacoustic modeling inference data 111 contains one or more audio inferences associated with the audio environment. In some examples, the neuroacoustic modeling inference data 111 may contain inferred values related to volume density, radiation, ambient audio volume density, reflected acoustic transfer function, transmitted acoustic transfer function, other inferred acoustic information, inferred material properties of structures or objects within the audio environment, and / or other information. In some examples, the inferred material properties may include inferred absorption characteristics (e.g., inferred absorption coefficients, etc.) of the corresponding material type associated with structures or objects within the audio environment. In some examples, the response or sensitivity of a user device and / or audio equipment such as a microphone or speaker can be modified in real time based on the determined location within the audio environment corresponding to the potential coding contained in the audio rendering model 105, inferred from the neuroacoustic modeling data 111.
[0091] Figure 5AAn example audio environment 502 according to one or more embodiments of the present disclosure is shown. Audio environment 502 may be an indoor environment, an outdoor environment, a room, a conference room, a meeting hall, an auditorium, a performance hall, a broadcasting environment, an arena (e.g., a stadium), a virtual environment, or another type of audio environment. Audio environment 502 includes a capture device 102a capable of capturing audio from an audio source 504. In some examples, capture device 102a is configured as an audio capture device, such as a microphone or microphone array. In some examples, capture device 102a is configured as a video capture device capable of capturing video data in addition to audio data.
[0092] In some examples, audio is reflected throughout the audio environment 502 to facilitate the generation of neural rendering models. For example, audio source 504 may emit audio (e.g., a broad-spectrum chirp, a sine sweep, etc.), and capture device 102a may be a directional microphone located near audio source 504 to capture the reflected audio. The audio emitted by audio source 504 may contain a frequency range corresponding to the desired frequency range of the neuroacoustic modeling inference data 111. For example, if it is desired to optimize the audio environment for human speech, the audio emitted by audio source 504 may contain a frequency range corresponding to 300 Hz to 3 kHz associated with human speech. In some examples, audio may be reflected from near-wall 506 and far-wall 508 of the audio environment before being captured by capture device 102a. However, it should be understood that audio emitted in audio environment 502 may be transmitted, reflected, refracted, diffracted, absorbed, and / or scattered within the audio environment before being captured by capture device 102a. Modeling of audio emitted in audio environment 502 can relate to reflection, transmission, refraction, diffraction, Doppler effect, resonance and / or one or more other acoustic phenomena.
[0093] Figure 5B An example audio environment 552 according to one or more embodiments of the present disclosure is shown. The audio environment 552 may be an indoor environment, an outdoor environment, a room, a conference room, a meeting hall, an auditorium, a performance hall, a broadcasting environment, an arena (e.g., a stadium), a virtual environment, or another type of audio environment. The audio environment 552 includes a capture device 102a capable of capturing audio from an audio source 554. In some examples, the capture device 102a is configured as an audio capture device, such as a microphone or microphone array. In some examples, the capture device 102a is configured as a video capture device capable of capturing video data in addition to audio data.
[0094] In some examples, audio is reflected throughout the audio environment 552 to facilitate neuroacoustic modeling for ray tracing. In some examples, ray direction sampling from audio source 554 can be provided to facilitate neuroacoustic modeling for ray tracing. Audio source 554 can emit audio rays from a source location. The audio rays can be broad-spectrum chirps, sinusoidal scans, or another type of audio. In some examples, the capture device 102a can sample the audio rays uniformly on a sphere centered at the source location associated with audio source 554. Alternatively, audio rays can be traced through the audio environment 552, where the corresponding energy of the audio rays is attenuated upon reflection from surfaces such as surface 555 or surface 556. In some examples, the model generation engine 112 can generate an energy histogram associated with the audio rays captured by the capture device 102a to facilitate the rendering process for the audio environment 502.
[0095] In some examples, the energy of a sampled audio ray can be attenuated based on the directionality of the audio source 554. For example, if the audio ray is sampled in the direction opposite to the directionality of the audio source 554, the calculated energy of the audio ray can be reduced. Alternatively, the directionality of the audio source 554 can be modeled by sampling the audio ray along the direction of the audio source 554. In some examples, the directionality of the capture device 102a (e.g., the directionality of the receiver of the capture device 102a) can be similarly modeled such that the energy of intersecting audio rays is reduced based on the direction of the audio ray relative to the direction of the pickup pattern used to capture the audio associated with the audio source 504.
[0096] Figure 6A model 600 according to one or more embodiments of the present disclosure is illustrated. In some examples, the audio rendering model is a machine learning model, an audio rendering model, an enhanced neural rendering model, an enhanced NeRF model, or another type of model. Model 600 may receive 3D coordinates 608 and a camera view direction 614 as input. The 3D coordinates 608 may include x, y, and z coordinates for a corresponding spatial location in the audio environment. The camera view direction 614 may include a corresponding view direction for a corresponding image captured in the audio environment. In some examples, the 3D coordinates 608 and the camera view direction 614 may be provided to a stack of fully connected linear layers separated by nonlinear activation functions. However, it should be understood that model 600 may additionally or alternatively include one or more other types of layers and / or functions. Based on the processing of the 3D coordinates 608 and the camera view direction 614 via fully connected linear layers and nonlinear activation functions, model 600 may output a volume density 610 and a view-dependent RGB color 616. The volume density 610 may indicate the amount of radiance and / or luminance associated with the corresponding 3D coordinates 608. The view-dependent RGB color 616 can indicate the RGB color information associated with the corresponding 3D coordinate 608. In some examples, model 600 can be an MLP model.
[0097] In some examples, model 600 includes a high-dimensional spatial mapping 602, layer 604, and layer 606. Layer 604 may be a 256-channel layer, and layer 606 may be a 128-channel layer. However, it should be understood that in some examples, layer 604 and / or layer 606 may contain different numbers of channels. Additionally, layer 604 and / or layer 606 may be neural network layers, fully connected layers, convolutional layers, and / or another type of layer that models a volumetric representation of the audio environment. The high-dimensional spatial mapping 602 may utilize position encoding to map 3D coordinates 608 into a high-dimensional space for processing by layer 604. The 3D coordinates 608 may represent the x, y, and z coordinate positions within the audio environment. Layer 604 may utilize the higher-dimensional 3D coordinates 608 to generate a volume density 610 and a feature vector 612. The volume density 610 may represent the differential probability that a ray terminates at 3D coordinates 608. Feature vector 612 may contain: encoding parameters, spatial coordinates, color information, radiometric information, reflectivity information, viewpoint information, density information, and / or other information related to the corresponding 3D coordinates in the audio environment. Feature vector 612 may be a 256-dimensional feature vector or a feature vector with different dimensions. Layer 606 may utilize feature vector 612 and camera viewing direction 614 to generate view-dependent RGB color 616. In some examples, volume density 610 and / or view-dependent RGB color 616 may be enhanced using audio encoding associated with audio data captured in the audio environment (e.g., audio data 106).
[0098] Figure 7 This illustration shows one or more embodiments of a device for use by... Figure 1 The neuroacoustic modeling system 104 enables neural rendering with a neural rendering stream 700. The neural rendering stream 700 includes a path 702 for creating a trained neural network and a path 704 for creating instructions from survey images for use by the installers of the audio system in the audio environment.
[0099] Using path 702, the simulation 706 of the audio environment generates an image 708 with parameters. Simulation 706 also generates an acoustic impulse response 710. Image 708 can be used for training 712 of the audio rendering model 714. In some examples, image 708 can undergo segmentation and / or classification 716 to facilitate the training 712 of the audio rendering model 714. In some examples, a domain expert 718 can facilitate supervised learning about image 708. The trained version of the audio rendering model 714 can associate the NeRF latent image with impulse responses, object categories, object placement, and / or object labels.
[0100] Using path 704, an image 720 of the audio environment can undergo photogrammetry 722 to determine the camera position and / or camera parameters associated with the image 720. In some examples, photogrammetry 722 can be used to estimate the relative 3D position (e.g., x, y, z coordinates) and orientation (e.g., pitch, yaw, and roll) of the image 720. Weights of an audio rendering model 714, associated with the volumetric probability density representing the corresponding location of structures within the audio environment, can then be trained via training 724 using the image 720 along with the position and orientation data. In some examples, the camera position and / or camera parameters can be used to train 724 of the audio rendering model 714 with a latent NeRF image 726. The latent NeRF image 726 can correspond to weights that represent the inherent encoding of the volumetric density of the audio environment (e.g., weights used for the audio rendering model 714). In some examples, the trained version of the audio rendering model 714 can be used to provide one or more inferences 728 associated with the audio environment.
[0101] Figure 8A This illustration shows one or more embodiments of a device for use by... Figure 1The neuroacoustic modeling system 104 enables a neural rendering stream 800 for ray-traced neuroacoustic modeling. The neural rendering stream 800 can utilize specular reflection to provide ray-traced neuroacoustic modeling. Audio rays generated by audio source 804 can reflect from a surface as they are traced through audio environment 802. When audio ray 810 intersects surface 812 via specular reflection, a new audio ray 814 with a reflection angle equal to the incident angle of audio ray 810 can be generated. Additionally, the energy of the new audio ray 814 can be based on the audio ray 810 at the impact location on surface 812, attenuated by the acoustic coefficient 821 of surface 812. In some examples, when audio ray 810 intersects surface 812 via diffuse reflection (e.g., scattering), the new audio ray 814 can be multiple new audio rays. For example, multiple new audio rays can be sampled from a hemispherical or conical shape along the reflection angle of audio ray 810.
[0102] In some examples, the acoustic coefficient 821 can be predicted by neural network processing 820 provided by model generation engine 112. For example, neural network processing 820 can utilize one or more MLP techniques to predict the acoustic coefficient 821. In some examples, neural network processing 820 receives 3D coordinates 822 as input. The 3D coordinates 822 can correspond to the impact position of the audio ray 810 on the surface 812. In some examples, the specular reflection coefficient 823 can be calculated based on the acoustic coefficient 821 via new ray generation 824 provided by model generation engine 112. The specular reflection coefficient 823 can be a specular reflection coefficient for a specular reflection scene associated with the audio ray 810. Alternatively, the specular reflection coefficient 823 can be a diffuse reflection coefficient for a diffuse reflection scene associated with the audio ray 810. In some examples, based on the specular reflection coefficient 823 and the energy 825 of the audio ray 810 at the impact position on the surface 812, the remaining energy after reflection can be predicted to generate a new ray 814. In some examples, the residual energy after reflection can be additionally predicted based on the incident angle 827 of the audio ray 810. For example, the reflected ray angle of the new ray 814 can be predicted based on the incident angle 827 and surface information of surface 812. Surface information can be determined based on the 3D geometry of the audio environment 802. For example, the 3D model of the audio environment 802 (e.g., NeRF, point cloud, mesh, voxel grid, etc.) can be queried to determine surface information, such as the surface normal of surface 812.
[0103] Figure 8B This illustration shows one or more embodiments of a device for use by... Figure 1The neuroacoustic modeling system 104 enables a neural rendering stream 830 for ray tracing-enabled neuroacoustic modeling. The neural rendering stream 830 can generate a set of impulse responses associated with audio rays in the audio environment 802. For example, a capture device 102a can capture a new audio ray 814. Additionally, a model generation engine 112 can determine the energy associated with the new audio ray 814 and add the energy associated with the new audio ray 814 to an energy histogram of the audio environment 802 via the neural rendering stream 850. The model generation engine 112 can also convert the energy histogram into a set of impulse responses associated with the audio environment 802. In some examples, the energy histogram may contain energy associated with one or more other audio rays (e.g., audio ray 834 reflected from a surface 832 of the audio environment 802). In some examples, the audio ray 834 may also undergo processing by a neural network processing 820 and / or a new ray generation 824. In some examples, the audio ray 834 may be the result of diffuse reflection associated with audio ray 810.
[0104] Figure 8C This illustration shows one or more embodiments of a device for use by... Figure 1The neuroacoustic modeling system 104 enables a neural rendering stream 850 for ray tracing-based neuroacoustic modeling. The neural rendering stream 850 can be associated with the ray-receiver intersection point between a new audio ray 814 and the capture device 102a. For example, the new audio ray 814 can intersect with the capture device 102a, and the remaining ray energy 854 can be added to the energy histogram 860 of the audio environment 802. In some examples, the value of the ray energy 854 can be modified based on the total ray length 852 of the new audio ray 814 to provide ray energy after energy loss due to the ray travel distance to the energy histogram 860. In some examples, the value of the ray energy 854 can also be modified based on the direction of arrival relative to the capture device 102a associated with the new audio ray 814. In some examples, the total ray length 852 can be used, additionally or alternatively, to calculate the total ray travel time of the new audio ray 814. In some examples, the total ray travel time of the new audio ray 814 can provide a time delay for the new audio ray 814, and this time delay can be additionally recorded in the energy histogram 860. In some examples, the energy histogram 860 can be a 2D histogram with intervals configured over the time delay and / or frequency range. In some examples, environmental conditions of the audio environment (e.g., temperature, humidity, etc.) may affect the travel time of the new audio ray 814 due to the influence of the speed of sound. Therefore, the remaining ray energy 854 can be added to the energy histogram 860 based on the environmental conditions and / or the associated speed of sound. In some examples, the time delay of the new audio ray 814 can be determined based on the environmental conditions. In some examples, the energy histogram 860 can be converted into an impulse response set of the audio environment 802. In some examples, the impulse response set can be synthesized via a Poisson distributed noise process.
[0105] Figure 9 The invention illustrates one or more embodiments of the invention. Figure 1 The neuroacoustic modeling system 104 enables a training stream 900. The training stream 900 includes a rendering process 902. The rendering process 902 can predict energy histograms based on source and receiver locations. In some examples, source and receiver locations are determined based on an impulse response set 904 associated with the audio environment. In some examples, the rendering process 902 utilizes a 3D environment model 906. The 3D environment model 906 can be a NeRF model, a point cloud model, a mesh model, a voxel raster model, or another type of 3D environment model of the audio environment. The 3D environment model 906 can be used to provide surface information, such as surface normals and / or intersecting surfaces in the audio environment. The 3D environment model 906 can also be used, or alternatively, to calculate the total ray travel distance. In some examples, the 3D environment model 906 can also be used, or alternatively, to calculate the time delay from the audio source to the audio capture device based on environmental conditions associated with the audio environment.
[0106] Rendering process 902 additionally or alternatively utilizes ray tracing 908. Ray tracing 908 can provide 3D coordinates associated with surface reflections of audio rays within the audio environment. Rendering process 902 additionally or alternatively utilizes neural network processing 820 to provide reflection coefficients associated with the 3D coordinates. Based on ray tracing 908, training stream 900 includes an energy histogram process 910 for generating an energy histogram associated with the audio environment. In some examples, in addition to or instead of ray tracing 908, rendering process 902 may utilize one or more other acoustic simulation techniques, such as, but not limited to: image-source acoustic simulation techniques, wave-based acoustic simulation techniques, diffuse rain acoustic simulation techniques (e.g., for modeling energy associated with diffuse reflection), and / or one or more other types of acoustic simulation techniques. In some examples, loss function 912 may be used to measure the error of the energy histogram based on a comparison with ground-based energy histograms included in the training samples. In some examples, loss function 912 may be used to adjust one or more weights of the neural network and / or generate one or more new ground-based energy histograms. For example, loss function 912 can calculate the error between the rendered energy histogram (e.g., energy histogram 860) and the ground-based energy histogram associated with the training dataset used for the audio environment. In some examples, training flow 900 is repeated until the training samples converge to the defined loss function value.
[0107] Figure 10 The invention illustrates one or more embodiments of the invention. Figure 1 The neuroacoustic modeling system 104 enables an inference stream 1000. The inference stream 1000 can be associated with an inference pipeline for neuroacoustic modeling used for ray tracing. Using the inference stream 1000, a 3D environment can be initialized via an inference substream 1004 using user-defined source and receiver locations 1002. The inference stream 1000 also includes a ray tracing process 1006 associated with the acoustic simulation of the 3D environment. In some examples, the ray tracing process 1006 can generate audio ray information, such as ray energy and / or total ray length. The ray tracing process 1006 can additionally or alternatively generate an energy histogram based on the audio ray information (e.g., ray energy and / or total ray length).
[0108] The inference stream 1000 also includes an inference substream 1008 that transforms the energy histogram into an impulse response set. Based on the impulse response set and one or more audio samples 1010 associated with captured audio in the audio environment, a convolution process 1012 can be performed to generate simulated reverberant audio samples 1014. In some examples, the convolution process 1012 can receive audio samples and impulse responses as input to produce audio samples that simulate sound propagating between two locations in the audio environment. The impulse responses can be used for source-receiver pairs from the acoustic simulation. If the sound of one or more audio samples 1010 propagates from a source to a receiver location in the audio environment, the simulated reverberant audio samples 1014 can simulate the sound that the audio associated with one or more audio samples 1010 will sound like. In some examples, visualizations related to surface material properties learned in the audio environment can be provided. In some examples, the learned surface material properties can be visualized by querying a trained network for each point on a surface observed through a visual rendering of the audio environment. In some examples, the visual properties of the surface (e.g., color) can be configured based on the surface's coefficients. In some examples, the neuroacoustic modeling inference data 111 includes simulated reverberant audio samples 1014 and / or learned surface material properties.
[0109] In some examples, the NeRF model can be leveraged or enhanced to provide simulations and / or renderings of audio for a specific 3D environment, also outputting surface material properties (e.g., material classification) of surfaces and / or objects in the environment. In some examples, a predefined lookup table is used to obtain known absorption coefficients from the predicted material classification. Using these absorption coefficients and geometries provided by the NeRF model, acoustic simulations and / or audio rendering techniques can be performed to synthesize the impulse response of source-receiver location pairs.
[0110] In some examples, the MLP of the NeRF model can be enhanced to output material classifications in addition to RGB and volume density (e.g., drywall, painted concrete, windows, etc.). Material classifications can be estimated for each 3D point and / or orientation in the audio environment.
[0111] In some examples, ground-based material classification can be obtained using a separate pre-trained network that segments collected images into material classifications. The ground-based material classification can be encoded as additional pixel channels. Given camera positions from the training set, the neural network can render pixels with these additional channels and / or compare the corresponding renders to ground-based pixels. In some examples, material types can be integrated with audio rays to enhance alpha synthesis of pixel-related color information.
[0112] In some examples, a lookup table can be queried to determine predictive material for inferences related to the audio environment. The lookup table can contain known materials and their corresponding absorption coefficients. Using these absorption coefficients and the 3D geometry of the room, a simulation of the audio environment can be provided to render one or more impulse responses related to the audio environment.
[0113] In some examples, the convolution process 1012 may query the audio rendering model 105 for each reflection point of each ray tracing through the 3D environment (e.g., each ray tracing through the 3D environment from source to receiver) during the ray tracing process 1006. In some examples, the energy of the ray may be attenuated in each frequency band based on inferred values provided by the audio rendering model 105. In some examples, the ray may contribute to a time delay in the energy reaching the receiver location, where the time delay is based on the path length of the ray and / or the speed of sound associated with the ray. In some examples, the path length may be used to attenuate the energy of each ray to account for the impedance of the air in the environment. In some examples, the energy of the ray may be transformed into an impulse response associated with the reflection point and / or the source-to-receiver pair.
[0114] In some examples, a user can select user-defined source and receiver locations 1002 via an electronic interface on a user device to indicate a source and receiver pair for simulating acoustics between the user-defined source and receiver locations 1002. Alternatively, a ray tracing process 1006 or another acoustic simulation process (e.g., an image source acoustic simulation process) can be performed to determine the ray path propagating between the user-defined source and receiver locations 1002. During propagation between the source and receiver locations, the ray can be reflected from one or more surfaces and / or objects within the audio environment. For example, the ray tracing process 1006 can determine a reflection point corresponding to a specific reflection from a surface and / or object within the audio environment.
[0115] For each ray, the position of each reflection point along the ray path can be encoded as a position embedding. In some examples, the corresponding encoded position of the corresponding reflection point can be fed as input to a machine learning model, such as audio rendering model 105. The output of the machine learning model can include the absorption coefficient and / or reflection coefficient for each frequency band. In some examples, the inferred absorption and / or reflection coefficients can be used to attenuate the energy associated with the ray. Additionally, the attenuated energy can be modified based on the time delay, path length, and / or impedance of the air associated with the ray. In some examples, the attenuated energy can be transformed into an energy histogram for the corresponding frequency band. Furthermore, the energy histogram can be transformed into an impulse response via inferred sub-current 1008.
[0116] In some examples, reflection point data of the audio path (e.g., a ray) associated with user-defined source and receiver locations 1002 can be determined via a ray tracing process 1006 or another acoustic simulation process associated with the audio environment. Additionally, a location embedding associated with the audio path can be determined based on the reflection point data. The location embedding can be provided as input to a machine learning model (e.g., an audio rendering model 105) configured to generate audio absorption characteristic data associated with the audio path. The audio absorption characteristic data can include absorption coefficients and / or reflection coefficients for each frequency band. Furthermore, a simulated reverberant audio sample 1014 associated with the audio environment can be determined based on the audio absorption characteristic data.
[0117] In some examples, energy data associated with the audio environment can be determined based on audio absorption characteristic data. The energy data may include one or more energy histograms. Additionally, audio data associated with the audio environment (e.g., one or more audio samples 1010) may be received. In some examples, simulated reverberation audio sample 1014 can be determined based on the energy data and the audio data.
[0118] In some examples, audio paths can be selected from multiple audio paths in an audio environment based on an association with the receiver location of the user-defined source and receiver location 1002. Additionally, a ray tracing process 1006 or another acoustic simulation process can be performed iteratively for multiple audio paths in the audio environment (e.g., one or more other audio paths among multiple audio paths). In some examples, the multiple audio paths can comprise hundreds, thousands, or another number of audio paths, which can be received by a capture device located at the receiver location of the user-defined source and receiver location 1002.
[0119] In some examples, the simulated reverberant audio sample 1014 can be output via an audio output device. The audio output device can be an audio mixer device, a DSP processing device, a smartphone, a tablet computer, a laptop computer, a personal computer, an audio workstation device, a wearable device, an augmented reality device, a virtual reality device, a recording device, a microphone, a headset, an earphone, a speaker, a haptic device, or another type of output device.
[0120] Embodiments of this disclosure are described below with reference to block diagrams and flowcharts. Therefore, it should be understood that each block of the block diagrams and flowcharts may be implemented as a computer program product, a complete hardware embodiment, a combination of hardware and computer program products, and / or a device, system, computing device / entity, computing entity, etc., which implement instructions, operations, steps, and similar terms (e.g., executable instructions, instructions for execution, program code, etc.) interchangeably used on a computer-readable storage medium for execution. For example, the retrieval, loading, and execution of code may be performed sequentially so that one instruction is retrieved, loaded, and executed at a time.
[0121] In some example embodiments, retrieval, loading, and / or execution can be performed in parallel, such that multiple instructions are retrieved, loaded, and / or executed together. Therefore, such embodiments can produce execution on a machine with a specific configuration for the steps or operations specified in the block diagram and flowchart description, and / or on a machine with a specific configuration for the steps or operations specified in the block diagram and flowchart description. Thus, the block diagram and flowchart description supports various combinations of embodiments for performing the specified instructions, operations, or steps.
[0122] Figure 11 It is based on, for example Figure 2 The flowchart illustrates an example process 1100 for providing neuroacoustic modeling for an audio environment using the neuroacoustic modeling device 202 shown. Through various operations of process 1100, the neuroacoustic modeling device 202 can enhance the quality and / or reliability of audio provided by an audio system.
[0123] Process 1100 begins with operation 1102, which (e.g., via model generation circuitry 208) receives audio and video data associated with an audio environment. The audio environment can be an indoor environment, an outdoor environment, a room, an auditorium, a performance hall, a broadcast environment, an arena (e.g., a stadium), a virtual environment, or another type of audio environment. In some examples, the audio and video data are captured via at least one microphone and at least one camera, respectively, of a capture device that scans the audio environment.
[0124] Process 1100 also includes operation 1104, which generates (e.g., via model generation circuitry 208) an image set comprising multiple images, each image associated with an audio sample representing the acoustic properties of the audio environment, based at least in part on audio and video data. In some examples, the corresponding image may be associated with a corresponding impulse response associated with the audio sample.
[0125] Process 1100 further includes operation 1106, which determines (e.g., via model generation circuitry 208) a camera property set based at least in part on audio samples, the camera property set including relative audio sample positions and camera orientations associated with the audio samples. In some examples, the camera property set includes: the corresponding position of the image set relative to the audio environment, the corresponding orientation of the image set relative to the audio environment, photogrammetric information associated with the image set, motion reconstruction structure information associated with the image set, point cloud information associated with the image set, and / or other information.
[0126] Process 1100 further includes operation 1108, which generates (e.g., via model generation circuitry 208) an audio rendering model based at least in part on an image set, audio samples, and / or a set of camera properties, wherein the audio rendering model includes a neural rendering volumetric representation of the audio environment augmented with audio coding. The audio rendering model may be an audio rendering model for an audio environment. In some examples, the audio rendering model is a NeRF model augmented with audio coding. In some examples, the image set, audio samples, and / or the set of camera properties are transformed into data vectors to be input into the audio rendering model. In some examples, the weights of the audio rendering model may be trained based on the volumetric probability density representing the corresponding location within the audio environment. The weights may be configured at least in part based on latent encodings of physical and acoustic information for the corresponding location within the audio environment.
[0127] In some examples, training data vectors are fed into the audio rendering model. These training data vectors may contain: image pixels of an image set augmented with camera property sets and audio coding, impulse responses of the image set augmented with camera property sets and audio coding, material acoustic properties of the image set augmented with camera property sets and audio coding, material type of the image set augmented with camera property sets and audio coding, and / or environmental measurements related to the audio environment. Material acoustic properties can be inferred from the acoustic classification network model. Alternatively, material type can be inferred from the image classification network model.
[0128] Process 1100 further includes operation 1110, which outputs (e.g., via model inference circuitry 210) one or more audio inferences associated with an audio environment, wherein the one or more audio inferences are generated at least in part based on an audio rendering model. The output or multiple audio inferences may include one or more portions of neuroacoustic modeling inference data (e.g., neuroacoustic modeling inference data 111) provided by an audio rendering model (e.g., audio rendering model 105). In some examples, outputting one or more audio inferences includes outputting one or more candidate audio component locations associated with an audio environment, wherein the one or more candidate audio component locations are generated at least in part based on an audio rendering model.
[0129] In some examples, the audio rendering model is used to: infer the location of an audio source within an audio environment; infer the acoustic properties of sound within or emanating from the audio environment; generate a digital twin of the audio environment; generate an audio heatmap of the audio environment; generate an audio simulation of the audio environment; generate a set of drawings or images of the audio environment and optimal locations and / or audio settings for audio / video equipment within the audio environment; infer the material properties of objects or surfaces within the audio environment; and / or control audio equipment within the audio environment. In some examples, material properties may include: sound reflection, absorption, transmission, refraction, diffraction, frequency correlation, resonance, nonlinearity, surface texture, material type, surface normal, surface geometry, absorption coefficient, and / or other properties.
[0130] In some examples, the inference from the audio rendering model can be used to generate new training data for training one or more other neural rendering models. For instance, the audio rendering model can be used to model reverberation in different audio environments based on impulse responses to position pairs in the audio environment associated with the audio rendering model. In another example, unprocessed audio and associated simulated reverberation output can be used to generate training samples for a machine learning model configured to perform de-reverberation on audio.
[0131] Figure 12 It is based on, for example Figure 2 The flowchart illustrates an example process 1200 for providing audio inference using neuroacoustic modeling for an audio environment in the neuroacoustic modeling device 202 shown. Through various operations of process 1200, the neuroacoustic modeling device 202 can enhance the quality and / or reliability of audio provided by an audio system.
[0132] Process 1200 begins with operation 1202, which receives (e.g., from model inference circuitry 210) a digital exploration request associated with an audio environment, the digital exploration request including at least an audio environment identifier for the audio environment. In some examples, the digital exploration request is received from a user device.
[0133] Process 1200 further includes operation 1204, which generates (e.g., via model inference circuitry 210) one or more audio inferences associated with an audio environment, at least in part, based on an audio rendering model associated with an audio environment identifier, wherein the audio rendering model includes a neural rendering volume representation of the audio environment enhanced with audio coding. In some examples, the audio rendering model is a NeRF model enhanced with audio coding. In some examples, process 1100 additionally or alternatively includes generating one or more candidate audio component locations based at least in part on the audio rendering model.
[0134] Process 1200 further includes operation 1206, which outputs (e.g., inferring circuitry 210 from a model) one or more candidate audio component locations associated with the audio environment. In some examples, process 1100 additionally or alternatively includes outputting one or more candidate audio component locations associated with the audio environment. In some examples, the one or more candidate audio component locations include candidate video component locations. In some examples, the one or more candidate audio component locations correspond to corresponding three-dimensional coordinates of an audio rendering model. In some examples, the optimal location of an audio component in the audio environment can be determined at least in part based on the audio rendering model.
[0135] In some examples, the audio rendering model is used to: render a visual representation of an audio environment via a user interface; render a digital twin of the audio environment via a user interface; render an audio heatmap of the audio environment via a user interface; generate an audio simulation of the audio environment; infer the acoustic properties of sounds emitted within the audio environment; and / or infer the material properties of objects or surfaces within the audio environment. In some examples, material properties may include: sound reflection, absorption, transmission, refraction, diffraction, frequency correlation, resonance, nonlinearity, surface texture, material type, surface normal, surface geometry, and / or other properties.
[0136] Although an example processing system has been described in the figures herein, the implementation of the subject matter and functional operations described herein can be implemented in other types of digital electronic circuit systems, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more of them.
[0137] The embodiments of the subject matter and operations described herein can be implemented in digital electronic circuit systems, or in computer software, firmware, or hardware, encompassing the structures disclosed herein and their structural equivalents, or in combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs encoded on a computer-readable storage medium to be executed by or control the operation of an information / data processing device, i.e., one or more modules of computer program instructions. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information / data for transmission to a suitable receiver device for execution by the information / data processing device. The computer-readable storage medium can be a computer-readable storage device, a computer-readable storage device substrate, a random or serial access memory array or device, or a combination thereof, or contained within a computer-readable storage device, a computer-readable storage device substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, while the computer-readable storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded as artificially generated propagating signals. Computer-readable storage media may also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices) or contained within said one or more separate physical components or media.
[0138] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or information / data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing portions of one or more modules, subroutines, or code). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected via a communication network.
[0139] The processes and logic flows described herein can be executed by one or more programmable processors, which execute one or more computer programs to perform actions by manipulating input information / data and generating output. Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and information / data from read-only memory, random access memory, or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data or operatively coupled to receive information / data from or to said one or more mass storage devices, or both. However, a computer does not need to have such devices. Suitable devices for storing computer program instructions and information / data include all forms of non-volatile memory, media, and memory devices, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by or incorporated into dedicated logic circuitry systems.
[0140] Unless otherwise specified, the term "or" is used herein in both the sense of substitution and associativity. The terms "illustrative," "example," and "exemplary" are used as examples that do not indicate a quality level. Similar numbers always refer to similar elements.
[0141] The term “comprising” means “including but not limited to” and should be interpreted in the manner commonly used in the patent context. For example, the use of broad terms such as “comprising,” “including,” and “having” should be understood to be supported by narrower terms such as “consisting of,” “mainly composed of,” and “substantially composed of.”
[0142] The phrases “in one embodiment”, “according to one embodiment”, etc., generally mean that the specific feature, structure or characteristic following the phrase can be included in at least one embodiment of this disclosure, and can be included in more than one embodiment of this disclosure (importantly, such phrases do not necessarily refer to the same embodiment).
[0143] While this specification contains numerous details of specific embodiments, these details should not be construed as limiting the scope of any disclosure or what may be claimed, but rather as descriptions of features specific to particular embodiments of the particular disclosure. Certain features described herein in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations, and even initially claimed so, in some cases one or more features of a claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0144] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or that all shown operations should be performed to achieve the desired result, unless otherwise described. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the above embodiments should not be construed as requiring such separation in all embodiments; rather, it should be understood that the described program components and systems can generally be integrated together in a product or packaged into multiple products.
[0145] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions listed in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result, unless otherwise described. In some embodiments, multitasking and parallel processing may be advantageous.
[0146] In the following text, various features will be highlighted in a set of numbered clauses or paragraphs. These features should not be construed as limiting the scope of this disclosure or the inventive concept, but are provided merely as a highlighting of some features as described herein, without implying a particular order of importance or relevance of such features.
[0147] Clause 1. A device comprising at least one processor and a memory storing instructions that, when executed by the processor, are operable to cause the device to: receive a digital exploration request associated with an audio environment, the digital exploration request including an audio environment identifier associated with the audio environment.
[0148] Clause 2. The device according to Clause 1, wherein the instructions are further operable to cause the device to: generate one or more candidate audio component locations based at least in part on an audio rendering model associated with the audio environment identifier.
[0149] Clause 3. The device according to any one of the preceding clauses, wherein the audio rendering model comprises a neural rendering volume representation of the audio environment enhanced with audio coding.
[0150] Clause 4. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: output the location of the one or more candidate audio components associated with the audio environment.
[0151] Clause 5. The device according to any one of the preceding clauses, wherein the audio rendering model includes an enhanced neural radiation field (NeRF) model.
[0152] Clause 6. The device pursuant to any of the preceding clauses, wherein the one or more candidate audio component locations include candidate video component locations.
[0153] Clause 7. The device according to any one of the preceding clauses, wherein the location of the one or more candidate audio components corresponds to the corresponding three-dimensional coordinates of the audio rendering model.
[0154] Clause 8. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine the optimal location of an audio component in the audio environment based at least in part on the audio rendering model.
[0155] Clause 9. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: cause a visual representation of the audio environment to be rendered via a user interface, at least in part, based on the audio rendering model.
[0156] Clause 10. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: cause a digital twin of the audio environment to be rendered via a user interface, at least in part, based on the audio rendering model.
[0157] Clause 11. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: render an audio heatmap of the audio environment via a user interface, at least in part, based on the audio rendering model.
[0158] Clause 12. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate an audio simulation of the audio environment based at least in part on the audio rendering model.
[0159] Clause 13. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate transfer function information for source / listener pairs within the audio environment, at least in part, based on the audio rendering model.
[0160] Clause 14. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate training data for one or more other models based at least in part on the audio rendering model.
[0161] Clause 15. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: infer the acoustic properties of sounds emitted within the audio environment based at least in part on the audio rendering model.
[0162] Clause 16. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: infer material properties of objects or surfaces within the audio environment based at least in part on the audio rendering model.
[0163] Clause 17. A computer-implemented method comprising steps relating to any of the preceding clauses.
[0164] Clause 18. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of the device, cause the one or more processors to perform one or more operations relating to any of the preceding clauses.
[0165] Clause 19. An apparatus comprising at least one processor and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to: receive audio data and image data associated with an audio environment.
[0166] Clause 20. The device according to Clause 19, wherein the instructions are further operable to cause the device to: generate an image set based at least in part on the audio data and the image data, the image set comprising a plurality of images, each image being associated with an audio sample representing the acoustic properties of the audio environment.
[0167] Clause 21. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine a set of camera properties based at least in part on the audio samples, the set of camera properties including relative audio sample positions and camera orientations associated with the audio samples.
[0168] Clause 22. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate an audio rendering model for the audio environment based at least in part on the image set, the audio samples, and / or the camera property set.
[0169] Clause 23. The device according to any one of the preceding clauses, wherein the audio rendering model comprises a neural rendering volume representation of the audio environment enhanced with audio coding.
[0170] Clause 24. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: output one or more candidate audio component locations associated with the audio environment.
[0171] Clause 25. The device according to any one of the preceding clauses, wherein the locations of the one or more candidate audio components are generated at least in part based on the audio rendering model.
[0172] Clause 26. The device according to any one of the preceding clauses, wherein the audio rendering model is an enhanced neural radiation field (NeRF) model enhanced with the audio encoding.
[0173] Clause 27. The device according to any one of the preceding clauses, wherein the audio data and the image data are captured via at least one microphone and at least one camera of a capture device for scanning the audio environment.
[0174] Clause 28. The device according to any one of the preceding clauses, wherein the camera property set includes the corresponding position and / or orientation of the image set relative to the audio environment.
[0175] Clause 29. The device according to any one of the preceding clauses, wherein the camera property set includes photogrammetric information about the image set.
[0176] Clause 30. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: transform the image set, the audio samples, and the camera property set into data vectors for input into the audio rendering model.
[0177] Clause 31. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: train the weights of the audio rendering model based on a volume probability density representing a corresponding location within the audio environment.
[0178] Clause 32. The device according to any one of the preceding clauses, wherein the weights are configured at least in part based on potential encoding of physical and acoustic information for the corresponding location in the audio environment.
[0179] Clause 33. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: train the weights of the audio rendering model based on color information relating to a corresponding location within the audio environment.
[0180] Clause 34. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: train the weights of the audio rendering model based on the material properties of one or more objects within the audio environment.
[0181] Clause 35. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input a training data vector to the audio rendering model, the training data vector comprising image pixels of the image set enhanced with the camera property set and the audio encoding.
[0182] Clause 36. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input a training data vector to the audio rendering model, the training data vector comprising the impulse response of the image set enhanced with the camera property set and the audio encoding.
[0183] Clause 37. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input a training data vector to the audio rendering model, the training data vector comprising material acoustic properties of the image set enhanced with the camera property set and the audio encoding.
[0184] Clause 38. The device according to any one of the preceding clauses, wherein the acoustic properties of the material are inferred from an acoustic classification network model.
[0185] Clause 39. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input a training data vector to the audio rendering model, the training data vector comprising material types of the image set enhanced with the camera property set and the audio encoding.
[0186] Clause 40. The device according to any one of the preceding clauses, wherein the material type is inferred from an image classification network model.
[0187] Clause 41. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input a training data vector to the audio rendering model, the training data vector including environmental measurements related to the audio environment.
[0188] Clause 42. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: infer the location of an audio source within the audio environment based at least in part on the audio rendering model.
[0189] Clause 43. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: infer, at least in part, the acoustic properties of sound emanating from or within the audio environment based on the audio rendering model.
[0190] Clause 44. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate a digital twin of the audio environment based at least in part on the audio rendering model.
[0191] Clause 45. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate an audio heatmap of the audio environment based at least in part on the audio rendering model.
[0192] Clause 46. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate an audio simulation of the audio environment based at least in part on the audio rendering model.
[0193] Clause 47. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate, at least in part, a set of drawings or images of the audio environment and optimal positions and / or audio settings of audio / video equipment within the audio environment based on the audio rendering model.
[0194] Clause 48. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: infer material properties of objects or surfaces within the audio environment based at least in part on the audio rendering model.
[0195] Clause 49. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: control audio equipment in the audio environment based at least in part on the audio rendering model.
[0196] Clause 50. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: receive a digital exploration request associated with the audio environment, the digital exploration request including an audio environment identifier associated with the audio environment.
[0197] Clause 51. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: identify the audio rendering model based at least in part on the audio environment identifier.
[0198] Clause 52. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate one or more audio inferences based at least in part on the audio rendering model.
[0199] Clause 53. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine one or more impulse responses associated with one or more audio sources in the audio environment, at least in part, based on the audio rendering model.
[0200] Clause 54. A computer-implemented method comprising steps relating to any of the preceding clauses.
[0201] Clause 55. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of the device, cause the one or more processors to perform one or more operations relating to any of the preceding clauses.
[0202] Clause 56. A device comprising at least one processor and a memory, the memory storing instructions that, when executed by the processor, are operable to cause the device to: receive audio data and image data associated with an audio environment.
[0203] Clause 57. The device according to Clause 56, wherein the instructions are further operable to cause the device to: generate an impulse response set based at least in part on the audio data and the image data, the impulse response set being associated with audio samples representing the acoustic properties of the audio environment.
[0204] Clause 58. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate an audio rendering model based at least in part on the impulse response set and the audio samples.
[0205] Clause 59. The device according to any one of the preceding clauses, wherein the audio rendering model comprises a neural rendering volume representation of the audio environment enhanced with audio coding.
[0206] Clause 60. A computer-implemented method comprising steps relating to any of the preceding clauses.
[0207] Clause 61. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of the device, cause the one or more processors to perform one or more operations relating to any of the preceding clauses.
[0208] Clause 62. A device comprising at least one processor and a memory, the memory storing instructions operable, when executed by the processor, to cause the device to: receive audio data associated with an audio environment.
[0209] Clause 63. The device according to Clause 62, wherein the instructions are further operable to cause the device to: generate an energy histogram associated with the audio data based on energy data associated with the audio data.
[0210] Clause 64. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input the energy histogram and the audio data into a machine learning model, the machine learning model being configured to generate simulated reverberant audio samples associated with the audio data.
[0211] Clause 65. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: output the analog reverberant audio sample via an audio output device.
[0212] Clause 66. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: query a three-dimensional (3D) model of the audio environment to generate surface information relating to the audio environment.
[0213] Clause 67. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: generate the energy histogram based on the surface information.
[0214] Clause 68. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input the energy histogram and the audio data into the machine learning model to generate a material classification of one or more surfaces within the audio environment.
[0215] Clause 69. A computer-implemented method comprising steps relating to any of the preceding clauses.
[0216] Clause 70. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of the device, cause the one or more processors to perform one or more operations relating to any of the preceding clauses.
[0217] Clause 71. An apparatus comprising at least one processor and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to: determine, at least in part, reflection point data for an audio path associated with a source location and a receiver location in the audio environment, based on an acoustic simulation process associated with the audio environment.
[0218] Clause 72. The device according to Clause 71, wherein the instructions are further operable to cause the device to: determine the position embedding associated with the audio path based at least in part on the reflection point data.
[0219] Clause 73. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: input the location embedding into a machine learning model, the machine learning model being configured to generate audio absorption characteristic data associated with the audio path.
[0220] Clause 74. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine, at least in part, an analog reverberant audio sample associated with the audio environment based on the audio absorption characteristics data.
[0221] Clause 75. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: output the analog reverberant audio sample via an audio output device.
[0222] Clause 76. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine energy data associated with the audio environment based at least in part on the audio absorption characteristics data.
[0223] Clause 77. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: receive audio data associated with the audio environment.
[0224] Clause 78. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine the analog reverberant audio sample based at least in part on the energy data and / or the audio data.
[0225] Clause 79. The device pursuant to any of the preceding clauses, wherein the machine learning model is a multilayer perceptron model.
[0226] Clause 80. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: determine the material classification of one or more surfaces within the audio environment based at least in part on the simulated reverberant audio samples.
[0227] Clause 81. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: query a three-dimensional (3D) model of the audio environment based on the simulated reverberant audio sample to determine surface information associated with the audio environment.
[0228] Clause 82. The device according to any one of the preceding clauses, wherein the instructions are further operable to cause the device to: select the audio path from a plurality of audio paths based on an association with respect to the location of the receiver.
[0229] Clause 83. A computer-implemented method comprising steps relating to any of the preceding clauses.
[0230] Clause 84. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of the device, cause the one or more processors to perform one or more operations relating to any of the preceding clauses.
[0231] Many modifications and other embodiments of the present disclosure will occur to those skilled in the art, taking advantage of the teachings presented in the foregoing description and the associated accompanying drawings. Therefore, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. While specific terminology is used herein, it is used only in a general and descriptive sense and is not intended for limiting purposes unless otherwise described.
Claims
1. A device comprising at least one processor and memory storing instructions that are operable, when executed by the processor, to cause the device to: receive audio data and image data associated with an audio environment; generate, based at least in part on the audio data and the image data, a set of images, the set of images comprising a plurality of images each associated with an audio sample representative of an acoustic property of the audio environment; and generate, based at least in part on the set of images and the audio sample, an audio rendering model for the audio environment, wherein the audio rendering model comprises a neural rendering volumetric representation of the audio environment augmented with an audio encoding.
2. The device of claim 1, wherein the audio rendering model comprises an augmented neural radiance field (NeRF) model augmented with the audio encoding.
3. The device of claim 1, wherein the audio data and the image data are captured via at least one microphone and at least one camera, respectively, of a capture device that scans the audio environment.
4. The device of claim 1, wherein the instructions are further operable to cause the device to: train weights of the audio rendering model according to a volumetric probability density representative of respective locations within the audio environment, wherein the weights are configured based at least in part on a latent encoding of physical information and acoustic information for respective locations of the audio environment.
5. The device of claim 1, wherein the instructions are further operable to cause the device to: generate, based at least in part on the audio sample, a set of camera properties, the set of camera properties comprising a relative audio sample location and a camera orientation associated with the audio sample; generate, based at least in part on the set of images, the audio sample, and the set of camera properties, the audio rendering model.
6. The device of claim 5, wherein the set of camera properties comprises respective locations of the set of images relative to the audio environment.
7. The device of claim 5, wherein the set of camera properties comprises respective orientations of the set of images relative to the audio environment.
8. The device of claim 5, wherein the instructions are further operable to cause the device to: input, to the audio rendering model, a training data vector comprising an impulse response of the set of images augmented with the set of camera properties and the audio encoding.
9. The device of claim 5, wherein the instructions are further operable to cause the device to: input, to the audio rendering model, a training data vector comprising material acoustic properties of the set of images augmented with the set of camera properties and the audio encoding.
10. The device of claim 1, wherein the instructions are further operable to cause the device to: determine, based at least in part on the audio rendering model, one or more impulse responses associated with one or more audio sources in the audio environment.
11. The device of claim 1, wherein the instructions are further operable to cause the device to: inference of a location of an audio source within the audio environment based at least in part on the audio rendering model.
12. The device of claim 1, wherein the instructions are further operable to cause the device to: inference of an acoustic property of a sound within or emanating from the audio environment based at least in part on the audio rendering model.
13. The device of claim 1, wherein the instructions are further operable to cause the device to: output one or more candidate audio component locations associated with the audio environment, wherein the one or more candidate audio component locations are generated based at least in part on the audio rendering model.
14. The device of claim 1, wherein the instructions are further operable to cause the device to: generation of a digital twin of the audio environment based at least in part on the audio rendering model.
15. The device of claim 1, wherein the instructions are further operable to cause the device to: generation of an audio heat map of the audio environment based at least in part on the audio rendering model.
16. The device of claim 1, wherein the instructions are further operable to cause the device to: generation of an audio simulation of the audio environment based at least in part on the audio rendering model.
17. The device of claim 1, wherein the instructions are further operable to cause the device to: control of audio equipment in the audio environment based at least in part on the audio rendering model.
18. The device of claim 1, wherein the instructions are further operable to cause the device to: inference of a material property of an object or surface within the audio environment based at least in part on the audio rendering model.
19. The device of claim 1, wherein the instructions are further operable to cause the device to: generation of a set of drawings or images of the audio environment and optimal locations or audio settings for audio equipment within the audio environment based at least in part on the audio rendering model.
20. The device of claim 1, wherein the instructions are further operable to cause the device to: receive a digital exploration request associated with the audio environment, the digital exploration request including an audio environment identifier associated with the audio environment; identify the audio rendering model based at least in part on the audio environment identifier; and generate one or more audio inferences based at least in part on the audio rendering model.
21. A computer-implemented method comprising: receiving audio data and image data associated with an audio environment; generating a set of images based at least in part on the audio data and the image data, the set of images including a plurality of images each associated with an audio sample representative of an acoustic property of the audio environment; and generating an audio rendering model for the audio environment based at least in part on the set of images and the audio samples, wherein the audio rendering model includes a neural rendering volumetric representation of the audio environment augmented with audio encoding.
22. A computer program product stored on a computer-readable medium, the computer program product comprising instructions that, when executed by one or more processors of a device, cause the one or more processors to: receive audio data and image data associated with an audio environment; generate, based at least in part on the audio data and the image data, a set of images, the set of images comprising a plurality of images each associated with an audio sample representative of an acoustic property of the audio environment; and generate, based at least in part on the set of images and the audio sample, an audio rendering model for the audio environment, wherein the audio rendering model comprises a neural rendered volumetric representation of the audio environment augmented with audio coding.
Citation Information
Cited By
Sound field reconstruction and new perspective audio synthesis method and system based on gaussian sputtering
CN122369422A