Generation of acoustics-matched audio for three-dimensional (3D) scene associated with metaverse application
Deep learning and cross-modal encoding techniques address the challenge of generating acoustics-matched audio in metaverse applications by processing multi-view images and integrating spatial characteristics with audio data, enhancing audio-visual coherence and immersion.
Patent Information
- Application Number
- PCT/IB2025/056635
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-03
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Conventional audio generation methods in virtual environments fail to accurately capture nuanced acoustic properties of complex 3D scenes and adapt to changes in the virtual scene or user interactions, leading to inconsistencies between visual and auditory cues in metaverse applications.
Utilizing deep learning and cross-modal encoding techniques to process multi-view images and integrate spatial characteristics with audio data, generating acoustics-matched audio through a cross-modal encoder model that considers both speaker and listener perspectives.
Enhances audio-visual coherence and immersion by providing accurate, contextually relevant audio that adapts to real-time changes in the 3D scene, offering a more realistic and dynamic audio experience.
Smart Images

Figure IB2025056635_08012026_PF_FP_ABST
Abstract
Description
GENERATION OF ACOUSTICS-MATCHED AUDIO FOR THREE-DIMENSIONAL (3D) SCENE ASSOCIATED WITH METAVERSE APPLICATIONCROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE
[0001] This application claims priority to Indian Provisional Application No. IN202411051082, filed July 3, 2024, which is hereby incorporated by reference in its entirety.FIELD
[0002] Various embodiments of the disclosure relate to Virtual Reality (VR). More specifically, various embodiments of the disclosure relate to an electronic device and a method to generate acoustics-matched audio for 3D scenes associated with metaverse applications using deep learning and cross-modal encoding techniques.BACKGROUND
[0003] Metaverse applications have gained significant popularity in recent years, offering immersive virtual environments where users may interact with digital content and other users in three-dimensional (3D) spaces. These applications span various domains, including entertainment, education, social networking, and business collaboration. As metaverse technologies have advanced, there has been an increasing focus on enhancing the realism and immersion of these virtual experiences.
[0004] One key aspect of creating realistic metaverse environments is the generation of accurate and contextually appropriate audio. Conventional approaches to audio generation in virtual environments often rely on pre-recorded sound effects, basic spatial audio techniques, or simplified acoustic models. While these methods may provide a basic level of audio immersion, they may not fully capture the nuanced acoustic properties ofcomplex 3D scenes or accurately represent the interactions between sound sources and virtual environments. Additionally, existing solutions may struggle to dynamically adapt audio characteristics based on changes in the virtual scene or user interactions, which may lead to inconsistencies between visual and auditory cues in the metaverse application.
[0005] Limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0006] An electronic device and method to generate acoustics-matched audio for three- dimensional (3D) scenes in metaverse applications is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a diagram that illustrates a network environment for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure.
[0009] FIG. 2 is a block diagram of an electronic device of FIG. 1 , in accordance with an embodiment of the disclosure.
[0010] FIG. 3 is a diagram that illustrates an exemplary execution pipeline for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure.
[0011] FIG. 4 is a diagram that illustrates an exemplary scenario of a metaverse engine, in accordance with an embodiment of the disclosure.
[0012] FIG. 5 is a diagram that illustrates an exemplary scenario of an audio processing system for a metaverse application, in accordance with an embodiment of the disclosure.
[0013] FIG. 6 is a diagram that illustrates an exemplary scenario for generation of acoustics-matched audio in a virtual environment, in accordance with an embodiment of the disclosure.
[0014] FIG. 7 is a diagram that illustrates an exemplary scenario for generation of acoustically matched audio in a metaverse environment, in accordance with an embodiment of the disclosure.
[0015] FIG. 8 is a diagram that illustrates an exemplary execution pipeline for generation of acoustics-matched audio, in accordance with an embodiment of the disclosure.
[0016] FIG. 9 is a diagram that illustrates an exemplary scenario for customization of audio in a virtual environment, in accordance with an embodiment of the disclosure.
[0017] FIG. 10 is a diagram that illustrates an exemplary scenario for processing of audio in a virtual environment, in accordance with an embodiment of the disclosure.
[0018] FIG. 11 is a diagram that illustrates an exemplary scenario of a spatial audio processing system, in accordance with an embodiment of the disclosure.
[0019] FIG. 12 is a diagram that illustrates an exemplary scenario for generation of listener perspective audio in a virtual environment, in accordance with an embodiment of the disclosure.
[0020] FIG. 13 is a flowchart of an example method for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure.DETAILED DESCRIPTION
[0021] The present disclosure relates to generation of acoustics-matched audio for three-dimensional (3D) scene associated with metaverse application. The following described implementation may be found in an electronic device and method that may be configured to generate acoustics-matched audio for three-dimensional (3D) scene associated with metaverse application. Exemplary aspects of the present disclosure may provide an electronic device that may be configured to receive, from a metaverse engine, multi-view images associated with a metaverse application. The multi-view images correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application. The 3D scene may include, but may not be limited to, virtual musical concert, virtual conference room, electronic learning or e-learning, or metaverse tourism. The electronic device may be further configured to apply a deep learning model on the multiview images and determine volume aggregation information associated with the 3D scene, based on the application of the deep learning model. For example, the volume aggregation information may refer to a comprehensive representation of spatial characteristics and acoustic properties derived from analysis of multiple views or perspectives of the 3D scene. The electronic device may receive a source audio associated with the first user and apply a cross-modal encoder model on the volume aggregation information and the source audio. The electronic device may further generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model, and control the display device to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio.
[0022] Conventional approaches to audio generation in virtual environments often rely on pre-recorded sound effects, basic spatial audio techniques, or simplified acoustic models. These methods may not fully capture the nuanced acoustic properties of complex 3D scenes or accurately represent the interactions between sound sources and virtual environments. Additionally, existing solutions may struggle to dynamically adapt audio characteristics based on changes in the virtual scene or user interactions, which may lead to inconsistencies between visual and auditory cues in the metaverse application. As metaverse technologies have advanced, there has been an increasing focus on enhancement of the realism and immersion of these virtual experiences, which may necessitate a more sophisticated approach to audio generation.
[0023] The disclosed technique may utilize deep learning and cross-modal encoding to generate acoustics-matched audio for 3D scenes in metaverse applications. Unlike traditional methods, this approach may process multi-view images of the virtual environment to determine volume aggregation information, which may be combined with source audio using a cross-modal encoder model. The cross-modal encoder model may facilitate an integration of visual and audio data, which may enable analysis and correlation of spatial characteristics of the 3D scene with the audio input. This may help to generate acoustically accurate, contextually relevant audio. The disclosed technique may help adapt real-time changes in the 3D scene, and enhanced immersion experience through better audio-visual coherence. The disclosed technique may have an ability to consider both speaker and listener perspectives, which may allow more realistic audio experiences in multi-user metaverse scenarios.
[0024] FIG. 1 is a diagram that illustrates a network environment for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shownan exemplary network environment 100. The network environment 100 includes an electronic device 102, a server 104, a database 106, and a communication network 108.The electronic device 102 is connected to the server 104 through the communication network 108. The server 104 may host the database 106. The electronic device 102 may include a metaverse engine 110, a deep learning model 112A, and a cross-modal encoder model 112B. FIG. 1 further shows multi-view images 114 and source audio 116, both of which may be stored on the server 104 and referenced on the database 106. The multiview images 114 and the source audio 116 may be associated with a metaverse application. FIG. 1 further shows an audio device 118 associated with the electronic device 102.
[0025] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the multi-view images 114 associated with the metaverse application. The multi-view images 114 may be received from the metaverse engine 110. The multi-view images 114 may correspond to a 3D scene associated with a first user of the metaverse application. The electronic device 102 may apply the deep learning model 112A on the multi-view images 114. The electronic device 102 may determine volume aggregation information associated with the 3D scene, based on the application of the deep learning model 112A. The volume aggregation information may include information related to the scene's geometry, materials, and spatial relationships. The electronic device 102 may receive the source audio 116, which may be associated with the first user. The electronic device 102 may apply the cross-modal encoder model 112B on the volume aggregation information and the source audio 116. The electronic device 102 may generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model 112B. The electronic device102 may control a display device to render the 3D for a second user (not shown in FIG. 1 ) of the metaverse application, based on the acoustics-matched audio.
[0026] The electronic device 102 may be a dedicated virtual reality (VR) or augmented reality (AR) device. The electronic device 102 may be configured to process and generate virtual environment data through the metaverse engine 110. The metaverse engine 110 may be a software component that, when executed by the electronic device 102, may render 3D environments, manage user interactions, and coordinate audio-visual experiences within the metaverse application. Examples of the electronic device 102 may include, but may not be limited to, Virtual Reality (VR) headsets, Augmented Reality (AR) glasses, motion controllers, haptic feedback devices, tracking systems, and display devices. The electronic device 102 may also be, for example, a desktop, a tablet, a television (TV), a laptop, a computing device, a smartphone, a cellular phone, a mobile phone, a consumer electronic (CE) device having a display.
[0027] The server 104 that may include suitable logic, circuitry, interfaces, and / or code configured to receive the multi-view images 114 from the electronic device 102. The server 104 may be configured to determine the volume aggregation information associated with the 3D scene and the acoustics-matched audio. The server 104 may further receive source audio 116 associated with the first user. The server 104 may be configured to train the deep learning model 112A to determine the volume aggregation information. The server 104 may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Example implementations of the servers 104 may include, but are not limited to, a database server, a file server, a web server, an application server, a mainframe server, a cloud computing server, or a combination thereof.
[0028] The server 104 may be a remote computing system that hosts various services and data required for the metaverse application. The server 104 may manage user accounts, store virtual assets, and coordinate multi-user experiences within the metaverse. The database 106 connected to the server 104 may store various types of data including the multi-view images 114 and the source audio 116. The multi-view images 114 may provide visual information about virtual environments from multiple perspectives, while the source audio 116 may contain audio data that may be processed and modified according to the virtual environment's characteristics.
[0029] The database 106 may include suitable logic, circuitry, interfaces, and / or code configured to store information, such as references (e.g., file paths or URLs) to the multiview images 114 and source audio 116. The database 106 may also store metadata associated with the 3D scene, metadata associated with each user, and metadata associated with a user experience associated with each user. For example, the database 106 may store volume aggregation information, acoustics-matched audio associated with the 3D scene. The database 106 may be stored or cached on a device or server, such as the server 104. The device storing the database 106 may be configured to query the database 106 for certain information such as, the multi-view images 114 and source audio 116, from the electronic device 102. The device storing the database 106 may be configured to retrieve the queried information (such as references to the multi-view images 114 and source audio 116) from the database 106 and may transmit the queried information to the electronic device 102.
[0030] In some embodiments, the database 106 may be hosted on a server located at the same or different locations. The operations of the database 106 may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or anapplication-specific integrated circuit (ASIC). In some other instances, the database 106 may be implemented using software.
[0031] The communication network 108 may include a communication medium through which the electronic device 102 and the server 104 may communicate with each other. The communication network 108 may be a wired or wireless communication network. The communication network 108 may facilitate data exchange between the electronic device 102 and the server 104, to enable real-time transmission of audio and visual information and allow delivery of acoustics-matched audio based on virtual environment parameters. Examples of the communication network 108 may include, but are not limited to, Internet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 108, in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity(Li-Fi), 802.16, IEEE 802.11 s, IEEE 802.11 g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0032] The metaverse engine 110 may be a software platform or framework to create, manage, and interact with virtual worlds and digital environments that collectively form the metaverse. However, the metaverse application may be a specific program or experiencebuilt using a metaverse engine and configured to execute by the electronic device 102. In an embodiment, the metaverse engine 110 associated with the electronic device 102 may execute the metaverse application. The metaverse engine 110 may include a 3D rendering engine to generate immersive virtual environments. The metaverse engine 110 may incorporate a spatial audio processing system to create acoustically matched audio experiences. The metaverse engine 110 may enable a multi-user interaction framework that may enable real-time collaboration in virtual spaces. Metaverse engine 110 may include a physics simulation module for realistic object interactions within the virtual world. The metaverse engine 110 may also include an asset management system for handling virtual objects, textures, and audio resources across different metaverse applications. The metaverse application may be a collective virtual shared space, created by the convergence of virtually enhanced physical reality and physically persistent virtual reality, including augmented reality (AR), virtual reality (VR), and other immersive digital experiences.
[0033] The deep learning model 112A may be configured to analyze and interpret complex visual data from the multi-view images 114 of virtual environments. The deep learning model may receive multi-view images as input. The deep learning model 112A may be configured for feature extraction and spatial information acquisition from the received multi-view images 114, which may enable accurate representation of the 3D scene. In some cases, the deep learning model 112A may undergo training on various types of 3D environments, which may allow generalization across different metaverse scenarios. The deep learning model 112A may be implemented as part of a processing pipeline that combines visual analysis with audio processing for generation of acoustically matched audio for virtual spaces. The deep learning model 112A may undergo fine-tuning or update based on user feedback or new data, which may lead to performanceimprovement over time in the creation of immersive metaverse experiences. The deep learning model 112A may output the volume aggregation information.
[0034] The cross-modal encoder model 112B may enable integration of spatial information derived from multiple modalities, such as, visual data with audio content. Implementation of the cross-modal encoder model 112B may allow for analysis of relationships between visual scenes and acoustic properties, which may facilitate more accurate audio transformations. The cross-modal encoder model 112B may include neural networks that perform learning to map between visual and auditory domains. The cross- modal encoder model 112B may be configured to generate the acoustics-matched audio that considers input of both the characteristics of the sound source and the properties of the virtual environment.
[0035] The deep learning model 112A and the cross-modal encoder model 112B may each be a neural network model having a plurality of layers with each layer forming a loop where the outputs of each element feed into the other elements. The plurality of layers of the neural network model may include an input layer, one or more hidden layers, and an output layer. The deep learning model 112A may receive multi-view images as input. The cross-modal encoder model 112B may receive the volume aggregation information as input. Each layer of the plurality of layers may include one or more nodes (or artificial neurons, represented by circles, for example). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network model. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network model. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network model. Such hyper-parameters may be set before, while training, or after training the neural network model on a training dataset. The deep learning model 112A may output the volume aggregation information. The cross-modal encoder model 112B may output the acoustics-matched audio.
[0036] Each node of the neural network model may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters, tunable during training of the neural network model. The set of parameters may include, for example, a weight parameter, a regularization parameter, and the like. Each node may use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the neural network model. All or some of the nodes of the neural network model may correspond to the same or a different mathematical function.
[0037] In training of the neural network model, one or more parameters of each node of the neural network model may be updated based on whether an output of the final layer for a given input (from the training dataset) matches a correct result based on a loss function for the neural network model. The above process may be repeated for the same or a different input until a minima of loss function is achieved, and a training error is minimized. Several methods for training are known in art, for example, gradient descent, stochastic gradient descent, batch gradient descent, gradient boost, meta-heuristics, and the like.
[0038] The neural network model may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The neural network model may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, such as, the electronic device 102. The neural network model may include code and routines configured to enable acomputing device to perform one or more operations. Additionally, or alternatively, the neural network model may be implemented using hardware including a processor, a microprocessor, a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network model may be implemented using a combination of hardware and software.
[0039] Examples of the deep learning model 112A may include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs) or long shortterm memory (LSTM) networks, transformer-based models, generative adversarial networks (GANs), autoencoders, graph neural networks, deep reinforcement learning models, deep learning models with attention mechanisms, variational autoencoders, and Siamese networks. The CNNs may processing multi-view images of virtual environments. The autoencoders may efficiently encode and decode audio and visual features. The graph neural networks (GNNs) may model complex spatial relationships in 3D environments. The deep reinforcement learning models may adapt audio processing based on user interactions. The attention mechanisms may focus on relevant features during cross-modal encoding. The variational autoencoders (VAEs) may generate diverse acoustic variations. The Siamese networks may compare acoustic properties between source audio and virtual environments.
[0040] Examples of the cross-modal encoder model 112B may include, but are not limited to, CNNs, RNNs, GANs, cross-model encoder model. The CNNs may be effective in capturing spatial and temporal patterns in the audio and visual data. The recurrent neural networks (RNNs) or long short-term memory (LSTM) networks may handle sequential audio data. The generative adversarial networks (GANs) may synthesize realistic acoustics-matched audio.
[0041] The audio device 118 may be associated with the electronic device 102 and may represent an audio input source within the virtual environment. In some cases, the audio device 118 may include be a microphone or other audio input device that captures the source audio 116 from a user of the metaverse application. In some cases, audio device 118 may represent a virtual audio source within the 3D scene of the metaverse environment. The audio device 118 may be implemented as a component of a VR headset worn by the first user to capture their voice or other sounds. In accordance with an embodiment, the audio device 118 may correspond to an array of multiple speakers used to capture spatial audio information from different positions within the real-world environment. The audio device 118 may serve as both an input device to capture a source audio and an output device for playback of an acoustics-matched audio to a user.
[0042] In operation, the electronic device 102 may be configured to execute the metaverse engine 110, the multi-view images 114 associated with a metaverse application. The multi-view images 114 may correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application. The multi-view images 114 may be a set of images captured from different viewpoints or angles around a particular scene or object. The multi-view images 114 may provide multiple perspectives of the same subject, which may allow more comprehensive analysis of the 3D scene. The multi-view images 114 may represent virtual environments or objects from various angles, which may be crucial for 3D reconstruction, rendering, and immersive experiences. The metaverse application may enable interaction with a virtual, immersive, and often interconnected digital environment. The metaverse may be a collective virtual shared space, created by the convergence of virtual reality, including augmented reality (AR), virtual reality (VR), and other immersive digital experiences. The reception of the multi-view images is described further, for example, in FIG. 3, FIG. 4, FIG.5, FIG. 8, and FIG. 13.
[0043] The electronic device 102 may be configured to apply the deep learning model 112A on the multi-view images 114. The application of the deep learning model is described further, for example, in FIG. 3. The electronic device 102 may be configured to determine volume aggregation information associated with the 3D scene, based on the application of the deep learning model 112A. The determination of the volume aggregation information may involve capture of multi-view images 114 of the 3D scene, application of the deep learning model 112A to the captured multi-view images 114, and calculation of a volume of objects or regions within the 3D scene. Considering, a virtual conference room scenario, the volume aggregation information might include data about the room's dimensions, the position and materials of walls and furniture, and the location of soundreflecting or absorbing surfaces. Such information may be crucial for accurate simulation of how sound would propagate and interact within the virtual space. The volume aggregation may include various steps such as data acquisition, feature extraction, 3D reconstruction, volume estimation, deep learning-based volume prediction, and aggregation and analysis. The determination of the volume aggregation information is described further, for example, in FIG. 3, FIG. 4, FIG.5, FIG. 8, and FIG. 13.
[0044] The electronic device 102 may be further configured to receive the source audio 116 associated with the first user. The source audio 116 may a user-generated audio, environmental sounds, a pre-recorded audio, a live audio, and the like. The audio device 118 connected to the electronic device 102 may represent an audio input source within the virtual environment. In some cases, the audio device 118 may be a microphone or other audio input device that may capture the source audio 116 from a user of the metaverse application. The reception of the source audio is described further, for example, in FIG. 4.
[0045] The electronic device 102 may be configured to apply the cross-modal encoder model 112B on the volume aggregation information and the source audio 116. The application of the cross-modal encoder model 112B to generate an acoustics-matched audio associated with the 3D scene may involve an integration of the volume aggregation information (visual data) and source audio 116 (auditory data) to produce audio that accurately reflects an acoustics of the 3D environment. The cross-modal encoder model 112B may correspond to a visual-acoustic multi-modal Al model including a first encoder model, a second encoder model, and a set of stacked decoder models or the stacked decoder blocks. The first encoder model may process the source audio 116. The source audio 116 may be processed using Convolution Neural Network (CNN) for spectrograms or Recurrent Neural Network (RNN) for raw audio waveforms to extract audio features, as a first encoded input. The first encoded input may be a feature vector that represents temporal and spectral properties of the audio data. The deep learning model 112A may be used to determine a spatial content including a head orientation and microphone position of the first user, as a second encoded input. The second encoded input may be a feature vector that represents the spatial context of the audio capture. The deep learning model 112A may be a simple feedforward neural network or a more complex architecture based on the complexity of the input data. For high complexity scenarios, such as applications requiring detailed and dynamic 3D scene reconstruction for real-time interaction, for example, processing high-resolution images to extract detailed features and integrating the high-resolution images to ensure accurate 3D scene reconstruction, supporting realtime updates and interactions for an immersive user experience. In contrast, for low complexity scenarios, such as applications needing basic 3D scene reconstruction for static environments, the deep learning model 112A may use a simple feedforward neural network to process lower-resolution images and extract basic features. The simplefeedforward neural network may support static environments with minimal updates, suitable for applications where real-time interaction is not critical. The first encoded input (i.e., the audio features) and the second encoded input (i.e., the spatial context) may be combined for multi-modal analysis using a concatenation or an attention mechanism. An output spatial audio may be determined based on an application the set of stacked decoder models on the first encoded input and the second encoded input. The output spatial audio may correspond to the acoustics-matched audio. The application of the cross-modal encoder model is described further, for example, in FIG. 8.
[0046] The electronic device 102 may be further configured to generate an acoustics- matched audio associated with the 3D scene, based on the application of the cross-modal encoder model 112B. The electronic device 102 may be configured to generate the acoustics-matched audio based on an up-sampled audio output. The up-sampled audio output generation may be performed on the audio output determined by applying the cross- modal encoder model. The generation of the acoustics-matched audio is described further, for example, in FIG. 8. The electronic device 102 may be configured to control the display device to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio. The control of an I / O device to render the 3D scene is described further, for example, in FIG. 12.
[0047] The metaverse application implemented by the various components of the network environment 100 may correspond to various immersive experiences. In some cases, the metaverse application may be a virtual musical concert, where users may experience realistic acoustics as if the user are present in a physical concert venue. In other cases, the metaverse application may be a virtual meeting or conference room, which may enable remote participants to interact in a spatially accurate audio environment. In some other cases, the metaverse application may be an electronic learning platform, ametaverse gaming application, a metaverse tourism application, or a metaverse shopping application, each of which may benefit from the acoustically matched audio to enhance a user immersion and interaction.
[0048] Unlike prior solutions that relied on pre-recorded sound effects or simplified acoustic models, the disclosed approach leverages deep learning and cross-modal modelbased encoding to generate acoustics-matched audio for 3D scenes in metaverse applications. The disclosed approach may offer an improved acoustic accuracy in complex virtual environments, real-time adaptation to changes in the 3D scene, and enhanced immersion through better audio-visual coherence. By processing multi-view images 114 and combining the processed images with source audio 116 using advanced Al models, the disclosed technique may create a more realistic and dynamic audio experience that adapts to the virtual environment and the user interactions.
[0049] FIG. 2 is a block diagram of an electronic device of FIG. 1 , in accordance with an embodiment of the disclosure. FIG. 2 is described in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown an exemplary block diagram 200. The block diagram 200 includes the electronic device 102 having a circuitry 202, memory 204, an input / output (I / O) device 206, and a network interface 208. The circuitry 202 may be communicatively coupled to the memory 204, the I / O device 206, and the network interface 208. The I / O device 206 may include a display device 206A.
[0050] The circuitry 202 may include suitable logic, circuitry, and interfaces that may be configured to execute program instructions associated with different operations to be executed by the electronic device 102. In operation, the circuitry 202 may execute instructions stored in the memory 204 to process data received through the network interface 208 or the I / O device 206. For example, the circuitry 202 may apply the deep learning model 112A to the multi-view images 114 received from the metaverse engine110. The circuitry 202 may also process the source audio 116 using the cross-modal encoder model 112B to generate acoustics-matched audio based on the application of the deep learning model 112A and the cross-modal encoder model 112B.
[0051] The circuitry 202 may include one or more specialized processing units. These units may be configured as a single integrated processor or a cluster of processors working together to perform specific functions. The circuitry 202 may be built using various processor technologies known in the industry. Examples of these technologies include x86-based processors, Graphics Processing Units (GPUs), Reduced Instruction Set Computing (RISC) processors, Application-Specific Integrated Circuits (ASICs), Complex Instruction Set Computing (CISC) processors, microcontrollers, Central Processing Units (CPUs), and other types of computing circuits.
[0052] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store the program instructions to be executed by the circuitry 202. The program instructions stored on the memory 204 may enable the circuitry 202 to execute operations of the circuitry 202 (and / or the electronic device 102). In at least one embodiment, the memory 204 may store the multi-view images 114 and the source audio 116 associated with the first user. In some cases, the memory 204 may store the deep learning model 112A and the cross-modal encoder model 112B. Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read- Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card. The memory 204 may be configured to store data and instructions for the electronic device 102. The memory 204 may include various types of storage, such as random-access memory (RAM), read-only memory (ROM), or non-volatile storage devices.
[0053] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive multimedia content and provide an output based on the received input. For example, the I / O device 206 may receive the multi-view images 114 and the source audio 116 as inputs. The I / O device 206 may be configured to render a 3D scene of the metaverse application and facilitate user interaction with the electronic device 102. The I / O device 206 may include input devices such as a keyboard, a mouse, a touchscreen, or a microphone, or a joystick, and output devices such as display devices, speakers, or haptic feedback systems.
[0054] The display device 206A may include suitable logic, circuitry, and interfaces that may be configured to present visual information to a user, such as through rendering of the 3D scenes of the metaverse application. The display device 206A may be a touch screen which may enable the user to provide a user-input via the display device 206A. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 206A may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 206A may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro- chromic display, an AR / VR headset, or a transparent display.
[0055] The network interface 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the electronic device 102 and the server 104 via the communication network 108. The network interface 208 may be implemented by use of various known technologies to support wired or wireless communication of the electronic device 102 with the communication network 108. Thenetwork interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0056] The network interface 208 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11 a, IEEE 802.11 b, IEEE 802.11g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0057] FIG. 3 is a diagram that illustrates an exemplary execution pipeline for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure. FIG. 3 is described in conjunction with elements from FIG. 1 and FIG. 2. With reference to FIG. 3, there is shown an exemplary execution pipeline 300. The execution pipeline 300 includes operations such as, a multi-view image reception 302, a deep learning model application 304, volume aggregation information determination 306, a source audio reception 308, a cross-modal encoder model application 310, an acoustics-matched audio generation 312, and a 3D scene rendering control 314. The execution pipeline 300 may be implemented on theelectronic device 102 and may be executed by the circuitry 202. The execution pipeline 300 may be configured to analyze the multi-view images 114 and the source audio 116 to generate acoustics-matched audio for 3D scenes in a metaverse application.
[0058] At 302, the reception of the multi-view images may be performed. The circuitry 202 may be configured to receive the multi-view images 114 from the metaverse engine 110. The multi-view images 114 may include visual data necessary for creation of an immersive 3D environment. The multi-view images 114 correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application. Such images may be captured from multiple angles or perspectives to provide a comprehensive view of the 3D scene.
[0059] For example, in a virtual musical concert scenario, the multi-view images 114 might include various angles of the stage, audience areas, and surrounding environment. This may allow users (for example, the first user) to experience the concert from different vantage points, such as front-row seats, balcony views, or even on-stage perspectives. The quality and quantity of these multi-view images 114 may impact the final rendered 3D scene. Higher resolution images or a greater number of viewpoints may lead to a more detailed and accurate 3D representation, which may enhance the user's sense of presence in the metaverse. In some cases, the circuitry 202 may dynamically adjust the number of viewpoints based on the complexity of the scene or the available computational resources. The complexity may refer to the level of details involved in the virtual musical concert scene. This may be influenced by several factors, such as the number of viewpoints, the resolution of images, and the overall richness of the environment. Considering an example of a simple stage setup with a few static audience members and basic lighting, the electronic device 102 may use fewer viewpoints and lower resolution images to render the simple stage setup, requiring less computational power. A detailed stage with dynamiclighting, moving audience members, and intricate decorations. The electronic device 102 may need to capture and process high-resolution images from multiple angles (front-row seats, balcony views, on-stage perspectives) to create a realistic and immersive 3D representation. This higher complexity may need more computational resources to maintain performance and visual quality.
[0060] At 304, the deep learning model 112A may be applied on the multi-view images 114. The circuitry 202 may be configured to apply the deep learning model 112A on the multi-view images 114. Based on the application of the deep learning model 112A, meaningful features and spatial information associated with the multi-view images 114 may be extracted. For instance, in a metaverse shopping application, the deep learning model 112A may analyze multi-view images 114 of a virtual store to identify product placements, store layout, and even crowd density. The analyses of the multi-view images 114 may be crucial for creation of an acoustically accurate representation of the 3D environment. The deep learning model 112A may be trained on various types of 3D environments, which may allow the deep learning model 112A to generalize across different metaverse scenarios. In some implementations, the deep learning model 112A may be fine-tuned or updated based on user feedback or new data, which may improve a performance of the deep learning model 112A over time.
[0061] At 306, the determination of the volume aggregation information may be performed. The circuitry 202 may be configured to determine the volume aggregation information associated with the 3D scene, based on the application of the deep learning model 112A. For example, the circuitry 202 may synthesize spatial characteristics of the 3D scene derived from the multi-view images 114. The volume aggregation information may include information related to the scene's geometry, materials, and spatial relationships.
[0062] For example, in a virtual conference room scenario, the volume aggregation information might include data about the room's dimensions, the position and materials of walls and furniture, and the location of sound-reflecting or absorbing surfaces. Such information may be crucial for accurate simulation of how sound would propagate and interact within the virtual space. The accuracy of this volume aggregation information may significantly impact the realism of the final acoustics-matched audio. In some implementations, the electronic device 102 may use probabilistic methods to handle uncertainties in the scene reconstruction, which may potentially lead to more robust audio generation in complex or partially occluded environments.
[0063] At 308, the reception of the source audio may be performed. The circuitry 202 may be configured to receive the source audio 116 associated with the first user. The circuitry 202 may receive an original audio content that may be required to be modified to match acoustic properties of the 3D scene. The circuitry 202 may receive the source audio 116 from the server 104 through the communication network 108. In some cases, the source audio 116 may be received from the audio device 118 associated with the first user.
[0064] In a virtual meeting scenario, the source audio 116 may include, for instance, a voice of a participant speaking. The source audio 116 may be required to be converted to a sound that may occur within the virtual meeting space, based on factors like room acoustics and the positions of other participants. The quality and characteristics of the source audio 116 may influence the final output. High-quality, clear audio inputs may generally lead to better results, but the circuitry 202 may be configured to handle various audio qualities and types, from speech to music to ambient sounds.
[0065] At 310, the application of the cross-modal encoder model may be performed. The circuitry 202 may be configured to apply the cross-modal encoder model 112B on the volume aggregation information and the source audio 116. The circuitry 202 may integratespatial characteristics and acoustic properties derived from the visual data with the audio content of the 3D metaverse environment to obtain the volume aggregation information. The electronic device may receive the source audio associated with the first user and apply the cross-modal encoder model on the volume aggregation information and the source audio. The electronic device may further generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model, and control the display device to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio. The cross-modal encoder model 112B may be analyze relationships between visual scenes and acoustic properties, which may allow more accurate audio transformations.
[0066] For example, in a metaverse gaming application, the cross-modal encoder model 112B may analyze both a layout of a virtual battlefield (from the volume aggregation information) and a sound of an explosion (from the source audio 116). The circuitry 202 may use the cross-modal encoder model 112B to determine how that explosion should sound from different positions within the battlefield, based on various factors like distance, obstacles, and environmental materials. A complexity of the cross-modal encoder model 112B may vary based on a type of the metaverse application. In some cases, the cross- modal encoder model 112B may include complex neural networks that learn to map between visual and auditory domains. In other cases, the cross-modal encoder model 112B may use more traditional signal processing techniques guided by the visual information. For high complexity scenarios, for example, a fully immersive VR game with real-time audio feedback, the model might include advanced neural networks like CNNs and LSTMs, use attention mechanisms or transformers for cross-modal mapping, and may require immediate processing to provide accurate and realistic audio that reflects the spatial and acoustic properties of 3D environment within the VR game. In contrast, for lowcomplexity scenarios, such as a basic AR application with simple audio cues, the cross- modal encoder model may rely on traditional signal processing techniques guided by visual information, use simpler neural networks like basic feedforward networks if needed, resulting in basic audio feedback without complex spatial characteristics.
[0067] At 312, the generation of the acoustics-matched audio may be performed. The circuitry 202 may be configured to generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model 112B. The generated acoustics-matched audio produced may match the acoustic properties of the 3D scene. The acoustics-matched audio may represent the source audio 116 transformed to sound that may occur within the virtual environment.
[0068] For instance, in a virtual tourism application, the circuitry 202 may make a tour guide's voice echo realistically in a virtual model of a grand cathedral building, or sound muffled and distant when the user (for example, the first user) virtually steps outside the building. The acoustics-matched audio may enhance the sense of presence and immersion in the virtual space. The quality of the acoustics-matched audio may depend on factors such as the accuracy of the volume aggregation information, the sophistication of the cross-modal encoder model 112B, and the computational resources available. In some implementations, the circuitry 202 may offer different levels of audio quality, which may allow users to balance realism against performance on less powerful devices.
[0069] At 314, control of 3D scene rendering may be performed. The circuitry 202 may be configured to control the display device 206A to render of the 3D scene for the second user of the metaverse application, based on the acoustics-matched audio. The control of the display device 206A to render of the 3D scene may include a coordination of the visual and auditory elements to create a cohesive and immersive experience for the user (forexample, the second user). The rendering instructions may be sent to the display device 206A and audio output instructions to the audio device 118.
[0070] For example, in an electronic learning platform within the metaverse, the circuitry 202 may ensure that a virtual lecturer's voice sounds appropriate to the lecturer’s position in a virtual lecture hall, while the visual elements of the hall and the lecturer's avatar may be simultaneously rendered. The electronic device 102 may need to dynamically adjust both for visual and audio rendering as users move or interact within the virtual space. The control of display device 206A to render visual elements of the hall and the lecturer's avatar may also involve optimization of the experience for different hardware setups. For instance, the electronic device 102 may be configured to adjust the audio output for stereo speakers, surround sound systems, or headphones, to ensure optimal spatial audio experience given the user's available audio output devices.
[0071] FIG. 4 is a diagram that illustrates an exemplary scenario of a metaverse engine, in accordance with an embodiment of the disclosure. FIG. 4 is described in conjunction with elements from FIG. 1 , FIG. 2, and FIG. 3. With reference to FIG. 4, there is shown an exemplary scenario 400. The scenario 400 includes an input module 402, a metaverse environment 404, an audio input 406, a visual image of metaverse 408, an acousticsmatching engine 410, a style transfer module 412, an output module 414, acoustics- matched audio 416, remixed audio 418, an acoustics-matched waveform 420, and an acoustics-matched and remixed waveform 422. In an embodiment, the operations of the input module 402, acoustics-matching engine 410, a style transfer module 412, an output module 414, may be performed by the circuitry 202 of the FIG. 2 or the electronic device 102 of the FIG. 1.
[0072] The input module 402 may be configured to receive inputs from the metaverse environment 404 (for example, image or text attributes) and the audio input 406 (forexample, audio attributes). The metaverse environment 404 may provide comprehensive information about the virtual environment, which may include detailed descriptions of the target environment in the form of high-resolution images, descriptive text, or the multi-view images 114 captured from various angles and perspectives. The metaverse environment 404 may include, but not limited to, images, text, or the multi-view images 114. The input for the metaverse environment 404 may be provided by the user (for example, the first user) or may be generated using a generative Artificial Intelligence (Al) model. For instance, in a virtual art gallery scenario, the metaverse environment 404 may provide floor plans, 3D models of exhibition spaces, and even information about the materials used in the construction of the virtual gallery walls and floors. The audio input 406 may correspond to the source audio 116 associated with the user (for example, the first user) in the metaverse application, which may range from voice recordings to ambient sounds or music tracks. The visual image of metaverse 408 may be of a virtual musical concert, a virtual meeting, a virtual conference room, an electronic learning platform, and the like. The input module 402 may transmit the received input to the metaverse engine 110. The inputs such as the metaverse environment 404 and the audio input 406 may be fed to the acousticsmatching engine 410 and the visual image of metaverse 408 may be fed to the style transfer module 412.
[0073] The acoustics-matching engine 410 may be connected to the input module 402 and may receive inputs from both the metaverse environment 404 and the audio input 406. The acoustics-matching engine 410 may be configured to process and analyze these inputs to create acoustically accurate audio outputs. The acoustics-matching engine 410 may be configured to operate in two modes: offline mode and online mode, each serving different purposes and offering unique advantages.
[0074] In the offline mode, the acoustics-matching engine 410 may process predefined inputs that include detailed descriptions of the target environment and the source audio 116. The detailed descriptions may include information about the size, shape, materials, and other acoustic properties of the virtual space. The offline mode may be used to prerender audio for static environments or prepare audio assets in advance. The audio assets may be processed and adjusted to sound as if the sound is being played within the virtual space, based on factors such as reverberation, echo, and sound diffusion. For example, in a virtual museum application, the offline mode may be used to pre-compute the acoustic properties of different exhibition halls, which may allow for quick loading and seamless audio transitions as users move through the virtual space. This approach may significantly reduce real-time computational requirements and may enhance the overall performance of the metaverse application.
[0075] In the online mode, the acoustics-matching engine 410 may receive inputs in real-time from the metaverse environment 404. This mode may allow dynamic audio generation that responds to changes in the virtual environment or user interactions. For instance, in a virtual conference room scenario, during the online mode, the acousticsmatching engine 410 may adapt the audio in real-time as participants move around the space, open virtual windows, or interact with objects in the room. This dynamic adaptation may create a more immersive and responsive audio experience, which may enhance a sense of presence for users in the metaverse.
[0076] The acoustics-matching engine 410 may generate the acoustics-matched audio 416 based on the inputs from the metaverse environment 404 (for example, image attributes) and the audio input 406 (for example, audio attributes). This process may involve techniques that may simulate sound propagation, reflection, and absorption based on the virtual environment's characteristics. The acoustics-matched audio 416 may berepresented as the acoustics-matched waveform 420, which may indicate the audio's characteristics such as amplitude and frequency over time.
[0077] The visual image of metaverse 408 may be received at the style transfer module 412, which may introduce an approach to audio customization based on visual cues. The style transfer module 412 may include an audio customization module (not shown in FIG. 4) that modifies the acoustics-matched audio 416 based on visual characteristics of the target environment. The style transfer module 412 may output updated audio attributes of the acoustics-matched audio 416 based on the visual image of metaverse 408. These visual characteristics may include, for example, a color tone, an exposure, a brightness, a contrast, and a saturation of the visual image of metaverse 408. The visual characteristics may allow a unique synergy between visual and auditory elements in the virtual space.
[0078] The audio customization module may employ various techniques to align audio characteristics with visual elements. For example, the audio customization module may vary the intensity of the audio based on the color tone of the target environment. In a virtual sunset scene, darker, warmer tones might result in more intense, richer audio output that may enhance an emotional impact of the experience. The audio customization module may also adjust the amplitude of the audio based on the exposure, brightness, and saturation of the target environment. In a bright, vibrant virtual marketplace, this may translate to louder, more energetic sounds, while a dimly lit virtual library may feature softer, more subdued audio.
[0079] The audio customization module may modify the frequency of the audio based on the contrast of the target environment. High-contrast scenes, such as a virtual cityscape with stark light-dark differences, might feature a wider range of audio frequencies, creating a more dynamic soundscape. This intricate relationship between visual and audioelements may significantly enhance the immersive quality of the metaverse experience, which may create a more cohesive and engaging virtual world.
[0080] The style transfer module 412 may generate the remixed audio 418 based on these modifications and effectively create a version of the audio that is not only acoustically matched to the virtual space but also stylistically aligned with its visual characteristics. The remixed audio 418 may be represented as the acoustics-matched and remixed waveform 422, visually indicating the combined effects of acoustics-matching and style transfer.
[0081] The output module 414 may be connected to both the acoustics-matching engine 410 and the style transfer module 412 and may serve as the final stage in the audio processing pipeline. This module may receive the acoustics-matched audio 416 and the remixed audio 418 and may output these audio signals for playback through the audio device 118 or other audio output devices. The output module 414 may be responsible for optimization of the audio output based on the user's hardware capabilities, which may ensure the best possible audio experience across different devices and setups.
[0082] In some cases, the scenario 400 may be implemented as part of various metaverse applications, to showcase its versatility and potential impact across different virtual experiences. For instance, in a virtual musical concert application, the acousticsmatching engine 410 may process the audio input 406 of a performer's voice or instrument, while the metaverse environment 404 may provide detailed information about the virtual concert venue, including its size, shape, and materials. The style transfer module 412 may modify the acoustics-matched audio 416 based on the visual characteristics of the concert venue, such as its lighting effects and stage design, to create a more immersive and emotionally resonant audio experience for users.
[0083] In an e-learning platform within the metaverse, the circuitry 202 may enhance the educational experience based on a creation of acoustically accurate virtual classrooms orlecture halls. The audio of a virtual lecturer may be processed to sound as if the voice is produced in a specific virtual space, with additional modifications based on visual cues like the classroom's decor or the time of day represented in the virtual environment. This may help to maintain student engagement and create a more immersive learning experience.
[0084] For virtual tourism applications, the circuitry 202 may recreate acoustic properties of famous landmarks or historical sites. Visitors to a virtual recreation of the Sistine Chapel, for example, may experience not only the visual grandeur of the space but also its unique acoustic properties, with the audio further enhanced based on the chapel's ornate visual details and lighting conditions.
[0085] The scenario 400 may enable the generation of realistic and immersive audio experiences in various metaverse applications by combining sophisticated acousticsmatching techniques with innovative visual-based audio customization. This approach may significantly enhance the overall user experience in virtual environments based on the acoustics-matched audio that not only matches the acoustic properties of the virtual space but also aligns with its visual characteristics. By creation of such deep integration between visual and auditory elements, the circuitry 202 may contribute to a more believable, engaging, and emotionally impactful metaverse experience, which may revolutionize how users interact with and perceive virtual worlds. It should be noted that the scenario 400 of FIG. 4 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0086] FIG. 5 is a diagram that illustrates an exemplary scenario of an audio processing system for a metaverse application, in accordance with an embodiment of the disclosure. FIG. 5 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, and FIG. 4. With reference to FIG. 5, there is shown an exemplary scenario 500. The scenario 500 includes a first user 502, a second user 518, the audio device 118, the source audio 116,a head orientation 504, the acoustics-matching engine 410, the acoustics-matched audio 416, an ambisonics conversion module 506, a sound field module 508, a perspective audio module 510, a game engine 512, a 3D scene 514, and listener variables 516. In an embodiment, the operations of the acoustics-matching engine 410, the ambisonics conversion module 506, the sound field module 508, perspective audio module 510, and the game engine 512 may be performed by the circuitry 202 of the FIG. 2 or the electronic device 102 of the FIG. 1.
[0087] The audio device 118 may serve as a producer of the source audio 116 within the metaverse environment 404 and may correspond to a user (e.g., the first user 502) who interacts with the electronic device 102. The audio device 118 may be fundamental to the audio processing pipeline, as the audio device 118 may provide the audio input 406 that may be transformed and adapted to the virtual space. The audio input 406 may be interchangeably referred as audio. The source audio 116 generated by the speaker may encompass a wide range of audio types, from voice communications to ambient sounds or music, based on the specific metaverse application. For instance, in a virtual classroom scenario, the source audio 116 might be a lecturer's voice, while in a virtual concert, the source audio 116 may be a live musical performance.
[0088] The head orientation 504 of the audio device 118 may be a crucial parameter that influences how sound is perceived and propagated within the virtual environment. Head orientation data may be captured through sensors in VR headsets or other input devices, which may allow the circuitry 202 to accurately represent a directionality of sound based on a speaker's position and facing direction. For example, if a user turns the user’s head while the user speaks in a virtual meeting room, the circuitry 202 may adjust an audio output to reflect this change in orientation, which may enhance a sense of spatial presence for other participants. The head orientation data may be captured using sensors embeddedin VR headsets or other input devices. The sensors may track the position and facing direction of the user’s head in real-time. The accurate representation of sound directionality may improve communication and interaction within the virtual environment. Participants may better understand where the sound is coming from, which may make the virtual experience more intuitive and engaging.
[0089] The acoustics-matching engine 410, which may be referred as a visual / 3D acoustics-matching Al engine, may play a pivotal role in creation of realistic and immersive audio experiences within the metaverse environment 404. The acoustics-matching engine 410 may process the source audio 116 in conjunction with environmental data provided by the game engine 512, such as information related to the 3D scene 514. The acousticsmatching engine 410 may employ advanced techniques that simulate complex acoustic phenomena like reflection, refraction, and absorption of sound waves. In a virtual museum scenario, for instance, the acoustics-matching engine 410 may analyze the dimensions of exhibition halls, the materials of walls and displays, and the presence of other visitors to accurately model how sound may propagate and interact within that space.
[0090] The game engine 512 and the 3D scene 514 may provide the foundational context for the entire audio processing system. The game engine 512 may continuously update the virtual environment's state, including user positions, object interactions, and environmental changes. This real-time information may be crucial to maintain audio-visual coherence within the metaverse environment 404. For example, in a dynamic virtual city environment, the game engine 512 may notify the acoustics-matching engine 410 about changes such as cars passing by, doors opening, or weather effects, to allow appropriate adjustments to the audio output.
[0091] The ambisonics conversion module 506 may execute spatial audio processing for virtual environments. Based on application of an Al-based ambisonics conversionmodule 506 to the head orientation 504 and acoustics-matched audio 416, ambisonics conversion module 506 may create a three-dimensional sound field that accurately represents how audio would be perceived from different positions and orientations within the virtual space. The ambisonics conversion module 506 may generate a sound field that surrounds a user. The sound field may change dynamically based on the user’s movements and changes in the head orientation 504, which may provide a highly immersive audio experience. In a virtual concert scenario, the ambisonics conversion module 506 may allow users to experience realistic changes in sound as the users may move around a venue. For example, the audio may be perceived differently when the user faces a stage versus when the user faces away from the stage. As users move around the virtual venue, the ambisonics conversion module 506 may continuously update the sound field to reflect changing positions and orientations of the users. This may result in realistic changes in how the audio may be perceived and may enhance a sense of presence and immersion.
[0092] The sound field module 508 may build upon the output of the ambisonics conversion module 506 to create a comprehensive representation of how sound propagates from the user’s position throughout the virtual environment. The operations of the sound field module 508 may be based on factors such as distance attenuation, occlusion by virtual objects, and the unique acoustic properties of different areas within the 3D scene 514. For instance, in a virtual office environment, the sound field module 508 might model how a user’s voice carries differently in an open plan area versus a small, enclosed meeting room.
[0093] The perspective audio module 510 may represent an audio processing model, which may consider a specific position and orientation of a listener (i.e. , the second user 518) within the virtual environment. The perspective audio module 510 may receivelistener variables 516 from the game engine 512. Based on the listener variables 516 provided by the game engine 512, the perspective audio module 510 may tailor an audio experience to each individual user. In a virtual theme park scenario, for example, the perspective audio module 510 would ensure that a user who stands near a virtual roller coaster hears the screams and mechanical sounds appropriately based on the user’s position, while another user farther away may hear a muted version of the same sounds. The listener variables 516 may help to adjust the audio output based on the listener’s position, orientation, and other factors. The listener variables 516 may represent the current position of the listener (for example, the second user) in the game world. The listener variables 516 may include, but is not limited to, a listener position, a listener orientation, a listener velocity, a listener environment, a listener Head-Related Transfer Function (HRTF), a listener occlusion and obstruction, and the like.
[0094] Based on factors such as environmental acoustics, speaker orientation, sound propagation, and listener position, the electronic device 102 may create a virtual sound field (i.e. , an audio environment around a user) that may closely mimic a real-world sound field. This audio realism significantly may enhance an overall immersion and sense of presence in virtual environments and make interactions feel more natural and engaging.
[0095] The ability of the metaverse environment 404 of the electronic device 102 may adapt in real-time for dynamic and responsive audio experiences. As users move through different virtual spaces or as the environment itself changes, the audio may adjust seamlessly and maintaining a consistent and believable auditory landscape. This real-time adaptability may be particularly valuable in interactive metaverse applications, such as virtual social gatherings, educational simulations, or complex multiplayer games, where the audio environment may need to react quickly to user actions and environmentalchanges. It should be noted that the scenario 500 of FIG. 5 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0096] FIG. 6 is a diagram that illustrates an exemplary scenario for generation of acoustics-matched audio in a virtual environment, in accordance with an embodiment of the disclosure. FIG. 6 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, and FIG. 5. With reference to FIG. 6, there is shown an exemplary scenario 600. The scenario 600 includes an environment image 602 (or 3D environment image), environment keywords 604, an environment prompt 606 or a 3D environment describing prompt, an environment description prompt 608, a scene generation module 610, a space scene generator 612, a visual-acoustics matcher 614, the source audio 116, and the acoustics-matched audio 416. The scenario 600 may include operations for generation of the acoustics-matched audio 416 in a virtual environment. The various operations of the scenario 600 may be executed by any computing device, such as, the electronic device 102 or the circuitry 202.
[0097] In some cases, the circuitry 202 may be configured to receive a user input associated with the 3D scene 514. This user input may correspond to at least one of the environment image 602 associated with the 3D scene 514, the environment keywords 604 associated with the 3D scene 514, or the environment prompt 606 associated with the 3D scene 514. Such inputs may serve as different ways to describe the target environment in the metaverse application.
[0098] The environment image 602 may be a visual representation of the desired 3D scene 514. The environment image 602 may be a single image or a series of images that depict the environment. The images may help in to analyze visual aspects and layout of the 3D scene 514. For example, in a virtual tourism application, the environment image 602 may be a photograph of a real-world location that needs to be recreated in themetaverse. The environment keywords 604 may be a set of descriptive terms associated with the 3D scene 514. In a virtual conference room scenario, the environment keywords 604 may include terms such as "spacious," "modem," "glass walls," or "circular table." The environment prompt 606 may be a first prompt, which may be a textual description of the desired 3D scene 514. For instance, in a metaverse gaming application, the environment prompt 606 may be a detailed description of a fantasy landscape, including information about terrain, vegetation, and atmospheric conditions.
[0099] In some cases, the circuitry 202 may be configured to generate a second prompt associated with the 3D scene 514 based on the environment image 602 and the environment keywords 604. The second prompt may be represented by the environment description prompt 608, which may combine and synthesize information from multiple input sources (e.g., the environment to create a comprehensive description of the target environment.
[0100] The circuitry 202 may be configured to apply a generative artificial intelligence (Al) model on the user input. This generative Al model may be implemented within the scene generation module 610. The scene generation module 610 may process the environment description prompt 608 and / or the environment prompt 606 to generate a detailed representation of the 3D scene 514. By leveraging the generative Al model, the scene generation module 610 may create a comprehensive and detailed 3D scene 514 that accurately reflects the information provided in the environment description prompt 608. This may include generation of various elements such as terrain, buildings, objects, lighting, and other environmental features.
[0101] Generative Al, or Gen Al, may refer to a type of artificial intelligence that is designed to create new content. This content may be in various forms, such as text, images, audio, or even video. The generative Al may produce original outputs based onthe patterns and information it has learned from existing data. The scene generation in generative Al models involves creating detailed and coherent visual scenes based on user inputs, such as text descriptions or other prompts. The user input provides a prompt, which may be a detailed text description of the scene to be generated. For example, "A sunny beach with palm trees and a sailboat in the distance". The gen Al model interprets the input using natural language processing (NLP) techniques to understand the key elements and context of the scene. This involves breaking down the text into meaningful components like objects, actions, and settings. The gen Al model may use the learned knowledge from training data to compose the scene. This step may involve arranging the elements in a coherent and visually appealing manner. For instance, placing the palm trees on the beach and the sailboat in the water. The gen Al model generates the image using techniques such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), or Diffusion Models. The technique may help creating realistic and high-quality images by refining the initial composition.
[0102] A discriminator model may be trained using embeddings associated with both sustainable materials and unsustainable materials included in the set of materials 116. The training may be such that the discriminator model may classify whether an output, generated by the generator model, is associated with a sustainable material. The generator model may be trained to generate an output for the queried material such that the discriminator model may accurately predict whether the queried material is sustainable. Thus, based on the training, the generative Al model may be configured to predict whether the queried material is sustainable. Examples of the generative Al model may include, but are not limited to, a Generative Adversarial Network (GAN) model, a variational autoencoder (VAE) mode, an auto-regressive model, a Generative Pre-trained Transformers (GPT) model, or a large language model (LLM).
[0103] The space scene generator 612 may be connected to the scene generation module 610. In some cases, the circuitry 202 may be configured to generate the 3D scene 514 based on the application of the generative Al model. The space scene generator 612 may create a virtual representation of the target environment based on the output from the scene generation module 610. The virtual representation may include detailed information about the geometry, materials, and spatial relationships within the 3D scene 514. In an example, the environment description prompt 608 may include, ‘picturesque landscape with a river and mountains in the background, multiple people aboard a boat on the river, and bright sunshine. The generative Al model may interpret the textual description of the environment and generate a detailed 3D representation of the space scene, based on the interpretation of the textual description of the environment. Once all elements as provided in the environment description prompt 608 are generated, the scene generation module 610 may render the 3D space scene. The 3D space scene may include high-quality visuals that accurately reflect the described environment, with realistic textures, lighting, and spatial relationships. The generated space scene may be presented to the user (for example, the second user) in various formats, such as a static image, an interactive 3D model, or a virtual reality experience. The second user may explore the 3D space scene, view the 3D space scene from different angles, and interact with the elements within the virtual space.
[0104] Examples of the 3D space scene generator 612 may include, but are not limited to, natural language processing models that interpret textual descriptions to create corresponding 3D environments. The 3D space scene generator 612 may include generative adversarial networks (GANs), transformer-based models, neural radiance field (NeRF) models, graph neural networks (GNNs), multi-modal learning systems, reinforcement learning models, language-guided 3D asset retrieval and placementsystems, semantic segmentation models, and text-to-3D mapping networks. The generative adversarial networks (GANs) may transform text inputs into detailed 3D scene layouts. The transformer-based models may generate 3D object placements and relationships based on textual descriptions. The neural radiance field (NeRF) models may produce 3D scenes with realistic lighting and textures from text prompts. The graph neural networks (GNNs) may construct scene graphs from text, which may be used to generate 3D environments. The multi-modal learning systems may combine text understanding with visual generation to create coherent 3D spaces. The reinforcement learning models may iteratively refine 3D scene generation based on text-derived goals and constraints. The language-guided 3D asset retrieval and placement systems may assemble scenes using pre-existing 3D models. The semantic segmentation models may interpret text to define spatial regions and their properties in generated 3D scenes. The text-to-3D mapping networks may directly translate textual features into 3D geometric and textural features.
[0105] The visual-acoustics matcher 614 may be connected to the space scene generator 612 and may receive the source audio 116 as an input. The visual-acoustics matcher 614 may analyze the generated 3D scene 514 from the space scene generator 612 and the source audio 116 to produce acoustically matched audio. This process may involve simulating how the source audio 116 would sound if the sound were produced within the generated 3D environment, based on factors such as room acoustics, material properties, and spatial relationships. In some cases, the generation of the 3D scene 514 may be further based on at least one of the environment prompt 606 or the environment description prompt 608. This approach may allow for flexibility in how the target environment is specified, which may accommodate various input methods and levels of detail. It should be noted that the scenario 600 of FIG. 6 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0106] FIG. 7 is a diagram that illustrates an exemplary scenario for generation of acoustically matched audio in a metaverse environment, in accordance with an embodiment of the disclosure. FIG. 7 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, and FIG. 6. With reference to FIG. 7, there is shown an exemplary scenario 700. The scenario 700 includes the metaverse environment 404, the game engine 512, a metaverse scene 702, an acoustics-matching engine 410, and the source audio 116. In an embodiment, the operations of the game engine 512 and the acoustics-matching engine 410 may be performed by the circuitry 202 of the FIG. 2 or the electronic device 102 of the FIG. 1 .
[0107] The metaverse environment 404 may represent a virtual space where the users may interact within the metaverse application. The metaverse environment 404 may encompass various virtual settings, such as virtual conference rooms, virtual classrooms, or virtual entertainment venues. The metaverse environment 404 may provide the context and spatial information necessary for generation of acoustically matched audio.
[0108] The game engine 512 may be coupled with the metaverse environment 404. The game engine 512 may be responsible for creation and management of virtual environments. Also, the game engine 512 may control the visual aspects of the metaverse environment 404. The game engine 512 may process the input from the metaverse environment 404 using various modules and techniques. The input processing may include interpretation of textual descriptions, analysis of images, and analysis of user interactions. Generative Al models may be employed to enhance the interpretation and generation of complex scenes based on the provided inputs. In some cases, the game engine 512 may process the spatial data from the metaverse environment 404 to create a detailed metaverse scene (e.g., the metaverse scene 702).
[0109] In an example, the environment description prompt 608 may be ‘a futuristic city with towering skyscrapers, flying cars, neon lights, and central part with a fountain’. The keywords for this environment description prompt 608 may include, for example, futuristic city, skyscrapers, flying cars, and the like. Sensors associated with the electronic device 102 (for example VR device) may collect sensor data associated with a user using the metaverse application. The sensor data may include head orientation 504 of the user, position of the user, hand movements of the user, and interactions with virtual objects within the virtual environment. Environmental data (such as time of day, weather conditions, and the like) may be collected. The pre-rendered audio or the source audio 116 that matches the acoustics of different areas within the city may be selected. For example, an echo of footsteps in narrow alleys and an ambient noise of the central park. The game engine 512 may process the environment description prompt 608 and keywords using generative Al model to analyze the desired elements and layout of the virtual city. The game engine 512 may generate 3D models of towering skyscrapers with futuristic designs, flying cars that move along designed paths, and neon lights that illuminate the streets. Spatial audio may be processed to match the user’s position and orientation. For example, the sound of flying cars may change as they pass by, and the ambient noise of the central park may be heard more prominently when the user is near the fountain. The game engine 512 may render the virtual city scene in real-time and provide high-quality visuals that may be displayed to the user through the electronic device 102 (for example, VR device). The scene may be continuously updated based on the user’s movements and interactions, which may ensure a smooth and immersive experience.
[0110] The metaverse scene 702 may be a comprehensive representation of the current state of the metaverse environment 404. This scene may include information about the geometry, materials, and spatial relationships of objects within the virtual space. Themetaverse scene 702 may be crucial to analyze how sound propagates and interacts within the virtual environment.
[0111] The acoustics-matching engine 410 may be connected to both the game engine 512 and the metaverse scene 702. The acoustics-matching engine 410 may receive the metaverse scene 702 and the source audio 116 as inputs. In some cases, the acousticsmatching module 708 may analyze the characteristics of the metaverse scene 702 and apply acoustic transformations to the source audio 116 to generate the acoustics-matched audio 416.
[0112] The circuitry 202 may be configured to determine a virtual distance between the first user 502 and the second user 518 in the metaverse application. This virtual distance may be calculated based on the positions of the users' avatars within the metaverse scene 702. The virtual distance may be a crucial parameter for accurate simulation of how sound travels between users in the virtual space.
[0113] In some cases, the circuitry 202 may be configured to apply an Al-based ambisonics model on the virtual distance based on the source audio 116. The Al-based ambisonics model may simulate how sound waves propagate and interact with the virtual environment over the calculated distance. The Al-based ambisonics model may operate based on factors such as sound attenuation over distance, reflections off virtual surfaces, and potential obstacles in the sound path.
[0114] The circuitry 202 may be further configured to determine a speaker sound field associated with the metaverse application for the first user 502 based on the application of the Al-based ambisonics model. The speaker sound field may represent how the source audio 116 from the first user 502 would propagate through the virtual space. The speaker sound field may be determined based on acoustic properties of the metaverse scene 702, such as the materials of virtual walls or the presence of sound-absorbing objects.
[0115] The acoustics-matching engine 410 may use the determined speaker sound field to generate the final acoustics-matched audio 416. The acoustics-matched audio 416 may be tailored to sound as if the source audio 116 is actually produced within the metaverse environment 404, based on the virtual distance between users and the unique acoustic properties of the virtual space.
[0116] In some cases, the acoustics-matching engine 410 may also include the listener variables 516 associated with the second user 518. The listener variables 516 may include the second user's position, orientation, and even simulated hearing characteristics within the metaverse environment 404. By incorporation of the listener variables 516, the acoustics-matching engine 410 may further refine the audio output to provide a highly personalized and immersive audio experience for the second user 518. It should be noted that the scenario 700 of FIG. 7 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0117] FIG. 8 is a diagram that illustrates an exemplary execution pipeline for generation of acoustics-matched audio, in accordance with an embodiment of the disclosure. FIG. 8 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, and FIG. 7. With reference to FIG. 8, there is shown an exemplary execution pipeline 800. The execution pipeline 800 includes the game engine 512, multi-view images or 3D data 802, an encoder 804, an image feature volume 806, a multi-view feature aggregation network 808, an audio 810, an audio encoder 812, a cross-modal encoder model 814, and an up-sampling module 816.
[0118] The game engine 512 may be configured to generate and manage the virtual environment of the metaverse application. In some cases, the game engine 512 may provide real-time updates of the 3D scene 514, including user positions, object interactions, and environmental changes. The game engine 512 may output the multi-viewimages or 3D data 802, which may represent different perspectives of the 3D scene 514. The multi-view images or 3D data 802 may provide comprehensive visual information about the scene, including spatial layout, object positions, and surface materials.
[0119] The multi-view images or 3D data 802 may be processed through the encoder 804. The encoder 804 may extract relevant features from the visual data or the multi-view images or 3D data 802. For example, the encoder 804 may process the multi-view images or 3D data 802 to extract important visual features, such as edges, textures, and depth information. The encoder 804 may converts the raw image data into a more compact and meaningful representation known as the image feature volume 806. Thus, the encoder 804 may output the image feature volume 806.
[0120] The image feature volume 806 may be a comprehensive representation of the visual features extracted from the multi-view images or 3D data 802. The image feature volume 806 may include spatial information about the 3D scene 514, which may be crucial for accurate acoustics-matching. The image feature volume 806 may be a 3D tensor representation of spatial features extracted from multi-view images 114 of the virtual environment. The image feature volume 806 may correspond to hierarchical feature maps that capture low-level textures, mid-level shapes, and high-level semantic information from the 3D scene and may include voxelized feature grids that encode geometric and appearance information of the 3D space. In an example, the image feature volume 806 may be an attention-based feature aggregations that highlights salient aspects of the virtual scene across multiple views.
[0121] The image feature volume 806 may then be processed by the multi-view feature aggregation network 808. The multi-view feature aggregation network 808 may combine and synthesize information from multiple viewpoints and create a unified representation of the 3D scene 514 based on an integration of spatial information from various angles andperspectives. The output of the multi-view feature aggregation network 808 may provide a rich spatial information of the virtual environment, which may be essential for generation of acoustics-matched audio 416. For example, the multi-view feature aggregation network 808 may employ techniques such as view pooling, attention mechanisms, or graph-based reasoning to effectively fuse features across different viewpoints. Based on aggregation of multi-view information, the multi-view feature aggregation network 808 may enable a comprehensive analysis of the 3D scene's geometry, occlusions, and spatial relationships, which may be crucial for generation of accurate acoustics-matched audio 416.
[0122] In parallel to the visual processing pipeline, the audio 810 may be processed through the audio encoder 812. The audio 810 may correspond to the source audio 116 associated with a user in the metaverse application. The audio encoder 812 may extract relevant features from the audio signal, such as frequency content, amplitude, and temporal characteristics.
[0123] The cross-modal encoder model 814 may combine the processed visual information output from the multi-view feature aggregation network 808 with the encoded audio features output from the audio encoder 812. The cross-modal encoder model 814 may correspond to a visual-acoustic Al model that includes the encoder 804 and a decoder model (not shown in FIG. 8). The cross-modal encoder model 814 may employ techniques such as attention mechanisms, transformer-based models, or neural cross-modal embedding to establish relationships between visual scenes and acoustic properties. The cross-modal encoder model 814 may learn to map between visual and auditory domains, which may enable the generation of acoustics-matched audio 416 that accurately reflects the spatial and material characteristics of the virtual environment.
[0124] The up-sampling module 816 may up-sample the output of the cross-modal encoder model 814 to generate the acoustics-matched audio 416. The up-samplingmodule 816 may increase the resolution or quality of the audio output generated by the cross-modal encoder model 814. The up-sampling module 816 may employ techniques such as transposed convolutions, sub-pixel convolution, or interpolation methods to enhance the detail and fidelity of the acoustics-matched audio 416. The up-sampling module 816 might also incorporate residual connections or attention mechanisms to preserve fine-grained acoustic information while the temporal or frequency resolution of the audio signal may be expanded.
[0125] For example, in case of a virtual concert hall, the game engine 512 may capture the multi-view images 114 of the stage, seating area, and surrounding architecture. The game engine 512 may capture images from different angles, including the stage, audience area, and walls. The multi-view images or 3D data 802 may be encoded to extract features such as shape of the stage, position of the speakers, and materials of the walls. The encoded features may form a comprehensive representation of the concert hall’s visual characteristics. The multi-view feature aggregation network 808 may combine features from all viewpoints to create unified model of the concert hall. The audio encoder 812 may process the initial audio or the source audio (for example, music or speech) to determine audio features, and the cross-modal encoder model 814 may integrate the audio features with the visual features to align the audio with the scene’s acoustics to generate acoustics- matched audio 416. The source audio may refer to an original audio recording that is used as the input for the process of generating acoustics-matched audio. The source audio may be be music, speech, or any other sound that needs to be acoustically matched to a target environment or the virtual concert hall in this example. The circuitry 202 may generate audio that reflects the concert hall’s acoustics, such as the reverberation of sound off the walls and the spatial distribution of sound from the speakers. The audio may be up- sampled to ensure high quality and clarity. The final audio output may be played back tothe listener (for example, the second user 518), which may provide an immersive experience where the sound matches the visual environment of the concert hall.
[0126] FIG. 9 is a diagram that illustrates an exemplary scenario for customization of audio in a virtual environment, in accordance with an embodiment of the disclosure. FIG. 9 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, FIG. 7, and FIG. 8. With reference to FIG. 9, there is shown an exemplary scenario 900. The scenario 900 includes a metaverse user 902, a visual environment 904, an image processor 906, an audio processor 908, an audio modification model 910, the acoustics- matched audio 416, and a customized audio output 912. The scenario 900 may illustrate a system for customization of audio based on visual characteristics of a virtual environment in a metaverse application.
[0127] The metaverse user 902 may interact with the visual environment 904 through the electronic device 102, which may be a virtual reality headset or another suitable device that may be capable of interaction with the metaverse application. The visual environment 904 may represent the 3D scene 514 generated by the metaverse engine 110. The visual environment 904 may encompass various visual elements such as virtual objects, landscapes, and other users' avatars. The visual environment 904 may serve as a foundation for the audio customization process and provide crucial visual cues for the audio customization.
[0128] The image processor 906 may analyze the visual environment 904 and extracting relevant image attributes. The image processor 906 may employ various computer vision techniques to determine characteristics such as color tone, exposure level, brightness, contrast, and saturation of the multi-view images 114. For example, the image processor 906 may use semantic segmentation to identify and classify different elements within the virtual environment. The image processor 906 may use convolutional neural networks(CNNs) or vision transformers to extract relevant visual features and analyze characteristics like color tone, exposure level, brightness, contrast, and saturation of the multi-view images 114. For instance, in a virtual museum scenario, the image processor 906 may analyze lighting conditions of different exhibition halls and identify subtle variations in ambient light that may influence the visitor's perception of artworks. Such detailed visual analysis may enable the circuitry 202 to create audio that complements the visual experience and enhance an overall immersive experience for a user.
[0129] In parallel to the operation of the image processor 906, the audio processor 908 may receive the acoustics-matched audio 416 from the acoustics-matching engine 410. The audio processor 908 may analyze the acoustics-matched audio 416 to determine key attributes such as intensity, amplitude, and frequency levels. The audio processor 908 may employ various digital signal processing techniques to extract the key attributes and may provide a comprehensive analysis of the audio properties. Such an analysis may be crucial for identification of aspects of the audio that may be modified to better align with the visual environment 904. For example, the audio processor 908 may apply Fourier transforms and spectral analysis to extract key audio attributes like intensity, amplitude, and frequency levels from the acoustics-matched audio 416. The audio processor 908 may apply adaptive filtering, reverberation modeling, and spatial audio processing methods to analyze and manipulate the audio characteristics (e.g., the extracted key audio attributes) for customization based on the visual attributes.
[0130] The audio modification model 910 may include an audio customization process, which may integrate inputs from both the image processor 906 and the audio processor 908. Various machine learning techniques may be used to establish correlations between visual attributes and appropriate audio modifications. For example, in a virtual concert scenario, the audio modification model 910 may learn to associate vibrant, high-contrastvisuals with more energetic and dynamic audio characteristics. Such an adaptive approach may allow creation of audio experiences that feel natural and intuitive within the context of the virtual environment. For example, the audio modification model 910 may employ dynamic equalization techniques to adjust frequency content based on visual characteristics, such as boosting bass frequencies for darker scenes or enhancing treble for brighter environments. Adaptive reverberation and spatial audio processing methods may be used to modify the perceived size and reflectivity of the virtual space, which may align the audio's spatial characteristics with the visual cues from the 3D environment.
[0131] The customized audio output 912 may represent an audio that has been tailored to complement various visual elements of the metaverse environment 404. Such an output may incorporate various modifications, such as adjustments to reverb levels to match the perceived size of virtual spaces, or changes in tonal balance to align with the color palette of the scene. The result may be a more cohesive sensory experience that enhances the user's sense of presence within the virtual world. For example, the customized audio output 912 may include dynamically adjusted reverberation, equalization, and spatial cues that closely match the visual characteristics of the virtual environment. Such an output may include elements such as position-dependent sound attenuation, frequency-specific reflections based on virtual surface materials, and ambient sound textures that correspond to the visual mood and lighting of the 3D scene.
[0132] In an embodiment, the electronic device 102 may receive a user selection of a specific visual environment (e.g., the visual environment 904) from a list of available options. The visual environment 904 may be, for example, a bustling cityscape, a serene forest, or a futuristic space station. The selected environment may be represented by one or more images that may capture its visual characteristics. The image processor 906 may analyze the images to extract key visual features such as spatial layout, object positions,textures, and materials. This selected environment may convert the image data into a structured representation that may be used for further processing. The audio processor 908 may take an audio input 406, which may be a generic sound clip, music, or any other audio relevant to the selected environment. The audio processor 908 may analyze the initial audio to extract important features such as frequency, amplitude, and temporal patterns. The audio modification model 910 may combine the visual features extracted by the image processor 906 with the audio features extracted by the audio processor 908. This integration may ensure that the audio is contextually aligned with the visual environment 904.
[0133] In an example, the user may select “medieval castle” environment from a virtual reality application. The circuitry 202 may be configured to use images of the medieval castle, including the grand hall, stone walls, and surrounding courtyards. The image processor 906 may extract features such as the layout of the castle, the texture of the stone walls, and the spatial arrangement of rooms and corridors. The audio input 406 may be ambient sounds like footsteps, echoes, and distant conversations. The audio processor 908 may extract features such as frequency and amplitude of the ambient sounds. The audio modification model 910 may combine the visual features of the castle with the audio features of the ambient sounds. The audio modification model 910 may simulate the acoustics of the medieval castle, including the reverberation of sound in the grand hall and the muffled echoes in the stone corridors. The audio modification model 910 may modify the initial audio to match the castle’s acoustics, add reverb, adjust echo, and enhance the frequency response to reflect the stone walls and open spaces. The modified audio features may be synthesized to create the final customized audio, which may accurately reflect the medieval castle’s acoustics. The audio may be up sampled and refined to ensure high quality and clarity.
[0134] The disclosed technique may have an ability to dynamically adapt audio based on visual cues may be useful in various metaverse application scenarios. For instance, in a virtual educational platform, the audio may automatically adjust to support different learning environments, such as creation a focused atmosphere for a virtual lecture hall or a more relaxed ambiance for a study group in a virtual library. This context-aware audio customization may help maintain user engagement and improve information retention by creating a more immersive and appropriate auditory backdrop for different activities.
[0135] The disclosed technique may allow real-time adjustments as users navigate through different areas of the metaverse environment 404. As a user transitions from one virtual space to another, the audio may also smoothly adapt to reflect these changes, which may maintain a seamless and believable experience. This capability may be particularly valuable in large-scale metaverse environments where users may encounter a wide variety of virtual spaces and scenarios. It should be noted that the scenario 900 of FIG. 9 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0136] FIG. 10 is a diagram that illustrates an exemplary scenario for processing of audio in a virtual environment, in accordance with an embodiment of the disclosure. FIG. 10 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, FIG. 7, FIG. 8, and FIG. 9. With reference to FIG. 10, there is shown an exemplary scenario 1000. The scenario 1000 includes the audio device 118, the source audio 116, the head orientation 504, the Al-based ambisonics conversion module 506, and a (virtual speaker) sound field module 1002. The scenario 1000 illustrates a system for processing audio in a virtual environment, which may be implemented by the electronic device 102 and executed by the circuitry 202. This system may be configured to create a realistic andimmersive audio experience for users in a metaverse application based on both the characteristics of the audio source and the spatial properties of the virtual environment.
[0137] The audio device 118 in Figure 10 may represent a user in the virtual audio environment. The user may serve as an origin point for sound within the metaverse application. The user may wear the electronic device 102, for example, the VR device including an avatar. The avatar within the virtual environment may be controlled by the user. The virtual environment may also include a non-player character, or even an environmental sound source. The speaker's position and orientation within the virtual space may be paramount in determination of how audio may be processed and perceived by other users. For instance, in a virtual concert scenario, the audio device 118 or the user may represent a performer on stage, and their position and orientation may directly influence how the audience perceives the music.
[0138] The source audio 116 may be associated with the audio device 118. The source audio 116 may encompasses the original audio signal produced within the virtual environment. The source audio 116 may vary widely in nature, from speech in a virtual meeting to ambient sounds in a simulated natural environment. The ability of the electronic device 102 to handle diverse audio types enhances its versatility across different metaverse applications. For example, in a virtual language learning environment, the source audio 116 may include pronunciations of words or phrases, which may require high fidelity to ensure accurate learning outcomes.
[0139] The head orientation 504 may provide critical spatial information about the speaker's facing direction. This data may be instrumental in creation of a realistic audio experience that responds dynamically to user movements. The electronic device 102 may utilize sensors, such as sensors of VR headsets, to accurately track and translate real- world head movements into the virtual space. This feature may become particularlyvaluable in scenarios like virtual sports simulations, where the direction of a player's voice may provide important cues to teammates.
[0140] The ambisonics conversion module 506 may transform the source audio 116 into a spatially aware format. The transformation may be based on the head orientation 504 and may ensure that the audio is correctly positioned relative to the speaker's location and direction. The use of ambisonics technology may allow for a full-sphere surround sound environment and may be capable of representation of sound sources above and below the listener as well as around them to provide a 3D audio experience. The comprehensive spatial audio approach significantly enhances the immersion factor in virtual environments. The audio experience may be of any dimension and not limited to only a 3D experience.
[0141] The ambisonics conversion process may employ techniques (such as, deep learning models) that simulate complex acoustic phenomena. Such algorithms may operate based on environmental factors within the virtual space, such as room size, material properties of surfaces, and the presence of obstacles. For instance, in a virtual architectural design application, the ambisonics conversion module 506 may simulate how sound would propagate in a proposed building design, which may allow architects and clients to experience the acoustic properties of spaces before they are built.
[0142] The (virtual speaker) sound field module 1002 may simulate a propagation of sound from a speaker's (i.e., the audio device 118) position throughout the virtual environment, based on factors such as distance attenuation, reflections, and occlusions. The (virtual speaker) sound field module 1002 may generate a 3D representation of how the source audio interacts with the virtual space and enable realistic spatial audio rendering from different listener positions within the metaverse environment. This feature may particularly be effective in large-scale virtual environments, such as open-worldgames or virtual cities, where users may experience realistic changes in sound as they navigate through different areas.
[0143] The ability of the electronic device 102 to generate a realistic virtual speaker sound field may have significant implications for various metaverse applications. In educational settings, the electronic device 102 may create more engaging and interactive learning experiences based on accurate simulation of acoustic environments relevant to the subject matter. For virtual conferences or collaborative workspaces, the electronic device 102 may provide a more natural communication environment, which may reduce fatigue associated with long-duration virtual meetings. It should be noted that the scenario 1000 of FIG. 10 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0144] FIG. 11 is a diagram that illustrates an exemplary scenario of a spatial audio processing system, in accordance with an embodiment of the disclosure. FIG. 11 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, FIG. 7, FIG. 8, FIG. 9, and FIG. 10. With reference to FIG. 11 , there is shown an exemplary scenario 1100. The scenario 1100 includes an audio 810, a first audio encoder 1102, stacked decoder blocks 1112, an orientation sensor 1104, a positional encoder 1106 (based on orientation), a microphone position 1108, a second encoder 1110, and a spatial audio output 1114. The scenario 1100 may illustrate a comprehensive spatial audio processing system configured to create immersive and realistic audio experiences within metaverse applications. This system may be implemented as part of the electronic device 102 and may be executed by the circuitry 202, by leveraging the capabilities of the metaverse engine 110 to enhance the overall user experience in virtual environments.
[0145] The audio 810 may serve as the primary input to the system (such as electronic device 102 or the VR device) and may represent the raw audio signal that may be capturedfrom various sources within the metaverse environment 404. In some cases, the audio 810 may correspond to the source audio 116 associated with a user or an object in the virtual space. For example, in a virtual concert scenario, the audio 810 may represent the live performance of a musician, which may capture the nuances and dynamics of their instrument or voice.
[0146] The first audio encoder 1102 may process and encode the raw audio signal (i.e. , the audio 810). The first audio encoder 1102 may employ advanced signal processing techniques to analyze and extract key features from the audio, such as frequency content, amplitude variations, and temporal characteristics. The first audio encoder 1102 may prepare the audio data (i.e., the audio 810) for further processing based on a conversion of the audio data into a format that may be efficiently manipulated by the stacked decoder blocks 1112.
[0147] The orientation sensor 1104 may provide real-time data about the user's head orientation 504 within the virtual space. The orientation sensor 1104 may be part of a VR headset or other input device connected to the electronic device 102. The orientation sensor 1104 may capture subtle movements and rotations of the user's head, which may allow the electronic device 102 to adjust the audio output accordingly. For instance, in a virtual meeting room, the orientation sensor 1104 may enable the electronic device 102 to accurately represent the directionality of different speakers' voices as the user turns their head.
[0148] The positional encoder 1106 may be connected to the orientation sensor 1104 and may be responsible for translation of the raw orientation data into a format that may be integrated with the audio processing pipeline. The positional encoder 1106 may convert the spatial information into a representation that may be easily combined with the audio features. The positional encoder 1106 may enable the electronic device 102 to create acohesive spatial audio experience that responds dynamically to the user's movements and orientation within the virtual environment.
[0149] The microphone position 1108 may correspond to information about a location of a virtual microphone or listening point within the metaverse environment 404. The microphone position 1108 may be particularly useful in scenarios where the user's avatar or listening position may be different from their physical orientation. For example, in a virtual cinema application, the microphone position 1108 may represent the optimal listening point within the virtual theater, regardless of how the user may be physically oriented in their real-world space.
[0150] The second encoder 1110 may be connected to both the positional encoder 1106 and the microphone position 1108. The second encoder 1110 may combine the spatial information from both the positional encoder 1106 and the microphone position 1108 with the processed audio data (from the first audio encoder 1102), at the stacked decoder blocks 1112. The second encoder 1110 may employ techniques (such as, deep learning models) to integrate these diverse inputs and create a unified representation that captures both the audio content and its spatial characteristics within the virtual environment.
[0151] The stacked decoder blocks 1112 may be connected to the first audio encoder 1102 and the second encoder 1110 and may play a crucial role in the spatial audio processing pipeline. The stacked decoder blocks 1112 may include multiple layers of neural networks or other advanced techniques that may be configured to decode and transform the encoded audio information. The stacked decoder blocks 1112 may employ techniques such as transposed convolutions, attention mechanisms, or recurrent layers to reconstruct and enhance the spatial audio characteristics based on the combined inputs from the first audio encoder 1102 and orientation and positional information. In some cases, the stacked decoder blocks 1112 may progressively refine the audio representationand incorporate spatial information and acoustic properties of the virtual environment to create a more immersive sound experience.
[0152] The spatial audio output 1114 may represent an output of the stacked decoder blocks 1112 and may deliver an immersive and spatially accurate audio experience to the user. The spatial audio output 1114 may be tailored to various playback systems, such as headphones, multi-speaker setups, or specialized spatial audio devices. The spatial audio output 1114 may provide users with a sense of presence and directionality that closely mimics real-world acoustic experiences, enhancing the overall realism of the metaverse application.
[0153] In operation, the electronic device 102 may process the audio 810 through the first audio encoder 1102 and the second encoder 1110 while spatial information from the orientation sensor 1104 and microphone position 1108 may be simultaneously incorporated. The stacked decoder blocks 1112 may then integrate these processed inputs to generate the spatial audio output 1114. This output may be dynamically adjusted in real-time based on changes in the user's orientation or position within the virtual environment, ensuring a consistently immersive audio experience.
[0154] The spatial audio processing system illustrated in scenario 1100 may significantly enhance the auditory aspects of metaverse applications. Based on a combination of advanced audio processing techniques with real-time spatial information, the electronic device 102 may create highly realistic and responsive sound environments. This may lead to improved user engagement and a stronger sense of presence within virtual spaces and may benefit applications ranging from virtual conferences and educational platforms to immersive gaming and entertainment experiences. It should be noted that the scenario 1100 of FIG. 11 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0155] FIG. 12 is a diagram that illustrates an exemplary scenario for generation of listener perspective audio in a virtual environment, in accordance with an embodiment of the disclosure. FIG. 12 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, FIG. 7, FIG. 8, FIG. 9, FIG. 10, and FIG. 11. With reference to FIG. 12, there is shown an exemplary scenario 1200. The scenario 1200 includes the audio device 118, the game engine 512, the listener variables 516, a sound field 1202, and a listener perspective audio 1204. The scenario 1200 may illustrate a system for generation of listener perspective audio 1204 in a virtual environment, which may be implemented within the electronic device 102 and executed by the circuitry 202. This system may be configured to create a personalized and immersive audio experience for users in a metaverse application by considering the listener's position and orientation within the virtual space.
[0156] The game engine 512 may serve as the core component of the virtual environment and may manage the overall state of the metaverse application. In some cases, the game engine 512 may render the visual elements of the virtual world, handle user interactions, and coordinate the various subsystems that contribute to the immersive experience. For example, in a virtual concert scenario, the game engine 512 may manage the placement of virtual instruments on stage, the movement of audience avatars, and the dynamic lighting effects that respond to the music.
[0157] The game engine 512 may be connected to the listener variables 516, which may represent a set of parameters that define the listener's position and orientation within the virtual environment. The listener variables 516 may include data such as the listener's coordinates in 3D space, the head orientation 504, and even simulated characteristics of the listener's virtual ears or hearing apparatus. In some cases, the listener variables 516 may be updated in real-time as the user moves and interacts within the virtual space. Forinstance, in a virtual museum application, the listener variables 516 may continuously update as the user walks through different exhibition halls, turns to examine artifacts, or leans in to hear audio descriptions of artworks.
[0158] The sound field 1202 may represent a comprehensive model of how audio propagates within the virtual environment. The sound field 1202 may be based on the acoustic properties of the virtual space, such as room size, material reflectivity, and the presence of obstacles that may affect sound propagation. The sound field 1202 may be generated based on inputs from the game engine 512 and the audio device 118 and may be continuously updated to reflect changes in the virtual environment. For example, in a virtual conference room scenario, the sound field 1202 may adapt in real-time if a user opens a virtual window, changes the room's furniture arrangement, or if more participants join the meeting, alters the acoustic dynamics of the space.
[0159] The listener perspective audio 1204 may be a final output of the electronic device 102, which may represent the personalized audio experience tailored to the listener's position and orientation within the virtual environment. The listener perspective audio 1204 may be determined based on information from the sound field 1202 with the listener variables 516. The listener perspective audio 1204 may be an audio that accurately reflects what the user would hear if they were physically present in the simulated space. In some cases, the listener perspective audio 1204 may incorporate advanced audio processing techniques such as binaural rendering or head-related transfer functions (HRTFs) to create a highly realistic spatial audio experience.
[0160] In operation, the electronic device 102 may receive update information from the game engine 512 about the current state of the virtual environment. The received update information may be used to update the sound field 1202 and ensure that sound field 1202 accurately represents the acoustic properties of the virtual space. Simultaneously, theelectronic device 102 may receive real-time updates of the listener variables 516, which may be captured through sensors in the user's VR headset or other input devices.
[0161] The sound field 1202 and listener variables 516 may then be processed together to generate the listener perspective audio 1204. This processing may involve calculations that simulate how sound waves would interact with the virtual environment and reach the listener's ears based on their position and orientation. For example, in a virtual sports stadium, the electronic device 102 may calculate how the roar of the crowd, the announcer's voice, and the sounds of the game would be perceived differently as the user moves from the stands to the sidelines or turns their head to focus on different areas of the field.
[0162] The listener perspective audio 1204 may be continuously updated in real-time, allowing for a dynamic and responsive audio experience that closely mirrors real-world hearing. This real-time processing may be particularly valuable in interactive scenarios, such as virtual multiplayer games, where the audio environment needs to adapt quickly to rapidly changing game states and player movements.
[0163] In some cases, the electronic device 102 may also incorporate additional features to enhance the realism of the audio experience. For instance, the listener perspective audio 1204 may include simulations of audio occlusion and obstruction, where sounds are muffled or blocked by virtual objects in the environment. The electronic device 102 may also use Doppler effect information for moving sound sources, to create realistic pitch shifts for objects that are approaching or receding from the listener.
[0164] The implementation of this listener perspective audio 1204 system may significantly enhance the immersion and realism of metaverse applications. Based on the provision of audio that accurately responds to the user's position and movements within the virtual space, the electronic device 102 may create a more believable and engagingexperience across a wide range of applications, from virtual social gatherings and educational simulations to complex multiplayer games and professional collaboration tools. It should be noted that the scenario 1200 of FIG. 12 is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0165] FIG. 13 is a flowchart of an example method for generation of acoustics-matched audio for three-dimensional (3D) scenes in metaverse applications, in accordance with an embodiment of the disclosure. FIG. 13 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, FIG. 7, FIG. 8, FIG. 9, FIG. 10, FIG. 11 , and FIG. 12. With reference to FIG. 13, there is shown an exemplary flowchart 1300. The flowchart 1300 includes steps 1302 to 1316 for generation of acoustics-matched audio for three- dimensional (3D) scenes in metaverse applications. The various operations of the flowchart 1300 may be performed by any computing device, such as, the electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2. The flowchart 1300 starts from 1302 and proceeds to 1304.
[0166] At 1304, multi-view images associated with a metaverse application may be received from a metaverse engine, where the multi-view images correspond to a 3D scene associated with a first user of metaverse application. The circuitry 202 may be configured to receive the multi-view images 114 from the metaverse engine 110 associated with metaverse application. In an embodiment, the multi-view images 114 may correspond to the 3D scene associated with the first user 502 of the metaverse application. For example, in a virtual museum scenario, the multi-view images 114 may represent different perspectives of an exhibition hall, including views from various angles and positions within the virtual space. The reception of multi-view images is described further, for example, inFIG. 3.
[0167] At 1306, a deep learning model 112A may be applied on the multi-view images. The circuitry 202 may be configured to apply the deep learning model 112A on the multiview images 114. The application of the deep learning model 112A may involve processing the received multi-view images 114 to extract relevant features and spatial information. For instance, in a virtual conference room application, the deep learning model 112A may analyze the layout, lighting conditions, and object placements within the virtual space. The application of the deep learning model 112A is described further, for example, in FIG. 3 and FIG. 8.
[0168] At 1308, volume aggregation information associated with the 3D scene may be determined based on the application of the deep learning model 112A. The circuitry 202 may be configured to determine the volume aggregation information. The determination of the volume aggregation information may involve synthesis of the spatial characteristics of the 3D scene 514 derived from the multi-view images 114. For example, in a virtual concert venue, the volume aggregation information may include data about the stage dimensions, seating arrangements, and acoustic properties of the virtual space. The determination of volume aggregation information is described further, for example, in FIG. 3 and FIG. 8.
[0169] At 1310, a source audio associated with the first user may be received. The circuitry 202 may be configured to receive the source audio 116 associated with the first user 502. The circuitry 202 may acquire the original audio content (e.g., the source audio 116) that will be modified to match the acoustic properties of the 3D scene 514. For instance, in a virtual language learning environment, the source audio 116 may be a recording of a native speaker's pronunciation. The reception of source audio is described further, for example, in FIG. 3 and FIG. 4.
[0170] At 1312, a cross-modal encoder model may be applied on the volume aggregation information and the source audio. The circuitry 202 may be configured toapply the cross-modal encoder model 814 on the volume aggregation information and the source audio 116. The application of the cross-modal encoder model 814 may involve integration of the spatial information derived from the visual data with the audio content. For example, in a virtual gaming environment, the cross-modal encoder model 814 may combine the layout of a virtual battlefield with the sound of in-game actions to create a cohesive audio-visual experience. The application of the cross-modal encoder model is described further, for example, in FIG. 3 and FIG. 8.
[0171] At 1314, an acoustics-matched audio associated with the 3D scene may be generated based on the application of the cross-modal encoder model. The circuitry 202 may be configured to generate the acoustics-matched audio 416 associated with the 3D scene 514. The circuitry 202 may produce a final audio output that has been modified to match the acoustic properties of the 3D scene 514. For instance, in a virtual architectural design application, the generated acoustics-matched audio 416 may simulate how conversations would sound in different rooms of a proposed building design. The generation of acoustics-matched audio is described further, for example, in FIG. 3, FIG. 4, and FIG. 8.
[0172] At 1316, the display device 206A may be controlled to render the 3D scene for a second user of the metaverse application based on the acoustics-matched audio. The circuitry 202 may be configured to control the display device 206A to render the 3D scene 514 for the second user based on the acoustics-matched audio 416. The control of display device 206A to render the 3D scene may involve coordination of the visual and auditory elements to create a cohesive and immersive experience for the user. For example, in a virtual tourism application, the control of display device 206A to render the virtual waterfall scene may ensure that the sound of a virtual waterfall changes appropriately as the user moves closer or farther from the sound of a virtual waterfall within the virtual environment.The control of the display device 206A to render 3D scene 514 is described further, for example, in FIG. 3 and FIG. 12. Control may pass to end.
[0173] Although the exemplary method is illustrated as discrete operations, such as 1302, 1304, 1306, 1308, 1310, 1312, 1314, and 1316, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.
[0174] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions by a machine and / or a computer to operate an electronic device (for example, the electronic device 102 of FIG. 1 ). Such instructions may cause the electronic device 102 to perform operations that may include reception, from a metaverse engine (e.g., the metaverse engine 110), of multi-view images (e.g., the multi-view images 114) associated with a metaverse application. The multi-view images 114 correspond to a three- dimensional (3D) scene associated with a first user of the metaverse application. The operations may further include application of a deep learning model (e.g., the deep learning model 112A) on the multi-view images 114. The operations may further include determination of volume aggregation information associated with the 3D scene, based on the application of the deep learning model 112A. The operations may further include receipt of a source audio (e.g., the source audio 116) associated with the first user and application of a cross-modal encoder model on the volume aggregation information and the source audio 116. The operations may further include generation of an acoustics- matched audio associated with the 3D scene, based on the application of the cross-modal encoder model. The operations may further include control the display device 206A torender the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio.
[0175] Exemplary aspects of the disclosure may provide an electronic device (such as, the electronic device 102 of FIG. 1 ) that includes circuitry (such as, the circuitry 202). The circuitry 202 may be configured to receive, from a metaverse engine (e.g., the metaverse engine 110), multi-view images (e.g., the multi-view images 114) associated with a metaverse application. The multi-view images 114 correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application. The circuitry 202 may be configured to apply a deep learning model (e.g., the deep learning model 112A) on the multi-view images 114. The circuitry 202 may be configured to determine volume aggregation information associated with the 3D scene, based on the application of the deep learning model 112A. The circuitry 202 may be configured to receive a source audio (e.g., the source audio 116) associated with the first user and application of a cross-modal encoder model on the volume aggregation information and the source audio 116. The circuitry 202 may be configured to generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model. The circuitry 202 may be configured to control the display device 206A to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio.
[0176] In an embodiment, the circuitry 202 may be configured to receive a user input associated with the 3D scene, apply a generative artificial intelligence (Al) model on the user input, and generate the 3D scene based on the application of the generative Al model. The user input may correspond to at least one of an image associated with the 3D scene, keywords associated with the 3D scene, or a first prompt associated with the 3D scene. The circuitry 202 may be further configured to generate a second prompt associated withthe 3D scene, based on the image and the keywords. The generation of the 3D scene may be further based on at least one of the first prompt or the second prompt.
[0177] In an embodiment, the circuitry 202 may be configured to receive head orientation information associated with the first user and apply an Al-based ambisonics model on the head orientation information, based on the acoustics-matched audio associated with the 3D scene. The circuitry 202 may be configured to determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model. The circuitry 202 receive, from the metaverse engine 110, listener variables associated with the metaverse application for the second user and determine a listener-based audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables. The control of display device 206A to render the 3D scene for the second user may be further based on the listener-based audio.
[0178] In an embodiment, the circuitry 202 may be configured to determine an audio output based on the application of the cross-modal encoder model, and up-sample the audio output. The generation of the acoustics-matched audio may be further based on the up-sampled audio output. The circuitry 202 may be further configured to determine image attributes associated with the multi-view images 114 and determine audio attributes associated with the acoustics-matched audio. The circuitry 202 may update the audio attributes of the acoustics-matched audio based on the image attributes associated with the multi-view images 114 and generate a remixed audio associated with the 3D scene, based on the updated audio attributes. The control the display device 206A to render the 3D scene may be further based on the remixed audio.
[0179] In an embodiment, the circuitry 202 may be configured to receive head orientation information associated with the first user and apply an Al-based ambisonicsmodel on the head orientation information, based on the source audio 116 associated with the first user. The circuitry 202 may determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model, and control the display device 206A to render the 3D scene for the first user of the metaverse application, based on the speaker sound field.
[0180] In an embodiment, the circuitry 202 may be configured to determine a virtual distance between the first user and the second user in the metaverse application and apply an Al-based ambisonics model on the virtual distance, based on the source audio 116. The circuitry 202 may determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model, and receive, from the metaverse engine 110, listener variables associated with the metaverse application for the second user. The circuitry 202 may determine a listenerbased audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables. The control of display device 206A to render the 3D scene for the second user may be further based on the listener-based audio.
[0181] In an embodiment, the cross-modal encoder model may correspond to a visualacoustic Al model including a first encoder model, a second encoder model, and a set of stacked decoder models. The circuitry 202 may be configured to apply the first encoder model on the source audio 116, determine a first encoded input based on the application of the first encoder model, and apply the second encoder model on a head orientation of the first user and a microphone position of the first user. The circuitry 202 may determine a second encoded input based on the application of the second encoder model and apply the set of stacked decoder models on the first encoded input and the second encoded input to determine an output spatial audio. The output spatial audio corresponds to the acoustics-matched audio.
[0182] In an embodiment, the metaverse application may correspond to at least one of: a virtual musical concert, a virtual meeting, a virtual conference room, an electronic learning platform, a metaverse gaming application, a metaverse tourism application, or a metaverse shopping application.
[0183] The present disclosure may also be positioned in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.
[0184] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:1 . An electronic device, comprising: circuitry configured to: receive, from a metaverse engine, multi-view images associated with a metaverse application, wherein the multi-view images correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application; apply a deep learning model on the multi-view images; determine volume aggregation information associated with the 3D scene, based on the application of the deep learning model; receive a source audio associated with the first user; apply a cross-modal encoder model on the volume aggregation information and the source audio; generate an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model; and control a display device to render the 3D scene to a second user of the metaverse application, based on the acoustics-matched audio.
2. The electronic device according to claim 1 , wherein the circuitry is further configured to: receive a user input associated with the 3D scene; apply a generative artificial intelligence (Al) model on the user input; and generate the 3D scene based on the application of the generative Al model.
3. The electronic device according to claim 2, wherein the user input corresponds to at least one of: an image associated with the 3D scene, keywords associated with the 3D scene, or a first prompt associated with the 3D scene.
4. The electronic device according to claim 3, wherein the circuitry is further configured to: generate a second prompt associated with the 3D scene, based on the image and the keywords, wherein the generation of the 3D scene is further based on at least one of the first prompt or the second prompt.
5. The electronic device according to claim 1 , wherein the circuitry is further configured to: receive head orientation information associated with the first user; apply an Al-based ambisonics model on the head orientation information, based on the acoustics-matched audio associated with the 3D scene; determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; receive, from the metaverse engine, listener variables associated with the metaverse application for the second user; and determine a listener-based audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables, whereinthe control of the display device to render the 3D scene for the second user is further based on the listener-based audio.
6. The electronic device according to claim 1 , wherein the circuitry is further configured to: determine an audio output based on the application of the cross-modal encoder model; and up-sample the audio output, wherein the generation of the acoustics-matched audio is further based on the up- sampled audio output.
7. The electronic device according to claim 1 , wherein the circuitry is further configured to: determine image attributes associated with the multi-view images; determine audio attributes associated with the acoustics-matched audio; update the audio attributes of the acoustics-matched audio based on the image attributes associated with the multi-view images; and generate a remixed audio associated with the 3D scene, based on the updated audio attributes, wherein the control of the display device to render of the 3D scene is further based on the remixed audio.
8. The electronic device according to claim 7, wherein the image attributes include at least one of: a color tone of an image of the multi-view images, an exposure level of the image,a brightness level of the image, a contrast level of the image, or a saturation level of the image.
9. The electronic device according to claim 7, wherein the image attributes include at least one of: an intensity level of the acoustics-matched audio, an amplitude level of the acoustics-matched audio, or a frequency level of the acoustics-matched audio.
10. The electronic device according to claim 1 , wherein the circuitry is further configured to: receive head orientation information associated with the first user; apply an Al-based ambisonics model on the head orientation information, based on the source audio associated with the first user; determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; and control the display device to render the 3D scene for the first user of the metaverse application, based on the speaker sound field.
11. The electronic device according to claim 1 , wherein the cross-modal encoder model corresponds to a visual-acoustic Al model including a first encoder model, a second encoder model, and a set of stacked decoder models.
12. The electronic device according to claim 11 , wherein the circuitry is further configured to:apply the first encoder model on the source audio; determine a first encoded input based on the application of the first encoder model; apply the second encoder model on a head orientation of the first user and a microphone position of the first user; determine a second encoded input based on the application of the second encoder model; and apply the set of stacked decoder models on the first encoded input and the second encoded input to determine an output spatial audio, wherein the output spatial audio corresponds to the acoustics-matched audio.
13. The electronic device according to claim 1 , wherein the circuitry is further configured to: determine a virtual distance between the first user and the second user in the metaverse application; apply an Al-based ambisonics model on the virtual distance, based on the source audio; determine a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; receive, from the metaverse engine, listener variables associated with the metaverse application for the second user; and determine a listener-based audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables, wherein the control of display device to render the 3D scene to the second user is further based on the listener-based audio.
14. The electronic device according to claim 1 , wherein the metaverse application corresponds to at least one of: a virtual musical concert, a virtual meeting, a virtual conference room, an electronic learning platform, a metaverse gaming application, a metaverse tourism application, or a metaverse shopping application.
15. A method, comprising: in an electronic device: receiving, from a metaverse engine, multi-view images associated with a metaverse application, wherein the multi-view images correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application; applying a deep learning model on the multi-view images; determining volume aggregation information associated with the 3D scene, based on the application of the deep learning model; receiving a source audio associated with the first user; applying a cross-modal encoder model on the volume aggregation information and the source audio; generating an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model; and controlling a display device to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio.
16. The method according to claim 15, further comprising: receiving head orientation information associated with the first user; applying an Al-based ambisonics model on the head orientation information, based on the acoustics-matched audio associated with the 3D scene; determining a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; receiving, from the metaverse engine, listener variables associated with the metaverse application for the second user; and determining a listener-based audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables, wherein the control of display device to render of the 3D scene for the second user is further based on the listener-based audio.
17. The method according to claim 15, further comprising: determining image attributes associated with the multi-view images; determining audio attributes associated with the acoustics-matched audio; updating the audio attributes of the acoustics-matched audio based on the image attributes associated with the multi-view images; and generating a remixed audio associated with the 3D scene, based on the updated audio attributes, wherein the control of the display device to render of the 3D scene is further based on the remixed audio.
18. The method according to claim 15, further comprising:receiving head orientation information associated with the first user; applying an Al-based ambisonics model on the head orientation information, based on the source audio associated with the first user; determining a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; and controlling the display device to render the 3D scene for the first user of the metaverse application, based on the speaker sound field.
19. The method according to claim 15, further comprising: determining a virtual distance between the first user and the second user in the metaverse application; applying an Al-based ambisonics model on the virtual distance, based on the source audio; determining a speaker sound field associated with the metaverse application for the first user, based on the application of the Al-based ambisonics model; receiving, from the metaverse engine, listener variables associated with the metaverse application for the second user; and determining a listener-based audio associated with the metaverse application for the second user, based on the speaker sound field and the listener variables, wherein the control of the display device to render the 3D scene for the second user is further based on the listener-based audio.
20. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising:receiving, from a metaverse engine, multi-view images associated with a metaverse application, wherein the multi-view images correspond to a three-dimensional (3D) scene associated with a first user of the metaverse application; applying a deep learning model on the multi-view images; determining volume aggregation information associated with the 3D scene, based on the application of the deep learning model; receiving a source audio associated with the first user; applying a cross-modal encoder model on the volume aggregation information and the source audio; generating an acoustics-matched audio associated with the 3D scene, based on the application of the cross-modal encoder model; and controlling a display device to render the 3D scene for a second user of the metaverse application, based on the acoustics-matched audio.
Citation Information
Patent Citations
IN202411051082A