Sound Source Association in Augmented Reality via 3D Model Fragments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current augmented reality systems face challenges in accurately rendering computer-generated 3D items in real-time within real-world scenes, particularly due to limitations in computing resources and the difficulty in consistently translating real-world coordinates to virtual coordinates, which hinders immersive experiences and real-time sound source localization.
Innovation Solution
A sound location system utilizing an array of microphones to capture 3D sound data, a 3D polyhedron model of the scene, and a sound source separation module to identify and localize sound sources by comparing theoretical positions with 3D model fragments, enabling precise positioning of sound sources within the virtual environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If computer generated 3D items are rendered in real-time within real-life scenes, then immersion and interactivity are improved, but computing resource requirements increase and translation accuracy between coordinate systems deteriorates
Solution Approach 1:
The patent introduces 3D model fragments as intermediary objects that serve as reference markers between the real-world coordinate system and the virtual coordinate system. These fragments are distributed throughout the scene and provide stable, pre-computed reference points that facilitate accurate translation between coordinate systems without requiring complex real-time computation.
Solution Approach 2:
The patent performs preliminary actions by pre-computing and distributing 3D model fragments throughout the real-world scene before the actual rendering occurs. These fragments are prepared in advance with known positions and orientations, allowing the system to quickly reference them during real-time rendering without performing complex computations at the moment of rendering.
2Measurement precision
If sound sources are localized in 3D space, then spatial audio immersion is improved, but measurement precision and computational complexity increase
Solution Approach 1:
The patent uses 3D model fragments as intermediary reference objects for sound source localization. Instead of performing complex acoustic field analysis throughout the entire 3D space, the system localizes sound sources relative to these pre-positioned fragments, which provide known spatial references that simplify the computation while maintaining precision.
Solution Approach 2:
The patent segments the 3D space into regions associated with individual 3D model fragments. Each fragment serves as a localized reference frame, allowing the system to perform sound source localization independently within each fragment's region rather than performing a single complex global analysis, thereby reducing overall computational complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables accurate and immersive real-time rendering of 3D items and sound sources within augmented reality, enhancing user interaction and experience by overcoming resource limitations and translation challenges.
Implementation Method 1
an array of microphones that capture sound of a three-dimensional (3D), real-world scene
Data Source
AI summary
Multiple Holocam Orbs observe a real-life environment and generate an artificial reality representation of the real-life environment. Depth image data is cleansed of error due to LED shadow by identifying the edge of a foreground object in an (near infrared light) intensity image, identifying an edge in a depth image, and taking the difference between the start of both edges. Depth data error due to parallax is identified noting when associated text data in a given pixel row that is progressing in a given row direction (left-to-right or right-to-left) reverses order. Sound sources are identified by comparing results of a blind audio source localization algorithm, with the spatial 3D model provided by the Holocam Orb. Sound sources that corresponding to identifying 3D objects are associated together. Additionally, types of data supported by a standard movie data container, such as an MPEG container, is expanding to incorporate free viewpoint data (FVD) model data. This is done by inserting FVD data of different individual 3D objects at different sample rates into a single video stream. Each 3D object is separately identified by a separately assigned ID.


