Sound Source Association in Augmented Reality via 3D Model Fragments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality systems face challenges in accurately rendering computer-generated 3D items in real-time within real-world scenes, particularly due to limitations in computing resources and the difficulty in consistently translating real-world coordinates to virtual coordinates, which hinders immersive experiences and real-time sound source localization.

Innovation Solution

A sound location system utilizing an array of microphones to capture 3D sound data, a 3D polyhedron model of the scene, and a sound source separation module to identify and localize sound sources by comparing theoretical positions with 3D model fragments, enabling precise positioning of sound sources within the virtual environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If computer generated 3D items are rendered in real-time within real-life scenes, then immersion and interactivity are improved, but computing resource requirements increase and translation accuracy between coordinate systems deteriorates

Engineering Contradiction:
Improvetranslation accuracyVSAvoidcomputing resources
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces 3D model fragments as intermediary objects that serve as reference markers between the real-world coordinate system and the virtual coordinate system. These fragments are distributed throughout the scene and provide stable, pre-computed reference points that facilitate accurate translation between coordinate systems without requiring complex real-time computation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary actions by pre-computing and distributing 3D model fragments throughout the real-world scene before the actual rendering occurs. These fragments are prepared in advance with known positions and orientations, allowing the system to quickly reference them during real-time rendering without performing complex computations at the moment of rendering.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If sound sources are localized in 3D space, then spatial audio immersion is improved, but measurement precision and computational complexity increase

Engineering Contradiction:
Improvesound source position accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses 3D model fragments as intermediary reference objects for sound source localization. Instead of performing complex acoustic field analysis throughout the entire 3D space, the system localizes sound sources relative to these pre-positioned fragments, which provide known spatial references that simplify the computation while maintaining precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the 3D space into regions associated with individual 3D model fragments. Each fragment serves as a localized reference frame, allowing the system to perform sound source localization independently within each fragment's region rather than performing a single complex global analysis, thereby reducing overall computational complexity.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables accurate and immersive real-time rendering of 3D items and sound sources within augmented reality, enhancing user interaction and experience by overcoming resource limitations and translation challenges.

Implementation Method 1

an array of microphones that capture sound of a three-dimensional (3D), real-world scene

Methodology Applied
Scientific EffectSound wave detection: Sound

Data Source

PatentUS10158939B2Sound Source association
Publication Date: 2018.12.18 SEIKO EPSON CORP
  • US10158939B2 patent drawing
  • US10158939B2 patent drawing
  • US10158939B2 patent drawing

AI summary

Multiple Holocam Orbs observe a real-life environment and generate an artificial reality representation of the real-life environment. Depth image data is cleansed of error due to LED shadow by identifying the edge of a foreground object in an (near infrared light) intensity image, identifying an edge in a depth image, and taking the difference between the start of both edges. Depth data error due to parallax is identified noting when associated text data in a given pixel row that is progressing in a given row direction (left-to-right or right-to-left) reverses order. Sound sources are identified by comparing results of a blind audio source localization algorithm, with the spatial 3D model provided by the Holocam Orb. Sound sources that corresponding to identifying 3D objects are associated together. Additionally, types of data supported by a standard movie data container, such as an MPEG container, is expanding to incorporate free viewpoint data (FVD) model data. This is done by inserting FVD data of different individual 3D objects at different sample rates into a single video stream. Each 3D object is separately identified by a separately assigned ID.