Audio-Visual Floorplan Reconstruction via Sparse Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for generating virtual models of three-dimensional spaces from digital images are inefficient, inflexible, and inaccurate, requiring large amounts of data and being limited to depicted areas, unable to adapt to sparse inputs or infer unviewed portions.

Innovation Solution

The use of an audio-visual floorplan reconstruction machine learning model that combines visual and audio information from sparse digital videos to generate two-dimensional floorplans, including unviewed areas, through a multi-modal encoder-decoder framework with self-attention layers, allowing for the inference of room structures and semantic labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems use high-volume pixel-dense images to model three-dimensional spaces, then modeling accuracy is improved, but computational resource requirements and data storage demands increase exponentially

Engineering Contradiction:
Improvemodeling accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines visual features from digital images with audio features from digital audio clips to form multi-modal representations. This merging allows the system to infer spatial information from audio data, reducing dependence on large volumes of visual data while maintaining or improving modeling accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces audio information as an intermediary modality that complements visual data. Audio features serve as a mediator that provides additional spatial and environmental information, enabling accurate space modeling with fewer visual inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If conventional systems require specific sequences of digital photos to generate virtual models, then model completeness within viewed areas is improved, but system flexibility and adaptability to various input types deteriorate

Engineering Contradiction:
Improvemodel completenessVSAvoidinput flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal modeling system that can process multiple input types including digital photos, digital videos, and digital audio clips. The multi-modal framework allows the same system to handle diverse input formats and capture methods, significantly improving adaptability while maintaining model completeness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs dynamic feature extraction and alignment processes that adapt to different input types and sequences. The system can dynamically adjust to various capture methods and input qualities, making the modeling process flexible and adaptable rather than rigid and sequence-dependent.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If conventional systems generate models limited to viewed areas in digital images, then processing complexity is reduced, but model accuracy and completeness of the entire space deteriorate

Engineering Contradiction:
Improveprocessing complexityVSAvoidspace representation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent adds the audio modality dimension to the traditional visual-only approach. By incorporating audio information as an additional dimension of data, the system can infer information about unviewed areas without significantly increasing visual processing complexity, thereby improving overall space representation accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11810354B2Generating digital floorplans from sparse digital video utilizing an audio-visual floorplan reconstruction machine learning model
Publication Date: 2023.11.07 META PLATFORMS INC
  • US11810354B2 patent drawing
  • US11810354B2 patent drawing
  • US11810354B2 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing multiple modalities to generate accurate two-dimensional floorplans based on sparse digital videos depicting three-dimensional space. In particular, in one or more embodiments, the disclosed systems extract both visual and audio information from sparse digital video coverage of portions of a three-dimensional space and utilize the extracted visual and audio information to generate a two-dimensional floorplan representing both viewed and unviewed portions of the three-dimensional space. For example, the disclosed systems utilize self-attention layers of a specialized machine learning model to maintain and leverage bi-directional relationships among sequences of visual and audio features to generate floorplan predictions associated with the three-dimensional space. The disclosed systems then combine the predictions to generate the two-dimensional floorplan including a geometric layout and one or more semantic room labels.