Audio-Visual Floorplan Reconstruction via Sparse Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for generating virtual models of three-dimensional spaces from digital images are inefficient, inflexible, and inaccurate, requiring large amounts of data and being limited to depicted areas, unable to adapt to sparse inputs or infer unviewed portions.
Innovation Solution
The use of an audio-visual floorplan reconstruction machine learning model that combines visual and audio information from sparse digital videos to generate two-dimensional floorplans, including unviewed areas, through a multi-modal encoder-decoder framework with self-attention layers, allowing for the inference of room structures and semantic labels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional systems use high-volume pixel-dense images to model three-dimensional spaces, then modeling accuracy is improved, but computational resource requirements and data storage demands increase exponentially
Solution Approach 1:
The patent combines visual features from digital images with audio features from digital audio clips to form multi-modal representations. This merging allows the system to infer spatial information from audio data, reducing dependence on large volumes of visual data while maintaining or improving modeling accuracy.
Solution Approach 2:
The patent introduces audio information as an intermediary modality that complements visual data. Audio features serve as a mediator that provides additional spatial and environmental information, enabling accurate space modeling with fewer visual inputs.
2Manufacturing precision
If conventional systems require specific sequences of digital photos to generate virtual models, then model completeness within viewed areas is improved, but system flexibility and adaptability to various input types deteriorate
Solution Approach 1:
The patent creates a universal modeling system that can process multiple input types including digital photos, digital videos, and digital audio clips. The multi-modal framework allows the same system to handle diverse input formats and capture methods, significantly improving adaptability while maintaining model completeness.
Solution Approach 2:
The patent employs dynamic feature extraction and alignment processes that adapt to different input types and sequences. The system can dynamically adjust to various capture methods and input qualities, making the modeling process flexible and adaptable rather than rigid and sequence-dependent.
3Device complexity
If conventional systems generate models limited to viewed areas in digital images, then processing complexity is reduced, but model accuracy and completeness of the entire space deteriorate
Solution Approach 1:
The patent adds the audio modality dimension to the traditional visual-only approach. By incorporating audio information as an additional dimension of data, the system can infer information about unviewed areas without significantly increasing visual processing complexity, thereby improving overall space representation accuracy.
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing multiple modalities to generate accurate two-dimensional floorplans based on sparse digital videos depicting three-dimensional space. In particular, in one or more embodiments, the disclosed systems extract both visual and audio information from sparse digital video coverage of portions of a three-dimensional space and utilize the extracted visual and audio information to generate a two-dimensional floorplan representing both viewed and unviewed portions of the three-dimensional space. For example, the disclosed systems utilize self-attention layers of a specialized machine learning model to maintain and leverage bi-directional relationships among sequences of visual and audio features to generate floorplan predictions associated with the three-dimensional space. The disclosed systems then combine the predictions to generate the two-dimensional floorplan including a geometric layout and one or more semantic room labels.


