3D Room Analysis Using Audio-Guided Computer Vision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computer vision systems struggle to accurately identify objects of interest in physical spaces, such as windows, doors, and furniture, due to the use of generic classifiers that often misclassify or fail to identify specific objects, leading to inefficient processing and resource usage.

Innovation Solution

The system prompts users to describe the physical space during scanning, using speech cues to guide computer vision algorithms, allowing for more accurate object identification by focusing on specific objects and reducing the need for multiple machine learning models, thereby improving processing efficiency and resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If generic classifiers are used to identify objects in physical spaces, then the system can process all objects uniformly, but the accuracy of identifying specific objects such as windows, doors, and furniture deteriorates

Engineering Contradiction:
Improveability to process all objectsVSAvoidaccuracy of object identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system segments the object identification process by dividing physical spaces into distinct regions (e.g., windows, doors, furniture areas) and applying specialized machine learning models to each region based on audio cues, rather than using a single generic classifier for all objects

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different quality levels of processing to different regions by using audio cues to identify areas of interest and applying more sophisticated, specialized machine learning models specifically to those regions, while using simpler processing for other areas

Inventive Principle:
Principle #3Local quality

2Measurement precision

If multiple machine learning models are applied to identify different objects, then the accuracy of object identification improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveaccuracy of object identificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by using audio processing to identify areas of interest before applying machine learning models, which allows it to focus computational resources only on relevant regions rather than processing the entire space uniformly

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by selectively applying machine learning models only to identified areas of interest rather than to the entire physical space, reducing overall processing time while maintaining accuracy for important objects

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If the system processes the entire physical space uniformly, then the coverage is complete, but the efficiency of processing deteriorates

Engineering Contradiction:
Improvecompleteness of processingVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary audio processing to identify areas of interest before visual processing, which guides subsequent focused processing of only those regions, maintaining reliability for important objects while improving overall efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies different processing quality levels to different regions by focusing detailed analysis on audio-identified areas of interest while using coarser processing for other regions, balancing completeness with efficiency

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11521376B1Three-dimensional room analysis with audio input
Publication Date: 2022.12.06 AMAZON TECH INC
  • US11521376B1 patent drawing
  • US11521376B1 patent drawing
  • US11521376B1 patent drawing

AI summary

System and methods are provided that generate a three-dimensional model from a physical space. While a user is scanning and/or recording the physical space with a user computing device, user speech describing the physical space is recorded. A transcript is generated from the audio captured during the scan and/or image recording of the physical space. Keywords from the transcript are used to improve computer-vision object identification, which is incorporated in the three-dimensional model.