Automatic Room Capture Using Neural Network Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing room capture technologies in extended reality (XR) require manual user intervention, leading to low capture efficiency and accuracy in mixed reality (MR) scenarios.

Innovation Solution

A full-automatic capture method and apparatus that uses a camera to acquire RGB images and depth information, processing this data with a capture model to automatically identify and position objects in a room, and displaying a 3D model of the captured objects in real time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual capture mode is used, then user control over capture process is maintained, but capture efficiency is low and capture accuracy is poor

Engineering Contradiction:
Improvecapture efficiencyVSAvoidautomation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The capture system performs automatic object detection, classification, and 3D model generation without requiring user intervention. The system captures images, processes them through neural networks to identify objects and their categories, automatically generates 3D models, and completes the entire capture workflow autonomously, eliminating the need for manual capture operations.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical capture operations with an automated computational system. Instead of users manually positioning and capturing objects, the system uses cameras to acquire images, neural networks to process and identify objects, and automated algorithms to generate 3D models, substituting human manual operations with automated computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual capture mode is used, then user control is maintained, but measurement precision and manufacturing precision of capture results are insufficient

Engineering Contradiction:
Improvecapture accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces neural networks as an intermediary between image acquisition and object identification. The neural networks process captured images, automatically identify objects, classify them by category, and determine their positions, serving as a sophisticated mediator that enhances capture accuracy without requiring direct manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system automatically generates 3D model copies of captured objects based on processed image data. Instead of manual modeling, the system creates accurate digital replicas of physical objects by analyzing captured images through neural networks, automatically generating precise 3D representations that maintain geometric fidelity.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250054256A1Method, device and medium of a full-automatic capture for room
Publication Date: 2025.02.13 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250054256A1 patent drawing
  • US20250054256A1 patent drawing
  • US20250054256A1 patent drawing

AI summary

Embodiments of the present disclosure provide a method, apparatus, device and medium of a full-automatic capture for a room, and the method comprises: acquiring an RGB image of a room to be captured, depth information of the RGB image and camera pose information and inputting them into a capture model to obtain capture information of an object in the RGB image, wherein the capture information of the object comprises a category of the object and position information of the object. A capture box of the object is displayed in a VST image of the room according to the capture information of the object, a 3D model of the object is added in a 3D model of the room, and the 3D model of the room is generated according to a pre-determined size ratio for the captured object in the room.