Multimodal Smart Eyewear for Natural Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality technologies face challenges in efficiently connecting offline real scenes with online virtual information for natural and intuitive human-machine interaction, particularly in detecting and recognizing objects without manual labeling, which limits user experience and interaction flexibility.

Innovation Solution

A smart eyewear apparatus that utilizes multimodal inputs such as image, voice, touch, and sensing information to generate operation commands through comprehensive logical analysis, enabling users to interact with both real and virtual environments in a more natural and intuitive manner, with a split-mount control device processing the core logic to reduce device size and heat issues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If manual labeling technologies (2D code, NFC, WiFi positioning) are used to detect offline objects, then object detection can be achieved, but deployment and maintenance costs increase, and interaction becomes less natural

Engineering Contradiction:
Improveobject detection capabilityVSAvoiddeployment and maintenance cost
Core Design Contradiction:
Difficulty of detecting and measuringVSEase of manufacture

Solution Approach 1:

The patent extracts the labeling requirement from the interaction system by using natural picture recognition technology that can automatically identify objects without any manual labeling. The system processes image data acquired by a camera to automatically determine object identity, category, and spatial posture, eliminating the need for 2D codes, NFC tags, or WiFi positioning infrastructure.

Inventive Principle:
Principle #2Taking out (Extraction)

2Difficulty of detecting and measuring

If manual labeling technologies are used, then object detection is possible, but the interaction loses naturalness and intuitiveness

Engineering Contradiction:
Improveobject detection capabilityVSAvoidinteraction naturalness
Core Design Contradiction:
Difficulty of detecting and measuringVSEase of operation

Solution Approach 1:

The system enables self-service object recognition where the computer automatically identifies and tracks objects in the real scene without requiring users to manually label or prepare objects. The natural picture recognition technology performs intelligent analysis of image data to automatically determine object identity, category, and spatial posture, making the interaction as natural as human visual perception.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If all processing functions are integrated in the smart eyewear apparatus, then interaction functionality is complete, but device size increases and heat dissipation becomes problematic

Engineering Contradiction:
Improveinteraction functionalityVSAvoiddevice weight and size
Core Design Contradiction:
Adaptability or versatilityVSWeight of moving object

Solution Approach 1:

The patent segments the smart eyewear system into a lightweight wearable apparatus that captures images and a separate processing system that performs comprehensive logic analysis. The smart eyewear apparatus primarily acquires image data and transmits it to an external device for processing, thereby reducing the computational burden and hardware requirements of the wearable component while maintaining complete interaction functionality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10664060B2Multimodal input-based interaction method and device
Publication Date: 2020.05.26 HISCENE INFORMATION TECH CO LTD
  • US10664060B2 patent drawing
  • US10664060B2 patent drawing
  • US10664060B2 patent drawing

AI summary

An object of the present disclosure is to provide a method for interacting based on multimodal inputs, which enables a higher approximation to user natural interaction, comprising: acquiring a plurality of input information from at least one of a plurality of input modules; performing comprehensive logic analysis of the plurality of input information so as to generate an operation command, wherein the operation command has operation elements, the operation elements at least including an operation object, an operation action, and an operation parameter; and performing a corresponding operation on the operation object based on the operation command.