Voice and Gesture Object Search Using Image Area Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current object search methods are inflexible, requiring users to clearly describe search criteria, which can lead to poor search results when criteria cannot be accurately defined, such as searching for objects by color or shape.

Innovation Solution

The method involves receiving voice input and gesture input from users to determine the name and characteristic category of a target object, extracting characteristic information from an image area selected via gesture input, and using this information to search for the object, allowing for a more flexible search process without requiring precise criterion description.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users are required to clearly describe search criteria using preset options or direct input, then search structure is simplified and ease of operation is improved, but search accuracy and adaptability deteriorate when users cannot accurately describe their search intent

Engineering Contradiction:
Improveease of operationVSAvoidadaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary system that includes image processing modules and characteristic extraction algorithms. This intermediary translates visual image data into structured search criteria automatically, mediating between the user's visual selection and the search system without requiring the user to manually describe criteria

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical interaction of manual criterion selection and description with automated image processing and characteristic extraction. The system automatically analyzes selected image regions, extracts relevant characteristics (color, shape, texture), and converts them into search parameters, substituting manual mechanical operations with automated computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If multiple input methods (voice and gesture) are integrated, then adaptability and user flexibility are improved, but device complexity increases

Engineering Contradiction:
ImproveadaptabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple input modalities (voice recognition and gesture recognition) into a unified search system. The system integrates these different input methods to work together, allowing users to switch between or combine them based on their needs, thereby improving adaptability while managing complexity through integrated design

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If search criteria must be precisely defined, then measurement precision is improved, but ease of operation and adaptability worsen when dealing with subjective or irregular characteristics

Engineering Contradiction:
Improvemeasurement precisionVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements self-service functionality where the system automatically performs characteristic extraction and criterion formulation. Instead of requiring users to manually define precise criteria, the system autonomously analyzes the selected image region, extracts relevant characteristics, and generates search criteria automatically, serving itself to bridge the gap between imprecise user input and precise search requirements

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10311115B2Object search method and apparatus
Publication Date: 2019.06.04 HUAWEI TECH CO LTD
  • US10311115B2 patent drawing
  • US10311115B2 patent drawing
  • US10311115B2 patent drawing

AI summary

An object search method and apparatus, where the method includes receiving voice input and gesture input that are of a user; determining, according to the voice input, a name of a target object for which the user expects to search and a characteristic category of the target object; extracting characteristic information of the characteristic category from an image area selected by the user by means of the gesture input; and searching for the target object according to the extracted characteristic information and the name of the target object. The solutions provided in the embodiments of the present disclosure can provide a user with a more flexible search manner, and reduce a restriction on an application scenario during a search.