Voice and Gesture Object Search Using Image Area Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current object search methods are inflexible, requiring users to clearly describe search criteria, which can lead to poor search results when criteria cannot be accurately defined, such as searching for objects by color or shape.
Innovation Solution
The method involves receiving voice input and gesture input from users to determine the name and characteristic category of a target object, extracting characteristic information from an image area selected via gesture input, and using this information to search for the object, allowing for a more flexible search process without requiring precise criterion description.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If users are required to clearly describe search criteria using preset options or direct input, then search structure is simplified and ease of operation is improved, but search accuracy and adaptability deteriorate when users cannot accurately describe their search intent
Solution Approach 1:
The patent introduces an intermediary system that includes image processing modules and characteristic extraction algorithms. This intermediary translates visual image data into structured search criteria automatically, mediating between the user's visual selection and the search system without requiring the user to manually describe criteria
Solution Approach 2:
The patent replaces the mechanical interaction of manual criterion selection and description with automated image processing and characteristic extraction. The system automatically analyzes selected image regions, extracts relevant characteristics (color, shape, texture), and converts them into search parameters, substituting manual mechanical operations with automated computational processes
2Adaptability or versatility
If multiple input methods (voice and gesture) are integrated, then adaptability and user flexibility are improved, but device complexity increases
Solution Approach 1:
The patent merges multiple input modalities (voice recognition and gesture recognition) into a unified search system. The system integrates these different input methods to work together, allowing users to switch between or combine them based on their needs, thereby improving adaptability while managing complexity through integrated design
3Measurement precision
If search criteria must be precisely defined, then measurement precision is improved, but ease of operation and adaptability worsen when dealing with subjective or irregular characteristics
Solution Approach 1:
The patent implements self-service functionality where the system automatically performs characteristic extraction and criterion formulation. Instead of requiring users to manually define precise criteria, the system autonomously analyzes the selected image region, extracts relevant characteristics, and generates search criteria automatically, serving itself to bridge the gap between imprecise user input and precise search requirements
Data Source
AI summary
An object search method and apparatus, where the method includes receiving voice input and gesture input that are of a user; determining, according to the voice input, a name of a target object for which the user expects to search and a characteristic category of the target object; extracting characteristic information of the characteristic category from an image area selected by the user by means of the gesture input; and searching for the target object according to the extracted characteristic information and the name of the target object. The solutions provided in the embodiments of the present disclosure can provide a user with a more flexible search manner, and reduce a restriction on an application scenario during a search.


