Gesture Path Recognition via Attribute Invariant Data Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer systems are limited in their ability to efficiently recognize and interpret human gestures in three-dimensional space, as existing user interfaces rely on two-dimensional interactions, which are inadequate for conveying complex gestures and additional information effectively.
Innovation Solution
A method using artificial intelligence to identify path assemblies by extracting attributes such as position, size, and direction from point groups along paths, transforming the data to be attribute-invariant, and inputting this data into an AI system to distinguish between gesture meanings and additional information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If two-dimensional user interface is used for gesture recognition, then device complexity is reduced, but gesture recognition accuracy and reliability deteriorate
Solution Approach 1:
The patent transitions from two-dimensional screen-based gesture recognition to three-dimensional spatial gesture recognition. The system tracks hand positions, movements, and gestures in 3D space using cameras and sensors, enabling accurate recognition of complex gestures that cannot be effectively captured in 2D. This dimensional expansion resolves the contradiction by maintaining low device complexity while dramatically improving gesture recognition reliability through spatial awareness.
2Speed
If simple positional data is used for path recognition, then processing speed is improved, but gesture meaning discrimination accuracy deteriorates
Solution Approach 1:
The patent segments gesture data into multiple attribute categories: positional information, movement velocity, acceleration, direction, and temporal characteristics. Each attribute is processed independently through dedicated algorithms, allowing the system to maintain high processing speed while capturing comprehensive gesture features. This segmentation enables parallel processing of different gesture aspects, resolving the contradiction between speed and accuracy.
Solution Approach 2:
The system transforms raw gesture data into multiple derived parameters including position, velocity, acceleration, and direction. By changing the parameter representation from simple coordinates to comprehensive motion descriptors, the system enhances gesture discrimination accuracy without significantly increasing processing time, as each parameter can be computed efficiently from the raw data.
3Measurement precision
If complex gesture attributes are processed, then gesture interpretation accuracy is improved, but computational energy consumption increases
Solution Approach 1:
The patent implements a hierarchical processing approach where only essential gesture attributes are fully processed based on the recognition context. For common gestures, simplified processing is used; for complex or ambiguous gestures, full attribute analysis is activated. This partial processing strategy maintains high accuracy for critical cases while reducing overall computational energy consumption.
Data Source
AI summary
The shape or movement of a gesturing body or portion thereof in two- or three-dimensional space is ascertained from the path of the outline of the shape or the path of the movement. The method disclosed involves receiving input of data that represents a path and using artificial intelligence to recognize the meaning of the path, i.e., to recognize which of a plurality of pre-prepared meanings is the meaning of a gesture. As pre-processing for inputting location data for a point group along the path to the artificial intelligence, at least one attribute from among the location, size, and direction of the entire point group is extracted, and location data for the point group is converted to attribute invariant location data that is relative to the extracted attribute(s) but not dependent on the extracted attribute(s). Then data that includes the attribute invariant location data and the extracted attribute(s) is inputted to the artificial intelligence as input data. The result is efficient and effective processing of gesture-based dialogue between a user and a computer.


