Camera-Based Gesture Screening for Self-Ordering Customer Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing self-ordering systems struggle with accurately identifying the intended customer for order input, often mistakenly recognizing gestures from bystanders or accompanying individuals rather than the actual customer.

Innovation Solution

An information processing device and program that utilizes a camera to detect and identify a first gesture, then classify and designate a candidate based on region and posture, using facial distance measurements to enhance accuracy in identifying the intended customer.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If gesture recognition is used for self-ordering, then ordering speed is improved, but identification accuracy deteriorates due to false recognition of bystander gestures

Engineering Contradiction:
Improveordering speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transitions from two-dimensional image data to three-dimensional spatial information by introducing depth maps and distance calculations. The camera system captures both image data and depth information, enabling the system to determine the actual distance between the customer and the terminal, as well as the distance to bystanders. This dimensional enhancement allows accurate differentiation between intended and unintended gesture makers based on spatial positioning.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an identification unit as an intermediary component that mediates between the gesture detection and order processing. This unit uses multiple criteria (gesture detection, region determination based on distance, and candidate selection) to filter and verify the intended customer before allowing order input. The intermediary layer prevents false recognition by validating that the detected gesture originates from the correct customer.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If simple gesture detection is implemented, then system complexity is reduced, but reliability deteriorates due to erroneous order inputs from wrong customers

Engineering Contradiction:
Improvesystem complexityVSAvoidorder input reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the customer identification process into distinct functional units: a detection unit for gesture recognition, a measurement unit for distance calculation, a region determination unit for spatial validation, and an identification unit for final customer verification. This segmentation allows each component to perform its specific function with high reliability while maintaining manageable system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors multiple parameters (gesture presence, distance measurements, region validation) and adjusts its identification decisions based on this feedback. The identification unit receives feedback from the detection and measurement units to verify whether the detected gesture should be acted upon, ensuring reliable order input by confirming the customer's intent through multiple validation checks.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12354410B2Information processing device and information processing program
Publication Date: 2025.07.08 TOSHIBA TEC KK
  • US12354410B2 patent drawing
  • US12354410B2 patent drawing
  • US12354410B2 patent drawing

AI summary

According to one embodiment, an information processing device for a gesture-based ordering taking system or the like includes a processor. The processor is configured to detect a person making a particular gesture in image data from a camera, identify each person making the particular gesture in the image data as a first type person, then identify each first type person meeting a candidate condition in the image data as a candidate. The processor then designates a candidate as a subject according to a subject selection condition, and then detects gestured-based input from the subject in additional image data from the camera after the subject has been designated.