Machine-Learning Gesture Recognition with PPG and Accelerometer Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices struggle to accurately detect user gestures without relying on touch input, particularly in scenarios where touch input is not feasible or desirable.

Innovation Solution

Utilizing a machine-learning based approach that combines outputs from biosignal and non-biosignal sensors, such as PPG and accelerometers, to predict user gestures through a neural network model trained on a general population, which can be personalized for specific users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If touch input is used for gesture detection, then detection accuracy is improved, but usability is reduced in scenarios where touch is not feasible

Engineering Contradiction:
Improvegesture detection accuracyVSAvoidusability without touch
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces mechanical touch-based gesture detection with biosignal-based detection using optical sensors (PPG) and motion sensors (accelerometers). This substitution enables gesture recognition through non-contact means, resolving the contradiction by maintaining detection accuracy while eliminating the need for physical touch in scenarios where it is not feasible.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces biosignal sensors and motion sensors as intermediary devices between the user's gestures and the electronic device's processing system. These sensors capture physiological and motion data that serves as a mediator for gesture recognition, enabling accurate detection without direct touch contact.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple sensors are combined for gesture recognition, then recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvegesture recognition accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple sensor types (biosignal sensors and motion sensors) into a unified gesture recognition system. By combining the complementary data from these sensors and processing it through a machine learning model, the system achieves improved recognition accuracy while managing complexity through integrated processing rather than separate independent systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-functional sensor system where both biosignal sensors and motion sensors serve the common purpose of gesture recognition. This universal approach allows the same processing architecture to handle diverse gesture types using different sensor modalities, improving accuracy without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If machine learning model is trained on general population, then adaptability to multiple users is improved, but precision for individual users may be reduced

Engineering Contradiction:
Improvemulti-user capabilityVSAvoidindividual user gesture accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements a dynamic training approach where the machine learning model can adapt between general population data and user-specific data. The system starts with a general model for multi-user capability and can dynamically refine its parameters with individual user data when available, balancing adaptability and precision through flexible learning mechanisms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent performs preliminary training on general population data to create a universal gesture recognition model that can immediately serve multiple users. This preliminary action establishes a solid baseline that can then be fine-tuned with minimal user-specific data, achieving both broad adaptability and individual precision through staged learning.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables gesture recognition across multiple users without prior knowledge of individual characteristics, allowing for gesture-powered control of electronic devices and seamless interaction through wearable sensors.

Implementation Method 1

a smartwatch may be equipped with one or more biosignal sensors (e.g., a photoplethysmogram (PPG) sensor)

Methodology Applied
Scientific EffectPhotoplethysmography:

Implementation Method 2

other types of sensors (e.g., a motion sensor, an optical sensor, an audio sensor and the like)

Methodology Applied
Scientific EffectAcceleration: Accelerometer

Data Source

PatentEP4038477B1Machine-learning based gesture recognition using multiple sensors
Publication Date: 2025.07.09 APPLE INC
  • EP4038477B1 patent drawingFigure 1
  • EP4038477B1 patent drawingFigure 2
  • EP4038477B1 patent drawingFigure 3

AI summary

A device implementing a system for machine-learning based gesture recognition includes at least one processor configured to, receive, from a first sensor of the device, first sensor output of a first type, and receive, from a second sensor of the device, second sensor output of a second type that differs from the first type. The at least one processor is further configured to provide the first sensor output and the second sensor output as inputs to a machine learning model, the machine learning model having been trained to output a predicted gesture based on sensor output of the first type and sensor output of the second type. The at least one processor is further configured to determine the predicted gesture based on an output from the machine learning model, and to perform, in response to determining the predicted gesture, a predetermined action on the device.