Real-Time Skill Assessment via Hand Tracking and Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current skill assessment systems using video cameras are inadequate for complex tasks with multiple sub-tasks, as they require manual annotation and are not automated, limiting their ability to monitor and improve task workflow designs effectively.

Innovation Solution

The implementation of deep learning methods, including a bottom-up approach using Convolutional Neural Networks (CNN) and Temporal Convolutional Nets (TCNs), to detect body parts, group sequential frames into sub-tasks, and evaluate task completion and order correctness, providing an automated skill assessment system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation and analysis methods are used for skill assessment, then measurement precision can be maintained, but device complexity and loss of time increase significantly

Engineering Contradiction:
Improveskill assessment accuracyVSAvoidtime for manual annotation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by automatically annotating and assessing skills through computer vision and machine learning models. The camera system captures video, the ML model automatically segments and labels sub-tasks, and the system evaluates skill completion without requiring manual annotation, thereby eliminating time loss while maintaining assessment accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual annotation process with an automated computer vision system. Instead of manually reviewing and annotating video clips, the system uses camera-based detection, optical flow analysis, and deep learning models to automatically identify sub-tasks, segment video, and assess skill completion, significantly reducing time requirements

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated camera systems are implemented for task monitoring, then productivity improves, but device complexity increases

Engineering Contradiction:
Improvetask workflow monitoring efficiencyVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex task monitoring function into manageable components: camera capture, optical flow computation, sub-task detection, video segmentation, and skill assessment. Each component handles a specific aspect of the overall system, making the complex productivity-monitoring task achievable through modular automated systems

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The camera system serves multiple functions simultaneously: capturing video, detecting body parts, tracking hand movements, identifying sub-tasks, and assessing skill completion. This multi-functionality reduces the need for separate specialized devices, thereby improving productivity while controlling device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If deep learning models are applied for real-time analysis, then measurement precision and productivity improve, but use of energy increases

Engineering Contradiction:
Improvereal-time skill assessment capabilityVSAvoidcomputational energy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by processing only the necessary portions of video data through deep learning models. Instead of analyzing every frame in detail, the system uses optical flow to identify significant changes, segments video into relevant clips, and applies ML models only to those segments, reducing computational energy consumption while maintaining real-time assessment capability

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes parameters dynamically by adjusting the level of computational analysis based on task complexity and real-time requirements. The deep learning models process images at varying resolutions and frame rates depending on the specific sub-task being performed, optimizing the balance between real-time productivity and energy consumption

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11093886B2Methods for real-time skill assessment of multi-step tasks performed by hand movements using a video camera
Publication Date: 2021.08.17 FUJIFILM BUSINESS INNOVATION CORP
  • US11093886B2 patent drawing
  • US11093886B2 patent drawing
  • US11093886B2 patent drawing

AI summary

Example implementations described herein are directed to systems and methods for skill assessment, such as hand washing compliance in hospitals, or assembling products in factories. Example implementations involve body part tracking (e.g., hands), skeleton tracking and deep neural networks to detect and recognize sub-tasks and to assess the skill on each sub-task. Furthermore, the order of the sub-tasks is checked for correctness. Beyond monitoring individual users, example implementations can be used for analyzing and improving workflow designs with multiple sub-tasks.