Neural Network Virtual Coach for Real-Time Exercise Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current virtual assistants lack the ability for sophisticated visual interaction, including understanding video data, spatial and temporal relations, and providing effective real-time feedback for personal coaching, due to limitations in determining training data and subjective labeling processes.

Innovation Solution

A neural network system configured to process real-time camera streams for providing feedback, utilizing a backbone network and head networks for activity classification, exercise scoring, and event detection, which generates feedback inferences such as exercise scores, calorie estimation, and form feedback, and includes a method for generating a feedback model through video sample labeling and gradient-based optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human coaches provide fitness training, then coaching quality and personalization are improved, but cost increases significantly

Engineering Contradiction:
Improvecoaching qualityVSAvoidcost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates a virtual coach that copies and simulates human coaching capabilities through neural networks and computer vision technology. The system captures video of exercise performances and uses AI models to provide coaching feedback, replacing expensive human coaches with affordable automated virtual assistants while maintaining coaching quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of human coaches with an automated computer vision system. The neural network processes video data and generates coaching feedback automatically, eliminating the need for human physical presence while delivering consistent coaching quality at low cost

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If automated virtual assistants use voice-based interaction, then device complexity is reduced, but interaction sophistication and visual understanding capability deteriorate

Engineering Contradiction:
Improvesimplicity of interactionVSAvoidvisual interaction capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent merges voice-based interaction with visual interaction capabilities in the virtual coach system. The neural network processes both audio commands and video data simultaneously, allowing the assistant to understand spatial relationships, exercise form, and provide comprehensive coaching through multiple modalities while maintaining device simplicity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The virtual coach system is designed with multi-functionality, handling both voice commands and visual analysis through a single integrated neural network architecture. This universal system can perform activity recognition, pose estimation, and coaching feedback delivery through various output modalities including visual and audio channels

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If video data is labeled by human reviewers for training, then measurement precision is improved, but productivity and scalability deteriorate

Engineering Contradiction:
Improvelabeling accuracyVSAvoiddata labeling speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses pre-labeled training datasets to train the neural network models before deployment. The system performs preliminary training with expert-labeled data to achieve high measurement precision, then uses the trained model for automated inference at scale without requiring continuous human labeling, thus resolving the productivity bottleneck

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The trained neural network model performs self-service by automatically labeling and analyzing video data without requiring ongoing human reviewer involvement. Once trained on precision-labeled data, the system independently processes unlimited video inputs with consistent accuracy, achieving both high precision and unlimited productivity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230082953A1System and Method for Real-Time Interaction and Coaching
Publication Date: 2023.03.16 QUALCOMM INC
  • US20230082953A1 patent drawing
  • US20230082953A1 patent drawing
  • US20230082953A1 patent drawing

AI summary

Methods and systems are described for real-time instruction and coaching using a virtual assistant for interaction with a user. Users may receive feedback inferences provided generally in real-time after collection of video samples from the user device. Neural network architectures and layers may be used to determine motion patterns and temporal aspects of the video samples, as well as detect activities of the foreground user despite background noise. The methods and systems may have various capabilities, including but not limited to live feedback on performed exercise activities, exercise scoring, calorie estimation, and repetition counting.