Multi-Modal Process Tracking for Real-Time User Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users often struggle with performing complex processes due to difficulty in understanding and following instructions, and determining if they are performing the process correctly, leading to errors and inefficiencies.

Innovation Solution

A computing system that tracks a user's process using multi-modal sensor information to recognize the process being performed, provides domain-specific instructions, and offers real-time guidance based on the user's actions and environment, using a head-mounted display device to synchronize user and world states with a digital twin model for accurate feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multi-modal sensor information is used to track user actions and provide real-time guidance, then the accuracy and reliability of process execution is improved, but the device complexity increases due to multiple sensors and processing systems

Engineering Contradiction:
Improveprocess execution accuracyVSAvoidsensor system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the complex process tracking into separate functional modules: world state tracking module, user state tracking module, process recognition module, and guidance generation module. Each module processes specific sensor data independently before integrating results, making the complex system more manageable and maintainable while improving reliability through specialized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The computing system is designed as a universal platform that can track multiple types of processes across different domains using the same core architecture. The system handles diverse sensor modalities (visual, auditory, tactile) and process types through a unified state tracking and recognition framework, reducing overall system complexity despite the multi-functional capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If real-time tracking and multi-modal sensor fusion is implemented, then the responsiveness and adaptability of guidance is improved, but the computational resources and processing time required increase

Engineering Contradiction:
Improveguidance adaptabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary processing of sensor data by continuously maintaining updated world state and user state representations before process recognition is needed. Pre-computed spatial relationships and object models are ready for immediate use during process execution, reducing real-time computational burden while maintaining high adaptability through pre-positioned processing capabilities.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts the level of processing based on process complexity and user needs. During simple tasks, the system uses lighter processing modes with fewer sensor modalities activated. When complex processes are detected, the system automatically engages more comprehensive sensor fusion and multi-modal analysis, optimizing computational resource consumption according to actual task requirements rather than operating at maximum capacity continuously.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If comprehensive process tracking is provided, then the completeness of process monitoring is improved, but the difficulty of detecting and measuring user actions increases

Engineering Contradiction:
Improveprocess information completenessVSAvoiduser action detection complexity
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The system merges multiple sensor modalities (visual, auditory, tactile) into a unified state representation that comprehensively captures user actions. By fusing data from different sensor types, the system achieves complete process monitoring while the integrated representation simplifies action detection compared to analyzing each sensor stream separately, reducing overall detection complexity through synthesis.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a virtual copy or digital twin of the physical process state through the world state and user state models. This virtual representation mirrors real-world actions, allowing comprehensive monitoring without directly measuring every physical detail. The digital twin approach simplifies detection by working with abstracted, processed information rather than raw sensor data, reducing measurement complexity while maintaining information completeness.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12469402B2Multi-modal sensor based process tracking and guidance
Publication Date: 2025.11.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12469402B2 patent drawing
  • US12469402B2 patent drawing
  • US12469402B2 patent drawing

AI summary

Examples are disclosed that relate to computer-based tracking of a process performed by a user. In one example, multi-modal sensor information is received via a plurality of sensors. A world state of a real-world physical environment and a user state in the real-world physical environment are tracked based on the multi-modal sensor information. A process being performed by the user within a working domain is recognized based on the world state and the user state. A current step in the process is detected based on the world state and the user state. Domain-specific instructions directing the user how to perform an expected action are presented via a user interface device. A user action is detected based on the world state and the user state. Based on the user action differing from the expected action, domain-specific guidance to perform the expected action is presented via the user interface device.