Robot Action Recognition Using On-Screen and Off-Screen Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional AI systems struggle to accurately recognize human actions, especially when the subject deviates from the camera's view, limiting their effectiveness in real-world applications such as robotics, where interaction and service provision require comprehensive situational awareness.
Innovation Solution
The proposed solution involves generating both on-screen and off-screen labels using image data and sensor data respectively, enabling self-supervised learning to recognize human actions even when they are outside the camera's view, thereby enhancing action recognition performance and enabling robots to interact effectively with humans in various scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If action recognition relies solely on camera-based visual data, then recognition accuracy is high when the subject is within view, but recognition performance drops significantly when the subject deviates from the camera's view
Solution Approach 1:
The patent combines multiple data sources including on-screen camera data, off-screen sensor data (accelerometer, gyroscope, magnetometer), and audio data to create a comprehensive action recognition system. This fusion of heterogeneous data sources allows the system to maintain high recognition accuracy whether the subject is within or outside the camera's field of view.
Solution Approach 2:
The patent introduces an intermediary processing layer that generates pseudo-labels from sensor data when visual data is unavailable or insufficient. This intermediary mechanism bridges the gap between on-screen and off-screen recognition, enabling continuous action monitoring regardless of camera orientation or subject position.
2Device complexity
If the system uses only on-screen camera data for action recognition, then the system complexity is low, but the system cannot recognize actions when the subject is outside the camera's view
Solution Approach 1:
The patent segments the action recognition system into distinct modules: on-screen camera processing, off-screen sensor processing, and data fusion. This segmentation allows each module to operate independently with optimized complexity, while the overall system achieves high reliability through their coordinated operation.
Solution Approach 2:
The patent adds a temporal dimension to action recognition by continuously processing sequential data from multiple sources over time. This multi-dimensional approach (spatial from camera + sensor + audio, and temporal from continuous processing) enables reliable recognition even when any single source temporarily fails or is unavailable.
3Measurement precision
If the system processes both on-screen and off-screen data simultaneously, then action recognition performance improves, but the computational load and processing time increase
Solution Approach 1:
The patent implements partial processing by selectively applying different processing intensities based on situation. When the subject is clearly within camera view, the system relies primarily on efficient visual processing. When the subject moves outside view, it activates sensor-based pseudo-label generation. This partial action approach maintains high performance while avoiding unnecessary computational overhead in all scenarios.
Solution Approach 2:
The system employs periodic action by alternating between different data sources and processing modes based on real-time conditions. Rather than continuously processing all data sources at maximum intensity, the system dynamically switches between camera-primary and sensor-primary modes, reducing overall computational load while maintaining recognition accuracy.
Data Source
AI summary
Disclosed are an artificial intelligence learning method and an operating method of a robot using the same. An on-screen label is generated based on image data acquired through a camera, an off-screen label is generated based on data acquired through other sensors, and the on-screen label and the off-screen label are used in learning for action recognition, thereby raising action recognition performance and recognizing a user's action even in a situation in which the user deviates from a camera's view.


