Motion Recognition via Spatial-Temporal Feature Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing motion recognition technologies face challenges in accurately recognizing continuous motions, such as hand signals of a traffic officer, due to overlap and occlusion between key points, leading to degraded recognition accuracy and high computational requirements.
Innovation Solution
A method and system that utilize both spatial and temporal features from image data and key point data to recognize motions, employing reshaping of time-series data, extraction of spatial and temporal features, and integration of these features for efficient motion recognition on low-power edge devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If skeleton-based motion recognition using key points is employed, then motion recognition can be performed, but recognition accuracy degrades when key points overlap or occur occlusion
Solution Approach 1:
The patent segments the motion recognition task into two independent pathways: a slow pathway for spatial feature extraction and a fast pathway for temporal feature extraction. This segmentation allows each pathway to specialize in specific aspects, with the slow pathway handling spatial relationships and the fast pathway handling temporal changes, thereby maintaining accuracy even when key points are occluded or overlapping.
Solution Approach 2:
The patent introduces an intermediary mechanism that combines outputs from both slow and fast pathways through a fusion layer. This intermediary fusion process integrates spatial and temporal features, allowing the system to compensate for inaccuracies in one pathway using information from the other, thus maintaining recognition accuracy under occlusion conditions.
2Measurement precision
If slow fast method with two pathways is used to handle spatial and temporal features, then motion recognition accuracy improves, but computational time and power consumption increase significantly
Solution Approach 1:
The patent applies partial action by implementing the slow pathway only when spatial feature extraction is needed and the fast pathway only when temporal feature extraction is needed. Rather than always running both pathways, the system selectively activates the appropriate pathway based on the specific motion recognition task requirements, thereby reducing overall computational load and power consumption while maintaining accuracy.
Solution Approach 2:
The patent implements dynamic resource allocation by adjusting the activation state of different pathways based on real-time needs. The system dynamically switches between slow and fast pathways or combines them in different proportions depending on the motion recognition task, allowing flexible adaptation to varying computational requirements and power constraints.
3Measurement precision
If both slow and fast pathways are implemented to extract spatial and temporal features, then motion recognition performance improves, but device complexity increases
Solution Approach 1:
The patent merges the slow and fast pathways into a unified motion recognition framework where both pathways share common components such as the fusion layer and subsequent motion recognition modules. This merging approach reduces overall system complexity by avoiding complete duplication of processing stages while still benefiting from the complementary strengths of both pathways.
Solution Approach 2:
The patent creates a universal motion recognition system where both slow and fast pathways can handle various motion recognition tasks. The same fusion layer and subsequent processing modules can work with outputs from either pathway or both pathways combined, making the system multi-functional and adaptable to different scenarios without requiring separate specialized systems for each task type.
Data Source
AI summary
There is provided a deep learning-based motion recognition method and system using multiple feature information. A motion recognition method according to an embodiment reshapes time-series image data obtained by shooting a target object to a type of image data of a spatial domain, extracts spatial features from the reshaped image data, reshapes the image data from which the spatial features are extracted to a type of time-series image data, integrates the time-series image data and time-series key point data of the target object, extracts temporal features from the integrated time-series data, and recognizes motions of the target object based on the extracted temporal features. Accordingly, motions can be more stably recognized even when there are a plurality of objects at the same time and an overlap, occlusion frequently occur, and lots of computations are not required and motion recognition can be performed in a small low-power edge device having relatively low computing power.


