First-Person Action Recognition for Accurate Virtual Terminal Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual terminal devices for extended reality technologies face limitations in application scope and low identification accuracy due to third-person perspective video capture, which hinders effective body movement identification and control.
Innovation Solution
Capture video from a first-person perspective to identify a target center point and associated body key points, utilizing three-dimensional coordinates and landmark detection algorithms to accurately model human body movements for improved control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If body movement identification is performed on a video captured from a third-person perspective, then the application scope is limited, but the identification accuracy is low
Solution Approach 1:
The patent inverts the conventional third-person perspective video capture approach by adopting a first-person perspective. The video capture device is worn by the user on the head, capturing videos from the user's own perspective rather than observing the user from outside. This inversion enables accurate identification of body movements by directly capturing the user's field of view and head movements, thereby improving identification accuracy while expanding application scope to include scenarios like augmented reality and autonomous driving assistance.
2Measurement precision
If a third-person perspective video capture is used, then the device complexity is reduced, but the control accuracy is low
Solution Approach 1:
The patent introduces an intermediary processing system that includes video capture devices worn by the user, video transmission modules, and central processing units. This intermediary system bridges the gap between simple video capture and complex movement analysis by automatically processing the first-person perspective videos to extract body movement information, thereby improving control accuracy without requiring the end user to manually manage the complexity of the system.
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
Provided in the embodiments of the present disclosure are an action recognition method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: in response to an action recognition request, which is triggered for a virtual terminal device, acquiring a target video, which is collected from a first-person perspective, wherein the first-person perspective is the perspective of an operator of the virtual terminal device; according to image frames in the target video, recognizing a target central point of the operator and limb key points associated with the target central point; and according to the target central point and the limb key points, determining a limb action which is conducted by the operator. In the embodiments of the present disclosure, image collection is performed from the first-person perspective, and limb recognition is then performed, so as to quickly and accurately control the virtual terminal device, thereby solving the problem of the action recognition precision of the virtual terminal device being low.