First-Person Action Recognition for Accurate Virtual Terminal Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual terminal devices for extended reality technologies face limitations in application scope and low identification accuracy due to third-person perspective video capture, which hinders effective body movement identification and control.

Innovation Solution

Capture video from a first-person perspective to identify a target center point and associated body key points, utilizing three-dimensional coordinates and landmark detection algorithms to accurately model human body movements for improved control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If body movement identification is performed on a video captured from a third-person perspective, then the application scope is limited, but the identification accuracy is low

Engineering Contradiction:
Improveidentification accuracyVSAvoidapplication scope
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent inverts the conventional third-person perspective video capture approach by adopting a first-person perspective. The video capture device is worn by the user on the head, capturing videos from the user's own perspective rather than observing the user from outside. This inversion enables accurate identification of body movements by directly capturing the user's field of view and head movements, thereby improving identification accuracy while expanding application scope to include scenarios like augmented reality and autonomous driving assistance.

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If a third-person perspective video capture is used, then the device complexity is reduced, but the control accuracy is low

Engineering Contradiction:
Improvecontrol accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing system that includes video capture devices worn by the user, video transmission modules, and central processing units. This intermediary system bridges the gap between simple video capture and complex movement analysis by automatically processing the first-person perspective videos to extract body movement information, thereby improving control accuracy without requiring the end user to manually manage the complexity of the system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4668234A1Action recognition method and apparatus, and electronic device, medium and computer program product
Publication Date: 2025.12.24 BEIJING ZITIAO NETWORK TECH CO LTD
  • EP4668234A1 patent drawingFigure 1~2
  • EP4668234A1 patent drawingFigure 3~4
  • EP4668234A1 patent drawingFigure 5~6

AI summary

Provided in the embodiments of the present disclosure are an action recognition method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: in response to an action recognition request, which is triggered for a virtual terminal device, acquiring a target video, which is collected from a first-person perspective, wherein the first-person perspective is the perspective of an operator of the virtual terminal device; according to image frames in the target video, recognizing a target central point of the operator and limb key points associated with the target central point; and according to the target central point and the limb key points, determining a limb action which is conducted by the operator. In the embodiments of the present disclosure, image collection is performed from the first-person perspective, and limb recognition is then performed, so as to quickly and accurately control the virtual terminal device, thereby solving the problem of the action recognition precision of the virtual terminal device being low.