Virtual Human Model Dimension Adaptation for AR Action Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing human action recognition technologies in augmented reality applications face challenges due to inaccuracies in recognizing hand actions because of differences between the shape and dimension of hands in training data and three-dimensional models, leading to incorrect key point determination.

Innovation Solution

Establish a virtual human object model with dimensions matching the real hand, using machine learning to align spatial points and orientations accurately, and iteratively refine the action model to improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional action models are used for human action recognition, then the system can operate with existing models, but the recognition accuracy deteriorates due to dimension mismatches between training data and three-dimensional models

Engineering Contradiction:
Improveaction recognition accuracyVSAvoidkey point determination accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transforms the action model from using fixed three-dimensional model dimensions to dynamically adapting to the dimension information of real hands in input images. This parameter change allows the model to match the actual dimensions of hands in augmented reality scenes, resolving the dimension mismatch problem that caused inaccurate key point determination and improving both recognition accuracy and reliability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a dynamic dimension adaptation mechanism where the action model adjusts its dimensional parameters based on the detected hand size in the input image. This dynamic adjustment enables the model to adapt to different hand sizes and distances, maintaining high recognition accuracy across varying conditions rather than relying on static three-dimensional model dimensions

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If the action model uses fixed three-dimensional model dimensions, then the model structure remains simple, but the adaptability to different hand sizes and distances deteriorates

Engineering Contradiction:
Improveadaptability to different hand dimensionsVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary dimension detection of the hand in the input image before conducting action recognition. By obtaining the dimension information of the real hand first, the system can then adjust the action model's dimensional parameters to match, enabling adaptability to different hand sizes and distances without requiring a completely complex model structure

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the dimensional parameters of the action model based on the detected hand dimensions in the input image. This parameter adaptation allows the model to handle various hand sizes and distances effectively, improving versatility while maintaining relatively simple model architecture through targeted parameter adjustments rather than structural complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4575935A1Method and apparatus for recognizing human body action, and device and medium
Publication Date: 2025.06.25 BEIJING ZITIAO NETWORK TECH CO LTD
  • EP4575935A1 patent drawingFigure 1~2
  • EP4575935A1 patent drawingFigure 3~4
  • EP4575935A1 patent drawingFigure 5~6

AI summary

Methods, apparatuses, devices, and medium for recognizing human action are provided. In a method, in response to receiving an input image including a human object, a target action associated with the human object in the input image is determined based on an action model. The action model describes an association relationship between an image including a human object and an action of the human object in the image. The action is represented by a set of spatial points in a virtual human object corresponding to the human object, and a dimension of the virtual human object matches a dimension of the human object. The input image is updated with the virtual human object performing the target action. With the example embodiment of the disclosure, the accuracy of the action model may be improved, such that the action model may recognize the human action in a more accurate and effective manner.