A Method and System for Testing Gesture Recognition Performance Based on Extended Reality Terminals

By combining a bionic robotic arm and a reinforcement learning model with a high-speed camera, we have achieved performance testing of gesture recognition on extended reality terminals, solving the problems of repeatability and accuracy in gesture recognition testing in existing technologies, and providing a standardized evaluation system.

CN122131907APending Publication Date: 2026-06-02CHINA ACADEMY OF INFORMATION & COMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ACADEMY OF INFORMATION & COMM
Filing Date
2026-01-20
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The lack of standardized gesture recognition performance testing systems and methods in existing technologies leads to insufficient repeatability, consistency and accuracy of gesture recognition tests, which affects the performance evaluation and optimization of gesture recognition systems in extended reality terminals.

Method used

By using a bionic robotic arm to simulate hand gestures, combined with a reinforcement learning model and a high-speed camera, and by detecting hand gestures through an extended reality terminal, the system simultaneously analyzes the hand gesture execution results and recognition results to achieve fully automated hand gesture recognition performance testing.

Benefits of technology

It improves the objectivity and repeatability of gesture recognition testing, provides a standardized performance evaluation system, and ensures the accuracy and reliability of the gesture recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131907A_ABST
    Figure CN122131907A_ABST
Patent Text Reader

Abstract

This invention proposes a method and system for testing gesture recognition performance based on an extended reality terminal, relating to the field of human-computer interaction technology. The method includes: issuing gesture control commands to a bionic robotic arm; the bionic robotic arm simulating gestures using a reinforcement learning model and recording the gesture execution results according to the gesture control commands; detecting the gestures using an extended reality terminal and determining the gesture recognition results; capturing an image displayed on the extended reality terminal; and simultaneously analyzing the gesture execution results corresponding to the bionic robotic arm in the captured image and the gesture recognition results corresponding to the image displayed on the extended reality terminal to obtain gesture recognition performance test results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and more particularly to a method and system for testing gesture recognition performance based on an extended reality terminal. Background Technology

[0002] Compared to interaction methods like controllers, which require holding positioning devices with both hands, gesture interaction is a more natural way to interact and frees up the hands, making it a promising direction for extended reality (XR) interaction. With the rapid development of XR technology, gesture recognition, as an important human-computer interaction method, has been widely applied in virtual reality, augmented reality, and other fields. Gesture recognition technology captures and analyzes users' hand movements to achieve natural and intuitive interactive control, providing users with a more immersive experience. However, in terms of gesture recognition performance testing, existing technologies lack standardized testing systems and methods. Using real people for gesture action testing makes it difficult to guarantee the repeatability, consistency, and objectivity of the tests, resulting in insufficient accuracy and reliability of the test results. Secondly, in terms of gesture recognition algorithms, existing technologies mainly use PID control or simple imitation learning methods, which suffer from stiff joint movements and poor timing coordination, making it difficult to achieve human-like control of finger movements and affecting the accuracy and naturalness of gesture recognition. These technical problems limit the performance evaluation and optimization of extended reality terminal gesture recognition systems, hindering the promotion and development of this technology in practical applications.

[0003] In summary, there is an urgent need for a technical solution that can overcome the above-mentioned shortcomings and objectively and accurately test the performance of gesture recognition. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention proposes a method and system for testing gesture recognition performance based on an extended reality terminal. This invention enables standardized testing of gesture interaction performance, improves the anthropomorphism of finger movements and precise multi-joint collaborative control, thereby achieving objective and accurate testing of gesture recognition performance.

[0005] In a first aspect of the present invention, a method for testing gesture recognition performance based on an extended reality terminal is proposed, the method comprising: Send gesture control commands to the bionic robotic arm; The bionic robotic arm simulates hand gestures and records the execution results through a reinforcement learning model based on hand gesture control commands. The hand gesture recognition result is determined by detecting hand gestures using an augmented reality terminal. Capture the image displayed on the extended reality terminal; By simultaneously analyzing the gesture execution results of the bionic robotic arm in the captured images and the gesture recognition results of the images displayed on the extended reality terminal, the gesture recognition performance test results are obtained.

[0006] In a second aspect of the present invention, a gesture recognition performance testing system based on an extended reality terminal is proposed, the system comprising: The test control module is used to issue gesture control commands to the bionic robotic arm; A bionic robotic arm is used to simulate hand gestures and record the execution results through a reinforcement learning model based on hand gesture control commands. An augmented reality terminal is used to detect hand gestures and determine the hand gesture recognition results. A high-speed camera is used to capture images displayed on the extended reality terminal. The performance testing module is used to simultaneously analyze the gesture execution results of the bionic robotic arm in the captured images and the gesture recognition results of the images displayed on the extended reality terminal, so as to obtain the gesture recognition performance test results.

[0007] In a third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a gesture recognition performance testing method based on an extended reality terminal.

[0008] In a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements a gesture recognition performance testing method based on an extended reality terminal.

[0009] In a fifth aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements a gesture recognition performance testing method based on an extended reality terminal.

[0010] The present invention proposes a gesture recognition performance testing method and system based on extended reality terminals. By simulating highly human-like gestures with a bionic robotic arm and simultaneously collecting recognition feedback from the extended reality terminal using a high-speed camera, the method achieves fully automated and highly consistent gesture recognition performance testing. The overall solution utilizes a reinforcement learning model to make the robotic arm's movements natural and coordinated, overcoming the problems of poor repeatability and strong subjectivity in manual testing. At the same time, it establishes an objective and quantitative gesture recognition evaluation system, providing reliable and efficient technical support for the research and development and standardization of gesture interaction in extended reality terminals. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of a gesture recognition performance testing method based on an extended reality terminal according to an embodiment of the present invention.

[0013] Figure 2 This is a schematic diagram of the training method of a reinforcement learning model according to an embodiment of the present invention.

[0014] Figure 3 This is a schematic diagram of the gesture recognition performance testing system architecture based on an extended reality terminal according to an embodiment of the present invention.

[0015] Figure 4 This is a schematic diagram of the system architecture of a specific embodiment of the present invention.

[0016] Figure 5 This is a schematic diagram of the data flow for a gesture interaction performance test according to a specific embodiment of the present invention.

[0017] Figure 6 This is a schematic diagram of a hand model according to a specific embodiment of the present invention.

[0018] Figure 7 This is a schematic diagram of the training process for a simulated humanoid hand based on the SAC and DTW algorithms according to an embodiment of the present invention.

[0019] Figure 8 This is a schematic diagram of a computer device structure according to an embodiment of the present invention. Detailed Implementation

[0020] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0021] Those skilled in the art will recognize that embodiments of the present invention can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0022] According to an embodiment of the present invention, a method and system for testing gesture recognition performance based on an extended reality terminal are proposed, relating to the field of human-computer interaction technology.

[0023] The principles and spirit of the present invention will be explained in detail below with reference to several representative embodiments.

[0024] Figure 1 This is a schematic flowchart of a gesture recognition performance testing method based on an extended reality terminal according to an embodiment of the present invention. Figure 1 As shown, the method includes: S101 sends gesture control commands to the bionic robotic arm; S102, the bionic robotic arm simulates hand gestures and records the execution results through a reinforcement learning model according to the hand gesture control command; S103, detects hand gestures using an extended reality terminal and determines the hand gesture recognition result; S104, Capture the image displayed by the extended reality terminal; S105 simultaneously analyzes the gesture execution results corresponding to the bionic robotic arm in the captured image and the gesture recognition results corresponding to the image displayed on the extended reality terminal to obtain the gesture recognition performance test results.

[0025] This invention uses commands to control a bionic robotic arm to simulate hand gestures, and utilizes an extended reality terminal for recognition and image capture, simultaneously analyzing the execution and recognition results to achieve fully automated testing from action input to recognition feedback. This method significantly improves the objectivity and repeatability of the test, solves the consistency problem that is difficult to guarantee with manual testing, and provides a standardized performance evaluation system for gesture recognition systems.

[0026] To provide a clearer explanation of the above-mentioned gesture recognition performance testing method based on extended reality terminals, each step will be explained in detail below.

[0027] In one embodiment, for S101, a gesture control command is issued to the bionic robotic arm.

[0028] Set the preset frequency and preset number of times; The system sends hand gesture control commands to the bionic robotic arm according to the preset frequency and preset number of times.

[0029] By setting preset frequencies and repetitions to control the issuance of gesture commands, the testing process is made procedural and configurable. This control method can simulate gesture interaction scenarios with different rhythms and repetitions, making the testing closer to real-world usage. It also enhances the repeatability and comparability of results, providing a standardized control method for performance evaluation of gesture recognition systems under different stress conditions.

[0030] In practical applications, a preset frequency and number of repetitions are first set. The testing tool then sends gesture control commands to the bionic robotic arm according to these preset frequencies and repetitions. Specifically, the frequency is set to 120Hz, with each type of gesture repeated 20 times, covering 10 typical gestures such as pinching, extending, grasping, thumbs-up, and OK. The control system of the PC-based testing tool sends control commands containing information such as gesture type, execution timestamp, and motion parameters to the motion controller of the bionic robotic arm.

[0031] In one embodiment, for S102, the bionic robotic arm simulates hand gestures and records the execution results through a reinforcement learning model according to the hand gesture control command.

[0032] The bionic robotic arm simulates the gesture control commands based on a reinforcement learning model. The reinforcement learning model employs a combination of dynamic time warping and reinforcement learning algorithms, and undergoes human-like training through a hierarchical reinforcement learning training method.

[0033] This invention introduces a human-like training model based on dynamic time warping and hierarchical reinforcement learning, enabling the bionic robotic arm to simulate more natural and coordinated hand gestures. Through optimization of the reinforcement learning model, the hand gestures more closely resemble the movement characteristics of real people in both time and space, thereby improving the realism and effectiveness of the test and ensuring the accuracy of performance evaluation of the gesture recognition system when faced with human-like gestures.

[0034] For details, please refer to Figure 2 The training method for the reinforcement learning model is as follows: S201 collects real-person gesture trajectories, constructs a reward function based on dynamic time warping algorithm and reinforcement learning algorithm, and trains single-joint control and multi-joint coordination in stages; S202 calculates the similarity between the image trajectory of a hand gesture and the trajectory of a real person's hand gesture using a dynamic time warping algorithm, and then optimizes the recognition accuracy by combining it with a reinforcement learning algorithm.

[0035] The specific training method for the reinforcement learning model was clarified. By training single-joint and multi-joint collaboration in stages and constructing a reward function in conjunction with the DTW algorithm, the learning of gestures from simple to complex was achieved. This method not only improves the realism of the motion imitation but also enhances training efficiency and model generalization ability, providing an efficient and human-like motion generation mechanism for bionic hands.

[0036] In practical applications, after receiving a gesture control command, the bionic robotic arm simulates the corresponding gesture action based on a reinforcement learning model. This reinforcement learning model employs a combination of dynamic time warping and reinforcement learning algorithms, using a hierarchical reinforcement learning training method for human-like training.

[0037] The training process of the reinforcement learning model includes: first, collecting the trajectory of real human hand gestures; constructing a reward function based on dynamic time warping and reinforcement learning algorithms; and training single-joint control and multi-joint coordination in stages. The similarity between the image trajectory of the hand gesture and the trajectory of the real human hand gesture is calculated using dynamic time warping, and the recognition accuracy is optimized by combining reinforcement learning algorithms.

[0038] Specifically, a hand model with 24 skeletal key points was used, and trajectory data of real people performing 10 typical gestures was collected using 8 infrared cameras. Each gesture was repeated 20 times, accumulating more than 5,000 valid trajectories. The training process was divided into two stages: Stage 1 was single-joint trajectory tracking, mastering fine control of a single joint and achieving an RMSE of less than 0.5mm at the end of a single joint; Stage 2 was multi-joint collaborative control, achieving multi-joint synchronization of complex gestures and achieving a multi-joint collaborative RMSE of less than 2mm.

[0039] During the execution of hand gestures, the bionic robotic arm records the execution results in real time, including parameters such as the action completion timestamp, joint angle changes, and end-effector coordinates, and reports the results to the statistical module of the testing tool.

[0040] In one embodiment, the bionic robotic arm is the key execution component of this invention, employing a combination structure of a 6-axis robotic arm and a 16-DOF highly simulated five-fingered dexterous hand. The robotic arm has an arm span of 0.638m, and its joint range of motion closely approximates the motion characteristics of a human arm. The robotic arm integrates a controllable motion platform, enabling the arm to execute 6DoF motion trajectories and achieve complex movements such as arm extension. The dexterous hand adopts a lightweight design, highly integrating brushless servo motors, reducers, encoders, and drive controllers. Each finger is equipped with a fingertip force sensor, enabling fine gesture movements such as finger pinching and finger spreading. Based on received gesture control commands, the bionic robotic arm simulates gesture movements and records the execution results through a reinforcement learning model. This reinforcement learning model adopts a hierarchical deep reinforcement learning architecture based on the SAC algorithm, comprising three components: a policy network, a double-Q evaluation network, and a reward function. The policy network uses a 4-layer neural network structure. The input layer receives 72-dimensional hand state data, which is processed through three hidden layers of 256 neurons each, and the output layer generates 48-dimensional motion distribution parameters. The dual-Q evaluation network employs the same network architecture to estimate action value. The reward function comprehensively considers DTW (Dynamic Time Warping) trajectory similarity, smoothness constraints, convergence, and contact force control, with weight ratios of 0.6, 0.2, 0.2, and 0.2, respectively. Through this reinforcement learning control method, the bionic robotic arm can achieve trajectory tracking accuracy at the 1mm level, with an end-effector RMSE of less than 2mm, ensuring highly human-like hand gestures.

[0041] In one embodiment, for S103, the gesture action is detected by the extended reality terminal, and the gesture action recognition result is determined.

[0042] The extended reality terminal detects the hand gestures performed by the bionic robotic arm through its built-in gesture recognition function, runs a gesture recognition algorithm to analyze the gesture features, and determines the gesture recognition result. The recognition result includes information such as gesture type, recognition success flag, and recognition completion timestamp, and is sent to the test tool's result statistics module.

[0043] In one embodiment, for S104, an image displayed by the extended reality terminal is captured.

[0044] A high-speed camera integrated inside the head model of the extended reality terminal is used to capture images presented by the extended reality terminal at a preset sampling frame rate, thereby obtaining images including the gestures.

[0045] By using a high-speed camera built into the head model to capture images from the extended reality terminal at a fixed frame rate, the synchronization and timeliness of image acquisition are ensured. High frame rate shooting can capture subtle changes in the image before and after gesture recognition, providing a high-quality, time-aligned image data foundation for subsequent calculations of recognition latency and success rate, thus enhancing the accuracy and reliability of test results.

[0046] In practical applications, a high-speed camera integrated into the head model of the extended reality terminal captures images displayed on the terminal at a preset sampling frame rate, obtaining images including hand gestures. The high-speed camera is a customized 2.3-megapixel global exposure camera from Lingyu Technology, with a sampling frame rate set to 120Hz. Power and data transfer are provided to the system camera via a USB 3.0 cable. The PC-based testing tool uses the high-speed camera SDK to capture photos at a frequency of 120Hz, and then sends the captured image data to the gesture recognition algorithm library for processing.

[0047] In one embodiment, for S105, the gesture execution results corresponding to the bionic robotic arm in the captured image and the gesture action recognition results corresponding to the image displayed by the extended reality terminal are analyzed simultaneously to obtain the gesture recognition performance test results.

[0048] The gesture recognition performance test results shall include at least the gesture recognition success rate and recognition latency; Among them, the gesture recognition success rate is calculated as the ratio of the number of times the gesture action is correctly recognized to the number of times the bionic robotic arm records the gesture execution result. Specifically, for the recognition delay, the difference between the timestamp of the extended reality terminal determining the gesture recognition result and the timestamp of the bionic robotic arm simulating the gesture is calculated as the recognition delay.

[0049] This invention establishes performance metrics for gesture recognition (gesture recognition success rate and recognition latency), enabling a quantitative evaluation of gesture recognition performance. By comparing the timestamps of the robotic arm's execution and the terminal's recognition, the system's recognition latency can be accurately measured; by counting the number of correct recognitions, the recognition success rate can be calculated. This provides clear and quantifiable output metrics for performance testing, supporting objective and comprehensive performance analysis and comparison of gesture recognition systems.

[0050] In practical application scenarios, the test tool's statistics module simultaneously analyzes the gesture execution results corresponding to the bionic robotic arm in the captured images and the gesture recognition results corresponding to the images displayed on the extended reality terminal, thereby obtaining the gesture recognition performance test results.

[0051] The performance test results for gesture recognition should include at least two key indicators: gesture recognition success rate and recognition latency.

[0052] The gesture recognition success rate is calculated as follows: the ratio of the number of correctly recognized gestures to the number of times the bionic robotic arm records the gesture execution results. The specific formula is: Gesture recognition success rate = Number of correctly recognized gestures / Number of gesture executions by the bionic robotic arm.

[0053] For calculating the recognition delay: the difference between the timestamp of the extended reality terminal determining the gesture recognition result and the timestamp of the bionic robotic arm simulating the gesture is calculated as the recognition delay. The accuracy of the timestamp is ensured through the synchronization function between the robotic arm and the high-speed camera. The recognition delay is statistically analyzed in milliseconds.

[0054] The above method enables an objective and accurate evaluation of the gesture recognition performance of extended reality terminals, providing reliable test data support for the optimization of gesture interaction technology. This method effectively improves trajectory tracking accuracy, effectively simulates real-person gestures, and ensures the accuracy and reliability of test results.

[0055] It should be noted that although the operation of the method of the present invention has been described in a specific order in the above embodiments and figures, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0056] After introducing the method of exemplary embodiments of the present invention, the following references are made. Figure 3 This paper introduces a gesture recognition performance testing system based on an extended reality terminal, according to an exemplary embodiment of the present invention.

[0057] The implementation of the gesture recognition performance testing system based on an extended reality terminal can refer to the implementation of the above method, and repeated details will not be elaborated further. The terms "module" or "unit" used below can refer to a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0058] Based on the same inventive concept, this invention also proposes a gesture recognition performance testing system based on an extended reality terminal, such as... Figure 3 As shown, the system includes: a test control module, a bionic robotic arm, an extended reality terminal, a high-speed camera, and a performance testing module; the test control module and the performance testing module can be deployed on a computer.

[0059] The test control module is used to issue gesture control commands to the bionic robotic arm.

[0060] In one embodiment, the test control module serves as the command center of the entire system, responsible for issuing gesture control commands to the bionic robotic arm. This module is implemented using a PC-based testing tool and features a unified gesture recognition interface. It can send various pre-defined gesture commands to the robotic arm at a set frequency and number of times, including 10 typical gesture actions such as pinching, extending, grasping, thumbs-up, and OK gestures. The test control module connects to the robotic arm's motion controller via a USB interface to achieve real-time command transmission and status monitoring.

[0061] A bionic robotic arm is used to simulate hand gestures and record the execution results through a reinforcement learning model based on hand gesture control commands.

[0062] In one embodiment, a bionic robotic arm is the key execution component, employing a combination structure of a 6-axis robotic arm and a 16-DOF highly simulated five-fingered dexterous hand. The robotic arm has an arm span of 0.638m, and its joint range of motion closely approximates the motion characteristics of a human arm. The robotic arm integrates a controllable motion platform, enabling the arm to execute 6DoF motion trajectories and achieve complex movements such as arm extension. The dexterous hand features a lightweight design, highly integrating brushless servo motors, reducers, encoders, and drive controllers. Each finger is equipped with a fingertip force sensor, enabling fine gestures such as finger pinching and finger spreading. Based on received gesture control commands, the bionic robotic arm simulates gesture movements and records the execution results through a reinforcement learning model. This reinforcement learning model adopts a hierarchical deep reinforcement learning architecture based on the SAC algorithm, comprising a policy network, a double-Q evaluation network, and a reward function. The policy network uses a 4-layer neural network structure; the input layer receives 72-dimensional hand state data, which is processed through three hidden layers of 256 neurons each, and the output layer generates 48-dimensional action distribution parameters. The double-Q evaluation network uses the same network architecture to estimate the action value. The reward function comprehensively considers trajectory similarity, smoothness constraints, convergence, and contact force control of DTW (Dynamic Time Warping) algorithm, with weight ratios of 0.6, 0.2, 0.2, and 0.2, respectively. Through this reinforcement learning control method, the bionic robotic arm can achieve trajectory tracking accuracy at the 1mm level, with an end-effector RMSE of less than 2mm, ensuring highly human-like hand gestures.

[0063] An extended reality terminal is used to detect hand gestures and determine the recognition results.

[0064] In one embodiment, the extended reality terminal, as the device under test, has gesture recognition capabilities, enabling it to detect gestures performed by the bionic robotic arm and determine the corresponding gesture recognition result. When the extended reality terminal detects a gesture, it executes a corresponding response action, causing a change in the displayed screen. For example, when a pinch gesture is recognized, the terminal's display screen switches from an unrecognized state to a successfully recognized state, resulting in a significant change in the screen content.

[0065] A high-speed camera is used to capture images displayed on the extended reality terminal.

[0066] In one embodiment, a high-speed camera, serving as the system's image acquisition device, is integrated into an artificial head model. The head model's size closely approximates a real human head, enabling it to be worn and held in place by various XR devices. The high-speed camera utilizes a custom-designed global exposure camera from Lingyu Technology, boasting 2.3 megapixels and a sampling frame rate of 120Hz. Power and data transmission are provided via a USB 3.0 cable. The high-speed camera is specifically designed to capture images displayed on the extended reality terminal, continuously acquiring photos at a frequency of 120Hz to capture real-time changes in the terminal's screen.

[0067] The performance testing module is used to simultaneously analyze the gesture execution results of the bionic robotic arm in the captured images and the gesture recognition results of the images displayed on the extended reality terminal, so as to obtain the gesture recognition performance test results.

[0068] In one embodiment, the performance testing module is responsible for the data analysis and result statistics of the entire system. This module uses a high-speed camera SDK to acquire photo data at a frequency of 120Hz and sends the images to a gesture recognition algorithm library for processing. Once the gesture recognition algorithm library recognizes a gesture, it notifies the test tool's result statistics module. Simultaneously, the performance testing module receives gesture execution result data reported by the robotic arm. By synchronously analyzing the gesture execution results corresponding to the bionic robotic arm in the captured images and the gesture recognition results corresponding to the images displayed on the extended reality terminal, the performance testing module can accurately calculate the gesture recognition success rate and recognition latency, ultimately obtaining complete gesture recognition performance test results.

[0069] The entire system's workflow is as follows: the test control module sends gesture commands to the bionic robotic arm, which then executes the corresponding gestures using a reinforcement learning model. The extended reality terminal detects the gestures and generates a response, while a high-speed camera simultaneously captures changes in the terminal's image. The performance testing module comprehensively analyzes all data and outputs the test results. The robotic arm and the high-speed camera have a synchronization function to ensure the consistency of data acquisition timing, thereby enabling objective and accurate testing of the extended reality terminal's gesture recognition performance.

[0070] It should be noted that although several modules of the gesture recognition performance testing system based on extended reality terminals are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0071] The following describes the gesture recognition performance testing method and system based on an extended reality terminal according to the present invention with reference to a specific embodiment.

[0072] This invention uses a bionic robotic arm to simulate human hand gestures. After a hand gesture is detected by a test device (extended reality terminal) with gesture recognition capabilities, the device performs the corresponding action, changing the displayed image. The hand gesture performance is then tested by capturing the image displayed by the test device using a high-speed camera.

[0073] In terms of hardware, it mainly includes two robotic arms, two dexterous hands, and a head model (with a built-in high-speed camera). Equipped with a motion controller for the robotic arms, it sends commands to control arm movement and gestures. The robotic arms are 6-axis with an arm span of 0.638m, and their arm span and range of motion are very close to those of a human arm. The robotic arms integrate a controllable motion platform, making the arm's movement trajectory closely resemble that of a human arm.

[0074] We have custom-designed and developed a lightweight, 16-DOF (DoF) highly realistic five-finger dexterity hand, with a highly anthropomorphic design that closely approximates the dexterity of a human hand. This dexterity hand integrates a brushless servo motor, reducer, encoder, drive controller, and fingertip force sensor, making it easy to operate.

[0075] Employing a bionic hand and a highly realistic 3D hand mold, integrated with an automated control system, it enables various hand gestures such as pinching, extending the arm, and spreading the fingers. The bionic hand's gestures closely resemble those of a real person, thus expanding the capabilities of real-world devices for recognition. The highly realistic living surface of the hand is achieved using materials such as silicone and resin, resulting in a visually striking resemblance to a human hand.

[0076] The head model is nearly the same size as a human head, supporting the wearing and clamping of various XR devices. It has a built-in high-speed camera that captures images at over 120 frames per second. The high-speed camera captures images from the VR glasses, and image algorithms are used to measure gesture recognition accuracy and latency.

[0077] The artificial head model features a built-in high-speed camera, custom-selected, with global exposure and 2.3 megapixels. Key technical specifications include: sampling frame rate of 120Hz; 2.3 megapixels; and a built-in high-speed camera USB 3.0 data cable for power supply and data transfer.

[0078] In terms of software design, the approach adopted is to adapt the same test APK to different XR headsets, and the testing tool has a unified gesture recognition interface. The camera detects changes in the image before and after a gesture is made.

[0079] refer to Figure 4 This is a schematic diagram of the system architecture of a specific embodiment of the present invention. Figure 4 As shown, the PC is connected to the robotic arm's control system. The testing tool can control the hand to perform 6DoF motion trajectories and pinching movements. Pre-set gestures such as pinching are supported. Through feedback data from the robotic arm's control system, the testing tool automatically recognizes the hand gestures of the robotic handle. The robotic arm and high-speed camera have synchronization capabilities.

[0080] The PC-side deployment includes an image detection module, testing tool applications, and a robotic arm and camera synchronization module; external devices include a high-speed camera, the device under test (XR head-mounted display gesture recognition), and a robotic arm.

[0081] The testing tool application acts as the command initiator, sending gesture control commands to the robotic arm to drive it to execute specified gestures. Upon receiving the commands, the robotic arm simulates the target gesture. The robotic arm-camera synchronization module plays a crucial role in timing coordination, synchronizing the robotic arm's movements, the high-speed camera's image acquisition, and the XR headset's gesture recognition process. This ensures consistency in actions and data acquisition across all stages, preventing timing deviations from affecting test results. The high-speed camera captures images containing gestures (corresponding to the images displayed on the XR terminal in the test) and transmits these images to the image detection module on the PC. The device under test (XR headset gesture recognition) recognizes the robotic arm's gestures, and its recognition data interacts with the robotic arm-camera synchronization module, serving as the source of the test's recognition results. The image detection module receives the images captured by the high-speed camera and, in conjunction with the testing tool application, performs data analysis to ultimately obtain the test results for gesture recognition performance (success rate, latency).

[0082] refer to Figure 5 This is a schematic diagram of the data flow for a gesture interaction performance test according to a specific embodiment of the present invention. Figure 5 As shown, the PC-based testing tool uses a high-speed camera SDK to capture photos at a frequency of 120Hz and sends them to the gesture recognition algorithm library. After the gesture recognition algorithm library recognizes the gesture, it notifies the test tool's result statistics module. The test tool then sends gesture commands to the robotic arm according to a set frequency and number of times. After the robotic arm completes the gesture action, it reports the execution result to the test tool's statistics module. The test tool's statistics module calculates the gesture recognition success rate and recognition latency based on the execution result reported by the robotic arm and the recognition result of the gesture recognition algorithm library.

[0083] In one specific embodiment, the mechanism for simulating hand gestures with a bionic robotic arm involves using reinforcement learning based on collected hand movement data from real people to make the bionic hand's movements closely resemble those of a real person. The reinforcement learning model combines Dynamic Time Warping (DTW algorithm, used to calculate the similarity between two time series) and a reinforcement learning algorithm (SAC algorithm, softened Actor-Critic), employing a hierarchical reinforcement learning training method for anthropomorphic training.

[0084] The bionic robotic arm uses a hand model with 24 skeletal key points (such as...). Figure 6 (as shown in the figure), the specific relationships are shown in Table 1.

[0085] Table 1

[0086] refer to Figure 7 This is a schematic diagram of the training process for a simulated humanoid hand based on the SAC and DTW algorithms according to an embodiment of the present invention. Figure 7 As shown, the specific methods include: S701, Data acquisition of human hand gestures. S702, Data preprocessing. S703, Hierarchical deep reinforcement learning training. S704, Deployment to the bionic hand robotic arm.

[0087] The following is a detailed explanation of each step.

[0088] S701, real-person gesture data acquisition.

[0089] Hardware configuration: 8 infrared cameras (resolution 1280×1024, frame rate 100Hz), spatial accuracy 0.1mm, time synchronization error <1ms; Data collection scenario: Real people perform 10 typical gestures (pinch, extend, grasp, thumbs up, OK gesture, etc.), each type is repeated 20 times, and a total of 5000+ valid trajectories are collected.

[0090] S702, Data Preprocessing.

[0091] Normalization: Map the trajectory coordinates to the joint workspace of the bionic hand (with the palm root set as the origin and finger length scaled according to the size of the bionic hand), as shown in the following formula: , where L 真人手 The distance from key point 0 to key point 21 is when the palm is fully open; L 仿生手 The distance from key point 0 to key point 21 is when the palm is fully open; The revised indicator; These are the original indicators.

[0092] Data augmentation: Data is augmented by random translation (±5mm) and rotation (±1°) to simulate minute hand displacements and changes in perspective, thereby improving the model's generalization ability.

[0093] S703, hierarchical deep reinforcement learning training.

[0094] Deep reinforcement learning consists of a policy network, a double-Q evaluation network, and a reward function.

[0095] 1) Policy Network: Refer to Table 2 for the hierarchical relationship of the policy network.

[0096] Table 2

[0097] 2) Double-Q Critic Network: Refer to Table 3 for the hierarchical relationship of the double-Q evaluation network.

[0098] Table 3

[0099] Note: There are two dual-Q evaluation models, both of which are MLP networks with model structure related.

[0100] 3) Reward function: By using DTW to guide trajectory similarity, smoothness to constrain movement naturalness, and convergence to accelerate end-effector localization, combined with staged weight adjustments, the training goal of the bionic hand can be achieved from "single-joint precision" to "multi-joint coordination".

[0101] The formula for the total reward function based on the segmented strategy is as follows:

[0102] in, Let w1, w2, and w3 represent the total reward function; w1, w2, and w3 are 0.6, 0.2, and 0.2 respectively; w4, w5, w6, and w7 are 0.4, 0.2, 0.2, and 0.2 respectively. , , , Let represent the DTW trajectory similarity reward function, smoothness reward function, convergence reward, and contact force reward, respectively.

[0103] 3.1) DTW trajectory similarity reward function ( ): Trajectory definition: Let the trajectory of a real person be... Bionic trajectory (These are all 72-dimensional vector sequences, corresponding to the three-dimensional coordinates of 24 skeletal keypoints); Given m real-person trajectory points, There are n biomimetic trajectory points.

[0104] Distance matrix construction: Calculate the Euclidean distance between trajectory points to form a matrix. ; , These are the i-th trajectory point in the real person's trajectory and the j-th trajectory point in the bionic trajectory, respectively.

[0105] Dynamic Time Warping (DTW): This method uses dynamic programming to find the minimum distance path. The recursive formula is as follows:

[0106] in, This represents the minimum cumulative distance between two sequences; , , These represent the cumulative distance from the previous step, corresponding to three different time-series alignments.

[0107] Normalized reward: Based on the maximum DTW distance MaxDTW in the sample data, it is mapped to the [0,1] interval:

[0108] in, This represents the similarity score between the biomimetic trajectory and the real person's trajectory; It is a biomimetic gesture trajectory T b With real human gesture trajectory T r The cumulative DTW distance between them (the smaller the distance, the more similar the two trajectories are); MaxDTW It is the maximum possible value of the DTW cumulative distance in the scenario (as a normalization benchmark).

[0109] 3.2) Smoothness reward function ( ): Design purpose: To punish sudden changes in joint angles and improve the naturalness of movement; Mathematical expression:

[0110] in, Let be the angle of the i-th joint at time t (there are 24 controllable joints in total); The weight is indicated by, for example, the distal joints (such as the distal phalanx of the thumb, the distal phalanx of the index finger, etc.) are set to 0.8, and the other joints are set to 0.2.

[0111] 3.3) Convergent Rewards ( ): Design objective: To encourage the bionic hand's end-effector to rapidly approach the target position, thereby accelerating training convergence; Mathematical expression:

[0112] in, , is the Euclidean distance between the end key point (such as the fingertip of the index finger) and the target position; the exponential decay form ensures that the closer the distance, the higher the reward (the maximum reward is 1, and the reward approaches 0 when the distance is ≥50mm).

[0113] 3.4) Contact Force Reward ( Enabled in Phase 2:

[0114] in, Real-time readings (in N) from 5 fingertip force sensors; A reasonable contact force of 0.5N is required. <5N, otherwise, the punishment is too lenient or too severe.

[0115] 4) Layered training method: Phase 1: Single-joint trajectory tracking (basic capability building).

[0116] Training goal: To master fine control of a single joint (such as thumb flexion and index finger extension); Trajectory example: Index finger extension trajectory (non-linear convergence, simulating the characteristics of real human movement).

[0117] Reward configuration: DTW weight = 0.6, contact force reward disabled; Training indicators: RMSE at the distal end of a single joint < 0.5 mm, and joint angular velocity change rate < 5° / ms.

[0118] Phase 2: Multi-joint collaborative control (reproduction of complex gestures).

[0119] Training objective: To achieve multi-joint synchronization of complex gestures (such as pinching and grasping); Trajectory example: Pinch gesture (the distance between the thumb and index finger changes dynamically to simulate an interactive scenario).

[0120] Reward configuration: DTW weight = 0.4, enable contact force reward (weight 0.2); Training indicators: Multi-joint coordination RMSE < 2mm, contact force control error < 5%.

[0121] The solution has achieved a trajectory tracking accuracy of 1mm in actual testing (RMSE at the end position < 2mm), and the complete training cycle takes about 72 hours (NVIDIA V100 GPU).

[0122] S704, deployed to a bionic hand robotic arm.

[0123] The trained policy network is deployed to the bionic hand robotic arm to achieve precise control that conforms to the gestures of a self-heating human.

[0124] Based on the aforementioned inventive concept, such as Figure 8 As shown, the present invention also proposes a computer device 800, including a memory 810, a processor 820, and a computer program 830 stored in the memory 810 and executable on the processor 820. When the processor 820 executes the computer program 830, it implements the aforementioned gesture recognition performance testing method based on an extended reality terminal.

[0125] Based on the aforementioned inventive concept, the present invention proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned gesture recognition performance testing method based on an extended reality terminal.

[0126] Based on the aforementioned inventive concept, this invention proposes a computer program product, which includes a computer program that, when executed by a processor, implements a gesture recognition performance testing method based on an extended reality terminal.

[0127] Compared to existing technologies, this invention makes improvements in at least the following aspects: 1. It establishes a complete software and hardware technical methodology system for the gesture performance testing system of a robotic arm-bionic hand. 2. It establishes a reinforcement learning training method for bionic hand motion following; wherein, a multi-dimensional reward system is established: integrating DTW trajectory similarity, motion smoothness, end-effector convergence, and contact force constraints to achieve a unity of "human-likeness and physical rationality." Hierarchical reinforcement learning: a two-stage training strategy, from single-joint fine control to multi-joint coordination, improves training efficiency by 40%.

[0128] The present invention proposes a gesture recognition performance testing method and system based on extended reality terminals. By simulating highly human-like gestures with a bionic robotic arm and simultaneously collecting recognition feedback from the extended reality terminal using a high-speed camera, the method achieves fully automated and highly consistent gesture recognition performance testing. The overall solution utilizes a reinforcement learning model to make the robotic arm's movements natural and coordinated, overcoming the problems of poor repeatability and strong subjectivity in manual testing. At the same time, it establishes an objective and quantitative gesture recognition evaluation system, providing reliable and efficient technical support for the research and development and standardization of gesture interaction in extended reality terminals.

[0129] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.

[0130] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This invention is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for testing gesture recognition performance based on an extended reality terminal, characterized in that, The method includes: Send gesture control commands to the bionic robotic arm; The bionic robotic arm simulates hand gestures and records the execution results through a reinforcement learning model based on hand gesture control commands. The hand gesture recognition result is determined by detecting hand gestures using an augmented reality terminal. Capture the image displayed on the extended reality terminal; By simultaneously analyzing the gesture execution results of the bionic robotic arm in the captured images and the gesture recognition results of the images displayed on the extended reality terminal, the gesture recognition performance test results are obtained.

2. The gesture recognition performance testing method based on an extended reality terminal according to claim 1, characterized in that, Sending gesture control commands to the bionic robotic arm, including: Set the preset frequency and preset number of times; The system sends hand gesture control commands to the bionic robotic arm according to the preset frequency and preset number of times.

3. The gesture recognition performance testing method based on an extended reality terminal according to claim 1, characterized in that, The bionic robotic arm simulates hand gestures and records the execution results based on hand gesture control commands using a reinforcement learning model, including: The bionic robotic arm simulates the gesture control commands based on a reinforcement learning model. The reinforcement learning model employs a combination of dynamic time warping and reinforcement learning algorithms, and undergoes human-like training through a hierarchical reinforcement learning training method.

4. The gesture recognition performance testing method based on an extended reality terminal according to claim 3, characterized in that, The training methods for the reinforcement learning model include: Collect real-person gesture trajectories, construct a reward function based on dynamic time warping algorithm and reinforcement learning algorithm, and train single-joint control and multi-joint coordination in stages; The similarity between the image trajectory of a hand gesture and the trajectory of a real person's hand gesture is calculated by using a dynamic time warping algorithm, and the recognition accuracy is optimized by combining it with a reinforcement learning algorithm.

5. The gesture recognition performance testing method based on an extended reality terminal according to claim 1, characterized in that, Capturing images displayed on the extended reality terminal includes: A high-speed camera integrated inside the head model of the extended reality terminal is used to capture images presented by the extended reality terminal at a preset sampling frame rate, thereby obtaining images including the gestures.

6. The method for testing the gesture recognition performance based on an extended reality terminal according to claim 1, characterized in that, The gesture recognition performance test results are obtained by simultaneously analyzing the gesture execution results corresponding to the bionic robotic arm in the captured images and the gesture recognition results corresponding to the images displayed on the extended reality terminal, including: The gesture recognition performance test results shall include at least the gesture recognition success rate and recognition latency; Among them, the gesture recognition success rate is calculated as the ratio of the number of times the gesture action is correctly recognized to the number of times the bionic robotic arm records the gesture execution result. Specifically, for the recognition delay, the difference between the timestamp of the extended reality terminal determining the gesture recognition result and the timestamp of the bionic robotic arm simulating the gesture is calculated as the recognition delay.

7. A gesture recognition performance testing system based on an extended reality terminal, characterized in that, The system includes: The test control module is used to issue gesture control commands to the bionic robotic arm; A bionic robotic arm is used to simulate hand gestures and record the execution results through a reinforcement learning model based on hand gesture control commands. An augmented reality terminal is used to detect hand gestures and determine the hand gesture recognition results. A high-speed camera is used to capture images displayed on the extended reality terminal. The performance testing module is used to simultaneously analyze the gesture execution results of the bionic robotic arm in the captured images and the gesture recognition results of the images displayed on the extended reality terminal, so as to obtain the gesture recognition performance test results.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.