Central venous catheter medical operation training method and system based on virtual reality and gesture tracking

By collecting users' hand movement data in real time, combining layered fusion gesture recognition algorithms and virtual reality visual scene information, multimodal haptic feedback is generated and quantitatively evaluated through an evaluation engine. This solves the problems of inaccurate gesture recognition, insufficient haptic feedback, and subjective evaluation methods in existing technologies, achieving high-precision gesture recognition and multimodal haptic feedback, enhancing the immersion and realism of operation, and improving the scientific nature and efficiency of training.

CN122116718APending Publication Date: 2026-05-29SHUNDE HOSPITAL SOUTHERN MEDICAL UNIV (THE FIRST PEOPLES HOSPITAL OF SHUNDE FOSHAN)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHUNDE HOSPITAL SOUTHERN MEDICAL UNIV (THE FIRST PEOPLES HOSPITAL OF SHUNDE FOSHAN)
Filing Date
2026-04-07
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing virtual reality technology has problems in training for central venous catheterization procedures, such as inaccurate gesture recognition, insufficient multimodal tactile feedback, subjective assessment methods, and lack of team collaboration simulation, resulting in poor training effectiveness.

Method used

By collecting users' hand movement data in real time, combining layered fusion gesture recognition algorithms and virtual reality visual scene information, multimodal haptic feedback is generated, and quantitative evaluation is performed through an evaluation engine to achieve high-precision gesture recognition, enhanced immersion, and objective evaluation.

Benefits of technology

It achieves high-precision gesture recognition and multimodal haptic feedback, enhancing the immersion and realism of the operation, and provides multi-dimensional automated evaluation, improving the scientific nature and efficiency of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116718A_ABST
    Figure CN122116718A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical education, in particular to a central venous catheter medical operation training method and system based on virtual reality and gesture tracking, which comprises the following steps: collecting user hand action data in real time, identifying standard central venous catheter operation actions performed by a user based on the hand action data through a hierarchical fusion gesture recognition algorithm, analyzing the current operation intention of the user in combination with virtual reality visual scene information, generating and outputting corresponding multi-modal tactile feedback signals according to the identified operation actions and the current operation intention, quantitatively evaluating the operation of the user through an evaluation engine based on the hand action data, the operation action recognition result and the current operation intention, and generating a comprehensive evaluation report. Through high-precision gesture tracking, intelligent intention understanding, multi-modal tactile feedback and automatic quantitative evaluation, the application realizes high immersion of the training process and objective and accurate evaluation of operation skills, and effectively improves the authenticity, safety and teaching efficiency of medical operation training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of medical education, and in particular to a training method and system for central venous catheterization based on virtual reality and gesture tracking. Background Technology

[0002] Central venous catheterization is a critical clinical procedure, and traditional apprenticeship training suffers from high risks and limited resources. While virtual reality technology offers a new approach to skills simulation training, existing systems still have significant shortcomings in practical applications. First, at the gesture recognition and interaction level, existing solutions (such as data gloves or ordinary visual recognition) are limited by bulky equipment, susceptibility to occlusion or lighting conditions, and difficulty in achieving high-precision, robust hand movement tracking. More importantly, they lack the ability to deeply understand the continuous and complex clinical procedures such as "disinfection-puncture-catheter insertion," failing to accurately interpret the trainee's operational intentions, resulting in insufficient intelligence in human-computer interaction.

[0003] Secondly, in terms of immersion and realism, most existing systems severely lack realistic multimodal tactile feedback. Trainees cannot feel the real-time resistance changes when puncturing different tissues, the "loose feeling" of blood vessels, or other key mechanical feedback. This "seeing but not touching" experience makes it difficult for trainees to form correct muscle memory and tactile feel, resulting in a significant gap in skill transfer from virtual training to real clinical operation, greatly weakening the effectiveness of training.

[0004] Finally, regarding skills assessment and training models, existing systems mostly employ crude and subjective assessment methods, relying heavily on outcome judgments or instructor experience. They lack the capacity for automated, data-driven quantitative analysis of process quality aspects such as operational fluency, technical precision, and aseptic technique adherence. Furthermore, training models are often single-person, single-device based, failing to simulate the crucial multi-role team collaboration scenarios in real clinical settings, resulting in a disconnect between training content and actual workflows. Therefore, the shortcomings of existing technologies in terms of immersion, intelligent assessment, and team training hinder the in-depth application of virtual reality training in advanced clinical skills education. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a training method and system for central venous catheterization medical procedures based on virtual reality and gesture tracking.

[0006] The above-mentioned objective of this application is achieved through the following technical solution:

[0007] A training method for central venous catheterization based on virtual reality and gesture tracking, comprising the following steps:

[0008] Real-time acquisition of user hand movement data, including three-dimensional coordinates of key hand points, joint angles, and movement trajectory information;

[0009] Based on the hand movement data, a hierarchical fusion gesture recognition algorithm is used to identify the standard central venous catheterization operation performed by the user, and the user's current operation intention is analyzed in combination with virtual reality visual scene information;

[0010] Based on the identified operational actions and current operational intentions, corresponding multimodal tactile feedback signals are generated and output to simulate resistance, vibration, and temperature changes in real-world operations.

[0011] Based on the hand motion data, operation action recognition results, and current operation intention, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report.

[0012] By adopting the above technical solution, real-time acquisition of user hand movement data provides a basis for quantitative analysis of the entire system. Then, through layered fusion of gesture recognition algorithms, clinically significant standardized operational actions are accurately identified from the original movements, and their intentions are analyzed, solving the problem that traditional videos or simple VR observations cannot achieve semantic understanding of actual operations. Next, multimodal tactile feedback is generated based on the recognition results, transforming visual and intentional information into perceptible force, vibration, and temperature signals, greatly enhancing the immersion and realism of the operation and compensating for the lack of tactile feedback in traditional VR training. Finally, quantitative evaluation is performed based on the entire process data through an assessment engine, achieving objective and multi-dimensional automated assessment of the operator's skill level. This overcomes the problems of inconsistent standards and low efficiency caused by traditional reliance on subjective evaluation by instructors. This method organically combines data collection, intelligent recognition, immersive feedback, and objective evaluation, forming a complete and quantifiable high-level clinical skills simulation training solution.

[0013] In a preferred embodiment, this application can be further configured such that the real-time acquisition of user hand movement data specifically includes:

[0014] Color and depth images of the hand were acquired using a binocular depth camera, and the two-dimensional pixel coordinates of 21 key points of the hand were extracted based on an image recognition algorithm.

[0015] Infrared reflection data of the hand surface is acquired by an infrared sensor array, and the three-dimensional spatial coordinates of key points of the hand are calculated by combining the depth image.

[0016] The angular velocity, acceleration, and orientation data of the hand joints are collected by an inertial measurement unit worn on the hand.

[0017] The two-dimensional pixel coordinates, three-dimensional spatial coordinates, and inertial measurement data are spatiotemporally aligned and fused to generate a complete hand motion data sequence containing timestamps.

[0018] By employing the aforementioned technical solution, specifically through the collaborative operation of a binocular depth camera, an infrared sensor, and an inertial measurement unit (IMU), multimodal data acquisition and fusion of hand movements were achieved. Binocular vision and infrared sensing provided high-precision spatial coordinates, while the IMU effectively compensated for data loss during rapid movement or under occlusion. Spatiotemporal alignment and fusion of data from different sources generated a complete and stable sequence of hand movement data, including timestamps. This technique effectively overcomes the susceptibility of single-vision sensors to illumination and occlusion, as well as the cumulative drift issues inherent in pure inertial sensors. It provides high-precision and robust underlying data support for subsequent gesture recognition and motion analysis, ensuring the accuracy and reliability of the entire system's input information.

[0019] In a preferred embodiment, this application can be further configured such that: the step of identifying the standard central venous catheterization procedure performed by the user based on the hand movement data using a hierarchical fusion gesture recognition algorithm specifically includes:

[0020] Based on the coordinates of key skeletal points in the hand motion data, the angle changes and relative positional relationships of each finger joint are calculated through the underlying skeletal tracking layer.

[0021] The hand motion data of consecutive frames is constructed into a spatiotemporal graph structure, and the temporal evolution characteristics of the hand motion sequence are analyzed by spatiotemporal graph convolutional network.

[0022] By combining the knowledge base of central venous catheterization procedures, a high-level intent understanding layer maps action features into specific clinical operational intents.

[0023] Based on the matching degree between the operation intention and the action sequence, the operation parameters are identified, and the completion status of the operation steps is marked.

[0024] By employing the aforementioned technical solution, a hierarchical recognition strategy combining low-level skeletal tracking, mid-level spatiotemporal graph convolutional network analysis, and high-level intent understanding is adopted. Low-level tracking extracts geometric features such as joint angles from coordinates; the mid-level spatiotemporal graph convolutional network, by analyzing spatiotemporal graphs constructed from consecutive frames, effectively captures the temporal evolution patterns of dynamic operations such as punctures and catheter insertions, accurately identifying standard operational action segments; high-level intent understanding combines virtual scene context and a clinical knowledge base to map action sequences to specific clinical operation stages and goals. This hierarchical fusion approach can not only recognize static gestures but also understand complex, continuous operational procedures with clear clinical objectives, greatly improving the semantic depth and accuracy of action recognition. This allows the system to truly "understand" what the trainee is doing, rather than simply "seeing" their hands moving.

[0025] In a preferred embodiment, this application can be further configured as follows: The step of generating and outputting corresponding multimodal tactile feedback signals based on the identified operational actions and current operational intentions to simulate resistance, vibration, and temperature changes in real-world operations specifically includes:

[0026] Based on the operation actions and current operation intentions, and combined with the virtual soft tissue mechanical model, the magnitude and direction of the real-time force generated when the virtual needle tip interacts with the virtual blood vessels and tissues are calculated.

[0027] The calculated real-time force is mapped into a controllable current signal, a vibration frequency signal, and a temperature control signal, and the controllable current signal, vibration frequency signal, and temperature control signal are sent to the corresponding tactile feedback device.

[0028] The device generates and outputs corresponding multimodal tactile feedback signals based on the tactile feedback device.

[0029] By adopting the above technical solution, a closed loop from virtual interaction to physical perception is achieved. The system first calculates the real-time force between the virtual needle tip and the tissue based on the identified operation and intention, combined with a virtual soft tissue mechanics model that simulates the physical properties of different tissues (skin, blood vessel walls). Then, the calculated force vector is intelligently mapped into control signals that drive specific hardware devices, including current signals to control the electromagnetic brake to generate resistance, frequency signals to drive the vibrator to simulate the sensation of breakthrough, and signals to adjust the temperature of the semiconductor cooling chip. This multimodal feedback mechanism, based on a physical model and closely linked to the operational context, can realistically reproduce key tactile sensations such as changes in puncture resistance, the feeling of blood vessel loss, and the cooling sensation of disinfection, significantly improving the immersion and realism of the training and helping trainees develop correct muscle memory and tactile feedback.

[0030] In a preferred embodiment, this application can be further configured as follows: based on the hand motion data, the operation action recognition result, and the current operation intention, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report, specifically including:

[0031] Based on the hand movement data, the hand tremor frequency, movement speed, and acceleration changes are calculated to evaluate the operation fluency dimension score;

[0032] The identified operational actions are compared with the standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the standardization of aseptic operation, and an operation standardization score is generated.

[0033] The virtual needle tip trajectory data is quantified based on the current operational intent, and the puncture angle deviation, puncture point position error and catheter insertion length accuracy are calculated to generate a technical accuracy score.

[0034] Based on the operation fluency score, operation standardization score, and technical accuracy score, the data are input into a pre-trained evaluation and analysis model to generate a comprehensive evaluation report.

[0035] By adopting the above technical solution, this application establishes a multi-dimensional, data-driven intelligent quantitative evaluation system. This evaluation engine not only quantitatively analyzes changes in hand tremor frequency, movement speed, and acceleration based on raw hand movement data to assess the fluency and stability of the operation, but also automatically checks the completeness, correctness of sequence, and aseptic operation standardization by comparing the identified standard operation sequence with a standard procedure database, generating an objective operation standardization score. Furthermore, by combining the parsed operational intent, it quantitatively analyzes the movement trajectory of the virtual needle tip, accurately calculating the puncture angle deviation, position error, and catheter insertion length accuracy, generating a technical accuracy score. Finally, these multi-dimensional scores are input into a pre-trained evaluation analysis model to generate a comprehensive and objective evaluation report. This method abandons the limitations of traditional methods that rely on subjective observation by instructors and single-result evaluation, achieving comprehensive, fine-grained data deconstruction and automated evaluation of the operation process, providing trainees with accurate and quantifiable skill feedback, and significantly improving the scientific nature of training and evaluation efficiency.

[0036] In a preferred embodiment, this application can be further configured as follows: comparing the identified operational actions with a standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the aseptic operation compliance, and generating an operation compliance score, specifically including:

[0037] Based on the operation actions, identify whether the key steps of disinfection, draping, puncture point selection and catheter fixation are performed in the preset order, and calculate the number of times the steps are missing or the order is incorrect;

[0038] During the virtual ultrasound-guided puncture phase, the duration of use and scanning range of the virtual ultrasound probe during the operation are analyzed to determine whether they meet the standard requirements for covering the target blood vessel area.

[0039] During the guidewire and catheter insertion phase, based on the complication judgment markers recorded in the operation action record, the number of virtual complication alarms triggered by improper operation is counted.

[0040] An operational standardization score is generated based on the number of missing or incorrect steps, the number of virtual complication alarms, and the judgment results.

[0041] By employing the aforementioned technical solution, key operational actions such as disinfection, draping, puncture, and fixation are automatically compared with the preset standard procedure sequence, accurately calculating the number of missing steps and incorrect sequences. During ultrasound-guided puncture, the usage time of the virtual probe and the coverage of the scanning trajectory are analyzed to objectively determine whether the operation conforms to the system's scanning specifications. During catheter insertion, the number of virtual complication alarms automatically triggered by improper operation is monitored, allowing for a reverse assessment of the operational compliance risk. Finally, a compliance score is generated by combining all the above quantitative indicators. This assessment mechanism transforms abstract clinical operational standards (such as aseptic principles and standard procedures) into a series of detectable, statistically significant, and deductible quantitative indicators, making the assessment of the soft requirement of "compliance" rigorous, objective, and traceable. This effectively guides trainees to establish a strong awareness of standardized operation from the initial stages of simulation training, overcoming the shortcomings of traditional assessments that struggle to quantify such requirements.

[0042] In a preferred embodiment, this application can be further configured as follows: quantifying the virtual needle tip trajectory data according to the current operational intent, calculating the puncture angle deviation, puncture point position error, and catheter insertion length accuracy, and generating a technical accuracy score specifically includes:

[0043] Based on the current operational intent, a puncture angle sequence is obtained, and the average and maximum deviations between the actual needle insertion angle and the preset standard needle insertion angle are calculated based on the puncture angle sequence.

[0044] According to the current operational intent, a puncture depth sequence is obtained. Based on the puncture depth sequence, the depth value when the needle tip touches the target blood vessel wall is identified, and the absolute error between the needle tip and the preset standard anatomical depth value is calculated.

[0045] The movement trajectory of the virtual medical device is obtained according to the current operation intention. Based on the movement trajectory of the virtual medical device, the path of the needle tip before reaching the target blood vessel is analyzed, and the maximum distance of its deviation from the preset ideal puncture path is calculated.

[0046] Based on the above calculation results, a technical accuracy score is generated.

[0047] By employing the aforementioned technical solution, after identifying the intention to perform procedures such as "puncture and needle insertion," the system tracks and records the angle sequence, depth sequence, and spatial trajectory of the virtual needle tip in real time. The accuracy and stability of angle control are assessed by calculating the average and maximum deviations of the needle insertion angle; the precision of depth control is evaluated by comparing the measured depth of the needle tip entering the blood vessel with the standard anatomical depth; and the straightness and stability of the operational path are assessed by analyzing the maximum distance the needle tip trajectory deviates from the ideal puncture path. Finally, a technical accuracy score is generated through weighted calculation of these multi-dimensional error data. This refined analysis method based on full-process trajectory data accurately reveals the trainee's true skill level in key technical aspects such as puncture angle, depth control, and hand stability, providing tiered, diagnostic feedback far exceeding a binary "success / failure" judgment, thus helping trainees to improve their skills in a targeted and efficient manner.

[0048] Secondly, the above-mentioned inventive objective of this application is achieved through the following technical solutions:

[0049] A central venous catheterization medical operation training system based on virtual reality and gesture tracking, the central venous catheterization medical operation training system based on virtual reality and gesture tracking includes:

[0050] The multimodal operation data acquisition module is used to collect user hand movement data in real time. The hand movement data includes the three-dimensional coordinates of key hand points, joint angles, and movement trajectory information.

[0051] The operation intent recognition module is used to identify the standard operation actions of central venous catheterization performed by the user based on the hand movement data through a hierarchical fusion gesture recognition algorithm, and to analyze the user's current operation intent in combination with virtual reality visual scene information;

[0052] The haptic feedback generation module is used to generate and output corresponding multimodal haptic feedback signals based on the identified operation actions and current operation intentions, simulating resistance, vibration and temperature changes in real operation;

[0053] The operation evaluation module is used to quantitatively evaluate the user's operation based on the hand movement data, operation action recognition results, and current operation intention through the evaluation engine, and generate a comprehensive evaluation report.

[0054] By adopting the above technical solution,

[0055] Thirdly, the above-mentioned objectives of this application are achieved through the following technical solutions:

[0056] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described training method for central venous catheterization based on virtual reality and gesture tracking.

[0057] Fourthly, the above-mentioned objectives of this application are achieved through the following technical solutions:

[0058] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described training method for central venous catheterization based on virtual reality and gesture tracking.

[0059] In summary, this application includes at least one of the following beneficial technical effects:

[0060] 1. By collecting real-time user hand movement data, a foundation for quantitative analysis is provided for the entire system. Then, through a layered fusion gesture recognition algorithm, clinically significant standardized operational actions are accurately identified from the raw movements, and their intentions are analyzed. This solves the problem that traditional videos or simple VR observations cannot achieve semantic understanding of actual operations. Next, multimodal tactile feedback is generated based on the recognition results, transforming visual and intentional information into perceptible force, vibration, and temperature signals, greatly enhancing the immersion and realism of the operation and compensating for the lack of tactile feedback in traditional VR training. Finally, based on the full-process data, a quantitative evaluation engine is used to achieve objective and multi-dimensional automated assessment of the operator's skill level. This overcomes the problems of inconsistent standards and low efficiency caused by traditional reliance on subjective evaluation by instructors. This method organically combines data collection, intelligent recognition, immersive feedback, and objective evaluation to form a complete and quantifiable high-level clinical skills simulation training solution.

[0061] 2. A hierarchical recognition strategy combining low-level skeletal tracking, mid-level spatiotemporal graph convolutional network (ST-GCN) analysis, and high-level intent understanding is employed. Low-level tracking extracts geometric features such as joint angles from coordinates; the mid-level ST-GCN, by analyzing spatiotemporal graphs constructed from consecutive frames, effectively captures the temporal evolution patterns of dynamic operations such as puncture and catheter insertion, accurately identifying standard operational action segments; high-level intent understanding combines virtual scene context and a clinical knowledge base to map action sequences to specific clinical operation stages and objectives. This hierarchical fusion approach not only recognizes static gestures but also understands complex, continuous operational procedures with clear clinical purposes, greatly improving the semantic depth and accuracy of action recognition. This allows the system to truly "understand" what the trainee is doing, rather than simply "seeing" their hands moving.

[0062] 3. A closed loop from virtual interaction to physical perception has been achieved. The system first calculates the real-time force between the virtual needle tip and the tissue based on the identified operation and intention, combined with a virtual soft tissue biomechanical model that simulates the physical properties of different tissues (skin, blood vessel walls). Then, the calculated force vector is intelligently mapped into control signals that drive specific hardware devices, including current signals to control the electromagnetic brake to generate resistance, frequency signals to drive the vibrator to simulate the sensation of breakthrough, and signals to adjust the temperature of the semiconductor cooling chip. This multimodal feedback mechanism, based on a physical model and closely linked to the operational context, can realistically reproduce key tactile sensations such as changes in puncture resistance, the feeling of loss of blood flow, and the cooling sensation of disinfection, significantly improving the immersion and realism of the training and helping trainees develop correct muscle memory and tactile feedback.

[0063] 4. Based on raw hand movement data, quantitative analysis of hand tremor frequency, movement speed, and acceleration changes is performed to assess the smoothness and stability of the operation. Furthermore, by comparing the identified standard operation sequence with a standard procedure database, the completeness, sequential correctness, and aseptic technique compliance of the steps are automatically checked, generating an objective operational compliance score. Further, by combining the analyzed operational intent, the trajectory of the virtual needle tip is quantitatively analyzed to accurately calculate puncture angle deviation, positional error, and catheter insertion length accuracy, generating a technical accuracy score. Finally, these multi-dimensional scores are input into a pre-trained evaluation and analysis model to generate a comprehensive and objective evaluation report. This method overcomes the limitations of traditional methods that rely on subjective observation by instructors and single-result evaluation, achieving comprehensive and fine-grained data deconstruction and automated evaluation of the operation process. It provides trainees with accurate and quantifiable skill feedback, significantly improving the scientific nature of training and evaluation efficiency. Attached Figure Description

[0064] Figure 1 This is a flowchart of a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application;

[0065] Figure 2 This is a flowchart illustrating the implementation of step S10 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0066] Figure 3 This is a flowchart illustrating the implementation of step S20 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0067] Figure 4 This is a flowchart illustrating the implementation of step S30 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0068] Figure 5 This is a flowchart illustrating the implementation of step S40 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0069] Figure 6 This is a flowchart illustrating the implementation of step S42 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0070] Figure 7 This is a flowchart illustrating the implementation of step S43 in a central venous catheterization medical operation training method based on virtual reality and gesture tracking in one embodiment of this application.

[0071] Figure 8 This is a principle block diagram of a central venous catheterization medical operation training system based on virtual reality and gesture tracking in one embodiment of this application;

[0072] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0073] The present application will be further described in detail below with reference to the accompanying drawings.

[0074] In one embodiment, such as Figure 1 As shown, this application discloses a training method for central venous catheterization based on virtual reality and gesture tracking, which specifically includes the following steps:

[0075] S10: Real-time acquisition of user hand movement data, including three-dimensional coordinates of key hand points, joint angles, and movement trajectory information.

[0076] In this embodiment, the purpose of this step is to provide high-precision, multimodal raw input data for the entire virtual operation training system. Hand movement data is the core information reflecting the trainee's operation skills and stability. It includes not only static hand posture (described by key point coordinates and joint angles) but also dynamic movement process (described by movement trajectory). High-precision data acquisition is the basis for subsequent accurate action recognition, realistic feedback, and objective evaluation.

[0077] Specifically, a multi-sensor fusion data acquisition array was deployed. Trainees wear VR headsets to enter a virtual operating room environment, with their hands positioned within a capture area comprised of a binocular depth camera and an infrared sensor array. First, the binocular depth camera simultaneously acquires color and depth images of the hands. Using image recognition algorithms, it extracts the two-dimensional pixel coordinates of 21 key points on both hands (such as fingertips, knuckles, and palm base points) in real time. Simultaneously, the infrared sensor emits invisible light and receives reflected signals from the hand surface. Combined with the depth information provided by the depth camera, and using triangulation principles, the two-dimensional pixel coordinates are upgraded to millimeter-level precision three-dimensional spatial coordinates. To further capture hand posture during rapid movements and when obscured, trainees also wear lightweight inertial measurement units on their wrists or the backs of their hands. These units incorporate gyroscopes and accelerometers, directly acquiring angular velocity, acceleration, and direction (Euler angles or quaternions) data for the hand and its major joints. The central processing unit receives these three data streams and ensures that all data has a unified time reference through hardware synchronization signals and software timestamp alignment algorithms. Finally, data fusion algorithms (such as Kalman filtering or complementary filtering) fuse visual coordinates with inertial data to generate a complete sequence of hand motion data frames containing timestamps, 3D coordinates of 21 key points, angles of each finger joint, and smooth motion trajectories calculated from consecutive frames.

[0078] S20: Based on the hand movement data, the user's standard central venous catheterization operation is identified by a hierarchical fusion gesture recognition algorithm, and the user's current operation intention is analyzed in conjunction with virtual reality visual scene information.

[0079] In this embodiment, the aim is to transform raw, low-level hand movement data into high-level operational semantics with clinical significance. Layered fusion means that the recognition process is not completed in one step, but rather proceeds from basic skeletal motion analysis to continuous action pattern recognition, and finally, understanding the operational purpose in conjunction with the scene context. This solves the problem that simple gesture recognition cannot understand complex, continuous operational procedures with clear clinical purposes.

[0080] Specifically, the recognition and analysis process involves three collaborative layers. First, at the bottom-level skeletal tracking layer, the system receives hand motion data generated in step S10. Based on the 3D coordinates of 21 key points, it calculates in real time the angles between finger joints, the relative distances between fingers, and the orientation vector of the palm, forming a precise description of the hand's static posture and simple dynamics (such as clenching a fist, opening a fist). Next, at the middle-level motion analysis layer, the system constructs a spatiotemporal graph structure from multiple consecutive frames (e.g., 30 frames, corresponding to 1 second) of hand key point data. Nodes represent key points, and edges represent the spatial connections and temporal continuity between key points. This spatiotemporal graph is input into a spatiotemporal graph convolutional network model pre-trained using a large number of central venous catheterization procedure videos. The model can learn and recognize standard operational motion segments such as "holding an ultrasound probe," "fan-shaped scanning," "holding a puncture needle," "inserting a needle at a specific angle," "inserting a guidewire," and "fixing a dressing," and assigns a motion label and confidence level to each segment. Finally, at the high-level intent understanding layer, the system combines the identified current action label sequence with real-time rendered virtual reality visual scene information (such as whether the virtual instrument has been picked up, which layer of virtual tissue model the virtual needle tip is in contact with, and whether the target blood vessel is displayed on the ultrasound image). Through an intent reasoning model based on a clinical operation standard knowledge base (e.g., disinfection must be performed before puncture, and the blood vessel must be confirmed by ultrasound before needle insertion), the system comprehensively determines which stage of the entire central venous catheterization process the trainee is currently in (e.g., "preoperative preparation stage," "puncture positioning stage," "catheter insertion stage"), and predicts the trainee's next reasonable operational goal. For example, when the system recognizes the "holding the puncture needle" action and the needle tip is above the disinfected virtual skin model, the resolved current operational intent is "preparing to perform venous puncture."

[0081] S30: Based on the identified operation action and current operation intention, generate and output corresponding multimodal tactile feedback signals to simulate resistance, vibration and temperature changes in real operation.

[0082] In this embodiment, the aim is to enhance the immersion and realism of virtual training by converting visual recognition results into physical perception through tactile feedback. Multimodal feedback overcomes the limitations of single vibration feedback by simulating changes in resistance (force), vibration (tactile), and temperature (thermal), allowing trainees to experience key mechanical events and physiological feedback during operation, such as changes in resistance when penetrating different tissues and the feeling of missing when puncturing a blood vessel.

[0083] Specifically, haptic feedback generation is a process of "virtual computation-physical mapping." First, virtual force calculation: Based on the operational action (e.g., "puncture and needle insertion") and intention ("needle tip is penetrating the skin") identified in step S20, the system invokes a virtual soft tissue mechanical model. This model is constructed based on a mass-spring-damped system or the finite element method, defining the elastic modulus, damping coefficient, and rupture threshold of different tissues such as virtual skin, subcutaneous tissue, and blood vessel walls. When a collision and penetration between the virtual needle tip and these tissue models is detected, the model calculates the magnitude and direction of the reaction force on the needle tip in real time. For example, the force is greater when penetrating the skin, decreases when entering subcutaneous fat, and there is a sudden drop in force when piercing the blood vessel wall (simulating a feeling of falling). Next, signal mapping and output: The calculated virtual force vector is mapped to a set of control signals driving the physical devices. 1) Controllable current signal: The magnitude of the force is mapped to the current value sent to the "electromagnetic brake" inside the operating handle. The larger the current, the greater the resistance generated by the brake, simulating the mechanical resistance encountered by the needle tip. 2) Vibration frequency signal: Sudden force characteristics (such as breaking through a blood vessel) or specific events (such as accidentally touching an arterial wall) are mapped to signals of specific frequencies and waveforms, driving the "linear resonant actuator (LRA)" or "piezoelectric vibrator" within the handle to produce instantaneous vibrations. 3) Temperature control signal: Depending on the operation stage, such as simulating disinfection with a sterile swab, the system generates a cooling command, controlling the "thermal cooling chip (TEC)" integrated into the handle grip to briefly cool down, simulating the cooling sensation of alcohol evaporation. Finally, these signals are transmitted via a low-latency wireless protocol to the force feedback handle in the student's hand, driving the corresponding devices to operate, thus allowing the student's hand to experience multimodal haptic feedback synchronized with the virtual operation.

[0084] S40: Based on the hand motion data, operation action recognition results, and current operation intention, the user operation is quantitatively evaluated through the evaluation engine to generate a comprehensive evaluation report.

[0085] In this embodiment, the aim is to provide an objective, comprehensive, and quantitative evaluation of trainees' performance, overcoming the limitations of traditional training that relies on subjective instructor scoring. The evaluation engine comprehensively utilizes process data (hand movements), semantic recognition results (operational actions and intentions), and virtual scene states to construct a scoring system from multiple dimensions.

[0086] Specifically, the evaluation process involves multi-dimensional parallel analysis. First, the fluency score is evaluated: the raw hand movement data collected in step S10 is directly analyzed to calculate the tremor frequency (obtained through frequency domain analysis of accelerometer data) and average movement speed during key operational periods (such as needle holding and insertion). The fluency score is inversely proportional to the tremor amplitude and directly proportional to the stability of the speed. Second, the standardization score is evaluated: the sequence of "standard operating actions" identified in step S20 (such as "disinfection-draping-ultrasound positioning-puncture…") is compared temporally with the "Standard Operating Procedure (SOP)" stored in the database. This checks for omissions and correct sequence. Simultaneously, by analyzing virtual scene information, aseptic operation standards are checked, such as whether the virtual hand touches non-sterile areas and whether the sterile drape coverage is sufficient. Any violation will result in a deduction of points. Next, the accuracy of the technique is assessed: Combining the "current operational intent" analyzed by S20 (e.g., intent to "puncture the internal jugular vein"), the virtual needle tip trajectory data within the time period of this intent is extracted. The deviation between the actual puncture angle and the standard needle insertion angle (e.g., 45°), the error between the puncture depth and the preset depth of the virtual patient's blood vessel model, and the appropriateness of the catheter insertion length are calculated. Smaller deviations and errors result in a higher accuracy score. Finally, a comprehensive evaluation report is generated: The scores from the above three dimensions, along with other indicators extracted from the operational data (e.g., total operation time, number of virtual complication triggers), are input into a pre-trained evaluation analysis model (e.g., a gradient boosting decision tree model). This model outputs a comprehensive score out of 100 and generates a structured report. The report includes not only the total score and scores for each sub-item but also lists specific advantages and errors in the operation (e.g., when the puncture angle was too large), and may include a replay video clip of the operational trajectory for trainees to review.

[0087] In this embodiment, real-time acquisition of user hand movement data provides a foundation for quantitative analysis of the entire system. Then, through a layered fusion gesture recognition algorithm, clinically significant standardized operational actions are accurately identified from the original movements, and their intentions are analyzed. This solves the problem that traditional videos or simple VR observations cannot achieve semantic understanding of actual operations. Next, multimodal tactile feedback is generated based on the recognition results, transforming visual and intentional information into perceptible force, vibration, and temperature signals, greatly enhancing the immersion and realism of the operation and compensating for the lack of tactile feedback in traditional VR training. Finally, quantitative evaluation is performed based on the entire process data through an assessment engine, achieving objective and multi-dimensional automated assessment of the operator's skill level. This overcomes the problems of inconsistent standards and low efficiency caused by traditional reliance on subjective evaluation by instructors. This method organically combines data acquisition, intelligent recognition, immersive feedback, and objective evaluation to form a complete and quantifiable high-level clinical skills simulation training solution.

[0088] In one embodiment, such as Figure 2 As shown, in step S10, the user's hand movement data is collected in real time, specifically including:

[0089] S11: Acquire color and depth images of the hand using a binocular depth camera, and extract the two-dimensional pixel coordinates of 21 key hand points based on an image recognition algorithm.

[0090] In this embodiment, the binocular depth camera simulates the human eye by using two side-by-side cameras and calculates depth information through parallax.

[0091] Specifically, the system calls the binocular camera SDK to simultaneously capture RGB color images and corresponding 16-bit depth images from the left and right viewpoints at a rate of 30 frames per second (FPS). Each pixel value in the depth image represents the distance from that point to the camera. Subsequently, the color images are input into a lightweight hand keypoint detection neural network (such as a cropped MediaPipe Hands model). This network outputs the two-dimensional pixel coordinates (u, v) of 21 keypoints on each hand (including the wrist, knuckles, and fingertips) in the color image coordinate system. Depth information in this step is primarily used for subsequent 3D reconstruction and does not participate in keypoint detection.

[0092] S12: Obtain infrared reflection data of the hand surface through an infrared sensor array, and calculate the three-dimensional spatial coordinates of key points of the hand by combining the depth image.

[0093] In this embodiment, the aim is to elevate two-dimensional pixel coordinates to three-dimensional spatial coordinates, which is crucial for achieving high-precision gesture tracking. Using RGB images alone results in the loss of depth information, while combining them with active infrared sensing can improve robustness and accuracy in low light or complex backgrounds.

[0094] Specifically, the infrared sensor array emits specific infrared structured light or speckle patterns and receives the patterns reflected from the hand's surface. Due to the uneven shape of the hand, the reflected patterns are distorted. The processor compares the emitted and received infrared patterns and combines them with the initial depth map provided by the binocular system in S11. Using triangulation or time-of-flight methods, it performs precise calculations to ultimately determine the three-dimensional spatial coordinates (x, y, z) corresponding to each of the 21 hand keypoints identified in S11, typically in millimeters. These coordinates are relative to the camera coordinate system.

[0095] S13: Collect angular velocity, acceleration, and direction data of the hand joints through an inertial measurement unit worn on the hand.

[0096] In this embodiment, visual tracking is prone to loss or jitter during rapid movements or when the hand is partially obscured (such as when holding a device). The IMU provides high-frequency self-motion data that is unaffected by occlusion.

[0097] Specifically, the IMU device worn by trainees incorporates a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The gyroscope directly outputs the angular velocity (degrees per second) of the hand's rotation around the three axes. The accelerometer outputs linear accelerations along the three axes (including gravitational acceleration). The magnetometer provides directional information relative to the Earth's magnetic north pole. This sensor data is acquired at frequencies up to several hundred Hz, pre-filtered by an embedded processor, and then transmitted to the host computer via low-latency wireless technologies such as Bluetooth 5.0. This data is used to track subtle hand movements such as rapid flips and tremors.

[0098] S14: The two-dimensional pixel coordinates, three-dimensional spatial coordinates and inertial measurement data are spatiotemporally aligned and fused to generate a complete hand motion data sequence containing timestamps.

[0099] In this embodiment, the goal is to unify data from different sensors, frequencies, and coordinate systems into a single, complete, consistent, high-frame-rate data stream describing hand movements. Spatiotemporal alignment is a prerequisite for ensuring data validity, while data fusion leverages the advantages of multiple sensors, compensating for each other's shortcomings.

[0100] Specifically, each frame of visual data (from S11 and S12) and each packet of IMU data (from S13) is timestamped with the clock. Since the IMU frequency is higher than the camera frequency, an interpolation algorithm is used to align the visual data to the higher IMU timeline. Next, coordinate system 1 is performed, transforming the visually calculated 3D coordinates (camera coordinate system) to the virtual world coordinate system with the VR headset as the origin using a pre-calibrated transformation matrix. The IMU data is also transformed to the same coordinate system through a similar calibration. Finally, a sensor fusion algorithm (such as an extended Kalman filter) is applied. This algorithm uses visual data as observations of absolute position and attitude (high accuracy but low frequency, potentially jittery), and IMU data as input for motion prediction (high frequency, good instantaneous dynamics but prone to drift). Through a prediction-correction loop, it outputs a "complete hand motion data sequence" with the same frequency as the IMU, smooth and stable, containing the 3D coordinates of 21 keypoints, the overall hand orientation (calculated using IMU orientation data), and joint angles (calculated from the keypoint coordinates). This sequence forms the basis for all subsequent processing.

[0101] In one embodiment, such as Figure 3 As shown, in step S20, the standard central venous catheterization procedure performed by the user is identified based on the hand movement data using a hierarchical fusion gesture recognition algorithm, specifically including:

[0102] S21: Based on the coordinates of the skeletal key points in the hand motion data, calculate the angle changes and relative positional relationships of each finger joint through the bottom-level skeletal tracking layer.

[0103] In this embodiment, this step is the preprocessing and feature extraction stage of gesture recognition. It is difficult to understand the action directly from coordinate points. Transforming it into geometric relationships such as joint angles and relative positions can better represent the hand's posture and intention.

[0104] Specifically, the complete hand motion data sequence generated in step S14 is read. For each frame of data, 21 key points are traversed, and based on the hand topology, the flexion / extension angles of joints such as the metacarpophalangeal joints, proximal interphalangeal joints, and distal interphalangeal joints are calculated. For example, the joint angle is obtained by calculating the angle between the vectors formed by adjacent key points (such as fingertips, knuckles, and metacarpophalangeal points). At the same time, the Euclidean distance between the tip of the thumb and the tip of the index finger (used to determine the pinching action) or the palm normal vector (used to determine the palm orientation) is calculated. These calculated angles and distances constitute a low-dimensional feature vector describing the hand posture of the current frame.

[0105] S22: Construct a spatiotemporal graph structure from continuous frames of hand motion data, and analyze the temporal evolution characteristics of the hand motion sequence through a spatiotemporal graph convolutional network.

[0106] In this embodiment, this step is the core of recognizing continuous actions. The pose of a single frame is static; only by combining the pose changes (spatiotemporal information) of consecutive frames can dynamic operations such as "puncture" and "guidewire delivery" be recognized. Spatiotemporal graph convolutional networks are effective deep learning models for processing this type of spatiotemporal graph data.

[0107] Specifically, the hand pose features of the most recent T frames (e.g., T=30) are cached. The hand keypoints of these T frames are connected to form a spatiotemporal graph: 21 keypoints in each frame constitute the nodes of the graph; keypoints within the same frame are connected according to the hand skeleton to form spatial edges; and the same keypoint is connected between adjacent frames to form temporal edges. This graph is fed into a pre-trained ST-GCN model. This model automatically learns and extracts deep spatiotemporal features that can represent specific maneuvers (such as "rotating an ultrasound probe" or "smoothly advancing a puncture needle") by alternately performing graph convolution operations in the spatial dimension (between different keypoints within the same frame) and the temporal dimension (between different keypoints in different frames). Finally, the model outputs the label and confidence score of the most likely one or more "standard maneuvers" in the current time period through a fully connected classification layer.

[0108] S23: Combining the knowledge base of central venous catheterization procedures, the action features are mapped to specific clinical operation intentions through a high-level intent understanding layer.

[0109] In this embodiment, after identifying the action of "moving a long, thin object forward," it is necessary to determine whether it is "puncture" or "guidewire insertion" based on the context. High-level intent understanding associates low-level actions with high-level clinical goals through rules or models.

[0110] Specifically, the system maintains a knowledge base of central venous catheterization procedures, which defines the steps, sequence, and conditions of the standard procedure in the form of flowcharts or state machines. For example, the knowledge base stipulates that the "puncture" action must occur after "ultrasound confirmation of the vessel"; the "guidewire delivery" action must be performed after the virtual event of "aspiration of venous blood". The intent understanding layer receives a continuous sequence of action labels from S22, while querying the scene context provided by the virtual reality engine (such as the type of virtual instrument currently held, indicators on the virtual patient's body, and UI prompts). Based on a Bayesian network or recurrent neural network-based intent understanding model, it integrates action history and the current scene to infer the trainee's current macroscopic operational stage (such as "establishing a sterile area") and the specific next operational intent (such as "the next intent is to perform local anesthesia"). For example, even if S22 recognizes a "slowly advancing forward" action, if the system detects that the trainee is not holding any virtual instruments, it may determine the intent as "hand movement and positioning" rather than "puncture".

[0111] In one embodiment, such as Figure 4 As shown, in step S30, based on the identified operation action and current operation intention, a corresponding multimodal tactile feedback signal is generated and output to simulate the resistance, vibration, and temperature changes in real operation, specifically including:

[0112] S31: Based on the operation action and the current operation intention, and combined with the virtual soft tissue mechanical model, calculate the magnitude and direction of the real-time force generated when the virtual needle tip interacts with the virtual blood vessels and tissues.

[0113] In this embodiment, this step is the "virtual physics engine" for haptic feedback. It determines the realism of the feedback. The virtual soft tissue biomechanical model is a mathematical simulation of the physical properties of human tissue, ensuring that the interaction between the virtual device and the tissue conforms to basic mechanical laws.

[0114] Specifically, when S20 detects a "puncture" operation intended to target specific tissue, the system activates a high-precision collision detection module to calculate in real time the penetration depth and contact area between the virtual needle tip (usually represented by a ray or a small sphere) and a multi-layered meshed tissue model of virtual skin, fat, muscle, and blood vessel walls. These tissue models are assigned physical parameters such as mass, stiffness, and damping. The virtual soft tissue mechanical model (such as a simplified version of position-based dynamics (PBD) or finite element method (FEM)) calculates a feedback force in real time based on the penetration depth, contact area, and tissue parameters. This force typically includes: an elastic drag proportional to the penetration depth, a viscous damping force proportional to the penetration velocity, and a sudden force (force instantaneously decreasing) indicating tissue puncture when the penetration depth exceeds the tissue's "strength threshold." The calculated force is a three-dimensional vector containing magnitude and direction (usually the opposite direction of the needle tip's direction).

[0115] S32: Map the calculated real-time force into a controllable current signal, a vibration frequency signal, and a temperature control signal, and send the controllable current signal, vibration frequency signal, and temperature control signal to the corresponding tactile feedback device.

[0116] S33: Generate and output corresponding multimodal tactile feedback signals based on the tactile feedback device.

[0117] In this embodiment, the calculated abstract mechanical vector is converted into understandable electronic control signals that can drive specific hardware actuators (brakes, vibrators, and cooling coils).

[0118] Specifically, for resistance feedback: the calculated force vector magnitude (scalar) is linearly or non-linearly mapped to the duty cycle of a PWM (Pulse Width Modulation) signal or an analog current value (0-20mA). The greater the force, the greater the current or duty cycle, and the greater the resistance generated after being sent to the electromagnetic brake. For vibration feedback: when a sudden change in force is detected (such as tissue puncture) or a specific pattern is reached (such as the tiny vibration of a needle tip touching the blood vessel wall), a pre-designed vibration waveform file (such as an 80Hz sine wave lasting 100ms) is triggered. This waveform file is converted into a digital audio stream and sent to the linear resonant actuator (LRA) via an audio decoding chip or directly through a digital interface. For temperature feedback: depending on the stage of operation, for example, when the system determines that the student is wiping their skin with a virtual disinfectant swab, a "cooling" command is generated, which is converted into a signal controlling the thermoelectric cooler (TEC) driver board to operate for several seconds. All these signals are sent to the force feedback handle in the student's hand via a low-latency wireless transmission module (such as a proprietary 2.4G protocol).

[0119] Furthermore, inside the handle, upon receiving a current signal, the electromagnetic brake's coil generates a magnetic field, which interacts with the permanent magnet, creating a magnetic resistance that hinders the rotation of the handle's motor (or directly hinders manual operation). When pushing the virtual needle tip, the learner needs to apply additional force to overcome this resistance, experiencing the mechanical sensation of "penetrating tissue." Upon receiving a vibration signal, the linear resonant actuator's internal spring and mass system resonates under electromagnetic force, generating vibrations at a specific frequency. This vibration is transmitted to the learner's palm through the handle's outer shell, simulating a "feeling of emptiness" or "accidental touch." Upon receiving a cooling command current, the semiconductor cooling chip's cold end absorbs heat, causing the temperature of the handle grip area, which is in close contact with the cold end, to drop by several degrees Celsius within seconds, providing a noticeable "cool" sensation. These three feedback modes can be triggered individually or in combination as needed, collectively creating a realistic tactile experience.

[0120] In one embodiment, such as Figure 5 As shown, in step S40, based on the hand motion data, operation action recognition results, and current operation intention, the user operation is quantitatively evaluated by the evaluation engine to generate a comprehensive evaluation report, specifically including:

[0121] S41: Calculate the hand tremor frequency, movement speed, and acceleration changes based on the hand movement data, and evaluate the operation fluency dimension score.

[0122] In this embodiment, this step assesses the trainee's basic skills and stability, which is an important indicator for evaluating the quality of operation. Hand tremors are a sign of tension or lack of proficiency, while unstable speed and acceleration indicate clumsy or hesitant operation.

[0123] Specifically, from the hand movement data in step S10, triaxial acceleration data of key wrist points are extracted during the critical operation window (e.g., the period from needle positioning to needle insertion completion). A Fast Fourier Transform (FFT) is performed on this acceleration data to analyze its spectrum, identifying frequency components with significant energy (typically in the 2-12Hz range), whose amplitude is quantified as the "tremor index." Simultaneously, the average hand movement velocity and its standard deviation are calculated during this period. The fluency score is calculated using a predefined formula, for example: Fluency Score = A * (1 / (1 + Tremor Index)) + B * (1 / Velocity Standard Deviation) + C * (Whether the average velocity is within the ideal range). Here, A, B, and C are weighting coefficients obtained through regression analysis of expert scoring data.

[0124] S42: Compare the identified operational actions with the standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the standardization of aseptic operation, and generate an operation standardization score.

[0125] In this embodiment, this step assesses whether the trainee has followed standard clinical operating procedures. In medical procedures, especially aseptic procedures, the completeness and sequence of steps are crucial.

[0126] Specifically, a structured Standard Operating Procedure (SOP) database is built-in, breaking down a standard central venous catheterization into a series of necessary atomic steps (such as "washing hands," "opening the sterile pack," "donning sterile gloves," "disinfecting and draping," ... "puncture," ... "fixing the dressing"). The evaluation engine matches the sequence of "standard operating actions" identified by S20 throughout the entire procedure with the ideal sequence in the SOP database using string matching or state machine comparison. Missing steps and steps with incorrect order (such as puncture before disinfection) are recorded. Simultaneously, by analyzing the virtual scenario logs, "aseptic operation compliance" is checked. For example, after identifying the "donning sterile gloves" action, the system monitors whether the virtual hand model touches non-sterile areas (such as one's own clothing or unsterilized virtual skin) in subsequent steps, recording the number of violations. The compliance score is calculated based on the number of missing steps, incorrect order, and aseptic violations, with a maximum score of 100 points. Each error deducts points according to its weight.

[0127] S43: Quantify the virtual needle tip trajectory data according to the current operational intent, and calculate the puncture angle deviation, puncture point position error and catheter insertion length accuracy to generate a technical accuracy score.

[0128] In this embodiment, this step assesses the trainee's skill level and is central to quantifying the accuracy of their operation. The angle, location, and depth of the puncture are key technical parameters that determine the success of the operation and the likelihood of complications.

[0129] Specifically, when S20 resolves the "puncture and needle insertion" intention, the evaluation engine begins recording the trajectory of the virtual needle tip. Puncture angle deviation: At the moment of needle insertion, the angle between the virtual needle tip axis and the virtual skin surface normal is calculated, and the difference is taken from the standard needle insertion angle specified in the SOP (e.g., 45° is commonly used for internal jugular vein puncture), and the absolute value is recorded. Puncture point position error: The virtual coordinates of the needle tip's first contact with the skin are recorded, and the two-dimensional Euclidean distance between them and the standard target point coordinates determined by the ultrasound simulation module or surface markers is calculated on the skin surface. Catheter insertion length accuracy: At the end of the "catheter insertion" intention phase, the position of the virtual catheter tip is recorded, and its virtual length along the vascular path from the puncture point is calculated. This is compared with the recommended insertion length calculated based on the patient's height and weight to obtain the error value. The accuracy score is a function of the aforementioned angle deviation, position error, and length error; the smaller the error, the higher the score. A piecewise linear or exponential decay function is typically used for mapping.

[0130] S44: Based on the operation fluency score, operation standardization score, and technical accuracy score, input them into the pre-trained evaluation and analysis model to generate a comprehensive evaluation report.

[0131] In this embodiment, this step aims to generate a comprehensive, intuitive, and instructive final evaluation result. A simple weighted average may not accurately reflect the overall level of performance; the pre-trained model can learn the comprehensive evaluation logic of experts on multi-dimensional indicators.

[0132] Specifically, the scores calculated from S41, S42, and S43, along with other derived indicators (such as total time and number of virtual complication triggers), form a feature vector, which is then input into a pre-trained machine learning model. This model (e.g., a gradient boosting tree model) is trained using a large amount of student performance data (containing the same features and comprehensive scores given by experienced clinical mentors). The model learns the complex non-linear relationship between each dimension indicator and the final overall performance level. The model outputs a comprehensive score from 0 to 100, along with a concise rating such as Excellent, Good, Satisfactory, or Needs Further Training. Finally, the system automatically generates a visually appealing evaluation report.

[0133] Furthermore, the report includes an overall score and grade; radar charts of scores for each dimension; a list of key errors and their occurrence times; personalized training suggestions for weak areas; and optional video replays of key steps in the operation process, recorded based on the scenario, for students and instructors to review and analyze together.

[0134] In one embodiment, such as Figure 6 As shown, in step S42, the identified operational actions are compared with the standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the standardization of aseptic operation, generating an operational standardization score, which specifically includes:

[0135] S421: Based on the operation actions, identify whether the key steps of disinfection, draping, puncture point selection, and catheter fixation are performed in the preset order, and calculate the number of times the steps are missing or the order is incorrect.

[0136] In this embodiment, this sub-step focuses on ensuring the correct sequence of the core procedures. Central venous catheterization follows a strict sequence logic; violating this sequence may lead to serious consequences such as infection or puncture failure.

[0137] Specifically, from the entire sequence of procedures, action tags representing several key milestones are selected, such as "Start Disinfection," "Lay Sterile Drapes," "Ultrasound Probe Positioning," "Needle Insertion," and "Catheter Fixation." Then, the preset order of these key steps is searched in the SOP database. For example, the standard order must be: Disinfection - Draping - Positioning - Puncture - ... - Fixation. The system iterates through the identified key action sequences, checking for any reversed order (e.g., "Puncture" appearing before "Disinfection") or missing necessary steps (e.g., the "Draping" action is not identified at all). For each incorrect or missing step, the "Sequence Error Counter" or "Missing Step Counter" is incremented by one.

[0138] S422: During the virtual ultrasound-guided puncture phase, analyze the usage time and scanning range of the virtual ultrasound probe during the operation to determine whether it meets the standard requirements for covering the target blood vessel area.

[0139] In this embodiment, this sub-step assesses the standardization of modern ultrasound-guided puncture. Standardization requires the operator to confirm the location, course, and relationship with surrounding tissues of the blood vessel through system scanning, rather than simply "pointing it out."

[0140] Specifically, when the system recognizes that the trainee is holding a "virtual ultrasound probe" and enters the "scanning" action mode, it begins recording the duration and probe movement trajectory of this phase. The system divides the virtual patient's neck region into several regions of interest (ROI) grids. The evaluation engine analyzes the probe's movement trajectory and calculates the percentage of the number of grids it covers relative to the total number of grids in the target area (such as a preset area including the internal jugular vein and carotid artery), which is used as the "scan coverage rate." Simultaneously, the scan duration is recorded. Standard requirements typically include "scan coverage of major vascular areas" and "observation time of at least N seconds to confirm the absence of vascular variations." The system determines whether the "scan coverage rate" reaches a preset threshold (e.g., 80%) and whether the "scan duration" exceeds the minimum requirement (e.g., 10 seconds). If both are met, this is considered compliant; otherwise, it is recorded as non-compliant.

[0141] S423: During the guidewire and catheter insertion phase, based on the complication judgment markers recorded in the operation action record, count the number of virtual complication alarms triggered by improper operation.

[0142] In this embodiment, this sub-step uses the results to infer the compliance of the operation. The system's built-in physiological engine simulates complications caused by non-compliant operation, and triggering an alarm is itself a serious violation of compliance standards.

[0143] Specifically, in the virtual system, improper operation will trigger predefined "complication judgment markers." For example, if the "guidewire advancement" action is performed too quickly or with too much amplitude, it may trigger the "arrhythmia" marker; if the "catheter insertion" length far exceeds the standard, it may trigger the "point displacement" marker; if the virtual instrument accidentally touches the "brachial plexus" model area during the operation, it may trigger the "nerve stimulation" marker. The evaluation engine does not directly judge the action itself, but monitors the number of these "complication alarm" events recorded in the system log throughout the entire operation. Each alarm clearly corresponds to an operational error that may endanger the patient, and therefore directly serves as a serious deduction item in the standardization assessment. The more alarms, the lower the standardization score.

[0144] S424: Generate an operational standardization score based on the number of missing or incorrect steps, the number of virtual complication alarms, and the judgment results.

[0145] Specifically, the system presets a base score of 100. Each error deducts points based on its severity. For example, missing a core step (such as disinfection) may result in a direct deduction of 20 points; an incorrect sequence (such as puncture before draping) deducts 15 points; improper ultrasound scanning (S422 not meeting standards) deducts 5 points; and each triggered complication alarm deducts 10 points. Deductions can be accumulated, but the minimum score is usually set to zero. The calculation formula is: Standardization Score = 100 - (Σ(Number of Errors * Corresponding Deduction Value)). The final score is between 0 and 100, directly reflecting the trainee's adherence to operating procedures.

[0146] In one embodiment, such as Figure 7 As shown, in step S43, the virtual needle tip trajectory data is quantified according to the current operational intent, and the puncture angle deviation, puncture point position error, and catheter insertion length accuracy are calculated to generate a technical accuracy score, specifically including:

[0147] S431: Obtain the puncture angle sequence according to the current operation intention, and calculate the average and maximum deviations between the actual needle insertion angle and the preset standard needle insertion angle based on the puncture angle sequence.

[0148] In this embodiment, this sub-step precisely quantifies the accuracy of the puncture angle. The needle insertion angle is crucial in determining the puncture depth and path; too large or too small an angle may damage adjacent tissue or lead to puncture failure.

[0149] Specifically, when the operation intent resolved by S20 remains "puncture and needle insertion," the system samples the virtual needle tip's posture at a high frequency (e.g., 10 times per second). Based on the needle tip direction vector and the surface normal vector at the virtual skin puncture point, the instantaneous needle insertion angle is calculated in real time for each sampling moment. This results in a "puncture angle sequence." After the entire sequence is completed, the average of all angle values ​​in the sequence is calculated as the "actual average needle insertion angle." Then, the absolute value of the difference between this average and the "preset standard needle insertion angle" is calculated to obtain the "average deviation." Simultaneously, the entire sequence is traversed to find the value with the largest difference from the standard angle, obtaining the "maximum deviation." The average deviation reflects the overall angle control level, while the maximum deviation reflects the moment of most severe angle loss of control.

[0150] S432: Obtain the puncture depth sequence according to the current operation intention, identify the depth value when the needle tip touches the target blood vessel wall based on the puncture depth sequence, and calculate the absolute error between it and the preset standard anatomical depth value.

[0151] In this embodiment, this sub-step assesses the accuracy of the puncture depth. Ideally, the puncture should be "spot on the first try"; too deep and it may penetrate the posterior wall of the blood vessel, while too shallow and it may not be able to enter the lumen of the blood vessel.

[0152] Specifically, during the "puncture and needle insertion" intention phase, the system records the depth of the virtual needle tip from the skin puncture point in real time, forming a "puncture depth sequence." Simultaneously, the virtual physiological engine detects collision events between the needle tip and the "target blood vessel (e.g., the internal jugular vein) inner wall" model. When the first "entry into the blood vessel" collision event occurs (synchronized by a "feeling of emptiness" tactile feedback marker), the system records the depth value at that moment as the "measured blood vessel depth." Based on the current virtual patient's body shape parameters (set at the start of training), the system invokes an anatomical depth prediction model to provide a "preset standard anatomical depth value" (e.g., for a standard-sized adult male, the right internal jugular vein depth is approximately 1.5-2.5 cm). The absolute error between the measured depth and the standard depth is calculated. The smaller this error value, the more accurate the trainee's control over the needle insertion depth.

[0153] S433: Obtain the motion trajectory of the virtual medical device according to the current operation intention, analyze the path of the needle tip before reaching the target blood vessel based on the motion trajectory of the virtual medical device, and calculate the maximum distance of its deviation from the preset ideal puncture path.

[0154] In this embodiment, this sub-step assesses the stability of the puncture path. The ideal path is a straight line from the skin puncture point to the target vessel. In practice, hand tremors or adjustments may cause the needle tip to "row" or follow a curve under the skin.

[0155] Specifically, the three-dimensional spatial coordinate sequence of the virtual needle tip, i.e., the "needle tip trajectory," is recorded from needle insertion to the first entry into the blood vessel. Simultaneously, a "preset ideal puncture path" (a straight line) is defined in virtual space from the skin puncture point to the blood vessel target point. For each sampling point on the trajectory, the vertical distance from that point to this ideal straight line is calculated. By iterating through all points, the maximum value of this vertical distance is found, which is the "maximum deviation distance." This value reflects the furthest deviation of the needle tip from the ideal straight line during its journey, and is an important indicator for evaluating operational stability and the ability to achieve "one-shot success."

[0156] S434: Based on the above calculation results, generate a technical accuracy score.

[0157] Specifically, the technical accuracy score uses a weighted scoring method. First, the raw error values ​​calculated from S431 to S433 (average angle deviation, maximum angle deviation, absolute depth error, and maximum path deviation distance) are normalized and mapped to a range of 0-1 (0 represents perfect accuracy, and 1 represents exceeding the maximum tolerance error). Then, a weight coefficient is assigned to each error indicator. These weight coefficients are determined through expert consultation or historical data statistical analysis, with depth and angle typically having higher weights. Finally, the accuracy score is calculated, with the final score ranging from 0 to 100 points; a higher score indicates more precise technical operation.

[0158] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0159] In one embodiment, a central venous catheterization medical operation training system based on virtual reality and gesture tracking is provided. This training system corresponds one-to-one with the central venous catheterization medical operation training method based on virtual reality and gesture tracking described in the above embodiments. Figure 8 As shown, this central venous catheterization medical operation training system based on virtual reality and gesture tracking includes a multimodal operation data acquisition module, an operation intention recognition module, a tactile feedback generation module, and an operation evaluation module. Detailed descriptions of each functional module are as follows:

[0160] The multimodal operation data acquisition module is used to collect user hand movement data in real time. The hand movement data includes the three-dimensional coordinates of key hand points, joint angles, and movement trajectory information.

[0161] The operation intent recognition module is used to identify the standard operation actions of central venous catheterization performed by the user based on the hand movement data through a hierarchical fusion gesture recognition algorithm, and to analyze the user's current operation intent in combination with virtual reality visual scene information;

[0162] The haptic feedback generation module is used to generate and output corresponding multimodal haptic feedback signals based on the identified operation actions and current operation intentions, simulating resistance, vibration and temperature changes in real operation;

[0163] The operation evaluation module is used to quantitatively evaluate the user's operation based on the hand movement data, operation action recognition results, and current operation intention through the evaluation engine, and generate a comprehensive evaluation report.

[0164] Specific limitations regarding the training system for central venous catheterization based on virtual reality and gesture tracking can be found in the above section on the limitations of the training method for central venous catheterization based on virtual reality and gesture tracking, and will not be repeated here. Each module in the aforementioned training system for central venous catheterization based on virtual reality and gesture tracking can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the electronic device, or stored in the memory of the electronic device as software, so that the processor can call and execute the corresponding operations of each module.

[0165] In one embodiment, an electronic device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores user hand movement data and haptic feedback signals. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a central venous catheterization medical operation training method based on virtual reality and gesture tracking.

[0166] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0167] Real-time acquisition of user hand movement data, including three-dimensional coordinates of key hand points, joint angles, and movement trajectory information;

[0168] Based on the hand movement data, a hierarchical fusion gesture recognition algorithm is used to identify the standard central venous catheterization operation performed by the user, and the user's current operation intention is analyzed in combination with virtual reality visual scene information;

[0169] Based on the identified operational actions and current operational intentions, corresponding multimodal tactile feedback signals are generated and output to simulate resistance, vibration, and temperature changes in real-world operations.

[0170] Based on the hand motion data, operation action recognition results, and current operation intention, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report.

[0171] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0172] Real-time acquisition of user hand movement data, including three-dimensional coordinates of key hand points, joint angles, and movement trajectory information;

[0173] Based on the hand movement data, a hierarchical fusion gesture recognition algorithm is used to identify the standard central venous catheterization operation performed by the user, and the user's current operation intention is analyzed in combination with virtual reality visual scene information;

[0174] Based on the identified operational actions and current operational intentions, corresponding multimodal tactile feedback signals are generated and output to simulate resistance, vibration, and temperature changes in real-world operations.

[0175] Based on the hand motion data, operation action recognition results, and current operation intention, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report.

[0176] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0177] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0178] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A training method for central venous catheterization based on virtual reality and gesture tracking, characterized in that, The training method for central venous catheterization based on virtual reality and gesture tracking includes the following steps: Real-time acquisition of user hand movement data, including three-dimensional coordinates of key hand points, joint angles, and movement trajectory information; Based on the hand movement data, a hierarchical fusion gesture recognition algorithm is used to identify the standard central venous catheterization operation performed by the user, and the user's current operation intention is analyzed in combination with virtual reality visual scene information; Based on the identified operational actions and current operational intentions, corresponding multimodal tactile feedback signals are generated and output to simulate resistance, vibration, and temperature changes in real-world operations. Based on the hand motion data, operation action recognition results, and current operation intention, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report.

2. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 1, characterized in that, The real-time collection of user hand movement data specifically includes: Color and depth images of the hand were acquired using a binocular depth camera, and the two-dimensional pixel coordinates of 21 key points of the hand were extracted based on an image recognition algorithm. Infrared reflection data of the hand surface is acquired by an infrared sensor array, and the three-dimensional spatial coordinates of key points of the hand are calculated by combining the depth image. The angular velocity, acceleration, and orientation data of the hand joints are collected by an inertial measurement unit worn on the hand. The two-dimensional pixel coordinates, three-dimensional spatial coordinates, and inertial measurement data are spatiotemporally aligned and fused to generate a complete hand motion data sequence containing timestamps.

3. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 1, characterized in that, The process of identifying standard central venous catheterization procedures performed by the user based on the hand movement data using a hierarchical fusion gesture recognition algorithm specifically includes: Based on the coordinates of key skeletal points in the hand motion data, the angle changes and relative positional relationships of each finger joint are calculated through the underlying skeletal tracking layer. The hand motion data of consecutive frames is constructed into a spatiotemporal graph structure, and the temporal evolution characteristics of the hand motion sequence are analyzed by spatiotemporal graph convolutional network. By combining the knowledge base of central venous catheterization procedures, a high-level intent understanding layer maps action features into specific clinical operational intents. Based on the matching degree between the operation intention and the action sequence, the operation parameters are identified, and the completion status of the operation steps is marked.

4. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 1, characterized in that, The process involves generating and outputting corresponding multimodal tactile feedback signals based on the identified operational actions and current operational intentions, simulating resistance, vibration, and temperature changes during real-world operation. Specifically, this includes: Based on the operation actions and current operation intentions, and combined with the virtual soft tissue mechanical model, the magnitude and direction of the real-time force generated when the virtual needle tip interacts with the virtual blood vessels and tissues are calculated. The calculated real-time force is mapped into a controllable current signal, a vibration frequency signal, and a temperature control signal, and the controllable current signal, vibration frequency signal, and temperature control signal are sent to the corresponding tactile feedback device. The device generates and outputs corresponding multimodal tactile feedback signals based on the tactile feedback device.

5. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 4, characterized in that, Based on the hand motion data, operation action recognition results, and current operation intent, the evaluation engine performs a quantitative evaluation of the user's operation and generates a comprehensive evaluation report, specifically including: Based on the hand movement data, the hand tremor frequency, movement speed, and acceleration changes are calculated to evaluate the operation fluency dimension score; The identified operational actions are compared with the standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the standardization of aseptic operation, and an operation standardization score is generated. The virtual needle tip trajectory data is quantified based on the current operational intent, and the puncture angle deviation, puncture point position error and catheter insertion length accuracy are calculated to generate a technical accuracy score. Based on the operation fluency score, operation standardization score, and technical accuracy score, the data are input into a pre-trained evaluation and analysis model to generate a comprehensive evaluation report.

6. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 5, characterized in that, The process involves comparing the identified operational actions with a standard operating procedure database to check the completeness of the steps, the correctness of the sequence, and the compliance with aseptic operation standards, generating an operational compliance score. This specifically includes: Based on the operation actions, identify whether the key steps of disinfection, draping, puncture point selection and catheter fixation are performed in the preset order, and calculate the number of times the steps are missing or the order is incorrect; During the virtual ultrasound-guided puncture phase, the duration of use and scanning range of the virtual ultrasound probe during the operation are analyzed to determine whether they meet the standard requirements for covering the target blood vessel area. During the guidewire and catheter insertion phase, based on the complication judgment markers recorded in the operation action record, the number of virtual complication alarms triggered by improper operation is counted. An operational standardization score is generated based on the number of missing or incorrect steps, the number of virtual complication alarms, and the judgment results.

7. The central venous catheterization medical operation training method based on virtual reality and gesture tracking according to claim 5, characterized in that, The process of quantifying virtual needle tip trajectory data based on the current operational intent, calculating puncture angle deviation, puncture point position error, and catheter insertion length accuracy, and generating a technical accuracy score specifically includes: Based on the current operational intent, a puncture angle sequence is obtained, and the average and maximum deviations between the actual needle insertion angle and the preset standard needle insertion angle are calculated based on the puncture angle sequence. According to the current operational intent, a puncture depth sequence is obtained. Based on the puncture depth sequence, the depth value when the needle tip touches the target blood vessel wall is identified, and the absolute error between the needle tip and the preset standard anatomical depth value is calculated. The movement trajectory of the virtual medical device is obtained according to the current operation intention. Based on the movement trajectory of the virtual medical device, the path of the needle tip before reaching the target blood vessel is analyzed, and the maximum distance of its deviation from the preset ideal puncture path is calculated. Based on the above calculation results, a technical accuracy score is generated.

8. A central venous catheterization medical operation training system based on virtual reality and gesture tracking, characterized in that, The central venous catheterization medical operation training system based on virtual reality and gesture tracking includes: The multimodal operation data acquisition module is used to collect user hand movement data in real time. The hand movement data includes the three-dimensional coordinates of key hand points, joint angles, and movement trajectory information. The operation intent recognition module is used to identify the standard operation actions of central venous catheterization performed by the user based on the hand movement data through a hierarchical fusion gesture recognition algorithm, and to analyze the user's current operation intent in combination with virtual reality visual scene information; The haptic feedback generation module is used to generate and output corresponding multimodal haptic feedback signals based on the identified operation actions and current operation intentions, simulating resistance, vibration and temperature changes in real operation; The operation evaluation module is used to quantitatively evaluate the user's operation based on the hand movement data, operation action recognition results, and current operation intention through the evaluation engine, and generate a comprehensive evaluation report.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the central venous catheterization medical operation training method based on virtual reality and gesture tracking as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the central venous catheterization medical operation training method based on virtual reality and gesture tracking as described in any one of claims 1 to 7.