A post training and cultivating system and method based on AR virtual vision
The AR virtual vision job training system utilizes intelligent AR terminals and training platforms to construct 3D environmental point cloud maps, integrates virtual devices, and provides real-time feedback on operational errors. This solves the problems of high equipment costs, significant safety hazards, and strong subjectivity in evaluation during university students' practical training, achieving efficient skills training and personalized development.
Patent Information
- Application Number
- CN202511221564.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-08-29
AI Technical Summary
College students lack practical work experience during the job search process. Physical equipment is costly to operate and poses safety hazards. Traditional virtual simulation systems lack real-world interaction and cannot train hand-eye coordination. Manual guidance is inefficient and highly subjective in evaluation.
An AR-based virtual vision-based job training system is adopted, including an intelligent AR terminal and a training system platform. Through environmental scanning, data collection, feedback modules, and adaptive evaluation modules, a three-dimensional environmental point cloud map is constructed, virtual devices are integrated, operational errors are fed back in real time, and personalized training programs are generated.
It achieves immersive training with zero physical loss, improves the sense of engagement and efficiency of training, corrects operational errors in real time, quantifies and evaluates skills, and reduces the subjectivity of human guidance.
Smart Images

Figure CN120726869B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vocational education and training technology, and more specifically, to a job training system and method based on AR virtual vision. Background Technology
[0002] There are many difficulties in the current employment process for college students. College students have little practical work experience. If companies provide actual operation, it will be costly. Practical training for college students relies on physical equipment, which is costly and poses safety hazards, such as mechanical operation. At the same time, existing traditional virtual simulation systems lack real environment interaction and cannot train hand-eye coordination. If traditional teaching is used, one mentor can only guide 5-10 students, and the evaluation is highly subjective. The training program only provides one-way operation prompts and lacks real-time error correction and quantitative skill assessment.
[0003] Therefore, in view of the many limitations and practical problems existing in the current background, in order to solve the three core problems of the shortage of real equipment, uncontrollable operational risks and low efficiency of manual guidance in college students' practical training, an AR job training system and adaptive training method based on multimodal virtual and real fusion are designed. The system achieves immersive training with zero physical loss through AR-environment fusion technology. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a job training system and method based on AR virtual vision to solve the problems existing in the above-mentioned background technology.
[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution: a job training system based on AR virtual vision, including an intelligent AR terminal worn by trainees and a training system platform set on the system terminal;
[0006] The intelligent AR terminal includes: an environment scanning module, used to scan and transmit data of the training environment where the trainees are located;
[0007] The data collection module is used to collect the trainees' operation data and transmit it to the training system platform.
[0008] The data feedback module is used to receive feedback data transmitted by the training system platform and provide feedback to the trainees through tactile and visual means.
[0009] The training system platform includes an environment perception construction module, which is used to construct a three-dimensional environment point cloud map model based on the real-time SLAM algorithm according to the data of the training scene environment where the trainees are located transmitted by the intelligent AR terminal.
[0010] The virtual-real construction module integrates and overlays virtual devices into the real scene based on the constructed 3D environment point cloud map model;
[0011] The adaptive evaluation module performs adaptive evaluation and operation data feedback based on the data fed back by the intelligent AR terminal and the dynamic model, and transmits the operation data feedback to the intelligent AR terminal.
[0012] The culture program generation module generates a corresponding culture program based on the adaptive evaluation transmitted by the adaptive evaluation module.
[0013] Optionally, the environment scanning module includes an RGB-D camera for collecting environmental data of the environment in which the trainee is located;
[0014] LiDAR scanners are used to scan environmental depth data to provide data for building point cloud map models.
[0015] Optionally, the data collection module includes:
[0016] Infrared hand tracking sensors are used to capture and transmit the trainee's joint pose data;
[0017] An eye tracker is used to record and transmit data on the trainee's gaze.
[0018] Pressure sensors are used to record and transmit the force with which trainees grip the tools.
[0019] Optionally, the data feedback module includes:
[0020] The visual feedback unit projects the corresponding visual signal within the AR field of view based on the transmitted feedback data.
[0021] The haptic feedback unit outputs haptic signals through a wristband-type vibration device based on the transmitted feedback data.
[0022] The spatial audio feedback unit outputs auditory signals based on the transmitted feedback data through sound conduction detection.
[0023] Optionally, the environment perception construction module includes:
[0024] The data synchronization unit is used to receive data on the training scene environment where the trainees are located, transmitted by the intelligent AR terminal, and output it after performing spatiotemporal synchronization.
[0025] The dynamic SLAM pose estimation unit, based on an improved pose graph optimization model, is used to extract and construct point cloud data from the training scene environment data after scene recognition through multi-source weighted nonlinear optimization.
[0026] The dynamic object filtering unit, based on the motion consistency test model, removes moving object point clouds by time-series velocity filtering, while retaining the construction point cloud data of static environment structure;
[0027] The point cloud fusion and feature extraction unit, based on the processed constructed point cloud data, performs data preprocessing and then uses a feature extraction algorithm to extract the geometric features from the constructed point cloud data;
[0028] The map model building unit is used to semantically bind the extracted geometric features, and then construct an octagonal map based on the semantically bound geometric features, thereby obtaining a 3D environmental point cloud map model.
[0029] Optionally, the virtual-real construction module includes:
[0030] The virtual device generation unit is used to generate corresponding virtual devices based on the training content and project them into the virtual reality environment according to the point cloud coordinates of the 3D environment point cloud map model.
[0031] The virtual-real rendering unit calculates the occlusion relationship between real objects and virtual devices in real time, and renders the shadow relationship and physical relationship to the virtual reality environment accordingly.
[0032] Optionally, the adaptive evaluation module includes:
[0033] The real-time operation data feedback unit, based on the set cloud skill database, generates feedback data in real time and transmits it to the smart AR terminal based on the real-time transmission of operation data;
[0034] The operation trajectory analysis subunit is used to perform time-series segmentation and separation of the overall operation data based on the transmission, decomposing it into operation data for the preparation stage, execution stage, and verification stage.
[0035] The timing logic scoring subunit is used to set weight coefficients for the preparation, execution, and verification phases, and output evaluation scores for each phase.
[0036] The multi-dimensional skill map generation unit, based on the evaluation score output by the temporal logic scoring subunit, outputs a radar chart that includes spatial positioning accuracy, operational fluency, and anomaly handling capability, and then fits and generates an adaptive evaluation.
[0037] Optionally, the culture scheme generation module includes:
[0038] The training program generation unit, based on the generated adaptive evaluation, uses a deep learning model to dynamically arrange and generate training programs for corresponding trainees.
[0039] A training method for a job training system based on AR virtual vision, as described above, is characterized by comprising:
[0040] S1. Environment binding: Trainees wear smart AR terminals and scan the overall data of the training environment through the environment scanning module on the smart AR terminal. Then, they construct a three-dimensional environment point cloud map model through the environment perception construction module of the training system platform.
[0041] S2. Virtual-Real Integration: The training system platform is based on a constructed 3D environment point cloud map model. It integrates and overlays virtual devices in the real scene, binds the virtual device model to the real desktop coordinate system, and calculates the occlusion relationship between real objects and virtual devices in real time.
[0042] S3. Operation detection: Trainees wear smart AR terminals and, during simulated practical operations, the data collection module on the smart AR terminal collects the trainees' operation data and transmits it to the training system platform.
[0043] S4. Operation feedback: Based on the transmitted operation data, after real-time feedback calculation, the feedback signal is transmitted to the smart AR terminal, which then converts it into a prompt signal to alert the student.
[0044] S5. Dynamic evaluation: The training system platform receives and transmits operational data, performs operational evaluation based on the approved evaluation criteria, and outputs adaptive evaluation.
[0045] S6. Training program generation: Based on adaptive evaluation and operational data, a personalized training program is generated by matching the dynamic strategy library, and the corresponding training program for each trainee is dynamically arranged.
[0046] In summary, the present invention has the following beneficial effects:
[0047] 1. Utilizing AR virtual reality, this system innovatively employs dynamic environment binding technology. An RGB-D camera scans training scenarios in real time, such as laboratories and factory workstations. A SLAM algorithm is used to construct a 3D environmental point cloud map, automatically aligning virtual devices with the real desktop coordinate system. Virtual devices can be "placed" on the real desktop, subject to physical occlusion constraints. For example, when a hand passes through a virtual device, the occluded portion is automatically hidden. Compared to traditional AR projection technology, this is more realistic, providing a stronger sense of immersion and training, thus enhancing the effectiveness of training and shortening the job adaptation cycle.
[0048] 2. A multimodal operation evaluation mechanism was adopted, which dynamically compared the operation trajectory with the standard path and generated a deviation heatmap in real time. The operation steps were decomposed into three stages: preparation, execution and verification, and scored separately to further realize training scoring. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the system implementation steps of the present invention;
[0050] Figure 2 This is a schematic diagram of the implementation steps of the method of the present invention;
[0051] Figure 3 This is a schematic diagram of the implementation steps of the data collection module of the present invention;
[0052] Figure 4 This is a schematic diagram of the feedback unit architecture of the real-time data feedback unit of the present invention;
[0053] Figure 5 This is a flowchart illustrating the virtual interactive operation steps of the present invention;
[0054] Figure 6 This is a flowchart illustrating the three-stage weight allocation model of the present invention. Detailed Implementation
[0055] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein.
[0056] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature.
[0057] In this invention, unless otherwise expressly specified and limited, "above" or "below" a second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of a second feature includes the first feature being directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" of a second feature includes the first feature being directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature. The terms "vertical," "horizontal," "left," "right," "above," "below," and similar expressions are for illustrative purposes only and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed or operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0058] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0059] This invention provides a job training system based on AR virtual vision, such as... Figure 1 As shown, it includes a smart AR terminal worn by trainees and a practical training system platform set on the system terminal;
[0060] The intelligent AR terminal includes: an environment scanning module, used to scan and transmit data of the training environment where the trainees are located;
[0061] The data collection module is used to collect the trainees' operation data and transmit it to the training system platform.
[0062] The data feedback module is used to receive feedback data transmitted by the training system platform and provide feedback to the trainees through tactile and visual means.
[0063] The training system platform includes an environment perception construction module, which is used to construct a three-dimensional environment point cloud map model based on the real-time SLAM algorithm according to the data of the training scene environment where the trainees are located transmitted by the intelligent AR terminal.
[0064] The virtual-real construction module integrates and overlays virtual devices into the real scene based on the constructed 3D environment point cloud map model;
[0065] The adaptive evaluation module performs adaptive evaluation and operation data feedback based on the data fed back by the intelligent AR terminal and the dynamic model, and transmits the operation data feedback to the intelligent AR terminal.
[0066] The culture program generation module generates a corresponding culture program based on the adaptive evaluation transmitted by the adaptive evaluation module.
[0067] Furthermore, the environment scanning module includes an RGB-D camera for collecting environmental data of the environment in which the trainee is located;
[0068] LiDAR scanners are used to scan environmental depth data to provide data for building point cloud map models.
[0069] In the specific implementation process, the RGB-D camera can be oriented to a location, such as a fixed point on a roof truss, and can rotate 360 degrees based on the fixed point to capture the simulated training scene location for trainees; the LiDAR scanner is set at a high position and performs a global scan every 2 hours. When a moving object is detected, it performs a localized encrypted scan to achieve area focusing and further global scanning of the training scene.
[0070] Furthermore, the data collection module includes:
[0071] Infrared hand tracking sensors are used to capture and transmit the trainee's joint pose data;
[0072] An eye tracker is used to record and transmit data on the trainee's gaze.
[0073] Pressure sensors are used to record and transmit the force with which trainees grip the tools.
[0074] In the specific implementation process, trainees perform operations, and the data collection module collects the corresponding operation data, which is then transmitted to the training system platform for identification and real-time feedback. The data collection module needs to identify the virtual devices operated by the trainees, the operation methods (e.g., grasping, rotating, pressing), and the operation parameters (e.g., force, angle, speed). Figure 3 As shown;
[0075] Virtual devices are overlaid in a virtual scene. Each virtual device has a unique spatial activation area. Let the position of the hand joint be... The device activation area is The device activation area is activated, indicated as ;
[0076] The system acquires student action data using infrared hand tracking sensors, eye trackers, and pressure sensors. Based on whether the virtual device is activated, the data is fused using a data association model and preprocessed using a Kalman filter model, resulting in the following representation:
[0077]
[0078] After the operational data is transmitted to the training system platform, the real-time data feedback unit of the training system platform provides feedback according to the feedback unit architecture, such as... Figure 4 As shown;
[0079] In one type of feedback, when the operation data is determined to have an operation space error, that is, when the virtual device is used improperly, the smart AR terminal generates a three-dimensional guide grid, projects directional indicator arrows, and provides tactile vibration prompts.
[0080] Feedback results are as follows: positional error > 5mm: display a semi-transparent green grid for guidance; angle error > 8°: display a red directional arrow; angle error > 15°: superimpose wrist vibration.
[0081] One type of feedback is divided into three different operational stages when the operational data is determined to be in the case of improper operation;
[0082] During the preparation phase, if the wrong virtual tool is selected, the correct virtual tool will be highlighted, and a voice prompt will be given indicating that the wrong tool has been selected.
[0083] During the execution phase, if an operational error occurs, the current step will be demonstrated at 0.5x speed and the trainee will be prompted; if the operation times out, a voice prompt will say "Please speed up the operation"; if the operation sequence is incorrect, the system will force a rollback to the previous step, with a red flashing warning and vibration alert.
[0084] During the result verification phase, the implementation results are verified. If they do not meet the specifications, the operation is restarted and the system is reset to its original position.
[0085] During operation, a multimodal feedback matrix can also be used, represented as:
[0086]
[0087] This will further improve the scenario adaptability during actual simulation operations.
[0088] Furthermore, the data feedback module includes:
[0089] The visual feedback unit projects the corresponding visual signal within the AR field of view based on the transmitted feedback data.
[0090] The haptic feedback unit outputs haptic signals through a wristband-type vibration device based on the transmitted feedback data.
[0091] The spatial audio feedback unit outputs auditory signals based on the transmitted feedback data through sound conduction detection.
[0092] Furthermore, the environment perception construction module includes:
[0093] The data synchronization unit is used to receive data on the training scene environment where the trainees are located, transmitted by the intelligent AR terminal, and output it after performing spatiotemporal synchronization.
[0094] The dynamic SLAM pose estimation unit, based on an improved pose graph optimization model, is used to extract and construct point cloud data from the training scene environment data after scene recognition through multi-source weighted nonlinear optimization.
[0095] The dynamic object filtering unit, based on the motion consistency test model, removes moving object point clouds by time-series velocity filtering, while retaining the construction point cloud data of static environment structure;
[0096] The point cloud fusion and feature extraction unit, based on the processed constructed point cloud data, performs data preprocessing and then uses a feature extraction algorithm to extract the geometric features from the constructed point cloud data;
[0097] The map model building unit is used to semantically bind the extracted geometric features, and then construct an octagonal map based on the semantically bound geometric features, thereby obtaining a 3D environmental point cloud map model.
[0098] In the specific implementation process, when environmental data is input into the environmental perception construction module, the data synchronization unit preprocesses the input environmental data and performs data spatiotemporal synchronization. Let the sensor set be represented as... Its data stream timestamp alignment model is as follows:
[0099]
[0100] in For sensors The transformation matrix at time t, It is represented as the clock error covariance matrix of each sensor. By solving for the time offset that minimizes the pose difference, hardware-level synchronization of multiple sensors is achieved.
[0101] The dynamic SLAM pose estimation unit, through an improved pose graph optimization model, namely a multi-source tightly coupled optimization model, is expressed as: ;
[0102] in Represented as visual residuals, denoted as , Represented as a camera projection function, Let j be the j-th map point, corresponding to the real coordinates. Represented as the observed pixel coordinates in frame i;
[0103] in Represented as the laser residual term, using the point-to-surface ICP error model, it is expressed as:
[0104]
[0105] in Represented as the current frame point, Hehe This is represented as the point and normal vector corresponding to the target frame, and further optimized through dynamic point filtering constraints, as follows:
[0106]
[0107] in Represented as IMU residuals, optimized using a pre-integral model, and expressed as...
[0108] in Represented as IMU residuals, optimized using a pre-integral model, and expressed as...
[0109]
[0110] in Represented as a time interval, This is represented as a zero-difference change, which is then used to obtain the pose point cloud data;
[0111] The dynamic object filtering unit performs motion consistency detection on the pose point cloud data, using a motion consistency test model, expressed as follows:
[0112]
[0113] in Represented as point cloud velocity, Represented as a dynamic threshold, it removes moving object point clouds based on temporal velocity filtering while preserving static environmental structures, thereby eliminating dynamic interference point clouds and improving mapping accuracy.
[0114] The point cloud fusion and feature extraction unit receives the processed point cloud data and constructs a model. Specifically, the implementation process includes:
[0115] Set the depth camera point cloud as LiDAR point cloud is ;
[0116] The point cloud fusion model is used for fusion, which is represented as follows: ;in and These are respectively represented as the transformation matrices from the pre-calibrated camera / radar to the common coordinate system. and The weighting coefficients are dynamically allocated using a weighting function, expressed as follows:
[0117]
[0118] and The measurement uncertainties of depth cameras and lidar are respectively determined to further adapt to actual application scenarios. By using variance inverse weighted fusion of multi-source point clouds, the reconstruction accuracy is improved, the data loss problem of depth cameras on reflective surfaces is solved, the integrity of point cloud data is improved, and the fusion error generated during fusion is reduced.
[0119] The fused point cloud data undergoes normal vector estimation and plane segmentation to elucidate geometric features, forming the main framework of the point cloud data model, represented as PCA-based normal vector calculation: ,in Represented as a point The covariance matrix of the neighborhood of ;
[0120] Plane detection uses the RANSAC optimization model, represented as: , The planar distance threshold is used; the surface geometric properties are estimated through principal component analysis to provide a physical property basis for virtual object collision detection, thereby improving the accuracy of planar recognition and reducing the error of normal vector estimation.
[0121] The map model building unit performs semantic annotation and attribute binding on the identified point cloud data, and uses point cloud semantic labels to transmit the model, represented as... ,in Represented as a projection function, Mapping 3D points to image coordinates Computation is performed using the YOLOv5 network, represented as follows: , It serves as the backbone network for CSPDarknet.
[0122] An attribute binding extended model is used to qualitatively characterize some data in point cloud data, representing them as follows: ,in Represented as an attribute vector, such as electrical conductivity and heat capacity. This is a predefined attribute dictionary; through the above model definition, the visual recognition results are associated with point cloud coordinates, and environmental physical attributes are assigned to achieve advanced interaction and improve the device recognition accuracy.
[0123] Based on the semantically bound point cloud data, a hierarchical octree map is constructed, represented as follows:
[0124]
[0125] Each voxel storage ,in Represented as point cloud density, the tree-like data structure of the constructed octree map enables efficient map storage and querying, supports real-time collision detection, and ultimately generates the final 3D environment point cloud map model.
[0126] Furthermore, the virtual-real construction module includes:
[0127] The virtual device generation unit is used to generate corresponding virtual devices based on the training content and project them into the virtual reality environment according to the point cloud coordinates of the 3D environment point cloud map model.
[0128] The virtual-real rendering unit calculates the occlusion relationship between real objects and virtual devices in real time, and renders the shadow relationship and physical relationship to the virtual reality environment accordingly.
[0129] In the specific implementation process, the 3D environment point cloud map model is used as the base model, namely the octree map, which is represented as:
[0130]
[0131] in, Represented as voxel center coordinates, Represented as a normal vector, it is used to determine the surface orientation. Represented as point cloud density, it is used as a basis for collision detection. These are represented as semantic tags, such as specific objects like "workbench" or "power distribution cabinet". These are expressed as physical properties, such as coefficient of friction and electrical conductivity.
[0132] When integrating virtual devices into a real-world environment, the target location can be selected based on the platform's settings or by having the user specify a placement area in the real-world scene using gestures. The system then queries the corresponding voxel and normal vector for that location and calculates the pose matrix, represented as follows:
[0133]
[0134] in This indicates that the Z-axis of the virtual device is aligned with the surface normal; t indicates that the center of the virtual device is placed at... h represents the device height, which binds the virtual device to the physical attribute of this voxel attribute and onto the virtual device.
[0135] A depth testing algorithm and edge anti-aliasing optimization are used to reflect the handling of virtual and real occlusion between real objects and virtual devices;
[0136] An integrated physics engine, employing a collision detection model, is represented as follows:
[0137]
[0138] in, Represented as the radius of the sphere bounded by the virtual device. Represented as a set of neighborhood voxels;
[0139] When actual operations result in a collision, rigid body dynamics is used to calculate the reaction force based on physical properties, expressed as:
[0140]
[0141] When actually doing it, such as Figure 5 As shown, this is done to provide the same response as interacting with a real device, thereby further increasing the realism of the simulation.
[0142] Furthermore, the adaptive evaluation module includes:
[0143] The real-time operation data feedback unit, based on the set cloud skill database, generates feedback data in real time and transmits it to the smart AR terminal based on the real-time transmission of operation data;
[0144] The operation trajectory analysis subunit is used to perform time-series segmentation and separation of the overall operation data based on the transmission, decomposing it into operation data for the preparation stage, execution stage, and verification stage.
[0145] The timing logic scoring subunit is used to set weight coefficients for the preparation, execution, and verification phases, and output evaluation scores for each phase.
[0146] The multi-dimensional skill map generation unit, based on the evaluation score output by the temporal logic scoring subunit, outputs a radar chart that includes spatial positioning accuracy, operational fluency, and anomaly handling capability, and then fits and generates an adaptive evaluation.
[0147] In the specific implementation process, the operation trajectory analysis subunit decomposes the operation data according to the time sequence based on the operation of each time stage, thereby achieving different scores for operations at different stages;
[0148] The temporal logic scoring subunit constructs a three-stage weight allocation model, defining the stages and weight allocation, as follows: Figure 6 As shown, in the actual operation process, the scoring for each stage is set as follows:
[0149]
[0150] The three-dimensional ability score is expressed as follows:
[0151]
[0152] In the preparation stage, the evaluation is based on three different dimensions: correctness of tool selection, accuracy of parameter setting, and safety compliance check. After weighted fusion, the score is calculated. In this embodiment, the correctness of tool selection is 0.4, the accuracy of parameter setting is 0.4, and the safety compliance check is 0.2. The preparation stage score is output after fusion and weighting.
[0153] When scoring during the execution phase, a dynamic evaluation model is used, represented as follows:
[0154]
[0155] The spatial accuracy score is as follows: ,in This is represented by the attenuation coefficient, and by the control error sensitivity, which is 0.05 in this embodiment. t represents the dynamic time warping distance, which is a measure of temporal spatial trajectory difference;
[0156] Timing smoothness score: ,in This is expressed as standard operating time. This represents the actual operation time. It is a hyperbolic tangent function used to compress the time ratio to the [0,1] interval and eliminate the influence of extreme values;
[0157] The anti-interference score is: ,in This represents the number of correct operations. E represents the total number of operations, and E represents the number of erroneous operations.
[0158] Based on the above, the trainee's performance score is then obtained;
[0159] The verification phase is scored using four indicators: result accuracy, inspection time ratio, problem tracing depth, and effectiveness of improvement measures. The score for the verification phase is obtained by weighting and integrating the four indicators.
[0160] The adaptive evaluation uses a weakness identification algorithm to compare the results with the actual threshold and then derive an adaptive evaluation for the current test participant.
[0161] The basic evaluation criteria are expressed as follows: A score below 65 indicates insufficient spatial awareness.
[0162] Smoothness of operation A score below 70 indicates unfamiliarity with the operating procedures.
[0163] Exception handling capabilities A value less than 75 indicates a weak ability to respond to emergencies.
[0164] The composite evaluation condition is expressed as spatial positioning accuracy. Less than 65, and with smooth operation A score below 70 indicates poor hand-eye coordination.
[0165] Smoothness of operation Less than 70, and exception handling capability A value less than 75 indicates poor performance under pressure.
[0166] The multi-dimensional skill map generation unit converts the above scores into three-dimensional scores and then draws a capability radar chart. The specific implementation process includes...
[0167] Three-dimensional capability normalization, represented as ,in The radar value, represented as the k-th dimension of capability, is used to standardize capability scores, allowing managers to visually see the skill distribution map. Represented as raw capability values, used to represent raw values of space / fluidity / emergency capabilities, preserving individual differences; It is represented as the group mean, used to indicate the average level of trainees in the current period, in order to establish a dynamic baseline; The standard deviation is represented by the group standard deviation, which indicates the dispersion of students and identifies the top 10% of students. 20 is a constant, which is represented by the scaling factor to control the size of the radar chart and optimize the visualization effect. 60 is represented by the baseline offset to ensure that the scores are within a reasonable range and to avoid negative values.
[0168] Based on the above, a radar chart is drawn and visualized to display the data of the current test trainees and the average level of the trainees in the current period, thereby further improving the training efficiency.
[0169] Furthermore, the culture program generation module includes:
[0170] The training program generation unit, based on the generated adaptive evaluation, uses a deep learning model to dynamically arrange and generate training programs for corresponding trainees.
[0171] In the specific implementation process, the personalized generation of the training program uses a decision matrix for decision-making, as shown below:
[0172]
[0173] For example, in the application scenario of training students in the process of underwear making, college students wear smart AR terminals to transmit the students' operation data to the training system platform;
[0174] The data for the trainee's preparation phase are as follows: correct tool selection, score 95; parameter setting deviation of 0.3MPa, score 92; one safety procedure check was missed, score 2, total score 82; preparation phase score is 90.
[0175] The data for the student's execution phase are as follows: spatial accuracy (DTW) is 1.8mm, with a score of 91; timing smoothness is 0.88, with a score of 88; anti-interference performance is 92%, with a score of 92; and the overall execution phase score is 92.
[0176] The data for the student verification phase are as follows: accuracy of results: 98%, score: 98; inspection time: +10%, score: 90%; problem tracing depth: level three, score: 75; verification phase score: 88.
[0177] After three-dimensional capability calculation, the radar chart and evaluation output are as follows:
[0178] Radar chart features:
[0179] Spatial positioning accuracy: 89.5 (better than 85% of other trainees in the same period)
[0180] Operational smoothness: 89.6 (better than 82% of other students in the same period)
[0181] Exception handling ability: 88.5 (better than 78% of the same period's trainees)
[0182] Adaptive evaluation:
[0183] The trainees performed well in basic operations, but their ability to trace problems in depth was insufficient.
[0184] The recommended training program is as follows: Enhancement is suggested.
[0185] 1. Fault root cause analysis training × 6 times (enhancing level 3 tracing capabilities)
[0186] 2. Simulate sudden complex failures four times (to enhance the depth of emergency response).
[0187] 3. Increase time pressure on the verification process (currently +10% → target -5%)
[0188] Through a closed-loop technology system of phased weight allocation, multi-dimensional competency assessment, and adaptive scheme generation, the system achieves precise transformation from operational data to training value, providing a new generation of intelligent evaluation paradigm for vocational education.
[0189] A training method for a job training system based on AR virtual vision, as described above, is characterized by comprising:
[0190] S1. Environment binding: Trainees wear smart AR terminals and scan the overall data of the training environment through the environment scanning module on the smart AR terminal. Then, they construct a three-dimensional environment point cloud map model through the environment perception construction module of the training system platform.
[0191] S2. Virtual-Real Integration: The training system platform is based on a constructed 3D environment point cloud map model. It integrates and overlays virtual devices in the real scene, binds the virtual device model to the real desktop coordinate system, and calculates the occlusion relationship between real objects and virtual devices in real time.
[0192] S3. Operation detection: Trainees wear smart AR terminals and, during simulated practical operations, the data collection module on the smart AR terminal collects the trainees' operation data and transmits it to the training system platform.
[0193] S4. Operation feedback: Based on the transmitted operation data, after real-time feedback calculation, the feedback signal is transmitted to the smart AR terminal, which then converts it into a prompt signal to alert the student.
[0194] S5. Dynamic evaluation: The training system platform receives and transmits operational data, performs operational evaluation based on the approved evaluation criteria, and outputs adaptive evaluation.
[0195] S6. Training program generation: Based on adaptive evaluation and operational data, a personalized training program is generated by matching the dynamic strategy library, and the corresponding training program for each trainee is dynamically arranged.
[0196] This invention discloses an AR virtual vision-based job training system and method. Utilizing AR virtual reality, it innovatively employs dynamic environment binding technology. An RGB-D camera scans the training scene in real time, such as a laboratory or factory workstation. A SLAM algorithm is used to construct a 3D environmental point cloud map. Virtual devices automatically align with the real desktop coordinate system, allowing them to be "placed" on the real desktop. Physical occlusion constraints, such as automatically hiding obscured parts when a hand passes through a virtual device, contribute to a more realistic and immersive learning experience compared to traditional AR projection technology. This enhances the effectiveness of training and shortens the job adaptation period. A multimodal operation evaluation mechanism is employed, dynamically comparing the operation trajectory with the standard path and generating a deviation heatmap in real time. The operation steps are broken down into three stages: preparation, execution, and verification, each scored separately, further realizing training evaluation.
[0197] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A job training system based on AR virtual vision, characterized in that, This includes smart AR terminals worn by trainees and a practical training system platform set on the system terminal; The intelligent AR terminal includes: an environment scanning module, used to scan and transmit data of the training environment where the trainees are located; The data collection module is used to collect the trainees' operation data and transmit it to the training system platform. The data feedback module is used to receive feedback data transmitted by the training system platform and provide feedback to the trainees through tactile and visual means. The training system platform includes an environment perception construction module, which is used to construct a three-dimensional environment point cloud map model based on the real-time SLAM algorithm according to the data of the training scene environment where the trainees are located transmitted by the intelligent AR terminal. The virtual-real construction module integrates and overlays virtual devices into the real scene based on the constructed 3D environment point cloud map model; The adaptive evaluation module performs adaptive evaluation and operation data feedback based on the data fed back by the intelligent AR terminal and the dynamic model, and transmits the operation data feedback to the intelligent AR terminal. The cultivation scheme generation module generates a corresponding cultivation scheme based on the adaptive evaluation transmitted by the adaptive evaluation module. The environmental scanning module includes an RGB-D camera, used to collect environmental data of the environment in which the trainee is located; LiDAR scanners are used to scan environmental depth data to provide data for building point cloud map models; The environment perception construction module includes: The data synchronization unit is used to receive data on the training scene environment where the trainees are located, transmitted by the intelligent AR terminal, and output it after performing spatiotemporal synchronization. The dynamic SLAM pose estimation unit, based on an improved pose graph optimization model, is used to extract and construct point cloud data from the training scene environment data after scene recognition through multi-source weighted nonlinear optimization. The dynamic object filtering unit, based on the motion consistency test model, removes moving object point clouds by time-series velocity filtering, while retaining the construction point cloud data of static environment structure; The point cloud fusion and feature extraction unit, based on the processed constructed point cloud data, performs data preprocessing and then uses a feature extraction algorithm to extract the geometric features from the constructed point cloud data; The map model building unit is used to semantically bind the extracted geometric features, and then build an octagon map based on the semantically bound geometric features to obtain a 3D environmental point cloud map model. When environmental data is input into the environmental perception construction module, the data synchronization unit preprocesses the input environmental data and performs spatiotemporal synchronization. Let the sensor set be represented as... Its data stream timestamp alignment model is as follows: ; in For sensors The transformation matrix at time t, It is represented as the clock error covariance matrix of each sensor. By solving for the time offset that minimizes the pose difference, hardware-level synchronization of multiple sensors is achieved. The dynamic SLAM pose estimation unit, through an improved pose graph optimization model, is represented as: ; in Represented as visual residuals, denoted as , Represented as a camera projection function, Let j be the j-th map point, corresponding to the real coordinates. Represented as the observed pixel coordinates in frame i; in Represented as the laser residual term, using the point-to-surface ICP error model, it is expressed as: ; in Represented as the current frame , and This is represented as the point and normal vector corresponding to the target frame, and further optimized through dynamic point filtering constraints, as follows: ; in Represented as IMU residuals, optimized using a pre-integral model, it is expressed as: ; in Represented as a time interval, This is represented as a zero-difference change, which is then used to obtain the pose point cloud data; The point cloud fusion and feature extraction unit receives the processed point cloud data and constructs a model. Specifically, the implementation process includes: Set the depth camera point cloud as LiDAR point cloud is ; The point cloud fusion model is used for fusion, which is represented as follows: ;in and These are the transformation matrices from the pre-calibrated camera and radar to the common coordinate system, respectively. and The weighting coefficients are dynamically allocated using a weighting function, expressed as follows: ; and The measurement uncertainties of depth camera and lidar are respectively; the normal vector of the fused point cloud data is estimated and the plane is segmented, and then the geometric features are extracted to form the main framework of the point cloud data model. Principal component analysis is used to estimate the surface geometric properties, providing a physical property basis for virtual object collision detection. The map model building unit performs semantic annotation and attribute binding on the identified point cloud data, and uses point cloud semantic labels to transmit the model, represented as... ,in Represented as a projection function, Mapping 3D points to image coordinates Computation is performed using the YOLOv5 network, represented as follows: , It serves as the backbone network for CSPDarknet. Furthermore, an attribute binding extension model is adopted to qualitatively characterize some data in the point cloud data, representing it as follows: ,in Represented as an attribute vector, This is a predefined attribute dictionary.
2. The job training system based on AR virtual vision according to claim 1, characterized in that, The data collection module includes: Infrared hand tracking sensors are used to capture and transmit the trainee's joint pose data; An eye tracker is used to record and transmit data on the trainee's gaze. Pressure sensors are used to record and transmit the force with which trainees grip the tools.
3. The job training system based on AR virtual vision according to claim 1, characterized in that, The data feedback module includes: The visual feedback unit projects the corresponding visual signal within the AR field of view based on the transmitted feedback data. The haptic feedback unit outputs haptic signals through a wristband-type vibration device based on the transmitted feedback data. The spatial audio feedback unit outputs auditory signals based on the transmitted feedback data through sound conduction detection.
4. The job training system based on AR virtual vision according to claim 1, characterized in that, The virtual-real construction module includes: The virtual device generation unit is used to generate corresponding virtual devices based on the training content and project them into the virtual reality environment according to the point cloud coordinates of the 3D environment point cloud map model. The virtual-real rendering unit calculates the occlusion relationship between real objects and virtual devices in real time, and renders the shadow relationship and physical relationship to the virtual reality environment accordingly.
5. The job training system based on AR virtual vision according to claim 1, characterized in that, The adaptive evaluation module includes: The real-time operation data feedback unit, based on the set cloud skill database, generates feedback data in real time and transmits it to the smart AR terminal based on the real-time transmission of operation data; The operation trajectory analysis subunit is used to perform time-series segmentation and separation of the overall operation data based on the transmission, decomposing it into operation data for the preparation stage, execution stage, and verification stage. The timing logic scoring subunit is used to set weight coefficients for the preparation, execution, and verification phases, and output evaluation scores for each phase. The multi-dimensional skill map generation unit, based on the evaluation score output by the temporal logic scoring subunit, outputs a radar chart that includes spatial positioning accuracy, operational fluency, and anomaly handling capability, and then fits and generates an adaptive evaluation.
6. The job training system based on AR virtual vision according to claim 1, characterized in that, The culture program generation module includes: The training program generation unit, based on the generated adaptive evaluation, uses a deep learning model to dynamically arrange and generate training programs for corresponding trainees.
7. A training method for a job training system based on AR virtual vision as described in any one of claims 1-6, characterized in that... include: S1. Environment binding: Trainees wear smart AR terminals and scan the overall data of the training environment through the environment scanning module on the smart AR terminal. Then, they construct a three-dimensional environment point cloud map model through the environment perception construction module of the training system platform. S2. Virtual-Real Integration: The training system platform is based on a constructed 3D environment point cloud map model. It integrates and overlays virtual devices in the real scene, binds the virtual device model to the real desktop coordinate system, and calculates the occlusion relationship between real objects and virtual devices in real time. S3. Operation detection: Trainees wear smart AR terminals and, during simulated practical operations, the data collection module on the smart AR terminal collects the trainees' operation data and transmits it to the training system platform. S4. Operation feedback: Based on the transmitted operation data, after real-time feedback calculation, the feedback signal is transmitted to the smart AR terminal, which then converts it into a prompt signal to alert the student. S5. Dynamic evaluation: The training system platform receives and transmits operational data, performs operational evaluation based on the approved evaluation criteria, and outputs adaptive evaluation. S6. Training program generation: Based on adaptive evaluation and operational data, a personalized training program is generated by matching the dynamic strategy library, and the corresponding training program for each trainee is dynamically arranged.
Citation Information
Patent Citations
Display method and apparatus
CN107066082A
Physical education method and system based on multimedia and holographic AR fusion
CN120335618A
Visual SLAM method and device, equipment and storage medium
CN120526112A