An artificial intelligence-based automatic examination and evaluation method for practical training process

CN122819980APending Publication Date: 2026-09-25BEIJING JIYUAN ZHIHANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610901641.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]鉴于现有技术的上述缺点、不足,本申请提供一种基于人工智能的实训过程自动考评方法,其解决了现有技术中实训过程评价依赖单一数据源、视频信息与操作终端数据之间缺乏有效关联导致考评结果准确性与客观性不足的技术问题

Benefits of technology

[0008]本申请实施例提供的基于人工智能的实训过程自动考评方法,通过对用户实训过程中的视频数据与操作终端数据进行协同采集与联合分析,能够有效克服单一数据源在行为刻画上的局限性,提高对用户实际操作过程的完整感知能力。进一步地,通过对视频数据进行关键帧提取,并结合目标检测与跟踪及时序动作识别处理,将连续的视频信息转化为结构化的用户操作动作信息,从而实现对用户操作行为的精准解析与时序表达,即使在光照变化、遮挡或动作连续性较强的情况下,仍能够保持对关键操作行为的稳定识别能力。同时,通过对操作终端数据进行数据清洗、标准化处理及一致性验证处理,能够有效剔除冗余、异常及不一致数据,提高终端操作信息的可靠性,并保证不同来源数据在时间维度与语义层面的对齐一致性,为后续融合分析提供稳定的数据基础。在此基础上,通过将用户操作动作信息与终端操作信息进行关联匹配,生成结构化操作序列,实现了多源异构数据在统一结构下的表达与融合,使得用户实训过程由非结构化的多模态信息转化为可计算、可比对的序列化操作数据,从而显著提升行为建模的规范性与可分析性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819980A_ABST
    Figure CN122819980A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, in particular to an artificial intelligence-based automatic examination and evaluation method for practical training processes, which comprises the following steps: collecting video data and operation terminal data in a user's practical training process; extracting key frames from the video data, performing target detection and tracking based on the extracted key frames, and performing time sequence action recognition based on the target detection and tracking results to obtain user operation action information; performing data cleaning, standardization processing and consistency verification processing on the operation terminal data to obtain terminal operation information; performing association matching based on the user operation action information and the terminal operation information to generate a structured operation sequence corresponding to the user's practical training process; obtaining operation step examination results, operation standardization examination results and result correctness examination results based on the structured operation sequence; and obtaining a user practical training examination and evaluation score according to the operation step examination results, the operation standardization examination results and the result correctness examination results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an automatic evaluation method for practical training processes based on artificial intelligence. Background Technology

[0002] In existing technologies, the assessment of practical training processes mainly relies on manual observation and recording or automatic scoring methods based on a single data source, such as evaluation through on-site teacher scoring, video playback analysis, or simple rule matching based on operation logs. These methods generally suffer from problems such as high subjectivity, low efficiency, and poor evaluation consistency, making it difficult to meet the needs of large-scale practical training for efficient, objective, and standardized assessment. Summary of the Invention

[0003] (a) Technical problems to be solved

[0004] In view of the above-mentioned shortcomings and deficiencies of the prior art, this application provides an automatic evaluation method for the training process based on artificial intelligence, which solves the technical problems in the prior art where the evaluation of the training process relies on a single data source and the lack of effective correlation between video information and operation terminal data leads to insufficient accuracy and objectivity of the evaluation results.

[0005] (II) Technical Solution

[0006] To achieve the above objectives, the main technical solution adopted in this application includes: This application provides an automatic evaluation method for practical training processes based on artificial intelligence, comprising: collecting video data and operational terminal data during user practical training; extracting keyframes from the video data, performing target detection and tracking based on the extracted keyframes, and performing temporal action recognition based on the target detection and tracking results to obtain user operation action information; performing data cleaning, standardization, and consistency verification processing on the operational terminal data to obtain terminal operation information; performing association matching between the user operation action information and the terminal operation information to generate a structured operation sequence corresponding to the user practical training process; obtaining operation step assessment results, operation standardization assessment results, and result correctness assessment results based on the structured operation sequence; and obtaining the user practical training evaluation score based on the operation step assessment results, operation standardization assessment results, and result correctness assessment results.

[0007] (III) Beneficial Effects

[0008] The AI-based automatic evaluation method for practical training provided in this application effectively overcomes the limitations of single data sources in behavioral characterization by collaboratively collecting and analyzing video data and operational terminal data during user training, thus improving the ability to fully perceive the user's actual operation process. Furthermore, by extracting keyframes from video data and combining them with target detection, tracking, and temporal action recognition, continuous video information is transformed into structured user operation action information. This enables accurate analysis and temporal representation of user operation behavior, maintaining stable recognition of key operational behaviors even under conditions of changing lighting, occlusion, or strong action continuity. Simultaneously, by performing data cleaning, standardization, and consistency verification on operational terminal data, redundant, abnormal, and inconsistent data are effectively eliminated, improving the reliability of terminal operation information and ensuring alignment consistency of data from different sources in both time and semantic dimensions, providing a stable data foundation for subsequent fusion analysis. Based on this, by associating and matching user operation information with terminal operation information, a structured operation sequence is generated, realizing the expression and fusion of multi-source heterogeneous data under a unified structure. This transforms the user training process from unstructured multimodal information into computable and comparable serialized operation data, thereby significantly improving the standardization and analyzability of behavior modeling. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating an AI-based automatic evaluation method for practical training processes according to an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of the structure of an AI-based automatic evaluation system for practical training processes according to an embodiment of this application. Detailed Implementation

[0011] In related technologies, the automatic evaluation and grade generation for practical training or assessment processes typically relies on action recognition and scoring schemes based on single video analysis. These schemes primarily use cameras to capture video of students' training processes, employing object detection, pose estimation, or behavior recognition models to directly identify operational actions and compare them with standard operating procedures to generate a score. While this approach can automate evaluation to some extent, the lack of synchronous constraints on data from the operating terminal (such as device commands, parameter settings, and execution feedback) makes it prone to action recognition ambiguities when relying solely on visual information. For example, misjudgments can occur due to occlusion, changes in perspective, or similar tools. Furthermore, it is difficult to accurately determine whether the operation was actually effective, thus affecting the reliability of the evaluation results. To address this, the AI-based automatic evaluation method for practical training provided in this application introduces a collaborative processing mechanism for video data and operational terminal data. This enables simultaneous acquisition of multi-source information during the data collection phase. Furthermore, it constructs refined user operation information through keyframe extraction, target detection, and temporal action recognition. Simultaneously, it performs data cleaning, standardization, and consistency verification processing on terminal operation data to generate highly reliable terminal operation information. Further, by associating and matching visual and terminal information, it generates structured operation sequences, achieving a unified expression of the practical training process. Finally, based on the structured operation sequences, it comprehensively evaluates the process from three dimensions: operation steps, operation standardization, and result correctness, generating evaluation scores. This effectively improves the consistency and accuracy of practical training evaluation results, enabling stable and reliable automated evaluation even in complex training environments.

[0012] Figure 1 This is a flowchart illustrating an AI-based automatic evaluation method for practical training processes according to one embodiment of this application. Figure 1 As shown, this AI-based automatic evaluation method for practical training includes:

[0013] S1 collects video data and data from the user's operating terminal during practical training. For example, in industrial equipment assembly or CNC equipment operation training, students need to complete operations such as equipment inspection, parameter setting, startup, and result verification in sequence according to standard operating procedures. During actual training, video data of the student's operating area is continuously collected through a camera, while the user's command input, parameter adjustment, and equipment feedback information on the equipment control interface are recorded in real time through the operating terminal, thus forming multi-source data containing visual information and terminal interaction information.

[0014] S2, keyframe extraction is performed on the video data. Target detection and tracking are then performed based on the extracted keyframes, and temporal action recognition is performed based on the target detection and tracking results to obtain user operation action information. For example, when a student switches from "selecting tools" to "installing parts" or "adjusting equipment parameters," frames with significant changes are automatically selected as keyframes. Target detection and tracking are then performed based on these keyframes, such as recognizing changes in the position of the wrench, screwdriver, workpiece, and the operator's hand movements. Subsequently, these actions are modeled using time-series information to identify continuous operation sequences such as "taking tools—aligning with the installation point—tightening bolts."

[0015] S3, perform data cleaning, standardization, and consistency verification on the operation terminal data to obtain terminal operation information. Specifically, in terms of operation terminal data processing, the device log data is cleaned and standardized, such as removing duplicate instruction records, correcting abnormal timestamps, and performing consistency verification on device status feedback. For example, when a user enters the command "start device" on the terminal, it is checked whether the device has actually entered the operating state. If there is a delay or abnormality, it is marked and corrected, thereby obtaining more reliable terminal operation information.

[0016] S4, Based on the user operation action information and the terminal operation information, perform association matching to generate a structured operation sequence corresponding to the user training process;

[0017] S5. Based on the structured operation sequence, obtain the operation step assessment results, operation standardization assessment results, and result correctness assessment results. For example, in the "device parameter setting" step, if the user visually completes the operation, but the terminal data does not show that the parameters are written correctly, it will be determined that there is a result correctness deviation in this step; in the "tool use" process, if the action sequence is correct but the operation action is not standardized (such as not inserting the tool at the standard angle), the operation standardization score will be reduced; and if there are missing steps or redundant operations, points will be deducted in the operation step assessment.

[0018] S6. Based on the assessment results of the operation steps, the operation standardization, and the correctness of the results, obtain the user's practical training evaluation score. For example, in a piece of equipment assembly training, although a student completed all the operation steps, there were two instances of improper tool use and one instance of incorrect parameter setting. The corresponding points will be deducted, and a more objective practical training evaluation score will be output in the end.

[0019] In this embodiment of the application, keyframes are extracted from the video data, target detection and tracking are performed based on the extracted keyframes, and temporal action recognition is performed based on the target detection and tracking results to obtain user operation action information, specifically including:

[0020] The video data is parsed to extract the video frame sequence. For example, continuously recorded video data during user training is decoded at fixed time intervals or frame-by-frame to obtain a set of image frames arranged in chronological order. For instance, in a CNC machine tool training operation, if the camera records the trainee's operation at a frequency of 30 frames per second, a 1-minute video will be broken down into 1800 static images. Each frame contains a corresponding timestamp and image information, such as hand position, tool status, and changes in the equipment interface, thus providing basic data for subsequent keyframe selection and motion recognition.

[0021] Based on preset keyframe filtering rules, keyframes are selected from the video frame sequence. In one specific embodiment of this application, during welding training, the image may remain stable for most of the time, with significant changes only occurring when the welding torch contacts the workpiece, initiates the arc, or moves. Keyframes such as the moment the welding torch ignites and the moment the weld forms can be selected by using image feature differences (such as changes in structural similarity index (SSIM)) or changes in target detection results, thereby reducing computational load and highlighting key operational nodes.

[0022] Based on the keyframes, the training tools, objects being operated on, and user actions are identified, and these actions are continuously tracked to obtain target detection and tracking results. For example, in electronic assembly training, targets such as "soldering iron," "circuit board," "solder wire," and "operating hand" can be identified, and these targets can be correlated across frames using algorithms such as Kalman filtering or DeepSORT to form continuous behavioral trajectories such as "soldering iron moving to solder joint" and "hand holding component placement."

[0023] Based on the target detection and tracking results, an action time series is constructed, and the sequence of actions, duration of actions, and changes in action state are identified to obtain temporal action recognition results. Specifically, the process of constructing an action time series based on the target detection and tracking results and identifying the sequence of actions, duration of actions, and changes in action state refers to the structured organization of discrete target trajectory information according to the time dimension, thereby reconstructing the temporal logic of the operation process. For example, in medical simulation injection training, actions such as "disinfecting the skin → drawing out the medication → expelling air → injecting" can be identified, and a time axis can be constructed based on the start and end times of each action. At the same time, the duration of each action (e.g., the disinfection action lasts 8 seconds, the injection action lasts 5 seconds) and whether there are interruptions, repetitions, or sequence errors in the state changes can be analyzed.

[0024] User operation action information is generated based on the time-series action recognition results. For example, in industrial robot programming training, the final output user operation action information may include: "Action 1: Start equipment, time 00:01:05, sequence 1, status normal; Action 2: Import program, time 00:02:10, sequence 2, status abnormal (parameter error); Action 3: Execute, time 00:03:20, sequence 3, status normal." This information uses a unified structure to represent the action category, occurrence time, execution sequence, and execution status, thereby providing standardized input for the subsequent scoring model. The user operation action information includes the action category, action occurrence time, action execution sequence, and action execution status.

[0025] Optionally, in some embodiments of this application, keyframes are selected from the video frame sequence based on preset keyframe filtering rules, specifically including:

[0026] Feature extraction is performed on each video frame using a pre-trained visual feature extraction network to obtain corresponding video frame feature vectors. These feature vectors characterize the visual features of the training tools, objects being manipulated, and areas of human movement within the video frame. Specifically, an example is given of "electronic soldering training (soldering chips and resistors on a circuit board)." After frame-by-frame analysis of the soldering training video, each frame is input into a pre-trained visual model (e.g., a feature extraction network based on ResNet or VisionTransformer) to transform the image into a high-dimensional feature vector. In this example, when the camera captures a student holding a soldering iron close to the circuit board, the feature vector of that frame will significantly contain information such as "metal tool outline," "hand movement area," and "circuit board texture." During the static waiting phase, the feature vectors primarily reflect the background and static structural features of the equipment. These feature vectors provide the basic representation for subsequent change analysis.

[0027] The visual change intensity Vt is obtained by calculating the feature difference based on the feature vectors of adjacent video frames. The visual change intensity Vt is obtained by calculating and normalizing the distance between the feature vectors of the t-th frame and the (t-1)-th frame. For example, during the soldering process, if the trainee has just moved the soldering iron from the stand to the solder joint area between the 100th and 101st frames, the feature vectors of the two frames will have a large difference, and a higher Vt value will be obtained by calculating the Euclidean distance or cosine distance. In the static observation stage, such as when the hand is not moving, the features of adjacent frames are almost the same, so Vt is close to 0, thus reflecting whether "action has occurred".

[0028] The training tools and objects in each video frame are detected using a target detection model to obtain target category information, target location information, and target state information. The target category information represents the category to which the target belongs, the target location information represents the target's spatial position in the video frame, and the target state information represents the target's current state attribute. Specifically, a target detection model (such as the YOLO series model) is used to identify specific objects in the image. In this example, targets such as "soldering iron," "solder wire," "circuit board," and "hand" can be identified, and their position coordinates (e.g., bounding box) and state information (e.g., "heating / not heating" for the soldering iron, "in contact / not in contact" for the solder wire) can be output. For example, when the soldering iron contacts the solder joint, the state information changes from "suspended" to "in contact," thus providing a basis for judging the change in action.

[0029] The target position and state information of the same target in adjacent video frames are compared. When the target position or state information changes, a target state change flag Ot is generated, where Ot is set to 1; otherwise, Ot is set to 0. For example, between frames 120 and 121, if the soldering iron is detected to move from the left to the solder joint area, or the state changes from "not in contact" to "in contact," then Ot is set to 1; if there is no significant change in the target position or state, then Ot is 0. In welding scenarios, this flag can effectively capture the "moment of critical operation."

[0030] Based on the operation command records and execution feedback records corresponding to the timestamps of each video frame in the operation terminal data, the operation command trigger type, operation command trigger time, and execution feedback status information are extracted and normalized to obtain the action trigger signal strength At. The execution feedback status information includes successful operation status, failed operation status, and abnormal status. For example, during soldering, when a trainee clicks the "Heat Soldering Iron" button on the terminal, the trigger time of the command and the execution feedback (success or failure) are recorded. If both "heating command + success feedback" occur at a certain moment, At is higher; if there is no command or execution fails, At is lower. This allows for the association and enhancement of "actual operation behavior signals" with "visual actions."

[0031] The process involves obtaining the action recognition results output by a pre-trained action recognition model for the video frame sequence. The pre-trained action recognition model is a deep learning-based temporal action recognition model used to identify action categories and predict action time intervals within the video frame sequence. In this embodiment, obtaining the action recognition results output by the pre-trained action recognition model for the video frame sequence refers to using a temporal deep learning model (such as TSN, SlowFast, or Transformer structure) as the action recognition model to classify and segment the entire video. For example, in welding training, the model might identify actions such as "preparing to weld," "heating the soldering iron," "welding contact," and "cooling end," and output the start and end times of each action (e.g., "heating the soldering iron: 00:01:10–00:01:20") and a confidence score (e.g., 0.92) to describe the overall structure of the action in the time dimension. The action recognition results include the action category, action start time, action end time, and action confidence score.

[0032] Obtain standard annotation data, which is a dataset formed by manually annotating training videos based on standard operating procedures. The standard annotation data includes at least the standard action category, the standard action sequence, and the standard action time interval. For example, in this soldering training, the standard procedure might specify "Step 1: Check equipment → Step 2: Heat the soldering iron → Step 3: Apply solder → Step 4: Complete the solder joint." This data not only includes action categories but also a strict standard execution sequence and time interval, serving as a benchmark reference for model evaluation.

[0033] Based on the alignment and comparison of the action recognition results with the standard labeled data, the action recognition accuracy Pt is calculated, where Pt is the ratio of the number of correctly recognized actions to the total number of recognized actions. The process of calculating the action recognition accuracy Pt involves determining whether the model's predicted actions match the standard actions one-to-one. For example, if the action recognition model identifies 10 actions, and 8 of them match the standard actions, then Pt = 8 / 10 = 0.8. If there are misidentifications, such as misidentifying "moving a soldering iron" as "soldering," these are counted as incorrect identifications, thus affecting the accuracy.

[0034] Based on the alignment and comparison between the action recognition results and the standard labeled data, the action recognition recall rate Rt is calculated, where Rt is the ratio of the number of correctly recognized actions to the total number of standard actions in the standard labeled data. For example, if the standard process specifies 12 actions, but the action recognition model only recognizes 10, of which 8 are correct matches, then Rt = 8 / 12 ≈ 0.67, which is used to reflect the missed detection situation.

[0035] Based on the visual change intensity Vt, target state change identifier Ot, action trigger signal intensity At, action recognition precision Pt, and action recognition recall Rt, the keyframe importance score St corresponding to each video frame is calculated, where the keyframe importance score St corresponding to the t-th frame satisfies:

[0036] St = α × (wv × Vt + wo × Ot + wa × At) + β × (Pt × Rt) / (Pt + Rt); where α and β are weighting coefficients and satisfy α + β = 1; wv, wo, and wa are the weighting coefficients corresponding to the intensity of visual change, the indicator of target state change, and the intensity of action trigger signal, respectively.

[0037] Video frames whose keyframe importance score St is greater than or equal to a preset threshold T are identified as keyframes. Specifically, this means automatically selecting the most representative operation frames based on a set threshold. For example, in the entire soldering video, only a few frames, such as "soldering iron contacting the solder joint," "the moment of applying tin," and "solder joint formation," have a St exceeding the threshold T and are therefore selected as keyframes, while a large number of static or repetitive frames are filtered out. This achieves data compression and retention of key actions, ensuring the efficiency and accuracy of subsequent action recognition and evaluation.

[0038] In practical applications of this application, based on the keyframe recognition training tool, the object being operated on, and user actions, and by continuously tracking the training tool, the object being operated on, and user actions, target detection and tracking results are obtained, specifically including:

[0039] Multi-scale feature extraction is performed on the keyframes to obtain target feature information, which is used to characterize the visual features of the training tool, the object being operated on, and the area corresponding to the user's actions. Specifically, the training process of "an industrial robot grasping a motor and completing assembly" is analyzed. A camera records the entire operation process of the trainee, while the operating terminal records the robot's control commands and execution feedback information. Specifically, feature encoding is performed on the keyframe images at different scales to simultaneously capture local details and overall structural information. For example, in the keyframe of grasping the motor, small-scale features can be used to identify details such as screw holes and gripper edges, while large-scale features can be used to identify the overall posture of the robotic arm and the workstation layout, thus forming a more complete visual representation that can distinguish between "grasping actions" and "empty movement actions."

[0040] Based on the target feature information, a pre-trained target detection model is used to identify and locate the training tool, the object being operated on, and the corresponding area of ​​the user's actions in each keyframe, obtaining target detection results. These results include target category information, target location information, and target state information. For example, a target detection network (such as YOLO or Faster R-CNN) can be used to identify targets in keyframes. For instance, in a certain frame, the output might include: the robotic arm position (x1, y1, x2, y2), the motor component position (x3, y3, x4, y4), and the gripper state (open / closed), identifying the current action state as "approaching the target material." This information collectively constitutes the target detection result, describing "who is in what position and what state."

[0041] Based on the target detection results, the target position information and target state information corresponding to the same target in adjacent keyframes are compared to obtain target cross-frame change information, which includes changes in target position information and changes in target state information. For example, between frame 50 and frame 51, the robotic arm end effector moves from "20cm from the motor" to "5cm from the motor", and at the same time, the gripper state changes from "open" to "semi-closed". Based on this, cross-frame change information is generated and marked as "significantly approaching the target + change in gripping preparation state".

[0042] Based on the target detection results and the target cross-frame change information, the correspondence of the same target in the time series is correlated and matched to generate target trajectory information. This target trajectory information includes the target position trajectory and the target state trajectory, used to characterize the continuous change process of the target in the time dimension. For example, a robotic arm forms a continuous trajectory throughout the grasping process: from the initial position → moving above the motor → descending → gripping → lifting. This trajectory includes not only the position change path but also the state change (opening → closing → stable gripping), thus constructing a complete behavioral chain of "target position trajectory + state trajectory".

[0043] Based on the target detection results and the target trajectory information, a visual detection result is generated. The visual detection result is used to characterize the temporal visual detection information after the target detection results and target trajectory information are fused. The visual detection result includes target category sequence, target position trajectory, target state trajectory, and action time information. For example, in the process of grasping a motor, the output within a certain time period is: target category is "motor", position trajectory is "moved from pallet to assembly table", state trajectory is "not gripped → gripped → stable", and time information is attached (such as 00:01:10–00:01:25), forming a complete temporal visual detection result, thereby fully describing the operation process.

[0044] The system acquires operation terminal data corresponding to the action time information and performs matching verification on the visual detection results based on the operation terminal data. The matching verification includes: consistency comparison between the target category sequence and the operation instruction record, consistency comparison between the target state trajectory and the execution feedback record, and synchronization error verification between the action time information and the corresponding time information of the operation terminal data.

[0045] In this embodiment, the process of acquiring the operation terminal data corresponding to the action time information and matching and verifying the visual detection result based on the operation terminal data refers to aligning and verifying the video recognition result with the robot control system log. For example, when the vision system recognizes the time of "robotic arm gripping motor" as 00:01:15, and the operation terminal record shows that the "gripper closing command" was executed at that moment, with the execution feedback being "successful," then the two are consistent. If there is a discrepancy between visual recognition and command (such as the vision showing the gripper closing but the terminal not issuing a command), it is marked as abnormal or to be corrected. The matching verification includes a consistency comparison between the target category sequence and the operation command record, a consistency comparison between the target state trajectory and the execution feedback record, and a synchronization error verification process between the action time information and the corresponding time information of the operation terminal data. This refers to confirming consistency from three dimensions. For example: ① The target category sequence should be consistent with commands such as "grip motor / move robotic arm"; ② The gripper state trajectory should be consistent with the "opening and closing feedback"; ③ The error between the visual action time and the command time should be controlled within a range such as ±200ms, thereby ensuring reliable synchronization of multi-source data.

[0046] If the matching verification passes, the visual detection result undergoes a time-based structured reconstruction process to obtain a reconstructed visual detection result. The structured reconstruction process includes:

[0047] Establishing a time correspondence between target trajectory information and operation terminal data based on action time information means binding the robotic arm's movement trajectory with control commands point by point in time, such as "00:01:10 corresponds to movement command A" and "00:01:15 corresponds to gripper closing command B", thereby establishing a unified time coordinate system.

[0048] Establishing a correspondence between target category information and operation instruction records refers to semantically binding visually recognized categories such as "motor" and "robotic arm" with terminal instructions such as "grab motor" and "move to assembly position," thereby achieving a one-to-one mapping between "visual objects and control instructions."

[0049] To complete and correct the mismatched target trajectory information and corresponding time information, it means that when the vision loses a frame of the robotic arm position due to occlusion, the position is inferred by using the previous and next trajectories and terminal control commands. For example, the missing frame can be recovered by linear interpolation or trajectory prediction to make the trajectory continuous and unbroken.

[0050] Based on the matching results, the target position trajectory, target state trajectory, and motion timing information are corrected to obtain the corrected structured visual inspection results. For example, during a robot's grasping of a motor, visual detection may cause a temporary gap in the position of the end effector of the robotic arm due to occlusion, while the corresponding moment in the terminal data contains a clear motion control command. In this case, the missing visual trajectory is completed based on the terminal command trajectory, and the time offset is synchronously corrected to ensure that the visual trajectory and the control trajectory are consistent under the same time reference. Furthermore, based on the matching results, the target position trajectory, target state trajectory, and motion timing information are corrected to obtain the corrected structured visual inspection results. This correction process includes not only spatial position correction but also alignment correction of state change moments. For example, the "gripper closure state" is adjusted from a visual detection lag time to be consistent with the control command trigger time, thereby improving the consistency of motion timing.

[0051] Based on the corrected structured visual detection results, information is extracted and structured encapsulated to generate target detection and tracking results. Specifically, action category information, action start time, action end time, and action status information are extracted from the corrected structured visual detection results and organized according to a unified data structure. For example, in a motor grasping task, the final output result is: action category is "grab motor", action start time is 00:01:10, action end time is 00:01:25, action status is "execution successful", and the corresponding target trajectory information represents the entire process of the robotic arm moving from the initial position to the assembly station and completing the gripping.

[0052] The target detection and tracking results are used to characterize the unified temporal behavior results after fusion and correction based on visual detection information and operation terminal data, including action category information, action start time, action end time, and action state information.

[0053] Optionally, in some embodiments of this application, an action time series is constructed based on the target detection and tracking results, and the sequence of action occurrence, duration of action, and changes in action state are identified to obtain a temporal action recognition result, specifically including:

[0054] Obtain the action category information, action start time, action end time and action status information corresponding to each action in the target detection and tracking results, wherein the action status information includes at least the action completion status, action abnormal status and action interruption status.

[0055] In this embodiment, taking "students performing circuit breaker wiring operations in an electrical engineering training course at a university" as an example, video data from the training classroom camera and instruction log data from the training platform's operating terminal are collected simultaneously for automatic evaluation and analysis of the student's operation process. For example, after fusing video recognition with terminal logs, it is identified that the student sequentially performed actions such as "removing the circuit breaker," "fixing the guide rail," "connecting the input terminal," "tightening the wiring," and "power-on test." Among them, the "removing the circuit breaker" action started at 10:01:03 and ended at 10:01:08, with a completed status; the "tightening the wiring" action started at 10:02:15, but the screwdriver was detected to have slipped midway, and it was marked as an "abnormal state"; the "power-on test" action was marked as an "interrupted state" because the equipment did not respond during execution. This information together constitutes the structured target detection and tracking results.

[0056] The corresponding action time interval is calculated based on the start and end times of each action. All actions are then sorted by time using the start time as the time base to generate an action time sequence. For example: 10:01:03 Circuit breaker activation → 10:01:20 Rail fixing → 10:01:45 Input terminal connection → 10:02:15 Wiring tightening → 10:03:00 Power-on test, thus forming a complete action time sequence. The duration of each action is also calculated simultaneously. For example, "wiring tightening" lasts approximately 40 seconds, while the standard operation should be around 25 seconds, providing a basis for subsequent deviation analysis.

[0057] Based on the action time sequence and action category information, the temporal relationship between each action is determined, and an action execution sequence is generated to represent the actual action execution order. For example, the standard procedure requires "fixing the guide rail → connecting the input terminal → tightening the wiring → power-on test," but in practice, it was detected that students delayed "connecting the input terminal" for a long time after "fixing the guide rail," and in some cases, partial tightening was performed before connection. Therefore, the generated execution sequence is marked with "sequence offset."

[0058] Obtain standard action information corresponding to the structured operation sequence, wherein the standard action information includes standard action category information and standard action time constraint information in the structured operation sequence. Specifically, during the "obtaining standard action information" process, the standard procedure corresponding to the experiment is retrieved from a preset training standard library, including standard action categories (5 steps in total), standard action sequence constraints, and standard time ranges (e.g., wiring tightening should be completed within 20-30 seconds, and power-on testing should be performed after all connections are completed). This standard action information serves as a reference for subsequent comparison.

[0059] The action time series is matched and analyzed with the standard action information to generate action matching relationship and matching deviation results. The matching analysis includes consistency matching based on action category information and standard action category information, and calculation of time deviation between action time interval and standard action time constraint.

[0060] Based on the matching relationship and the matching deviation result, the sequence deviation value, execution deviation value, missing action identifier and redundant action identifier of each action are calculated. The execution deviation value is calculated based on the action status information and time deviation.

[0061] The sequence action recognition result is generated based on the sequence deviation value, execution deviation value, missing action identifier and redundant action identifier.

[0062] The temporal action recognition results include: action sequence correctness information, which characterizes the consistency between the actual action execution order and the standard action order in the structured operation sequence, and is quantitatively represented based on the sequence deviation value; action standardization information, which characterizes the standardization of a single action during execution, and is quantitatively calculated based on the execution deviation value and action status information; and process integrity information, which characterizes the coverage consistency between the actual set of executed actions and the standard set of actions in the structured operation sequence, and is calculated based on the number of missing actions and the number of redundant actions.

[0063] In this embodiment, during the process of "matching and analyzing the action time sequence with the standard action information," taking the "wiring tightening" action as an example, firstly, the standard execution time interval of the action is obtained according to the preset standard operating procedure, and the midpoint reference time of this interval is used as the standard benchmark value. Simultaneously, the actual occurrence time of the action during user training is determined based on the actual detection results, and compared with the standard benchmark time to obtain the time deviation value of the action. Since the actual execution time of the action is significantly later than the standard benchmark time, this time difference is recorded as a positive offset.

[0064] Building upon this, the time deviation is further normalized by calculating the ratio of the actual deviation to the preset maximum allowable deviation range, thus obtaining a standardized time deviation value. Simultaneously, the continuity and stability of the action are analyzed by incorporating detected state changes during execution. For example, when pauses, repetitive operations, or discontinuous actions are detected, this persistent anomaly is converted into a state deviation index, which is then fused with the time deviation index to obtain the comprehensive execution deviation value for the action. This comprehensive execution deviation value characterizes the overall degree of deviation of the action in both the time and execution quality dimensions.

[0065] In the sequence deviation calculation process, the standard execution order of each action is first determined according to the standard operating procedure, and the actual identified action sequence is aligned and compared with the standard sequence. The difference between the position of each action in the actual sequence and its position in the standard sequence is calculated, and normalized by incorporating the total number of standard actions, thus obtaining the sequence deviation value for each action. Simultaneously, when an action is detected to be executed concurrently with other actions or its order is partially reversed during actual execution, an additional sequence penalty factor is introduced to reflect the degree of structural interference between actions. The final sequence deviation value is used to characterize the overall deviation between the actual operating procedure and the standard procedure.

[0066] In the process of identifying redundant and missing actions, the actual set of detected actions is compared and analyzed with the standard set of actions. When there are additional actions in the actual set that are not defined in the standard process, these actions are marked as redundant, and a redundancy index is calculated based on the ratio of the number of redundant actions to the total number of standard actions. When some actions in the standard set are not detected during actual execution, they are identified as missing actions, and a missingness index is calculated. By comprehensively statistically analyzing the redundancy and missing situations, the degree of complete execution of the standard process by the user during practical training can be further reflected.

[0067] In the process of "generating sequential action recognition results based on the sequence deviation value, execution deviation value, and missing and redundant action identifiers," the three types of indicators are normalized, and corresponding evaluation information is generated. Specifically, the action sequence correctness information is obtained by inverse mapping of the sequence deviation value, used to characterize the consistency between the actual action execution sequence and the standard process; the action standardization information is obtained by inverse mapping of the execution deviation value, used to characterize the standardization of a single action during execution; and the process integrity information is calculated by combining the proportion of missing actions and the proportion of redundant actions, used to characterize the coverage consistency between the actual set of executed actions and the standard set of actions.

[0068] Finally, the three types of evaluation information are integrated and processed to form a complete temporal action recognition result, thereby achieving a unified quantitative evaluation of the user's training operation in three dimensions: action execution sequence, execution standardization, and process integrity, and providing a structured input basis for subsequent training assessment.

[0069] Optionally, in some embodiments of this application, generating user operation action information based on the timing action recognition result specifically includes:

[0070] The sequence action recognition results are used to obtain the correctness information of the action sequence, the standardization information of the action, and the completeness information of the process for each action. In this embodiment, taking the "circuit breaker control circuit wiring operation in an electrical engineering training task" as an example, the video analysis and terminal data fusion processing have been completed in the previous steps, and the sequence action recognition results containing multiple actions are output, such as standard operation actions like "retrieving the circuit breaker", "installing the guide rail", "connecting the input terminal", "tightening the wiring", and "power-on test" and their deviation analysis results. First, in the process of "obtaining the correctness information of the action sequence, the standardization information of the action, and the completeness information of the process for each action in the sequence action recognition results", structured information is extracted for each identified action. For example, for the action "connecting the input terminal", its sequence correctness information is 0.92, indicating that the action basically conforms to the standard sequence but has a slight prematureness; its action standardization information is 0.88, indicating that the operation process is relatively standard but has a slight pause; its completeness information is 1, indicating that the action has no missing or redundant problems in the overall process. Similarly, for the "wiring tightening" action, the sequence correctness information is 0.75, the standardization information is 0.70, and the integrity information is 1.

[0071] Based on the start and end times of the actions used to construct the action time sequence from the time-series action recognition results, the occurrence time of each action is determined, where the occurrence time is the start time of the corresponding action. Specifically, the start time of each action is directly used as the occurrence time. For example, if the action "connecting input terminals" starts at 10:02:10, its occurrence time is recorded as 10:02:10; if "tightening wiring" starts at 10:02:55, its occurrence time is recorded as 10:02:55. In this way, all actions are uniformly mapped onto the timeline for subsequent sequential and timing analysis.

[0072] Based on the action execution sequence generated from the time-series action recognition results, the execution order of each action is determined. All actions are sorted according to their occurrence time. For example, the actual recognition order is: "Take circuit breaker → Install rail → Connect input terminal → Tighten wiring → Power-on test". Further comparison with the standard operating sequence reveals a slight overlap between "Connect input terminal" and "Tighten wiring," but the overall order remains largely consistent. Therefore, each action is assigned a corresponding execution sequence number (e.g., steps 1 to 5) for subsequent consistency checks.

[0073] Based on the correctness information of the action sequence, the standardization information of the action, and the integrity information of the process, the execution status of each action is determined, and the corresponding action execution status is generated. For example, for the "connect input terminal" action, because its sequence correctness is high and its standardization score is higher than the set threshold, it is determined to be in a normal execution state. However, for the "tighten wiring" action, because there is a significant operation delay and repeated tightening behavior occurs during the action, it is determined to be in an abnormal execution state.

[0074] The execution states of the action include normal execution state, abnormal execution state, omitted execution state, redundant execution state, and interrupted execution state, specifically:

[0075] When the action standardization information meets a preset standardization threshold, and the sequence deviation value corresponding to the action sequence correctness information meets the sequence consistency requirement, the corresponding action execution state is determined to be a normal execution state. For example, the action of "installing guide rails" has a standardized execution process and a correct sequence, so it is determined to be a normal execution state.

[0076] When the action standardization information is lower than the preset standard threshold, or the sequence deviation value corresponding to the action sequence correctness information does not meet the sequence consistency requirement, the action execution state of the corresponding action is determined to be an abnormal execution state. For example, the "wiring tightening" action is determined to be an abnormal execution state because the operation time exceeds the standard and there is a misoperation.

[0077] When it is determined, based on the process integrity information in the timing action recognition result, that a corresponding standard action exists in the structured operation sequence but is not matched in the action time sequence, the action execution status of the corresponding action is determined to be an omitted execution status. For example, if the "insulation detection" step is not executed, then the standard action is marked as an omitted execution status;

[0078] When the process integrity information determines that there are actions in the action time sequence that do not belong to the standard action information in the structured operation sequence, the corresponding action's execution state is determined to be a redundant execution state. For example, the "repeated screw pre-tightening" operation is determined to be a redundant execution state.

[0079] When the action status information in the timing action recognition result includes an action interruption status, the corresponding action's execution status is determined to be an interrupted execution status. For example, if a "power-on test" is prematurely terminated due to a device not responding, then the action is determined to be in an interrupted execution status.

[0080] The action category, action occurrence time, action execution sequence, and action execution status are correlated and integrated to generate user operation action information. For example, the following user operation action information records are generated: "Connecting input terminal: Occurrence time 10:02:10, execution sequence step 3, execution status normal"; "Tightening wiring: Occurrence time 10:02:55, execution sequence step 4, execution status abnormal"; "Power-on test: Occurrence time 10:03:40, execution sequence step 5, execution status interrupted". Through the above integration, a structured operation description data that fully reflects the entire process of user training behavior is finally formed, thus providing a unified data foundation for subsequent training scoring, capability assessment, and error analysis.

[0081] To illustrate further, consider the "Low-voltage power distribution line wiring experiment" in a certain electrical engineering training course. The standard operating procedure includes six actions in sequence: "installing the guide rail," "fixing the circuit breaker," "connecting the input terminals," "connecting the output terminals," "tightening the wiring," and "power-on test." The actual action time sequence is obtained through the aforementioned steps, and further information on the correctness of the action sequence, the standardization of the actions, and the completeness of the process is obtained for each action.

[0082] Taking the "connecting output terminals" action as an example, the standard procedure requires this action to be performed after "connecting input terminals" is completed, and the standard operation time should be controlled between 20 and 35 seconds. It also requires the wire to be fully inserted into the wiring hole before the clamping operation. The timing action recognition results show that the actual execution sequence of this action is consistent with the standard sequence, therefore the action sequence correctness score is 0.96; the action duration is 28 seconds, during which no abnormalities such as wire detachment or repeated adjustments were detected, therefore the action standardization score is 0.93; furthermore, this action is a necessary step in the standard procedure and there was no repeated execution, therefore the procedure integrity score is 1. The preset action standardization threshold is 0.85, and the sequence consistency threshold is 0.90. When both of these scores are higher than the corresponding thresholds and the procedure integrity is normal, the action is determined to meet the normal execution conditions, and the execution status corresponding to the "connecting output terminals" action is marked as the normal execution status.

[0083] Taking the "tightening wiring" action as an example, the standard procedure requires tightening after connecting the input and output terminals, with a standard duration of 15 to 25 seconds. Although the execution sequence was correct, achieving a sequence correctness score of 0.91, the student repeatedly slipped the screwdriver and tightened it repeatedly, causing the action duration to reach 48 seconds. Furthermore, the number of tool vibrations exceeded a preset threshold, resulting in an action standardization score of only 0.68, below the standardization threshold of 0.85. At this point, although the process integrity is normal, the action is judged as an abnormal execution state because the standardization score does not meet the requirements, and an anomaly marker is generated in the corresponding action record.

[0084] Furthermore, taking the "insulation test" action as an example, the standard operating procedure stipulates that an insulation test must be performed after the wiring is tightened. However, in the actual action sequence, the use of the insulation tester was not detected, nor was a corresponding test instruction record found in the operation terminal log. Therefore, the process integrity analysis results show that this action is missing. At this point, the "insulation test" step was found in the standard action set, but no corresponding action record was found in the actual action time sequence. Therefore, the execution state corresponding to this action was identified as an omitted execution state, and a missing action identifier was generated. Additionally, during the actual operation, after tightening the wiring, the student repeated the "re-tighten screws" action. Action category matching revealed that this action did not belong to the prescribed steps in the current standard operating procedure, and there was no corresponding item in the standard action set. Therefore, the process integrity analysis results determined that this action was an extra action. At this point, this action was identified as a redundant execution state, and a redundant action identifier was generated for subsequent scoring deductions. For example, during the "power-on test," after a student presses the start button, a wiring error triggers the equipment's protection device. The test start time is detected as 10:05:32, but the equipment automatically stops running at 10:05:36, and the operation terminal feedback log shows a "startup failed" status code. Combined with the video showing the student stopping the operation, it is determined that the action was not completed correctly, and the action status information includes an action interruption flag. Therefore, even if the action has already started, its execution status is still determined to be an interrupted execution state.

[0085] Finally, the above analysis results are integrated and unified. For example, the following user operation information is generated: "Connect output terminal": Action occurrence time 10:03:05, execution sequence step 4, execution status is normal execution status; "Tighten wiring": Action occurrence time 10:03:42, execution sequence step 5, execution status is abnormal execution status; "Insulation detection": Standard procedure requires it, but it was not actually executed, execution status is omitted execution status; "Retighten screws": Action occurrence time 10:04:35, not part of the standard procedure, execution status is redundant execution status; "Power-on test": Action occurrence time 10:05:32, execution sequence step 6, interrupted due to equipment malfunction, execution status is interrupted execution status. Through the above method, based on the correctness of the action sequence, the standardization of the action, and the integrity of the process, the execution status of each action can be finely judged, thus forming user operation information that includes action category, action occurrence time, action execution sequence, and action execution status, providing a reliable data foundation for subsequent structured operation sequence construction and automatic evaluation of training.

[0086] Optionally, in some embodiments of this application, the operation terminal data is subjected to data cleaning, standardization, and consistency verification to obtain terminal operation information, specifically including:

[0087] Acquire the operation terminal data generated by the operation terminal during the training process, wherein the operation terminal data includes operation instruction records, parameter setting records, device status records, and execution feedback records.

[0088] In this embodiment, student Zhang conducted a wiring training exercise for an intelligent power distribution cabinet, using a PLC control console and a host computer monitoring system as his operating terminal. During the training, data generated by the operating terminal was continuously collected. This included operation command records such as "turn on power module," "set output voltage to 24V," and "start detection program"; parameter setting records such as voltage parameters, operating mode parameters, and sampling period parameters; device status records such as relay status, output voltage status, and indicator light status; and execution feedback records such as "execution successful," "execution failed," "parameter out of bounds," and "communication abnormality." For example, at 10:05:32, the student issued the command "set output voltage to 24V"; at 10:05:33, the device responded "parameter setting successful"; at 10:05:35, the output voltage status changed to 24.1V; and at 10:05:36, the relay status switched to the closed state. All of the above data was saved as the original operating terminal data.

[0089] The operation terminal data undergoes data cleaning to remove duplicate, abnormal, and invalid data, and missing data is completed to obtain cleaned operation terminal data. For example, at 10:05:32, due to the network retransmission mechanism, the host computer continuously records two identical "set output voltage 24V" commands. The timestamps of the two records are only 20ms apart and their contents are completely identical, so the duplicate record is automatically deleted. Simultaneously, at 10:05:40, the device generates an abnormal data value of 500V due to communication interference. Since the device's normal operating range is 0-30V, this data is identified as abnormal and removed. Furthermore, an execution feedback record is found at 10:05:33, but the corresponding device status record is missing. Therefore, based on the status changes at adjacent time points 10:05:32 and 10:05:35, linear interpolation and status continuation rules are used to fill in the missing data, indicating the device was in a "voltage establishment" state at 10:05:33, thus obtaining complete cleaned operation terminal data.

[0090] The cleaned operation terminal data undergoes field unification and format conversion processing. The operation instruction records, parameter setting records, device status records, and execution feedback records are standardized to obtain standardized operation terminal data. For example, the time format in the PLC record is "2026 / 06 / 18 10:05:32," while the time format in the host computer log is "10:05:32.213," both are uniformly converted to millisecond-level Unix timestamp format; "execution success" is represented as "OK," "SUCCESS," and "1" in different devices, respectively, and is uniformly converted to the standard field "Success"; voltage values ​​in parameter fields are sometimes represented as "24V" and sometimes as "24000mV," and are uniformly converted to a data format with units of V. After the above processing, the operation instruction records, parameter setting records, device status records, and execution feedback records all have a unified data structure, forming standardized operation terminal data.

[0091] Based on the standardized operation terminal data, the operation instruction records, parameter setting records, device status records, and execution feedback records are associated in chronological order and sorted based on timestamps to generate terminal operation information. For example, the operation instruction "Set output voltage 24V" is recorded at 10:05:32.100, the execution feedback "Parameter setting successful" is recorded at 10:05:33.200, the device output voltage reaches 24.1V at 10:05:35.500, and the relay is closed at 10:05:36.000. By establishing a chronological association, the above records are integrated into a single operation event chain and arranged chronologically to generate terminal operation information. At this point, a single terminal operation message simultaneously includes: operation instruction record: "Set output voltage 24V", parameter setting record: target voltage = 24V, device status record: output voltage = 24.1V, relay status = closed, execution feedback record: execution successful, corresponding to a time range of 10:05:32 to 10:05:36. In this way, the originally scattered data is transformed into terminal operation information with temporal correlation.

[0092] Based on the terminal operation information, the consistency of the correspondence between the operation instruction record and the execution feedback record is verified, and the consistency of the correspondence between the parameter setting record and the device status record is verified.

[0093] The consistency verification includes: determining whether there is a corresponding execution feedback record between the operation instruction record and the execution feedback record; determining whether the parameter setting record and the device status record meet the consistency requirements of parameter setting changes and device status changes; and verifying the time sequence and time continuity between adjacent records in the terminal operation information.

[0094] Based on the consistency verification results, operation records that do not meet the consistency conditions are corrected or completed, and the terminal operation information is updated.

[0095] Specifically, after generating terminal operation information, consistency verification begins. First, it verifies the consistency between the operation command record and the execution feedback record. For example, at 10:06:10, the student executes the command "Start detection program," but receives no corresponding execution feedback. It is found that the command lacks a feedback record, so this event is marked as an "inconsistent event." The cache and log database are then queried, and a feedback message "Start successful" is found at 10:06:11 that was not read in time due to network latency. This feedback record is automatically added, and the consistency relationship is restored. Next, the system verifies the consistency between parameter setting records and device status records. For example, the student sets the output voltage to 24V at 10:08:20, but the device status record shows the actual output voltage is only 12V. The parameter change is calculated as ΔU = |24−12| = 12V. Since the preset allowable error range is ±1V, it is determined that the parameter setting change and the device status change are inconsistent. Further inspection reveals that the power module has not switched to high-voltage output mode, so this state is marked as abnormal, and the message "Parameter setting not effective" is recorded. Next, the time sequence and continuity of the terminal operation information are verified. For example, the normal sequence should be: voltage setting → successful execution feedback → output voltage establishment → relay closure. However, the log shows: 10:09:00 output voltage establishment; 10:09:03 voltage setting; 10:09:04 successful execution feedback. An abnormal time sequence is detected, so the time difference between adjacent records is calculated: Δt1 = 10:09:03 − 10:09:00 = -3s. Since the negative time interval does not meet the time continuity requirement, a time discrepancy is determined. The time is further reordered based on the cached logs and the device's internal time to restore the correct sequence. Based on the consistency verification results, abnormal records are corrected or supplemented. For example, a "start detection program" command is detected at 10:10:20, but no corresponding device status change is detected. Reading the device cache reveals that the relay was already closed at 10:10:22, so the missing status record is automatically supplemented. Meanwhile, for parameter anomaly events, "24V setting → 12V output" was corrected to "Parameter setting failed," and an anomaly feedback record was added. After correction, the event chain was updated to: Operation command: Set output voltage to 24V, Parameter setting: 24V, Execution feedback: Setting failed, Device status: Output voltage 12V, Reason for anomaly: High voltage mode not switched. Finally, after data cleaning, standardization, and consistency verification, stable and reliable terminal operation information was obtained. For example, for the operation "Set output voltage to 24V," the final terminal operation information is as follows: Start time: 10:05:32, Operation command record: Set output voltage to 24V, Parameter setting record: Target voltage 24V, Device status record: Output voltage 24.1V, Relay closed; Execution feedback record: Execution successful; Status indicator: Normal; End time: 10:05:36.The terminal operation information has eliminated duplicate records, abnormal data, and time discrepancies, and established a correspondence between operation instructions, device status, and execution feedback. This allows it to serve as a basis for subsequent time alignment and correlation matching with user operation information on the video side, further generating structured operation sequences and enabling automatic evaluation of the training process. This demonstrates that this application, through cleaning, standardizing, and verifying the consistency of the operation terminal data, can effectively improve the integrity, accuracy, and temporal consistency of the terminal data, providing a reliable data foundation for subsequent multi-source information fusion and automatic scoring, and is practically feasible.

[0096] In some embodiments of this application, a structured operation sequence corresponding to the user training process is generated by associating and matching the user operation information with the terminal operation information, specifically including:

[0097] The system obtains the action category, action occurrence time, action execution sequence, and action execution status from the user operation action information. For example, in the intelligent power distribution cabinet wiring training, students complete operations such as "connecting input terminals," "tightening wiring," "setting output voltage," and "starting the detection program" in sequence according to the training requirements. After the aforementioned video analysis and processing, the user operation action information is obtained. Among them, the action category corresponding to the "connecting input terminals" action is "connecting input terminals," the action occurrence time is 10:05:25, the action execution sequence is step 2, and the action execution status is normal execution status; the action occurrence time corresponding to the "tightening wiring" action is 10:05:42, the action execution sequence is step 3, and because the operation time exceeds the standard time, the action execution status is judged to be abnormal execution status; the action occurrence time corresponding to the "setting output voltage" action is 10:06:10, the action execution sequence is step 4, and the action execution status is normal execution status; the action occurrence time corresponding to the "starting the detection program" action is 10:06:40, the action execution sequence is step 5, and the action execution status is normal execution status. This allows us to obtain user operation information, including action type, action occurrence time, action execution order, and action execution status.

[0098] The system acquires operation instruction records, parameter setting records, device status records, and execution feedback records from the terminal operation information. For example, it acquires terminal operation information after cleaning, standardization, and consistency verification through a host computer, PLC controller, and device status acquisition module. For instance, at 10:05:24, an operation instruction of "Input Terminal Connection Confirmation" is recorded; at 10:05:43, an operation instruction of "Tightening Complete Confirmation" and device status information showing a contact resistance of 0.3Ω are recorded; at 10:06:11, an operation instruction of "Set Output Voltage 24V" and device status information showing an output voltage of 24.1V are recorded, and successful execution feedback is received; at 10:06:41, an instruction of "Start Detection Program" and relay closure status are recorded, and successful program startup feedback is received. This forms the terminal operation information, including operation instruction records, parameter setting records, device status records, and execution feedback records.

[0099] The user operation information and the terminal operation information are time-aligned based on timestamps to establish a correspondence between the time of the action and the corresponding time of the recorded operation instruction. For example, video analysis shows that the "connect input terminal" action occurred at 10:05:25, while the corresponding "input terminal connection confirmation" instruction in the terminal operation information was recorded at 10:05:24. The time difference between the two is 1 second, which is less than the preset time matching window of 3 seconds, thus confirming a correspondence. Similarly, the "set output voltage" action occurred at 10:06:10, and the corresponding terminal recording time is 10:06:11, with a time difference of 1 second, thus also establishing a correspondence. After time alignment, a one-to-one correspondence between each action and the terminal operation is established.

[0100] Based on the aforementioned correspondence, the action categories are matched with the operation instruction records, and the consistency of the action execution status, execution feedback records, and device status records are verified to obtain the matching relationship between user operation action information and terminal operation information. For example, for the action of "connecting input terminals," the video side identifies the action execution status as normal, while the terminal side records that the input terminal is conducting, and the execution feedback result is successful. Therefore, it is determined that the video action information and the terminal feedback information are consistent, thus confirming a successful match. For the action of "tightening wiring," video analysis reveals that the duration of this action is significantly longer than the standard time. Therefore, the action execution status is determined to be an abnormal execution status. Although the terminal side returns feedback that tightening is complete, it detects that the contact resistance value is slightly higher than the standard value. Therefore, it is considered that the abnormal status on the video side is consistent with the device status on the terminal side, thus confirming that the action has an abnormal operation. For the action of "setting output voltage," the terminal record shows that the target voltage setting value is 24V, and the actual output voltage of the device is 24.1V, with an error value of only 0.1V, which is less than the allowable error range. Therefore, it is determined that the action execution status is consistent with the device feedback result. After the above verification, the matching relationship between user operation information and terminal operation information is obtained.

[0101] Based on the matching relationship, user operation information with corresponding relationships is fused with terminal operation information to generate operation records in a structured operation sequence. For example, for the "set output voltage" operation, the action category "set output voltage", action occurrence time 10:06:10, action execution sequence step 4, and normal execution status are fused with the terminal-side "set output voltage 24V" operation instruction record, output voltage 24.1V device status record, and execution success feedback information to form a complete operation record; for the "tighten wiring" operation, the action occurrence time 10:05:42, action execution sequence step 3, and abnormal execution status are fused with the corresponding tightening completion instruction, contact resistance status, and execution feedback record on the terminal-side to generate another operation record. In this way, each operation record simultaneously includes the action category, action occurrence time, action execution sequence, action execution status, and corresponding operation instruction record, device status record, and execution feedback record.

[0102] Each of the operation records includes an action category, action occurrence time, action execution sequence, action execution status, and corresponding operation instruction record, device status record, and execution feedback record;

[0103] The operation records are sorted based on the time of the actions, and verified in conjunction with the execution order of the actions to establish the sequential relationship between the operation records and generate a structured operation sequence. For example, operation records such as "connect input terminals," "tighten wiring," "set output voltage," and "start detection program" are arranged in chronological order, corresponding to steps 2, 3, 4, and 5, respectively. When the time order of an operation record is found to be inconsistent with the execution order of the actions, for example, the occurrence time of "start detection program" is earlier than that of "set output voltage," it is considered that there is a sequence conflict, and the sequential relationship between the operation records is corrected by using the execution order information, thereby ensuring that the entire operation process meets the standard process requirements. In this embodiment, the structured operation sequence formed includes operation records such as "connect input terminals," "tighten wiring," "set output voltage," and "start detection program," where each operation record contains an action category, action occurrence time, action execution order, action execution status, and corresponding operation instruction record, device status record, and execution feedback record. By employing the above method, user behavior information from the video side and device operation information from the terminal side are uniformly integrated to form a structured operation sequence that accurately reflects the entire training process. This provides a reliable data foundation for subsequent assessments of operation steps, operational compliance, and result correctness, thereby achieving automated evaluation of the training process. The structured operation sequence includes multiple operation records arranged in chronological order.

[0104] In detail, based on the structured operation sequence, the operation step assessment results, operation standardization assessment results, and result correctness assessment results are obtained, specifically including:

[0105] Multiple operation records are obtained from the structured operation sequence. Each operation record includes an action category, action occurrence time, action execution order, action execution status, and corresponding operation instruction record, equipment status record, and execution feedback record. In this embodiment, after the structured operation sequence is constructed, the student's operation process is automatically scored, taking a "low-voltage distribution cabinet wiring training task" as an example. First, multiple operation records are extracted from the structured operation sequence, such as operation records for "connecting input terminals," "tightening wiring," "setting output voltage," and "starting the detection program." Each operation record includes an action category, action occurrence time, action execution order, action execution status, and corresponding operation instruction record, equipment status record, and execution feedback record. For example, in the "tightening wiring" operation record, the action execution order is step 3, the action execution status is an abnormal execution status, the corresponding equipment status record shows a contact resistance of 0.32Ω, and the execution feedback record indicates successful completion.

[0106] Obtain the pre-configured training evaluation rule package corresponding to the current training task. The training evaluation rule package is pre-built and stored during the system initialization phase and is used to uniformly evaluate the structured operation sequence.

[0107] The training assessment rule package includes rules for evaluating operational steps, rules for evaluating operational standardization, and rules for evaluating the correctness of results. This rule package is pre-built and stored based on standard operating procedures during system initialization and is not regenerated during runtime; instead, it is directly called for unified scoring. Specifically, the rule package comprises three parts: rules for evaluating operational steps, rules for evaluating operational standardization, and rules for evaluating the correctness of results. For example, the standard operating sequence for this training task is defined as: "Power off check → Connect input terminals → Tighten wiring → Set output voltage → Start detection program," and the allowable time range, allowable error range, and standard equipment feedback status for each step are predefined.

[0108] The operational step assessment rules are pre-constructed based on standard operational sequences and their sequential relationships. These rules define the mapping relationship and sequential constraints between each standard operational step in the standard operational sequence and the actual operational record. For example, in this embodiment, the standard step "connecting the input terminal" corresponds to step 2 in the standard sequence. In the student's actual operation, this step is correctly identified as step 2, thus the mapping is successful. However, if a student executes "setting the output voltage" before "tightening the wiring," a sequence deviation value will be calculated based on the sequential constraints. For example, by comparing the standard sequence index difference Δs = |iactual - istandard|, and combining this with the offset count statistics, an operational step deviation score is generated to identify sequence errors or offsets. Simultaneously, standard steps not appearing in the structured operational sequence are marked as operational omissions, and operations that are repeated or have no corresponding standard steps are marked as operational redundancy.

[0109] The operational standardization assessment rules are pre-built based on the execution requirements of standard operating procedures. These rules define the standard execution constraints and allowable deviation ranges for each operation record during execution. For example, in this embodiment, the "tightening wiring" step requires an operation duration between 8 and 12 seconds, and the equipment contact resistance should be less than 0.30Ω. If, in actual operation, a student is found to have taken 18 seconds for this step and the contact resistance is 0.32Ω, then the operation is deemed to have a standardization deviation. The execution deviation value is calculated based on the deviation magnitude; for example, time deviation Δt = 18 − 10 = 8 seconds, resistance deviation ΔR = 0.02Ω. A comprehensive calculation using a preset weighting function yields a decrease in the standardization score. Simultaneously, abnormal action execution states (such as jitter, repetitive operations, etc.) are also included in the standardization constraints for score deduction.

[0110] The result correctness assessment rules are pre-constructed based on standard operation results and their corresponding standard execution status information. These rules define the consistency judgment relationship and deviation grading rules between the execution feedback record and the standard result status corresponding to the device status record. For example, in this embodiment, the standard result for the "set output voltage" step is that the output voltage should be stable within the range of 24V±1V. When the student actually operates the device, the device status record shows an output voltage of 24.1V, and the execution feedback record indicates success. The system calculates the voltage error ΔU=|24.1−24|=0.1V, which is less than the allowable range. Therefore, the result of this step is judged to be completely consistent. If the output voltage is 26V in a certain operation, the system will classify it into a first-level deviation or a second-level deviation based on the deviation range and score it according to the preset deduction strategy in the rule package.

[0111] Based on the aforementioned operation step assessment rules, the operation records in the structured operation sequence are matched with the standard operation sequence based on time order constraints. Operation sequence deviations, omissions, and redundancies are identified. An operation step assessment score is generated according to the preset scoring and deduction strategy in the operation step assessment rules, forming the operation step assessment result. Specifically, firstly, based on the operation step assessment rules, the structured operation sequence and the standard operation sequence are aligned and matched. The system identifies one slight sequence deviation (delayed execution of wiring tightening), no omitted steps, and no redundant steps in this training session. Therefore, the calculated operation step assessment score is 92 points. Secondly, based on the operation standardization assessment rules, constraint analysis is performed on the execution process of each operation record. The "wiring tightening" step has time limits exceeding limits and slight resistance deviations, resulting in a large deduction for standardization, ultimately yielding an operation standardization assessment score of 85 points. Finally, based on the result correctness assessment rules, the final equipment state of all operations is compared for consistency. Except for a few minor deviations, the overall result is correct, thus yielding a result correctness assessment score of 96 points.

[0112] Based on the operational standardization assessment rules, the action execution status and execution process of each operation record in the structured operation sequence are analyzed for compliance with the standardization rules. Operation records that do not meet the standard execution constraints are identified, and operational standardization assessment scores are generated according to the preset scoring and deduction strategies in the operational standardization assessment rules, thus forming operational standardization assessment results.

[0113] Based on the aforementioned result correctness assessment rules, the execution feedback records and device status records of each operation record in the structured operation sequence are compared with the standard execution status information corresponding to the standard operation result to determine the result deviation type and deviation level. A result correctness assessment score is then generated according to the preset scoring and deduction strategy in the result correctness assessment rules, forming the result correctness assessment result. Specifically, the operation step assessment result characterizes the degree of consistency between the actual operation sequence and the standard operation sequence; the operation standardization assessment result characterizes the degree of standardization compliance of each operation record during execution; and the result correctness assessment result characterizes the degree of consistency between the actual operation execution result and the standard operation result.

[0114] Specifically, based on the assessment results of the operation steps, the assessment results of the operation standardization, and the assessment results of the result correctness, the user's practical training evaluation score is obtained. This includes: inputting the assessment results of the operation steps, the assessment results of the operation standardization, and the assessment results of the result correctness into a pre-trained and deployed practical training evaluation fusion model. The practical training evaluation fusion model is constructed based on historical practical training data samples during the system training phase and is used to comprehensively evaluate the multi-dimensional assessment results.

[0115] The integrated evaluation model for practical training assessment includes a task semantic evaluation channel, a capability and behavior evaluation channel, and an execution environment consistency evaluation channel. The task semantic evaluation channel uses job task description information corresponding to the standard operation sequence to semantically vectorize task requirements and calculates a task matching score based on an attention weighting mechanism. The capability and behavior evaluation channel uses user operation behavior features corresponding to the structured operation sequence to feature map user operation capabilities and calculates capability evaluation scores through a nonlinear mapping network. The execution environment consistency evaluation channel uses standard environment information corresponding to the practical training task and operation terminal information generated during user operation to perform correlation modeling for execution environment consistency and calculates an environment consistency score. The task matching score, capability evaluation score, and environment consistency score are weighted and integrated to obtain a comprehensive semantic evaluation score. Based on the comprehensive semantic evaluation score, and combined with the operation step assessment results, operation standardization assessment results, and result correctness assessment results, a weighted summary calculation is performed to generate the user's practical training assessment score. The user's practical training assessment score is used to represent the comprehensive evaluation result of the user's operation step execution, operation standardization, and operation result correctness during the practical training process.

[0116] In one specific embodiment, taking the "electronic equipment assembly training task" as an example, users need to complete the operation steps of "installing the motherboard → connecting the power cord → fixing the heat dissipation module → powering on for self-test" according to the standard procedure. The system first uses "operation step assessment results, operation standardization assessment results, and result correctness assessment results" as input data. These results exist in the form of structured scores, for example, an operation step score of 0.82, an operation standardization score of 0.76, and a result correctness score of 0.88. These are then input into a pre-trained and deployed integrated training assessment model. This model is built based on a large number of historical training samples during the system training phase. Each sample contains the operation video features of different students, terminal operation logs, and manual scoring labels, enabling the model to learn the mapping relationship between different operation modes and the final score. In this model, the task semantic evaluation channel first semantically vectorizes the standard operation sequence of the "electronic device assembly task." For example, "installing the motherboard, connecting the power cord, fixing the heat dissipation module, and powering on for self-test" are mapped into high-dimensional semantic vectors. Combined with task requirements in the job description, such as "standard assembly and avoiding electrostatic damage," an attention weighting mechanism is used to calculate the importance weight of each operation step. For instance, the system might determine that "powering on for self-test" has a higher weight in the overall task, thus assigning a higher influence coefficient to the matching degree of this step, ultimately resulting in a task matching score of 0.85. Simultaneously, the capability and behavior evaluation channel, based on the user's actual generated structured operation sequence—for example, the user exhibiting a repetitive plugging and unplugging action in the "connecting the power cord" step, or taking an excessively long time to tighten screws in the "fixing the heat dissipation module" step—is transformed into behavioral feature vectors (such as action stability, sequence consistency, and execution efficiency). These vectors are then mapped and calculated using a nonlinear mapping network (such as a multilayer perceptron structure), resulting in a capability evaluation score of 0.78, used to characterize the user's actual operational proficiency and standardization level. Furthermore, the execution environment consistency evaluation channel correlates the standard environmental information of the training task (e.g., standard voltage 220V, initial equipment state as unloaded, tools complete) with the operation terminal information collected during user operation. For example, if the system detects a "not fully grounded" message in the device status record before power-on, it calculates the environmental deviation through consistency modeling and outputs an environment consistency score of 0.90, reflecting whether the operating environment meets the standard training conditions and whether it affects the operation results. Subsequently, the system weights and fuses the task matching score (0.85), the ability evaluation score (0.78), and the environment consistency score (0.90). For example, if the weighting coefficients are set to 0.4, 0.4, and 0.2, the comprehensive semantic evaluation score is calculated as: 0.85×0.4+0.78×0.4+0.90×0.2=0.834. This score reflects the overall performance after the fusion of "correct task understanding + operational ability level + environmental consistency".

[0117] Finally, the system weights and summarizes the comprehensive semantic evaluation score with the aforementioned three basic assessment results. For example, setting the weight of operation steps to 0.3, operation standardization to 0.3, result correctness to 0.2, and comprehensive semantic evaluation to 0.2, the final evaluation score is: 0.82×0.3+0.76×0.3+0.88×0.2+0.834×0.2≈0.8148, thus generating the user's final training score of 0.815. Through the above-mentioned integrated evaluation mechanism, this application embodiment can upgrade the traditional evaluation method that relies solely on operation results or manual scoring to a multi-dimensional evaluation system that integrates "task semantic understanding, behavioral ability characterization, and execution environment consistency analysis." On the one hand, the task semantic evaluation channel enhances the quantitative ability of understanding the degree of operation objectives, so that the scoring no longer depends solely on the correctness of the result; on the other hand, the ability and behavior evaluation channel can characterize the user's actual execution quality during the operation process, realizing fine-grained evaluation of the operation process; in addition, the execution environment consistency evaluation channel can effectively eliminate the interference of environmental factors on the scoring, improving the fairness and stability of the evaluation.

[0118] It should be noted that, in some embodiments of this application, before performing temporal action recognition based on the target detection and tracking results, a multimodal validity assessment is performed on the target detection and tracking results to determine whether they meet the conditions for subsequent processing. Specifically, this includes:

[0119] The target detection and tracking results are obtained, and a visual detection result is generated based on the target detection and tracking results. The visual detection result includes a target category sequence, a target position trajectory sequence, a target state change sequence, and action time correlation information. Based on the visual detection result, the precision and recall of visual target detection are calculated, and the visual detection F1 score is calculated based on the precision and recall to obtain the visual detection score Fv, where: Fv=2×Pv×Rv / (Pv+Rv), where Pv is the precision of visual target detection and Rv is the recall of visual target detection.

[0120] Based on the visual detection results and the operation records in the terminal operation information, time alignment and semantic matching are performed. The number of matches Cm in the visual detection results that are consistent with the terminal operation records and the total number of terminal operation records Ct are counted. The matching degree Mo between the visual and terminal data is calculated, where Mo = Cm / Ct. Based on the visual detection score Fv and the matching degree Mo, a weighted fusion is performed according to a preset weight coefficient γ to obtain the multimodal target detection comprehensive score Fm, where Fm = γ × Fv + (1 − γ) × Mo. Here, γ is the weight coefficient of the visual detection results, which is used to adaptively adjust according to different training types. When the multimodal target detection comprehensive score Fm is greater than or equal to a preset threshold Tm, it is determined that the visual detection results meet the validity requirements, and the visual detection results are used as input data for subsequent temporal action recognition and structured operation sequence generation.

[0121] When the comprehensive score Fm of the multimodal target detection is less than the preset threshold Tm, it is determined that the visual detection result does not meet the validity requirements, and a data re-acquisition or manual review process is triggered to improve the reliability of the target detection result.

[0122] In a specific embodiment, taking the "Industrial Power Distribution Cabinet Wiring Training Task" as an example, visual output information is first extracted from the target detection and tracking results to generate visual detection results. These visual detection results include a target category sequence, a target position trajectory sequence, a target state change sequence, and action time correlation information. For example, during the "connecting input terminals" operation, the target category is identified as "wire terminal + screwdriver," and its spatial position change trajectory in the video frame is continuously tracked. Simultaneously, the state change sequence from "not connected" to "inserted" is recorded, and combined with the timestamp, the action time correlation information between 10:05:20 and 10:05:28 is formed. Similarly, during the "tightening wiring" process, the target state change sequence shows a change from "loose → initial fixation → complete tightening," thus forming a complete visual detection result. After obtaining the visual detection results, the precision and recall of the visual target detection are calculated based on manually labeled data or historical verification data. For example, in this embodiment, a total of 10 valid operation target events were detected, of which 9 were consistent with the true annotations, so the precision Pv = 9 / 10 = 0.9; at the same time, there were 10 target events that should be identified in the standard annotations, and 9 were correctly identified, so the recall Rv = 9 / 10 = 0.9. Based on this, the visual detection F1 score, i.e., the visual detection score Fv, is further calculated. Its calculation formula is Fv = 2 × Pv × Rv / (Pv + Rv), which gives Fv = 2 × 0.9 × 0.9 / (0.9 + 0.9) = 0.9, thus obtaining the overall reliability score of the visual detection. Subsequently, the visual detection results are time-aligned and semantically matched with the operation records in the terminal operation information to calculate the degree of consistency between the visual and terminal data. For example, in the "tightening wiring" operation, the visual detection time is 10:05:40, while the terminal operation record time is 10:05:41. The time difference is 1 second, which is less than the preset time window of 3 seconds, so it is determined that the time match is successful. At the same time, the action category recognized by the vision is "tightening screws," and the terminal record is "screw tightening instruction," which semantically belong to the same operation type, so it is counted as a valid match. Assuming that there are a total of 10 terminal operation records in the entire training process, of which 9 can be matched by vision detection, the number of matches Cm=9, the total number of terminal operation records Ct=10, so the matching degree Mo=9 / 10=0.9, which is used to characterize the consistency between the visual recognition result and the actual operation behavior. After obtaining the visual detection score Fv and the matching degree Mo, the multimodal target detection comprehensive score Fm is calculated by weighted fusion according to the preset weight coefficient γ. For example, in this training task, since the video quality is high but the stability of terminal data is more critical, γ=0.6 is set, then Fm=0.6×0.9+0.4×0.9=0.9, thus obtaining a final multimodal consistency score of 0.9.When the multimodal target detection comprehensive score Fm is greater than or equal to the preset threshold Tm (e.g., Tm=0.85), the current visual detection result is deemed to meet the validity requirements, indicating a high degree of consistency between the visual and terminal data. Therefore, this visual detection result is used as input data for subsequent temporal action recognition and structured operation sequence generation, ensuring that subsequent analysis is based on a reliable data source. Conversely, when Fm is less than the preset threshold Tm, for example, due to insufficient lighting or occlusion leading to numerous false visual detections, causing Fv to drop to 0.7 and Mo to drop to 0.65, ultimately resulting in Fm=0.68, the current visual detection result is deemed unreliable. This automatically triggers a data re-acquisition process (e.g., adjusting the camera frame rate or re-acquiring video) or initiates a manual review process to improve the accuracy of the target detection result.

[0123] Optionally, in some embodiments of this application, after the generation step of the temporal action recognition result, a comprehensive score is performed on the temporal action recognition result to quantitatively evaluate the effectiveness of the temporal action recognition result. Specifically, this includes: calculating an action sequence correctness score Sorder based on the action sequence correctness information in the temporal action recognition result, wherein the action sequence correctness score is calculated based on the edit distance between the standard operation sequence and the structured operation sequence; calculating an action normalization score Snorm based on the action normalization information in the temporal action recognition result and combined with the comprehensive score of multimodal target detection corresponding to each action, wherein the action normalization score is determined based on the average of each action normalization score; and calculating a process integrity score Scomplete based on the process integrity information in the temporal action recognition result, wherein the process integrity score is based on... The ratio of the actual number of completed actions to the number of required actions in the standard operation sequence is determined. The action sequence correctness score Sorder, action standardization score Snorm, and process integrity score Scomplete are weighted and fused according to preset weight coefficients to obtain the comprehensive score Sseq for temporal action recognition, where: Sseq = λ × Sorder + μ × Snorm + ν × Scomplete; where λ, μ, and ν are weight coefficients, and satisfy λ + μ + ν = 1. When the comprehensive score Sseq for temporal action recognition is greater than or equal to the preset threshold Tseq, the temporal action recognition result is determined to meet the validity requirements and is used for subsequent training evaluation score calculation. When the comprehensive score Sseq for temporal action recognition is less than the preset threshold Tseq, the temporal action recognition result is determined to not meet the validity requirements, and a manual review or data reprocessing process is triggered.

[0124] In one specific embodiment, taking the "industrial equipment wiring training task" as an example, after completing the timing action recognition, a timing action recognition result is generated for the student's entire operation process. This result includes information on the correctness of the action sequence, the standardization of the action, and the completeness of the process. Furthermore, a comprehensive score is given to the recognition result to quantitatively determine its reliability and usability for subsequent grade calculation. First, in the calculation of the action sequence correctness score (Sorder), the structured operation sequence actually performed by the student is compared with the standard operation sequence. For example, the standard operation sequence is "power off check → connect input terminals → tighten wiring → insulation test → power on test," while the student's actual execution sequence is "power off check → tighten wiring → connect input terminals → insulation test → power on test." The edit distance is calculated based on the difference in sequence between the two. For example, if a "swap operation" (reversing the order of connecting input terminals and tightening wiring) is required, the edit distance is recorded as 1. With a standard sequence length of 5, the edit distance is normalized, for example, Sorder = 1 - (1 / 5) = 0.8, resulting in an action sequence correctness score of 0.8, which characterizes the degree of consistency between the student's operation sequence and the standard procedure. Secondly, in calculating the action standardization score (Snorm), the standardization score for each action in the temporal action recognition results is evaluated in conjunction with the comprehensive score from multimodal object detection. For example, the student's action standardization score is 0.9 in the "connecting input terminals" operation (stable action with no obvious jitter), 0.7 in the "tightening wiring" process due to excessive operation time and large tool angle deviation, and 0.85 in the "power-on test" process. The corresponding multimodal object detection scores are 0.92, 0.88, and 0.95, respectively. The system averages the scores for each action's standardization, for example, Snorm = (0.9 + 0.7 + 0.85) / 3 ≈ 0.817, resulting in an action standardization score of 0.817, which characterizes the overall level of standardized execution of the operation process. Furthermore, in calculating the process completeness score Scomplete, the system compares the set of actions actually completed by the student with the set of required actions in the standard operation sequence. For example, if the standard operation sequence contains 5 required actions, and the student actually completes all 5 actions without omissions or redundancies, then the number of completed actions is 5, and the number of required actions is 5. Therefore, Scomplete = 5 / 5 = 1.0, indicating that the process completeness fully meets the standard requirements; if there are any missing actions, this percentage will decrease accordingly.Subsequently, the system weights and fuses the action sequence correctness score (Sorder), action standardization score (Snorm), and process completeness score (Scomplete) according to preset weight coefficients. For example, setting λ=0.4, μ=0.3, and ν=0.3, the comprehensive score is calculated as: Sseq=0.4×0.8+0.3×0.817+0.3×1.0≈0.8651, resulting in a comprehensive temporal action recognition score of approximately 0.865. Finally, when the comprehensive score Sseq is greater than or equal to the preset threshold Tseq (e.g., Tseq=0.8), the current temporal action recognition result is deemed valid, indicating that the recognition result can accurately reflect the student's operation process and can be directly used for subsequent practical training assessment score calculation. Conversely, when Sseq is lower than the threshold, for example, due to video occlusion causing some actions to be misrecognized and Sorder to drop to 0.6, the recognition result is deemed unreliable, and a manual review or data reprocessing process is automatically triggered to re-perform target detection, action recognition, or data alignment, thereby improving the overall assessment accuracy.

[0125] In this embodiment of the application, the operational standardization assessment step involves standardization scoring of the operational standardization assessment results, specifically including:

[0126] Obtain multiple operation records from the structured operation sequence, wherein each operation record includes action category, action occurrence time, action execution order, action execution status, operation instruction record, device status record, and execution feedback record;

[0127] Based on the operational standardization assessment rules, the standard execution requirements and allowable deviation ranges of each operation record defined in the operational standardization assessment rules are analyzed to determine the standardization evaluation index corresponding to each operation record; cross-consistency analysis is performed on the visual detection results corresponding to the terminal operation information and the video data to obtain the operational standardization sub-scores, where: Sterminal=(Ccorrect / Ctotal)×(1−Edev); where Ccorrect is the number of operation records that are consistent between the operation instruction record and the execution feedback record, Ctotal is the total number of operation records, and Edev is the average execution deviation value of each operation record; Svisualnorm=Σ(Fmi×Wi), where i=1~N, N is the number of operation records corresponding to the structured operation sequence, and Wi is the number of operation records of the i-th operation record in the standard operation sequence. The importance weights are determined, and ΣWi=1; based on the operational standardization assessment rules, the standardization sub-scores are weighted and fused to obtain the operational standardization comprehensive score Snorm, where: Snorm=ω×Sterminal+(1−ω)×Svisualnorm; where ω is the operational terminal information weight coefficient, used for adaptive adjustment according to the training type; when the operational standardization comprehensive score Snorm is greater than or equal to the preset threshold Tnorm, it is determined that the operational standardization assessment result meets the standardization requirements, and an operational standardization assessment score is generated; when the operational standardization comprehensive score Snorm is less than the preset threshold Tnorm, it is determined that the operational standardization assessment result does not meet the standardization requirements, and the corresponding abnormal operation record is recorded; where the operational standardization assessment result is used to characterize the degree of standardization of each operation record during the execution process.

[0128] In a specific embodiment, taking the "Industrial Electrical Wiring Training Task" as an example, after completing the construction of the structured operation sequence, multiple operation records are obtained. These records include items such as "Power Off Confirmation," "Connect Input Terminals," "Wire Connection Tightening," "Insulation Detection," and "Power-On Test." Each operation record includes the action category, action occurrence time, execution sequence, execution status, and corresponding operation instruction record, equipment status record, and execution feedback record. For example, in the "Wire Connection Tightening" operation record, the operation instruction record is "Execute Tightening Instruction," the equipment status record shows that the torque did not reach the standard value, and the execution feedback record is "Execution Delay and Insufficient Torque," thus forming a complete structured operation record information. During the analysis of the normative evaluation indicators, based on the pre-configured operational normative assessment rules, the standard execution requirements and allowable deviation ranges of each operation record are analyzed. For example, for the "Wire Connection Tightening" operation, the rules stipulate that the standard execution requirement is "Complete tightening within 10 seconds after the input terminal connection is completed, and the torque reaches 2.5 N·m ± 0.2 N·m." Based on this, the normative evaluation indicators for the operation record are determined to include the allowable range of time deviation, the range of torque deviation, and the consistency requirements of execution feedback, thus forming quantifiable normative evaluation constraints. During the cross-consistency analysis between the terminal and vision systems, the terminal operation information is compared with the visual inspection results. For example, the visual inspection results show that the actual occurrence time of the "wiring tightening" action is T=120s, while the corresponding instruction time recorded in the terminal operation information is T=110s, resulting in a 10-second delay. Simultaneously, the equipment status shows that the torque is not up to standard, consistent with the visual inspection finding that "the operation did not complete the standard tightening posture." Based on this, the consistency quantity Ccorrect is calculated. For example, if the instructions and feedback are consistent in 4 out of 5 operation records, then Ccorrect = 4 and Ctotal = 5. Simultaneously, the average execution deviation value Edev is calculated using time deviation and state deviation. For example, if the deviations for each operation record are 0.05, 0.10, 0.20, 0.08, and 0.12 respectively, then Edev = 0.11. Therefore, Sterminal = (4 / 5) × (1 − 0.11) = 0.8 × 0.89 = 0.712. During the visual standardization score calculation, a visual standardization score Fmi is generated for each operation record. For example, the visual score for the "Power Off Confirmation" action is 0.95, "Connect Input Terminal" is 0.88, "Wire Tightening" is 0.70, "Insulation Detection" is 0.92, and "Power On Test" is 0.96. These are weighted according to the importance weight Wi of the standard operation sequence. For example, the tightening operation has the highest weight of 0.3, and the weights of the other operations are 0.2, 0.2, 0.15, and 0.15 respectively. Therefore: Svisualnorm = 0.95 × 0.2 + 0.88 × 0.2 + 0.70 × 0.3 + 0.92 × 0.15 + 0.96 × 0.15 ≈ 0.8685.Subsequently, the terminal weight coefficient ω is set according to the training type. For example, if ω = 0.6, the comprehensive standardization score is calculated as: Snorm = 0.6 × 0.712 + 0.4 × 0.8685 ≈ 0.796. When this Snorm is compared with the preset threshold Tnorm (e.g., 0.75), since 0.796 ≥ 0.75, the operation standardization assessment result is determined to meet the standardization requirements, and the corresponding standardization assessment score is generated. If, in a training session, Edev rises to 0.25 due to a severe timeout in "tightening the wiring," Snorm will drop below the threshold. At this time, the corresponding abnormal operation record is automatically recorded, such as "time exceeded" or "insufficient torque," and used for subsequent teaching feedback and problem localization analysis. Through the above method, this embodiment realizes multi-source fusion quantitative evaluation of operation standardization, so that terminal data consistency, visual execution standardization, and operation deviation can all be uniformly modeled, thereby improving the objectivity and interpretability of the training process evaluation.

[0129] Example 2

[0130] This embodiment provides an intelligent evaluation system for practical training processes based on multimodal artificial intelligence. The system adopts a five-layer modular architecture, with each layer operating independently yet collaboratively to achieve automatic data collection, structured analysis, and quantitative evaluation of the entire practical training process. (See attached...) Figure 2 As shown, the practical training assessment system in this embodiment achieves full-process perception and intelligent evaluation of students' practical training process based on multi-view video acquisition and structured modeling of the operation area. Figure 2As shown, the 6-camera headband-style acquisition device 1 (head-mounted multi-camera array) is worn on the student's head to collect real-time video data related to the student's hand operations, tool interactions, and line of sight from a first-person perspective, thereby obtaining highly spatiotemporally consistent subjective operational perspective information. This data serves as one of the important data sources for visual change intensity calculation (Vt) and action trigger recognition in the AI ​​processing layer. The third-person perspective camera 2 is located outside the training space to monitor the entire training process from a fixed external perspective, focusing on collecting the student's hand movement trajectory, tool usage status, and changes in the overall operation process. This perspective data is used to assist in the calculation of target state change identifier (Ot) in the target detection module and to provide external verification basis for multimodal target detection scoring (Fm). The training operation area 3 in the figure is the core space for students to conduct specific operation training. All operational behaviors (such as tool picking, parameter setting, and action execution) are completed in this area. This area corresponds to the keyframe extraction and temporal action modeling object in the system's AI processing layer and is the core data carrier for forming structured action sequences. The control panel 4 houses the training equipment and the input interface for the operating terminal. Students use this panel to input parameters, trigger commands, and control the equipment. The output data from the operating terminal is used to generate At (action trigger signal), Stationary (terminal compliance score), and Mo (cross-modal consistency verification with visual detection results). The third-view camera bracket structure 5 is used to fix and adjust the shooting angle and coverage of the third-view camera, ensuring stable coverage of the entire training operation area 3 and guaranteeing stable global visual information under different heights or operating postures, thereby improving the stability and accuracy of multimodal target detection. Figure 2 The multi-view acquisition structure shown allows this system to simultaneously acquire collaborative video data from both the first-view (headband 1) and third-view (camera 2 and bracket 5) perspectives. Combined with terminal data from the operating console 4 and spatial behavior information from the operating area 3, this system achieves synchronous acquisition and time-aligned processing of multimodal data. This provides a unified data foundation for subsequent keyframe filtering, target detection, temporal action recognition, and calculation of standardized scoring formulas, ultimately enabling a structured evaluation of the entire training process. Specifically, the system first connects to general-purpose high-definition camera equipment and student operating terminals in the training scenario through a data acquisition layer, synchronously acquiring video data and operational interaction data throughout the training process. Video data characterizes students' visual operational behavior during training, while terminal data records device commands, parameter inputs, and status feedback. To ensure consistency of multi-source data in subsequent analysis, the acquired multi-source data undergoes timestamp alignment processing, establishing a one-to-one mapping between each video frame and its corresponding operational event, thus forming a unified multimodal temporal data stream. This data stream serves as the input foundation for the subsequent AI processing layer.

[0131] In the AI ​​processing layer, the aforementioned multimodal temporal data stream is first standardized. This involves sequentially extracting keyframes, detecting targets, and recognizing temporal actions from the video data, while the data from the operating terminal undergoes format cleaning and semantic standardization. This results in structured student operation sequence data. This structured sequence describes the content, execution order, and standardized characteristics of the student's operations throughout the training process and serves as a unified input for subsequent scoring calculations.

[0132] During keyframe extraction, a keyframe importance score is calculated for each video frame t using the following formula: St = α × (wv × Vt + wo × Ot + wa × At) + β × (Pt × Rt) / (Pt + Rt); where α and β are weighting coefficients used to balance the contribution ratio between multimodal feature terms and model recognition reliability terms, and satisfy α + β = 1. Preferably, α is 0.7 and β is 0.3 to emphasize the dominant role of visual and behavioral changes in keyframe selection, while also considering model detection stability. Further, wv, wo, and wa represent the weighting coefficients corresponding to the intensity of visual changes, target state changes, and action trigger signal intensity, respectively, and all three satisfy wv + wo + wa = 1, where wv = 0.4, wo = 0.3, and wa = 0.3, respectively, corresponding to the different contribution levels of visual information changes, object state changes, and action trigger signals in keyframe determination. Wherein, Vt represents the intensity of visual change of frame t relative to the previous frame t-1. This value is obtained by calculating the pixel-level difference between adjacent frames and normalizing it using the L2 norm. Its value ranges from 0 to 1, and the larger the value, the more significant the visual change between frames. Ot represents the target state change indicator output by the target detection model. It takes a value of 1 when a change in the position, category, or attribute of the training tool or the object being operated is detected, and a value of 0 otherwise. At represents the intensity of the action trigger signal obtained by normalizing the event triggered by the student's operating terminal. Its value ranges from 0 to 1 and is used to reflect the trigger clarity and intensity of the operation behavior in the time dimension. Meanwhile, Pt and Rt represent the prediction precision and recall of the pre-trained action recognition model on frame t, respectively. Pt is the proportion of correctly identified positive samples out of all predicted positive samples, and Rt is the proportion of correctly identified positive samples out of all actual positive samples. Both are used to characterize the reliability and completeness of the action recognition results for that frame, and their combined term (Pt×Rt) / (Pt+Rt) is used to constrain the overall recognition quality of the model on that frame. After obtaining the keyframe importance score St, it is compared with a preset threshold T. When St is greater than or equal to T, the video frame is determined to be a keyframe and retained for subsequent target detection and time-series action recognition processes; when St is less than the threshold, the frame is filtered to reduce subsequent computational redundancy. The threshold T is set to 0.6 by default and can be dynamically adjusted according to the action density of different training scenarios. For example, in action-intensive training, T can be reduced to 0.5 to retain more fine-grained process information, while in training scenarios with relatively sparse actions or slow operation rhythm, T can be increased to 0.7 to further improve the selectivity and computational efficiency of keyframe selection.

[0133] Subsequently, for the selected keyframes and their corresponding operation events, multimodal target detection and cross-modal consistency verification are performed to obtain a multimodal target detection score, which is calculated as follows: Fm=γ×Fv+(1−γ)×Mo; where γ is the weight coefficient of the visual detection result, used to adjust the contribution ratio of the visual detection result and the terminal data verification result in the overall score. Its preferred value is 0.6, and it can be dynamically adjusted according to the characteristics of different training scenarios. For example, in mechanical operation training, since visual action features are more obvious, γ can be appropriately increased to 0.7 to enhance the dominant role of visual information; while in software operation training, since the operation behavior is mainly reflected at the terminal instruction level, γ can be reduced to 0.5 to increase the weight of terminal data.

[0134] Here, Fv represents the F1 score of visual object detection, used to measure the overall recognition performance of the visual detection model on the current keyframe. It is calculated as the harmonic mean of precision (Pv) and recall (Rv), specifically expressed as Fv = 2 × Pv × Rv / (Pv + Rv), where Pv is the proportion of correctly identified targets out of all detected targets, and Rv is the proportion of correctly identified targets out of the actual number of existing targets, reflecting the accuracy and completeness of the model's object detection. Mo represents the consistency matching degree between the operation terminal data and the visual detection results, reflecting the degree of correspondence between the visual recognition results and the actual operation behavior. It is calculated as Mo = Cm / Ct, where Ct represents the total number of operation nodes reported by the operation terminal, and Cm represents the number of operation nodes that successfully match the visual detection results. This is the number of nodes that can be matched one-to-one after timestamp alignment and operation identifier consistency verification, thus quantifying the consistency between cross-modal data. After completing the multimodal fusion calculation, the obtained Fm is compared with a preset threshold Tm. When Fm is greater than or equal to Tm, the multimodal target detection result corresponding to the keyframe is deemed valid and is used as a trusted action node input into the subsequent temporal action recognition module. When Fm is less than Tm, it is considered that there is a significant inconsistency between the visual detection and the terminal data or that the detection confidence of the keyframe is insufficient. In this case, the system will automatically trigger a data re-acquisition process or a manual review mechanism to ensure the accuracy and reliability of subsequent action recognition and scoring calculation. The threshold Tm is preferably set to 0.8 and can be dynamically adjusted according to the accuracy requirements of different training scenarios. For example, it can be increased to 0.85 in high-precision medical training, while the default value can be kept unchanged in conventional skills training.

[0135] After completing the multimodal consistency verification, the verified keyframes and their corresponding operation events are reconstructed into a sequence of trusted action nodes. Based on this sequence, temporal action recognition and modeling are performed to obtain a comprehensive score for temporal action recognition. The calculation formula is as follows:

[0136] Sseq = λ × Sorder + μ × Snorm + ν × Scomplete; where λ, μ, and ν are weighting coefficients used to balance the influence of the correctness of the action sequence, the standardization of the action, and the completeness of the process on the overall score, and satisfy λ + μ + ν = 1. In the preferred case, λ is 0.4, μ is 0.4, and ν is 0.2. In specific application scenarios, when the training task emphasizes the strictness of the operation sequence, such as software process operations, λ can be appropriately increased to 0.5. In medical or fine operation training, in order to strengthen the evaluation of the standardization of single-step actions, μ can be appropriately increased to 0.5. Here, Sorder represents the action sequence correctness score, which measures the degree of deviation between the student's actual action sequence and the standard operating procedure at the sequence level. Its calculation is based on the edit distance D, specifically calculated as Sorder = 1 - (D / Lmax), where D represents the minimum number of edit operations required to convert the student's action sequence into a standard operating sequence, including insertion, deletion, or replacement operations, and Lmax represents the length of the standard operating sequence, used to normalize the edit distance, thus limiting the value of Sorder to between 0 and 1. Snorm represents the action standardization score, reflecting the degree of standardization of each independent action during execution. This indicator is obtained by statistically averaging the multimodal target detection score Fmi corresponding to each action node, specifically calculated as Snorm = (1 / N)ΣFmi, where N represents the total number of actions actually completed by the student, and Fmi represents the comprehensive score obtained by the i-th action node after multimodal detection and consistency verification. This score integrates visual detection results and operational terminal data, thus reflecting the overall execution quality of a single step action. Scomplete represents the process integrity score, used to measure whether students have fully executed all required steps in the standard operating procedure. It is calculated as Scomplete = Cvalid / Crequired, where Cvalid represents the number of required steps actually completed by the student, and Crequired represents the total number of required steps specified in the standard operating procedure, thus reflecting the degree to which the operation covers the standard procedure. After calculating the above three sub-indicators, a combined temporal action recognition score Sseq is obtained and compared with a preset threshold Tseq. When Sseq is greater than or equal to Tseq, the temporal action recognition result is deemed valid, and a structured standard action sequence is output, which can be directly used for subsequent comprehensive score calculation. When Sseq is less than Tseq, the system considers the action sequence to have significant deviations in terms of sequential rationality, standardization, or completeness. Therefore, the student's training process is marked as "pending review" and pushed to the teacher for manual review and correction to ensure the accuracy and reliability of the final evaluation results.The threshold Tseq is preferably set to 0.75, and can be dynamically adjusted according to the rigor of different training projects. It can be increased to 0.8 in high-precision scenarios, while the default value can be kept unchanged in basic skills training scenarios.

[0137] After obtaining the structured action sequence, the system further integrates the data output from the operation terminal to evaluate the standardization of the operation, and obtains a comprehensive score for the standardization of the operation. The calculation formula is as follows:

[0138] Snormop = ω × Sterminal + (1 − ω) × Svisualnorm; where ω is the terminal data weighting coefficient, used to adjust the contribution ratio of terminal data and visual motion data in the standardization evaluation. A value of 0.5 is preferred, and it can be dynamically adjusted according to different training types. For example, in software operation training, since operation instructions are mainly directly reflected by terminal input, ω can be appropriately increased to 0.6 to enhance the weight of terminal data; while in mechanical operation or physical equipment operation training, since motion standardization is more reflected in visual behavior, ω can be appropriately decreased to 0.4 to increase the influence weight of visual evaluation. Here, Staten represents the standardization score of the operation terminal data, which is used to measure the accuracy of students' instruction execution and the rationality of parameter settings in the operation terminal. It is calculated as follows: Staten = (Ccorrect / Ctotal) × (1−Eavg), where Ctotal represents the total number of instructions or parameters fed back by the operation terminal, Ccorrect represents the number of instructions or parameters judged to be correctly executed, and this correctness judgment is obtained by comparing the system rule base with the standard operation process; Eavg represents the average relative error of parameter settings. This error is obtained by calculating the difference between the actual input parameters and the standard parameters and then normalizing it to the range of 0 to 1. It is used to reflect the degree of parameter deviation. The smaller the value, the closer the operation is to the standard requirements. Svisualnorm represents the visual action standardization score, used to characterize the degree of standardization of students' actions during actual operation. This index is calculated by weighting the multimodal object detection scores Fmi of each action node in the temporal action recognition results. The calculation method is: Svisual_norm=∑(Fmi×Wi), where N represents the total number of action nodes, Fmi represents the comprehensive score obtained by the i-th action node after completing multimodal object detection and cross-modal consistency verification. This score has integrated the visual detection results and the matching results of the operation terminal, so it can reflect the true execution quality of a single action; Wi represents the importance weight of the i-th action in the standard operation process, and satisfies that the sum of all weights is 1, which is used to reflect the difference in the impact of different actions in the overall process.

[0139] After completing the fusion calculation of the aforementioned visual and terminal data, the system obtains a comprehensive operational standardization score, Snormop, and compares it with a preset threshold, Tnorm. When Snormop is greater than or equal to Tnorm, the student's operation process is deemed to meet the standardization requirements, and the system can proceed to the subsequent comprehensive score calculation process. When Snormop is less than Tnorm, the system determines that there is a standardization deviation in the operation process. In this case, the system will automatically record abnormal operation nodes, such as incorrect parameter settings, non-standard operation actions, or deviations from the standard process, and synchronously feed the abnormal information back to the interactive display layer. On the one hand, targeted improvement suggestions are pushed to the student to guide them in correcting specific operational problems; on the other hand, the teacher marks this training session as a key review target to support further manual review and correction of the scoring results. The threshold Tnorm is preferably set to 0.7 and can be dynamically adjusted according to the standardization requirements of different training projects. For example, it can be increased to 0.8 in high-precision operation scenarios such as medical applications, while the default value can be kept unchanged in basic skills training scenarios. After completing the scoring calculations at each level, the system enters the comprehensive scoring stage, integrating keyframe scores, multimodal detection results, temporal action recognition results, and operational standardization scores. A final training score is generated using a combination of deductions and additions, and output to the interactive display layer. The student side displays the scoring results and improvement suggestions, while the teacher side views the overall class learning analysis and individual problem distribution, and supports manual review and score correction. Furthermore, the system constructs a modular rule package mechanism through an evaluation configuration layer. Each training project corresponds to an independent rule package, including modules for operational steps assessment, operational standardization assessment, and result correctness assessment. By adjusting the weight parameters and threshold parameters of each module (including T, Tm, Tseq, Tnorm, etc.), the system achieves adaptive evaluation requirements for different training scenarios, enabling it to adapt to different types of training tasks, such as medical, mechanical, and software training.

[0140] Through the above technical solutions, this embodiment realizes a full-process intelligent evaluation mechanism that integrates multi-source data synchronous acquisition, cross-modal consistency verification, temporal action structured modeling, and multi-level score fusion. This effectively improves the objectivity, consistency, and interpretability of training evaluation and realizes end-to-end automated processing from video data to structured scoring results.

[0141] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make modifications, alterations, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. An automatic evaluation method for practical training processes based on artificial intelligence, characterized in that, include: Collect video data and terminal data during user training; Keyframes are extracted from the video data, target detection and tracking are performed based on the extracted keyframes, and temporal action recognition is performed based on the target detection and tracking results to obtain user operation action information. The terminal data is cleaned, standardized, and validated for consistency to obtain terminal operation information. Based on the association and matching of the user operation information and the terminal operation information, a structured operation sequence corresponding to the user's training process is generated; Based on the structured operation sequence, obtain the operation step assessment results, operation standardization assessment results, and result correctness assessment results; Based on the assessment results of the operation steps, the operation standardization, and the correctness of the results, the user's practical training evaluation score is obtained.

2. The automatic evaluation method for practical training based on artificial intelligence according to claim 1, characterized in that, Keyframes are extracted from the video data; target detection and tracking are performed based on the extracted keyframes; and temporal action recognition is performed based on the target detection and tracking results to obtain user operation action information, specifically including: The video data is parsed to extract the video frame sequence; Based on preset keyframe filtering rules, keyframes are filtered out from the video frame sequence; Based on the keyframe recognition training tool, the operation object, and the user action, and continuously track the training tool, the operation object, and the user action, the target detection and tracking results are obtained. Based on the target detection and tracking results, an action time series is constructed, and the sequence of action occurrence, duration of action, and changes in action state are identified to obtain the time-series action recognition results. User operation action information is generated based on the time sequence action recognition results; The user operation information includes the action category, the time of action occurrence, the order of action execution, and the action execution status.

3. The automatic evaluation method for practical training based on artificial intelligence according to claim 2, characterized in that, Based on preset keyframe filtering rules, keyframes are filtered from the video frame sequence, specifically including: Each video frame is processed by a pre-trained visual feature extraction network to extract features and obtain corresponding video frame feature vectors. The video frame feature vectors are used to characterize the visual features of the training tools, the objects being operated on, and the human action areas in the video frame. The visual change intensity Vt is obtained by calculating the feature difference based on the video frame feature vectors corresponding to adjacent video frames. The visual change intensity Vt is obtained by calculating the distance between the video frame feature vectors of the t-th frame and the t-1-th frame and normalizing them. The training tools and objects in each video frame are detected based on the target detection model to obtain target category information, target location information and target state information. The target category information is used to represent the category to which the target belongs, the target location information is used to represent the spatial location of the target in the video frame, and the target state information is used to represent the current state attribute of the target. The target location information and target status information corresponding to the same target in adjacent video frames are compared. When the target location information or the target status information changes, a target status change identifier Ot is generated, where Ot is 1 and Ot is 0 otherwise. Based on the operation instruction records and execution feedback records corresponding to the timestamps of each video frame in the operation terminal data, the operation instruction trigger type, operation instruction trigger time and execution feedback status information are extracted and normalized to obtain the action trigger signal strength At. The execution feedback status information includes operation success status, operation failure status and abnormal status. Obtain the action recognition results output by the pre-trained action recognition model for the video frame sequence, wherein the pre-trained action recognition model is a deep learning-based temporal action recognition model used to identify action categories and predict action time intervals for the video frame sequence. The action recognition results include action category, action start time, action end time, and action confidence. Obtain standard annotation data, wherein the standard annotation data is a dataset formed by manually annotating training videos based on standard operating procedures, and the standard annotation data includes at least the standard action category, standard action sequence, and standard action time interval; Based on the alignment and comparison between the action recognition results and the standard annotation data, the action recognition accuracy Pt is calculated, where Pt is the ratio of the number of correctly recognized actions to the total number of recognized actions; Based on the alignment and comparison between the action recognition results and the standard annotation data, the action recognition recall rate Rt is calculated, where Rt is the ratio of the number of correctly recognized actions to the total number of standard actions in the standard annotation data; Based on the visual change intensity Vt, target state change identifier Ot, action trigger signal intensity At, action recognition precision Pt, and action recognition recall Rt, the keyframe importance score St corresponding to each video frame is calculated, where the keyframe importance score St corresponding to the t-th frame satisfies: St=α×(wv×Vt+wo×Ot+wa×At)+β×(Pt×Rt) / (Pt+Rt); Where α and β are weighting coefficients, and satisfy α+β=1; wv, wo, and wa are the weighting coefficients corresponding to the intensity of visual change, the identifier of target state change, and the intensity of action trigger signal, respectively. Each video frame whose keyframe importance score St is greater than or equal to a preset threshold T is identified as a keyframe.

4. The automatic evaluation method for practical training based on artificial intelligence according to claim 3, characterized in that, Based on the keyframe recognition training tool, the object being operated on, and the user's actions, and by continuously tracking the training tool, the object being operated on, and the user's actions, target detection and tracking results are obtained, specifically including: Multi-scale feature extraction is performed on the keyframes to obtain target feature information, wherein the target feature information is used to characterize the visual features of the training tool, the operation object, and the area corresponding to the user's action. Based on the target feature information, a pre-trained target detection model is used to identify and locate the training tools, operation objects, and user action corresponding areas in each keyframe to obtain target detection results. The target detection results include target category information, target location information, and target state information. Based on the target detection results, the target position information and target state information corresponding to the same target in adjacent keyframes are compared to obtain the target cross-frame change information, wherein the target cross-frame change information includes changes in target position information and changes in target state information; Based on the target detection results and the target cross-frame change information, the correspondence of the same target in the time series is matched to generate target trajectory information, wherein the target trajectory information includes target position trajectory and target state trajectory, which are used to characterize the continuous change process of the target in the time dimension; Based on the target detection result and the target trajectory information, a visual detection result is generated. The visual detection result is used to characterize the temporal visual detection information after the target detection result and the target trajectory information are fused. The visual detection result includes target category sequence, target position trajectory, target state trajectory and action time information. Obtain the operation terminal data corresponding to the action time information, and perform matching verification on the visual detection result based on the operation terminal data; If the matching verification passes, the visual detection result undergoes a time-based structured reconstruction process to obtain a reconstructed visual detection result. The structured reconstruction process includes: Establish a time correspondence between target trajectory information and operation terminal data based on action time information; Establish a correspondence between target category information and operation instruction records; Complete and correct any unmatched target trajectory information and corresponding time information; The target location trajectory, target state trajectory, and action time information are corrected based on the matching results; Based on the reconstructed visual detection results, target detection and tracking results are generated; The target detection and tracking results include action category information, action start time, action end time, and action status information.

5. The automatic evaluation method for practical training based on artificial intelligence according to claim 4, characterized in that, Based on the target detection and tracking results, an action time series is constructed, and the sequence of action occurrence, duration of action, and changes in action state are identified to obtain the time-series action recognition results, specifically including: The action category information, action start time, action end time and action status information corresponding to each action in the target detection and tracking results are obtained. The action status information includes at least the action completion status, action abnormal status and action interruption status. Calculate the corresponding action time interval based on the start time and end time of each action, and sort all actions by time using the start time of the action as the time base to generate an action time sequence. Based on the action time sequence and action category information, the temporal sequence relationship between each action is determined, and an action execution order sequence is generated to represent the actual action execution order. Obtain standard action information corresponding to the structured operation sequence, wherein the standard action information includes standard action category information and standard action time constraint information in the structured operation sequence; The action time series is matched and analyzed with the standard action information to generate action matching relationship and matching deviation results. The matching analysis includes consistency matching based on action category information and standard action category information, and calculation of time deviation between action time interval and standard action time constraint. Based on the matching relationship and the matching deviation result, the sequence deviation value, execution deviation value, missing action identifier and redundant action identifier of each action are calculated. The execution deviation value is calculated based on the action status information and time deviation. Based on the sequence deviation value, execution deviation value, and missing action identifiers and redundant action identifiers, a temporal action recognition result is generated; The timing action recognition results include: Action sequence correctness information is used to characterize the degree of consistency between the actual action execution sequence and the standard action sequence in the structured operation sequence, and is quantitatively represented based on the sequence deviation value; Action standardization information is used to characterize the degree of standardization of a single action during execution, and is quantitatively calculated based on the execution deviation value and action state information; Process integrity information is used to characterize the degree of coverage consistency between the actual set of executed actions and the standard set of actions in the structured operation sequence, and is calculated based on the number of missing actions and the number of redundant actions.

6. The automatic evaluation method for practical training based on artificial intelligence according to claim 5, characterized in that, User operation action information is generated based on the temporal action recognition results, specifically including: Obtain the action sequence correctness information, action standardization information, and process integrity information corresponding to each action in the time-series action recognition results; Based on the start time and end time of the action used to construct the action time sequence in the time sequence action recognition result, the action occurrence time corresponding to each action is determined, wherein the action occurrence time is the action start time of the corresponding action; Based on the action execution sequence generated from the temporal action recognition results, the action execution order corresponding to each action is determined. Based on the correctness information of the action sequence, the standardization information of the action, and the integrity information of the process, the execution status of each action is determined, and the corresponding action execution status is generated. The execution states of the action include normal execution state, abnormal execution state, omitted execution state, redundant execution state, and interrupted execution state, specifically: When the action standardization information meets the preset standardization threshold and the sequence deviation value corresponding to the action sequence correctness information meets the sequence consistency requirement, the action execution state of the corresponding action is determined to be the normal execution state. When the action standardization information is lower than the preset standardization threshold, or when the sequence deviation value corresponding to the action sequence correctness information does not meet the sequence consistency requirement, the action execution state of the corresponding action is determined to be an abnormal execution state. When it is determined, based on the process integrity information in the time sequence action recognition result, that there is a corresponding standard action in the structured operation sequence but no corresponding action is matched in the action time sequence, the action execution status of the corresponding action is determined to be the omission execution status. When it is determined, based on the process integrity information, that there are actions in the action time sequence that do not belong to the standard action information in the structured operation sequence, the action execution state of the corresponding action is determined to be a redundant execution state. When the action status information in the timing action recognition result includes an action interruption status, the action execution status of the corresponding action is determined to be an interrupted execution status. The user operation action information is generated by associating and integrating the action category, action occurrence time, action execution order, and action execution status.

7. The automatic evaluation method for practical training based on artificial intelligence according to claim 6, characterized in that, The terminal data is cleaned, standardized, and validated for consistency to obtain terminal operation information, specifically including: Acquire the operation terminal data generated by the operation terminal during the training process, wherein the operation terminal data includes operation instruction records, parameter setting records, device status records, and execution feedback records; The operation terminal data is cleaned to remove duplicate, abnormal, and invalid data, and missing data is filled in to obtain cleaned operation terminal data. The cleaned operation terminal data is processed for field unification and format conversion. The operation instruction records, parameter setting records, device status records and execution feedback records are standardized to obtain standardized operation terminal data. Based on the standardized operation terminal data, the operation instruction records, parameter setting records, device status records, and execution feedback records are associated in chronological order and sorted based on timestamps to generate terminal operation information. Based on the terminal operation information, the consistency of the correspondence between the operation instruction record and the execution feedback record is verified, and the consistency of the correspondence between the parameter setting record and the device status record is verified. The consistency verification includes: Perform a consistency check between the operation instruction record and the execution feedback record to determine whether there is a corresponding execution feedback record; Determine whether the parameter setting records and device status records meet the consistency requirement of parameter setting changes and device status changes; The temporal order and temporal continuity between adjacent records in the terminal operation information are verified; Based on the consistency verification results, operation records that do not meet the consistency conditions are corrected or completed, and the terminal operation information is updated.

8. The automatic evaluation method for practical training based on artificial intelligence according to claim 7, characterized in that, Based on the association and matching of the user operation information and the terminal operation information, a structured operation sequence corresponding to the user's training process is generated, specifically including: Obtain the action category, action occurrence time, action execution order, and action execution status from the user operation action information; The operation instruction records, parameter setting records, device status records, and execution feedback records in the terminal operation information are obtained. The user operation information and the terminal operation information are time-aligned based on the timestamp to establish a correspondence between the time of the action and the corresponding time of the operation instruction record; Based on the correspondence, the action category is matched with the operation instruction record, and the consistency of the action execution status, execution feedback record and device status record is verified to obtain the matching relationship between user operation action information and terminal operation information. Based on the matching relationship, user operation information with corresponding relationships is fused with terminal operation information to generate operation records in a structured operation sequence. Each of the operation records includes an action category, action occurrence time, action execution sequence, action execution status, and corresponding operation instruction record, device status record, and execution feedback record; The operation records are sorted based on the time of the action, and the execution order of the actions is verified to establish the sequential relationship between the operation records and generate a structured operation sequence. The structured operation sequence includes multiple operation records arranged in chronological order.

9. The automatic evaluation method for practical training based on artificial intelligence according to claim 8, characterized in that, Based on the structured operation sequence, the evaluation results for operation steps, operation standardization, and result correctness are obtained, specifically including: Obtain multiple operation records from the structured operation sequence, wherein each operation record includes an action category, action occurrence time, action execution order, action execution status, and corresponding operation instruction record, device status record, and execution feedback record; Obtain the pre-configured training evaluation rule package corresponding to the current training task. The training evaluation rule package is pre-built and stored and is used to uniformly evaluate the structured operation sequence. The training evaluation rule package includes operation step assessment rules, operation standardization assessment rules, and result correctness assessment rules. Based on the operation step assessment rules, the operation records in the structured operation sequence are matched with the standard operation sequence based on time order constraints to identify operation order deviations, operation omissions and operation redundancies, and operation step assessment scores are generated according to the preset scoring and deduction strategies in the operation step assessment rules to form operation step assessment results. Based on the operational standardization assessment rules, the action execution status and execution process of each operation record in the structured operation sequence are analyzed for compliance with the standardization rules. Operation records that do not meet the standard execution constraints are identified, and operational standardization assessment scores are generated according to the preset scoring and deduction strategies in the operational standardization assessment rules, thus forming operational standardization assessment results. Based on the result correctness assessment rules, the execution feedback records and device status records of each operation record in the structured operation sequence are compared with the standard execution status information corresponding to the standard operation results to determine the result deviation type and deviation level, and the result correctness assessment score is generated according to the preset scoring and deduction strategy in the result correctness assessment rules to form the result correctness assessment result. Wherein: the operation step assessment result is used to characterize the degree of consistency between the actual operation sequence and the standard operation sequence; the operation standardization assessment result is used to characterize the degree of standardization compliance of each operation record during the execution process; the result correctness assessment result is used to characterize the degree of consistency between the actual operation execution result and the standard operation result.

10. The automatic evaluation method for practical training processes based on artificial intelligence according to claim 9, characterized in that, Based on the assessment results of the described operation steps, the assessment results of the operation standardization, and the assessment results of the correctness of the results, the user's practical training evaluation score is obtained, specifically including: The results of the operation steps assessment, the results of the operation standardization assessment, and the results of the result correctness assessment are input into a pre-trained and deployed integrated evaluation model of practical training assessment. The integrated evaluation model of practical training assessment is constructed based on historical practical training data samples during the training phase and is used to comprehensively evaluate the multi-dimensional assessment results. The training assessment integration evaluation model includes a task semantic evaluation channel, a capability and behavior evaluation channel, and an execution environment consistency evaluation channel. The task semantic evaluation channel uses job task description information corresponding to the standard operation sequence to represent task requirements semantically in vector form, and calculates task matching score based on attention weight mechanism. The capability and behavior evaluation channel performs feature mapping on the user's operational capabilities based on the user's operational behavior characteristics corresponding to the structured operation sequence, and calculates the capability evaluation score through a nonlinear mapping network. The execution environment consistency evaluation channel performs correlation modeling on the execution environment consistency based on the standard environment information corresponding to the training task and the operation terminal information generated during the user operation, and calculates the environment consistency score. The task matching score, ability evaluation score, and environmental consistency score are weighted and fused to obtain the comprehensive semantic evaluation score. Based on the comprehensive semantic evaluation score, and combined with the operation step assessment results, the operation standardization assessment results, and the result correctness assessment results, a weighted summary calculation is performed to generate the user training evaluation score. The user training evaluation score is used to represent the comprehensive evaluation result of the user's execution of operation steps, the degree of standardization of operation, and the correctness of operation results during the training process.