Railway operation real operation examination system based on AI image motion recognition
By using an AI image and motion recognition system, initial assessment scores and reaction evaluation coefficients for operational actions are generated, and a three-dimensional skill assessment model is constructed. This solves the problems of refined motion assessment and skill assessment in railway operations, and enables scientific and comprehensive personnel assessment and training support.
Patent Information
- Application Number
- CN202511590147.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-01-30
AI Technical Summary
Existing technologies are insufficient for the precise and quantitative assessment of railway workers' actions, especially the standardized assessment of critical safety actions, and they neglect the workers' skill level and reaction speed.
By using an AI image and motion recognition system, an initial assessment score for the work action is generated. Combining the motion recognition response evaluation coefficient and skill dimensions, a three-dimensional skill assessment model is constructed to analyze the efficiency of the work action assessment and set skill assessment standards at different levels.
It enables precise assessment of railway workers' movements, comprehensively considering both movement accuracy and time factors, providing scientific and comprehensive personnel assessment results, and supporting intelligent grading and training optimization.
Smart Images

Figure CN121436775A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, in particular to a railway operation implementation examination system based on AI image action recognition. BACKGROUND
[0002] The AI image action recognition can capture and recognize the key task actions in the communication and protection operation process, and automatically compare the standard actions in the protection operation process of the protection personnel to perform intelligent judgment.
[0003] In the training and examination process, the high-definition camera is used for action capture, the data is collected through the action capture system program, and the computer vision technology is used for action recognition. At the same time, the protective special clothes and protective articles of the examinee can also be recognized, and automatic and scientific scoring and examination can be realized according to the standard requirements of the current protection operation stage. In the prior art, the actions of the operation personnel are diverse and complex, and the action standards in different operation scenarios also differ. The existing image recognition technology cannot be directly applied to the railway operation implementation examination, and needs to be specially optimized and improved according to the characteristics of the railway operation. The existing technology often focuses on the integrity of the process, but ignores the skill level stratification and reaction speed evaluation of the operation personnel, and it is difficult to finely and quantitatively capture and analyze the key operation actions, especially for the standardization of the medium key safety actions in the operation.
[0004] In view of the above technical defects, a solution is proposed. SUMMARY
[0005] The purpose of the present application is to adjust the operation action time sequence according to the action time consumption, generate an initial operation action examination score, obtain the time information of the target personnel when the gesture feature sequence is collected, generate an action recognition interval data set, deeply analyze the action performance of the operation personnel from the time dimension, obtain an action recognition reaction evaluation coefficient based on the time reaction comprehensive evaluation, stratify the skill level, set different skill dimensions and action complexity, and construct a three-dimensional skill evaluation model. The initial operation action examination score, the action recognition reaction evaluation coefficient and the quantitative evaluation index of each skill dimension are combined to analyze the operation action examination efficiency, and to provide a reliable basis for intelligent grading of the implementation examination of the target personnel.
[0006] In order to achieve the above purpose, the present application adopts the following technical scheme: a railway operation implementation examination system based on AI image action recognition, comprising an AI action recognition module, an action matching module, a skill analysis module, a skill modeling module and an examination decision module. The AI action recognition module is configured to capture real-time actions of the target personnel through the camera, obtain an action capture data vector, and use a multi-scale 3D convolutional neural network to collect gesture feature sequences of the workers based on the action capture data vector to generate work assessment feature data. The action matching module is configured to identify action features in the assessment scene based on the work assessment feature data and action data in the standard work process, obtain an action recognition matching index, adjust the order of work action time based on the action time consumption, and generate an initial work action assessment score. The skill analysis module is configured to obtain the time when the gesture feature sequences of the target personnel are collected, obtain time points of each action recognition item based on different action recognition items, generate an action recognition interval data set, and perform comprehensive evaluation of time reaction to obtain an action recognition reaction evaluation coefficient. The skill modeling module is configured to perform skill level stratification based on the work assessment feature data, set skill dimensions and action complexities at different levels, construct a three-dimensional skill evaluation model, and generate a quantitative evaluation index of each skill dimension. The assessment decision module is configured to combine the initial work action assessment score, the action recognition reaction evaluation coefficient, and the quantitative evaluation index of each skill dimension to analyze the work action assessment efficiency, obtain a comprehensive assessment result of the work assessment, and intelligently grade the target personnel based on the work assessment.
[0007] Further, the gesture feature sequences of the workers are collected to generate work assessment feature data, and the specific process is as follows: S100, in the standard work process, predefined key action nodes, each node corresponds to a specific coordinate range in three-dimensional space, through camera calibration, 2D pixel coordinates are converted into 3D world coordinates, to determine the key points in the work assessment; S101, capture the video of the assessment site through an industrial-grade camera, synchronously obtain color images and depth information, detect 2D coordinates of key points, combine depth information to generate 3D coordinate vectors through triangulation, generate coordinate vectors in each time frame, and superimpose time stamps to form a time series data stream; Data preprocessing: apply Kalman filter to smooth the key point trajectory and eliminate jitter noise; scale the coordinates to the [0, 1] interval through Min-Max normalization to eliminate scale differences; use cubic spline interpolation to complete the missing frames to ensure data continuity and generate real-time action capture data vectors; S102, feature extraction is performed on the real-time action capture data vector, the whole process action is segmented through a sliding window, and each window outputs a feature vector. These vectors are spliced in chronological order to form a gesture feature sequence to generate work assessment feature data.
[0008] Further, the action data in the job scene of the constructed standard operation process is used to perform action feature recognition in the examination scene, and the specific process is as follows: S200, a standard action template library containing typical operations is constructed, each template containing spatial features, time features and semantic features; S201, different recognition standards are set according to different recognition scenes, and the scenes are divided according to the complexity of the action to obtain a standard template library in different scenes; S202, the real-time captured gesture feature sequence is matched with the standard template library in multiple scales, a two-stage matching strategy is used for dynamic matching, based on the aligned sequence, the spatial trajectory similarity, the time rhythm matching degree and the semantic logic consistency are calculated, and the action recognition matching index is generated by weighting.
[0009] Further, the action time is sorted according to the action time, and the initial examination score of the action is generated, and the specific process is as follows: S300, the time stamp of each key action node is extracted from the real-time action sequence, the continuous action frames are spliced in time sequence to form an action sequence containing time dimension; S301, the timeliness weight of different action nodes is allocated, the absolute deviation rate of actual time consumption and standard time consumption is calculated, the overtime action node is identified and the subsequent action sequence is reordered by using dynamic programming algorithm; S302, the action recognition matching index is mapped to the initial score, the matching index initial score and the time efficiency adjustment score are added, and the initial examination score of the action is generated.
[0010] Further, the time reaction comprehensive evaluation is performed to obtain the action recognition reaction evaluation coefficient, and the specific process is as follows: S400, based on the action recognition project defined by the implementation examination, the starting and ending time points of each project are automatically identified, for each action recognition project, the interval between adjacent time points is calculated, the standardized interval data is stored according to the project to form a structured interval data set; S401, the mean and standard deviation of all intervals of each project are calculated, and the deviation rate of actual time consumption and standard time consumption is calculated according to the key action point as the key action delay index; S402, the mean and standard deviation and the key action delay index are linearly combined, and different proportion coefficients are given to generate a comprehensive reaction evaluation coefficient.
[0011] Further, a three-dimensional skill evaluation model is constructed to generate quantitative evaluation indexes of each skill dimension, and the specific process is as follows: S500: Based on the performance evaluation feature data, extract key performance evaluation features from the preprocessed data, and classify the skill level according to the length of employment of the target personnel into initial skills, intermediate skills and advanced skills. S501. Based on the key assessment features of the degree of freedom of action, the distance of key point movement, and the allowed operation time, the complexity level is defined as three levels: low, medium, and high. The three complexity levels correspond to three different assessment scenarios. S502, the three-dimensional skill assessment model uses skill level, skill dimension, and motion complexity as three-dimensional coordinates to construct a dynamic assessment space. The characteristic data of the operation assessment is used as the input data of the three-dimensional skill assessment model. Data analysis is carried out through the dynamic assessment space to analyze the trajectory deviation rate, time deviation index, and safe operation rate. The score of each skill dimension under the corresponding skill level and motion complexity is calculated by weighted summation to generate a three-dimensional scoring matrix, and the quantitative assessment index of each skill dimension is obtained.
[0012] Furthermore, an efficiency analysis of the work action assessment is conducted to obtain a comprehensive assessment result, which includes the following: The initial assessment score of the task action and the action recognition response evaluation coefficient are combined with the quantitative evaluation indicators of each skill dimension to calculate the comprehensive efficiency score. The obtained comprehensive efficiency score is used to generate the comprehensive assessment result of the task. The system uses an intelligent grading system based on the overall efficiency score to determine the level of a student. The levels are divided into excellent, good, passable, and unsatisfactory based on the score. Establish a skills degradation early warning system. By analyzing the continuous assessment data of target personnel, identify the fluctuation pattern of skill level and make intelligent classification judgments. When the comprehensive score is ≥R, it is judged as excellent; when P≤comprehensive score≤R, it is judged as good; when T≤comprehensive score≤P, it is judged as qualified; and when the comprehensive score≤T, a re-examination is required. At the same time, an abnormal action detection mechanism is embedded to mark the operation that has the same type of error three times in a row.
[0013] Furthermore, the system also includes a feedback and optimization module, which generates motion correction suggestions by comparing the three-dimensional posture reconstruction results of the operator's actual movements with those of the standard movements, and overlays a visual comparison of the skeletal models of the standard movements and the trainee's movements.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The system of the present invention can accurately acquire motion capture data vectors by capturing the movements of target personnel, effectively collect the hand gesture feature sequences of operators, and generate work assessment feature data that accurately reflects the movement characteristics of operators.
[0015] 2. The work actions are sorted and adjusted according to their time consumption to generate an initial assessment score. This not only focuses on the accuracy of the actions but also takes into account the time factor of the actions, making the score more comprehensive and reasonable, and able to accurately reflect the performance of the workers in actual work.
[0016] 3. By acquiring the time information of the target personnel during the collection of gesture feature sequences, an action recognition interval dataset is generated. The action performance of the workers is analyzed in depth from the time dimension. Based on the comprehensive evaluation of time response, the action recognition response evaluation coefficient is obtained, which can accurately measure the time response ability of workers in different action recognition projects.
[0017] 4. The system stratifies skill levels, sets different skill dimensions and motion complexity for each level, and constructs a three-dimensional skill assessment model. This model can comprehensively assess the skill level of operators from multiple perspectives, avoiding the limitations of single-dimensional assessment. It generates quantitative assessment indicators for each skill dimension, providing specific and clear standards for assessment and training. By combining the initial assessment score of the operation, the action recognition response assessment coefficient, and the quantitative assessment indicators for each skill dimension, the system conducts an efficiency analysis of the operation assessment. It comprehensively considers the accuracy of the action, time factors, and skill dimension information, resulting in a more scientific and comprehensive overall assessment result. This provides a reliable basis for the intelligent grading and determination of the target personnel's practical assessment. Attached Figure Description
[0018] Figure 1 A schematic diagram of the overall system steps of the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1 like Figure 1 As shown, the railway operation practice assessment system based on AI image action recognition includes an AI action recognition module, an action matching module, a skill analysis module, a skill modeling module, and an assessment decision module. The AI motion recognition module is used to capture the real-time motion of target personnel through a camera, obtain motion capture data vectors, and use a multi-scale 3D convolutional neural network to collect the gesture feature sequence of the workers from the motion capture data vectors in order to generate work assessment feature data. The action matching module, based on the task assessment feature data, uses the action data in the task scenario of the constructed standard task process to identify the action features in the assessment scenario, obtain the action recognition matching index, and adjust the task action time sorting according to the action consumption time to generate the initial task action assessment score. The skills analysis module is used to acquire the time when the target person's gesture feature sequence is collected. Based on different action recognition projects, the time point of each action recognition project is obtained, an action recognition interval dataset is generated, and a comprehensive evaluation of time response is performed to obtain the action recognition response evaluation coefficient. The skills modeling module, based on the characteristics of the job assessment, stratifies the skill level to set different skill dimensions and action complexity for each level, constructs a three-dimensional skills assessment model, and generates quantitative assessment indicators for each skill dimension. The assessment decision module combines the initial assessment score of the work action with the action recognition response evaluation coefficient and the quantitative evaluation indicators of each skill dimension to conduct an efficiency analysis of the work action assessment, obtain a comprehensive assessment result, and make intelligent grading judgments for the target personnel's practical assessment.
[0021] The process of collecting hand gesture feature sequences from workers to generate performance evaluation feature data is as follows: S100. Predefine key action nodes in the standard operating procedure. Each node corresponds to a specific coordinate range in three-dimensional space. Convert 2D pixel coordinates into 3D world coordinates through camera calibration to determine the key points in the operation assessment. S101. Capture video of the assessment location using an industrial-grade camera, simultaneously acquire color images and depth information, perform 2D coordinate detection of key points, and generate 3D coordinate vectors by triangulation in combination with depth information. Generate 3D coordinate vectors for each time frame and superimpose timestamps to form a time-series data stream. Data preprocessing: Kalman filtering is applied to smooth the key point trajectory and eliminate jitter noise; Min-Max normalization is used to scale the coordinates to the [0,1] interval to eliminate scale differences; cubic spline interpolation is used to complete missing frames to ensure data continuity and generate real-time motion capture data vectors; Input devices: Use a camera such as ZED 2 or RGB-D sensor to capture color images and depth maps at 60FPS; Key point detection: Detect 21 key points of the hand, such as the arm, and output 3D coordinates; OpenPose was used to detect 18 key points on the torso to generate full-body skeletal data; Data preprocessing: Coordinate normalization, transforming the coordinates of key points to a local coordinate system with the center of the shoulder as the origin, to eliminate the influence of the camera's perspective; Time alignment: Unifying the timing of data at different frame rates using interpolation algorithms; Output: Generate motion capture data vectors, including 21 hand keypoints and 18 body keypoints, each with 3 dimensions; Multi-scale 3D convolutional feature extraction, network structure, input layer: receives motion capture data vector data such as T×39×3, where T is the length of the time window; 3D convolutional blocks: using 3×3×3 convolutional kernels to progressively extract local spatiotemporal features; Feature fusion: Multi-scale features are merged through 1×1×1 convolution to output a gesture feature sequence and generate assignment assessment feature data; S102. Extract features from the real-time motion capture data vectors, segment the entire process of motion through a sliding window, output a feature vector for each window, and concatenate these vectors in chronological order to form a gesture feature sequence and generate task assessment feature data.
[0022] By using action data from the established standard operating procedures in the work scenario, action feature recognition is performed in the assessment scenario. The specific process is as follows: The operator demonstrates the standard operating procedure, and a high-precision motion capture system is used to collect three-dimensional motion data. Key motion nodes are labeled for each standard motion sequence data, and the three-dimensional coordinate range, trajectory feature vector and timestamp of each motion node are extracted. Features are extracted using the same 3D convolutional network for standard actions and stored as D={F 标准,1 ,...,F 标准,J}; S200. Construct a standard action template library containing typical operations. Each template contains spatial features, temporal features, and semantic features. S201. Set different recognition standards according to different recognition scenarios, divide the scenarios by the complexity of the actions, and obtain a standard template library for different scenarios. Coarse matching: By classifying actions into basic operations, complex operations, and safety confirmations, it quickly locates potentially matching template clusters; Fine matching: Within the located cluster, cosine similarity is used to calculate the similarity between the real-time feature vector and each template, thereby selecting the corresponding standard template library; S202. The real-time captured gesture feature sequence is matched with the standard template library at multiple scales. A two-stage matching strategy is used for dynamic matching. Based on the aligned sequence, spatial trajectory similarity, temporal rhythm matching degree, and semantic logic consistency are calculated and weighted to generate an action recognition matching index.
[0023] The work actions are sorted and adjusted according to their duration to generate an initial assessment score. The specific process is as follows: S300: Extract the timestamp of each key action node from the real-time action sequence, and splice the continuous action frames in chronological order to form an action sequence containing the time dimension. S301. Assign timeliness weights to different action nodes, calculate the absolute deviation rate between actual time consumption and standard time consumption, identify time-out action nodes, and use dynamic programming algorithm to reorder subsequent action sequences. Dynamic temporal warping matching: Construct an N*M cost matrix and calculate the minimum cumulative distance between the evaluation feature and the standard feature. ; The assessment dataset is {F} 考核,1 ,...,F 考核,I}, the standard dataset is {F 标准,1 ,...,F 标准,J}; Matching index: Normalized distance to [0,1], where 1 indicates a perfect match; Action time statistics: Record the start frame Q1 and end frame Q2 for each action in the assessment; Efficiency weighting: For timeout actions Q1-Q2 exceeding the set threshold, the matching index weight is reduced. The threshold is set by combining the matching index and time weight, and is adjusted based on historical thresholds. S302. Map the action recognition matching index to the initial score, add the initial score of the matching index to the time efficiency adjustment score, and generate the initial assessment score of the operation action. ; Piecewise linear function is used: matching index ≥ 0.9 → 90-100 points (excellent); 0.8 ≤ Match Index < 0.9 → 80-89 points (Good); 0.7 ≤ Match Index < 0.8 → 70-79 points (Pass); A matching index of <0.7 (0-69 points) indicates a failure to meet the requirements and the need for improvement. The initial score is adjusted based on the time consumption deviation rate. If the time consumption deviation rate is ≤10%, the score remains unchanged; if 10% < deviation rate ≤20%, 5 points are deducted; if the deviation rate >20%, 10 points are deducted and a skill degradation warning is triggered.
[0024] A comprehensive evaluation of time-response time is performed to obtain the action recognition response evaluation coefficient. The specific process is as follows: S400: Based on the action recognition items defined in the practical assessment, automatically identify the start and end time points of each item, calculate the interval between adjacent time points for each action recognition item, and store the standardized interval data by item to form a structured interval dataset. Based on predefined action recognition items such as "gesture start", "operation execution", "safety confirmation" and "end reset", the start and end time points of each item are automatically identified through key point trajectory analysis algorithms such as peak detection based on speed threshold. For example, "gesture start" is defined as the time point when the key points of the hand go from a static state to the first significant acceleration. For each action recognition item, calculate the interval between adjacent time points, such as the "starting gesture - execution" interval and the "execution - confirmation" interval. For example, if the "gesture start" time is t1=1.2s and the "operation execution" time is t2=3.8s, then the interval Δt=2.6s. S401. Calculate the mean and standard deviation of all intervals for each project, and calculate the deviation rate between the actual time and the standard time based on the key action points, as the key action delay index. S402. Linearly combine the mean J, standard deviation B, and key action delay index C, assign different proportion coefficients, and generate a comprehensive response evaluation coefficient G. G = 0.3 * J + 0.3 * B + 0.4 * C.
[0025] A three-dimensional skills assessment model is constructed, and quantitative assessment indicators for each skills dimension are generated. The specific process is as follows: S500: Based on the performance evaluation feature data, extract key performance evaluation features from the preprocessed data, and classify the skill level according to the length of employment of the target personnel into initial skills, intermediate skills and advanced skills. S501. Based on the key assessment features of the degree of freedom of action, the distance of key point movement, and the allowed operation time, the complexity level is defined as three levels: low, medium, and high. The three complexity levels correspond to three different assessment scenarios. S502, the three-dimensional skill assessment model uses skill level, skill dimension, and motion complexity as three-dimensional coordinates to construct a dynamic assessment space. The characteristic data of the operation assessment is used as the input data of the three-dimensional skill assessment model. Data analysis is carried out through the dynamic assessment space to analyze the trajectory deviation rate, time deviation index, and safe operation rate. The scores of each skill dimension under the corresponding skill level and motion complexity are calculated by weighted summation to generate a three-dimensional scoring matrix, and the quantitative assessment index of each skill dimension is obtained. Action complexity classification: simple action: single-step operation, time ≤ 2 seconds, key point movement distance ≤ 0.1m; Medium-sized actions: multi-step sequence (such as changing parts), taking 2-10 seconds, with key point movement distance of 0.1-0.5m; Complex actions: multi-step parallel operation + conditional branching, time ≥ 10 seconds, key point movement distance ≥ 0.5m; Dynamic mapping of 3D models: Step 1: Filter input data based on action complexity (Z-axis) and compare the complexity of the actions; Step 2: Calculate quantitative indicators on the skill dimension (Y-axis), including trajectory deviation rate, time deviation index, and safety compliance; Step 3: Set indicator thresholds according to skill level (X-axis), such as TDI ≤ 20% for intermediate skill level and trajectory deviation rate ≤ 10% for expert level; Step 4: Generate a three-dimensional scoring matrix. The score for each skill dimension under the corresponding skill level and action complexity is calculated by weighted summation.
[0026] Efficiency analysis of work actions is conducted to obtain comprehensive work assessment results, including the following: The initial assessment score of the task action and the action recognition response evaluation coefficient are combined with the quantitative evaluation indicators of each skill dimension to calculate the comprehensive efficiency score. The obtained comprehensive efficiency score is used to generate the comprehensive assessment result of the task. The system uses an intelligent grading system based on the overall efficiency score to determine the level of a student. The levels are divided into excellent, good, passable, and unsatisfactory based on the score. Establish a skills degradation early warning system. By analyzing the continuous assessment data of target personnel, identify the fluctuation pattern of skill level and make intelligent classification judgments. When the comprehensive score is ≥R, it is judged as excellent; when P≤comprehensive score≤R, it is judged as good; when T≤comprehensive score≤P, it is judged as qualified; and when the comprehensive score≤T, a re-examination is required. At the same time, an abnormal action detection mechanism is embedded to mark the operation that has the same type of error three times in a row.
[0027] It includes a feedback and optimization module, which generates motion correction suggestions by comparing the 3D posture reconstruction results of the operator's actual movements with those of the standard movements, and overlays a visual comparison of the skeletal models of the standard movements and the trainees' movements. Generate structured suggestions such as a voice prompt: "Adjust the swing of your right arm to shoulder height"; Visual guidance: By overlaying virtual benchmark lines using AR, the correct shooting angle can be determined; Difference heatmap: Joint position deviations are marked using color mapping; Trajectory Comparison: Synchronously plays the skeletal motion trajectories of standard movements and trainees' movements, supporting frame-by-frame comparison; Error labeling: Automatically identifies typical errors such as "kicking the ball with the toes" and highlights the error location.
[0028] The size of the interval and threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by those skilled in the art for each set of sample data; as long as it does not affect the ratio between the parameter and the quantized value.
[0029] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation. In the two embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways; for example, the system embodiments described above are merely illustrative, for example, the division of modules is merely a logical functional division, and there may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed; another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms; The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1.A railway operation practice assessment system based on AI image action recognition, characterized in that, Comprise AI action recognition module, action matching module, skill analysis module, skill modeling module, examination decision module; The AI action recognition module is used for real-time action capture of the target personnel through the camera, obtains an action capture data vector, and adopts a multi-scale 3D convolutional neural network to collect a gesture feature sequence of the operation personnel, so as to generate operation examination feature data; The action matching module is based on the operation examination feature data, and the action data in the operation scene of the standard operation process is constructed to identify the action features in the examination scene, obtain an action recognition matching index, adjust the operation action time sequence according to the action time consumption, and generate an operation action initial examination score; The skill analysis module is used for acquiring the time when the gesture feature sequence of the target personnel is collected, obtaining the time point of each action recognition item based on different action recognition items, generating an action recognition interval data set, and comprehensively evaluating the time reaction to obtain an action recognition reaction evaluation coefficient; The skill modeling module is based on the operation examination feature data, and the skill level is layered to set different levels of skill dimensions and action complexity, construct a three-dimensional skill evaluation model, and generate a quantitative evaluation index of each skill dimension; The examination decision module combines the operation action initial examination score, the action recognition reaction evaluation coefficient, and the quantitative evaluation index of each skill dimension to analyze the operation action examination efficiency, obtains an operation examination evaluation comprehensive result, and intelligently grades and judges the target personnel according to the actual examination. 2.The railway operation practice assessment system based on AI image action recognition according to claim 1, wherein The gesture feature sequence of the operation personnel is collected to generate operation examination feature data, and the specific process is as follows: S100, in the standard operation process, the key action nodes are predefined, each node corresponds to a specific coordinate range in the three-dimensional space, the 2D pixel coordinates are converted into 3D world coordinates through camera calibration to determine the key points in the operation examination; S101, capture the video of the examination site through an industrial-grade camera, synchronously acquire color images and depth information, detect 2D coordinates of key points, generate 3D coordinate vectors through triangulation combined with depth information, generate coordinate vectors in each time frame, and superimpose time stamps to form time series data flow; Data preprocessing: apply Kalman filter to smooth the key point trajectory and eliminate jitter noise; Scale the coordinates to the [0, 1] interval through Min-Max normalization to eliminate scale differences; Complete the missing frames by cubic spline interpolation to ensure data continuity and generate real-time action capture data vectors; S102, feature extraction is performed on the real-time action capture data vector, the whole process action is segmented through a sliding window, and each window outputs a feature vector. These vectors are spliced in time sequence to form a gesture feature sequence to generate operation examination feature data. 3.The railway operation practice assessment system based on AI image action recognition according to claim 1, characterized in that, The action data in the operation scene of the standard operation process is constructed to identify the action features in the examination scene, and the specific process is as follows: S200, a standard action template library containing typical operations is constructed, each template contains spatial features, time features and semantic features; S201, different recognition standards are set according to different recognition scenes, the scenes are divided by the complexity of the action, and a standard template library under different scenes is obtained; S202, the real-time captured gesture feature sequence is matched with the standard template library in multiple scales, a two-stage matching strategy is adopted for dynamic matching, based on the aligned sequence, the spatial trajectory similarity, the time rhythm matching degree and the semantic logic consistency are calculated, and an action recognition matching index is generated by weighting. 4.The railway operation practice assessment system based on AI image action recognition of claim 1, wherein The action time sequence is adjusted according to the action time consumption, and an initial assessment score of the action is generated. The specific process is as follows: S300, the time stamp of each key action node is extracted from the real-time action sequence, the continuous action frames are spliced in time sequence to form an action sequence containing time dimension; S301, the time efficiency weight is allocated to different action nodes, the absolute deviation rate of actual time consumption and standard time consumption is calculated, the overtime action node is identified, and the subsequent action sequence is reordered by using dynamic programming algorithm; S302, the action recognition matching index is mapped to the initial score, the matching index initial score is added to the time efficiency adjustment score, and the initial assessment score of the action is generated. 5.The railway operation practice assessment system based on AI image action recognition according to claim 1, wherein, The time reaction comprehensive evaluation is carried out, and an action recognition reaction evaluation coefficient is obtained. The specific process is as follows: S400, based on the action recognition items defined in the implementation assessment, the starting and ending time points of each item are automatically identified, the interval between adjacent time points is calculated for each action recognition item, the standardized interval data is stored according to the items, and a structured interval data set is formed; S401, the mean and standard deviation of all intervals of each item are calculated, and the deviation rate of actual time consumption and standard time consumption is calculated according to the key action points as the key action delay index; S402, the mean and standard deviation and the key action delay index are linearly combined, different proportion coefficients are given, and a comprehensive reaction evaluation coefficient is generated. 6.The railway operation practice assessment system based on AI image action recognition according to claim 1, wherein A three-dimensional skill evaluation model is constructed, and a quantitative evaluation index of each skill dimension is generated. The specific process is as follows: S500, based on the characteristics of the implementation assessment, the key assessment characteristics are extracted from the preprocessed data, the skill level is stratified according to the length of the on-the-job time of the target personnel, and is divided into initial skill, intermediate skill and advanced skill; S501, based on the key assessment characteristics of action freedom, key point moving distance and allowed operation time, the complexity level is defined, which is divided into low, medium and high levels, and the three complexity levels correspond to three different assessment scenes; S502, the three-dimensional skill evaluation model takes skill level, skill dimension and action complexity as three-dimensional coordinates to construct a dynamic evaluation space, takes the implementation assessment characteristics as input data of the three-dimensional skill evaluation model, analyzes the data in the dynamic evaluation space, analyzes the trajectory deviation rate, time deviation index and safe operation rate, calculates the score of each skill dimension under the corresponding skill level and action complexity by weighted summation, generates a three-dimensional score matrix, and obtains a quantitative evaluation index of each skill dimension. 7.The railway operation practice assessment system based on AI image action recognition according to claim 1, wherein, The efficiency of the action assessment is analyzed, and the comprehensive result of the action assessment is obtained, which includes the following: The initial assessment score of the operation action is combined with the action recognition reaction evaluation coefficient and the quantitative evaluation index of each skill dimension to calculate a comprehensive efficiency score, and the obtained comprehensive efficiency score is used to generate a comprehensive evaluation result of the operation assessment; The intelligent grading is determined according to the comprehensive efficiency score to obtain the level, and the level is divided based on the score, and the level is divided into excellent, good, pass and unqualified; The skill degradation early warning is established, the skill level fluctuation mode is identified by analyzing the continuous assessment data of the target personnel, the intelligent grading is determined, when the comprehensive score is greater than or equal to R, it is determined as excellent, when P is less than or equal to the comprehensive score and greater than or equal to R, it is determined as good, when T is less than or equal to the comprehensive score and greater than or equal to P, it is determined as qualified, and when the comprehensive score is less than or equal to T, re-examination is needed, and at the same time, the abnormal action detection mechanism is embedded, and the same type of failure operation appearing for three times is specially marked. 8.The railway operation practice assessment system based on AI image action recognition of claim 1, wherein, The system further comprises a feedback and optimization module, which generates action correction suggestions by comparing the three-dimensional posture reconstruction results of the actual action of the operation personnel and the standard action, and superimposes the skeleton model comparison visualization diagram of the standard action and the action of the student.