ST-GCN-based Monitoring Method and Device for the Operation Normativity of Rail Transit Drivers

Through the improved ST-GCN model, combined with dynamic topology and domain prior fusion layer, real-time monitoring of rail transit driver operations is solved, and the accuracy and real-time problems of traditional methods in complex scenarios are improved, and the intelligence and safety of driver operation supervision is improved.

CN119832527BActive Publication Date: 2025-07-08天津致新轨道交通运营有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411894836.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-07-08
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Traditional methods are difficult to fully respond to the complex scenarios and diversified challenges in rail transit driver operation supervision, especially in the identification of continuous behaviors, which affects the real-time and accuracy of driver operation norms and leads to safety operation risks.

Method used

Using the improved ST-GCN model, the dynamic topology and domain prior fusion layer are constructed, combined with the joint reliability weighting module, driver operations are monitored in real time, and normative evaluation and early warning are output.

Benefits of technology

It has improved the intelligence and automation level of driver operation supervision, ensured that the operation meets safety requirements, reduced human errors and safety hazards, and improved the efficiency of rail transit safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832527B_ABST
    Figure CN119832527B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of event detection and behavior norm supervision, and discloses a monitoring method and device for the operation standardization of rail transit drivers based on an improved ST-GCN model. The method includes: determining a variety of key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrating the key joint points corresponding to each key standard operation; collecting video data of the driver performing the various key standard operations in the vehicle and external environments, and evaluating and annotating the compliance of the driver's operation behavior in the video data; constructing an ST-GCN model based on a dynamic topology and domain prior fusion layer, and training the model using the evaluated and annotated video data; and using the trained ST-GCN model based on dynamic topology and domain prior fusion to monitor the driver's operation in real time and output a prediction result. The device and method can improve the real-time performance and accuracy of supervision and ensure the safe operation of rail transit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of event detection and behavior norm supervision, and particularly relates to a method and device for monitoring the operation standardization of rail transit drivers based on ST-GCN. Background Art

[0002] In the rail transit system, the supervision of the operation standardization of train drivers is crucial because every operation decision of the driver directly affects the safety and operation stability of the train. Therefore, ensuring that the driver strictly complies with the operation specifications, especially in the operations of before driving, during driving, and after parking, is crucial for avoiding accidents and improving the safety of train operation.

[0003] To supervise the operation standardization of drivers, traditional methods mainly rely on manual supervision and regular inspections. However, manual supervision has certain limitations. It is difficult to achieve real-time and comprehensive monitoring and is easily affected by human factors. With the continuous expansion of the rail transit system and the increasing operation pressure, the effectiveness of manual supervision gradually decreases and it is difficult to meet the high requirements of modern rail transit for safety monitoring. In recent years, automated behavior monitoring methods based on image data have been gradually applied to other industries, such as the safety monitoring of workers at construction sites, providing a reference for the operation supervision of rail transit drivers. These methods usually adopt pose estimation techniques, such as OpenPose, combined with spatio-temporal graph convolutional networks (ST-GCN) for behavior recognition and standardization detection. However, when the traditional OpenPose+ST-GCN method is applied to the operation monitoring of rail transit drivers, many challenges still exist. Most of the existing monitoring methods perform behavior recognition based on single-frame images. Although they can accurately identify objects and actions, it is difficult to capture the temporal changes of behaviors. For example, the inspection actions before driving by the driver usually consist of multiple consecutive actions, and a single image cannot reflect the coherence of these actions, resulting in difficulty in judging whether the operation is complete and standardized.

[0004] The pose estimation method based on a single moment is difficult to understand the context relationship of behaviors, affecting the comprehensive evaluation of the operation standardization of drivers. Complex operation processes require the system to have a profound understanding of the action sequence and association, while traditional methods perform poorly in this regard.

[0005] In addition, data collection methods based on sensors, such as wearable sensors, although they can provide real-time physiological data and behavior information, also have limitations. The complexity of multi-sensor data fusion and the inconvenience to drivers make these methods difficult to be widely applied to the operation supervision of rail transit drivers. Especially in the recognition of continuous behaviors, it is very difficult to accurately judge multiple instantaneous actions (such as checking whether the door is closed, looking down at the screen, etc.) without temporal analysis.

[0006] Therefore, due to the above limitations, it is difficult for traditional methods to comprehensively handle complex scenarios and diverse challenges in the operation supervision of rail transit drivers. There is an urgent need for a monitoring solution that can more effectively capture continuous actions and temporal information to improve the real-time performance and accuracy of supervision and ensure the safe operation of rail transit. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, the present invention provides a method and device for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model, aiming to solve the technical problems that traditional methods are difficult to comprehensively handle complex scenarios and diverse challenges in the operation supervision of rail transit drivers, difficult to achieve real-time, accurate, and comprehensive monitoring and evaluation of the operation standardization of drivers, and bring risks to the safe operation of rail transit.

[0008] The object of the present invention can be achieved through the following technical solutions:

[0009] In a first aspect, the present invention provides a method for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model, including:

[0010] Determine various key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrate the key joint points corresponding to each key standard operation;

[0011] Collect video data of the driver performing the various key standard operations in the vehicle interior and exterior environments, and evaluate and annotate the compliance of the driver's operation behavior in the video data;

[0012] Construct an ST-GCN model based on the fusion layer of dynamic topology and domain prior, and train the model using the evaluated and annotated video data;

[0013] Use the trained ST-GCN model based on the fusion of dynamic topology and domain prior to monitor the driver's operation in real time and output the prediction results.

[0014] Preferably, the various key standard operations are five, including: whether the door is fully closed, whether there are items or people stuck in the carriage, safety inspection of the vehicle interior and exterior environments, inspection of the train braking system, and standardized confirmation of the train start operation.

[0015] Preferably, the constructing an ST-GCN model based on the fusion layer of dynamic topology and domain prior and training the model using the evaluated and annotated video data includes:

[0016] Perform human pose estimation and skeleton data extraction on the evaluated and annotated video data, where the skeleton data includes the key joint point coordinates and the corresponding joint point confidence levels;

[0017] Dynamically adjust the weights of each joint point using the joint confidence weighted module according to the joint point confidence and domain prior knowledge to obtain a weighted bone sequence;

[0018] Input the weighted bone sequence into the ST-GCN model for training, including:

[0019] Perform domain adaptive skeleton normalization on the weighted bone sequence, adjust the skeleton proportion based on the driver's body type parameters to obtain a standardized bone sequence;

[0020] Input the normalized standardized bone sequence into the ST-GCN model for feature extraction. During the ST-GCN feature extraction process, through a learnable dynamic topology mechanism, adaptively adjust the joint connection relationship according to the current time series segment and candidate operation types, embed a domain prior fusion layer in the middle layer, fuse the normalized standardized bone sequence with the canonical operation prior vector to guide the features to be close to the standard operation mode in the rail transit field; among the 9 convolutional blocks of ST-GCN, a residual learning structure is retained after each convolutional block, and a dynamic topology and domain prior fusion layer is inserted between some blocks to dynamically optimize feature extraction;

[0021] Finally, calculate the probability that the input activity belongs to a specific canonical operation through the SoftMax function, calculate the loss with the label truth value, and complete the model training.

[0022] Preferably, wherein, the human pose estimation and skeleton data extraction for the evaluated and annotated video data include:

[0023] Use OpenCV to convert the evaluated and annotated video frame by frame into an image sequence and perform preprocessing;

[0024] Input each preprocessed image into a pre-trained convolutional neural network based on VGG-19 to extract the high-dimensional convolutional feature map of the image;

[0025] Input the high-dimensional convolutional feature map into the multi-stage network of OpenPose to generate the skeleton data of the human body in each frame of the image. The skeleton data includes: the coordinate information of each key joint point of the human body and the confidence score of the detection result of each key node. These key node coordinate information includes but is not limited to the two-dimensional position coordinates of the shoulders, elbows, wrists, knees, and ankles.

[0026] Preferably, the dynamically adjusting the weights of each joint point using the joint confidence weighted module according to the joint point confidence and domain prior knowledge to obtain a weighted bone sequence includes:

[0027] Obtain the key joint point coordinates and joint point confidence output by OpenPose;

[0028] According to the pre-calibrated key joints and domain prior importance, dynamic weights are assigned in combination with the joint confidence; among them, higher weights are assigned to the joints that play a key role and have high confidence in the current standard operation, and lower weights are assigned to the joints with low confidence or irrelevance.

[0029] According to the coordinates of the key joints and the dynamically assigned weights, a weighted skeleton sequence is constructed and output.

[0030] Preferably, the multi-stage network of OpenPose adopts a multi-stage, two-branch structure, where Branch 1 is used to estimate the probability distribution of the existence of human key joints and generate a confidence map for each key joint point; Branch 2 is used to calculate the part affinity field between joints to determine which joints belong to the same person and reconstruct the skeleton topology; through multiple iterative stages, OpenPose continuously optimizes the confidence map and the part affinity field, and finally outputs the coordinates of the key joint points and the corresponding joint confidence scores of each frame of image.

[0031] Preferably, the dynamic topology mechanism specifically includes:

[0032] A learnable graph attention layer is introduced before or in the spatio-temporal convolution block for feature extraction in the spatio-temporal graph convolutional network.

[0033] According to the current time series segment and the candidate operation type, the adjacency relationship weights between joint points are adaptively calculated; higher weights are given to the connections of joint points with high relevance in key operations, and the edges of joint points irrelevant to this operation are weakened.

[0034] Preferably, the method includes:

[0035] In the domain prior fusion layer, a prior vector corresponding to each standard operation is defined in advance based on expert knowledge.

[0036] The spatio-temporal features extracted from the middle layer of ST-GCN are fused with the prior vector to make the features closer to the standard action patterns in the rail transit field.

[0037] Preferably, the real-time monitoring of the driver's operation using the trained ST-GCN model based on dynamic topology and domain prior fusion and outputting the prediction result includes:

[0038] The trained ST-GCN model outputs the operation probability corresponding to the driver's operation in real time.

[0039] Convert the operation probability into a standard operation score.

[0040] Compare the standard operation score with a preset safety threshold. When the score is lower than the threshold, give a warning prompt, record the score at the same time, and formulate a personalized incentive strategy based on the historical trend.

[0041] In a second aspect, the present invention also provides a monitoring device for the operation standardization of rail transit drivers based on an improved ST-GCN model, including:

[0042] An operation definition unit, configured to determine various key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrate the key joint points corresponding to each key standard operation;

[0043] A video acquisition unit, configured to acquire video data of the driver performing the various key standard operations in the vehicle interior and exterior environments, and evaluate and annotate the compliance of the driver's operation behavior in the video data;

[0044] A model construction unit, constructing an ST-GCN model based on a dynamic topology mechanism and a domain prior fusion layer, and training the model using the evaluated and annotated video data;

[0045] A result output unit, configured to use the trained ST-GCN model based on dynamic semantic topology and domain prior fusion to monitor the driver's operation in real time and output a prediction result.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] 1. The joint credibility weighted module SCWM of the present invention dynamically weights the OpenPose results and combines a spatio-temporal graph convolutional network (ST-GCN) based on dynamic topology to detect the operation behavior of rail transit train drivers in real time, and can accurately evaluate whether the driver complies with the operation specifications, thereby improving the intelligence and automation level of operation supervision;

[0048] 2. The present invention monitors and outputs the probability of the driver's standard operation in real time, converts it into a standard operation score, and classifies it according to the set threshold, and can give a warning or reward in time to ensure that the driver's operation meets the safety requirements and reduce human errors and potential safety hazards;

[0049] 3. The present invention uses domain prior information and a personalized normalization strategy, which can not only accurately monitor the compliance of actions, but also maintain high discrimination and universality among different drivers, and finally improve the overall efficiency and level of rail transit safety management. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0051] Figure 1 It is a schematic flowchart of a method for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model provided by an embodiment of the present invention;

[0052] Figure 2 It is a network architecture diagram of a method for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model provided by an embodiment of the present invention;

[0053] Figure 3 It is an architecture diagram of OpenPose of the present invention;

[0054] Figure 4 It is an architecture diagram of the ST-GCN based on the fusion of dynamic topology and domain prior of the present invention;

[0055] Figure 5 It is a unit module diagram of a device for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model provided by an embodiment of the present invention. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0058] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the recited number, and above, below, within, etc. are understood as including the recited number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0059] Please refer to Figures 1-4 , an embodiment of the present invention provides a method 100 for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model, including:

[0060] S110: Determine various key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrate the key key points corresponding to each key standard operation;

[0061] In this step S110, first, determine the standard operation actions and the importance of joints. In this embodiment, by comprehensively considering the operation specifications of urban rail transit train drivers and the suggestions of traffic safety experts, various common key standard operations taken by the driver during the train start and stop processes are determined. Preferably, five common key standard operations are adopted, including: whether the doors are fully closed, whether there are objects or people stuck inside the car, safety inspections of the internal and external environments of the car, inspections of the train braking system, and standardized confirmation of the train start operation.

[0062] These five common key standard operations are the key links to ensure the safety, accuracy, and efficiency of the driver's operations. The following are the specific operation specifications:

[0063] 1. Before starting the train, check whether the doors are fully closed. The driver confirms whether all doors are fully closed and locked to avoid accidental door opening during the start process and ensure the safety of passengers.

[0064] 2. Confirm that there are no objects or people stuck inside the car. Check whether there are objects or people stuck inside the car to ensure that the doors close normally and avoid being affected by stuck objects or people during the start.

[0065] 3. Conduct a final safety inspection of the internal and external environments of the car. The driver checks whether there are safety hazards inside and outside the car to ensure that there are no obstacles or other threats around the train and guarantee the safe start of the train.

[0066] 4. Check and ensure that the train braking system is normal. The driver checks whether the braking system of the train is working properly to ensure that it can stop quickly and safely when needed.

[0067] 5. Start the train according to the standard operation process. The driver confirms all safety measures step by step according to the regulations to ensure that the train starts in accordance with the operation specifications and avoid any safety loopholes.

[0068] When determining these five key operations, the expert group simultaneously pre-calibrates the key key points and action patterns involved in each operation. For example, for the operation involving door inspection, the key degrees of the arm and head key points can be pre-calibrated to provide a domain prior basis for subsequent joint credibility weighting. Specifically, the movements of the upper limbs (shoulder, elbow, wrist) and the head (neck joint), the direction of the gaze, and the hand postures are particularly important; the "environmental safety inspection" step pays more attention to the rotation of the driver's head and neck joints and the unique patterns generated during the touch or indication process of the upper limbs. These prior knowledge will provide reference for the later data processing stage (such as joint credibility weighting and dynamic topology construction).

[0069] In this step, five operations are clearly associated with the importance of specific joint points, so as to be targeted in the subsequent weighted processing of skeleton data.

[0070] S120: Collect video data of the driver performing the above-mentioned multiple key specification operations in the vehicle interior and exterior environments, and evaluate and annotate the compliance of the driver's operation behaviors in the video data;

[0071] This step S120 includes two main steps:

[0072] (1) Experimental environment construction and data collection

[0073] Build an experimental environment for train stopping and starting when approaching a station, which is divided into two areas: inside the vehicle and outside the vehicle. Specifically, for the inside-vehicle area: install high-resolution cameras at positions such as in front of, on the side of, and on the top of the train driver's cab, so that the driver's upper body, upper limbs, facial orientation, and arm movements can be clearly captured; arrange auxiliary landmark points around the instrument panel, door controller, and seat to facilitate the determination of the correspondence between joint movements and specific control operations in the later stage. For the outside-vehicle area: arrange one camera on each of the side, front, and rear directions of the train body to record the driver's actions of poking his head to check the platform conditions or the outside-vehicle environment from different angles. When the driver needs to turn his head, reach out to indicate, or use the observation window to check the environment, these cameras will provide supplementary perspectives for subsequent OpenPose to extract skeleton features and reduce the single-perspective occlusion problem.

[0074] Recruit train drivers to participate in the experiment, and ensure that the participants have three or more years of driving experience. Before each driver enters the experimental environment, collect and register basic physical sign data such as height and arm length to support subsequent personalized adaptive normalization, specifically Domain Adaptive Skeleton Normalization (DASN).

[0075] During the experiment, each driver needs to perform operations in sequence according to the five operation specifications defined in step S110. To improve the diversity and robustness of the data: each driver conducts multiple rounds of experiments, and takes a three-minute break after each round of experiments to reduce operation inaccuracies caused by fatigue; in some experiments, the interior lighting of the vehicle can be adjusted (simulating different platform environments), the camera position can be slightly adjusted, or the marker layout in the outside-vehicle area can be changed to test the adaptability of the model to environmental changes; each round of experiment is video-recorded in AVI format at a frame rate of 30fps, with a resolution of 532×300. The video duration is about 20 - 30 seconds, covering the complete start-stop preparation process and specification operation steps.

[0076] (2) Evaluation and annotation of video data

[0077] Rail transit safety experts evaluate and annotate the compliance of drivers' operation behaviors based on the recorded videos, providing high-quality supervision labels for model training.

[0078] Specifically, experts review the action quality and compliance degree of each driver in five operations frame by frame, and give scores and labels (such as "fully compliant", "partially compliant", "non-compliant") for each standard operation.

[0079] During the expert annotation process, the joint point features related to the standard operations are further recorded. For example, in the "door inspection" operation, the joint points most relevant to this step and their typical movement trajectories are noted. These annotation data will provide additional guidance for subsequent domain prior fusion (Domain Prior Fusion Layer (DPFL)) and dynamic topology (DT) construction.

[0080] After the data annotation is completed, quality control and statistical analysis are performed on all the data. The average duration of each driver's video is about 22.3 seconds, and the average duration of each single operation is 3.5 seconds. This can ensure that in subsequent training, ST-GCN can fully learn the spatio-temporal features of continuous skeletal sequences. In case of situations where it is difficult to identify key joints due to excessive occlusion of the driver's posture or atypical actions in the video, experts will add a "low confidence" label to this section of the video to guide the joint confidence weighting module to dynamically reduce the influence of this joint point on feature modeling during training.

[0081] Through the above refined data collection and annotation strategy, it provides a high-quality, multi-perspective information source with rich prior annotations for subsequent methods (including joint confidence weighting, personalized skeleton normalization, dynamic topology construction, and domain prior fusion), enabling the model to more accurately and robustly identify the standard operation behaviors of rail transit drivers.

[0082] S130: Build an ST-GCN model based on dynamic topology and domain prior fusion layer and use the evaluated and annotated video data to train the model;

[0083] In this step S130, it includes the following two main steps:

[0084] S131: OpenPose and skeleton data processing:

[0085] First, use OpenCV to convert the captured video clips into images frame by frame and input them into the OpenPose architecture (using the VGG-19 algorithm to extract features). Through multiple stages of convolution and part affinity field (PAF) estimation, OpenPose outputs the human joint positions, their confidence maps, and joint connection information in each frame. In the present invention, the skeleton data output by OpenPose is not directly used for ST-GCN processing. Instead, a "joint credibility weighted module" (SCWM) is introduced on this basis. This module dynamically adjusts the joint weights according to the OpenPose confidence of each joint point and the previously defined action domain prior. For the joint points required for key operations (such as the arm and head joints when checking the car door), if the confidence is high, their feature weights are increased; the influence on non-critical joint points or low-confidence points is weakened, and the weighted skeleton sequence is obtained as the input for the downstream ST-GCN.

[0086] Specifically, the OpenPose of the present invention uses an architecture based on the VGG-19 algorithm, and its working process is as Figure 3 shown, step S131 includes the following sub-steps:

[0087] S1311: Perform human pose estimation and skeleton data extraction on the evaluated and annotated video data. Among them, the skeleton data includes the coordinates of key joint points and the corresponding joint point confidences, specifically including:

[0088] S13111: Use OpenCV to convert the evaluated and annotated video into an image sequence frame by frame and perform preprocessing;

[0089] This step S13111 is used to implement video frame extraction and preprocessing. Specifically, before inputting into the OpenPose architecture, use OpenCV to convert the captured original video data (AVI format, 30fps, resolution 532×300) into an image sequence frame by frame. To ensure the stability of pose estimation, perform slight preprocessing on the input images, including size scaling, brightness normalization, and noise suppression, etc.

[0090] S13112: Input each preprocessed image into a pre-trained VGG-19-based convolutional neural network to extract the high-dimensional convolutional feature map of the image;

[0091] This step S13112 is used to implement OpenPose model input and feature extraction: Use the OpenPose architecture to perform human pose estimation on each frame of image. The front end of OpenPose uses a VGG-19-based convolutional neural network to extract features from the input image and outputs the feature map F.

[0092] Let the input image be , after passing through the VGG-19 convolutional and pooling layers, a feature map is obtained , where H and W are the height and width of the feature map, and C is the number of channels.

[0093] S13113: Input the high-dimensional convolutional feature map into the multi-stage network of OpenPose to generate the skeletal data of the human body in each frame of the image. The skeletal data includes: the coordinate information of each key joint point of the human body and the confidence score of the detection result of each key node. These key node coordinate information includes, but is not limited to, the two-dimensional position coordinates of the shoulders, elbows, wrists, knees, and ankles.

[0094] This step S13113 is used to implement multi-stage inference and part affinity fields (PAF) estimation:

[0095] Among them, the multi-stage network of OpenPose adopts a multi-stage, two-branch structure, where:

[0096] Branch 1 (confidence map): Estimate the probability distribution of the existence of human key joints and generate a confidence map for each joint point , where k = 1,..., K, and K is the number of joint points.

[0097] Branch 2 (PAFs): Calculate the affinity field between joints to determine which joints belong to the same person and reconstruct the skeletal topology. The output of PAFs is , where M is the number of bone connections, and there are 2 components corresponding to each connection to represent the vector field.

[0098] Through multiple iterative stages, OpenPose continuously optimizes the confidence map and PAFs, and finally outputs the key joint coordinates and corresponding confidence scores in each frame:

[0099] Among them, is the coordinate position of joint k, is the confidence score of this joint.

[0100] S1312: Dynamically adjust the weights of each joint point using the Skeletal Confidence Weighting Module (SCWM) according to the joint point confidence and domain prior knowledge to obtain a weighted skeletal sequence, including:

[0101] After obtaining the joint coordinates and confidence levels output by OpenPose, this embodiment further introduces a joint confidence weighted module to perform secondary processing on these raw skeleton data. The design idea of SCWM is to combine domain knowledge with the joint confidence levels given by OpenPose, so as to adaptively weight the influence of each joint.

[0102] The process of dynamically adjusting the weights of each joint point using the joint confidence weighted module (SCWM) includes:

[0103] S13121: Obtain the key joint coordinates and joint confidence levels output by OpenPose;

[0104] In this step S13121, it is used to obtain the key joint coordinates and joint confidence levels output by OpenPose in step S1311.

[0105] S13122: Calculate the joint confidence weights according to the pre-calibrated key joints and domain prior importance;

[0106] Specifically, this step S13122 includes:

[0107] (1) Domain prior and joint point importance mapping: When determining the standard operations (the five key operations defined in step S110), the expert group has qualitatively labeled the correlation between each operation and the joint points. For example, for the "door closing inspection" operation, the weights of the upper arm and head joint points are greater; for the "environmental safety inspection" operation, the neck and torso joint points are more critical. Define the domain importance vector , where represents the importance degree of joint k in a specific standard operation scenario. This vector is obtained based on expert suggestions and operation process analysis.

[0108] (2) Joint confidence weight calculation: For each joint k, combine the OpenPose confidence with the domain prior importance to calculate the joint confidence weight, that is, the weighting coefficient . A feasible weighting formula is the linear combination plus regularization method:

[0109] where, and are adjustable parameters used to balance the influence of confidence and domain prior on the weights. The denominator is the normalization term to ensure that the sum of all weights is 1.

[0110] S13123: Construct a weighted skeleton feature based on the key joint coordinates and the joint confidence weights, and dynamically adjust the weights of the joint confidence of the weighted skeleton feature; among them, higher weights are assigned to the joint points that play a key role and have high confidence in the current standard operation, and lower weights are assigned to the joint points with low confidence or irrelevance.

[0111] Specifically, this step S13123 includes:

[0112] (1) Construction of weighted skeleton data: Perform weighted processing on the original joint coordinates to obtain a weighted skeleton representation. The weighting is directly used for the joint point coordinates to construct a weighted skeleton feature , where:

[0113] Considering that the joint point features will be used as the node features of the graph input in the subsequent ST-GCN, then is the initial feature of joint point k. In the feature representation, the joint point coordinate information with higher weights has more influence in the subsequent convolution and graph structure processing.

[0114] (2) Handling of low-confidence anomalies: If the of some joint points is too low (such as less than a certain threshold ), then special processing can be performed on it:

[0115] If this joint is not critical in this operation ( is low), then the influence of this point can be considered to be directly weakened, such as being forced to decrease. If this joint is very critical in the operation ( is high) but has low confidence, then compensation can be attempted through multi-frame temporal smoothing or multi-view fusion strategies in subsequent steps.

[0116] S13124: Output a sequence of weighted skeleton features.

[0117] After being processed by SCWM, each frame corresponds to a sequence of weighted skeleton features . These features have comprehensively considered the confidence of joint detection and domain prior knowledge, laying a more stable and targeted input foundation for the subsequent ST-GCN model to perform action analysis in the spatio-temporal dimension.

[0118] Through the processing process of this step S131, the joint point information originally solely relying on the output of OpenPose is weighted and integrated based on domain knowledge and confidence information. In this way, in the subsequent spatio-temporal graph convolution network processing, the model can focus more on the key joint points that truly affect the recognition of standard operations, thereby improving the accuracy and stability of action recognition.

[0119] S132: ST-GCN Model Inference and Training:

[0120] The weighted skeletal feature sequence obtained after joint confidence weighted (SCWM) processing is input into an improved ST-GCN model for training and inference, as Figure 4 shown. This improvement introduces Domain Adaptive Skeletal Normalization (DASN), Dynamic Topology (DT) construction mechanism, and Domain Prior Fusion Layer (DPFL), thus obtaining a spatio-temporal graph convolutional network (ST-GCN) model based on dynamic topology and domain prior fusion layer, so as to improve the accuracy and robustness of recognizing the standard operations of rail transit drivers.

[0121] After inputting the skeleton sequence corresponding to consecutive image frames, first, through the Domain Adaptive Skeletal Normalization (DASN) layer, the skeleton is scaled according to the individual characteristics of the driver, making the skeletal feature distributions of different drivers more comparable. Then, during the feature extraction process of ST-GCN, a Dynamic Topology (DT) construction mechanism and a Domain Prior Fusion Layer (DPFL) are introduced. DT adaptively adjusts the joint connection relationship according to the candidate operation types of the current time series segment, thus highlighting the joint group most relevant to the current operation; DPFL fuses the normalized skeletal features with the predefined standard operation prior vector, guiding the features to be close to the standard operation mode in the rail transit field. In the 9 convolutional blocks of ST-GCN, a residual learning structure is retained after each block, and DT and DPFL layers are inserted between some blocks to dynamically optimize feature extraction. Finally, the probability that the input activity belongs to a specific standard operation is calculated through the SoftMax function, and the loss is calculated with the label ground truth to complete the model training. Among them,

[0122] (1) Domain Adaptive Skeletal Normalization (DASN):

[0123] Input Data and Normalization Target: The weighted skeletal features obtained from step S13124 (K is the number of joints), forming an input sequence for the continuous frame sequence . The height and arm length ratios of each driver vary greatly. Directly inputting into ST-GCN may cause the model to be affected by individual differences. In DASN, according to the pre-measured driver body type parameters (such as height H driver and arm span L driver ), scaling is performed. The normalization mapping can be defined as:

[0124] Among them, is the standardized skeletal sequence obtained after normalization processing, and It can be obtained based on the average joint position and scale statistics of the driver, and is used to normalize the skeleton coordinate distribution to a spatial range with a similar scale at the individual level. Through DASN, the bone feature distributions of different drivers are made more comparable, laying a unified benchmark for subsequent spatio-temporal feature extraction.

[0125] (2)Dynamic Topology (DT) Construction Mechanism

[0126] Conventional ST-GCN uses a fixed skeleton topology, that is, a fixed adjacency matrix to represent the connectivity relationship between joint points. However, in different standard operations, the proportion of the roles of key joint points and unimportant joint points changes, and it is difficult for a fixed topology to pay targeted attention to specific operations.

[0127] DT Adaptive Adjacency Matrix Generation: In the embodiments of the present invention, a graph attention module is inserted before each spatio-temporal convolution block of ST-GCN to adaptively generate a dynamic adjacency matrix according to the current temporal segment features , and the method based on dot-product attention can be adopted:

[0128]

[0129] where is the feature representation of the i-th joint point at time t, is the feature representation of the j-th joint point at time t (obtained from the output of the previous layer or processed by DASN), and are trainable projection matrices. By calculating and normalizing the correlation scores between different joint points, the dynamic adjacency matrix at time t is obtained . For the joint point pairs (i, j) with a high degree of association with the current candidate operation type, will be larger, thus strengthening the influence of this edge connection; while for the connections between irrelevant or unimportant joint points, is assigned a lower value to weaken its influence.

[0130] This dynamic topology (DT) mechanism can adaptively calculate the weight of the adjacency relationship between joint points according to the current temporal segment and candidate operation type; it gives higher weights to the joint point connections with a high degree of relevance in key operations and weakens the edge connections of joint points irrelevant to this operation. Through the dynamic topology (DT), the connection relationship between joint points is adaptively adjusted, strengthening the connectivity between key joint points according to the normative requirements of the current operation, and the generated dynamic adjacency matrix is used to guide the subsequent training of bone features.

[0131] Domain Prior Fusion Layer (DPFL)

[0132] Domain prior feature vector: In the previous expert annotation and statistical analysis, a corresponding prior vector has been defined for each standard operation , where d is the feature dimension. It reflects the typical skeletal motion features, key joint point layouts, and temporal patterns of this operation.

[0133] In addition, in order to align in terms of dimension, the weighted skeleton coordinates obtained in the previous step need to be dimensionality-increased to the matched feature space. First, map X to skeletal features through a linear layer , and then perform fusion:

[0134] Feature fusion strategy: Insert a DPFL layer in the middle of ST-GCN to fuse the currently extracted skeletal features (d is the number of hidden layer channels) with the operation prior . A linear mapping and weighted fusion method can be adopted: ; where, is the dilation vector, and is broadcast to K joint points; is a hyperparameter used to control the prior information fusion ratio. If is smaller, the prior has a greater impact on the current feature, making the feature closer to the standard operation mode in the rail transit field.

[0135] Therefore, in the domain prior fusion layer (DPFL), the prior vector corresponding to each standard operation is predefined based on expert knowledge; the spatio-temporal features (normalized skeletal features) extracted from the middle layer of ST-GCN are fused with the prior vector, making the features closer to the standard action mode in the rail transit field; reducing the generalization uncertainty of the pure data-driven method and improving the recognition accuracy and interpretability of specific standard operations.

[0136] (4) Residual learning and 9-convolution block architecture

[0137] Basic structure of ST-GCN: The model contains 9 spatio-temporal convolution blocks, and each block includes spatio-temporal graph convolution and a non-linear activation function. Through dynamic topology and fused features input, the spatio-temporal graph convolution operation can be expressed as: ; where p traverses the spatial subset (such as the subset division of the joint point itself, neighboring nodes, and distant neighboring nodes), are trainable parameters, is the activation function (such as ReLU). represents the block adjacency matrix assigned to different subset topologies.

[0138] Residual learning: Add residual connections between every two consecutive blocks to ensure the stability of deep network training and the smoothness of gradient transmission: Through residual learning, the degradation problem caused by the increase in depth can be alleviated.

[0139] (5) Loss calculation and model training

[0140] Finally, during the training process, SoftMax classification is performed for each time period, and the probability distribution of the time period belonging to a specific standard operation is output . Compare it with the true label and use cross-entropy loss to update the parameters: ; where t represents the time step and c represents the class index. Optimize the model parameters through backpropagation and gradient descent, including the normalization parameters in DASN, the graph attention weights in DT, the prior fusion ratio in DPFL, and the convolution kernel parameters of each layer in ST-GCN, so that the entire set of models can dynamically adapt to different drivers and different behavior scenarios, and improve the robustness and accuracy of standard operation recognition.

[0141] The above steps and formulas describe the innovation points of the improved ST-GCN in aspects such as dynamic topology construction, domain prior fusion, and domain adaptive normalization. This multi-level improvement enables the model to better capture the semantic and spatio-temporal features of the rail transit driver's operation process, thus providing more reliable and intelligent operation standard monitoring in practical applications.

[0142] S140: Use the trained ST-GCN model based on dynamic topology and domain prior fusion to monitor the driver's operation in real time and output the prediction result.

[0143] Specifically, this step S140 includes:

[0144] The trained ST-GCN model outputs the operation probability corresponding to the driver's operation in real time; convert the operation probability into a standard operation score; compare the standard operation score with a preset safety threshold, and issue a warning prompt when the score is lower than the threshold. At the same time, record the score and formulate a personalized incentive strategy based on the historical trend.

[0145] Specifically, during the driver's operation, the trained improved ST-GCN model is used to monitor the train driver in real time. The model outputs the probability that the operation in a certain time period belongs to a specific specification category. This probability is converted into a driver's specification operation score and recorded, and the score is compared with a preset safety threshold. If the operation score is lower than the set safety threshold, the system immediately issues a real-time reminder, which is fed back to the driver through voice or display, indicating that there are non-standard or potential safety hazards in the operation being performed, and corresponding improvement suggestions are provided, such as "Please check whether the door is fully closed" or "Please confirm that the braking system is normal". This real-time feedback mechanism can help the driver adjust the operation in time and reduce safety hazards. In addition, the present invention uses the driver's score record to formulate a personalized incentive strategy to guide the driver to consciously improve the operation behavior, thereby improving the overall safety management level in the long-term operation.

[0146] Specifically, the specification operation scores recorded in the real-time monitoring are analyzed and a historical trend model is built, and a personalized incentive strategy is formulated according to the driver's long-term performance. For example, positive incentives are provided to drivers with continuously improving scores, and targeted training or safety reminders are given to drivers with frequently low scores. Through such long-term incentive and management strategies, the subjective initiative of the driver for standard operation is improved, thereby reducing the probability of accidents as a whole and improving the operation safety level of rail transit.

[0147] In summary, a method for monitoring the operation standardization of rail transit drivers based on an improved ST-GCN model according to an embodiment of the present invention determines the standard operation as the experimental subject through the above steps 110-140, determines the key standard operations that the driver should take during the train start and stop process, and builds an experimental environment to complete the data collection for training the model; uses the OpenPose architecture (based on the VGG-19 algorithm) to convert the collected video frame by frame into images, and then extracts the human joint positions, joint connection relationships and confidence levels in each frame, and combines the joint credibility weighting module (SCWM) to dynamically adjust the weights of each joint point to generate a weighted skeleton sequence. Subsequently, these processed skeleton data are input into an improved model based on the spatio-temporal graph convolutional network (ST-GCN) to achieve the classification and recognition of the driver's operation, as Figure 2 shown; uses the trained model to monitor the driver's operation in real time and output the probability of operation specification, and gives a timely reminder when it is lower than the threshold; formulates a safety incentive policy using the monitored driver operation specification score to improve the driver's subjective initiative and reduce the risk of accidents.

[0148] The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through the device embodiment. Figure 5It shows a unit module diagram of a monitoring device 200 for the operation standardization of rail transit drivers based on an improved ST-GCN model according to an embodiment of the present invention.

[0149] As Figure 5 shown, a monitoring device 200 for the operation standardization of rail transit drivers based on an improved ST-GCN model provided by the present invention includes:

[0150] An operation definition unit 201, configured to determine various key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrate the key joint points corresponding to each key standard operation;

[0151] A video acquisition unit 202, configured to acquire video data of the driver performing the various key standard operations in the vehicle interior and exterior environments, and evaluate and annotate the compliance of the driver's operation behavior in the video data;

[0152] A model construction unit 203, constructing an ST-GCN model based on a dynamic topology mechanism and a domain prior fusion layer, and training the model using the evaluated and annotated video data;

[0153] A result output unit 204, configured to use the trained ST-GCN model based on dynamic semantic topology and domain prior fusion to monitor the driver's operation in real time and output a prediction result.

[0154] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0155] The methods and the above reaction device of this embodiment will not be elaborated herein for other identical contents.

[0156] The above has described in detail an embodiment of the present invention, but the described content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention shall still fall within the scope covered by the claims of the present invention.

Claims

1. A monitoring method for the operation standardization of rail transit drivers based on an improved ST-GCN model, characterized in that, Including: Determine various key standard operations involved in the start and stop processes of urban rail transit train drivers, and pre-calibrate the key key points corresponding to each key standard operation; Collect video data of the driver performing the various key standard operations in the vehicle and external environments, and evaluate and annotate the compliance of the driver's operation behaviors in the video data; Construct an ST-GCN model based on the dynamic topology and domain prior fusion layer, and use the evaluated and annotated video data to train the model, including: Perform human pose estimation and skeleton data extraction on the evaluated and annotated video data, where the skeleton data includes the coordinates of key key points and the corresponding joint confidence; Dynamically adjust the weights of each joint point using the joint credibility weighting module according to the joint confidence and domain prior knowledge to obtain a weighted skeleton sequence; Input the weighted skeleton sequence into the ST-GCN model for training, including: Perform domain self-adaptive skeleton normalization processing on the weighted skeleton sequence, adjust the skeleton ratio based on the driver's body size parameters to obtain a standardized skeleton sequence; Input the normalized standardized skeleton sequence into the ST-GCN model for feature extraction. During the ST-GCN feature extraction process, through a learnable dynamic topology mechanism, adaptively adjust the joint connection relationship according to the current time series segment and the candidate operation type, embed a domain prior fusion layer in the middle layer, and fuse the normalized standardized skeleton sequence with the standard operation prior vector to guide the features to be close to the standard operation mode in the rail transit field; among the 9 convolutional blocks of ST-GCN, a residual learning structure is retained after each convolutional block, and a dynamic topology and domain prior fusion layer is inserted between some blocks to dynamically optimize feature extraction; Finally, calculate the probability that the input activity belongs to a specific standard operation through the SoftMax function, calculate the loss with the label ground truth, and complete the model training; Among them, the performing human pose estimation and skeleton data extraction on the evaluated and annotated video data includes: Use OpenCV to convert the evaluated and annotated video frame by frame into an image sequence and perform preprocessing; Input each preprocessed frame image into a pre-trained convolutional neural network based on VGG-19 to extract the high-dimensional convolutional feature map of the image; Input the high-dimensional convolutional feature map into the multi-stage network of OpenPose to generate the skeleton data of the human body in each frame image. The skeleton data includes: the coordinate information of each key joint point of the human body and the confidence score of the detection result of each key node. These key node coordinate information includes but is not limited to the two-dimensional position coordinates of the shoulders, elbows, wrists, knees, and ankles; The multi-stage network of OpenPose adopts a multi-stage, two-branch structure. Among them, Branch 1 is used to estimate the probability distribution of the existence of human key joints and generate a confidence map for each key joint point; Branch 2 is used to calculate the part affinity field between joints to determine which joints belong to the same person and reconstruct the skeleton topology; through multiple iterative stages, OpenPose continuously optimizes the confidence map and the part affinity field, and finally outputs the coordinates of the key joint points and the corresponding joint point confidence scores of each frame of image. Use the trained ST-GCN model based on the fusion of dynamic topology and domain prior to monitor the driver's operation in real time and output the prediction result.

2. The method according to claim 1, wherein The multiple key standard operations are five kinds, including: whether the car door is fully closed, whether there are items or people stuck in the car, safety inspection of the internal and external environment of the car, inspection of the train braking system, and standardized confirmation of the train start operation.

3. The method according to claim 2, characterized in that The dynamic adjustment of the weights of each joint point by using the joint credibility weighting module according to the joint point confidence and domain prior knowledge to obtain the weighted skeleton sequence includes: Obtain the coordinates of the key joint points and the joint point confidence output by OpenPose. Calculate the joint credibility weight according to the pre-calibrated key joint points and the importance of the domain prior. According to the coordinates of the key joint points and the joint credibility weight, construct a weighted skeleton feature and dynamically adjust the weight of the joint credibility of the weighted skeleton feature; among them, higher weights are assigned to the joint points that play a key role and have high confidence in the current standard operation, and lower weights are assigned to the joint points with low confidence or irrelevance. Output the weighted skeleton feature sequence.

4. The method according to claim 3, characterized in that The dynamic topology mechanism specifically includes: Introduce a learnable graph attention layer before or in the spatio-temporal convolution block for feature extraction in the spatio-temporal graph convolution network. According to the current time series segment and the candidate operation type, adaptively calculate the adjacency relationship weight between joint points; give higher weights to the connections of joint points with high relevance in key operations and weaken the edges of joint points irrelevant to this operation.

5. The method according to claim 4, wherein The method includes: In the domain prior fusion layer, pre-define the prior vector corresponding to each standard operation based on expert knowledge. Perform a fusion operation on the spatio-temporal features extracted from the middle layer of ST-GCN and the prior vector to make the features closer to the standard action mode in the rail transit field.

6. The method according to claim 5, wherein The use of the trained ST-GCN model based on the fusion of dynamic topology and domain prior to monitor the driver's operation in real time and output the prediction result includes: The trained ST-GCN model outputs the operation probability corresponding to the driver's operation in real time. Convert the operation probability into a standard operation score. Compare the standard operation score with a preset safety threshold. When the score is lower than the threshold, give a warning prompt, and at the same time record the score and formulate a personalized incentive strategy according to the historical trend.

7. A monitoring device for the operation standardization of rail transit drivers based on an improved ST-GCN model, which is used to implement the method for monitoring the operation standardization of rail transit drivers based on the improved ST-GCN model as described in any one of claims 1-6, and is characterized in that, Include: An operation definition unit for determining multiple key standard operations involved in the start and stop process of urban rail transit train drivers and pre-calibrating the key joint points corresponding to each key standard operation. The video acquisition unit is used to acquire video data of the driver performing the multiple key standard operations in the vehicle interior and exterior environments, and evaluate and annotate the compliance of the driver's operation behaviors in the video data; The model construction unit constructs an ST-GCN model based on a dynamic topology mechanism and a domain prior fusion layer, and trains the model using the evaluated and annotated video data; The result output unit is used to use the trained ST-GCN model based on dynamic semantic topology and domain prior fusion to monitor the driver's operations in real time and output prediction results.

Citation Information

Patent Citations

  • Vehicle-mounted pedestrian detection method and system

    CN101714213A

  • Action recognition method, device and equipment and computer readable storage medium

    CN114863325A