Motor train unit driver state identification method and system

By integrating multimodal data fusion and spatiotemporal attention networks, the system can identify violations by EMU drivers in real time and automatically match personalized training courses. This solves the problem of data isolation between real trains and simulation equipment in existing technologies, and achieves efficient training synchronization and improved management efficiency.

CN121767969APending Publication Date: 2026-03-31CHENGDU YUNDA TECH CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The existing EMU driver training system suffers from data isolation between real trains and simulation equipment, resulting in low efficiency in violation identification and a lack of targeted course delivery, leading to a disconnect between training content and actual conditions.

Method used

Employing a multimodal data fusion and improved spatiotemporal attention network, the system collects and analyzes driver video, audio, and operational signal data in real time to identify violations and automatically match personalized reinforcement courses.

Benefits of technology

It enables simultaneous training of real vehicle driving and simulation equipment, improves the accuracy and efficiency of violation identification, reduces human intervention, supports simulated driving of various EMU models, and has good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767969A_ABST
    Figure CN121767969A_ABST
Patent Text Reader

Abstract

The invention discloses a motor train unit driver state identification method and system. The method comprises the following steps: collecting multi-modal data; receiving the synchronized multi-modal data to recalculate the running scene of the vehicle; the illegal behavior of the driver is automatically identified through an improved identification model of the space-time attention network; the identified illegal behaviors are classified; automatically matching a corresponding reinforcement training course according to the type of the identified illegal behavior and associated scene information, and pushing the reinforcement training course to the simulated driving equipment to break a data barrier between a real vehicle and a simulator, so as to realize accurate training based on a real vehicle driving problem; multi-modal data fusion and a space-time attention network are applied, so that the accuracy and efficiency of violation recognition are improved; performing dynamic course matching based on violation classification and historical data; an evaluation report is automatically generated, manual intervention is reduced, and the management efficiency is improved; the system supports simulation driving of various motor vehicle models, and has good expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent rail transit technology, and in particular to a method and system for identifying the status of a high-speed train driver. Background Technology

[0002] Traditional high-speed train driver training mainly relies on a combination of driving simulators and actual trains, but this approach suffers from high costs and risks associated with actual train training, as well as a disconnect between simulation training and real-world operation.

[0003] To solve the above problems, the existing technology mainly adopts the following methods:

[0004] First, there are simulation training systems based on adaptive strategies, such as the standard EMU driving simulation training system and method based on adaptive strategies disclosed in patent CN112102681B. This system can match training courses for trainees according to factors such as train line type and signal type. Second, there are hardware technologies for simulated driving equipment, such as a Fuxing intelligent EMU simulated driving training device disclosed in patent CN215868248U. This device simulates the vibration of the train line through a motion platform, improving the realism of the training. Third, there are EMU linkage simulation training systems, such as the system disclosed in patent CN113823145A, which can realize multi-job linkage training.

[0005] In summary, most existing driving simulators use preset scenarios for training, failing to acquire and utilize real-time data from actual driving, resulting in a disconnect between training content and real-world conditions. This is primarily manifested in the fact that existing EOAS systems rely on simple rules or manual replay analysis of violations, leading to low efficiency and high subjectivity. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for identifying the status of high-speed train drivers. In view of the problems existing in the current high-speed train driver training system, such as data isolation between the real vehicle and the simulation equipment, low efficiency of violation identification, and lack of targeted course delivery, this invention provides a method that can realize the synchronization of real vehicle driving and simulation equipment, intelligently identify violations based on deep learning, and automatically push personalized reinforcement courses.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] In a first aspect, this application discloses a method for recognizing the state of a high-speed train driver, comprising the following steps: S1, during the actual operation of the high-speed train, collecting multimodal data, including video data, audio data, and train operation signal data of the driver's work, and synchronizing the multimodal data in time; S2, receiving the synchronized multimodal data, parsing the driving operation instructions and operating environment information therein, and driving the simulated driving equipment based on the parsing results to reproduce the actual train operation scenario; S3, using an improved spatiotemporal attention network recognition model, performing cross-modal feature fusion and analysis on the video data, audio data, and operation signal data in the synchronized multimodal data to automatically identify the driver's violations; S4, classifying the identified violations; S5, automatically matching corresponding supplementary training courses according to the type of the identified violations and their associated scenario information, and pushing the supplementary training courses to the simulated driving equipment.

[0009] Furthermore, before step S3, a data preprocessing step is included, which includes jitter removal and / or illumination correction processing of video data, noise reduction processing of audio data, and unification of timestamps from various data sources.

[0010] Furthermore, when the improved spatiotemporal attention network performs feature extraction, it uses a 3D CNN network to extract spatiotemporal features of the video, a CNN network based on Mel spectrograms to extract audio features, and an LSTM network to extract operational signal sequence features.

[0011] Furthermore, the cross-modal feature fusion is an attention mechanism fusion, which includes: using video features as query vectors, calculating attention weights for audio features and operational signal features respectively, and performing weighted fusion.

[0012] Furthermore, the types of violations are classified according to preset rules, including key violations involving driving safety, ordinary violations affecting safety, and general operational violations.

[0013] Furthermore, the method also includes the S6 assessment report generation step: regularly collect statistics on the occurrence of various violations by drivers, and automatically generate analysis reports that include the violation level, the time of occurrence, and / or the route segment dimension.

[0014] Its effects include breaking down the data barriers between real vehicles and simulators, enabling precise training based on real vehicle driving problems; applying multimodal data fusion and spatiotemporal attention networks to improve the accuracy and efficiency of violation identification; dynamic course matching based on violation classification and historical data; fully automatic generation of assessment reports, reducing manual intervention and improving management efficiency; and the system supports simulated driving of various train models and has good scalability.

[0015] Secondly, this application discloses a state recognition system for EMU drivers, comprising: a real-vehicle data acquisition and synchronization module, used to acquire multimodal data during the actual operation of the EMU, including video data, audio data, and train operation signal data of the driver's work, and to synchronize the multimodal data in time; a simulated driving equipment synchronization and driving module, used to receive the synchronized multimodal data, parse the driving operation instructions and operating environment information therein, and drive the simulated driving equipment based on the parsing results to reproduce the real-vehicle operation scenario; a multimodal violation recognition module, used to perform cross-modal feature fusion and analysis on the video data, audio data, and operation signal data in the synchronized multimodal data through an improved spatiotemporal attention network recognition model, so as to automatically identify the driver's violation behavior; a data storage and analysis center, used to classify the identified violation behavior; and an adaptive course push module, used to automatically match the corresponding reinforcement training course according to the type of the identified violation behavior and its associated scenario information, and push the reinforcement training course to the simulated driving equipment. Attached Figure Description

[0016] Figure 1 This is a simplified flowchart illustrating a method for identifying the status of a high-speed train driver according to an embodiment of this application.

[0017] Figure 2 This is a simplified schematic diagram of a train driver status recognition system according to an embodiment of this application. Detailed Implementation

[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] refer to Figure 1 and Figure 2 Understand a method and system for identifying the status of a high-speed train driver according to embodiments of this application.

[0020] like Figure 1The method and system for identifying the state of a high-speed train driver according to embodiments of this application include the following steps: S1. During the actual operation of the high-speed train, multimodal data is collected, including video data, audio data, and train operation signal data of the driver's work, and the multimodal data is synchronized in time; S2. The synchronized multimodal data is received, the driving operation instructions and operating environment information therein are parsed, and the simulation driving equipment is driven based on the parsing results to reproduce the actual train operation scenario; S3. Through an improved spatiotemporal attention network recognition model, cross-modal feature fusion and analysis are performed on the video data, audio data, and operation signal data in the synchronized multimodal data to automatically identify the driver's violations; S4. The identified violations are classified; S5. According to the type of the identified violations and its associated scenario information, the corresponding supplementary training courses are automatically matched, and the supplementary training courses are pushed to the simulation driving equipment. Multimodal data refers to various types of data from different sensors. In this method, it specifically includes driver cab video data collected by multi-angle cameras, audio data collected by microphones, and driver operation data such as ATP (Automatic Train Protection System), CIR (Comprehensive Radio Interchange Equipment), and driving operation data collected by the EOAS host (Train Operation Status Monitoring Host). The simulated driving equipment refers to the ground simulation system, including a simulation console, a visual simulation system, and a sound simulation system, which is used to replicate the real vehicle operating environment.

[0021] During actual vehicle operation, the EOAS host collects driver operation data such as ATP, CIR, and driving control information; multi-angle cameras collect video data from the driver's cab; and microphones collect audio data. The collected multimodal data is aligned using high-precision time synchronization algorithms (such as PTP and NTP) to ensure data consistency. Key data is transmitted to the ground server in real time via a 5G network; complete data is uploaded to the locomotive depot server after the driver's shift ends. After receiving the actual vehicle data, the ground server parses driving operation commands (such as driver control gear positions and braking levels) and operating environment information (such as track gradient and signal status). The simulated driving equipment receives this information and recreates the actual vehicle operation scenario: the simulated control console dynamically responds to the actual vehicle signals; the visual simulation system operates in the same track environment as the actual vehicle; and the sound simulation system simulates the corresponding operating sounds. This synchronous driving method achieves real-time mapping between synchronous simulation and actual vehicle operation. An improved spatiotemporal attention network is used to analyze multimodal data. This network first extracts features from video, audio, and operational signals separately, then fuses these features through a cross-modal attention mechanism, and finally uses a temporal attention network to identify sequences of violations. The identified violations are categorized according to maintenance production management regulations, including: major violations (serious violations affecting driving safety, such as using a mobile phone while driving, unauthorized absence from duty, etc.), Class I violations (violations affecting safety but of lesser severity, such as failure to use arm call, speeding, etc.), Class II violations (general operational violations, such as improper timing of operations, non-standard call responses, etc.), and Class III violations (minor violations, such as improper record filling, etc.). Based on the violation type, frequency of occurrence, and associated scenarios, the system automatically matches supplementary courses using a course matching algorithm. The matching algorithm comprehensively considers factors such as the severity of the violation, historical training effectiveness, and course relevance. Personalized course tasks are pushed to drivers through the human-machine interface of the simulated driving equipment. The course content covers the entire process of a standardized crew operation, including departure and arrival procedures, depot operations, departure and en route operation, abnormal driving, and fault handling. After drivers complete the training, the system automatically evaluates the training effectiveness. If the evaluation is passed, the violation is marked as corrected; if it is not passed, the course difficulty or content is adjusted and the course is pushed again.

[0022] In some embodiments, a data preprocessing step is included before step S3. This preprocessing step includes jitter removal and / or illumination correction of the video data, noise reduction of the audio data, and unification of timestamps from various data sources. Specifically, before the multimodal data stream is input into the improved spatiotemporal attention network, data preprocessing is performed, including jitter removal and illumination correction of the video data; noise reduction and enhancement of the audio data; unification of timestamps from various data sources to ensure temporal consistency; and extraction of keyframes to reduce computational complexity.

[0023] Furthermore, in some embodiments, the method also includes an S6 performance evaluation report generation step: periodically compiling statistics on the occurrence of various violations by drivers and automatically generating analysis reports that include violation level, occurrence time, and / or route segment dimensions. Specifically, the number of various violations by each driver is counted monthly, and an EOAS record performance evaluation deduction report is automatically generated. The report is statistically analyzed according to dimensions such as violation level, occurrence time, and route segment, providing decision support for locomotive depot management personnel.

[0024] The following description will be illustrated with some specific examples.

[0025] Specifically, the multimodal data input in S3 is a real vehicle multimodal data stream, including at least video, audio, and operation signals, which in turn outputs personalized course push instructions in S5 and EOAS assessment reports in S6.

[0026] Before the multimodal data stream is input into the improved spatiotemporal attention network, data preprocessing is performed, including dummy and illumination removal for video data; noise reduction and enhancement for audio data; unification of timestamps from various data sources to ensure temporal consistency; and extraction of keyframes to reduce computational complexity.

[0027] For example, the input is a video sequence V∈R (N×H×W×3) Where N is the number of time steps (i.e., sequence length), H and W are the video height and width, respectively, 3 represents the number of RGB channels, and the audio sequence A∈R (N×D) D is the audio feature dimension, and the operation signal sequence O∈R (N×K) K is the dimension of the operation signal, and R is the set of real numbers.

[0028] Then, multimodal feature extraction is performed, including:

[0029] In this embodiment, video features Fv are extracted using a 3D CNN network to extract the spatiotemporal features of the driver's actions. For example, Fv∈R is extracted using a 3D CNN network. (N×Cv) Cv represents the video feature dimension; for audio features Fa, in this embodiment, Mel spectrograms and CNN networks are used to extract speech features, for example, a VGGish network (a pre-trained convolutional neural network) is used, where Fa∈R (N×Ca) Ca represents the audio feature dimension; the operation signal feature Fo is extracted using LSTM in this embodiment, for example, Fo∈R. (N×Co) Co represents the operational feature dimension.

[0030] Then, cross-modal attention feature fusion is performed, including:

[0031] The query matrix Qv = Fv·Wq, where Wq is the learnable parameter matrix.

[0032] The bond matrix Ka = Fa·Wk, Ko = Fo·Wk, and Wk is the learnable parameter matrix.

[0033] The value matrix Va = Fa·Wv, Vo = Fo·Wv, where Wv is the learnable parameter matrix.

[0034] Attention weight calculation:

[0035] α is the video-audio attention weight, and d is the feature dimension, used to scale the attention weight calculation;

[0036] β is the attention weight of the video-operation signal.

[0037] Fusion feature calculation:

[0038] F fused = αVa·Va + βVo·Vo + Fv;

[0039] Where Wq, Wk, and Wv are learnable parameter matrices, and d is the feature dimension.

[0040] In S3, the improved spatiotemporal attention network structure for identifying violation sequences using a temporal attention network includes:

[0041] Spatial attention module, input is fused feature F fused ∈R (N×C) Query vector Qs = F fused • Wsq, the query vector is a spatial query vector, and Wsq is a learnable parameter; key vector Ks = F fused ·Wsk, where the key vector is the spatial key vector, and Wsk is a learnable parameter; spatial attention weights As = softmax( ) .

[0042] Spatial augmentation feature: Fs = As·F fused .

[0043] The temporal attention module takes spatial augmentation features Fs∈R as input. (N×C) Temporal attention weights, At = softmax(LSTM(Fs)); Temporal augmentation features: Ft = At·Fs.

[0044] Then, the classification output is performed, and the probability of violation is P = softmax(MLP(Ft)); the loss function is: L = λ1·Lcls + λ2·Lregular; where Lcls is the focus loss, which solves the class imbalance problem; Lregular is the temporal consistency regularization term; λ1 and λ2 are weight parameters; and MLP represents multilayer perceptron.

[0045] A status recognition system for a high-speed train driver according to an embodiment of this application includes:

[0046] The real-vehicle data acquisition and synchronization module is used to collect multimodal data during the actual operation of the EMU, including video and audio data of driver operations and train operation signal data, and to synchronize the time of the multimodal data. Specifically, during the actual operation, the module collects driver operation data such as ATP (Automatic Train Protection System), CIR (Comprehensive Radio Interception), and driving operations through the EOAS host, video data from the driver's cab through multi-angle cameras, and audio data through microphones. The collected multimodal data is aligned using a high-precision time synchronization algorithm to ensure data consistency. Key data is transmitted to the ground server in real time via a 5G network; complete data is uploaded to the locomotive depot server after the driver leaves duty.

[0047] The synchronous drive module of the driving simulator receives the synchronized multimodal data, parses the driving operation commands and operating environment information, and drives the driving simulator based on the parsing results to reproduce the actual vehicle operation scenario. Specifically, after receiving the actual vehicle data from the ground server, the synchronous drive module parses the driving operation commands (such as the driver's control gear position and braking level) and operating environment information (such as the track gradient and signal status). After receiving this information, the driving simulator reproduces the actual vehicle operation scenario: the simulated control panel dynamically responds to the actual vehicle signals; the visual simulation system operates in a track environment consistent with the actual vehicle; and the sound simulation system simulates the corresponding operating sounds. Through this synchronous drive method, real-time mapping between synchronous simulation and actual vehicle operation is achieved.

[0048] The multimodal violation recognition module is configured to analyze multimodal data based on an improved spatiotemporal attention network. This network first extracts features from video, audio, and operational signals respectively, then fuses the features through a cross-modal attention mechanism, and finally uses a temporal attention network to identify violation sequences.

[0049] The data storage and analysis center categorizes identified violations according to aircraft maintenance production management regulations, including:

[0050] Key violations: Serious acts involving driving safety, such as using a mobile phone while driving or leaving one's post without authorization. Category I violations: Violations that affect safety but are of a lesser degree, such as failure to use the arm call function or speeding. Category II violations: General operational violations, such as improper timing of operations or non-standard call responses. Category III violations: Minor violations, such as improper record-keeping.

[0051] The system automatically matches supplementary courses based on the type of violation, frequency of occurrence, and associated scenarios, using a course matching algorithm. The matching algorithm comprehensively considers factors such as the severity of the violation, historical training effectiveness, and course relevance.

[0052] The adaptive course delivery module pushes personalized course tasks to drivers through the human-machine interface of the driving simulator. The course content covers the entire process of a standardized crew operation, including departure and arrival procedures, segment operations, departure and en route operation, abnormal driving, and troubleshooting. After the driver completes the training, the system automatically evaluates the training effectiveness. If the evaluation is passed, the violation is marked as corrected; if it is not passed, the course difficulty or content is adjusted and the course is pushed again.

[0053] The intelligent performance evaluation form generation module allows the system to statistically analyze the number of various violations committed by each driver monthly and automatically generate an EOAS (Emergency Evaluation and Assessment) report with deduction points. The report provides statistical analysis based on violation level, time of occurrence, and route section, offering decision support for locomotive depot management personnel.

[0054] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for recognizing a state of a driver of a trainset, characterized by, The method comprises the following steps: S1, during the operation of the EMU, collecting multi-modal data including video data, audio data and train operation signal data of the driver's operation, and time synchronizing the multi-modal data; S2, receiving the synchronized multi-modal data, analyzing the driving operation instructions and operating environment information therein, and driving the simulation driving device based on the analysis result to reproduce the real train operation scene; S3, through an improved spatio-temporal attention network recognition model, performing cross-modal feature fusion and analysis on the video data, audio data and operation signal data in the synchronized multi-modal data to automatically identify the driver's violation behavior; S4, classifying the identified violation behavior; S5, according to the type of the identified violation behavior and its associated scene information, automatically matching the corresponding reinforcement training course, and pushing the reinforcement training course to the simulation driving device.

2. The EMU driver state recognition method according to claim 1, characterized in that, before step S3, a data preprocessing step is further included, which comprises de-jittering and / or illumination correction processing of the video data, noise reduction processing of the audio data, and unifying the time stamps of each data source. When the improved spatio-temporal attention network extracts features, a 3D CNN network is used to extract the spatio-temporal features of the video, a CNN network based on Mel spectrum graph is used to extract the audio features, and an LSTM network is used to extract the operation signal sequence features.

3. The method of claim 1, wherein The cross-modal feature fusion is an attention mechanism fusion, which comprises:

4. The method of claim 3, wherein Taking the video features as the query vector, calculating the attention weights of the audio features and the operation signal features respectively, and performing weighted fusion. The types of the violation behaviors are classified according to preset rules, including key violations related to train safety, ordinary violations affecting safety, and general operation violations.

5. The method of claim 1, wherein The method further comprises a S6 evaluation report generation step:

6. The method of claim 1, wherein Periodically statistics the occurrence of each type of violation behavior of the driver, and automatically generates an analysis report containing the violation level, occurrence time and / or line section dimension. It comprises:

7. A system for recognizing the state of a driver of a trainset, characterized in that An actual train data acquisition and synchronization module for collecting multi-modal data including video data, audio data and train operation signal data of the driver's operation during the operation of the EMU, and time synchronizing the multi-modal data; A simulation driving device synchronization driving module for receiving the synchronized multi-modal data, analyzing the driving operation instructions and operating environment information therein, and driving the simulation driving device based on the analysis result to reproduce the real train operation scene; A multi-modal violation behavior recognition module for performing cross-modal feature fusion and analysis on the video data, audio data and operation signal data in the synchronized multi-modal data through an improved spatio-temporal attention network recognition model to automatically identify the driver's violation behavior; A data storage and analysis center for classifying the identified violation behavior; An adaptive course pushing module for automatically matching the corresponding reinforcement training course according to the type of the identified violation behavior and its associated scene information, and pushing the reinforcement training course to the simulation driving device. ​

Citation Information

Patent Citations

  • A Simulation Training System and Method for Standard EMU Driving Based on Adaptive Strategy

    CN112102681B

  • Motor train unit linkage simulation training system and method

    CN113823145A

  • Real vehicle and virtual environment highly-fused man-machine co-driving online evaluation system and method

    CN116738824A

  • Taking-over training method based on remote driving simulator and related device

    CN117765793A

  • Intelligent evaluation system for virtual control of motor train unit

    CN118965106A