A real-time abnormality sensing method for an automatic protection system of a high-speed train
By constructing an anomaly perception model for the ATP system and using LSTM and multi-head attention mechanisms to process the log information of the high-speed train ATP system, the problems of temporal correlation and concurrency in fault diagnosis were solved, and rapid and accurate fault location and handling were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SIGNAL & COMM RES INST OF CHINA ACAD OF RAILWAY SCI
- Filing Date
- 2023-01-12
- Publication Date
- 2026-05-05
AI Technical Summary
Existing fault diagnosis methods for Automatic Train Protection (ATP) systems suffer from problems such as long fault location time, low accuracy, and low automation. Furthermore, existing methods fail to simultaneously meet the requirements of temporal correlation and concurrency.
An anomaly detection model for the ATP system is constructed, employing an encoding network, a decoding network, and a multi-head attention mechanism layer. Long Short-Term Memory (LSTM) neural network is used to memorize log information, and the multi-head attention mechanism is combined to process logs from multiple devices, enabling rapid concurrent fault localization.
It enables real-time anomaly detection in the ATP system, accurately identifies fault locations, narrows the scope of fault impact, provides detailed anomaly handling solutions, and improves the speed and accuracy of fault location.
Smart Images

Figure CN116127395B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rail transit technology, and in particular to a real-time anomaly detection method for a high-speed train automatic protection system. Background Technology
[0002] The Automatic Train Protection (ATP) system is the core of ensuring the safe and efficient operation of high-speed trains, often referred to as the "nerve center" of the train. Installed at both ends of the high-speed train, this redundant system connects to external equipment such as the train and monitoring systems. It primarily consists of sub-devices including a Vital Computer (VC), a Speed & Distance Processing Unit (SDU), a Balise Transmission Module (BTM), a Track Circuit Reader (TCR), a Train Interface Unit (TIU), a Driver-Machine Interface (DMI), a GSM-Railway (GSM-R), and a Juridical Recorder Unit (JRU). These sub-devices work together to ensure train safety. A malfunction in any sub-device will affect the normal operation of the ATP system.
[0003] By the end of 2022, five types of ATP systems were in large-scale operation across the railway network: 300T, 300S, 300H, 200H, and 200C. After the systems went online, each type of sub-equipment exhibited several failure modes. For example, the 300T ATP system has 107 main components, and BTM-related failures alone can be categorized into five types: runtime BSA (Balise Service Available), startup BSA, invalid BTM port, BTM failing to correctly parse messages, and routine test failure. Therefore, each type of ATP system has tens of thousands of possible failure modes, making accurate fault location relatively difficult once a system failure occurs.
[0004] From the current operational status quo, fault diagnosis and location for various ATP systems almost entirely rely on manual labor, resulting in problems such as time-consuming fault location, low accuracy, and low automation. Furthermore, when ATP systems performing transportation tasks experience faults on the line, they are usually handled temporarily by the driver or onboard mechanic. Due to barriers or blind spots in professional skills, there are instances of unfamiliarity with fault handling procedures, inappropriate handling methods, and inability to handle complex scenarios, leading to significant delays and a wide-ranging impact.
[0005] Depending on the research object and diagnostic method, current fault diagnosis for ATP mainly falls into three categories: methods based on empirical knowledge, methods based on analytical models, and data-driven methods. With the development of new artificial intelligence technologies such as machine learning, and the increasing demands on the adaptability of fault diagnosis algorithms to ATP application scenarios, data-driven methods have become the trend and mainstream. In recent years, scholars have successively proposed methods such as Bayesian networks, Labeled-LDA (Latent Dirichlet Allocation), convolutional neural networks, and extreme gradient boosting (XGBoost) for ATP fault diagnosis and classification. Although experimental results show that these methods have some effectiveness, from the perspective of actual ATP application scenarios, there is a temporal correlation between the preceding and following operational data. Furthermore, operational data is generated simultaneously and independently by several devices, and fault localization requires parallel fusion of information from multiple aspects for comprehensive judgment, exhibiting concurrency. Unfortunately, the above methods do not simultaneously meet the temporal correlation and concurrency requirements of ATP fault diagnosis. Summary of the Invention
[0006] The purpose of this invention is to provide a real-time anomaly detection method for high-speed train automatic protection systems, which can simultaneously meet the time-series correlation, speed and concurrency requirements of ATP fault diagnosis, and accurately realize real-time anomaly detection of high-speed train automatic protection systems.
[0007] The objective of this invention is achieved through the following technical solution:
[0008] A method for real-time anomaly detection in a high-speed train automatic protection system includes:
[0009] An anomaly detection model for the ATP system is constructed, including: an encoding network, a decoding network, an attention mechanism layer, and a classifier; the ATP system is a high-speed train automatic protection system.
[0010] Training Phase: The log data of each sub-device in the ATP system at historical moments are encoded using an encoding network. At each moment, an internal state is used for linear, cyclical information transfer, memorizing information from previous moments. This internal state is then combined with the hidden state to calculate the hidden state for all moments. The hidden state of the last moment is used as the initial hidden state of the decoding network. Subsequently, the hidden state of the decoding network at each moment is determined by an attention function calculated using a multi-head attention mechanism layer, combining the hidden state of the previous moment with the hidden states of all moments in the encoding network. The hidden state of the last moment in the decoding network is input to the classifier, which outputs predicted information. The difference between the predicted and actual information is used to train the ATP system anomaly detection model. Each sub-device corresponds to one attention mechanism layer.
[0011] Perception Phase: Using the trained ATP system anomaly perception model, anomalies are perceived in the log data generated in real time during normal operation of the ATP system.
[0012] As can be seen from the technical solution provided by the present invention, by using the real-time operation logs (log data) of all ATP sub-devices as the analysis object, the ATP system log sequence memory and correlation analysis of log information before and after are realized through an encoding and decoding network. A multi-head attention mechanism is introduced to process the operation logs of each device in a hierarchical manner, solving the requirements of speed and concurrency when multiple devices generate logs simultaneously. Finally, the system aggregates the data to accurately capture the abnormal behavior of any sub-device in real time, and comprehensively and promptly determine the specific fault location at the smallest granularity. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart of a real-time anomaly detection method for a high-speed train automatic protection system provided in an embodiment of the present invention;
[0015] Figure 2 This is a schematic diagram of the ATP system anomaly sensing network provided in an embodiment of the present invention;
[0016] Figure 3 This is a schematic diagram of the ATP system anomaly detection model architecture provided in an embodiment of the present invention;
[0017] Figure 4 This is a schematic diagram of the ATP abnormality handling process provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0019] First, the following explanations are provided for the terms that may be used in this article:
[0020] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0021] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0022] The following is a detailed description of a real-time anomaly detection method for a high-speed train automatic protection system provided by the present invention. Contents not described in detail in the embodiments of the present invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer.
[0023] This invention provides a real-time anomaly detection method for a high-speed train automatic protection system (ATP). Using the real-time operation logs of all ATP sub-devices as the analysis object, it employs a cleverly designed hidden layer structure of a Long Short-Term Memory (LSTM) neural network to characterize the temporal correlation between operation logs. Then, a multi-head attention mechanism is introduced to process the operation logs of each device hierarchically, addressing the concurrency requirements of multiple devices simultaneously generating logs. Finally, the system aggregates the logs to accurately capture the abnormal behavior of any sub-device in real time. At the smallest granularity, it comprehensively and promptly determines the specific fault location. Based on the precise fault point, a finite state machine model for anomaly handling is constructed, providing detailed anomaly handling solutions to guide the driver or onboard mechanic in quickly handling anomalies and minimizing the impact of the fault. Figure 1As shown, the method mainly includes the following steps:
[0024] Step 1: Construct an anomaly perception model for the ATP system, including: an encoding network, a decoding network, an attention mechanism layer, and a classifier.
[0025] Step 2: Model training phase.
[0026] The training process is as follows: Log data from each sub-device in the ATP system at historical moments is encoded using an encoding network. At each moment, an internal state is used for linear, cyclical information transmission, memorizing information from previous moments. This internal state is then combined with the hidden state to calculate the hidden state of the encoding network at all moments. The hidden state of the last moment in the encoding network is used as the initial hidden state of the decoding network. Subsequently, the hidden state of the decoding network at each moment is determined by an attention function calculated using a multi-head attention mechanism layer (one attention mechanism layer for each sub-device), combining the hidden state of the previous moment with the hidden states of the encoding network at all moments. The hidden state of the last moment in the decoding network is input to the classifier, which outputs predicted information. The difference between the predicted information and the true information is used to train the ATP system anomaly perception model. In this part, different types of input log data correspond to different predicted information. For example, if the input log data is a log key sequence, the output is the conditional probability distribution of the logs; if the input log data is a sequence of parameter value vectors, the output is a predicted vector. Details will be explained later.
[0027] In this embodiment of the invention, the encoding network is implemented using a Long Short-Term Memory (LSTM) neural network. At each time step, based on the input log data, candidate states, forget gates, input gates, and output gates are calculated. The internal state is calculated using the forget gate, input gate, and candidate states, and the hidden state is calculated using the internal state and output gate. The hidden states at all time steps are recorded as follows: The input log data consists of log sequences generated by each sub-device in the ATP system at historical moments, where m represents the length of the log sequence, equal to the total number of moments. Let represent the hidden state at time i, where i = 1, ..., m. Specifically: the internal state is calculated using the forget gate, input gate, and candidate states, and then the hidden state is calculated using the internal state and output gate, represented as:
[0028]
[0029]
[0030] Among them, c i f represents the internal state at time i. i This represents the forget gate at time i, where i i This represents the input gate at time i. Let c represent the candidate state at time i, ⊙ represent the vector element-wise product, and c i-1 c represents the internal state at time i-1. When i=1, c i-1 The internal state is initialized and can generally be set to 0.
[0031] In this embodiment of the invention, log information generated by all sub-devices is typically aggregated in chronological order to form a log sequence. One log entry from the log sequence is processed at each time point; therefore, the length *m* of the log sequence equals the total number of time points. Furthermore, the log data includes two categories: log keys and vector parameters. Each category's corresponding log sequence is input into the encoding network to obtain the corresponding hidden state.
[0032] In this embodiment of the invention, the decoding network is also implemented using LSTM. The hidden state of the decoding network at each time step is determined by the attention function calculated by combining the hidden state of the previous time step with the hidden states of the encoding network at all time steps through a multi-head attention mechanism layer. Specifically, a multi-head attention mechanism is used, with one attention mechanism corresponding to each sub-device. For the s-th sub-device, at time 1, the hidden state corresponding to time 1-1 is used. As the query vector, it is passed through the attention mechanism layer corresponding to the s-th sub-device and the hidden states H of the encoding network at all times. enc Calculate the attention function at time l corresponding to the s-th sub-device, and concatenate the attention functions at time l corresponding to all devices as the attention function ca at time l of the decoding network. l Then, using the attention function ca at time l l The hidden state of the decoding network at time l-1 Calculate the hidden state of the decoding network at time l. The hidden state of the decoding network at each moment contains the hidden states of all sub-devices.
[0033] In this section, the principle behind the output hidden state of the decoding network is similar to that of the encoding network. The main difference is that the input to the encoding network is a log sequence, while the input to the decoding network is an attention function ca. l .
[0034] For time l, the attention function corresponding to the s-th sub-device Calculated using the following formula:
[0035]
[0036] Where att(.) represents the attention mechanism layer, H enc Representing the hidden states of the encoding network at all times, H enc Represented as key-value pairs (K enc V enc ), and To utilize the parameter matrix in the attention mechanism corresponding to the s-th sub-device and The calculated key vector matrix and value vector matrix are calculated as follows: Key vector matrix The i-th key vector, Value vector matrix The i-th value vector, This represents the hidden state of the encoding network at time i. This represents the hidden state of the s-th sub-device at time l-1. This represents the attention distribution corresponding to the s-th sub-device.
[0037] The calculation formula is as follows:
[0038]
[0039] Where s(.) represents the attention scoring function, calculated using a scaled dot product model. Key vector matrix The j-th key vector.
[0040] Then, the attention functions corresponding to all sub-devices are concatenated to obtain the attention function ca at time l of the decoding network. l .
[0041] Step 3, Perception Stage.
[0042] The trained ATP system anomaly perception model is used to detect anomalies in the log data generated in real time during normal operation of the ATP system. In other words, the prediction information is output in the same way as in the prediction phase, and the prediction information is used to classify the system as normal or abnormal.
[0043] In this embodiment of the invention, the ATP system anomaly detection model mainly includes a log key anomaly detection module and a parameter value anomaly detection module. The log key anomaly detection module and the parameter value anomaly detection module have the same structure, both including corresponding encoding networks, decoding networks, attention mechanism layers, and classifiers. Historical logs are pre-parsed to obtain two parts of log data: log keys and vector parameters. The log data composed of log keys is input to the log key anomaly detection module, which is trained using a training phase. Similarly, the log data composed of parameter value vectors is input to the parameter value anomaly detection module, which is trained using the aforementioned training phase method. During the detection phase, the logs generated in real-time during normal ATP system operation are parsed. The parsed log data composed of log keys and the log data composed of vector parameters are respectively input to the trained log key anomaly detection module and parameter value anomaly detection module. When either the log key anomaly detection module or the parameter value anomaly detection module outputs an anomaly, the ATP system is considered to be abnormal.
[0044] In this embodiment of the invention, the log key anomaly detection module, during the training phase, will... g Let K ∈ K be the conditional probability distribution of the target log key. A loss function is constructed based on the difference between this distribution and the actual log key at the corresponding time step to train the log key anomaly detection module, thereby updating the module's internal parameters. The goal is to learn a conditional probability distribution that maximizes the training log key sequence; where K = {k1, k2, ..., k...} u} represents a specified set of log keys in the ATP system, where u is the number of log keys; during the sensing phase, for the target log key km g Based on the input log key window w, calculate the conditional probability distribution Pt[km]. g |w]={k1:p1,k2:p2,…,k u :p u}, where w={km g-h ,…,km g-2 ,km g-1} Contains the target log key km g For the first h most recent log keys, any element in w belongs to set K, k1:p1, k2:p2, ..., k u :p u Let p1 represent the probability corresponding to log key k1, p2 represent the probability corresponding to log key k2, and p3 represent the probability corresponding to log key k4. u The corresponding probability is p u Log keys are sorted in descending order according to their probability, and the first q log keys are determined as corresponding candidate log keys. If a log key is within the first q candidate values, it is marked as normal; otherwise, it is marked as abnormal.
[0045] In this embodiment of the invention, during the training phase, the parameter value anomaly perception module takes a vector parameter as input and outputs a real-valued vector as output, which serves as the predicted value for the next input vector parameter. By dynamically adjusting the weights of the ATP system anomaly perception model, the error between the next input vector parameter and the predicted value is reduced. During the perception phase, the error between the input vector parameter and the predicted value is modeled as a Gaussian distribution. If the error is within the confidence interval of the predicted value, the input vector parameter is marked as normal; otherwise, it is abnormal.
[0046] In this embodiment of the invention, when the ATP system is identified as abnormal during the sensing phase, an early warning is issued by expanding the display of the DMI to guide the user to execute the abnormal handling process. Specifically, different abnormal handling processes are abstracted into finite state automata, and the state transitions and execution order of the finite state automata are used to characterize the system abnormal handling process. On this basis, a knowledge base is formed and reused by different operation and maintenance entities across the entire network for handling the same type of abnormality.
[0047] To more clearly demonstrate the technical solution and its effects provided by the present invention, the following describes in detail a real-time anomaly detection method for a high-speed train automatic protection system provided by the present invention, using specific embodiments.
[0048] I. Introduction to the overall principle.
[0049] like Figure 2 The diagram shows the overall framework of the real-time anomaly detection method for the ATP system. First, the log data generated during ATP operation is preprocessed. Based on the idea of common subsequences, the unstructured logs are parsed online into structured log keys and parameter vectors. Then, anomalies in the log keys and parameter vectors are detected using an anomaly detection model. Finally, based on the accurately located anomalies, the corresponding anomaly handling plan is automatically triggered to guide the resolution of the anomalies. The technical implementation methods of each part are explained below.
[0050] 1. ATP operation log analysis.
[0051] The ATP system continuously generates log messages from startup. These text-based log messages are unstructured, which is inconvenient for computer processing. Therefore, it is necessary to automatically parse the unstructured logs into a structured representation. This invention designs a log parser based on the Longest Common Subsequence (LCS) method to parse log entries generated by the system online in real time. The basic idea is to extract a log key and a vector parameter from each log entry. The log key of log entry e refers to the string constant k in the print statement in the source code, and the parameter in the log key is abstracted as the symbol *. These parameter values reflect the performance status of the ATP system and are represented by the vector v. Therefore, all log entries can be parsed into the log key k and the vector parameter v, thus allowing for log classification.
[0052] The core of designing the parser is to extract log keys through comparison. For example, a log entry e1 of a 300T ATP system: Profibus ATP to JRU connected.
[0053] Enter log entry e2 again: Profibus TI-H to TSG connected.
[0054] Then, by iterating through the list of log objects, we found an object whose log key property was: Profibus ATP to JRUconnected.
[0055] Therefore, the log key k is: Profibus<*>to<*>connected. The vector parameter v is: [TI-H,TSG].
[0056] 2. Log sequence memory and correlation analysis of log information before and after.
[0057] Long Short-Term Memory (LSTM) neural networks introduce new internal states and gating mechanisms in the hidden layers to control the information transmission path. An internal state (vector) c is used. i ∈R D It specifically performs linear cyclic information transfer, and then non-linearly outputs information to the external state (hidden state, vector) h of the hidden layer. i ∈R D State c i and h i The calculation formula is:
[0058]
[0059] h i =o i⊙tanh(c i (2)
[0060] Among them, f i This represents the forget gate at time i, where i i O represents the input gate at time i. i This represents the output gate at time i. Let c represent the candidate state at time i obtained through the tanh activation function, ⊙ represent the element-wise product of vectors, and c i-1 Let R represent the internal state at time i-1, R represent the symbol of the real number set, and D be the vector dimension.
[0061] Internal state c i Similar to a conveyor belt, information is transmitted along the belt but its state does not change. i The ability to capture key information at a certain moment and retain this key information for a certain time interval means that c i It can record all ATP operation history information up to the current moment.
[0062] Gating mechanisms use clever gate design to remove or add information to the internal state c. i Three doors i i i and o i The values are all between (0,1), indicating that only a certain percentage of information is allowed to pass.
[0063] The candidate states and the three gates are calculated as follows:
[0064]
[0065] Where, x i Let x be the input at time i. The inputs at all times can be viewed as a log sequence, denoted as x. 1:M =(x1,x2,..,x M M is the length of the sequence; h i-1 This represents the hidden state of the previous time step (i-1); W ∈R D×M The state input weight matrix; b∈R D Let be the bias vector. σ is the Logistic activation function, whose output interval is (0,1), defined as...
[0066] The above calculation process can be described as follows:
[0067] 1) Utilize the hidden state h from the previous time step i-1 and the input x at time i i Calculate f at time i i i i,o i ,
[0068] 2) Combining the forgetting gate f at time i i and input gate i i Update the internal state c at time i i .
[0069] 3) Combine the output gate o at time i i The information of the internal state at time i is passed to the hidden state h at time i. i .
[0070] Through the above steps, the entire network achieves long-distance temporal relationship dependencies, thus realizing the memory of ATP system log sequences and the correlation analysis of log information before and after.
[0071] 3. Attention mechanism.
[0072] The core of the attention mechanism is to monitor a larger set of information at each step, setting different weight parameters for the input, and obtaining the importance of each element during the learning process, extracting the more important and crucial information. The attention focusing process is reflected in the calculation of the weight parameters; the larger the weight, the more focused it is on the corresponding log information, meaning the weight represents the importance of the log. This allows the model to make more accurate and faster judgments.
[0073] The calculation of the attention mechanism consists of two steps: first, the attention distribution is calculated over all inputs X, and then the weighted average of the input information is calculated based on the attention distribution.
[0074] In order to obtain from the input vector X = [x1, x2, ..., x...] N In this context, we select information relevant to a specific task, denoted as the query vector q, where N is the number of input data points. The attention variable z∈[1,N] represents the index position of the selected information, and z=n indicates that the nth vector has been selected. Given input X and query vector q, what is the probability α of selecting the nth vector? n for:
[0075]
[0076] Where, α n Here, the attention distribution is defined, softmax is the activation function, and s(x,q) is the attention scoring function, typically calculated using a scaled dot product model.
[0077]
[0078] Where D is the vector dimension.
[0079] α nThis explains the degree of attention given a query vector q relevant to a specific task, indicating the importance of the nth vector. The attention function is a weighted average of the input information.
[0080]
[0081] If the input information uses more general key-value pairs (K,V)=[(k1,v1),(k2,v2),…,(k N ,v N The key vector matrix K is used to calculate the attention distribution α. n The value vector matrix V is based on α n To calculate aggregated information, the corrected attention function is:
[0082]
[0083] In the formula, n,j∈[1,2,…,N] are the positions of the output and input X vector sequence of the attention function, respectively, and α nj This represents the weight that the nth output focuses on the jth input. During the operation of the ATP system, devices such as VC, SDU, and BTM simultaneously generate operational log data. To process this information in parallel, the concept of multi-head attention is employed. Multiple query vectors Q = [q1, q2, ..., q...] are used. M The attention function selects multiple sets of key information from the input information in parallel, with each attention focusing on a different part of the input X. For example, one attention focuses more on the runtime information generated by VC, while another attention focuses more on the information generated by SDU. In this case, the attention function can be written as follows:
[0084]
[0085] In the formula, ⊕ represents vector concatenation, and the calculation of each sub-part is shown in formula (7).
[0086] In summary, the focusing characteristics of the attention mechanism and the parallel characteristics of multi-head attention solve the problems of speed and concurrency in abnormal perception of the ATP system.
[0087] II. Anomaly Detection Model of the ATP System.
[0088] Based on the above principles, an ATP system anomaly detection model can be constructed that simultaneously meets the requirements of temporal correlation, speed, and concurrency in ATP anomaly detection, taking into account the actual characteristics of ATP system application scenarios. Furthermore, it can narrow the impact range of high-speed rail ATP system faults and accurately locate system anomalies. Specifically:
[0089] The real-time log data generated during the operation of the ATP system can be viewed as a sequence of elements following certain patterns and grammatical rules. It is generated by a logically rigorous and well-structured computer program, closely resembling natural language. Deep learning methods are used to process this "natural language," extracting its informational value to achieve accurate anomaly detection. The challenges in this process are:
[0090] Temporal correlation requirements. When an ATP system executes a transportation task, the operational data is generated sequentially due to the continuity of the operational scenario, and the data from different points in time shows a clear correlation in business operations. Anomaly detection cannot rely on isolated judgments based on one or two log entries at the current moment; it needs to memorize information over a period of time and deeply analyze the correlations between operational information.
[0091] Speed is crucial. If the ATP system malfunctions, it needs to be detected immediately and responded to swiftly, intervening promptly in any ongoing attacks or anomalies. The anomaly detection solution must be both rapid and accurate, providing detailed and clear guidance and operational procedures for fault handling.
[0092] Concurrency requirements. The ATP system's operational information is generated simultaneously and independently by several critical devices. Log messages originate from several different threads or concurrently running tasks, and the final fault diagnosis and localization require comprehensive judgment and system evaluation by integrating all information from related devices. Furthermore, due to the different functions of each device, the operational information generated has a clear hierarchical structure within the business context. Anomaly detection needs to be performed in parallel across multiple devices, and the final results need to be organically integrated.
[0093] By employing a clever hidden layer structure design in a Long Short-Term Memory (LSTM) neural network to memorize operational information over a specific time period, the temporal correlation problem in anomaly detection is addressed. An attention mechanism is introduced, assigning different weight parameters to the input information to address the speed issue in anomaly detection. Furthermore, a multi-head attention mechanism is used to assign different attention functions to each sub-device of the ATP system, resolving the concurrency problem in anomaly detection. Finally, a real-time anomaly detection model for the ATP system based on the LSM neural network and attention mechanism is established. The following sections provide a detailed introduction to each part.
[0094] 1. Implement the framework and process.
[0095] like Figure 3 The diagram shown illustrates the architecture of the ATP system anomaly detection model. The essence of ATP system anomaly detection is to use a neural network to estimate the conditional probability p of the log. θ (x L |x 1:(L-1) ), p θ (x L |x 1:(L-1) ) indicates based on historical data x 1:(L-1)Predicting future data x L The conditional probability, 1:(L-1) represents the time from the first time step to the (L-1)th time step, and L represents the next time step. This is the idea of the deep sequence model. The direct way to realize sequence-to-sequence is to use an LSTM network for encoding and decoding. Figure 3 In the diagram, A represents the LSTM network of the encoder, and A′ represents the LSTM network of the decoder. The specific implementation ideas and methods are as follows.
[0096] (1) The encoding network consists of LSTM, and the input log data is the log sequence x. 1:m =(x1,x2,..,x m ), m is the length of the input log sequence. According to the above equations (1) and (2), the last hidden state h and the internal state c are obtained as the output (h,c) of the encoding network.
[0097] As shown in equation (2), the internal state information c of the coding network is contained in the hidden state h. For the convenience of subsequent descriptions, any output of the coding network is directly denoted as h. remember If we represent all the outputs of the encoding network, then
[0098] Those skilled in the art will understand that a neural network structure typically consists of three layers: an input layer, a hidden layer, and an output layer. LSTM is a type of neural network, where the input layer is primarily responsible for encoding log data (e.g., using one-hot encoding), the hidden layer is primarily responsible for calculating the hidden states, and the output layer is primarily responsible for outputting the hidden states.
[0099] (2) The decoding network is also composed of LSTM, and the hidden state of the network is denoted as... Initial state This means that the initial input is the last output of the encoding network. Because of the added attention mechanism, all states of the encoding network need to be preserved, and are calculated according to equation (4). And each The correlation, i.e., the attention distribution α i Because the encoding network has m output states, it will produce m α values, and one α value... i One Then calculate the weighted average of the m states.
[0100] (3) At time l in the decoding process, use the hidden state at time l-1. As a query vector, from the input sequence H enc Selecting useful information means that at each moment, relevant information is selected from the hidden states of all the encoding networks through an attention mechanism.
[0101] Furthermore, the ATP system comprises multiple sub-devices (e.g., VC, SDU, BTM, etc., S devices), each generating log information. Therefore, the parameter matrices in the attention mechanism differ for different sub-devices. Thus, a multi-head attention mechanism is employed to calculate the attention function for each sub-device separately, and then concatenate them to obtain the attention function ca. l The calculation of the attention function for each sub-device involves key vectors, value vectors (i.e., key-value pairs), and query vectors (i.e., the hidden state of each sub-device at the previous time step). Although the inputs to each attention mechanism are the same, the parameter matrices and the hidden states of different sub-devices at the previous time step are different. Therefore, the attention function for each sub-device can be calculated using the following formula, and the hidden state of each sub-device at each time step is included in the hidden state of the decoding network at the corresponding time step. In this embodiment of the invention, the specific number of multi-head attention mechanisms is related to the number of sub-devices. For example, if there are 4 sub-devices, a 4-head attention mechanism is used. However, this application does not limit the specific number of multi-head attention mechanisms.
[0102] Attention function corresponding to the s-th sub-device Calculate using the following formula:
[0103]
[0104] In the formula, l = 1, 2, ..., t, where t represents the total number of hidden states in the decoding network, i.e., the total number of time points in the decoding network; s = 1, 2, ..., S, where H... enc Represented as key-value pairs (K enc V enc ), and To utilize the parameter matrix in the attention mechanism corresponding to the s-th sub-device and The calculated key vector matrix and value vector matrix are calculated as follows: Key vector matrix The i-th key vector, Value vector matrix The i-th value vector; the query vector is That is, the hidden state of the s-th sub-device at time l-1 of the decoding network. Then, according to equation (7), equation (9) can be calculated to obtain the attention function corresponding to the s-th sub-device.
[0105] After calculating the attention functions corresponding to the S sub-devices using the above method, the attention functions corresponding to the S sub-devices are concatenated using the aforementioned equation (8) to obtain the attention function ca. l For the attention function shown in equation (9), the difference before and after using the attention mechanism is further explained:
[0106] On the one hand, calculate the attention function at any time l, and then for each ca l Each corresponds to a state of a decoding network. The hidden state of the decoding network without using an attention mechanism The previous update only depended on the previous state and did not consider the state of the encoding network. With the attention mechanism, the state update is related not only to the previous state but also to the current state (ca). l Related, meaning related to the state of the encoding network.
[0107] On the other hand, since the attention function includes inputs x1 to x m Complete information, and decoding the new state of the network. The reliance on the attention function indicates that the input information of the encoding network is passed to the decoding network, thus achieving information memorization.
[0108] Furthermore, attention-based sequence encoding can be viewed as a fully connected feedforward neural network, since the cache at each time step... l If they are all different, then it means that while each position in the l-th layer receives the output of the (l-1)-th layer position, the connection weights also change dynamically.
[0109] (4) Ca l As input to the decoding network at time l, the hidden state at time l is obtained. And repeat step (3) above until the hidden state at the last moment is obtained. The hidden state of the decoding network at each moment contains the hidden states of all sub-devices.
[0110] (5) Decode the hidden state at the last moment of the decoding network. The input classifier g(·) outputs predicted information. During the training phase, the predicted information can be used to train the network. During the perception phase, the predicted information is used to classify the data as normal or abnormal. In this part, if the input log data is a sequence of log keys, the output is the conditional probability distribution of the logs; if the input log data is a sequence of parameter value vectors, the output is the predicted vector, which will be explained in detail later.
[0111] 2. Ideas and methods for anomaly detection.
[0112] like Figure 3As shown, the anomaly detection model consists of two parts: a log key anomaly detection module and a parameter value anomaly detection module.
[0113] Model training phase: The training data consists of log files generated by each sub-device of the ATP system. After parsing, each log entry comprises a log key and a parameter value vector, arranged in chronological order. The log data composed of log keys (log key sequence) is input into the log key anomaly detection module, which is trained using the method described in the previous training phase. Similarly, the log data composed of parameter value vectors (parameter value vector sequence) is input into the parameter value anomaly detection module, which is trained using the method described in the previous training phase. Both the log key sequence and the parameter value vector sequence can be collectively referred to as the log sequence mentioned above.
[0114] Model Awareness Phase: The ATP system operates normally, generating a new log entry in real time, which is immediately parsed into a log key and a parameter value vector. The log key anomaly detection module checks if the incoming log key is normal. If so, the parameter value anomaly detection module checks the parameter value vector. If either the log key or the parameter value vector in the log entry is detected as abnormal, it is marked as such. At this point, the model provides a semantic warning on the Driver-Machine Interface (DMI), issuing an immediate alert and providing clear guidance and emergency response plans. If the anomaly flag is ultimately confirmed as a false alarm, the model is updated to adapt to the new mode.
[0115] The methods for detecting log key anomalies and parameter value anomalies are explained below.
[0116] (1) Log key anomaly detection.
[0117] Log data is generated by a specific program of the ATP system, and its type is constant. Let K = {k1,k2,…,k} u} represents a specific set of log keys for the ATP system, and the sequence of log keys reflects the business logic order in which the ATP system operates. Let km i This represents the log key at position i in a given log key sequence. Clearly, km i It is one of the u log keys in K, and largely depends on km. i The most recent log key.
[0118] Since the log type is constant, log key anomaly detection is abstracted as a multi-class classification problem, where different log keys are defined as different classes. The input to the multi-classifier is the historical record of the most recent log keys, assuming it is h recent log key windows w, i.e., w = {km}. g-h ,…,km g-2 ,kmg-1}, where any element in w belongs to set K, and g is the target log key k. a The sequence ID. The input layer encodes u possible log keys from set K into one-hot vectors. The output is the conditional probability distribution of the target log key, Pr[km], which is obtained by using a classification function to transform the hidden state of the decoding network at the last time step into the conditional probability distribution of u log keys. g =k a |w], where k a ∈K(a=1,2,…,u).
[0119] Training phase: The log key anomaly detection module will km g Let ∈K be the conditional probability distribution of the target log keys, used to update the model itself. The goal is to learn a conditional probability distribution that maximizes the training log key sequence. Here, the first q log keys in the conditional probability distribution are considered normal. q is a positive integer, and the q value adjusts the trade-off between the anomaly detection rate and the false alarm rate; q is a variable that needs to be adjusted based on the training results. For example: during initialization, a q value is manually set, for example, 8; various metrics of the log key anomaly detection module (e.g., precision, recall, etc.) are tested; then the q value is modified, for example, to 10; and the various metrics of the log key anomaly detection module are tested again. This process is repeated until a satisfactory q value is obtained.
[0120] Perception Phase: In order to perceive a target log key km g Whether an anomaly is detected is determined by sorting all possible log keys K. The conditional probability distribution Pt[km] is then obtained based on the input log key window w. g |w]={k1:p1,k2:p2,…,k u :p u}, where k1:p1,k2:p2,…,k u :p u Let p1 represent the probability corresponding to log key k1, p2 represent the probability corresponding to log key k2, and p3 represent the probability corresponding to log key k4. u The corresponding probability is p u Sort the log keys in descending order of probability and determine the first q log keys as the corresponding candidate log keys; if km g If there are q candidate log keys, then the target log key is km. g Mark as normal; otherwise, mark as abnormal.
[0121] During the perception phase, if a target log key in the log key sequence is marked as abnormal, it is determined that the corresponding sub-device is malfunctioning.
[0122] (2) Parameter value anomaly detection.
[0123] Log key sequences are very useful for anomaly detection in ATP systems. However, in some scenarios, ATP system anomalies do not manifest as execution sequence anomalies, but rather as irregular parameter values. These parameter value vectors v constitute a parameter value vector sequence, i.e., a parameter matrix. The parameter value vector sequence is used as a separate time series to train the parameter value anomaly detection module. Each column is a univariate time series, and the resulting matrix has multiple columns, which can be viewed as a multivariate time series. Based on the LSTM method, the input is a vector parameter sequence composed of recent historical records. From the perspective of each row, each time instance t... i It is the vector parameter in the timestamp; the output is: a real value vector, which serves as the predicted value of the next parameter value vector.
[0124] Training phase: Dynamically adjust the weights of the LSTM model to minimize the error between the predicted and observed values.
[0125] Perception phase: At each time step, the error between the input vector parameters and the predicted value is modeled as a Gaussian distribution. If the error is within the confidence interval of the predicted value, the input vector parameters are marked as normal; otherwise, they are marked as abnormal.
[0126] Similarly, if a vector parameter in the parameter value vector sequence is marked as abnormal, then the corresponding sub-device is considered to be malfunctioning.
[0127] 3. Abnormal handling plan.
[0128] After the anomaly detection model enables rapid and accurate fault location, the next question is how to handle it. Different abnormal states of the ATP system correspond to different workflows, and the anomaly handling methods vary.
[0129] Different exception handling processes are abstracted into Finite State Automatons (FSAs), and the state transitions and strict execution sequences of these automata characterize the system's exception handling process. Based on this, a knowledge base is formed, which can be reused by different operation and maintenance entities across the entire network for rapid handling of such exceptions. An example is shown below. Figure 4As shown, when the model detects a BTM software configuration anomaly, the handling procedure is as follows: the driver should restart the ATP system, and after the train enters the depot, maintenance personnel should check and test it. When the model detects a BSA anomaly, the handling procedure is as follows: after the train brakes to a stop, if the text disappears, the driver must confirm with the dispatcher that the route ahead is normal before proceeding; otherwise, the ATP system should be restarted. Different electrical engineering sections (or high-speed rail sections) across the entire railway network optimize and improve the ATP anomaly handling procedures within their jurisdictions, and then share them on a unified platform. If other electrical engineering sections encounter similar faults, they can directly reuse them or continue to improve them. This cycle forms a standardized fault handling procedure for the ATP system across the entire railway network, which is then abstracted into automata models and embedded into the anomaly detection model. When an anomaly is detected, the automaton is activated, automatically starting the corresponding handling procedure and solution. This can be achieved by expanding and enriching the prompts of the DMI, guiding the onboard mechanics and driver to quickly handle anomalies.
[0130] The above-mentioned solutions provided by the embodiments of the present invention mainly achieve the following beneficial effects:
[0131] 1) It realizes the memory of ATP operation information and the correlation analysis of time sequence.
[0132] 2) An encoder-decoder architecture incorporating an attention mechanism was designed and built, with both the encoder and decoder consisting of LSTM networks. The model focuses on key positions among numerous input log sequence matrices. The previous time-step hidden state of the decoder serves as the query vector. At each step, an attention function different from the encoder's hidden state is dynamically calculated, resulting in different connection weights that update the decoder's hidden state output. By setting different weights for the inputs, the speed requirement for anomaly detection in the ATP system is addressed.
[0133] 3) By utilizing a multi-head attention mechanism to abstract the scenario where multiple devices simultaneously generate operational information, the hierarchical characteristics of the information are characterized, and multiple sets of key information are selected from the input information in parallel. Each attention focuses on a different part of the input sequence, and then they are combined to form a unified attention function, which solves the concurrency problem of abnormal perception in the ATP system.
[0134] 4) A unified method is proposed to detect ATP system anomalies online, which can simultaneously meet the requirements of time correlation, speed and concurrency needed in the actual application scenarios of ATP.
[0135] 5) Train the model using historical log data to obtain optimal model parameters. Then use the optimal model to process log information online and detect anomalies, solving problems such as difficulty in fault location and improper emergency response by drivers.
[0136] 6) The handling procedures for different abnormal states of the ATP system are abstracted into finite state automata to form a unified knowledge base, which can be reused in other similar scenarios across the entire railway network. Drivers or onboard mechanics can accurately and quickly handle abnormalities, reducing fault delays and the scope of impact.
[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0138] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0139] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for real-time anomaly detection in a high-speed train automatic protection system, characterized in that, include: An anomaly detection model for the ATP system is constructed, including: an encoding network, a decoding network, an attention mechanism layer, and a classifier; the ATP system is a high-speed train automatic protection system. Training Phase: Log data from each sub-device in the ATP system at historical moments is encoded using an encoding network. At each moment, an internal state is used for linear, cyclical information transmission, remembering information from previous moments. The hidden state is then calculated by combining the internal state with the previous state, resulting in the hidden states for all moments. The encoding network is implemented using a Long Short-Term Memory (LSTM) neural network. At each moment, based on the input log data, candidate states, a forget gate, an input gate, and an output gate are calculated. The internal state is calculated using the forget gate, input gate, and candidate state, and then the hidden state is calculated using the internal state and the output gate. The hidden states for all moments are recorded as follows: The input log data is a log sequence formed by log data generated by several sub-devices in the ATP system, where m is the length of the input log sequence, equal to the total number of time points. Let i represent the hidden state at time i, where i = 1, ..., m. The hidden state at the last time step is used as the initial hidden state of the decoding network. Subsequently, the hidden state of the decoding network at each time step is determined by an attention function calculated by combining the hidden state at the previous time step with the hidden states of the encoding network at all time steps through a multi-head attention mechanism layer. The hidden state at the last time step of the decoding network is input into the classifier, which outputs predicted information. The difference between the predicted information and the real information is used to train the anomaly perception model of the ATP system. Each sub-device corresponds to one attention mechanism layer. Perception Phase: Using the trained ATP system anomaly perception model, anomalies are perceived in the log data generated in real time during normal operation of the ATP system.
2. The real-time anomaly detection method for a high-speed train automatic protection system according to claim 1, characterized in that, The internal state is calculated using the forget gate, input gate, and candidate states. The hidden state is then calculated using the internal state and output gate, and represented as follows: ; ; in, This represents the internal state at time i. This represents the forget gate at time i. This represents the input gate at time i. This represents the output gate at time i. Let i represent the candidate state at time i. Represents the element-wise product of vectors. This represents the internal state at time i-1. This represents the hidden state at time i.
3. The real-time anomaly detection method for a high-speed train automatic protection system according to claim 1, characterized in that, The hidden state of the decoding network at each time step is determined by the attention function calculated through a multi-head attention mechanism layer, which combines the hidden state of the previous time step with the hidden states of the encoding network at all time steps. The function includes: A multi-head attention mechanism is used, with one attention mechanism corresponding to each sub-device. For the s-th sub-device, at time l, the hidden state corresponding to time l-1 is utilized. As the query vector, it is obtained by connecting the attention mechanism layer corresponding to the s-th sub-device and the hidden states of the encoding network at all time points. Calculate the attention function at time l corresponding to the s-th sub-device, and concatenate the attention functions at time l corresponding to all sub-devices as the attention function at time l of the decoding network. Then, using the attention function at time l The hidden state of the decoding network at time l-1 Calculate the hidden state of the decoding network at time l. The hidden state of the decoding network at each moment contains the hidden states of all sub-devices.
4. The real-time anomaly detection method for a high-speed train automatic protection system according to claim 3, characterized in that, The attention function is calculated by combining the hidden state from the previous time step with the hidden states of the encoding network at all time steps through a multi-head attention mechanism layer, including: For time l, the attention function corresponding to the s-th sub-device Calculated using the following formula: ; Where att(.) represents the attention mechanism layer, Representing the hidden state at all times, Represented as key-value pairs , and To utilize the parameter matrix in the attention mechanism corresponding to the s-th sub-device and The calculated key vector matrix and value vector matrix are calculated as follows: , , , , Key vector matrix The i-th key vector, Value vector matrix The i-th value vector, This represents the hidden state of the encoding network at time i. This represents the hidden state of the s-th sub-device at time l-1. This represents the attention distribution corresponding to the s-th sub-device. The input log data is a log sequence formed by log data generated by several sub-devices in the ATP system, and m is the length of the input log sequence, which is equal to the total number of time points. The calculation formula is as follows: ; in, The attention scoring function is calculated using a scaled dot product model. Key vector matrix The j-th key vector; By concatenating the attention functions corresponding to all sub-devices, we obtain the attention function of the decoding network at time l. .
5. A real-time anomaly detection method for a high-speed train automatic protection system according to any one of claims 1 to 4, characterized in that, The ATP system anomaly detection model includes: a log key anomaly detection module and a parameter value anomaly detection module; the log key anomaly detection module and the parameter value anomaly detection module have the same structure, each containing a corresponding encoding network, decoding network, attention mechanism layer and classifier; The logs from historical moments are parsed in advance to obtain two parts of log data: log keys and vector parameters. The log data consisting of log keys is input into the log key anomaly detection module, which is trained using a training phase. The log data consisting of parameter value vectors is input into the parameter value anomaly detection module, which is trained using a training phase. During the perception phase, the logs generated in real time during the normal operation of the ATP system are parsed. The log data consisting of log keys and the log data consisting of vector parameters obtained from the parsing are respectively input to the trained log key anomaly perception module and parameter value anomaly perception module. When the output result of the log key anomaly perception module or the parameter value anomaly perception module is abnormal, the ATP system is identified as abnormal.
6. The real-time anomaly detection method for a high-speed train automatic protection system according to claim 5, characterized in that, The log key anomaly detection module, during the training phase, will... g Let K = {k1, k2, ..., k} be the conditional probability distribution of the target log key. A loss function is constructed based on the difference between this distribution and the actual log key at the corresponding time step to train the log key anomaly detection module, thereby updating the module's internal parameters. The goal is to learn a conditional probability distribution that maximizes the training log key sequence; where K = {k1, k2, ..., k}. u } represents a set of specified log keys in the ATP system, where u is the number of log keys; during training, the first q log keys in the conditional probability distribution are marked as normal; q is a positive integer used to balance the anomaly detection rate and the false alarm rate; During the perception phase, for the target log key km g Based on the input log key window w, calculate the conditional probability distribution Pt[km]. g |w]={k1:p1, k2:p2,…, k u :p u }, where w={km g-h ,…, km g-2 ,km g-1 } Contains the target log key km g For the first h most recent log keys, any element in w belongs to set K, k1:p1, k2:p2,…, k u :p u Let p1 represent the probability corresponding to log key k1, p2 represent the probability corresponding to log key k2, and p3 represent the probability corresponding to log key k4. u The corresponding probability is p u Sort the log keys in descending order of probability and determine the first q log keys as the corresponding candidate log keys; if km g If there are q candidate log keys, then the target log key is km. g Mark as normal; otherwise, mark as abnormal.
7. The real-time anomaly detection method for a high-speed train automatic protection system according to claim 5, characterized in that, During the training phase, the parameter value anomaly perception module takes a vector parameter as input and outputs a real value vector as output, which serves as the predicted value of the next input vector parameter. By dynamically adjusting the weights of the LSTM model, the error between the next input vector parameter and the predicted value of the next input vector parameter is reduced. During the perception phase, at each time step, the error between the input vector parameters and the predicted value is modeled as a Gaussian distribution. If the error is within the confidence interval of the predicted value, the input vector parameters are marked as normal; otherwise, they are marked as abnormal.
8. A real-time anomaly detection method for a high-speed train automatic protection system according to claim 5, characterized in that, When the sensing phase identifies an anomaly in the ATP system, it issues an early warning by expanding the display of the DMI, guiding the user to execute the anomaly handling procedure.
9. A real-time anomaly detection method for a high-speed train automatic protection system according to claim 8, characterized in that, Different exception handling processes are abstracted into finite state automata. The state transitions and execution order of the finite state automata are used to characterize the system exception handling process. Based on this, a knowledge base is formed, which can be reused by different operation and maintenance entities across the entire network to handle the same type of exceptions.
Citation Information
Patent Citations
Container abnormal behavior detection method of LSTM network based on attention mechanism
CN112905421A
Business process anomaly detection method based on attention mechanism
CN113807452A