A train current curve feature extraction and fault diagnosis method
By applying a neural network fault diagnosis network based on CNN and LSTM in urban rail transit, feature extraction and intelligent analysis of real-time current data of current collector boots is solved, and real-time high-precision diagnosis and early warning of current collector boots is realized.
Patent Information
- Application Number
- CN202410702287.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-02
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-06-02
AI Technical Summary
The traditional urban rail vehicle current collector boot fault diagnosis methods have problems such as insufficient real-time, low reliability, and complex algorithms, which are difficult to meet the needs of high-density and high-frequency operation of urban rail transit.
A fusion fault diagnosis network based on CNN and LSTM neural networks is adopted to pre-process, feature extraction and intelligent analysis of real-time current data of current collector boots, real-time online diagnosis of current collector boots removal faults is achieved.
Real-time high-precision diagnosis of current collector boot faults is realized, timeliness and reliability of fault detection is improved, labor costs and operation and maintenance costs are reduced, and the safe and efficient operation of urban rail transit is ensured.
Smart Images

Figure CN118779630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rail transit technology, and in particular to a train current curve feature extraction and fault diagnosis method. Background Art
[0002] As an important mode of transportation in modern cities, urban rail transit plays an increasingly important role in alleviating traffic congestion and reducing environmental pollution. The collector shoe is one of the key components of urban rail vehicles. It transmits current through physical contact with the third rail and provides traction power for the train. The safe and reliable operation of the collector shoe directly affects the normal operation of urban rail transit.
[0003] In actual operation, the collector shoe is prone to failures such as shoe removal and abnormal arcing, which will cause power outages in the train, not only reducing the train's operating efficiency but also threatening driving safety. Therefore, real-time online monitoring of the collector shoe's operating status and timely diagnosis and warning of shoe removal failures are of great significance to ensuring the safe and reliable operation of urban rail transit.
[0004] Traditional urban rail vehicle collector shoe fault diagnosis mainly relies on manual inspection and regular maintenance, which has problems such as slow response speed, high labor cost, and untimely detection. It is difficult to meet the needs of high-density and high-frequency operation of urban rail transit, and a real-time, efficient, and intelligent collector shoe fault online diagnosis technology is urgently needed. At present, there are related studies such as collector shoe video monitoring technology based on image processing, contact state evaluation technology based on arc spectrum analysis, and shoe removal detection technology based on current or acceleration signals. However, these methods generally have shortcomings such as lack of real-time performance, low reliability, and complex algorithms, making them difficult to be put into practical application.
[0005] The real-time current signal of the collector shoe contains rich status information. By extracting and intelligently analyzing the current signal, early diagnosis and early warning of the collector shoe fault can be achieved. However, affected by the complex environment and operating conditions of urban rail transit, the real-time current signal of the collector shoe presents complex characteristics such as non-stationary, multi-scale, and strong noise. Traditional signal processing methods are difficult to accurately extract fault features. Summary of the invention
[0006] The purpose of the present invention is to provide a train current curve feature extraction and fault diagnosis method in order to overcome the defects of the above-mentioned prior art. By combining CNN and LSTM neural network algorithms, the fault feature information in the output current signal of the collector shoe current loop is fully mined to realize real-time online diagnosis of collector shoe unbundling fault, which is beneficial to improving the safety, reliability and operational efficiency of the urban rail transit system.
[0007] The purpose of the present invention can be achieved by the following technical solution: A train current curve feature extraction and fault diagnosis method comprises the following steps:
[0008] S1, collecting the original current time series output when the collector shoe is in contact with the third rail and performing preprocessing;
[0009] S2. Process the preprocessed current time series in a sliding pane manner to construct a sample data set of shoe rail status labels:
[0010] Sampling: Let the sliding pane width be T 2 , the sliding window time step is t, and continuous current data is taken from the current time series to form a matrix sample containing strong spatiotemporal features at a certain moment. 2 ) The sampling data is:
[0011]
[0012] Rearrangement: First, the T of each current loop output 2 The current values are arranged in sequence into a matrix of size T×T, and then the four current loop matrices are stacked into a sample tensor of size T×T×4. The rearranged current sample space can be expressed as r∈R T×T×4 .
[0013] S3. Construct a fusion fault diagnosis network based on CNN and LSTM. After the input data is convolutionally operated to extract spatial features, a bidirectional LSTM and attention mechanism are used to capture temporal dependencies. The network structure specifically includes:
[0014] Input layer: used to input preprocessed sample data. During model training, the input data shape is (batch, T, T, 4), which represents a batch of sample data during model training. Batch is the batch size, which indicates the number of samples in one training or inference.
[0015] CNN layer: used to convolve the input samples and extract the spatial features of the current data, including the first convolution layer, the first pooling layer, the second convolution layer, the second pooling layer and the flattening layer. The feature map of each layer can be expressed as The flattening layer flattens the three-dimensional feature map output by the second pooling layer into a one-dimensional feature vector sequence of length (w×h×d).
[0016] Bidirectional LSTM layer: used to extract the temporal dependencies of current data. It includes two independent LSTM layers, the forward LSTM layer and the backward LSTM layer. The forward LSTM layer processes the input sequence in the order of time steps, and the backward LSTM layer processes the input sequence in the reverse order of time steps.
[0017] Attention mechanism layer: used to make the network focus on key time steps and suppress invalid information
[0018] Fully connected layer: receives c and h output by the attention mechanism layer t And as its input, perform operation y t =ReLU(w y ·[c,h t ]+b y ), ReLU(x)=max(0,x) is the activation function, and the number of hidden units is appropriately reduced.
[0019] Output layer: performs operations Output fault category probability distribution, is the activation function, and n is the number of fault categories.
[0020] S4, use the sample data set to train and optimize the feature extraction and fault diagnosis model, and dynamically adjust the network parameters through the back propagation algorithm;
[0021] S5. Deploy the fault diagnosis model, input the real-time current data of the collector shoe and output the multi-classification shoe failure diagnosis results to achieve real-time detection and early warning of the collector shoe failure.
[0022] In the present invention, the original current sequence collected in step S1 when the collector shoe is in contact with the third rail is 4-dimensional current data output by four current loops, and the matrix is represented as in, is the current characteristic vector collected by the collector for the first time, is the current characteristic vector obtained by sampling for the i-th time, is the current characteristic vector obtained by the last sampling;
[0023] Furthermore, the specific steps of preprocessing the original current time series in step S1 include:
[0024] S11. Missing value supplement: Use linear interpolation method to supplement missing values. The formula is:
[0025]
[0026] Among them, x 0 is the value before the missing value, x 1 is the value after the missing value, w a is the number of missing values in the vacancy, and μ is the order of the missing values to be filled in the vacancy.
[0027] S12, filtering and denoising: Use the moving average filtering method to smooth the completed current time series. Assuming the sliding window width is (2w+1), the smoothing formula for the current feature vector collected for the i-th time is:
[0028]
[0029] S13, normalization: use the Min-Max normalization method to normalize the current value of each feature dimension after filtering and map it to the range of [0,1]:
[0030]
[0031] in, Represents the current value of the cth feature dimension collected for the ith time after normalization.
[0032] In the present invention, the cell calculation process of the bidirectional LSTM layer in step S3 includes:
[0033] f t =σ(w f ·[h (t-1) ,p t ]+b f )
[0034] i t =σ(W f ·[h (t-1) ,p t ]+b i )
[0035]
[0036] o t =σ(w o ·[h (t-1) ,p t ]+b o )
[0037] h t =o t *tanh(C t )
[0038] Among them, input p t is the feature vector extracted by the CNN layer at a certain moment, and the output is h t Encodes the temporal dependencies of the entire process before the current moment, C t To record the hidden state of temporal dependence, σ represents the sigmoid activation function Tanh represents the activation function The operation “·” represents matrix multiplication, the operation “*” represents element-by-element matrix multiplication, the parameter W is the weight matrix, and the parameter b is the bias term.
[0039] In the present invention, the calculation process of the attention mechanism layer in step S3 includes:
[0040] e t =v T tanh(W h h t +W q q+b h )
[0041]
[0042] Among them, q is the query vector, v is the attention vector, and W h is the weight matrix, e t is the correlation between the current hidden state and the query vector, b h is the bias term, α t is the importance weight of the current hidden state, and c is the weighted sum of the input hidden sequence, highlighting the important feature information in the current temporal relationship.
[0043] In the present invention, the feature extraction and fault diagnosis model training in step S4 specifically includes the following steps:
[0044] Construction of fault diagnosis network integrating S41, CNN and LSTM;
[0045] S42, data set division: divide the labeled sample data set into a training set, a validation set and a test set according to a certain ratio;
[0046] S43, model training and optimization: Use the training set to train the model, use the cross entropy loss function to measure the model fault diagnosis effect in the iterative process, and use the Adam optimization algorithm to update the model trainable parameters;
[0047] S44, Model evaluation: At the end of each training epoch, the validation set is used to evaluate the model performance, and hyperparameters are tuned according to the changes in performance indicators (accuracy, F1 value, etc.) and loss functions. The tuning method uses random search. After the model training is completed, the generalization ability of the trained model is evaluated on the test set.
[0048] In the present invention, the model deployment and fault diagnosis process in step S5 specifically includes:
[0049] S51, deploy the trained CNN-LSTM fusion fault diagnosis model to the train on-board monitoring equipment;
[0050] S52, the device collects the current data output by the four current loops of the collector shoe in real time, and constructs a feature input sample in the same way as the data preprocessing;
[0051] S53, the model performs online reasoning on the input sample and outputs the boot-off fault diagnosis result and confidence level;
[0052] S54. Furthermore, the diagnosis results can be displayed in real time on the display screen in the cab, and when a fault occurs, the driver can be reminded to take emergency measures through sound and light alarms.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] (1) A novel train current curve feature extraction method is proposed. Using machine learning algorithms, especially the CNN-LSTM neural network model based on deep learning, the real-time current data of the collector shoe is automatically extracted and intelligently analyzed, overcoming the challenges of non-stationarity, multi-scale, and strong noise faced by traditional signal processing methods, and achieving real-time and high-precision diagnosis of the collector shoe failure, greatly improving the timeliness and reliability of fault detection, and ensuring driving safety.
[0055] (2) The intelligent and automated fault diagnosis of the collector shoe is realized. Through the CNN-LSTM neural network model, the real-time current data of the collector shoe can be automatically and continuously analyzed and diagnosed without human intervention. This intelligent diagnosis method can greatly reduce manpower input and reduce operating and maintenance costs. At the same time, the intelligent algorithm based on big data and deep learning can continuously improve its diagnostic performance with the continuous accumulation of data and continuous optimization of the model. It has good sustainability and self-improvement capabilities, and provides a new efficient, economical and reliable way for fault diagnosis of urban rail transit collector shoes.
[0056] (3) It has good versatility and scalability. On the one hand, the method of the present invention can be widely applied to various types of urban rail transit vehicles, such as subways, light rails, etc., without being restricted by vehicle models and operating environments; on the other hand, the technical framework and ideas of the present invention can be extended to other types of fault diagnosis, such as turnout switch machine failures, contact network failures, etc. In addition, the CNN-LSTM neural network used in the present invention can also be flexibly adjusted and optimized according to actual needs to adapt to different data characteristics and fault diagnosis tasks. Therefore, the present invention provides a universal, scalable, intelligent solution for equipment fault diagnosis in urban rail transit and even other fields, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0058] Figure 1 It is a flow chart of a collector shoe failure diagnosis method according to an embodiment of the present invention;
[0059] Figure 2 Schematic diagram of a sliding window sample tensor generation method according to an embodiment of the present invention;
[0060] Figure 3 2 is a diagram of a CNN-LSTM fusion fault diagnosis network architecture according to an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, other embodiments obtained by ordinary technicians in this field without creative work all belong to the protection scope of the present invention.
[0062] Example 1
[0063] In this embodiment, a train current curve feature extraction and fault diagnosis method is provided. Figure 1 As shown, the following steps are included:
[0064] (1) Collecting the original current time series output when the collector shoe is in contact with the third rail and performing preprocessing;
[0065] (2) The preprocessed current time series is sampled, rearranged and calibrated in a sliding pane manner to construct a sample dataset of shoe rail state labels;
[0066] (3) Construct a fusion fault diagnosis network based on CNN and LSTM;
[0067] (4) Use sample data sets to train and optimize feature extraction and fault diagnosis models;
[0068] (5) Deploy the fault diagnosis model, input the real-time current data of the collector shoe and output the shoe-off fault diagnosis results.
[0069] Through the above steps, feature extraction and boot-off fault diagnosis are performed based on the real-time current data output by the collector shoe current loop. Compared with the existing technology, this method can greatly improve the accuracy and real-time performance of boot-off fault detection by using the CNN-LSTM neural network to automatically extract current curve features and intelligently analyze them. At the same time, it can effectively reduce labor costs and operation and maintenance costs, improve the operation efficiency of urban rail transit, and ensure driving safety.
[0070] The following will be described in conjunction with a specific optional embodiment.
[0071] In this embodiment, a current collector is used to collect the current curve data of the four current loops output by the collector shoe of a subway train within one week, including the current data of normal and shoe-off failure.
[0072] (I) Collect the current time series of the four current loops output by the collector shoe of a subway line. The current values are expressed as I M ,I N ,I P ,I Q It means that the current data sampling frequency is 20Hz, that is, 20 sets of current data are collected per second, and the data is collected continuously for four hours, that is, a total of 288,000 sets of current time series are collected, forming a 4×288,000 current matrix in, is the current characteristic vector collected by the collector for the first time, is the current characteristic vector obtained by the i-th acquisition, is the current characteristic vector collected last time.
[0073] The collected data are then preprocessed, including missing value supplementation, filtering and normalization.
[0074] Preferably, a linear interpolation method is used to supplement the missing values, and the calculation formula is as follows:
[0075]
[0076] Preferably, the moving average method is used to smooth the current time series, with a window width of 5:
[0077]
[0078] Preferably, the Min-Max normalization method is used to normalize each current feature dimension after filtering:
[0079]
[0080] (ii) The preprocessed current time series is sampled, rearranged and calibrated in a sliding window manner, as follows: Figure 2 As shown:
[0081] Sampling: Assume that the sliding window width is 100 and the sliding window time step is 50, which means that every 5s of current data constitutes a sampling sample, and the sample interval is 2.5s;
[0082] Rearrangement: First, the 100 current values output by each current loop are arranged into a 10×10 matrix in sequence, and then the four current loop matrices are stacked into a 10×10×4 sample tensor. The rearranged current sample space can be expressed as r∈R 10 ×10×4 , the sample capacity of this embodiment is 5760.
[0083] Calibration: Calibrate the rearranged sample data according to established rules or relevant standards. In this embodiment, the collector shoe status labels are divided into four categories: normal, unshoe, fault, and other abnormalities, numbered 0, 1, 2, and 4 respectively, that is, the sample data set R for model training is obtained. 1 ,r 2 ,r 3 ,…,r 5760}.
[0084] (III) Construct a fusion fault diagnosis network based on CNN and LSTM, such as Figure 3 As shown, the specific architecture includes:
[0085] Input layer: used to input preprocessed sample data. During model training, the input data shape is (batch, T, T, 4), which represents a batch of sample data during model training. Batch is the batch size, which indicates the number of samples in one training or inference.
[0086] CNN layer: used to convolve the input samples and extract the spatial features of the current data, including the first convolution layer, the first pooling layer, the second convolution layer, the second pooling layer and the flattening layer. The feature map of each layer can be expressed as The flattening layer flattens the three-dimensional feature map output by the second pooling layer into a one-dimensional feature vector sequence of length (w×h×d).
[0087] Bidirectional LSTM layer: used to extract the temporal dependency of current data, including two independent LSTM layers, the forward LSTM layer and the backward LSTM layer. The forward LSTM layer processes the input sequence in the order of time steps, and the backward LSTM layer processes the input sequence in the reverse order of time steps. Each LSTM layer consists of multiple cells, and the calculation process of each cell is as follows:
[0088] f t =σ(w f ·[h (t-1) ,p t ]+bf )
[0089] i t =σ(w f ·[h (t-1) ,p t ]+b i )
[0090]
[0091] o t =σ(w o ·[h (t-1) ,p t ]+b o )
[0092] h t =o t *tanh(C t )
[0093] Among them, input p t is the feature vector extracted by the CNN layer at a certain moment, and the output is h t Encodes the temporal dependencies of the entire process before the current moment, C t To record the hidden state of temporal dependence, σ represents the sigmoid activation function Tanh represents the activation function The operation “·” represents matrix multiplication, the operation “*” represents element-by-element matrix multiplication, the parameter W is the weight matrix, and the parameter b is the bias term.
[0094] Attention mechanism layer: used to make the network focus on key time steps and suppress invalid information. The calculation process is as follows:
[0095] e t =v T tanh(W h h t +W q q+b h )
[0096]
[0097] Among them, q is the query vector, v is the attention vector, and W h is the weight matrix, e t is the correlation between the current hidden state and the query vector, b h is the bias term, α t is the importance weight of the current hidden state, and c is the weighted sum of the input hidden sequence, highlighting the important feature information in the current temporal relationship.
[0098] Fully connected layer: receives c and h output by the attention mechanism layert And as its input, perform operation y t =ReLU(w y ·[c,h t ]+b y ), ReLU(x)=max(0,x) is the activation function, and the number of hidden units is appropriately reduced.
[0099] Output layer: performs operations Output fault category probability distribution, is the activation function, and n is the number of fault categories.
[0100] The hyperparameter settings of the model architecture are shown in the following table. Preferably, random search is used to tune the hyperparameters.
[0101]
[0102] (IV) Training a fault diagnosis model based on the constructed CNN-LSTM fusion fault diagnosis network. Divide the calibration sample data set into a training set, a validation set, and a test set in a ratio of 8:1:1, and use the training set to train the fault classification model. Preferably, the training iteration process uses a cross entropy loss function and an Adam optimization algorithm to update the model hyperparameters.
[0103] Since this embodiment is a multi-classification problem, the cross entropy loss function is expressed as:
[0104]
[0105] Use the Adam optimizer to replace the traditional stochastic gradient descent optimization algorithm. The specific steps are:
[0106] 1) Calculate the exponentially weighted unbiased estimate of the gradient:
[0107] M t =β 1 *M t-1 +(1-β 1 )*g t
[0108] 2) Calculate the exponentially weighted unbiased estimate of the squared gradient:
[0109]
[0110] 3) Modify M t and V t To obtain unbiased estimates;
[0111] 4) According to the revised M t and V t Calculate parameter update value:
[0112]
[0113] θ t+1 =θ t +Δθ t
[0114] Among them, β 1 and β 2 is the exponentially weighted decay rate, controlling the influence of past gradients; g t is the current gradient; α is the learning rate; ε is a very small positive number to avoid the denominator being 0.
[0115] The model was trained with a learning rate of 0.001, epochs of 100, batch sizes of 16, and GPUs.
[0116] Preferably, after the model is trained, its performance is evaluated using accuracy and harmonic mean F1:
[0117]
[0118] Among them, TP (true positive), TN (true negative), FP (false positive), and FN (false negative) respectively represent the four situations of the prediction results of the fault diagnosis model.
[0119] (6) Deploy the trained CNN-LSTM fusion fault diagnosis model, input the preprocessed real-time current data of the collector shoe, and output the collector shoe failure diagnosis result.
[0120] Obviously, the above embodiments are only examples provided to clearly illustrate the principles of the present invention and should not be regarded as limitations on the implementation methods. For those of ordinary skill in the art, various modifications and variations can be made to the above embodiments according to specific needs without departing from the spirit and scope of the present invention. The embodiments listed herein cannot exhaust all possible changes and variations. Any modifications, equivalent substitutions, improvements, etc. made to the above embodiments based on the technical essence of the present invention should be included in the protection scope of the present invention.
Claims
1. A train current curve feature extraction and fault diagnosis method, characterized in that: The following steps are involved: S1, collecting the original current time series output when the collector shoe is in contact with the third rail and performing preprocessing; S2. Process the preprocessed current time series in a sliding pane manner to construct a sample data set of shoe rail status labels: Sampling: Let the sliding pane width be T 2 , the sliding window time step is t, and continuous current data is taken from the current time series to form a matrix sample containing strong spatiotemporal features at a certain moment. 2 ) The sampling data is: Rearrangement: First, the T of each current loop output 2 The current values are arranged in sequence into a matrix of size T×T, and then the four current loop matrices are stacked into a sample tensor of size T×T×4. The rearranged current sample space can be expressed as r∈R T×T×4 ; Calibration: Calibrate the rearranged sample data according to established rules or relevant standards. The status labels are numbered 0, 1, 2, ..., n according to the specific fault classification, and the sample data set for model training is obtained; S3. Construct a fusion fault diagnosis network based on CNN and LSTM. After the input data is convolutionally operated to extract spatial features, a bidirectional LSTM and attention mechanism are used to capture temporal dependencies. The network structure specifically includes: Input layer: used to input preprocessed sample data. The input data shape during model training is (batch, T, T, 4), which represents a batch of sample data during model training. Batch is the batch size, which indicates the number of samples in one training or inference. CNN layer: used to convolve the input samples and extract the spatial features of the current data, including the first convolution layer, the first pooling layer, the second convolution layer, the second pooling layer and the flattening layer. The feature map of each layer can be expressed as The flattening layer flattens the three-dimensional feature map output by the second pooling layer into a one-dimensional feature vector sequence of length (w×h×d). Bidirectional LSTM layer: used to extract the temporal dependency of current data. It includes two independent LSTM layers, the forward LSTM layer and the backward LSTM layer. The forward LSTM layer processes the input sequence in the order of time steps, and the backward LSTM layer processes the input sequence in the reverse order of time steps. Attention mechanism layer: used to make the network focus on key time steps and suppress invalid information; Fully connected layer: receives c and h output by the attention mechanism layer t And as its input, perform operation y t =ReLU(w y ·[c,h t ]+b y ), ReLU(x)=max(0,x) is the activation function, and the number of hidden units is appropriately reduced in dimension; Output layer: performs operations Output fault category probability distribution, is the activation function, n is the number of fault categories; S4, using the sample data set to train and optimize the feature extraction and fault diagnosis model; S5. Deploy the fault diagnosis model, input the real-time current data of the collector shoe and output the multi-classification shoe failure diagnosis results to achieve real-time detection and early warning of the collector shoe failure.
2. The train current curve feature extraction and fault diagnosis method according to claim 1 is characterized in that: The original current time series in step S1 is the four-dimensional current data output by the four current loops of the collector shoe, and the matrix is expressed as in, is the current characteristic vector collected by the collector for the first time, is the current characteristic vector obtained by sampling for the i-th time, is the current characteristic vector obtained by the last sampling.
3. The train current curve feature extraction and fault diagnosis method according to claim 1 is characterized in that: The data preprocessing in step S1 specifically includes the following steps: S11. Missing value supplement: Use linear interpolation method to supplement missing values. The formula is: Among them, x0 is the value before the missing value, x1 is the value after the missing value, and w a is the number of missing values in the vacancy, and μ is the order of the missing values to be filled in the vacancy; S12, filtering and denoising: Use the moving average filtering method to smooth the completed current time series. Assuming the sliding window width is (2w+1), the smoothing formula for the current feature vector collected for the i-th time is: S13, normalization: use the Min-Max normalization method to normalize the current value of each feature dimension after filtering and map it to the range of [0,1]: in, Represents the current value of the cth feature dimension collected for the ith time after normalization.
4. The train current curve feature extraction and fault diagnosis method according to claim 1 is characterized in that: The cell calculation process of the bidirectional LSTM layer in step S3 includes: f t =σ(w f ·[h (t-1) ,p t ]+b f ) i t =σ(w f ·[h (t-1) ,p t ]+b i ) the t =σ(w o ·[h (t-1) ,p t ]+b o ) h t =o t *tanh(C t ) Among them, input p t is the feature vector extracted by the CNN layer at a certain moment, and the output is h t Encodes the temporal dependencies of the entire process before the current moment, C t To record the hidden state of temporal dependence, σ represents the sigmoid activation function Tanh represents the activation function The operation "·" represents matrix multiplication, the operation "*" represents element-by-element matrix multiplication, the parameter W is the weight matrix, and the parameter b is the bias term.
5. The train current curve feature extraction and fault diagnosis method according to claim 1 is characterized in that: The calculation process of the attention mechanism layer in step S3 include: e t =v T tanh(W h h t +W q q+b h ) Among them, q is the query vector, v is the attention vector, and W h is the weight matrix, e t is the correlation between the current hidden state and the query vector, b h is the bias term, α t is the importance weight of the current hidden state, and c is the weighted sum of the input hidden sequence, highlighting the important feature information in the current temporal relationship.
Citation Information
Patent Citations
CNN-LSTM-based depth learning method and multi-attribute time sequence data fault diagnosis method
CN109814523A
Straddle type monorail shoe rail monitoring system and control method thereof
CN116086522A