Track health diagnosis method and device

By building a balanced dataset and using dynamic contrast loss function to train feature extraction models, the low classification accuracy and low robustness problems caused by sample imbalance in traditional orbital health diagnostic methods are solved, and higher diagnostic accuracy and robustness are achieved.

CN120067796APending Publication Date: 2025-05-30CHINA RAILWAY SIYUAN SURVEY & DESIGN GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126209.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional data-driven orbital health diagnosis methods are less classified and poorly robust because the cross-entropy loss function is susceptible to factors such as abnormal data, measurement noise and sample imbalance.

Method used

By obtaining the passing data of the train passing through the track, constructing the original data set, and constructing a balanced data set through downsampling, the feature extraction model is pre-trained by the classic contrast loss function, and the loss function is dynamically adjusted to adapt to the classification accuracy of different categories of samples to achieve re-training of the feature extraction model.

Benefits of technology

It improves the accuracy and robustness of orbital health diagnosis, reduces the adverse effects of unbalanced data on the model loss function, and enhances resistance to abnormal data and noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067796A_ABST
    Figure CN120067796A_ABST
Patent Text Reader

Abstract

The invention discloses a track health diagnosis method and device, and the method comprises the steps: obtaining vehicle passing data, marking the vehicle passing data based on a track state, and constructing an original data set; performing down-sampling on the original data set to construct a balanced data set, inputting samples in the balanced data set into the feature extraction model to output corresponding feature vectors, and pre-training the feature extraction model through a comparative learning method; selecting a first sample group from the original data set, inputting samples in the first sample group into a feature extraction model to obtain first feature vectors, and classifying the first feature vectors to obtain classification accuracy scores of various samples; adjusting a model loss function based on classification accuracy scores of various samples, training a feature extraction model through comparative learning, and continuing to dynamically adjust model parameters according to training and classification results; and repeating the steps until the parameters of the model converge to obtain a trained feature extraction model. According to the invention, the accuracy and robustness of track health diagnosis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of track detection, and particularly to a method and device for track health diagnosis. Background Art

[0002] Main components of the track structure, such as rails, fasteners, track slabs, and shear keys, will inevitably develop various types of diseases under the repeated action of long-term train loads, threatening the safety of train operation. Structural health monitoring technology has become an important means to ensure the safety of track operation and maintenance. How to utilize massive monitoring data to accurately evaluate the track status and identify the types of diseases is of great significance for track operation and maintenance. Traditional data-driven track health diagnosis methods usually adopt the cross-entropy function and directly train a deep neural network model using labeled data to achieve the purpose of classifying input signals. However, due to the fact that the cross-entropy loss function is easily affected by factors such as abnormal data, measurement noise, and sample imbalance, the classification accuracy of traditional methods is relatively low and the robustness is poor. Summary of the Invention

[0003] The present invention provides a method and device for track health diagnosis, which can reduce the adverse effects of unbalanced data on the loss function of the model and improve the accuracy and robustness of track health diagnosis.

[0004] According to one aspect of the present invention, there is provided a method for track health diagnosis, including:

[0005] Obtaining passing train data when a train passes through a track, and labeling the passing train data based on the track status to construct an original data set;

[0006] Performing downsampling on the original data set to construct a balanced data set, inputting samples in the balanced data set into a feature extraction model, outputting corresponding feature vectors, and pre-training the feature extraction model through a classical contrast loss function; the number of samples of each type of label in the balanced data set is the same; the feature extraction model is used to extract features of input samples and output feature vectors;

[0007] S1: Selecting a preset number of first sample groups from the original data set, inputting samples in the first sample groups into the feature extraction model, obtaining first feature vectors of each sample, and classifying each first feature vector by using a preset classification method to obtain classification accuracy scores of samples of each type of state in the first sample groups;

[0008] S2: Adjusting the contribution of each type of sample to the loss function based on the classification accuracy scores of each type of sample in the first sample groups, making samples with low accuracy contribute more and samples with high accuracy contribute less, and retraining the feature extraction model through a dynamic contrast loss to adjust the parameters of the feature extraction model;

[0009] Repeat steps S1 - S2 until the parameters of the feature extraction model converge, and obtain the trained feature extraction model.

[0010] Optionally, the obtaining of the passing train data when the train passes through the track and the annotation of the passing train data based on the track state to construct the original data set includes:

[0011] Obtain the monitoring data when the train passes through the track in a healthy state and the monitoring data when the train passes through the tracks in various types of disease states;

[0012] Intercept the passing train data in the monitoring data through a signal window of equal width, and annotate the passing train data based on the track state. The annotated states include the healthy state and various disease states;

[0013] Preprocess the passing train signal. The preprocessing includes one or more of downsampling, filtering, and standardization to construct an original data set without abnormal data.

[0014] Optionally, construct a balanced data set based on the downsampling of the original data, input the samples in the balanced data set into the feature extraction model to output the corresponding feature vectors, and pre - train the feature extraction model through a classical contrast loss function. It includes:

[0015] Perform data augmentation on the data samples in the original data set; the data augmentation includes one or more of adding noise, random scaling, and random translation.

[0016] Optionally, after the step of repeating steps S1 - S2 until the parameters of the feature extraction model converge and obtaining the trained feature extraction model, it further includes:

[0017] Construct a linear classification model. The input of the linear classifier is the feature vector, and the output is the probability that the track belongs to various states;

[0018] Jointly construct an orbit state diagnosis model by the trained feature extraction model and the linear classifier; the input of the orbit state diagnosis model is the passing train signal, and the output is the probability that the track belongs to various states;

[0019] Pre - train the orbit state diagnosis model using the balanced data set and the classical contrast loss function, and retrain the orbit state diagnosis model using the original data set and the dynamic contrast damage function to obtain the trained orbit state diagnosis model.

[0020] Optionally, the feature extraction model consists of a multi - branch convolutional neural network model and a self - attention mechanism;

[0021] The multi-branch convolutional neural network uses convolutional kernels of different sizes to extract different-scale information of the input samples; the feature dimensions of the different-scale information output by the multi-branch convolutional neural network are the same. After the different-scale information is concatenated, the multi-head self-attention mechanism is used for feature extraction, and multi-scale features are fused;

[0022] The fused features are output as a one-dimensional feature sequence through the non-linear transformation of a multi-layer neural network. The one-dimensional feature sequence is the feature vector of the input sample, and the length of the one-dimensional feature sequence is less than the length of the input data.

[0023] Optionally, to construct a balanced dataset, the samples in the balanced dataset are input into the feature extraction model to output corresponding feature vectors, and the feature extraction model is pre-trained through a classical contrast loss function. The classical contrast loss function is:

[0024]

[0025] In the formula, z · represents the feature vector output by the backbone network of the feature extraction model when the input sample is x · ; P(i) is the set of samples of the same class as the sample x i , and the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the set of samples in the first sample group; sim(·) is the similarity metric function; represents the temperature coefficient, which is used to adjust the distribution of similarity scores and is a hyperparameter;

[0026] Samples of the same class form positive sample pairs, and samples of different classes form negative sample pairs; the cosine similarity function is used for the similarity measurement between samples, as follows:

[0027]

[0028] Optionally, a preset classification method is used to classify each first feature vector to obtain the classification accuracy scores of various state samples in the first sample group, including:

[0029] Using the method of clustering or support vector machine to classify each first feature vector output by the feature extraction model;

[0030] According to the classification results of the first feature vectors corresponding to the first sample group, calculate the classification accuracy scores of various state samples in the first sample group.

[0031] Optionally, the calculation formula for the classification accuracy is:

[0032]

[0033] Among them, TPi is the true positive of the i-th type of sample, representing the number of samples correctly predicted as the positive class; FP i is the false positive of the i-th type of sample, representing the number of samples wrongly predicted as the positive class; FN i is the false negative of the i-th type of sample, representing the number of samples wrongly predicted as the negative class.

[0034] Optionally, adjust the contribution of each type of sample to the loss function based on the classification accuracy scores of each type of sample in the first sample group. The dynamic contrast loss function is as follows:

[0035]

[0036] In the formula, z · represents the feature vector output by the backbone network of the feature extraction model when the input sample is x · ; P(i) is the set of samples of the same class as the sample x i , and the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the sample set of the first sample group; sim(·) is the similarity metric function; represents the temperature coefficient, which is used to adjust the distribution of the similarity scores and is a hyperparameter; F i represents the classification accuracy of the i-th type of sample corresponding to the first sample group.

[0037] According to another aspect of the present invention, there is provided an orbit health diagnosis device, including:

[0038] A data acquisition unit, configured to acquire passing train data when the train passes through the orbit, and label the passing train data based on the orbit state to construct an original data set;

[0039] A pre-training unit, configured to downsample the original data set to construct a balanced data set, input the samples in the balanced data set into the feature extraction model, output the corresponding feature vectors, and pre-train the feature extraction model through a classical contrast loss function; the number of samples with each type of label in the balanced data set is the same; the feature extraction model is used to extract the features of the input samples and output feature vectors;

[0040] A feature extraction unit, configured to select a preset number of first sample groups from the original data set, input the samples in the first sample group into the feature extraction model, obtain the first feature vector of each sample, classify each first feature vector by using a preset classification method, and obtain the classification accuracy scores of each type of state samples in the first sample group;

[0041] The first training unit is configured to adjust the contribution of each type of sample to the loss function based on the classification accuracy scores of various samples in the first sample group, such that samples with low accuracy contribute more and samples with high accuracy contribute less, and retrain the feature extraction model through dynamic contrast loss to adjust the parameters of the feature extraction model;

[0042] The second training unit is configured to repeat the steps of the feature extraction unit and the training unit until the parameters of the feature extraction model converge, thereby obtaining the trained feature extraction model.

[0043] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0044] At least one processor; and

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the track health diagnosis method according to any embodiment of the present invention.

[0047] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the track health diagnosis method according to any embodiment of the present invention when executed by a processor.

[0048] The technical solution of the embodiments of the present invention uses a balanced dataset as input, pre-trains a feature extraction model by means of contrastive learning to obtain initial parameters of the feature extraction model. When training the feature extraction model with an unbalanced original dataset, samples in the original dataset are input into the feature extraction network in batches for feature extraction, and the classification accuracy needs to be calculated once after the feature vectors of each batch of samples are extracted, and the parameters of the loss function are adjusted based on the classification accuracy. The feature extraction model is trained by means of contrastive learning. Since the loss function used in each training adjusts the contribution of different types of samples in real time according to the classification accuracy of different types of samples, the problem of sample imbalance is solved.

[0049] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0051] Figure 1 is a flowchart of a track health diagnosis method provided in Embodiment 1 of the present invention;

[0052] Figure 2 The architecture diagram of the track health diagnosis model in a track health diagnosis method provided in an embodiment of the present invention;

[0053] Figure 3 is a flowchart of a track health diagnosis method provided in Embodiment 2 of the present invention;

[0054] Figure 4 is a structural diagram of a track health diagnosis device provided in Embodiment 2 of the present invention;

[0055] Figure 5 is a schematic structural diagram of an electronic device for implementing the track health diagnosis method of the embodiments of the present invention. Detailed Embodiments

[0056] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0058] Embodiment 1

[0059] Figure 1 A flow chart of a track health diagnosis method is provided for the first embodiment of the present invention. Figure 1 As shown, the method includes:

[0060] S101. Obtaining train passing data when a train passes through a track, and annotating the train passing data based on the track state to construct an original data set.

[0061] In order to detect the state of the rails, the data of trains passing through the rails in different states can be obtained. In this embodiment, the data of trains passing through the rails in healthy state and various types of disease states can be collected as training data. The various disease states can include shear hinge fracture, isolator support failure, spring bar fracture, etc.

[0062] In one embodiment, a supervised training method may be used, that is, each sample of the training data is labeled, and the vehicle passing data and the state label corresponding to the vehicle passing data are used as a sample to construct the original data set.

[0063] It should be noted that, generally speaking, the number of samples in healthy state among the collected samples is much larger than the number of samples in each type of disease state, that is, the original data set is usually an unbalanced data set, which includes a larger proportion of samples in healthy state and a smaller proportion of samples in each type of disease state.

[0064] S102, downsampling the original data set to construct a balanced data set, inputting samples in the balanced data set into a feature extraction model, outputting corresponding feature vectors, and pre-training the feature extraction model through a classic contrast loss function.

[0065] In this embodiment, the feature extraction model is used to extract the features of the input sample data and output a feature vector, so that each input sample data has a corresponding feature vector as a feature representation of the vehicle passing data.

[0066] In this embodiment, the sample data in the original data set can be downsampled. In order to ensure that the number of samples of each category in the constructed balanced data set is consistent, the redundant data in each category in the original data set can be removed, so as to obtain the same number of samples under each category label. The feature extraction model is used to extract the features of each sample data in each category of sample data to obtain the corresponding feature vector, and the feature extraction model is pre-trained by the classic contrast loss function. The classic contrast loss function can be used to find the feature vector with the greatest similarity to the feature vector of the current sample as a positive sample in the feature vector, thereby shortening the distance between positive samples and pushing the distance between negative samples, thereby learning the effective representation of data and preliminarily training the feature extraction model.

[0067] S103. Select a preset number of first sample groups from the original data set, input the samples in the first sample groups into the feature extraction model to obtain the first feature vector of each sample, and classify each first feature vector using a preset classification method to obtain the classification accuracy scores of various state samples in the first sample groups.

[0068] It should be noted that each time during training, a preset number of first sample groups can be extracted from the original data set for training the feature extraction model. Among them, the preset number can be set as needed. For example, 300 samples can be selected from the original data set as the first sample group, or 500 samples can be selected from the original data set as the first sample group.

[0069] After inputting the samples in the first sample group into the feature extraction model, feature vectors with the same number as the samples in the first sample group can be generated. Common classification methods, such as clustering or support vector machine, can be used to classify the feature vectors corresponding to the first sample group to obtain the classification results.

[0070] The classification results can be compared with the labels of the samples corresponding to each feature vector, and thus the classification accuracy of the classification results can be obtained. In this embodiment, the calculation method of the classification accuracy can use the F-score as the classification accuracy, and then the F-scores corresponding to various state samples can be determined respectively.

[0071] S104. Adjust the contributions of various samples to the loss function based on the classification accuracy scores of various samples in the first sample group, so that samples with low accuracy contribute more and samples with high accuracy contribute less, and retrain the feature extraction model through the dynamic contrast loss to adjust the parameters of the feature extraction model.

[0072] It should be noted that after determining the classification accuracy scores of each state type sample in the first sample group, the loss function of the feature extraction model can be updated according to the state type of the samples input into the feature extraction model, that is, each time a sample is input, the loss function can be updated once according to the state type of the sample. The model is trained by dynamically updating the loss function through the state type of the input samples, so that during the training process, the contributions of different category samples can be adjusted in real time according to the classification accuracy of different category samples. Samples with low accuracy contribute more and samples with high accuracy contribute less. Different loss functions are selected according to different types of sample data to train the feature extraction model to obtain a trained feature extraction model. The trained feature extraction model using the solution of this embodiment can solve the influence of sample imbalance on the classification accuracy.

[0073] S105. Repeat steps S103 - S104 until the parameters of the feature extraction model converge to obtain the trained feature extraction model.

[0074] Repeat steps S103 - S104, that is, repeatedly select a preset number of first sample groups from the original dataset multiple times, and each time use the selected first samples to train the feature extraction model. It should be noted that in this embodiment, a sample selection method with replacement or without replacement can be adopted. Repeat the training multiple times until the parameters of the feature extraction model converge, and then stop the training to obtain the trained feature extraction model.

[0075] In the technical solution of the embodiment of the present invention, using the balanced dataset as the input, the feature extraction model is pre-trained by the contrastive learning method to obtain the initial parameters of the feature extraction model. When training the feature extraction model with the unbalanced original dataset, the samples in the original dataset are input into the feature extraction network in batches for feature extraction, and the classification accuracy needs to be calculated once after the feature vectors of each batch of samples are extracted, and the parameters of the loss function are adjusted based on the classification accuracy. The feature extraction model is trained by the contrastive learning method. Since the loss function used in each training adjusts the contributions of different category samples in real time according to the classification accuracy of different category samples, the problem of sample imbalance is solved.

[0076] Embodiment Two

[0077] Figure 3 is a schematic diagram of an orbit health diagnosis method provided by the second embodiment of the present invention, Figure 3 which includes:

[0078] S301. Obtain the monitoring data when the train passes through the healthy orbit and the monitoring data when the train passes through the orbits in various disease states.

[0079] It should be noted that the monitoring data of the train passing through the orbit can be collected by setting strain sensors on the orbit. The states of the orbit include healthy state, shear hinge fracture state, isolator support failure state, spring clip fracture state, etc.

[0080] S302. Intercept the passing - train data in the monitoring data through an equi - width signal window and label the passing - train data, and the labeled states include healthy state and various disease states.

[0081] Since the monitoring data includes passing - train data and environmental noise signals, it is necessary to intercept the passing - train signals from the monitoring data when actually obtaining the passing - train data. In this embodiment, an equi - width signal window can be used to intercept the passing - train data in the monitoring data, and the passing - train data is manually labeled to label the state types of the passing - train data. Among them, the state types include healthy state, shear hinge fracture state, isolator support failure state, spring clip fracture state, etc.

[0082] S303. Preprocess the passing vehicle signal, where the preprocessing includes one or more of downsampling, filtering, and normalization to construct an original dataset without abnormal data.

[0083] It should be noted that the original passing vehicle signal is preprocessed, and the preprocessing methods include downsampling, filtering, and normalization. Moreover, the preprocessed passing vehicle data and the labels corresponding to the passing vehicle data can be used as a sample to construct an original dataset without abnormal data.

[0084] S304. Construct a feature extraction model, which is used to extract the features of the input data and output a feature vector.

[0085] In one embodiment, the feature extraction model can adopt a multi-scale sequence feature extraction model, which includes a multi-branch convolutional neural network and a self-attention mechanism. Among them, the multi-branch convolutional neural network uses convolutional kernels of different sizes to extract different-scale information of the input sample data. The large convolutional kernel is used to extract low-frequency information, and the small convolutional kernel is used to extract high-frequency information. The feature dimensions of the different-scale information output by the multi-branch convolutional neural network are the same. After splicing the different-scale information, feature extraction is performed through the multi-head self-attention mechanism, and multi-scale features are fused.

[0086] The fused features are output as a one-dimensional feature sequence through the non-linear transformation of the multi-layer neural network. The one-dimensional feature sequence is the feature vector of the monitoring data, and the length of the one-dimensional feature sequence is less than the input data.

[0087] The feature extraction network adopted in this embodiment can be like Figure 2 the structure of the multi-scale self-attention encoder shown in Figure 2 The red square in it represents the multi-scale convolutional layer, and the orange square represents the self-attention mechanism.

[0088] S305. Perform data augmentation on the data samples in the original dataset; the data augmentation includes one or more of adding noise, random scaling, and random translation.

[0089] Perform data augmentation on the sample data in the original dataset to generate new samples, so as to increase the number and diversity of samples. Among them, the data augmentation adopts a random combination of one or more methods of adding noise, random scaling, and random translation:

[0090] Among them, adding noise to the sample data can be randomly adding a certain proportion of Gaussian noise to the sample data. For example:

[0091]

[0092] where x is the original signal; The signal after data augmentation; G is Gaussian noise, and its variance is determined according to the variance of the original signal.

[0093] Random scaling can be to randomly scale the amplitude of the sample data. For example:

[0094]

[0095] Among them, s is the scaling factor, which follows a Gaussian distribution or a uniform distribution with a mean of 1.

[0096] Random translation can be to randomly translate the sample data c. along the time axis, and the missing data is filled with overflow values. For example:

[0097]

[0098] Among them, i is the translation step along the time axis.

[0099] S306. Divide the original data set into a training set, a test set, and a validation set; downsample the data in the training set to construct a balanced data set.

[0100] It should be noted that the training set includes samples of different state types, and each state type of sample contains samples obtained after data augmentation of the samples in the original data set, so as to construct a new original data set, and the newly generated original data set is richer in data than the original data set.

[0101] S307. Input the samples in the balanced data set into the feature extraction model, output the corresponding feature vectors, and pre-train the feature extraction model through a classical contrast loss function.

[0102] It should be noted that the sample data in the balanced data set is input into the feature extraction model to obtain the corresponding feature vectors, and then Info-NCE contrast learning is performed based on the data with the same label in the same batch of sample groups, so as to pre-train the feature extraction model.

[0103] Among them, the classical contrast loss function Inf-ONCE Loss (Noise Contrastive Estimation Loss) is a loss function for self-supervised learning, used for learning feature representations or representation learning. It is based on the idea of information theory and learns the model parameters by comparing the similarity between positive samples and negative samples.

[0104] One loss function for contrastive learning is the Inf-ONCE Loss. Contrastive learning can be regarded as a dictionary query task, that is, training an encoder to perform the dictionary query task. Suppose there is already an encoded query q (a feature), and a series of encoded samples k0, k1, k2... Then k0, k1, k2... can be regarded as the keywords in the dictionary. Suppose there is only one keyword in the dictionary, that is, k+ is matched with q, then k+ and q are positive sample pairs with each other, and the remaining keywords are negative samples of q.

[0105] Once the positive and negative sample pairs are defined, a contrastive learning loss function is needed to guide the model to learn. This loss function needs to meet these requirements, that is, when the query q is similar to the only positive sample and dissimilar to all other negative sample keywords, the value of this loss should be relatively low. Conversely, if q is not similar to the positive sample, or q is similar to the keywords of other negative samples, then the loss should be large, so as to punish the model and prompt the model to update the parameters.

[0106] In one embodiment, the loss function of the feature extraction model is:

[0107]

[0108] In the formula, z · represents the feature vector output by the backbone network of the feature extraction model when the input sample is x · ; P(i) is the set of samples of the same class as the sample x i , and the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the set of samples in the first sample group; sim(·) is the similarity metric function; represents the temperature coefficient, which is used to adjust the distribution of similarity scores and is a hyperparameter;

[0109] Samples of the same class form positive sample pairs, and samples of different classes form negative sample pairs; in contrastive learning, by calculating the similarity between the feature vector of the input sample x i and the feature vectors of other samples, the sample corresponding to the maximum similarity with the feature vector of the input sample x i is found as the positive sample of the input sample x i , and the cosine similarity function is used for the similarity measurement between samples, as follows:

[0110]

[0111] S308. Select a preset number of first sample groups from the original data set, input the samples in the first sample group into the feature extraction model, and obtain the first feature vector of each sample.

[0112] It should be noted that each time of training, a preset number of first sample groups can be extracted from the original dataset for training the feature extraction model. Among them, the preset number can be set as needed. For example, 300 samples can be selected from the original dataset as the first sample group, or 500 samples can be selected from the original dataset as the first sample group. After inputting the samples in the first sample group into the feature extraction model, feature vectors with the same number as the samples in the first sample group can be generated.

[0113] S309. Classify each first feature vector output by the feature extraction model by using a clustering or support vector machine method.

[0114] Common classification methods can be used, such as clustering or support vector machine and other methods, to classify the feature vectors corresponding to the first sample group to obtain a classification result.

[0115] S310. Calculate the classification accuracy scores of various state samples in the first sample group according to the classification result of the first feature vector corresponding to the first sample group.

[0116] The classification result can be compared with the label of the sample corresponding to each feature vector, and the classification accuracy of the classification result can be obtained. In this embodiment, the calculation method of the classification accuracy can use the F-score as the classification accuracy, and the F-scores corresponding to various samples are determined respectively.

[0117] In one embodiment, the solution formula of the F-score is:

[0118]

[0119] where TP i is the true positive of the i-th class of samples, representing the number of samples correctly predicted as the positive class; FP i is the false positive of the i-th class of samples, representing the number of samples wrongly predicted as the positive class; FN i is the false negative of the i-th class of samples, representing the number of samples wrongly predicted as the negative class. This formula only adds a term of -(1 - F i ) compared with the formula in the pre-training process, where F i is the F-value of the classification of the i-th class of samples, used to measure the classification effect of this class of samples, ranging from 0 to 1, and the larger the better.

[0120] S311. Based on the classification accuracy scores of various samples in the first sample group, adjust the contributions of various samples to the loss function, so that samples with low accuracy contribute more, and samples with high accuracy contribute less, and retrain the feature extraction model through dynamic contrast loss to adjust the parameters of the feature extraction model.

[0121] The final feature extraction model is obtained through improved Info-NCE contrastive learning based on the same label samples in the first sample group of the same batch, where Info-NCE contrastive learning is used to narrow the distance between positive samples and push the distance between negative samples, thereby learning the effective representation of data.

[0122] In one embodiment, the loss function of the feature extraction model is updated based on the state type of the input feature extraction model sample. In this embodiment, the classic contrast loss function is improved, and the improved dynamic contrast loss function is:

[0123]

[0124] In the formula, z · The backbone network of the feature extraction model is represented by the input sample x · The feature vector output when P(i) is the sample x i A set of samples of the same type, the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the sample set of the first sample group; sim(·) is the similarity metric function; τ represents the temperature coefficient, which is used to adjust the distribution of the similarity score and is a hyperparameter; F i represents the classification accuracy of the i-th sample corresponding to the first sample group, F i Dynamically adjust according to the state type of the sample input to the feature extraction model.

[0125] For example, when the current input sample state is a healthy sample, the loss function can be adjusted according to the classification accuracy of the healthy sample, that is, F in the loss function i The value is adjusted to the classification accuracy of the healthy state samples in the first sample group of this batch. Similarly, for samples in other states, the F in the loss function needs to be adjusted. i The value is adjusted to the classification accuracy of the corresponding state samples in this batch.

[0126] S112. Repeat steps S310-S311 until the parameters of the feature extraction model converge, thereby obtaining the trained feature extraction model.

[0127] Repeat steps S310-S311, that is, repeatedly select a preset number of first sample groups from the original data set, and each time use the selected first sample to train the feature extraction model. It should be noted that in this embodiment, a sample selection method with replacement or a sample selection method without replacement can be used. Repeat the training multiple times until the parameters of the feature extraction model converge, and then stop the training to obtain a trained feature extraction model.

[0128] S313. Construct a linear classification model. The input of the linear classifier is a feature vector, and the output is the probability that the track belongs to various states.

[0129] S314. Jointly construct an orbit state diagnosis model by the trained feature extraction model and the linear classifier; the input of the orbit state diagnosis model is the passing train signal, and the output is the probability that the orbit belongs to various states.

[0130] S315. Pre-train the orbit state diagnosis model using a balanced data set and a classical contrast loss function, and re-train the orbit state diagnosis model using the original data set and a dynamic contrast loss function to obtain the trained orbit state diagnosis model.

[0131] It should be noted that the orbit health diagnosis model is specifically the orbit health diagnosis model as Figure 2 shown. The classifier used in Figure 2 is a linear classifier, which consists of a single-layer or double-layer fully connected layer. The input dimension of the linear classifier is the same as the dimension of the feature vector output by the feature extraction model, and the output dimension of the linear classifier is the same as the number of orbit state categories. Combine the feature extraction model and the linear classification model to form an orbit health diagnosis model. This model takes the passing train signal in the original data set as the input and directly outputs the probability that the orbit is in each state. After training the feature extraction model in step S312, fix the parameters of the feature extraction model, and train the linear classification model based on the balanced training data and the cross-entropy loss function to obtain the final orbit health diagnosis model.

[0132] In one embodiment, the cross-entropy loss function of the orbit health diagnosis model is:

[0133] h(·) = f( g (·))

[0134] In the formula, g is the feature extraction model, f is the classifier, and h is the orbit health diagnosis model.

[0135] Furthermore, after pre-training the orbit state diagnosis model, a preset number of second sample groups can be selected from the original data set. Input the samples in the second sample groups into the orbit state diagnosis model to obtain the feature vector of each sample, and classify each feature vector using a preset classification method to obtain the classification accuracy scores of the samples of various states in the second sample group; adjust the contribution of each sample to the loss function based on the classification accuracy scores of various samples in the second sample group, so that the samples with low accuracy contribute more, and the samples with high accuracy contribute less. Re-train the orbit health diagnosis model through the dynamic contrast loss to adjust the parameters of the orbit health diagnosis model; repeat the above steps until the parameters of the orbit health diagnosis model converge to obtain the trained orbit health diagnosis model.

[0136] The present invention uses data augmentation technology to transform the original monitoring data, enhancing the sample quantity and diversity of the input data; secondly, constructs a multi-branch deep neural network model with an attention mechanism to extract high-order features of track monitoring data under different train speeds; then, constructs a dynamic contrast loss function to train the model, which adjusts the contributions of different category samples in real time during the training process according to the classification accuracy of different category samples, solving the problem of sample imbalance; constructs a balanced dataset through downsampling and trains a classifier to classify the high-order features of the monitoring data, realizing data-driven track health diagnosis; finally, by separating the feature extraction model from the classification model, for new disease types, only the classification model needs to be trained, with strong adaptability.

[0137] Embodiment III

[0138] Figure 4 It is a schematic structural diagram of a track health diagnosis device provided in Embodiment III of the present invention. As Figure 4 shown, the device includes:

[0139] A data acquisition unit 401, configured to acquire passing train data when a train passes through a track, and label the passing train data based on the track state to construct an original dataset;

[0140] A pre-training unit 402, configured to perform downsampling on the original dataset to construct a balanced dataset, input samples in the balanced dataset into a feature extraction model, output corresponding feature vectors, and pre-train the feature extraction model through a classical contrast loss function; the number of samples of each type of label in the balanced dataset is the same; the feature extraction model is used to extract features of input samples and output feature vectors;

[0141] A feature extraction unit 403, configured to select a preset number of first sample groups from the original dataset, input samples in the first sample groups into the feature extraction model, obtain first feature vectors of each sample, classify each first feature vector by using a preset classification method, and obtain classification accuracy scores of samples of each type of state in the first sample groups;

[0142] A first training unit 404, configured to adjust the contributions of different category samples to the loss function based on the classification accuracy scores of different category samples in the first sample groups, so that samples with low accuracy contribute more and samples with high accuracy contribute less, and re-train the feature extraction model through a dynamic contrast loss to adjust parameters of the feature extraction model;

[0143] A second training unit 405, configured to repeat the steps of the feature extraction unit and the training unit until the parameters of the feature extraction model converge, obtaining the trained feature extraction model.

[0144] The track health diagnosis device provided by the embodiment of the present invention can execute the track health diagnosis method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0145] Embodiment 4

[0146] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0147] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0148] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0149] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a track health diagnosis method.

[0150] In some embodiments, a track health diagnosis method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the track health diagnosis method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute a track health diagnosis method by any other suitable means (e.g., by means of firmware).

[0151] The various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0154] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0155] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0156] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0157] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0158] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A rail health diagnosis method, characterized in that: include: Obtaining train passing data when a train passes through a track, and annotating the train passing data based on the track state to construct an original data set; Downsampling the original data set to construct a balanced data set, inputting samples in the balanced data set into a feature extraction model, outputting corresponding feature vectors, and pre-training the feature extraction model through a classical contrast loss function; the number of samples of each type of label in the balanced data set is the same; the feature extraction model is used to extract features of the input samples and output feature vectors; S1: Selecting a preset number of first sample groups from the original data set, inputting samples in the first sample group into the feature extraction model, obtaining a first feature vector of each sample, classifying each first feature vector using a preset classification method, and obtaining classification accuracy scores of samples of each state in the first sample group; S2: adjusting the contribution of each type of sample to the loss function based on the classification accuracy scores of each type of sample in the first sample group, so that samples with low accuracy contribute more and samples with high accuracy contribute less, retraining the feature extraction model through dynamic contrast loss, and adjusting the parameters of the feature extraction model; Repeat steps S1-S2 until the parameters of the feature extraction model converge, thereby obtaining the trained feature extraction model.

2. The rail health diagnosis method according to claim 1, characterized in that: The obtaining of the train passing data when the train passes the track and labeling the train passing data based on the track state to construct an original data set includes: Obtain monitoring data when trains pass through healthy tracks, as well as monitoring data when trains pass through tracks with various types of disease conditions; The vehicle passing data in the monitoring data is intercepted through a signal window of equal width, and the vehicle passing data is marked based on the track state, wherein the marked state includes a healthy state and various disease states; The vehicle passing signal is preprocessed, wherein the preprocessing includes one or more of downsampling, filtering and standardization to construct an original data set without abnormal data.

3. The rail health diagnosis method according to claim 1, characterized in that: A balanced data set is constructed based on downsampling of the original data, samples in the balanced data set are input into a feature extraction model to output corresponding feature vectors, and the feature extraction model is pre-trained using a classic contrast loss function, which includes: Data enhancement is performed on the data samples in the original data set; the data enhancement includes one or more of adding noise, random scaling, and random translation.

4. The rail health diagnosis method according to claim 1, characterized in that: After repeating steps S1-S2 until the parameters of the feature extraction model converge, the trained feature extraction model is obtained, the method further includes: Constructing a linear classification model, wherein the input of the linear classifier is a feature vector, and the output is the probability of the track belonging to each state; The track state diagnosis model is constructed by the trained feature extraction model and the linear classifier; the input of the track state diagnosis model is the vehicle passing signal, and the output is the probability of the track belonging to each state; The track state diagnosis model is pre-trained using a balanced data set and a classical contrast loss function, and the track state diagnosis model is re-trained using an original data set and a dynamic contrast damage function to obtain a trained track state diagnosis model.

5. The rail health diagnosis method according to claim 1, characterized in that: The feature extraction model consists of a multi-branch convolutional neural network model and a self-attention mechanism; The multi-branch convolutional neural network uses convolution kernels of different sizes to extract information of different scales of input samples; the feature dimensions of the information of different scales output by the multi-branch convolutional neural network are the same, and after the information of different scales is spliced, feature extraction is performed through the multi-head self-attention mechanism, and multi-scale features are fused; The fused features are transformed nonlinearly by a multi-layer neural network to output a one-dimensional feature sequence, which is the feature vector of the input sample. The length of the one-dimensional feature sequence is less than the length of the input data.

6. The rail health diagnosis method according to claim 1, characterized in that: The balanced data set is constructed, samples in the balanced data set are input into the feature extraction model to output the corresponding feature vector, and the feature extraction model is pre-trained by the classic contrast loss function, and the classic contrast loss function is: In the formula, z · The backbone network of the feature extraction model is represented by the input sample x · The feature vector output when P(i) is the sample x i A set of samples of the same type, the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the sample set of the first sample group; sim(·) is the similarity metric function; τ represents the temperature coefficient, which is used to adjust the distribution of the similarity score and is a hyperparameter; The same type of samples constitute positive sample pairs, and the different types of samples constitute negative sample pairs. The cosine similarity function is used to measure the similarity between samples, as follows:

7. The rail health diagnosis method according to claim 1, characterized in that: Each first feature vector is classified using a preset classification method to obtain classification accuracy scores of each type of state samples in the first sample group, including: Using a clustering or support vector machine method to classify each first feature vector output by the feature extraction model; According to the classification result of the first feature vector corresponding to the first sample group, the classification accuracy scores of each type of state samples in the first sample group are calculated.

8. The rail health diagnosis method according to claim 7, characterized in that: The calculation formula of the classification accuracy is: Among them, TP i is the true example of the i-th class sample, indicating the number of samples correctly predicted as positive; FP i is the false positive example of the i-th class sample, indicating the number of samples incorrectly predicted as positive; FN i is the false negative example of the i-th class sample, indicating the number of samples incorrectly predicted as negative.

9. The rail health diagnosis method according to claim 1, characterized in that: Based on the classification accuracy scores of each type of samples in the first sample group 9., the contribution of each type of samples to the loss function is adjusted. The dynamic contrast loss function is: In the formula, z · The backbone network of the feature extraction model is represented by the input sample x · The feature vector output when P(i) is the sample x i A set of samples of the same type, the number of samples in the set is |P(i)|; I is the number of samples in the balanced sample set; A(i) is the sample set of the first sample group; sim(·) is the similarity metric function; τ represents the temperature coefficient, which is used to adjust the distribution of the similarity score and is a hyperparameter; F i Indicates the classification accuracy of the i-th category samples corresponding to the first sample group.

10. A track health diagnostic device, characterized in that: include: A data acquisition unit, used to acquire the passing data of a train when it passes through the track, and to mark the passing data based on the track state to construct an original data set; A pre-training unit is used to downsample the original data set to construct a balanced data set, input samples in the balanced data set into a feature extraction model, output corresponding feature vectors, and pre-train the feature extraction model through a classical contrast loss function; the number of samples of each type of label in the balanced data set is the same; the feature extraction model is used to extract features of the input samples and output feature vectors; a feature extraction unit, configured to select a preset number of first sample groups from the original data set, input samples in the first sample group into the feature extraction model, obtain a first feature vector of each sample, classify each first feature vector using a preset classification method, and obtain classification accuracy scores of samples of each state in the first sample group; A first training unit is used to adjust the contribution of each type of sample to the loss function based on the classification accuracy score of each type of sample in the first sample group, so that the sample with low accuracy contributes more and the sample with high accuracy contributes less, and the feature extraction model is trained again by dynamic contrast loss to adjust the parameters of the feature extraction model; The second training unit is used to repeat the steps of the feature extraction unit and the training unit until the parameters of the feature extraction model converge, thereby obtaining the trained feature extraction model.