Rail locomotive operation and maintenance data intelligent monitoring method and system and electronic equipment
By collecting key operating parameters on track locomotives, training intelligent annotation models using partial annotation and adversarial networks, generating decision-making suggestions, the shortcomings of manual inspections are solved, and the intelligence and accuracy of track locomotive operation and maintenance data are improved.
Patent Information
- Application Number
- CN202510536143.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-11
AI Technical Summary
In the operation and maintenance of rail locomotives, manual inspections are difficult to monitor the locomotive status in real time, resulting in missed and missed inspections. The cost of obtaining labeled data is high and timely, making it difficult to detect faults quickly and accurately.
The data acquisition module is used to obtain key operating parameters, train the intelligent annotation model through partial annotation and adversarial network, and generate decision suggestions using the decision tree algorithm to achieve rapid annotation and intelligent decision-making of abnormal data.
It reduces the workload of manual labeling, improves data utilization and model accuracy, can timely identify potential faults, and improves the intelligence and reliability of track locomotive operation and maintenance data monitoring.
Smart Images

Figure CN120296430A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent monitoring scheme design for the operation and maintenance data of rail locomotives, and particularly relates to an intelligent monitoring method and system for the operation and maintenance data of rail locomotives, and an electronic device. Background Art
[0002] In the operation and maintenance management of rail locomotives, it is crucial to monitor key operation parameters in real time and accurately to ensure the safe and stable operation of locomotives. Traditional operation and maintenance monitoring mainly rely on regular inspections and manual inspections, and this method has certain limitations. On the one hand, it is difficult for manual inspections to grasp the operation status of locomotives in real time, and it is easy to miss inspections and misjudge inspections; on the other hand, the inspection cycle is relatively long, and sudden faults and abnormal situations cannot be detected in time, which may lead to the occurrence of safety accidents.
[0003] With the development of sensor technology and data processing technology, by installing various types of sensors on rail locomotives, key operation parameter data of locomotives can be collected in real time. However, in the face of a large amount of data, how to extract valuable information from it, quickly and accurately detect faults and make intelligent decisions has become an urgent problem to be solved.
[0004] Currently, although there are already some methods based on machine learning and deep learning for the analysis of rail locomotive operation and maintenance data, in the process of model training, the acquisition of labeled data is still a difficult problem, and there are the following technical problems: High labeling cost and easy to make mistakes: Artificial intelligence technologies such as deep learning require a large amount of labeled data for training. In the field of rail locomotive operation and maintenance, data labeling requires professional knowledge and experience. The labeling process is cumbersome and easy to make mistakes, resulting in a relatively high labeling cost.
[0005] Timeliness of labeled data: The technology and operation status of rail locomotives are constantly changing, new fault modes and operation scenarios are constantly emerging, and the existing labeled data may not meet the model training requirements in new situations, and it is necessary to update and supplement the labeled data in a timely manner.
[0006] Therefore, the existing technology still needs to be further developed. Summary of the Invention
[0007] The purpose of the present invention is to overcome the above technical deficiencies, and provide an intelligent monitoring method and system for the operation and maintenance data of rail locomotives, and an electronic device to solve the problems existing in the prior art.
[0008] To achieve the above technical purpose, according to the first aspect of the present invention, the present invention provides an intelligent monitoring method for the operation and maintenance data of rail locomotives, including: S100. Use the data acquisition module installed on the rail locomotive to obtain the key operating parameter data of the rail locomotive, and preprocess the collected key operating parameter data according to a preset method; S200. Perform partial annotation on the collected key operating parameter data, annotate the reasons for the generation of abnormal data, record the annotated data set as the annotation set, record the unannotated data as the unannotated set, mix the annotation set and the unannotated data according to a preset ratio to form a training data set, and then train an intelligent annotation model; S300. When it is detected that the key operating parameters of the rail locomotive exceed the preset threshold, immediately trigger an emergency annotation process, use the intelligent annotation model to preferentially annotate the newly emerged abnormal data, and output a prompt signal regarding the annotation result.
[0009] Specifically, before the intelligent annotation model is trained, partial annotation is performed on the collected key operating parameter data by relevant management personnel. If the relevant management personnel determine that the operating state of the locomotive is normal based on the collected key operating parameter data, the corresponding label for the relevant data is normal; if it is determined that the operating state of the locomotive is abnormal, the corresponding label for the relevant data is annotated as the reason for the generation of abnormal data.
[0010] Specifically, the method further includes: Have relevant management personnel confirm whether the result annotated by the model is correct. If it is incorrect, have the relevant management personnel modify it to the correct label, and re-enter the modified key operating parameters and the corresponding correct labels into the intelligent annotation model for training.
[0011] Specifically, the data acquisition module includes at least one of the following sensors: A temperature sensor installed around the engine and / or transformer, and the temperature sensor is used to detect the component temperature signal of the engine and / or transformer; A pressure sensor installed in the hydraulic system and / or the brake pipeline, and the pressure sensor is used to detect the pressure state signal in the hydraulic system and / or the brake pipeline; A vibration sensor installed at the locomotive body and / or the axle position, and the vibration sensor is used to detect the vibration signal at the locomotive body and / or the axle position; A Hall current sensor installed on the motor power supply line, and the Hall current sensor is used to detect the current signal of the motor power supply line.
[0012] Specifically, the method further includes: The training of the intelligent annotation model is carried out by using a generative adversarial network. The generative adversarial network includes a generator and a discriminator. The input of the generator is a random noise vector, which is used to generate a fake data distribution similar to the real data through a multi-layer transposed convolutional layer. The discriminator is used to distinguish whether the input data comes from the real data set or the fake data generated by the generator, that is, to calculate the confidence of the fake data generated by the generator. When the probability that the fake data generated by the generator is judged as real data by the discriminator is greater than the preset confidence threshold, the data is regarded as pseudo-annotated data and added to the annotation set for subsequent model training.
[0013] Specifically, the method further includes: After the intelligent annotation model outputs the annotation result, an intelligent decision-making suggestion is generated according to the decision-making algorithm.
[0014] Specifically, after the intelligent annotation model outputs the annotation result, generating an intelligent decision-making suggestion according to the decision-making algorithm includes: Extract key features from the annotated data. The key features include the temperature change trend, the pressure fluctuation amplitude, the vibration frequency, and the cause of the fault. The key features also include the maintenance history data, which corresponds one-to-one with the cause of the fault. The maintenance history data is input by the user into the system database, and the maintenance history data corresponding to the cause of the fault is obtained from the system database. Use the decision tree algorithm to construct a decision tree model, with the extracted features as the input nodes, the cause of the fault and the maintenance history data as the intermediate nodes, and the decision-making suggestion as the leaf node. Through the training of the annotated data, learn the relationship between the features and the decision result. Use the leave-one-out method to quantitatively evaluate the effectiveness of the decision tree model. Each time, select 1 sample as the test set, and the remaining samples as the training set for training. Calculate the prediction accuracy rate each time, and finally take the average value as the evaluation result.
[0015] Specifically, the decision-making suggestions include parking for maintenance and replacing parts. The causes of the faults include engine overheating and cooling system failure that may be caused by cooling system failure.
[0016] According to the second aspect of the present invention, there is provided an intelligent monitoring system for the operation and maintenance data of a rail locomotive, including: A data acquisition module, which is arranged on the rail locomotive and is used to acquire the key operation parameter data of the rail locomotive. A control module is configured to preprocess the collected key operating parameter data according to a preset method; to partially label the collected key operating parameter data, label the causes of abnormal data, record the labeled data set as the labeled set, record the unlabeled data as the unlabeled set, mix the labeled set and the unlabeled data according to a preset ratio to form a training data set, and then train an intelligent labeling model; and to immediately trigger an emergency labeling process when it detects that the key operating parameters of the rail locomotive exceed the preset threshold, and use the intelligent labeling model to preferentially label the newly emerged abnormal data and output a prompt signal regarding the labeling result.
[0017] According to a third aspect of the present invention, there is provided an electronic device, including: a memory; and a processor, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the above-mentioned intelligent monitoring method for rail locomotive operation and maintenance data is implemented.
[0018] Beneficial effects: In the present invention, by partially labeling the collected key operating parameter data, labeling the causes of abnormal data, recording the labeled data set as the labeled set, recording the unlabeled data as the unlabeled set, mixing the labeled set and the unlabeled data according to a preset ratio to form a training data set, and then training an intelligent labeling model, when it detects that the key operating parameters of the rail locomotive exceed the preset threshold, it immediately triggers an emergency labeling process, and uses the intelligent labeling model to preferentially label the newly emerged abnormal data, which greatly reduces the workload of manual labeling, improves the utilization rate of data, enables the model to learn richer data features, and thus more accurately identifies various operating states, including potential fault hazards, and further greatly improves the intelligence, reliability, and accuracy of rail locomotive operation and maintenance data monitoring. Description of the Drawings
[0019] Figure 1 is a schematic flowchart of the intelligent monitoring method for rail locomotive operation and maintenance data provided in a specific embodiment of the present invention; Figure 2 is a schematic diagram of the system composition of the intelligent monitoring system for rail locomotive operation and maintenance data provided in a specific embodiment of the present invention. Detailed Embodiments
[0020] To enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Based on the embodiments in this application, other similar embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application. In addition, the directional terms mentioned in the following embodiments, in the preferred embodiments of the present invention, "up", "down", "left", "right", etc. are only in reference to the directions of the accompanying drawings. Therefore, the directional terms used are for illustration rather than to limit the present invention.
[0021] The present invention will be further described below in conjunction with the accompanying drawings and preferred embodiments.
[0022] Please refer to Figure 1 , the present invention provides an intelligent monitoring method for operation and maintenance data of rail locomotives, including: S100. Utilize the data acquisition module provided on the rail locomotive to obtain the key operation parameter data of the rail locomotive, and preprocess the collected key operation parameter data according to a preset method.
[0023] Specifically, the data acquisition module includes at least one of the following sensors: A temperature sensor disposed around the engine and / or transformer, and the temperature sensor is used to detect the component temperature signal of the engine and / or transformer.
[0024] A pressure sensor disposed in the hydraulic system and / or the brake pipeline, and the pressure sensor is used to detect the pressure state signal in the hydraulic system and / or the brake pipeline; A vibration sensor disposed at the locomotive body and / or the axle position, and the vibration sensor is used to detect the vibration signal at the locomotive body and / or the axle position; A Hall current sensor disposed on the motor power supply line, and the Hall current sensor is used to detect the current signal of the motor power supply line.
[0025] Specifically, the preprocessing of the collected key operation parameter data according to a preset method includes: Data cleaning: First, check whether there are missing values in the collected data. For a small number of missing values, the mean filling method is used for supplementation; for a large number of missing or data abnormal records, they are directly deleted to ensure the integrity of the data. In the preferred embodiment of the present invention, when the temperature data in a certain time period is missing, calculate the mean value of the temperature data before and after this time period, and fill the missing value with the mean value.
[0026] Denoising process: For the possible noise in the data collected by the sensor, a wavelet transform denoising algorithm is adopted. Specifically, the data is decomposed into sub-bands of different frequencies through wavelet transform. According to the characteristics of the noise in each sub-band, a threshold is set, and the coefficients below the threshold are set to zero. Then, the data is reconstructed through inverse wavelet transform to achieve the purpose of denoising. In the preferred embodiment of the present invention, the threshold is set to zero for the wavelet coefficients whose absolute value is less than 10 in a group of wavelet coefficients.
[0027] Format unification: The data from different sources and in different formats is uniformly converted into the standard JSON format to facilitate subsequent data processing and analysis. In the preferred embodiment of the present invention, the original binary data collected by the temperature sensor is converted into numerical data according to the specified coding rules, and is assembled together with other sensor data according to the predefined JSON structure.
[0028] Specifically, for the data from different sources of the rail locomotive, according to its data characteristics and application scenarios, a unique identifier is assigned to each data type, and this identifier is recorded during the annotation process, so that the source of the data can be accurately identified and traced during subsequent data processing and analysis.
[0029] The specific operations are as follows: Assign the identifier "TEMP" to the temperature sensor data, add the field "data_type" in the annotated data and record it as "TEMP", and at the same time record information such as the sensor installation location and acquisition time; The identifier of the pressure sensor data is "PRESS", and relevant fields are also added to record information; The identifier of the vibration sensor data is set to "VIB", and information such as the vibration direction and sampling frequency is recorded; The identifier of the current sensor data is "CUR", and information such as the circuit where the current sensor is located and the acquisition time is recorded.
[0030] S200. Partially annotate the collected key operating parameter data, annotate the reasons for the generation of abnormal data, record the annotated data set as the annotation set, record the unannotated data as the unannotated set, mix the annotation set and the unannotated data according to a preset ratio to form a training data set, and then train the intelligent annotation model.
[0031] Specifically, before the training of the intelligent annotation model is completed, part of the collected key operating parameter data is annotated by relevant management personnel. If the relevant management personnel determine that the operating state of the locomotive is normal according to the collected key operating parameter data, the corresponding label of the relevant data is normal; if it is determined that the operating state of the locomotive is abnormal, the corresponding label of the relevant data is annotated as the reason for the generation of abnormal data.
[0032] In a preferred embodiment of the present invention, when the temperature sensor detects that the engine temperature is too high, the management personnel judge, based on experience and relevant knowledge, that it is caused by a failure of the cooling system. At this time, the corresponding data is labeled as "cooling system failure".
[0033] Furthermore, the labeled data set is denoted as the labeled set, and the unlabeled data is denoted as the unlabeled set. The labeled set and the unlabeled data are mixed according to a preset ratio to form a training data set. In a preferred embodiment of the present invention, the ratio of the labeled set to the unlabeled set is set to 1:10. Such a ratio setting can ensure that the model learns enough labeled information while making full use of a large amount of unlabeled data to improve the generalization ability of the model.
[0034] Specifically, the method further includes: The relevant management personnel confirm whether the result labeled by the model is correct. If it is incorrect, the relevant management personnel modify it to the correct label, and re-enter the modified key operating parameters and the corresponding correct label into the intelligent labeling model for training. When the management personnel find that the model mislabels the data in a certain normal operating state as "cooling system failure", they correct it to "normal" and re-enter it into the model for training to improve the accuracy of the model.
[0035] Specifically, the method further includes: The intelligent labeling model is trained using an adversarial network. The adversarial network includes a generator and a discriminator. The input of the generator is a random noise vector, which is used to generate a fake data distribution similar to the real data through multiple transposed convolutional layers. The discriminator is used to distinguish whether the input data comes from the real data set or the fake data generated by the generator, that is, to calculate the confidence of the fake data generated by the generator. When the probability that the fake data generated by the generator is judged as real data by the discriminator is greater than the preset confidence threshold, the data is regarded as pseudo-labeled data and added to the labeled set for subsequent model training.
[0036] Furthermore, the method for training the intelligent labeling model using an adversarial network includes: Adversarial network structure: The intelligent labeling model is trained using an adversarial network. The adversarial network includes a generator and a discriminator.
[0037] Generator: The input of the generator is a random noise vector, which is used to generate a fake data distribution similar to the real data through multiple transposed convolutional layers. Specifically, the generator can adopt the generator structure in the Deep Convolutional Generative Adversarial Network (DCGAN). The input is a random noise vector with a dimension of 100, and through multiple transposed convolutional layers, it gradually generates a fake data distribution similar to the real data. Discriminator: The discriminator is used to distinguish whether the input data comes from the real data set or the fake data generated by the generator, that is, to calculate the confidence of the fake data generated by the generator. The discriminator can adopt the discriminator structure in DCGAN. By passing through multiple convolutional layers and fully connected layers, it judges the authenticity of the input data, and the output is a probability value between 0 and 1.
[0038] Training process: Define the loss function: The goal of the generator is to minimize the probability that the discriminator judges it as fake data, that is, to maximize the probability that the discriminator thinks it is real data; the goal of the discriminator is to correctly distinguish between real data and fake data generated by the generator, that is, to minimize the probability of misjudgment. The loss function adopts the cross-entropy loss function, which is specifically expressed as: The generator loss function is as follows: Discriminator loss function: Among them, represents the generator, represents the discriminator, represents the noise vector 's distribution, represents the distribution of real data.
[0039] Initialize the model parameters: The weights of the generator and the discriminator are initialized by using the random initialization method.
[0040] Iterative training: Mix the labeled data and unlabeled data, and form training batches according to a certain ratio (in the preferred embodiment of the present invention, the ratio of labeled data to unlabeled data is 1:10). In each iteration, first input the real data into the discriminator to calculate the discriminator loss , and then update the discriminator parameters through an optimization algorithm (preferably the Adam optimizer in the present invention, with the learning rate set to 0.0002 and the momentum set to 0.5); then input the random noise vector into the generator to generate fake data. The fake data passes through the discriminator to obtain the output result, and calculate the generator loss , and also update the generator parameters through the optimization algorithm. The iterative training process continues until the model converges (in the preferred embodiment of the present invention, the change in the loss function is less than 1e-6 or reaches the preset maximum number of training rounds of 10,000 rounds).
[0041] It is understandable that the selection of a 1:10 ratio is a trade-off between the accuracy of the labeled data and the utilization efficiency of the model for unlabeled data. If the proportion of labeled data is too high, although the model can learn more accurate labeled information, it may be limited by the finiteness of the labeled data and unable to fully explore the potential information in the unlabeled data; conversely, if the proportion of labeled data is too low, the model may not be able to learn enough prior knowledge, resulting in performance degradation.
[0042] It is understandable that this learning rate of 0.0002 is determined through a large number of experiments and tuning processes. Generally speaking, the selection of the initial learning rate needs to be adjusted according to the specific task, dataset, and model structure. For some complex deep learning models, a smaller learning rate is usually more conducive to the convergence of the model.
[0043] It is understandable that the momentum value is usually set between 0.5 and 0.9. 0.5 is a relatively common default value, which can achieve good results in most cases. If the momentum value is too large, the model may overly rely on previous gradient information, resulting in a slower convergence speed; if the momentum value is too small, the convergence speed of the model may not be fast enough.
[0044] It is understandable that during the training process, as the number of iterations increases, the value of the loss function will gradually decrease. When the change in the loss function is very small, it indicates that the model parameters are close to the optimal solution, and further iteration may not greatly improve the model performance. 1e-6 is a relatively small threshold, which means that the change in the loss function is small enough, and it can be considered that the model has converged.
[0045] It is understandable that in order to prevent the model from falling into an infinite loop during training or being unable to converge due to certain abnormal situations, a maximum number of training epochs is usually set. 10,000 epochs is an empirical value, which can be adjusted according to the specific task and dataset. If the model has converged before reaching the maximum number of training epochs, the training process will stop early; if the model has not converged after reaching the maximum number of training epochs, it may be necessary to further adjust the model structure or hyperparameters.
[0046] Generation and screening of pseudo-labeled data: When the probability that the discriminator judges the fake data generated by the generator as real data is greater than the preset confidence threshold, this data is regarded as pseudo-labeled data and added to the labeled set for subsequent model training. The preset confidence threshold can be adjusted according to the actual situation and is set to 0.8 in the preferred embodiment of the present invention. Such a threshold setting can improve the training efficiency of the model while ensuring the quality of the pseudo-labeled data.
[0047] It is understandable that the preset confidence threshold is used to determine whether the fake data generated by the generator can be added to the annotation set as pseudo-annotated data. Selecting 0.8 as the threshold is a trade-off between the accuracy of the model and the annotation efficiency. If the threshold is set too high, in the preferred embodiment of the present invention, 0.9 or higher, then only when the model has a very high confidence in the generated data will it be regarded as pseudo-annotated data, which will result in a small amount of data that can be used as pseudo-annotated data, insufficient training data for the model, and affect the training effect and generalization ability of the model; if the threshold is set too low, in the preferred embodiment of the present invention, 0.6 or lower, some incorrect data with low confidence may be added to the annotation set, thus introducing noise and reducing the accuracy of the model. After a large number of experiments and analyses, a threshold of 0.8 can make full use of the fake data generated by the generator on the premise of ensuring the quality of the annotated data, and improve the training efficiency and performance of the model.
[0048] It should be further noted that in the present invention, the initial model trained is used to predict the unannotated data, and the data with high prediction confidence is selected as pseudo-annotated data to expand the annotation set. Specifically, the confidence threshold is set to 0.8, that is, when the probability that the fake data generated by the generator is judged as real data by the discriminator is greater than 0.8, this data is regarded as pseudo-annotated data and added to the annotation set for subsequent model training.
[0049] It should be further noted that the present invention continuously repeats the above process of semi-supervised learning model training and pseudo-annotated data generation, continuously optimizes the annotation scheme, and reduces the manual annotation workload and annotation error. Specifically, in each iteration, the generative adversarial network is retrained according to the new annotation set until the performance metrics (accuracy, recall rate) of the model meet certain requirements (preferably the accuracy is greater than 95% in the present invention) or reach the preset number of iterations (preferably 10 times in the present invention). The above preferred values are obtained by those skilled in the art of the present invention through a large number of tests and can well implement the method described in the present invention. The above parameters can be specifically set according to the actual needs of the users of the present invention, and the present invention does not limit them here.
[0050] S300. When it is detected that the key operating parameters of the rail locomotive exceed the preset threshold, an emergency annotation process is immediately triggered, and the intelligent annotation model is used to preferentially annotate the newly emerged abnormal data, and a prompt signal regarding the annotation result is output.
[0051] Specifically, in this step, when it is detected that the key operating parameters of the rail locomotive exceed the preset threshold, an emergency annotation process is immediately triggered, the newly emerged abnormal data is preferentially annotated, and the annotation result is timely fed back to the model training module to quickly adjust the model parameters. Specifically: Set the thresholds of key operating parameters. The engine temperature threshold is set to 120 °C, the pressure threshold is set to 35 MPa, the vibration acceleration threshold is set to 5 g, and the current threshold is set to 400 A.
[0052] When the real-time collected data exceeds the corresponding threshold, the annotation module is notified through the message queue (preferably RabbitMQ in the present invention) for emergency annotation. The annotator manually annotates the reasons for the abnormal data (including "engine overheating may be due to a failure of the cooling system") and the fault types (including "cooling failure") according to the real-time data and the locomotive operating status.
[0053] After the annotation is completed, the annotation results are sent to the model training module through network transmission (including based on the TCP / IP protocol). The model training module retrains the model according to the new annotation data and adjusts the model parameters, such as updating the neural network weights, to adapt to the new operating status.
[0054] It can be immediately noted that the engine temperature threshold, pressure threshold, vibration acceleration threshold, and current threshold can be specifically set according to the actual needs of the users of the present invention, as long as they are applicable to the solution proposed in the present invention, and the present invention does not make any restrictions here. In the present invention, the engine temperature threshold is set to 120 °C, the pressure threshold is set to 35 MPa, the vibration acceleration threshold is set to 5 g, and the current threshold is set to 400 A, which are obtained by the technical personnel of the present invention through a large number of tests and can well implement the method proposed in the present invention.
[0055] Specifically, the method further includes: After the intelligent annotation model outputs the annotation results, generate intelligent decision-making suggestions according to the use of decision algorithms.
[0056] Specifically, after the intelligent annotation model outputs the annotation results, generating intelligent decision-making suggestions according to the use of decision algorithms includes: Extract key features from the annotated data. The key features include temperature change trends, pressure fluctuation amplitudes, vibration frequencies, and fault reasons; the key features also include maintenance history data, and the maintenance history data corresponds one-to-one with the fault reasons. The maintenance history data is input by the user into the system database, and the maintenance history data corresponding to the fault reasons is obtained from the system database; Use the decision tree algorithm to construct a decision tree model, with the extracted features as input nodes, the fault reasons and maintenance history data as intermediate nodes, and the decision-making suggestions as leaf nodes. Through the training of the annotated data, learn the relationship between the features and the decision results; Use the leave-one-out method to quantitatively evaluate the effectiveness of the decision tree model. Each time, select 1 sample as the test set, and the remaining samples as the training set for training. Calculate the prediction accuracy rate each time, and finally take the average value as the evaluation result.
[0057] Specifically, the decision-making suggestions include parking for maintenance and replacing parts, and the reasons for the failures include engine overheating possibly caused by the failure of the cooling system and the failure of the cooling system.
[0058] In a preferred embodiment of the present invention, by analyzing the data collected by the temperature sensor, the temperature change trend is calculated; by processing the data of the pressure sensor, the pressure fluctuation amplitude is obtained; by analyzing the data of the vibration sensor, the vibration frequency is obtained; and the maintenance history data corresponding to the failure cause is queried from the system database.
[0059] Further, after the intelligent annotation model outputs the annotation result, the specific process of generating intelligent decision-making suggestions using the decision-making algorithm is as follows: Data preprocessing: Data cleaning: The annotated data is cleaned to remove noise data, duplicate data, and missing values. In a preferred embodiment of the present invention, for the temperature change trend data, it is checked whether there are outliers that significantly deviate from the normal range. If so, they are corrected or deleted according to the specific situation.
[0060] Data standardization: The key feature data is standardized to make it have the same scale and distribution. In a preferred embodiment of the present invention, for numerical features such as temperature change trend, pressure fluctuation amplitude, and vibration frequency, the Z-score standardization method can be used to convert them into standard normal distribution data with a mean of 0 and a standard deviation of 1 for subsequent model training.
[0061] Key feature extraction: Temperature change trend extraction: By analyzing the temperature data in the time series, statistical quantities such as the change rate and slope of the temperature are calculated to describe the temperature change trend. In a preferred embodiment of the present invention, a linear regression method can be used to fit the curve of the temperature changing with time, and the slope of the regression line is obtained as a quantitative index of the temperature change trend.
[0062] Pressure fluctuation amplitude extraction: The standard deviation or range of the pressure data within a certain time window is calculated to measure the pressure fluctuation amplitude. In a preferred embodiment of the present invention, the pressure data at 10 consecutive time points is selected, and the range (the maximum value minus the minimum value) of these 10 data is calculated as the pressure fluctuation amplitude within this time window.
[0063] Vibration frequency extraction: The vibration signal is subjected to spectral analysis to extract the main vibration frequency components and their corresponding amplitudes. The fast Fourier transform (FFT) algorithm can be used to convert the vibration signal in the time domain into a frequency domain signal, and then the frequency components with higher energy and their amplitudes in the frequency domain signal are statistically analyzed.
[0064] Fault cause extraction: Determine possible fault causes based on labeled data and domain knowledge. In a preferred embodiment of the present invention, for an engine system, the fault causes may include cooling system failures, fuel supply problems, and mechanical component wear. By analyzing and summarizing the labeled data, establish association rules between fault causes and key features.
[0065] Maintenance history data extraction: Obtain maintenance history data corresponding to the fault causes from the system database. When the user inputs maintenance history data, classify and store it in the database according to the fault causes. When needed, retrieve the corresponding maintenance history data from the database based on the currently determined fault causes.
[0066] Decision tree model construction and training: Decision tree algorithm selection: Select one of the common decision tree algorithms such as ID3, C4.5, or CART. Here, the CART algorithm is used as an example for illustration. The CART algorithm uses the Gini Index as the criterion for feature selection, can handle continuous and discrete features, and can construct a binary decision tree.
[0067] Feature selection and node splitting: Calculate the Gini Index: For each feature, calculate its corresponding Gini Index. The formula for the Gini Index is as follows: where D represents the dataset, K represents the number of classes, and p k represents the proportion of samples in the dataset belonging to the k-th class.
[0068] Select the optimal feature: For each feature, split it according to different values, and calculate the Gini Index of the subsets after splitting. Select the feature with the smallest Gini Index as the splitting feature for the current node. In a preferred embodiment of the present invention, for the feature of temperature change trend, divide it into different intervals, calculate the Gini Index corresponding to each interval, and select the interval splitting method with the smallest Gini Index.
[0069] Node splitting: According to the selected optimal feature and its splitting value, divide the dataset into different subsets and create corresponding child nodes. Recursively repeat the above feature selection and node splitting process for each child node until the stopping condition is met. The stopping condition can be reaching the maximum tree depth, the number of samples in the node being less than a certain threshold, or all samples belonging to the same class, etc.
[0070] Model training steps: Initialize the decision tree: Create a root node and use the labeled dataset as the sample set of the root node.
[0071] Recursively construct the decision tree: For the current node, if the stopping condition is met, mark the node as a leaf node and determine the decision recommendation (stop for maintenance or replace parts) of the leaf node according to the majority class of the samples in the node.
[0072] Otherwise, calculate the Gini index for each feature and select the feature with the minimum Gini index as the splitting feature for the current node.
[0073] According to the different values of the splitting feature, divide the sample set of the current node into multiple subsets and create a child node for each subset.
[0074] Recursively execute the above steps for each child node until the stopping condition is met.
[0075] Model parameter adjustment: During the training process, the parameters of the decision tree, such as the maximum tree depth, minimum sample split number, etc., can be adjusted by methods such as cross-validation to optimize the performance of the model. In the preferred embodiment of the present invention, 5-fold cross-validation is adopted, the data set is divided into 5 subsets, 4 of them are selected as the training set each time, 1 subset is used as the validation set, the average validation accuracy under different parameter settings is calculated, and the parameter combination with the highest accuracy is selected as the final model parameters.
[0076] Model evaluation: Initialize evaluation metrics: Set a variable to record the number of correctly predicted samples, with an initial value of 0; set a variable to record the total number of samples, with an initial value of the total number of samples in the data set.
[0077] Perform leave-one-out evaluation in a loop: For each sample i in the data set (i = 1, 2, ⋯, n, where n is the total number of samples in the data set): Use the remaining n - 1 samples except sample i as the training set, and use the training set to train the decision tree model to obtain the trained model Ti.
[0078] Use sample i as the test set, input its key features into the trained model Ti, and obtain the predicted decision recommendation yi of the model.
[0079] Obtain the true decision recommendation zi of sample i (obtained from the labeled data).
[0080] Compare whether the predicted decision recommendation yi and the true decision recommendation zi are the same. If they are the same, increment the number of correctly predicted samples by 1.
[0081] Calculate the average prediction accuracy: After the loop ends, calculate the average prediction accuracy Accuracy, and the calculation formula is as follows: Basis for the selection of relevant values and thresholds: Basis for the selection of decision tree parameters: Maximum tree depth: The maximum tree depth determines the complexity of the decision tree. If the maximum tree depth is set too small, the decision tree may not be able to fully learn the features of the data, resulting in underfitting; if it is set too large, the decision tree may be too complex and prone to overfitting. Generally, through methods such as cross-validation, different values are tried within a certain range, and the value that makes the model perform best on the validation set is selected. In the preferred embodiment of the present invention, in this solution, different values such as 3, 5, 7, 10, etc. can be tried for the maximum tree depth first, and then the optimal value is selected according to the results of cross-validation.
[0082] Minimum sample split number: The minimum sample split number represents the minimum number of samples required for splitting at a node. If the minimum sample split number is set too small, it may cause the decision tree to be overly subdivided, increasing the complexity of the model; if it is set too large, some nodes may not be able to be effectively split, affecting the performance of the model. Similarly, appropriate values can be determined through methods such as cross-validation.
[0083] Basis for the selection of the prediction accuracy threshold: The selection of the prediction accuracy threshold needs to be determined according to the specific application scenario and requirements. In this solution, for the decision-making suggestions output by the intelligent annotation model, a high accuracy needs to be ensured to ensure the reliability of the maintenance decision. Generally, the prediction accuracy threshold can be set above 0.8, that is, it is required that the prediction accuracy of the model reaches more than 80% for the model to be considered effective. If the prediction accuracy is lower than the threshold, the reasons need to be further analyzed, such as checking the data quality, adjusting the model parameters, etc., to improve the performance of the model.
[0084] It should be noted here that in the present invention, by partially annotating the key operating parameter data collected, the reasons for the abnormal data are annotated, the annotated data set is recorded as the annotation set, the unannotated data is recorded as the unannotated set, the annotation set and the unannotated data are mixed according to a preset ratio to form a training data set, and then the intelligent annotation model is trained. When it is detected that the key operating parameters of the rail locomotive exceed the preset threshold, the emergency annotation process is immediately triggered, and the intelligent annotation model is used to preferentially annotate the newly emerging abnormal data, greatly reducing the workload of manual annotation, improving the utilization rate of data, enabling the model to learn richer data features, thus more accurately identifying various operating states, including potential fault hazards, and then greatly improving the intelligence, reliability and accuracy of the rail locomotive operation and maintenance data monitoring.
[0085] Please refer to Figure 2, the present invention provides another embodiment, which provides an intelligent monitoring system for the operation and maintenance data of rail locomotives. The intelligent monitoring system for the operation and maintenance data of rail locomotives includes: A data acquisition module 100, which is arranged on the rail locomotive and is used to acquire the key operation parameter data of the rail locomotive; A control module 200, which is used to preprocess the acquired key operation parameter data according to a preset method; used to partially label the acquired key operation parameter data, label the reasons for the generation of abnormal data, record the labeled data set as the labeled set, record the unlabeled data as the unlabeled set, mix the labeled set and the unlabeled data according to a preset ratio to form a training data set, and then train an intelligent labeling model; used to immediately trigger an emergency labeling process when it is detected that the key operation parameters of the rail locomotive exceed the preset threshold, and use the intelligent labeling model to preferentially label the newly emerged abnormal data and output a prompt signal regarding the labeling result.
[0086] It can be understood that the present invention partially labels the acquired key operation parameter data, labels the reasons for the generation of abnormal data, records the labeled data set as the labeled set, records the unlabeled data as the unlabeled set, mixes the labeled set and the unlabeled data according to a preset ratio to form a training data set, and then trains an intelligent labeling model. When it is detected that the key operation parameters of the rail locomotive exceed the preset threshold, an emergency labeling process is immediately triggered, and the intelligent labeling model is used to preferentially label the newly emerged abnormal data, greatly reducing the workload of manual labeling, improving the utilization rate of data, enabling the model to learn richer data features, thus more accurately identifying various operating states, including potential fault hazards, and further greatly improving the intelligence, reliability and accuracy of the operation and maintenance data monitoring of rail locomotives.
[0087] In a preferred embodiment, the present application also provides an electronic device, which includes: A memory; and a processor, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, the intelligent monitoring method for operation and maintenance data of rail locomotives is implemented. This computer device can generally be a server, a terminal, or any other electronic device with necessary computing and / or processing capabilities. In one embodiment, this computer device may include a processor, a memory, a network interface, a communication interface, etc. connected through a system bus. The processor of this computer device can be used to provide necessary computing, processing, and / or control capabilities. The memory of this computer device may include a non-volatile storage medium and an internal memory. An operating system, a computer program, etc. may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of this computer device can be used to connect and communicate with external devices through a network. When the computer program is executed by the processor, the steps of the method of the present invention are executed.
[0088] The present invention can be implemented as a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method of the embodiments of the present invention are caused to be executed. In one embodiment, the computer program is distributed on a plurality of network-coupled computer devices or processors, so that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, can be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations can be executed by one or more computer devices or processors, and one or more other method steps / operations can be executed by one or more other computer devices or processors. One or more computer devices or processors can execute a single method step / operation, or execute two or more method steps / operations.
[0089] Those of ordinary skill in the art can understand that the method steps of the present invention can be completed by a computer program instructing relevant hardware such as a computer device or a processor, and the computer program can be stored in a non-transitory computer-readable storage medium, and when the computer program is executed, the steps of the present invention are caused to be executed. Depending on the situation, any reference to a memory, storage, database, or other medium in this article may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.
[0090] It should be noted here that in the present invention, by partially annotating the collected key operating parameter data, the reasons for the generation of abnormal data are annotated. The annotated data set is recorded as the annotation set, the unannotated data is recorded as the unannotated set, and the annotation set and the unannotated data are mixed according to a preset ratio to form a training data set, and then an intelligent annotation model is trained. When it is detected that the key operating parameters of the rail locomotive exceed the preset threshold, an emergency annotation process is immediately triggered, and the intelligent annotation model is used to preferentially annotate the newly emerged abnormal data, greatly reducing the workload of manual annotation, improving the utilization rate of data, enabling the model to learn richer data features, and thus more accurately identifying various operating states, including potential fault hazards, and further greatly improving the intelligent level, reliability and accuracy of the rail locomotive operation and maintenance data monitoring.
[0091] The above-described technical features can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification as long as such a combination does not exist in contradiction.
[0092] The specific embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. An intelligent monitoring method for operation and maintenance data of rail locomotives, characterized in that, The method includes: S100. Using the data acquisition module provided on the rail locomotive, obtain the key operation parameter data of the rail locomotive, and preprocess the collected key operation parameter data according to a preset method; S200. Perform partial annotation on the collected key operation parameter data, annotate the reasons for the abnormal data, record the annotated data set as the annotation set, record the unannotated data as the unannotated set, mix the annotation set and the unannotated data according to a preset ratio to form a training data set, and then train the intelligent annotation model; S300. When it is detected that the key operation parameters of the rail locomotive exceed the preset threshold, immediately trigger an emergency annotation process, use the intelligent annotation model to preferentially annotate the newly emerged abnormal data, and output a prompt signal regarding the annotation result.
2. The intelligent monitoring method for operation and maintenance data of rail locomotives according to claim 1, wherein, Before the intelligent annotation model is trained, perform partial annotation on the collected key operation parameter data by relevant management personnel. If the relevant management personnel determine that the operation state of the locomotive is normal based on the collected key operation parameter data, then label the corresponding data as normal. If it is determined that the operation state of the locomotive is abnormal, then label the corresponding data as the reason for the abnormal data.
3. The intelligent monitoring method for operation and maintenance data of rail locomotives according to claim 1, wherein, The method further includes: Have the relevant management personnel confirm whether the result annotated by the model is correct. If it is incorrect, have the relevant management personnel modify it to the correct label, and re-enter the modified key operation parameters and the corresponding correct labels into the intelligent annotation model for training.
4. The intelligent monitoring method for the operation and maintenance data of rail locomotives according to claim 1, characterized in that The data acquisition module includes at least one of the following sensors: A temperature sensor provided around the engine and / or the transformer, and the temperature sensor is used to detect the component temperature signal of the engine and / or the transformer; A pressure sensor provided in the hydraulic system and / or the brake pipeline, and the pressure sensor is used to detect the pressure state signal in the hydraulic system and / or the brake pipeline; A vibration sensor provided at the locomotive body and / or the axle position, and the vibration sensor is used to detect the vibration signal at the locomotive body and / or the axle position; A Hall current sensor provided on the motor power supply line, and the Hall current sensor is used to detect the current signal of the motor power supply line.
5. The intelligent monitoring method for operation and maintenance data of rail locomotives according to claim 1, wherein The method further includes: Use a generative adversarial network for training the intelligent annotation model. The generative adversarial network includes a generator and a discriminator. The input of the generator is a random noise vector, which is used to generate a fake data distribution similar to the real data through multiple transposed convolutional layers. The discriminator is used to distinguish whether the input data comes from the real data set or the fake data generated by the generator, that is, to calculate the confidence of the fake data generated by the generator. When the probability that the fake data generated by the generator is judged as real data by the discriminator is greater than the preset confidence threshold, regard this data as pseudo-annotated data and add it to the annotation set for subsequent model training.
6. The intelligent monitoring method for the operation and maintenance data of rail locomotives according to claim 1, characterized in that, The method further includes: After the intelligent annotation model outputs the annotation result, generate an intelligent decision-making suggestion according to the use of a decision algorithm.
7. The intelligent monitoring method for operation and maintenance data of rail locomotives according to claim 6, characterized in that, After the intelligent annotation model outputs the annotation result, generating an intelligent decision-making suggestion according to the use of a decision algorithm includes: Extract key features from the labeled data. The key features include temperature change trend, pressure fluctuation amplitude, vibration frequency, and fault causes. The key features also include maintenance history data, which corresponds one-to-one with the fault causes. The maintenance history data is input by the user into the system database, and the maintenance history data corresponding to the fault cause is obtained from the system database. Use the decision tree algorithm to construct a decision tree model, with the extracted features as input nodes, fault causes and maintenance history data as intermediate nodes, and decision suggestions as leaf nodes. Through the training of the labeled data, learn the relationship between the features and the decision results. Use the leave-one-out method to quantitatively evaluate the effectiveness of the decision tree model. Each time, select 1 sample as the test set, and the remaining samples as the training set for training. Calculate the prediction accuracy rate each time, and finally take the average value as the evaluation result.
8. The intelligent monitoring method for the operation and maintenance data of rail locomotives according to claim 7, wherein, The decision suggestions include parking for maintenance and replacing components. The fault causes include engine overheating and cooling system failure that may be caused by the cooling system failure.
9. An intelligent monitoring system for operation and maintenance data of rail locomotives, characterized in that, It includes: A data acquisition module, which is set on the rail locomotive and is used to obtain the key operating parameter data of the rail locomotive. A control module, which is used to preprocess the collected key operating parameter data according to a preset method. It is used to partially label the collected key operating parameter data, label the causes of abnormal data, record the labeled data set as the labeled set, record the unlabeled data as the unlabeled set, mix the labeled set and the unlabeled data according to a preset ratio to form a training data set, and then train the intelligent labeling model. It is used to immediately trigger an emergency labeling process when it detects that the key operating parameters of the rail locomotive exceed the preset threshold, use the intelligent labeling model to preferentially label the newly emerged abnormal data, and output a prompt signal regarding the labeling result.
10. An electronic device, characterized in that, It includes: A memory; And a processor. The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the intelligent monitoring method for rail locomotive operation and maintenance data according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Maintenance decision tree / word vector-based fault remote diagnosis platform
CN106054857A
Voltage regulator fault diagnosis method based on self-training semi-supervised generative adversarial network
CN113884290A
Multi-mode anomaly detection method for high-speed rail train control system
CN118915692A
Intelligent detection and supervision system and detection method for wind power generation equipment
CN118998004A
Transform-based abnormal index prediction method for water-turbine generator set
CN119150063A
Cited By
Rail transit field operation and maintenance data security assessment method
CN120995358A
Rail transit field section operation and maintenance data security evaluation method
CN120995358B