Train operation data analysis method based on self-learning
Through the real-time data cleaning, analysis and transpose of the data acquisition and processing platform, the proportion of abnormal data is counted, and instructions are generated to adjust the data acquisition frequency or issue fault notifications, which solves the problem of low accuracy and reliability of data analysis in the existing technology, and improves the safety and maintenance efficiency of train operations.
Patent Information
- Application Number
- CN202510670345.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art cannot process and analyze the acquired vehicle data sets in real time, resulting in low accuracy and reliability of the data analysis process, affecting the safety and operation and maintenance efficiency of train operations.
Real-time data acquisition is carried out through several data sensors, and data is transmitted to the data processing platform using the on-board wireless transmission system, data cleaning, analysis and transpose, statistics of the number of normal and abnormal data, calculate the proportion of abnormal data, and generate corresponding instructions to adjust the data acquisition frequency or issue fault repair notifications.
Improve the accuracy and reliability of data analysis, ensure the safety and maintenance efficiency of train operation, and dynamically adjust the data analysis system parameters to adapt to equipment aging and environmental changes.
Smart Images

Figure CN120503852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of train operation monitoring, and in particular to a train operation data analysis method based on self-learning. Background Art
[0002] With the rapid development of global rail transportation (such as high-speed rail and urban rail transit), train safety, efficiency, and service quality are facing higher demands. At the same time, the prevalence of intelligent and networked train technologies (such as IoT-based train control systems and onboard sensor networks) has enabled the real-time collection of massive amounts of operational data. This real-time train operation data can be used for analysis to ensure efficient train operations and passenger safety.
[0003] Prior art Chinese patent publication number CN112678035B provides a train operation data analysis method, system, server, and computer-readable medium. This technical solution receives train operation data from a data communication module installed on a target train, where the data communication module obtains the train operation data from at least one subsystem of the train automatic control system on the target train. Fault information is then extracted from the train operation data. Based on the time of occurrence of the train fault indicated by the fault information and the state of the target train at the time of the fault, the method determines whether the fault information is invalid. In response to a received fault query instruction, the method displays valid fault information, excluding invalid fault information, in the fault information through a display interface, thereby improving the efficiency of train fault troubleshooting. However, this technical solution does not establish a "data-decision-feedback" closed loop. Instead, it extracts fault information from operation data based on fixed thresholds or a historical rule base. This method lacks self-learning capabilities and cannot adapt to the effects of equipment aging and environmental changes, thus affecting the efficiency of train operation and maintenance. Summary of the Invention
[0004] To this end, the present invention provides a train operation data analysis method based on self-learning to solve the problem that the existing technology is unable to process and analyze the acquired vehicle data set in real time and dynamically adjust the corresponding parameters or devices in the data analysis system based on the analysis results, resulting in low accuracy and reliability of the data analysis process, and thus low safety of train operation and low efficiency in the execution of train operation and maintenance strategies.
[0005] To achieve the above objectives, the present invention provides a train operation data analysis method based on self-learning, comprising:
[0006] Real-time data collection of the vehicle is performed through a number of data sensors and the collected raw data sets of the vehicle are transmitted to the data processing platform through the on-board wireless transmission system;
[0007] Preprocessing the original vehicle dataset using a data processing platform to obtain a processed vehicle dataset, wherein the preprocessing process includes data cleaning, data parsing, and data transposition;
[0008] Performing statistics on basic data information, categories, quantities, and time nodes corresponding to each data on the processed vehicle data set and combining data correlation analysis to determine the number of normal vehicle operation data and the number of abnormal vehicle operation data;
[0009] Calculating the sum of the number of normal vehicle data and the number of abnormal vehicle data, and recording the sum as the total vehicle data;
[0010] Calculating a ratio between the number of abnormal vehicle data and the total amount of vehicle data, and recording the ratio as the abnormal data ratio; determining whether the train data collection and analysis process meets the standards based on the abnormal data ratio; and generating corresponding instructions if the standards are not met;
[0011] Based on the instruction, a data collection frequency for the vehicle, a data interval value in the data cleaning, or a fault maintenance notice for the vehicle is determined.
[0012] Furthermore, the process of preprocessing the original vehicle dataset to obtain the processed vehicle dataset by the data processing platform includes:
[0013] Cleaning the original vehicle dataset to obtain a cleaned vehicle dataset;
[0014] Compiling a data parsing rule table corresponding to the vehicle;
[0015] parsing the cleaned vehicle dataset based on the data parsing rule table to obtain a vehicle parsed dataset;
[0016] The vehicle parsed dataset is transposed to obtain the processed vehicle dataset.
[0017] Furthermore, the process of determining whether the train data collection and analysis process meets the standards based on the abnormal data proportion includes:
[0018] Determining whether the train data collection and analysis process meets the standards is re-determined based on a comparison result of the abnormal data ratio with a preset abnormal data ratio, or based on a comparison result of the total amount of data in the processed vehicle data set with a critical data processing amount, wherein the critical data processing amount is the total amount of data that the data processing platform can process within a single detection cycle;
[0019] When it is determined that the train data collection and analysis process does not meet the standards, the reason for not meeting the standards is determined based on the difference between the abnormal data ratio and the preset abnormal data ratio.
[0020] Furthermore, the process of re-determining whether the train data collection and analysis process meets the standards based on the comparison result of the amount of data in the processed vehicle data set and the total amount of critical data processing includes:
[0021] Counting the number of data in the processed vehicle data set and recording it as the total amount of processed data;
[0022] Based on the comparison result of the total amount of processed data and the total amount of critical data processed, it is determined whether to increase the number of shards for node-wise processing of the total amount of processed data.
[0023] Furthermore, the process of determining whether to increase the number of shards for performing node-wise processing on the total amount of processed data based on the comparison result of the total amount of processed data with the total amount of critical data processed includes:
[0024] Calculating the difference between the total amount of processed data and the total amount of critical data processed, and recording the difference as a data processing amount difference;
[0025] The number of shards to be increased is determined based on a comparison result of the data processing volume difference with a preset data processing volume difference, and the increase in the number of shards is directly proportional to the data processing volume difference.
[0026] Furthermore, the process of determining the cause according to the difference between the abnormal data ratio and the preset abnormal data ratio includes:
[0027] Calculate the difference between the abnormal data ratio and the preset abnormal data ratio, and record the difference as the data ratio difference;
[0028] Determine the reason why the train data collection and analysis process does not meet the standards based on a comparison result of the data proportion difference and a preset data proportion difference;
[0029] The data interval value in the data cleaning is determined based on the cause, and the corresponding processing or issuance of a fault maintenance notice for the vehicle is determined based on the amount of data in the vehicle original data set within several detection cycles.
[0030] Furthermore, the process of determining the data interval value in the data cleaning includes:
[0031] Counting the number of data in the processed vehicle data set and recording it as the total amount of processed data, and counting the number of data in the original vehicle data set and recording it as the total amount of original data;
[0032] Calculating the difference between the total amount of original data and the total amount of processed data and recording the difference as a processed data difference;
[0033] The data interval value in the data cleaning is determined based on a comparison result of the processed data difference value and a preset processed data difference value.
[0034] Furthermore, the process of determining corresponding processing based on the amount of data in the vehicle original data set within the plurality of detection cycles includes:
[0035] Counting the number of raw data in the vehicle raw data set in several historical detection cycles and the number of raw data in the vehicle raw data set in the current detection cycle, and calculating the raw data variance based on the number of raw data;
[0036] Based on a comparison result of the raw data variance with a preset raw data variance, it is determined whether to increase the data collection frequency for the vehicle or to issue a fault maintenance notice for the vehicle.
[0037] Furthermore, the process of increasing the data collection frequency for the vehicle includes:
[0038] Calculating the difference between the original data variance and the preset original data variance and recording it as the variance difference;
[0039] Based on a comparison result between the variance difference value and a preset variance difference value, it is determined that the data acquisition frequency should be increased, and the increase range of the data acquisition frequency is proportional to the variance difference value.
[0040] Furthermore, after completing the increase in the data collection frequency for the vehicle, when it is determined that the train data collection and analysis process does not meet the standards based on the comparison result of the original data variance and the preset original data variance, a notification to increase the number of data collection sensors is issued based on the instruction.
[0041] Compared with the prior art, the train operation data analysis method based on self-learning of the present invention has the following advantages: the method collects real-time data from the vehicle through a plurality of data sensors and transmits the collected vehicle raw data set to a data processing platform through an on-board wireless transmission system; the vehicle raw data set is pre-processed by the data processing platform to obtain a processed vehicle data set; the processed vehicle data set is statistically analyzed in terms of data type and quantity to determine the number of normal vehicle operation data and the number of abnormal vehicle data; the abnormal data ratio is determined based on the number of normal vehicle operation data and the number of abnormal vehicle data, and based on the abnormal data ratio, it is determined whether the train data collection and analysis process meets the standard, and corresponding instructions are generated if it does not meet the standard; the data collection frequency for the vehicle, the data interval value in the data cleaning, or the issuance of a fault maintenance notice for the vehicle are determined based on the instructions. By determining whether the current train data collection and analysis process meets the normal standard based on the analysis results of the data, and determining whether there is a problem with the acquired vehicle operation data, the parameters or devices in the system can be dynamically re-determined and adjusted to ensure that the acquired operation data is accurate and reliable, thereby ensuring the accuracy of subsequent analysis and early warning based on the vehicle data, and improving the safety of train operation and maintenance efficiency.
[0042] Furthermore, the present invention determines whether the train data collection and analysis process meets the standards based on the comparison results of the abnormal data ratio and the preset abnormal data ratio, and when it is determined that it does not meet the standards, the reason for non-compliance with the standards can be determined based on the difference between the abnormal data ratio and the preset abnormal data ratio.
[0043] Furthermore, the present invention further re-determines whether the train data collection and analysis process meets the standards based on the comparison results of the total amount of processed data and the total amount of critical data processed, and performs a secondary judgment, thereby improving the accuracy of the judgment on the train data collection and analysis process.
[0044] Furthermore, the present invention can determine the number of shards to be increased based on the comparison results when comparing the total amount of processed data with the total amount of critical data processed, and determine the number of shards to be increased based on the comparison results of the difference in the amount of processed data with the difference in the preset data processing amount. By increasing the number of shards, the data processing platform's ability to process and analyze large amounts of real-time data is improved, thereby improving the processing efficiency of vehicle data.
[0045] Furthermore, the present invention can further determine the reasons why the train data collection and analysis process does not meet the standards based on the comparison results of the data proportion difference and the preset data proportion difference, and then determine the corresponding processing method according to the corresponding reason, thereby improving the accuracy of the train data collection and analysis process.
[0046] Furthermore, the present invention determines that the reason for non-compliance with the standards is that there is a problem in the data preprocessing process of the vehicle original data set, and then determines the data interval value for adjusting the data cleaning in the preprocessing process based on the comparison result of the processed data difference and the preset processed data difference, thereby improving the data processing platform's screening process for abnormal data, ensuring the number of data samples in the subsequent analysis process, and improving the accuracy of early warning analysis.
[0047] Furthermore, when the present invention determines that there is a problem in the data collection process under the current detection cycle based on the comparison result of the original data variance and the preset original data variance, it further determines to increase the data collection frequency based on the comparison result of the variance difference and the preset variance difference, thereby distributing the original data samples more densely, thereby improving the train's ability to collect data in real time, reducing data omissions, and ensuring the authenticity and reliability of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Schematic diagram of modules of a system for analyzing train operation data based on self-learning according to the present invention;
[0049] Figure 2 Schematic diagram of the flow of the train operation data analysis method based on self-learning of the present invention;
[0050] Figure 3 This is a logic decision diagram for determining whether the train data collection and analysis process meets the standards based on the proportion of abnormal data in the present invention;
[0051] Figure 4 This is a logical decision diagram for determining the reasons why the train data collection and analysis process does not meet the standards and the corresponding processing methods based on the data proportion difference in the present invention. DETAILED DESCRIPTION
[0052] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0053] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0054] It should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the term "connection" should be understood in a broad sense. For example, it can mean a fixed connection, a detachable connection, or an integral connection; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0055] The vehicle dataset obtained in this embodiment is collected for a single train, and the operation process of the single train is analyzed. It can be clearly seen that the analysis process is also applicable to other trains. The data sources include but are not limited to the door system, network system, air-conditioning system, traction system, auxiliary system, braking system, auxiliary system, power system, etc.
[0056] See also Figure 1 As shown, it is a module diagram of the system of the self-learning train operation data analysis method of this embodiment, and the system includes: a data acquisition module, an on-board wireless transmission system, a data processing platform, and an adjustment module. The data acquisition module is connected to the on-board wireless transmission system, and the data acquisition module includes a number of data sensors, which collect real-time data of the vehicle through the several data sensors and transmit the collected vehicle original data set to the data processing platform through the on-board wireless transmission system; the data processing platform is connected to the data acquisition module, and the data processing platform is used to pre-process the vehicle original data to obtain a processed vehicle data set, and the data processing platform is also used to perform statistics on the basic information, data category, data quantity and time nodes corresponding to each data of the processed vehicle data set and combine the data type correlation analysis to determine the number of normal vehicle operation data and the number of abnormal vehicle data, and calculate the number of normal vehicle operation data and the number of abnormal vehicle data. The sum of the data is recorded as the total amount of vehicle data, the ratio between the number of abnormal vehicle data and the total amount of vehicle data is calculated, and the ratio is recorded as the abnormal data ratio. Based on the abnormal data ratio, it is determined whether the train data collection and analysis process meets the standards, and corresponding instructions are generated when it does not meet the standards. Among them, the basic data information includes vehicle model, purchase date, etc. The data correlation analysis can identify abnormal data that does not match the vehicle model performance based on the combined statistics of the basic data information, type, quantity and each time node; the adjustment module is connected to the data acquisition module and the data processing platform respectively, and the adjustment module is used to determine the data collection frequency for the vehicle, the data interval value in the data cleaning or issue a fault maintenance notice for the vehicle based on the instructions.
[0057] Specifically, in this embodiment, real-time data collection is performed on the vehicle through a number of sensors. Then, the vehicle's on-board host transmits the vehicle's original data set through 4G, 5G, EUHT (Enhanced Ultra High Throughput) and other channels through the on-board wireless transmission system, and then transmits it to the data processing platform through the TCP protocol. The data processing platform includes a data receiving service program developed based on the Netty framework, which receives and preliminarily processes data in real time in a load balancing manner, including packet sticking, unpacking, response processing, heartbeat reply, and CRC verification to ensure data accuracy. The load balancing methods include HTTP redirection load balancing, DNS domain name resolution load balancing, reverse proxy load balancing (such as Nginx), IP load balancing, etc. The vehicle's original data set is then preprocessed by the data processing platform. The preprocessing process includes data cleaning, data parsing, and data transposition. The data processing platform also includes KafKa (distributed stream processing platform) and IoTDB (time series database). The processed vehicle data set is sent to KafKa and stored in IoTDB as data reserve for subsequent early warning and analysis. A determination is made based on the number of normal vehicle operation data and the number of abnormal vehicle operation data to determine the status of the train data collection and analysis process. When it is determined that the train data collection and analysis process does not meet the standards, the cause can be determined. Non-compliance here means that the proportion of the number of abnormal vehicle operation data in the total amount of vehicle operation data exceeds the preset standard, where the total amount of vehicle operation data is the sum of the number of normal vehicle operation data and the number of abnormal vehicle operation data. The data correlation analysis here can distinguish between normal vehicle operation data and abnormal vehicle operation data by setting device parameter thresholds. For example, data with a motor temperature of less than 60°C and a three-phase current deviation of less than 5% can be regarded as normal vehicle operation data, and data outside this range can be regarded as abnormal vehicle operation data. In this case, a single abnormal vehicle operation data will not have a significant impact on vehicle operation. A time efficiency index can also be set. For example, data with a plan fulfillment rate of ≥99 and a delay time of ≤2 minutes / 10,000 vehicle kilometers can be regarded as normal vehicle operation data, and data outside this range can be regarded as abnormal vehicle operation data. In this case, a single abnormal vehicle operation data will not have a significant impact on vehicle operation. When the cause is determined, corresponding instructions can be generated, and then the data collection frequency for the vehicle can be adjusted through the instructions, the data interval value in the data cleaning process can be re-determined, or a maintenance work order can be established and a maintenance notice for a single vehicle can be issued.This embodiment collects vehicle data and processes and analyzes it through a data processing platform. It can then determine whether the current train data collection and analysis process meets normal standards based on the data analysis results, and can determine whether there are problems with the acquired vehicle operation data. It can then adjust the data acquisition and processing process to ensure the accuracy of subsequent analysis and early warning based on vehicle data, thereby improving the safety of train operation and maintenance efficiency.
[0058] See also Figure 2 As shown in FIG, it is a flow chart of a train operation data analysis method based on self-learning in this embodiment. The process steps include:
[0059] S1: Real-time data collection of the vehicle is performed through a number of data sensors and the collected raw data set of the vehicle is transmitted to the data processing platform through the vehicle-mounted wireless transmission system;
[0060] S2: Preprocessing the original vehicle dataset using a data processing platform to obtain a processed vehicle dataset, wherein the preprocessing process includes data cleaning, data parsing, and data transposition;
[0061] S3: performing statistics on basic data information, categories, quantity, and time nodes corresponding to each data on the processed vehicle data set, and combining data correlation analysis to determine the number of normal vehicle operation data and the number of abnormal vehicle operation data;
[0062] S4: Calculate the sum of the number of normal vehicle data and the number of abnormal vehicle data, and record the sum as the total vehicle data;
[0063] S5: Calculate the ratio between the number of abnormal vehicle data and the total amount of vehicle data, and record the ratio as the abnormal data ratio; determine whether the train data collection and analysis process meets the standards based on the abnormal data ratio; and generate corresponding instructions if it does not meet the standards;
[0064] S6: Determine the data collection frequency for the vehicle, the data interval value in the data cleaning, or issue a fault maintenance notice for the vehicle based on the instruction.
[0065] Furthermore, the process of preprocessing the original vehicle dataset to obtain the processed vehicle dataset by the data processing platform includes:
[0066] The vehicle original data set is cleaned to obtain a cleaned vehicle data set; a data parsing rule table corresponding to the vehicle is compiled; the cleaned vehicle data set is parsed based on the data parsing rule table to obtain a vehicle parsed data set; and the vehicle parsed data set is transposed to obtain the processed vehicle data set.
[0067] Specifically, in this embodiment, the raw vehicle data can be cleaned to remove the corresponding data. A corresponding vehicle data parsing rule table can be compiled based on the communication protocol. The byte offset, bit position, and length corresponding to the fields described in the communication protocol are compiled into corresponding data codes (e.g., 0001 represents the first complete byte, 00011 represents the first bit of the first byte, and 000104 represents the first byte starting at 0 and having a length of 4 bits). The data codes are then configured with corresponding parsing rules (e.g., conversion to a decimal value, conversion to a floating-point type, conversion to an ASCII character, conversion to a specially formatted version number, etc.). Using the above rule table, data parsing can be performed to obtain a vehicle parsed dataset. The vehicle parsed dataset is then transposed to obtain a processed vehicle dataset. The processed vehicle dataset can adapt the data to the data format requirements of different scenarios.
[0068] See also Figure 3 As shown, it is a logical decision diagram for determining whether the train data collection and analysis process meets the standards based on the abnormal data ratio in this embodiment. The process of determining whether the train data collection and analysis process meets the standards based on the abnormal data ratio includes:
[0069] A judgment is made based on a comparison result between the abnormal data ratio and the preset abnormal data ratio, or a re-judgment is made as to whether the train data collection and analysis process meets the standards based on a comparison result between the total amount of data in the processed vehicle data set and the critical data processing total amount, wherein the critical data processing total amount is the total amount of data that the data processing platform can process within a single detection cycle; when it is determined that the train data collection and analysis process does not meet the standards, the reason for non-compliance is determined based on the difference between the abnormal data ratio and the preset abnormal data ratio.
[0070] Specifically, in this embodiment, in order to make the determination process more accurate, the preset abnormal data proportion B0 can be divided into a first preset abnormal data proportion B1 and a second preset abnormal data proportion B2. The vehicle's brake system data is analyzed and the preset abnormal data proportion standards B3=6%, B1=B3-0.5%, and B2=B3+0.5% are set. The specific comparison process based on the abnormal data proportion B with B1 and B2 is as follows:
[0071] If the abnormal data proportion B is less than or equal to the first preset abnormal data proportion B1, it means that the current number of vehicle abnormal data is relatively small, and the cumulative amount of abnormal data has not reached the stage of triggering a fault warning. At this time, it can be determined that the train data collection and analysis process meets the standards, and no corresponding processing measures are required within the detection cycle, which can effectively reduce the maintenance cost of the train. It is clear that B at this time is not infinitely small.
[0072] If the abnormal data proportion B is greater than the first preset abnormal data proportion B1 and less than or equal to the second preset abnormal data proportion B2, the amount of vehicle abnormal data is between relatively small and relatively large. At this time, in order to further determine whether the vehicle operation process meets the standards, the total amount of data in the processed vehicle data set can be obtained, and then the total amount of data is compared with the critical data processing amount based on the comparison result and the train data collection and analysis process is re-judged based on the comparison result.
[0073] If the abnormal data proportion B is greater than the second preset abnormal data proportion B2, it means that the current number of vehicle abnormal data is relatively large. If the fault warning critical value is set to 6.45%, it can be known that the number of vehicle abnormal data at this time exceeds the critical value that triggers a fault warning under normal circumstances. A fault warning should be issued at this time, but no warning is issued subsequently. At this time, it can be determined that the train data collection and analysis process does not meet the standards. It is necessary to compare the difference between the abnormal data proportion and the preset abnormal data proportion with the preset value to determine the reason why the train data collection and analysis process does not meet the standards. Based on the determined reason, the corresponding processing method is determined to adjust the train's data collection and analysis process to ensure the accuracy of the data. It is clear that B at this time is not infinite. It can be understood that the embodiment of the present invention does not impose any specific restrictions on the preset abnormal data proportion B0. In an optional embodiment, a new preset abnormal data proportion standard B31=5.5%, a new first preset abnormal data proportion B11=B31-0.5%, and a new second preset abnormal data proportion B21=B31+0.5% can be set. As long as the comparison between B and B11 and B21 can determine whether the train data collection and analysis process meets the standards and the reasons can be determined when it does not meet the standards, the specific comparison process can be referred to the above content.
[0074] Furthermore, the process of re-determining whether the train data collection and analysis process meets the standards based on the comparison result of the amount of data in the processed vehicle data set and the total amount of critical data processing includes:
[0075] The amount of data in the processed vehicle data set is counted and recorded as the total amount of processed data; based on the comparison result of the total amount of processed data and the total amount of critical data processing, it is determined whether to increase the number of shards for node processing of the total amount of processed data.
[0076] Specifically, in this embodiment, the comparison process based on the total amount of processed data D1 and the total amount of critical data processed D2 is as follows:
[0077] If the total amount of processed data D1 is less than or equal to the critical data processing amount D2, it means that the number of data blocks determined by the current data processing platform based on the total amount of processed data is appropriate, and the data processing platform can normally process data for the train data collection and analysis process based on the number of normal vehicle operation data and the number of vehicle abnormal data. At this time, if there is a situation where the abnormal data proportion B is greater than the first preset abnormal data proportion B1 and less than or equal to the second preset abnormal data proportion B2, it is determined that the train data collection and analysis process does not meet the standards. At this time, it is necessary to determine the reason for non-compliance with the standards. It is clear that D1 at this time is not infinitely small.
[0078] If the total amount of processed data D1 is greater than the critical data processing amount D2, it indicates that the current total amount of processed data is relatively large. In this case, tools such as Spark and Flink can be used to process the total amount of processed data in parallel across nodes. Dynamically increasing the number of shards in the data processing platform can shard the total amount of processed data to more cluster nodes for processing. This reduces the data processing burden of the data processing platform within a single detection cycle, improves the accuracy of analysis results, and thus provides more precise analysis of the train data collection and analysis process. For example, data from 1 to 10,000 can be stored in shard 1, and data from 10,001 to 20,000 can be stored in shard 2. The setting of the critical data processing amount is determined by the data processing system in the data processing platform. The data processing systems that can be used include relational database managers (RDBMs) or distributed file systems (Hadoop Distributed File Systems (HDFS)) for data processing and analysis. It is clear that D1 is not infinite in this case.
[0079] In other embodiments, the total amount of processed data can also be divided into blocks for processing, and the real-time data can be divided into blocks according to time windows (such as every minute) or event volume (such as every 1,000) to avoid excessive single processing load.
[0080] Furthermore, the process of determining whether to increase the number of shards for performing node-wise processing on the total amount of processed data based on the comparison result of the total amount of processed data with the total amount of critical data processed includes:
[0081] Calculate the difference between the total amount of processed data and the total amount of critical data processed, and record the difference as the data processing volume difference; determine the number of shards to be increased based on the comparison result of the data processing volume difference and the preset data processing volume difference, and the increase in the number of shards is proportional to the data processing volume difference.
[0082] Specifically, in this embodiment, the data processing platform uses the Flink tool. Flink supports dynamic scaling (such as Kubernetes autoscaling) and uses a distributed database. The number of shards can be automatically adjusted with the cluster size. The preset data processing volume difference W0 can be divided into a first preset data processing volume difference W1 and a second preset data processing volume difference W2. W1=40GB and W2=160GB are set. The total critical data processing volume in a single detection cycle is set to 80GB. The specific process of comparing the data processing volume difference W with W1 and W2 is as follows:
[0083] Set the original number of shards to 3100, and the data size that a single shard can process is 256MB. If the data processing volume difference W is less than or equal to the first preset data processing difference W1, and W is set to 40GB, other idle cluster nodes can be enabled to increase the number of shards by about 156.
[0084] If the data processing volume difference W is greater than the first preset data processing volume difference W1 and less than or equal to the second preset data processing volume difference W2, and W=160GB is set, other idle cluster nodes can be enabled to increase the number of shards by about 625.
[0085] If the data processing volume difference W is greater than the second preset data processing volume difference W2, and W is set to 180GB, other idle cluster nodes can be enabled to increase the number of shards by about 703. It is clear that the increase in the number of shards is proportional to the data processing difference. The larger the data processing difference, the more shards are increased. It should be noted that the increase in the number of shards is also related to the data size that a single shard can process and other related parameters. The relevant parameters include the comprehensive data scale, storage performance, and computing requirements. Therefore, the increase in the number of shards can also be set to other values. In this embodiment, for the convenience of example calculation, the data is set according to the decimal standard, 1TB = 1000GB, 1GB = 1000MB. It can be understood that the embodiment of the present invention does not impose specific restrictions on the preset data processing volume difference W0 and the critical data processing volume. In an optional embodiment, a new first preset data processing volume difference W11=35GB, a new second preset data processing volume difference W21=155GB, and a new critical data processing volume of 75GB can be set. As long as the increase in the number of shards can be determined by comparing W with W11 and W21 and combining the new critical data processing volume, the specific comparison process can be referred to the above content.
[0086] See also Figure 4 As shown, it is a logical decision diagram for determining the reasons why the train data collection and analysis process does not meet the standards and the corresponding processing methods based on the data proportion difference in this embodiment. The process of determining the reason based on the difference between the abnormal data proportion and the preset abnormal data proportion includes:
[0087] Calculate the difference between the abnormal data proportion and the preset abnormal data proportion, and record the difference as the data proportion difference; determine the reason why the train data collection and analysis process does not meet the standards based on the comparison result of the data proportion difference and the preset data proportion difference; determine the data interval value in the data cleaning based on the reason, and determine the corresponding processing or issue a fault maintenance notice for the vehicle based on the number of data in the vehicle original data set within several detection cycles.
[0088] Specifically, in this embodiment, the preset data proportion difference S0 can be divided into a first preset data proportion difference S1 and a second preset data proportion difference S2, with S1=0.6% and S2=1%. The comparison process based on the data proportion difference S with S1 and S2 is as follows:
[0089] If the data proportion difference S is less than or equal to the first preset data proportion difference S1, it is determined that the reason why the train data collection and analysis process does not meet the standards is that there is a problem in the data preprocessing process of the vehicle original data set. Problems in the data preprocessing process will increase the number of abnormal vehicle data that are screened out, thereby reducing the number of data samples provided for subsequent early warning and maintenance analysis processes, thereby affecting the accuracy of the early warning analysis. It is necessary to correct the data preprocessing process, including correcting the data interval value in data cleaning. It is clear that S at this time is not infinitely small.
[0090] If the data proportion difference S is greater than the first preset data proportion difference S1 and is less than or equal to the second preset data proportion difference, it is determined that the reason why the train data collection and analysis process does not meet the standards is that there is a problem with the data collection of the vehicle or the vehicle itself has a fault. At this time, it is necessary to obtain the number of vehicle original data within several detection cycles, and then determine whether it is necessary to increase the data collection frequency or issue a fault maintenance notice for the vehicle based on the number of vehicle original data to resolve the non-compliance with the train data collection and analysis process.
[0091] If the data proportion difference S is greater than the second preset data proportion difference S2, it can be directly determined that the reason why the train data collection and analysis process does not meet the standards is that the vehicle itself generates too much abnormal data. It is determined that there is a fault problem with the vehicle and it is necessary to establish a maintenance work order based on the abnormal data and issue a maintenance notice for the individual vehicle based on the maintenance work order. It can be clearly seen that S at this time is not infinite. It is understandable that the embodiment of the present invention does not impose specific restrictions on the preset data proportion difference S0. In an optional embodiment, a new first preset data proportion difference S11 = 0.7% and a new second preset data proportion difference S21 = 1.1% can also be set. As long as the comparison of S with S11 and S21 can determine the reason why the train data collection and analysis process does not meet the standards, the specific comparison process can be referred to the above content.
[0092] Furthermore, the process of determining the data interval value in the data cleaning includes:
[0093] The number of data in the processed vehicle data set is counted and recorded as the total amount of processed data, and the number of data in the original vehicle data set is counted and recorded as the total amount of original data; the difference between the total amount of original data and the total amount of processed data is calculated and recorded as the processed data difference; and the data interval value in the data cleaning is determined based on the comparison result of the processed data difference and the preset processed data difference.
[0094] Specifically, in this embodiment, taking the collection of vehicle-mounted temperature sensor data for data cleaning as an example, the initial interval is set to [-45°C, 145°C]. If the processed data difference is larger, it means that the difference between the total amount of processed data and the total amount of original data is larger. In this case, the temperature value interval can be appropriately increased; the preset processed data difference M0 can be divided into a first preset processed data difference M1 and a second preset processed data difference M2, with M1=3GB and M2=5GB. The specific comparison process based on the processed data difference M with M1 and M2 is as follows:
[0095] If the processed data difference M is less than or equal to the first preset processed data difference M1, the initial interval is expanded to [-48°C, 148°C]; if the processed data difference M is greater than the first preset processed data difference M1 and less than or equal to the second preset processed data difference M2, the initial interval is expanded to [-52°C, 152°C]; if the processed data difference M is greater than the second preset processed data difference M2, the initial interval is expanded to [-55°C, 155°C]; it can be clearly seen that the expansion range of the initial interval is reasonable and in line with the existing technology, and the expansion range value of the initial interval can also be set to other values.
[0096] Furthermore, the process of determining corresponding processing based on the amount of data in the vehicle original data set within the plurality of detection cycles includes:
[0097] The number of raw data in a vehicle raw data set within several historical inspection cycles and the number of raw data in a vehicle raw data set within a current inspection cycle are counted, and a raw data variance is calculated based on each raw data quantity; and based on a comparison result of the raw data variance with a preset raw data variance, it is determined whether to increase the data collection frequency for the vehicle or issue a fault maintenance notice for the vehicle.
[0098] Specifically, in this embodiment, taking the subway train power system as an example, where the power system is determined to be powered by a 1500V system and the rated current of the traction motor is 800A, the normal fluctuation range is 24A-40A, and the preset raw data variance R0 is set to 350; the comparison process based on the raw data variance R and the preset raw data variance R0 is as follows:
[0099] If the raw data variance R is less than or equal to the preset raw data variance R0, it is determined that the variation range between the numbers of raw data corresponding to each detection cycle within the preset time period is relatively small, that is, the overall numerical size of each raw data number is relatively close. When it is determined that the data collection of the historical detection cycle is normal, it can be determined that the data collection process of the current inspection cycle is normal. At this time, a maintenance work order can be established based on the abnormal data, and a maintenance notice for the vehicle can be issued based on the maintenance work order. It can be clear that R at this time is not infinitely small.
[0100] If the raw data variance R is greater than the preset raw data variance R0, it is determined that the variation between the raw data quantities corresponding to each detection cycle within the preset time period is relatively large, that is, the overall numerical values of the raw data quantities are relatively discrete. This indicates that there may be a problem with the data collection process during the current detection cycle, such as data omissions during data collection during a certain detection cycle, resulting in a large degree of variation in the raw data quantities collected during each detection cycle. In this case, the vehicle-mounted wireless transmission system can be adjusted to collect data from the vehicle. By increasing the data collection frequency, a more dense raw data distribution sample can be provided to avoid data omissions caused by excessively long collection time intervals. It is clear that R in this case is not infinite. It is understood that the embodiment of the present invention does not specifically limit the preset raw data variance R0. In an optional embodiment, the preset raw data variance R01 can also be set to 355, as long as the comparison between R and R01 determines whether to increase the data collection frequency for the vehicle or issue a vehicle fault maintenance notification. The specific comparison process is described above.
[0101] Furthermore, the process of increasing the data collection frequency for the vehicle includes:
[0102] The difference between the original data variance and the preset original data variance is calculated and recorded as the variance difference; based on the comparison result of the variance difference and the preset variance difference, it is determined to increase the data collection frequency, and the increase in the data collection frequency is proportional to the variance difference.
[0103] Specifically, in this embodiment, taking the motor current acquisition frequency in the power system under normal conditions as an example, the initial motor current acquisition frequency is set to 200 Hz; the preset variance difference value H0 can be divided into a first preset variance difference value H1 and a second preset variance difference value H2, with H1=20 and H2=30; the comparison process based on the variance difference value H with H1 and H2 is as follows:
[0104] If the variance difference H is less than or equal to the first preset variance difference H1, the data collection frequency is adjusted to 1.5 times the initial value by using the first frequency adjustment coefficient. It can be clearly seen that H at this time is not infinitely small; if the variance difference H is greater than the first preset variance difference H1 and less than or equal to the second preset variance difference H2, the data collection frequency is adjusted to 2 times the initial value by using the second frequency adjustment coefficient; if the variance difference H is greater than the second preset variance difference H2, the data collection frequency is adjusted to 2.5 times the initial value by using the third frequency adjustment coefficient. It can be clearly seen that H at this time is not infinitely large; it should be noted that the increase ratio of the data collection frequency can be set to other values according to the initial value, and the increase in the data collection frequency is consistent with the data collection of the vehicle. It can be understood that the embodiment of the present invention does not impose any specific restrictions on the preset variance difference value H0. In an optional embodiment, a new first preset variance difference value H11=22 and a new second preset variance difference value H21=32 can be set, as long as the increase ratio of the data acquisition frequency can be determined by comparing H with H11 and H21. For the specific comparison process, refer to the above content.
[0105] Furthermore, after completing the increase in the data collection frequency for the vehicle, when it is determined that the train data collection and analysis process does not meet the standards based on the comparison result of the original data variance and the preset original data variance, a notification to increase the number of data collection sensors is issued based on the instruction.
[0106] Specifically, in this embodiment, after increasing the data collection frequency for the vehicle, if it is still determined that the raw data variance R is greater than the preset raw data variance R0, a corresponding instruction can be generated through the data processing platform. The adjustment module issues a notification based on the instruction, and the maintenance personnel increases the number of data acquisition sensors on the vehicle or starts the number of spare data acquisition sensors on the vehicle according to the notification.
[0107] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0108] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A train operation data analysis method based on self-learning, characterized in that: include: Real-time data collection of the vehicle is performed through a number of data sensors and the collected raw data sets of the vehicle are transmitted to the data processing platform through the on-board wireless transmission system; Preprocessing the original vehicle dataset using a data processing platform to obtain a processed vehicle dataset, wherein the preprocessing process includes data cleaning, data parsing, and data transposition; Performing statistics on basic data information, categories, quantities, and time nodes corresponding to each data on the processed vehicle data set and combining data correlation analysis to determine the number of normal vehicle operation data and the number of abnormal vehicle operation data; Calculating the sum of the number of normal vehicle data and the number of abnormal vehicle data, and recording the sum as the total vehicle data; Calculating a ratio between the number of abnormal vehicle data and the total amount of vehicle data, and recording the ratio as the abnormal data ratio; determining whether the train data collection and analysis process meets the standards based on the abnormal data ratio; and generating corresponding instructions if the standards are not met; Based on the instruction, a data collection frequency for the vehicle, a data interval value in the data cleaning, or a fault maintenance notice for the vehicle is determined.
2. The train operation data analysis method based on self-learning according to claim 1, characterized in that: The process of preprocessing the original vehicle dataset by the data processing platform to obtain the processed vehicle dataset includes: Cleaning the original vehicle dataset to obtain a cleaned vehicle dataset; Compiling a data parsing rule table corresponding to the vehicle; parsing the cleaned vehicle dataset based on the data parsing rule table to obtain a vehicle parsed dataset; The vehicle parsed dataset is transposed to obtain the processed vehicle dataset.
3. The train operation data analysis method based on self-learning according to claim 2 is characterized in that: The process of determining whether the train data collection and analysis process meets the standards based on the abnormal data ratio includes: Determining whether the train data collection and analysis process meets the standards is re-determined based on a comparison result of the abnormal data ratio with a preset abnormal data ratio, or based on a comparison result of the total amount of data in the processed vehicle data set with a critical data processing amount, wherein the critical data processing amount is the total amount of data that the data processing platform can process within a single detection cycle; When it is determined that the train data collection and analysis process does not meet the standards, the reason for not meeting the standards is determined based on the difference between the abnormal data ratio and the preset abnormal data ratio.
4. The train operation data analysis method based on self-learning according to claim 3 is characterized in that: The process of re-determining whether the train data collection and analysis process meets the standards based on the comparison result of the data quantity in the processed vehicle data set and the total critical data processing amount includes: Counting the number of data in the processed vehicle data set and recording it as the total amount of processed data; Based on the comparison result of the total amount of processed data and the total amount of critical data processed, it is determined whether to increase the number of shards for node-wise processing of the total amount of processed data.
5. The train operation data analysis method based on self-learning according to claim 4 is characterized in that: The process of determining whether to increase the number of shards for node-wise processing of the total amount of processed data based on a comparison result of the total amount of processed data and the total amount of critical data processed includes: Calculating the difference between the total amount of processed data and the total amount of critical data processed, and recording the difference as a data processing amount difference; The number of shards to be increased is determined based on a comparison result of the data processing volume difference with a preset data processing volume difference, and the increase in the number of shards is directly proportional to the data processing volume difference.
6. The train operation data analysis method based on self-learning according to claim 3 is characterized in that: The process of determining the cause according to the difference between the abnormal data ratio and the preset abnormal data ratio includes: Calculate the difference between the abnormal data ratio and the preset abnormal data ratio, and record the difference as the data ratio difference; Determine the reason why the train data collection and analysis process does not meet the standards based on a comparison result of the data proportion difference and a preset data proportion difference; The data interval value in the data cleaning is determined based on the cause, and the corresponding processing or issuance of a fault maintenance notice for the vehicle is determined based on the amount of data in the vehicle original data set within several detection cycles.
7. The train operation data analysis method based on self-learning according to claim 6, characterized in that: The process of determining the data interval value in the data cleaning includes: Counting the number of data in the processed vehicle data set and recording it as the total amount of processed data, and counting the number of data in the original vehicle data set and recording it as the total amount of original data; Calculating the difference between the total amount of original data and the total amount of processed data and recording the difference as a processed data difference; The data interval value in the data cleaning is determined based on a comparison result of the processed data difference value and a preset processed data difference value.
8. The train operation data analysis method based on self-learning according to claim 6, characterized in that: The process of determining corresponding processing based on the amount of data in the vehicle original data set within a plurality of detection cycles includes: Counting the number of raw data in the vehicle raw data set in several historical detection cycles and the number of raw data in the vehicle raw data set in the current detection cycle, and calculating the raw data variance based on the number of raw data; Based on a comparison result of the raw data variance with a preset raw data variance, it is determined whether to increase the data collection frequency for the vehicle or to issue a fault maintenance notice for the vehicle.
9. The train operation data analysis method based on self-learning according to claim 8, characterized in that: The process of increasing the data collection frequency for the vehicle includes: Calculating the difference between the original data variance and the preset original data variance and recording it as the variance difference; Based on a comparison result between the variance difference value and a preset variance difference value, it is determined that the data acquisition frequency should be increased, and the increase range of the data acquisition frequency is proportional to the variance difference value.
10. The train operation data analysis method based on self-learning according to claim 9, characterized in that: After completing the increase in the data collection frequency for the vehicle, when it is determined that the train data collection and analysis process does not meet the standards based on the comparison result of the original data variance and the preset original data variance, a notification to increase the number of data collection sensors is issued based on the instruction.
Citation Information
Patent Citations
Train operation data analysis methods, systems, servers, and computer-readable media
CN112678035B
On-board diagnosis system data-based traffic condition analysis and early warning method
CN108769104A
Vehicle accessory abnormal information automatic detection processing method and related equipment thereof
CN115310542A
Abnormal vehicle data identification method, device and equipment
CN116403427A
Distributed task fragmentation method and system
CN117112212A