A train operation data analysis method based on self-learning

By collecting, cleaning, and analyzing data in real time, and dynamically adjusting the data collection frequency and cleaning interval values, the problem of low accuracy and reliability in existing train operation data analysis methods has been solved, thereby improving operation and maintenance efficiency and safety.

CN120503852BActive Publication Date: 2025-12-16BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510670345.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-12-16
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing train operation data analysis methods lack self-learning capabilities and cannot adapt to equipment aging and environmental changes, resulting in low accuracy and reliability of the data analysis process and affecting train operation and maintenance efficiency.

Method used

Vehicle data is collected in real time by several data sensors and transmitted to the data processing platform via an onboard wireless transmission system. The data is then cleaned, parsed, and transposed. The number of normal and abnormal data is counted, the percentage of abnormal data is calculated, and the data collection and analysis process is judged based on the percentage to determine whether it meets the standards. Corresponding instructions or notifications are generated, and the data collection frequency and cleaning interval are dynamically adjusted to ensure data accuracy.

Benefits of technology

It improves the accuracy and reliability of train operation data analysis, enhances operation and maintenance efficiency, ensures the accuracy of data analysis and early warning, reduces data omissions, and improves the safety of train operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120503852B_ABST
    Figure CN120503852B_ABST
Patent Text Reader

Abstract

The present application relates to train operation monitoring technical field, especially train operation data analysis method based on self-learning, the method is through sensor to vehicle data acquisition and through vehicle wireless transmission system to vehicle raw data set transmission to data processing platform, data processing platform to vehicle raw data set pre-processing to obtain after processing vehicle data set, to after processing vehicle data set data statistics to determine vehicle normal operation data quantity and vehicle abnormal data quantity and obtain abnormal data proportion, based on abnormal data proportion determination for train data acquisition and analysis process whether it is in line with standard and when not in line with standard generates instruction, based on instruction to determine the parameter of re-adjusting corresponding unit or send the fault maintenance notice of vehicle, through based on the analysis result of data dynamic to the parameter of system or device re-determine and adjust, improved train operation safety and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of train operation monitoring, and in particular to a train operation data analysis method based on self-learning. BACKGROUND

[0002] With the rapid development of global railway transportation (such as high-speed rail, urban rail transit, etc.), the safety, efficiency and service quality of train operation are facing higher requirements. At the same time, the popularization of train intelligentization and networking technology (such as train control system based on Internet of Things, vehicle-mounted sensor network, etc.) makes it possible to collect massive operation data in real time. The train can be analyzed based on real-time train operation data to ensure the operation and maintenance efficiency of the train and the travel safety of passengers.

[0003] The prior art Chinese patent CN112678035B provides a train operation data analysis method, system, server and computer readable medium. The technical solution receives train operation data from a data communication module arranged on a target train, wherein the train operation data is obtained by the data communication module from at least one subsystem included in the train automatic control system on the target train, then extracts fault information from the train operation data, determines whether the fault information is invalid fault information according to the occurrence time of the train fault represented by the fault information and the state of the target train when the train fault occurs, and responds to the received fault query instruction to display the valid fault information in the fault information through the display interface, thereby improving the efficiency of train fault troubleshooting. However, the technical solution does not construct a "data-decision-feedback" closed loop, and the technical solution extracts fault information from operation data based on fixed threshold or historical rule library, lacks self-learning ability, and cannot adapt to the influence of equipment aging and environmental changes, thereby affecting the execution efficiency of train operation and maintenance. SUMMARY

[0004] Therefore, the present application provides a train operation data analysis method based on self-learning to solve the problem that the prior art cannot process and analyze the acquired vehicle data set in real time and dynamically adjust the corresponding parameters or devices in the data analysis system based on the analysis result, resulting in low accuracy and reliability of the data analysis process, and thus low safety of train operation and execution efficiency of train operation and maintenance strategy.

[0005] To achieve the above-mentioned purpose, the present application provides a train operation data analysis method based on self-learning, comprising:

[0006] real-time data acquisition of the vehicle by a plurality of data sensors and transmission of the acquired vehicle original data set to the data processing platform through the vehicle-mounted wireless transmission system;

[0007] Preprocess the vehicle original data set through a data processing platform to obtain a processed vehicle data set, wherein the preprocessing process includes data cleaning, data parsing, and data transposition;

[0008] Statistically analyze the processed vehicle data set in terms of data basic information, category, quantity, and time node corresponding to each data, and combine the data correlation analysis to determine the quantity of vehicle normal operation data and the quantity of vehicle abnormal data;

[0009] Calculate the sum of the quantity of vehicle normal operation data and the quantity of vehicle abnormal data, and record the sum as the total quantity of vehicle data;

[0010] Calculate the ratio between the quantity of vehicle abnormal data and the total quantity of vehicle data, and record the ratio as the abnormal data proportion, determine whether the train data collection and analysis process meets the standard based on the abnormal data proportion, and generate a corresponding instruction when the standard is not met;

[0011] Determine the data collection frequency for the vehicle, the data interval value in the data cleaning, or issue a fault maintenance notice for the vehicle based on the instruction.

[0012] Further, the process of preprocessing the vehicle original data set through the data processing platform to obtain the processed vehicle data set includes:

[0013] Clean the vehicle original data set to obtain a cleaned vehicle data set;

[0014] Compile a data parsing rule table corresponding to the vehicle;

[0015] Parse the cleaned vehicle data set based on the data parsing rule table to obtain a vehicle parsed data set;

[0016] Transpose the vehicle parsed data set to obtain the processed vehicle data set.

[0017] Further, the process of determining whether the train data collection and analysis process meets the standard based on the abnormal data proportion includes:

[0018] Determine based on the comparison result of the abnormal data proportion and a preset abnormal data proportion, or re-determine whether the train data collection and analysis process meets the standard based on the comparison result of the total quantity of data in the processed vehicle data set and a critical data processing total quantity, wherein the critical data processing total quantity is the total quantity of data that the data processing platform can process in a single detection cycle;

[0019] When it is determined that the train data collection and analysis process does not meet the standard, a difference between the abnormal data proportion and the preset abnormal data proportion is determined to be the reason for not meeting the standard.

[0020] Further, the process of re-determining whether the train data collection and analysis process meets the standard based on a comparison result of the number of data in the processed vehicle data set and the critical data processing total amount includes:

[0021] The number of data in the processed vehicle data set is counted and recorded as a total amount of processed data.

[0022] Based on a comparison result of the total amount of processed data and the critical data processing total amount, it is determined whether to increase the number of shards for node-by-node processing of the total amount of processed data.

[0023] Further, the process of determining whether to increase the number of shards for node-by-node processing of the total amount of processed data based on a comparison result of the total amount of processed data and the critical data processing total amount includes:

[0024] The difference between the total amount of processed data and the critical data processing total amount is calculated, and the difference is recorded as a data processing amount difference;

[0025] Based on a comparison result of the data processing amount difference and a preset data processing amount difference, the number of shards is increased, and the increase in the number of shards is in a positive relationship with the data processing amount difference.

[0026] Further, the process of determining the reason according to the difference between the abnormal data proportion and the preset abnormal data proportion includes:

[0027] The difference between the abnormal data proportion and the preset abnormal data proportion is calculated, and the difference is recorded as a data proportion difference;

[0028] Based on a comparison result of the data proportion difference and a preset data proportion difference, the reason for the train data collection and analysis process not meeting the standard is determined;

[0029] Based on the reason, a data interval value in the data cleaning is determined, and a corresponding processing or a fault maintenance notice for the vehicle is issued based on the number of data in the vehicle original data set within a plurality of detection periods.

[0030] Further, the process of determining the data interval value in the data cleaning includes:

[0031] The number of data in the processed vehicle data set is counted and recorded as a total amount of processed data, and the number of data in the vehicle original data set is counted and recorded as a total amount of original data.

[0032] calculating a difference between the total amount of the original data and the total amount of the processed data and recording the difference as a processed data difference value;

[0033] determining the data interval value in the data cleaning based on a comparison result of the processed data difference value and a preset processed data difference value.

[0034] Further, the process of determining the corresponding processing based on the data quantity in the vehicle original data set in a plurality of detection periods comprises:

[0035] counting the original data quantity in the vehicle original data set in a plurality of historical detection periods and the original data quantity in the vehicle original data set in a current detection period and calculating an original data variance based on each original data quantity;

[0036] determining to increase the data collection frequency for the vehicle or issuing a fault maintenance notice for the vehicle based on a comparison result of the original data variance and a preset original data variance.

[0037] Further, the process of increasing the data collection frequency for the vehicle comprises:

[0038] calculating a difference between the original data variance and the preset original data variance and recording the difference as a variance difference value;

[0039] determining to increase the data collection frequency based on a comparison result of the variance difference value and a preset variance difference value, and the increase amplitude of the data collection frequency is in a positive correlation with the variance difference value.

[0040] Further, after the data collection frequency for the vehicle is increased, if the train data collection and analysis process does not meet the standard based on the comparison result of the original data variance and the preset original data variance, a notification of increasing the data collection sensor quantity is issued based on the instruction.

[0041] Compared with the prior art, the self-learning-based train operation data analysis method has the beneficial effects that the method collects real-time data of the vehicle through a plurality of data sensors and transmits the collected vehicle original data set to a data processing platform through a vehicle-mounted wireless transmission system; the data processing platform pre-processes the vehicle original data set to obtain a processed vehicle data set; the processed vehicle data set is counted in terms of data type and quantity to determine the number of vehicle normal operation data and the number of vehicle abnormal data; the proportion of abnormal data is determined based on the number of vehicle normal operation data and the number of vehicle abnormal data, and whether the train data collection and analysis process meets the standard is determined based on the proportion of abnormal data, and corresponding instructions are generated when the standard is not met; the data collection frequency of the vehicle, the data interval value in the data cleaning or the fault maintenance notice for the vehicle is determined based on the instructions. Through the analysis result of the data, it is determined whether the current train data collection and analysis process meets the normal standard, and whether the obtained vehicle operation data has a problem can be determined, and the parameters or devices in the system can be dynamically determined and adjusted to ensure that the obtained operation data is accurate and reliable, thereby ensuring the accuracy of subsequent analysis and early warning based on the vehicle data, and improving the safety and maintenance efficiency of the train operation.

[0042] Further, the application determines whether the train data collection and analysis process meets the standard according to the comparison result of the proportion of abnormal data and the preset proportion of abnormal data, and determines the reason for not meeting the standard based on the difference between the proportion of abnormal data and the preset proportion of abnormal data when it is determined that the standard is not met.

[0043] Further, the application further determines whether the train data collection and analysis process meets the standard based on the comparison result of the total amount of processed data and the critical data processing total amount, which is a secondary determination, thereby improving the accuracy of the determination of the train data collection and analysis process.

[0044] Further, the application can determine to increase the number of fragments according to the comparison result when comparing the total amount of processed data and the critical data processing total amount, and determine the number of increased fragments based on the comparison result of the quantity processing difference and the preset data processing difference, thereby improving the processing and analysis capability of the data processing platform for a large amount of real-time data, and improving the processing efficiency of the vehicle data.

[0045] Further, the application can further determine the reason why the train data collection and analysis process does not meet the standard based on the comparison result of the data proportion difference and the preset data proportion difference, and then determine the corresponding processing method according to the corresponding reason, thereby improving the accuracy of the train data collection and analysis process.

[0046] Further, the application determines that the reason for not meeting the standard is that the data preprocessing process of the vehicle original data set has a problem, and then determines the data interval value of the data cleaning in the preprocessing process according to the comparison result of the processed data difference value and the preset processed data difference value, thereby improving the screening process of the data processing platform for abnormal data, ensuring the number of data samples in the subsequent analysis process, and improving the accuracy of the early warning analysis.

[0047] Further, when the application determines that there is a problem in the data acquisition process in the current detection period based on the comparison result of the original data variance and the preset original data variance, the data acquisition frequency is further increased based on the comparison result of the variance difference value and the preset variance difference value, thereby obtaining more dense original data distribution samples, thereby improving the real-time data acquisition capability of the train, reducing data omission, and ensuring the authenticity and reliability of data analysis. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 A module schematic diagram of the system of the application based on the self-learning train operation data analysis method;

[0049] Figure 2 A flowchart of the application based on the self-learning train operation data analysis method;

[0050] Figure 3 A logic determination diagram for determining whether the train data acquisition and analysis process meets the standard based on the proportion of abnormal data;

[0051] Figure 4 A logic determination diagram for determining the reason for the train data acquisition and analysis process not meeting the standard and the corresponding processing method based on the data proportion difference. DETAILED DESCRIPTION

[0052] In order to make the purpose and advantages of the application more clear and obvious, the application will be further described below in combination with embodiments; it should be understood that the specific embodiments described herein are only used to explain the application, and do not limit the protection scope of the application.

[0053] The preferred embodiments of the application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the application, and are not intended to limit the protection scope of the application.

[0054] It should be noted that in the description of the present application, unless otherwise explicitly specified and limited, the term "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium, or internal communication of two elements. For those skilled in the art, the specific meaning of the above-mentioned term in the present application can be understood according to the specific circumstances.

[0055] The vehicle data set obtained in the embodiment is collected for a single train, and the running process analysis is performed for the single train. It can be clearly understood that the analysis process is also applicable to other trains. The data types and sources include, but are not limited to, door system, network system, air conditioning system, traction system, auxiliary system, braking system, auxiliary system, power system, etc.

[0056] Please refer to Figure 1 As shown in the figure, it is a module schematic diagram of the system of the train running data analysis method based on self-learning of the embodiment. The system comprises a data acquisition module, a vehicle-mounted wireless transmission system, a data processing platform, and an adjustment module. The data acquisition module is connected with the vehicle-mounted wireless transmission system. The data acquisition module comprises a plurality of data sensors. The real-time data of the vehicle is acquired through the plurality of data sensors, and the acquired vehicle original data set is transmitted to the data processing platform through the vehicle-mounted wireless transmission system. The data processing platform is connected with the data acquisition module. The data processing platform is used to pre-process the vehicle original data to obtain the processed vehicle data set. In addition, the data processing platform is also used to statistically analyze the data basis information, data type, data quantity, and time node corresponding to each data of the processed vehicle data set, and combine the data type correlation analysis to determine the number of vehicle normal running data and the number of vehicle abnormal data, calculate the sum of the number of vehicle normal running data and the number of vehicle abnormal data, and record the sum as the total amount of vehicle data. Calculate the ratio between the number of vehicle abnormal data and the total amount of vehicle data, and record the ratio as the proportion of abnormal data. Based on the proportion of abnormal data, it is determined whether the train data acquisition and analysis process meets the standard, and corresponding instructions are generated when it does not meet the standard. The data basis information includes vehicle model, purchase date, etc. The data correlation analysis can identify abnormal data that does not match the performance of the vehicle type according to the combination of data basis information, type, quantity, and each time node.

[0057] Specifically, in the present embodiment, the vehicle is subjected to real-time data collection by several sensors, and then the vehicle-mounted host of the vehicle transmits the vehicle raw data set through the vehicle-mounted wireless transmission system through channels such as 4G, 5G, EUHT (Enhanced Ultra High Throughput), and then transmits the data to the data processing platform through the TCP protocol. The data processing platform includes a data receiving service program developed based on the Netty framework, which receives and preliminarily processes data in a load balancing manner, including data sticking, unpacking, response processing, heartbeat reply, and CRC check to ensure data accuracy. The load balancing manner includes HTTP redirection load balancing, DNS domain name resolution load balancing, reverse proxy load balancing (such as Nginx), IP load balancing, etc. The data processing platform further preprocesses the vehicle raw data set, including data cleaning, data parsing, and data transposition. The data processing platform also includes KafKa (a distributed stream processing platform) and IoTDB (a time series database), which sends the processed vehicle data set to KafKa and stores it in IoTDB for data storage for subsequent early warning and analysis. Based on the number of vehicle normal operation data and the number of vehicle abnormal data, the train data collection and analysis process status is determined, and the reason is determined when the train data collection and analysis process is determined to be non-standard. The non-standard here refers to the proportion of vehicle abnormal data in the total amount of vehicle operation data exceeding the preset standard, wherein the total amount of vehicle operation data is the sum of the number of vehicle normal operation data and the number of vehicle abnormal data. The data correlation analysis here can distinguish between vehicle normal operation data and vehicle abnormal data by setting equipment parameter thresholds, for example, data with motor temperature < 60℃ and three-phase current deviation < 5% is considered as vehicle normal operation data, and data not within the range is considered as vehicle abnormal data. At this time, a single vehicle abnormal data will not have a significant impact on vehicle operation. Time efficiency indicators can also be set, for example, data with a plan fulfillment rate ≥ 99 and a delay time ≤ 2 minutes / 10,000 vehicle kilometers is considered as vehicle normal operation data, and data not within the range is considered as vehicle abnormal data. At this time, a single vehicle abnormal data will not have a significant impact on vehicle operation. When the reason is determined, corresponding instructions can be generated, and then the data collection frequency for the vehicle is adjusted, the data interval value in the data cleaning process is re-determined, or a maintenance work order is established and a maintenance notice for a single vehicle is issued.The embodiment can determine whether the current train data acquisition and analysis process is in line with the normal standard based on the analysis result of the data, and can determine whether the obtained vehicle operation data has a problem, and can adjust the data acquisition and processing process to ensure the accuracy of subsequent analysis and early warning based on vehicle data, thereby improving the safety and maintenance efficiency of train operation.

[0058] Referring to Figure 2 The flowchart of the train operation data analysis method based on self-learning is shown in FIG. 1.

[0059] S1: Real-time data acquisition of the vehicle is performed by a plurality of data sensors, and the collected vehicle original data set is transmitted to a data processing platform through a vehicle-mounted wireless transmission system.

[0060] S2: The vehicle original data set is preprocessed by the data processing platform to obtain a processed vehicle data set, wherein the preprocessing process includes data cleaning, data analysis and data transposition.

[0061] S3: The processed vehicle data set is statistically analyzed in terms of data basis information, category, quantity and time node corresponding to each data, and combined with data correlation analysis to determine the quantity of vehicle normal operation data and the quantity of vehicle abnormal data.

[0062] S4: The sum of the quantity of vehicle normal operation data and the quantity of vehicle abnormal data is calculated, and the sum is recorded as the total quantity of vehicle data.

[0063] S5: The ratio between the quantity of vehicle abnormal data and the total quantity of vehicle data is calculated, and the ratio is recorded as the abnormal data proportion. Based on the abnormal data proportion, it is determined whether the train data acquisition and analysis process meets the standard, and corresponding instructions are generated when the standard is not met.

[0064] S6: Based on the instructions, the data acquisition frequency of the vehicle, the data interval value in the data cleaning or the fault maintenance notice for the vehicle are determined.

[0065] Further, the process of preprocessing the vehicle original data set by the data processing platform to obtain the processed vehicle data set includes:

[0066] The vehicle original data set is cleaned to obtain a cleaned vehicle data set. A data analysis rule table corresponding to the vehicle is prepared. The cleaned vehicle data set is analyzed based on the data analysis rule table to obtain a vehicle analysis data set. The vehicle analysis data set is transposed to obtain the processed vehicle data set.

[0067] Specifically, in the present embodiment, the vehicle original data can be cleaned to eliminate the relative. The corresponding vehicle data analysis rule table can be compiled through the communication protocol. The byte offset, bit and length corresponding to the field described in the communication protocol are compiled into the corresponding data code (for example: 0001 represents the first complete byte, 00011 represents the first bit of the first byte, and 000104 represents the length of 4 bits from 0 of the first byte). Then the corresponding analysis rules are configured for the data code (for example: converted to decimal value, converted to floating point type, converted to character represented by ascii code, converted to version number information in special format, etc.). Through the above rule table, the data analysis processing can be realized to obtain the vehicle analysis data set. Then the vehicle analysis data set is transposed to obtain the processed vehicle data set. The processed vehicle data set can adapt the data format requirement in different scenes.

[0068] Please refer to Figure 3 The abnormal data proportion determination process for determining whether the train data collection and analysis process meets the standard is shown in the logic determination diagram based on the abnormal data proportion determination process for determining whether the train data collection and analysis process meets the standard. The abnormal data proportion determination process for determining whether the train data collection and analysis process meets the standard includes:

[0069] Based on the comparison result of the abnormal data proportion and the preset abnormal data proportion, or based on the comparison result of the total amount of data in the processed vehicle data set and the critical data processing total amount, it is determined whether the train data collection and analysis process meets the standard. The critical data processing total amount is the total amount of data that can be processed by the data processing platform in a single detection period. When it is determined that the train data collection and analysis process does not meet the standard, the difference between the abnormal data proportion and the preset abnormal data proportion is determined to be the reason for not meeting the standard.

[0070] Specifically, in the present embodiment, in order to make the determination process more accurate, the preset abnormal data proportion B0 can be divided into a first preset abnormal data proportion B1 and a second preset abnormal data proportion B2. The brake system data of the vehicle is analyzed. The preset abnormal data proportion standard B3 is set to 6%, B1 = B3-0.5%, and B2 = B3+0.5%. The comparison process of the abnormal data proportion B with B1 and B2 is as follows:

[0071] If the abnormal data proportion B is less than or equal to the first preset abnormal data proportion B1, it indicates that the number of current vehicle abnormal data is relatively small, and the cumulative amount of abnormal data has not reached the stage of triggering a fault warning. At this time, it can be determined that the train data collection and analysis process is in line with the standard, and no corresponding processing measures need to be taken in the detection period, which can effectively reduce the maintenance cost of the train. It can be clearly understood that B is not infinitely small at this time.

[0072] If the abnormal data proportion B is greater than the first preset abnormal data proportion B1 and less than or equal to the second preset abnormal data proportion B2, the number of vehicle abnormal data is between relatively small and relatively large. At this time, in order to further determine whether the vehicle operation process is in line with the standard, the total amount of data in the processed vehicle data set can be obtained, and then the total amount of data is compared with the critical data processing amount based on the comparison result, and the train data collection and analysis process is re-determined based on the comparison result.

[0073] If the abnormal data proportion B is greater than the second preset abnormal data proportion B2, it indicates that the number of current vehicle abnormal data is relatively large. If the fault warning threshold is set to 6.45%, it can be known that the number of vehicle abnormal data at this time exceeds the critical value of triggering a fault warning under normal circumstances. At this time, a fault warning should be issued, but no warning is issued subsequently. At this time, it can be determined that the train data collection and analysis process is not in line with the standard. The difference between the abnormal data proportion and the preset abnormal data proportion is compared with the preset value to determine the reason why the train data collection and analysis process is not in line with the standard. Based on the determined reason, a corresponding processing method is determined to adjust the train data collection and analysis process to ensure the accuracy of the data. It can be clearly understood that B is not infinitely large at this time. It can be understood that the preset abnormal data proportion B0 is not limited in the embodiment of the application. In optional embodiments, a new preset abnormal data proportion standard B31=5.5% can be set, a new first preset abnormal data proportion B11=B31-0.5%, and a new second preset abnormal data proportion B21=B31+0.5%. As long as the comparison of B with B11 and B21 can determine whether the train data collection and analysis process is in line with the standard and determine the reason when it is not in line with the standard, the specific comparison process is referred to the above content.

[0074] Further, the process of re-determining whether the train data collection and analysis process is in line with the standard based on the comparison result of the number of data in the processed vehicle data set and the critical data processing amount includes:

[0075] count the number of data in the processed vehicle data set and record it as a total amount of processed data; and determine whether to increase the number of shards for processing the total amount of processed data based on a comparison result of the total amount of processed data and the critical data processing total amount.

[0076] Specifically, in the embodiment, the comparison process of the total amount of processed data D1 and the critical data processing total amount D2 is as follows:

[0077] If the total amount of processed data D1 is less than or equal to the critical data processing total amount D2, it indicates that the number of data blocks determined by the data processing platform based on the total amount of processed data is appropriate, and the data processing platform can normally process the train data collection and analysis process based on the number of normal operation data and the number of abnormal data. If the proportion of abnormal data B is greater than the first preset proportion of abnormal data B1 and less than or equal to the second preset proportion of abnormal data B2, it is determined that the train data collection and analysis process does not meet the standard. At this time, the reason for not meeting the standard needs to be determined. It can be clearly seen that D1 is not infinitely small at this time.

[0078] If the total amount of processed data D1 is greater than the critical data processing total amount D2, it indicates that the current total amount of processed data is relatively large. At this time, tools such as Spark and Flink can be used to perform node parallel processing on the total amount of processed data, dynamically increase the number of shards of the data processing platform, and process the total amount of processed data in more cluster nodes to reduce the data processing burden of the data processing platform in a single detection period, improve the accuracy of the analysis result, and thus make the analysis of the train data collection and analysis process more accurate. For example, 1-10000 data in the processed data can be stored in shard 1, and 10001-20000 data can be stored in shard 2. The critical data processing total amount is determined according to the data processing system in the data processing platform. The data processing system can include a relational database manager (RDBM) or a distributed file system (HDFS) for processing and analyzing data. It can be clearly seen that D1 is not infinitely large at this time.

[0079] In other embodiments, the total amount of processed data can also be cut into blocks for processing. Real-time data can be divided into blocks according to time windows (such as every minute) or event quantities (such as every 1000 pieces) to avoid high single processing load.

[0080] Further, the process of determining whether to increase the number of shards for processing the total amount of processed data based on the comparison result of the total amount of processed data and the critical data processing total amount includes:

[0081] a difference between the total amount of processed data and the critical total amount of data processing is calculated, and the difference is recorded as a data processing amount difference; based on a comparison result of the data processing amount difference and a preset data processing amount difference, an increase in the number of shards is determined, and the increase in the number of shards is in a positive correlation with the data processing amount difference.

[0082] Specifically, in the present embodiment, the data processing platform uses Flink tool, Flink supports dynamic scaling (such as Kubernetes automatic scaling), and uses a distributed database, and the number of shards can be automatically adjusted according to the size of the cluster; the preset data processing amount difference W0 can be divided into a first preset data processing amount difference W1 and a second preset data processing amount difference W2, W1 = 40GB and W2 = 160GB are set, and the critical total amount of data processing in a single detection period is set to 80GB; based on the comparison of the data processing amount difference W and W1 and W2, the process is as follows:

[0083] The original number of shards is set to 3100, and the data size that can be processed by a single shard is 256MB; if the data processing amount difference W is less than or equal to the first preset data processing difference W1, and W = 40GB is set, other idle cluster nodes can be enabled, and the number of shards is increased by about 156.

[0084] If the data processing amount difference W is greater than the first preset data processing amount difference W1 and less than or equal to the second preset data processing amount difference W2, and W = 160GB is set, other idle cluster nodes can be enabled, and the number of shards is increased by about 625.

[0085] If the data processing amount difference W is greater than the second preset data processing amount difference W2, and W=180 GB is set, other idle cluster nodes can be enabled, and the number of shards is increased by about 703; it can be clearly understood that the increase in the number of shards is directly proportional to the data processing difference, the greater the data processing difference, the more the number of shards increases, and it should be noted that the increase in the number of shards is also related to the data size that a single shard can process and other related parameters, including the overall data size, storage performance, and computing requirements, and therefore the increase in the number of shards can also be set to other values. In this embodiment, for the convenience of example calculation and explanation, the data is set according to the decimal standard, 1 TB=1000 GB, and 1 GB=1000 MB. It can be understood that the preset data processing amount difference W0 and the critical data processing total amount are not specifically limited in the embodiment of the application, and in an optional embodiment, a new first preset data processing amount difference W11=35 GB, a new second preset data processing amount difference W21=155 GB, and a new critical data processing total amount of 75 GB can be set, as long as the increase in the number of shards can be determined by comparing W with W11 and W21 and combining the new critical data processing total amount, and the specific comparison process is described above.

[0086] Referring to FIG. 6, Figure 4 FIG. 6 is a logic determination diagram for determining the reason for non-compliance of the train data collection and analysis process with the standard and the corresponding processing mode based on the data proportion difference according to the embodiment. The process of determining the reason according to the difference between the abnormal data proportion and the preset abnormal data proportion includes:

[0087] The difference between the abnormal data proportion and the preset abnormal data proportion is calculated, and the difference is recorded as a data proportion difference; the reason for non-compliance of the train data collection and analysis process with the standard is determined based on the comparison result of the data proportion difference and the preset data proportion difference; the data interval value in the data cleaning is determined based on the reason, the number of data in the vehicle original data set within a plurality of detection periods is determined, and the corresponding processing or a fault maintenance notice for the vehicle is issued.

[0088] Specifically, in this embodiment, the preset data proportion difference S0 can be divided into a first preset data proportion difference S1 and a second preset data proportion difference S2, S1=0.6%, and S2=1%; the comparison process of the data proportion difference S with S1 and S2 is as follows:

[0089] If the data proportion difference S is less than or equal to a first preset data proportion difference S1, it is determined that the reason for the train data collection and analysis process not meeting the standard is that there is a problem in the data preprocessing process for the vehicle original data set. The problem in the data preprocessing process will increase the number of vehicle abnormal data excluded, and thus reduce the number of data samples provided for the subsequent early warning and maintenance analysis process, thereby affecting the accuracy of the early warning analysis. The data preprocessing process needs to be corrected, including correcting the data interval value in the data cleaning. It can be clearly understood that S is not infinitely small at this time.

[0090] If the data proportion difference S is greater than the first preset data proportion difference S1 and less than or equal to a second preset data proportion difference, it is determined that the reason for the train data collection and analysis process not meeting the standard is that there is a problem in the data collection of the vehicle or the vehicle itself has a fault. At this time, the number of vehicle original data in a plurality of detection periods needs to be obtained, and then whether the data collection frequency needs to be increased or a fault maintenance notice for the vehicle needs to be issued to solve the problem of the train data collection and analysis process not meeting the standard is determined based on the number of vehicle original data.

[0091] If the data proportion difference S is greater than the second preset data proportion difference S2, it can be directly determined that the reason for the train data collection and analysis process not meeting the standard is that the number of abnormal data generated by the vehicle itself is too large, and it is determined that the vehicle has a fault problem. A maintenance work order needs to be established based on the abnormal data, and a maintenance notice for a single vehicle needs to be issued based on the maintenance work order. It can be clearly understood that S is not infinitely large at this time. It can be understood that the preset data proportion difference S0 is not specifically limited in the embodiment of the application. In optional embodiments, a new first preset data proportion difference S11=0.7% and a new second preset data proportion difference S21=1.1% can also be set, as long as the reason for the train data collection and analysis process not meeting the standard can be determined by comparing S with S11 and S21. The specific comparison process is referred to the above content.

[0092] Further, the process of determining the data interval value in the data cleaning includes:

[0093] The number of data in the processed vehicle data set is counted and recorded as a total amount of processed data, and the number of data in the vehicle original data set is counted and recorded as a total amount of original data. The difference between the total amount of original data and the total amount of processed data is calculated and recorded as a processed data difference value. The data interval value in the data cleaning is determined based on the comparison result of the processed data difference value and a preset processed data difference value.

[0094] Specifically, in the embodiment, taking collecting vehicle-mounted temperature sensor data for data cleaning as an example, the initial interval is set to [-45℃, 145℃], if the processing data difference value is larger, it means that the total amount of processing data and the total amount of original data are more different, at this time, the numerical interval of temperature can be appropriately increased; the preset processing data difference value M0 can be divided into the first preset processing data difference value M1 and the second preset processing data difference value M2, M1 = 3GB, M2 = 5GB; the comparison process of the processing data difference value M and M1 and M2 is as follows:

[0095] If the processing data difference value M is less than or equal to the first preset processing data difference value M1, the initial interval is expanded to [-48℃, 148℃]; if the processing data difference value M is greater than the first preset processing data difference value M1 and less than or equal to the second preset processing data difference value M2, the initial interval is expanded to [-52℃, 152℃]; if the processing data difference value M is greater than the second preset processing data difference value M2, the initial interval is expanded to [-55℃, 155℃]; it can be clear that the expansion range of the initial interval is reasonable and conforms to the prior art, and the expansion range value of the initial interval can also be set to other values.

[0096] Further, the process of determining the corresponding processing based on the number of data in the vehicle original data set in a plurality of detection periods includes:

[0097] The number of original data in the vehicle original data set in a plurality of historical detection periods and the number of original data in the vehicle original data set in the current detection period are counted, and the original data variance is calculated based on each original data number; based on the comparison result of the original data variance and the preset original data variance, the data collection frequency for the vehicle is increased or the fault maintenance notice for the vehicle is sent.

[0098] Specifically, in the embodiment, taking the power system of a subway train determined to use a 1500V system for power supply, and the rated current of the traction motor being 800A as an example, the normal fluctuation interval is 24A-40A, and the preset original data variance R0 is set to 350; the comparison process of the original data variance R and the preset original data variance R0 is as follows:

[0099] If the original data variance R is less than or equal to the preset original data variance R0, it is determined that the change amplitude between each original data number corresponding to each detection period in the preset time period is relatively small, that is, the numerical size of each original data number is relatively close on the whole, and in the case that the data collection of the historical detection period is normal, it can be determined that the data collection process of the current detection period is normal, at this time, a maintenance work order can be established based on the abnormal data, and a maintenance notice for the vehicle can be sent based on the maintenance work order, it can be clear that R is not infinitely small at this time.

[0100] If the original data variance R is greater than the preset original data variance R0, it is determined that the change range between the original data quantities corresponding to each detection period in the preset time period is relatively large, that is, the numerical size of the original data quantities as a whole is relatively discrete, and it is determined that the data acquisition process in the current detection period may have a problem, for example, data is missing when data acquisition is performed in a certain detection period, thereby causing the original data quantities collected in each detection period to have a large change degree, at this time, the data acquisition frequency of the vehicle-mounted wireless transmission system for the vehicle can be adjusted, and a more intensive original data distribution sample can be provided by increasing the data acquisition frequency to avoid data missing due to too large acquisition time interval. It can be clearly understood that the preset original data variance R0 is not specifically limited in the embodiment of the application, and in an optional embodiment, the preset original data variance R01 can be set to 355, as long as the comparison between R and R01 can determine that the data acquisition frequency for the vehicle is increased or a fault maintenance notice for the vehicle is issued, and the specific comparison process is described above.

[0101] Further, the process of increasing the data acquisition frequency for the vehicle includes:

[0102] The difference between the original data variance and the preset original data variance is calculated and recorded as a variance difference; and the comparison result between the variance difference and the preset variance difference is used to determine that the data acquisition frequency is increased, and the increase amplitude of the data acquisition frequency is in a positive relationship with the variance difference.

[0103] Specifically, in this embodiment, taking the motor current acquisition frequency in the power system in the normal state as an example, the initial motor current acquisition frequency is set to 200 Hz; the preset variance difference H0 can be divided into a first preset variance difference H1 and a second preset variance difference H2, and H1 is set to 20 and H2 is set to 30; the comparison process between the variance difference H and H1 and H2 is specifically as follows:

[0104] If the variance difference H is less than or equal to a first preset variance difference H1, the data acquisition frequency is adjusted to 1.5 times of the initial value by using a first frequency adjustment coefficient, and it can be understood that H is not infinitely small at this time; if the variance difference H is greater than the first preset variance difference H1 and less than or equal to a second preset variance difference H2, the data acquisition frequency is adjusted to 2 times of the initial value by using a second frequency adjustment coefficient; if the variance difference H is greater than the second preset variance difference H2, the data acquisition frequency is adjusted to 2.5 times of the initial value by using a third frequency adjustment coefficient, and it can be understood that H is not infinitely large at this time; it should be noted that the increase rate of the data acquisition frequency can be set to other values according to the initial value, and the increase rate of the data acquisition frequency meets the data acquisition of the vehicle. It can be understood that the preset variance difference H0 is not specifically limited in the embodiment of the application, and in the optional embodiment, a new first preset variance difference H11=22 and a new second preset variance difference H21=32 can also be set, as long as the increase rate of the data acquisition frequency can be determined by comparing H with H11 and H21, and the specific comparison process is referred to the above content.

[0105] Further, after the data acquisition frequency for the vehicle is increased, if it is determined again based on the comparison result of the original data variance and the preset original data variance that the data acquisition and analysis process for the train does not meet the standard, a notification is sent to increase the number of data acquisition sensors based on the instruction.

[0106] Specifically, in the embodiment, after the data acquisition frequency for the vehicle is increased, if it is still determined that the original data variance R is greater than the preset original data variance R0, a corresponding instruction can be generated by the data processing platform at this time, and the maintenance personnel increases the number of data acquisition sensors on the vehicle or starts the redundant number of data acquisition sensors on the vehicle according to the notification sent by the adjustment module based on the instruction.

[0107] The technical solutions of the application have been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the application, and the technical solutions after the changes or replacements will fall within the protection scope of the application.

[0108] The above description is only the preferred embodiments of the application and is not used to limit the application; for those skilled in the art, the application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.

Claims

1. A train operation data analysis method based on self-learning, characterized in that, include: The vehicle collects real-time data through several data sensors and transmits the collected raw vehicle dataset to the data processing platform via an in-vehicle wireless transmission system. The original vehicle dataset is preprocessed using a data processing platform to obtain a processed vehicle dataset. The preprocessing process includes data cleaning, data parsing, and data transposition. The processed vehicle dataset is statistically analyzed for basic data information, categories, quantities, and corresponding time points, and combined with data correlation analysis to determine the number of normal vehicle data and the number of abnormal vehicle data. Calculate the sum of the number of normal vehicle data and the number of abnormal vehicle data, and record the sum as the total vehicle data. Calculate the ratio between the number of abnormal vehicle data and the total amount of vehicle data, and record the ratio as the abnormal data ratio. Based on the abnormal data ratio, determine whether the train data collection and analysis process meets the standard, and generate corresponding instructions when it does not meet the standard. Based on the instructions, determine the data collection frequency for the vehicle, the data interval value during data cleaning, or issue a fault repair notification for the vehicle. The process of determining whether the train data collection and analysis process meets the standards based on the aforementioned abnormal data ratio includes: The judgment is made based on the comparison result of the abnormal data ratio and the preset abnormal data ratio, or, based on the comparison result of the total amount of data in the processed vehicle dataset and the critical data processing total amount, the judgment is re-determined on whether the train data collection and analysis process meets the standard, wherein the critical data processing total amount is the total amount of data that the data processing platform can process in a single detection cycle. When it is determined that the train data collection and analysis process does not meet the standards, the reason for non-compliance is determined based on the difference between the abnormal data ratio and the preset abnormal data ratio. The process of re-determining whether the train data acquisition and analysis process meets the standards based on the comparison between the amount of data in the processed vehicle dataset and the total critical data processing volume includes: The total number of data points in the processed vehicle dataset is counted and recorded as the total processed data. Based on the comparison between the total amount of data processed and the critical total amount of data processed, it is determined whether to increase the number of shards for node-based processing of the total amount of data processed.

2. The train operation data analysis method based on self-learning according to claim 1, characterized in that, The process of preprocessing the original vehicle dataset using the data processing platform to obtain the processed vehicle dataset includes: The original vehicle dataset is cleaned to obtain a cleaned vehicle dataset. Compile a data parsing rule table corresponding to the vehicle; The cleaned vehicle dataset is parsed based on the data parsing rule table to obtain the vehicle parsing dataset. The vehicle parsing dataset is transposed to obtain the processed vehicle dataset.

3. The train operation data analysis method based on self-learning according to claim 1, characterized in that, The process of determining whether to increase the number of shards for node-based processing of the total processed data based on the comparison between the total processed data and the critical total processed data includes: Calculate the difference between the total amount of data processed and the critical total amount of data processed, and record the difference as the data processing volume difference. The number of shards is increased based on the comparison between the difference in data processing volume and the preset difference in data processing volume. The increase in the number of shards is directly proportional to the difference in data processing volume.

4. The train operation data analysis method based on self-learning according to claim 1, characterized in that, The process of determining the cause based on the difference between the percentage of abnormal data and the preset percentage of abnormal data includes: Calculate the difference between the percentage of abnormal data and the preset percentage of abnormal data, and record the difference as the data percentage difference; Based on the comparison results between the data ratio difference and the preset data ratio difference, the reasons for the non-compliance of the train data collection and analysis process with the standard are determined. Based on the aforementioned reasons, determine the data range values ​​during data cleaning, and based on the amount of data in the vehicle's original dataset within several detection cycles, determine the corresponding processing or issue a fault repair notification for the vehicle.

5. The train operation data analysis method based on self-learning according to claim 4, characterized in that, The process of determining the data interval values ​​in the data cleaning process includes: The number of data points in the processed vehicle dataset is counted and recorded as the total processed data, and the number of data points in the original vehicle dataset is counted and recorded as the total original data. Calculate the difference between the total amount of original data and the total amount of processed data, and record the difference as the processed data difference. The data interval value in the data cleaning process is determined based on the comparison result between the processed data difference and the preset processed data difference.

6. The train operation data analysis method based on self-learning according to claim 4, characterized in that, The process of determining the corresponding processing based on the amount of data in the original vehicle dataset within several detection periods includes: The number of raw data points in the vehicle raw datasets for several historical detection periods and the number of raw data points in the vehicle raw datasets for the current detection period are statistically analyzed, and the variance of the raw data is calculated based on the number of raw data points. Based on the comparison result between the original data variance and the preset original data variance, determine whether to increase the data collection frequency for the vehicle or issue a fault repair notification for the vehicle.

7. The train operation data analysis method based on self-learning according to claim 6, characterized in that, The process of increasing the data acquisition frequency for the vehicle includes: Calculate the difference between the original data variance and the preset original data variance and record it as the variance difference value; Based on the comparison result between the variance difference and the preset variance difference, the data acquisition frequency is increased, and the increase in the data acquisition frequency is directly proportional to the variance difference.

8. The train operation data analysis method based on self-learning according to claim 7, characterized in that, After increasing the data acquisition frequency for the vehicle, if the comparison between the original data variance and the preset original data variance determines that the train data acquisition and analysis process does not meet the standard, a notification to increase the number of data acquisition sensors is issued based on the instruction.

Citation Information

Patent Citations

  • Train operation data analysis methods, systems, servers, and computer-readable media

    CN112678035B

  • On-board diagnosis system data-based traffic condition analysis and early warning method

    CN108769104A

  • Vehicle accessory abnormal information automatic detection processing method and related equipment thereof

    CN115310542A