A method and system for governing train operation log data

By performing binary analysis and structured processing of train driving log data, combined with the hierarchical storage and multi-dimensional modeling of the big data platform, the problem of inefficient management of urban rail driving signal data is solved, and data support and intelligent operation and maintenance of urban rail operation and passenger services are realized.

CN115391302BActive Publication Date: 2025-07-18TRAFFIC CONTROL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110705828.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-25
Filing Date
2021-06-24
Publication Date
2025-07-18
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

The lack of systematic solutions in the prior art to comprehensively manage and analyze urban rail driving signal data, especially in terms of data collection, integration and unified modeling, resulting in inefficient data governance and analysis and inability to effectively support smart operations and passenger services.

Method used

By obtaining train driving log data, performing binary data analysis and structured processing, using the big data platform for layered storage and multi-dimensional modeling, and combining actual application scenarios for data analysis and mining, realizing effective governance of driving data.

Benefits of technology

It has improved the data support capabilities for urban rail operation, improved the quality of passenger service and intelligent station operation and maintenance, and realized efficient utilization and intelligent scheduling of driving data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391302B_ABST
    Figure CN115391302B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for governing train operation log data, including: acquiring train operation log data; performing binary data parsing and data structuring on the train operation log data to obtain processed data; importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; combining the data governance result with a preset actual application scenario and outputting data analysis and data mining results. The present invention realizes data support for the operation of urban rail, the tracking and improvement of passenger service quality, and the intelligent operation and maintenance of stations by proposing a data governance method for analyzing operation logs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rail transit big data processing, and particularly to a method and system for governing train operation log data. Background Art

[0002] With the popularization of smart city rail transit, the comprehensive system including smart passenger services, smart operation organizations, and smart equipment control has higher and higher requirements for the big data to be processed, and the key point is the big data governance technology.

[0003] During the operation of urban traffic, a large amount of multi-train signal data is generated. From the source, it is divided into data such as yard signal equipment, on-vehicle signal equipment, trackside signal equipment, gateway equipment, and signal system application log records; from the data structure, it is divided into structured, semi-structured, and unstructured data. The log data of the train signal system generally generates new records every 200 ms, with a huge amount of data and strong heterogeneity in data structure and business relationships. If big data analysis is to be performed on the log data of the train signal system, it is necessary to centrally collect, integrate, and unify the heterogeneous data of train operations, uniformly model the data from each source of the train, and establish multi-dimensional data and business models related to the train, and then construct a label and index system for urban rail train operations for the business model. Currently, the governance of train signal data in the urban rail field is in the initial data collection stage, lacking a systematic solution for comprehensive governance and analysis of train data, as well as relevant storage and architecture solutions.

[0004] Therefore, a new systematic method for governing urban rail data needs to be proposed. Summary of the Invention

[0005] The present invention provides a method and system for governing train operation log data to solve the defects existing in the prior art.

[0006] In a first aspect, the present invention provides a method for governing train operation log data, including:

[0007] Obtaining train operation log data;

[0008] Performing binary data parsing and data structuring on the train operation log data to obtain processed data;

[0009] Importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result;

[0010] Combining the data governance result with a preset actual application scenario to output data analysis and data mining results.

[0011] In one embodiment, the train log data is subjected to binary data parsing and data structuring to obtain processed data, including:

[0012] Read the latest data row in the train driving log data in real time;

[0013] Sending binary data of the latest data row to a distributed publish-subscribe messaging system;

[0014] The log processing program Flink reads the distributed publish-subscribe message system in real time and performs parsing and data structuring;

[0015] Real-time analysis is injected into the data warehouse tool HIVE and Hadoop distributed file system HDFS to obtain tabular text data.

[0016] In one embodiment, the processed data is imported into a big data platform for hierarchical storage processing, and a preset urban rail data multi-dimensional modeling model is used to analyze the processed data to obtain data governance results, including:

[0017] storing the processed data and performing initial calculation processing to obtain calculated data;

[0018] Divide the data interval according to the train stop data points and train departure data points in the calculated data, determine whether the end change is completed according to the head code number change data point, and check the first stop point in the reverse direction before the end change completion point to determine the stop time as the end change start time;

[0019] Sorting the data of the processing routes in the calculated data to obtain a sorting result;

[0020] Parsing the binary route electronic map in the calculated data, and calculating the link list containing the route in each route based on the binary route electronic map;

[0021] Based on the data interval and the sorting result, the TOP1 data in the route is searched, the TOP1 data is associated with the train data in the driving record to obtain associated data, the route length is calculated based on the TOP1 data and the link list, and the average speed of the train in the route is calculated from the associated data and the route length;

[0022] Generate driving fact key point dimensions and driving index labels according to the end-change start time and the average speed;

[0023] Based on the driving fact dimension and the driving indicator label, the data governance result is output.

[0024] In one embodiment, the train operation fact dimension and the train operation index labels include parking, departure, route setting, completion of route setting, and end-changing completion.

[0025] In one embodiment, the combination of the data governance result with a preset actual application scenario to output data analysis and data mining results includes:

[0026] Generating a data mining report through data mining of train arrival and departure delays;

[0027] Outputting the core points for optimizing the turnaround index through analysis of turnaround data.

[0028] In one embodiment, generating a data mining report through data mining of train arrival and departure delays includes:

[0029] Obtaining the train operation logs and planned operation diagrams on the main line, analyzing the train arrival and departure delay data, and obtaining the characteristics of the arrival and departure delay data;

[0030] Based on the characteristics of the arrival and departure delay data, analyzing the influencing factors and solutions for train operation, and outputting a data mining report.

[0031] In one embodiment, outputting the core points for optimizing the turnaround index through analysis of turnaround data includes:

[0032] Analyzing the relevant stages included in train turnaround, and calculating the label indicators for each stage;

[0033] Calculating the classic time used for the core index stage;

[0034] Comparing the time used for each stage and determining the key stages affecting turnaround;

[0035] Conducting multi-dimensional analysis on each stage, mining the core elements of turnaround, and obtaining the core operation elements.

[0036] In a second aspect, the present invention also provides a train operation log data governance system, including:

[0037] An acquisition module for acquiring train operation log data;

[0038] A first processing module for parsing the binary data of the train operation log data and structuring the data to obtain processed data;

[0039] A second processing module for importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result;

[0040] A third processing module for combining the data governance result with a preset actual application scenario to output data analysis and data mining results.

[0041] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned train operation log data governance methods are implemented.

[0042] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned train operation log data governance methods are implemented.

[0043] The train operation log data governance method and system provided by the present invention realize data support for the operation of urban rail transit, the tracking and improvement of passenger service quality, and the intelligent operation and maintenance of stations by proposing a data governance method for analyzing train operation logs. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 is a schematic flowchart of the train operation log data governance method provided by the present invention;

[0046] Figure 2 is a comparison flowchart between the solution of the present invention and the traditional log processing solution;

[0047] Figure 3 is a comparison schematic diagram of the data format parsing formats provided by the present invention;

[0048] Figure 4 is a specific description schematic diagram of the new data structure provided by the present invention;

[0049] Figure 5 is a comparison schematic diagram of the data storage layer provided by the present invention;

[0050] Figure 6 is a data governance process diagram provided by the present invention;

[0051] Figure 7 is a complete flowchart of the calculation and data generation scheme for relevant dimensions provided by the present invention;

[0052] Figure 8 is a multi-dimensional modeling scheme schematic diagram provided by the present invention;

[0053] Figure 9 It is a schematic diagram of mining and analyzing the data of train arrival and departure delays provided by the present invention;

[0054] Figure 10 It is a comparison chart of the time used in the reverse operation stage provided by the present invention:

[0055] Figure 11 It is a schematic diagram of the structure of the train operation log data governance system provided by the present invention;

[0056] Figure 12 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] In view of the problems existing in the prior art, the present invention proposes a method for governing train operation log data based on big data technology, Figure 1 It is a schematic flowchart of the method for governing train operation log data provided by the present invention, as Figure 1 shown, including:

[0059] S1. Obtain train operation log data;

[0060] S2. Perform binary data parsing and data structuring on the train operation log data to obtain processed data;

[0061] S3. Import the processed data into a big data platform for hierarchical storage processing, and analyze the processed data by using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result;

[0062] S4. Combine the data governance result with a preset actual application scenario, and output data analysis and data mining results.

[0063] Specifically, the method for governing train operation log data for urban rail data proposed by the present invention includes the following aspects:

[0064] 1) Obtain train operation log data of a rail transit line, usually including historical and real-time ATP data and VOBC operation logs;

[0065] 2) Process historical and real-time ATP (Automatic Train Protection) data and VOBC (vehicle on-board controller) train operation log data, perform binary data parsing and data structuring;

[0066] 3) Import into the big data platform, and perform data modeling and train operation data analysis based on the data hierarchy of ODS -> DWD -> DWS -> ADS;

[0067] 4) Implement data governance, and perform data analysis and data mining on the governed data. Realize the value of train operation data and provide data support for the intelligent service of urban rail transit.

[0068] The train operation log data governance method proposed by the present invention solves the problems of the acquisition, aggregation, parsing, and injection of binary log data of train operation signal systems such as ATP and VOBC. Through multi-dimensional modeling and big data technology support, it realizes the combination with train operation log data, establishes corresponding data models, and more effectively utilizes the log data of train operation signals such as ATP and VOBC. By adding events such as early and late train arrival analysis and train turnaround analysis in the multi-dimensional model establishment, it improves the ability of intelligent dispatching service and intelligent urban rail.

[0069] Based on the above embodiments, step S1 in the method includes:

[0070] Read the latest data row in the train operation log data in real time;

[0071] Send the binary data of the latest data row to the distributed publish-subscribe message system;

[0072] The log processing program Flink reads the distributed publish-subscribe message system in real time and performs parsing and data structuring;

[0073] Parse and inject into the data warehouse tools HIVE and the Hadoop distributed file system HDFS in real time to obtain tabular text data.

[0074] Specifically, the preparation of the train operation ATP data and VOBC data in the present invention is realized through the following steps:

[0075] First of all, it is clear that two technical points need to be broken through in the data preparation stage. First, the data in the urban rail industry are all binary data, which is very unfriendly to data analysis, so that data analysis work cannot be carried out at all; second, the problem of the way of data injection into big data. The traditional data injection big data solution is based on the log file method. The process needs to copy and parse the file, and then import it into the data warehouse tool HIVE and HDFS (Hadoop Distributed File System, based on Hadoop Distributed File System) big data platform, which makes the solution inefficient and the real-time performance is very poor. The present invention realizes the functions of real-time reading and real-time parsing of logs, separates the reading log and the parsing log from the structure, is not invasive to the traffic signal system, and the overall solution is very progressive in real-time and isolation protection. Figure 2 The figure is a flow chart comparing the solution of the present invention and the traditional log processing solution.

[0076] In addition, the present invention optimizes the data structure of the parsing result and realizes structured storage of complex data, wherein the complex data is the driving data containing array or object type data. In particular, the present invention adopts the json data structure to store nested data and object type data, such as Figure 3 shown.

[0077] The present invention also adopts a new Flink-based log processing program to convert binary driving log data into required tabular text data to support data operations including aggregation, combination, sorting and grouping. The specific description format is as follows: Figure 4 shown.

[0078] In the stage of preparing driving ATP and VOBC data, the present invention connects different professional data sources through data interface and ETL technology, realizes remote data collection, and realizes binary data conversion of driving signals to form a storage method for signal data and urban rail electronic map data.

[0079] Based on any of the above embodiments, step S3 in the method includes:

[0080] storing the processed data and performing initial calculation processing to obtain calculated data;

[0081] Divide the data interval according to the train stop data points and train departure data points in the calculated data, determine whether the end change is completed according to the head code number change data point, and check the first stop point in the reverse direction before the end change completion point to determine the stop time as the end change start time;

[0082] Sorting the data of the processing routes in the calculated data to obtain a sorting result;

[0083] Parse the binary line electronic map in the calculated data, and calculate the link list of each route containing the route based on the binary line electronic map;

[0084] Based on the data interval and the sorting result, search for the TOP1 data in the processed route, associate the TOP1 data with the train data in the train record to obtain associated data, calculate the route length based on the TOP1 data and the link list, and calculate the average speed of the train running in the route from the associated data and the route length;

[0085] Generate key dimensions of train operation facts and train operation index labels according to the end-changing start time and the average speed;

[0086] Output the data governance result based on the train operation fact dimensions and the train operation index labels.

[0087] Among them, the train operation fact dimensions and the train operation index labels include stop, departure, route processing, completion of route processing, and end-changing completion.

[0088] Specifically, the data modeling and data storage solutions proposed in the present invention are established based on the hierarchical scheme of the standard big data governance model: the source-attached data layer (ODS), the data detail layer (DWD), the data summary layer (DWS / DWT), the data application layer (ADS), and the dimension data layer (DIM).

[0089] The usage scenarios of each layer are as follows:

[0090] The ODS layer is the layer closest to the data in the data source. The data in the data source is extracted, cleaned, and transmitted and loaded into this layer. Generally, the data in this layer is the original data prepared for train operation ATP and VOBC. In the solution of the present invention, calculations will be performed on the original data, such as calculating the distance between stations and sections to facilitate calculating the average train operation speed, etc.;

[0091] The DWD layer and the DIM layer are key processes in the entire data governance process. The key points of the present invention are the dimension design of the train operation-related DIM dimension layer and the calculation scheme and calculation method during the process of data from the ODS layer to the DWD layer;

[0092] The functions of the DWS / DWT layer include: calculating urban rail indicators and preliminary governance of some data. In addition to supporting the standard indicators within the urban rail industry, the present invention also supports detailed time analysis of train operation reversals;

[0093] The ADS layer is the implementation of objectifying and aggregating indicators for data on top of the DWS layer, forming a wide table of train operation data.

[0094] It should be noted that, for exampleFigure 5 As shown in the figure, the present invention has the following characteristics when storing data:

[0095] In the ODS layer, some data calculation functions are performed on the original data, such as calculating the distance between stations to facilitate the more rapid calculation of the average train speed later;

[0096] In the ADS layer, Clickhouse is used to store the final analysis data, enabling fast multi-dimensional analysis, supporting ad-hoc analysis and distributed data query;

[0097] In data application, various basic data reports of arrival and departure delays are formed to comprehensively understand the operation situation; the train delay data in each dimension is mined, including finding the trains with the most delays, the stations with the most delays, and the trains with the most delays; the reasons for delays are mined based on dimension analysis; the duration of each operation is calculated, the train speed and the duration of handling related operations are calculated; the key time-consuming operations are analyzed, and the data at the turning-back key points is mined to find the core stage indicators of turning-back.

[0098] Data modeling is adopted in the entire process of train operation log data governance and train data mining. Similarly, data modeling is also adopted in the traditional solution. The present invention and the traditional solution both adopt the processing process as shown in Figure 6 the figure, Figure 6 The main work content shown in the figure is divided into three parts. Among them, the two steps of data modeling and data index labeling are to sort out the business process, distinguish the relationship mapping between facts, dimension data and business objects, and the main work is the writing of ETL scripts. The differences from the traditional data governance solution are reflected in the following two points:

[0099] 1. Dimension modeling, including sorting out the business and establishing the data domain, and dividing facts and dimension data;

[0100] 2. The establishment of index labels, including the process of label analysis and establishment of train operation data.

[0101] The traditional dimension modeling scheme includes the following dimensions: time, station, equipment, vehicle, etc. According to the characteristics of urban rail transit train operation related data, the train operation dimensions and indicators designed by the present invention include the following: parking, departure, route handling, completion of route handling, and completion of end change. The finally established data dimension scheme is shown in Table 1.

[0102] Table 1

[0103]

[0104] Based on the above DIM layer design, the present invention proposes a calculation and data generation scheme for related dimensions, and the implementation steps are as shown in Figure 7 the figure:

[0105] Input the original driving data in the ODS layer to obtain the calculated data. Divide the data interval according to the train stop data points and train departure data points in the calculated data. Determine whether the end change is completed according to the head code number change data point. Before the end change completion point, check the first stop point in the reverse direction to determine the stop time as the end change start time.

[0106] At the same time, the data of the route handling in the calculated data are sorted to obtain the sorting result, and the binary line electronic map in the calculated data is parsed, and the link list containing the route in each route is calculated based on the binary line electronic map;

[0107] Based on the data interval and the sorting result, the TOP1 data in the route is searched. The TOP1 data here is the data that ranks first in the route after calculation. Then the TOP1 data is associated with the train data in the driving record to obtain the associated data. Then the route length is calculated from the TOP1 data and the link list. Furthermore, the average speed of the train in the route is calculated from the associated data and the route length.

[0108] The key point dimensions of driving facts and driving indicator labels are comprehensively generated according to the start time of the end change and the average speed; finally, based on the key point dimensions of driving facts and driving indicator labels, the data governance results are output, and the relevant dimensions and indicator data of the completed train are output.

[0109] The calculation process finally forms Figure 8 The multi-dimensional modeling scheme shown completes the driving log modeling process.

[0110] The present invention realizes the combination with driving log data through multi-dimensional modeling and big data technology support, establishes a corresponding data model, and can more effectively utilize the log data of driving signals such as ATP and VOBC.

[0111] Based on any of the above embodiments, step S4 in the method includes:

[0112] Generate data mining reports through early and late driving data mining;

[0113] Through the analysis of the return data, the core points of the optimized return indicators are output.

[0114] Among them, through the early and late driving data mining, a data mining report is generated, including:

[0115] Obtain the train operation log and planned operation diagram on the main line, analyze the train early and late data, and obtain the characteristics of the early and late data;

[0116] Based on the characteristics of the early and late train data, the factors affecting train operation and the solutions are analyzed, and a data mining report is output.

[0117] Among them, through the analysis of turnback data, the core points of the optimized turnback index are output, including:

[0118] Analyze the relevant stages included in the train turnback, and calculate the label indexes of each stage;

[0119] Calculate the classical time used for the core index stage;

[0120] Compare the time used for each stage and determine the key stages affecting the turnback;

[0121] Conduct multi-dimensional analysis on each stage, excavate the core elements of the turnback, and obtain the core operation elements.

[0122] Specifically, at the practical application level, the present invention proposes two train operation data usage schemes: train arrival and departure time data mining and train turnback data analysis.

[0123] Train arrival and departure time data mining includes analyzing the arrival and departure time data of trains by obtaining the operation logs and planned operation diagrams of trains on the main line, obtaining the characteristics of arrival and departure time data, and then analyzing the factors and solutions affecting train operation during urban rail operation, and generating a report, as Figure 9 shown.

[0124] The turnback data analysis aims to calculate and compare various indexes through the governance results of train operation log data, and find the core points for optimizing the turnback indexes. For example: the length of each action time, train speed, operation duration, etc., to find and optimize the indexes of the core turnback stage, as Figure 10 shown. The specific implementation steps are as follows:

[0125] 1) Analyze the stages related to the turnback and calculate the label indexes of each stage;

[0126] 2) Calculate the classical time used for the core index stage, including:

[0127] a) Generate the operation and train operation time related to the turnback through the evaluation software;

[0128] b) Calculate the median of the long-term operation time;

[0129] 3) Compare the time used for each stage and find the stage affecting the turnback;

[0130] 4) Conduct multi-dimensional analysis on each stage to excavate the core elements of the turnback;

[0131] 5) Obtain the core operation elements and propose improvement solutions to facilitate the better scheduling and capacity improvement of the operation company.

[0132] The present invention realizes intelligent service scheduling and intelligent urban rail solutions by adding events such as morning and evening arrival analysis and train turnaround analysis in the establishment of a multi-dimensional model.

[0133] The train operation log data governance system provided by the present invention will be described below. The train operation log data governance system described below can be correspondingly referred to the train operation log data governance method described above.

[0134] Figure 11 is a schematic structural diagram of the train operation log data governance system provided by the present invention, as Figure 11 shown, including: an acquisition module 1101, a first processing module 1102, a second processing module 1103, and a third processing module 1104, where:

[0135] The acquisition module 1101 is used to acquire train operation log data; the first processing module 1102 is used to perform binary data parsing and data structuring on the train operation log data to obtain processed data; the second processing module 1103 is used to import the processed data into a big data platform for hierarchical storage processing, and analyze the processed data by using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; the third processing module 1104 is used to combine the data governance result with a preset actual application scenario and output data analysis and data mining results.

[0136] The present invention realizes data support for the operation of urban rail, the tracking and improvement of passenger service quality, and the intelligent operation and maintenance of stations by proposing a data governance system for analyzing train operation log data.

[0137] Figure 12 illustrates a schematic structural diagram of an electronic device, as Figure 12 shown, the electronic device may include: a processor 1210, a communication interface 1220, a memory 1230, and a communication bus 1240. Among them, the processor 1210, the communication interface 1220, and the memory 1230 complete mutual communication through the communication bus 1240. The processor 1210 can call logical instructions in the memory 1230 to execute the train operation log data governance method, which includes: acquiring train operation log data; performing binary data parsing and data structuring on the train operation log data to obtain processed data; importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data by using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; combining the data governance result with a preset actual application scenario and outputting data analysis and data mining results.

[0138] In addition, when the logical instructions in the above-mentioned memory 1230 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0139] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the train operation log data governance method provided by the above-mentioned various methods. The method includes: obtaining train operation log data; performing binary data parsing and data structuring on the train operation log data to obtain processed data; importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; combining the data governance result with a preset actual application scenario to output data analysis and data mining results.

[0140] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the train operation log data governance method provided by the above-mentioned various methods. The method includes: obtaining train operation log data; performing binary data parsing and data structuring on the train operation log data to obtain processed data; importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; combining the data governance result with a preset actual application scenario to output data analysis and data mining results.

[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for governing train operation log data, characterized in that, Including: Obtain train operation log data; Perform binary data parsing and data structuring on the train operation log data to obtain processed data; Import the processed data into a big data platform for hierarchical storage processing, and analyze the processed data using a preset multi-dimensional modeling model for urban rail data to obtain data governance results; Combine the data governance results with a preset actual application scenario and output data analysis and data mining results; Import the processed data into a big data platform for hierarchical storage processing, and analyze the processed data using a preset multi-dimensional modeling model for urban rail data to obtain data governance results, including: Store the processed data and perform initial calculation processing to obtain calculated data; Divide data intervals according to train stop data points and train departure data points in the calculated data, determine whether end-changing is completed based on head code number change data points, and backtrack to the first stop point in the reverse direction before the end-changing completion point to determine the stop time as the start time of end-changing; Sort the data for route handling in the calculated data to obtain a sorting result; Parse the binary line electronic map in the calculated data, and calculate the link list of the routes included in each route based on the binary line electronic map; Based on the data interval and the sorting result, find the TOP1 data in the processed routes, associate the TOP1 data with the train data in the train operation record to obtain associated data, calculate the route length based on the TOP1 data and the link list, and calculate the average speed of the train running within the route from the associated data and the route length; Generate a driving fact dimension and driving index labels based on the start time of end-changing and the average speed; Output the data governance results based on the driving fact dimension and the driving index labels; The driving fact dimension and the driving index labels include stop, departure, route handling, completion of route handling, and end-changing completion.

2. The train operation log data governance method according to claim 1, characterized in that Perform binary data parsing and data structuring on the train operation log data to obtain processed data, including: Read the latest data row in the train operation log data in real time; Send the binary data of the latest data row to a distributed publish-subscribe message system; The log processing program Flink reads the distributed publish-subscribe message system in real time and performs parsing and data structuring; Parse and inject into the data warehouse tools HIVE and the Hadoop Distributed File System HDFS in real time to obtain tabular text data.

3. The train operation log data governance method according to claim 1, wherein The combination of the data governance results with a preset actual application scenario and output of data analysis and data mining results includes: Generate a data mining report through train arrival / departure delay data mining; Output the core points for optimizing the reverse operation index through reverse operation data analysis.

4. The train operation log data governance method according to claim 3, wherein Generate a data mining report through train arrival / departure delay data mining, including: Obtain the train operation logs and planned operation diagrams on the main line, analyze the train arrival / departure delay data, and obtain the characteristics of arrival / departure delay data; Based on the characteristics of arrival / departure delay data, analyze the influencing factors and solutions for train operation, and output a data mining report.

5. The train operation log data governance method according to claim 3, wherein Output the core points of optimizing the turnback index through turnback data analysis, including: Analyze the relevant stages included in the train turnback and calculate the label indexes of each stage; Calculate the classical time used for the core index stage; Compare the time used for each stage and determine the key stage affecting the turnback; Conduct multi-dimensional analysis on each stage, excavate the core elements of the turnback, and obtain the core operation elements.

6. A train operation log data governance system applying the train operation log data governance method according to any one of claims 1-5, characterized in that, Including: An acquisition module for acquiring train operation log data; A first processing module for parsing binary data and structuring the train operation log data to obtain processed data; A second processing module for importing the processed data into a big data platform for hierarchical storage processing, and analyzing the processed data using a preset multi-dimensional modeling model for urban rail data to obtain a data governance result; A third processing module for combining the data governance result with a preset actual application scenario and outputting data analysis and data mining results.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the train operation log data governance method according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the train operation log data governance method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Big data fusion analysis method applied to massive logs of automatic train control system

    CN107256219A

  • Log management system and operation method thereof

    CN111190876A