Virtual marshalling car-mounted log data processing method, platform, device and storage medium
By monitoring remote cloud storage and automatically storing and parsing virtual trainset vehicle log data, a horizontally layered data model was established, which solved the problem of low processing efficiency of virtual trainset vehicle log data, realized automated pipeline operation and real-time analysis of data, and improved operational service capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, the processing efficiency of virtual train vehicle log data is low, and there is a lack of systematic solutions. The main problem is that data processing relies on manual operation, which is inefficient and cannot meet the needs of real-time analysis.
By monitoring remote cloud storage, raw log data is automatically stored and parsed into structured data, and data is processed according to business domains. Distributed file systems and big data technologies are used to automate the acquisition, parsing, and analysis of data, establish a horizontally layered data model, and realize automated pipeline operations for data.
It improves the processing efficiency of virtual trainset vehicle log data, realizes automated data acquisition, parsing and analysis, meets the needs of real-time data processing and business applications, and enhances operational service capabilities.
Smart Images

Figure CN116719789B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of intelligent rail transit, and particularly relates to a virtual marshalling vehicle-mounted log data processing method, platform, device and storage medium. BACKGROUND
[0002] With the increasing number of rail transit passengers and the continuous expansion of the line network, there is a dynamic change in passenger flow in the urban transportation system, and the uneven spatiotemporal distribution of urban rail transit passenger flow demand leads to a more prominent mismatch between the capacity provided by local areas and different time periods and the passenger flow demand. Virtual marshalling aims to achieve better precision matching of passenger flow and train flow, run a large train set in peak hours to improve the train frequency and shorten the train running interval, quickly relieve passenger flow pressure, run a small train set in off-peak and low-peak hours to save resources, reduce operating costs, and improve passenger service. A large amount of running data will be generated during the running of the virtual marshalling train, and these data are semi-structured text data, with a large amount of information and strong heterogeneity. The current analysis of virtual marshalling vehicle-mounted logs is still in the initial stage, and basically involves manually pulling data remotely and then analyzing the parsed data to manually draw charts and make tables, etc., which is low in data processing efficiency.
[0003] At present, there is no effective solution to the technical problem of low processing efficiency of virtual marshalling vehicle-mounted log data. SUMMARY
[0004] The present disclosure provides a virtual marshalling vehicle-mounted log data processing method, platform, device and storage medium.
[0005] According to a first aspect of the present disclosure, a virtual marshalling vehicle-mounted log data processing method is provided. The method comprises: monitoring a remote network disk and storing original log data uploaded to the remote network disk to a distributed file system, wherein the original log data is semi-structured text for recording virtual marshalling train running information, and the original log data comprises train set speed information; parsing the original log data into structured data and storing the structured data into a target data table of the distributed file system; and acquiring and processing, according to a business domain, structured data corresponding to the virtual marshalling and the business domain from the target data table to obtain business data corresponding to the virtual marshalling and the business domain, wherein the corresponding relationship between two trains in the virtual marshalling is determined according to a marshalling establishment condition and the train set speed information.
[0006] According to the aspect and any possible implementation manner described above, further provided is an implementation manner that, when the business domain is the running time length of the marshalling train in the interval, the detailed data corresponding to the virtual marshalling and the business domain is obtained, and the detailed data corresponding to the virtual marshalling and the business domain is analyzed and calculated according to the preset rule corresponding to the business domain to obtain the business data corresponding to the virtual marshalling and the business domain, including: determining the interval terminal station parking time of the train set in the target interval according to the train set speed information in the detailed data; taking two train sets with a difference in the interval terminal station parking time in the target interval within a specified time length as marshalling two trains, wherein the marshalling two trains include a front train and a rear train; determining the interval running time length of the front train in the target interval and the interval running time length of the rear train in the target interval according to the interval starting station platform departure time and the interval terminal station platform parking time of each train set in the marshalling two trains in the target interval.
[0007] According to the aspect and any possible implementation manner described above, further provided is an implementation manner that, when the business domain is the running time length of the marshalling train in the interval, the detailed data corresponding to the virtual marshalling and the business domain is obtained, and the detailed data corresponding to the virtual marshalling and the business domain is analyzed and calculated according to the preset rule corresponding to the business domain to obtain the business data corresponding to the virtual marshalling and the business domain, including: determining the interval terminal station parking time of the train set in the target interval according to the train set speed information in the detailed data; taking two train sets with a difference in the interval terminal station parking time in the target interval within a specified time length as marshalling two trains, wherein the marshalling two trains include a front train and a rear train; determining the interval running time length of the front train in the target interval and the interval running time length of the rear train in the target interval according to the interval starting station platform departure time and the interval terminal station platform parking time of each train set in the marshalling two trains in the target interval.
[0008] According to the aspect and any possible implementation manner described above, further provided is an implementation manner that, when the business domain is the running time length of the marshalling train in the interval, the detailed data corresponding to the virtual marshalling and the business domain is obtained, and the detailed data corresponding to the virtual marshalling and the business domain is analyzed and calculated according to the preset rule corresponding to the business domain to obtain the business data corresponding to the virtual marshalling and the business domain, including: determining the interval terminal station parking time of the train set in the target interval according to the train set speed information in the detailed data; taking two train sets with a difference in the interval terminal station parking time in the target interval within a specified time length as marshalling two trains, wherein the marshalling two trains include a front train and a rear train; determining the interval running time length of the front train in the target interval and the interval running time length of the rear train in the target interval according to the interval starting station platform departure time and the interval terminal station platform parking time of each train set in the marshalling two trains in the target interval.
[0009] According to a second aspect of the present disclosure, a virtual marshalling train-mounted log data processing platform is provided. The platform comprises:
[0010] a data bus configured to monitor a remote network disk and store raw log data uploaded to the remote network disk to a distributed file system, wherein the raw log data is semi-structured text used to record virtual marshalling train running information, and the raw log data includes train set speed information;
[0011] a raw layer configured to parse the raw log data into structured data and store the structured data into a target data table of the distributed file system, wherein the structured data includes a plurality of fields and a value corresponding to each field;
[0012] a unified data warehouse layer configured to obtain and process structured data corresponding to a virtual marshalling and a business domain from the target data table according to the business domain to obtain business data corresponding to the virtual marshalling and the business domain, wherein a corresponding relationship of marshalling two trains in the virtual marshalling is determined according to marshalling establishment conditions and train set speed information.
[0013] According to the aspect and any possible implementation manner as described above, an implementation manner is further provided, the unified data warehouse layer comprises:
[0014] a detailed data layer, configured to perform data deduplication, data cleaning and dimension processing on the structured data to obtain detailed data;
[0015] a summary data layer, configured to obtain the detailed data corresponding to the virtual grouping and the business domain, and perform analysis and calculation on the detailed data corresponding to the virtual grouping and the business domain according to preset rules corresponding to the business domain to obtain business data corresponding to the virtual grouping and the business domain.
[0016] According to the aspect and any possible implementation manner as described above, an implementation manner is further provided, the summary data layer is further configured to: determine the parking time of the train set at the terminal station in the target interval according to the train set speed information in the detailed data; take two train sets with a time difference of the parking time at the terminal station in the target interval within a specified time length as two trains in a grouping, wherein the two trains in a grouping comprise a front train and a rear train; in a case where the business domain is the interval running time of the train in a grouping, determine the interval running time of the front train in the target interval and the interval running time of the rear train in the target interval according to the departure time of each train set at the terminal station in the target interval and the parking time of each train set at the terminal station in the target interval.
[0017] According to the aspect and any possible implementation manner as described above, an implementation manner is further provided, the data platform further comprises:
[0018] an application layer, configured to obtain the business data from the unified data warehouse layer according to application requirements, and process the business data to provide data services corresponding to the application requirements.
[0019] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method as described above when executing the program.
[0020] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the method according to the first aspect and / or the second aspect of the present disclosure.
[0021] The present disclosure realizes the automatic acquisition, analysis and processing of the virtual grouping train log data by monitoring and automatically storing the original log data uploaded to the remote network disk, and analyzing the original log data into structured data and processing the data according to different business domains, thereby improving the processing efficiency of the virtual grouping train log data.
[0022] It is to be understood that the description in the summary is not intended to identify key or essential features of embodiments of the disclosure, nor is it intended to limit the scope of the disclosure. Other features of the disclosure will be apparent from review of the disclosure, which is as follows. BRIEF DESCRIPTION OF DRAWINGS
[0023] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent upon reading of the following detailed description, taken in conjunction with the accompanying drawings, wherein:
[0024] Figure 1 A flow chart of a virtual marshalling vehicle-mounted log data processing method according to an embodiment of the present disclosure is shown;
[0025] Figure 2 A block diagram of a virtual marshalling vehicle-mounted log data processing platform according to an embodiment of the present disclosure is shown;
[0026] Figure 3 A comparative flow chart of a conventional vehicle-mounted log integration and analysis scheme and the scheme of the present embodiment is shown;
[0027] Figure 4 A schematic diagram of raw semi-structured text data according to an embodiment of the present disclosure is shown;
[0028] Figure 5 Data after semi-structured text data structure analysis according to an embodiment of the present disclosure is shown;
[0029] Figure 6 A detailed diagram of a virtual marshalling vehicle-mounted log data storage model and data flow process according to an embodiment of the present disclosure is shown;
[0030] Figure 7 A speed curve analysis diagram of two vehicles before and after virtual marshalling in an operation process according to an embodiment of the present disclosure is shown;
[0031] Figure 8 A schematic diagram of two vehicles before and after virtual marshalling in an operation process according to an embodiment of the present disclosure is shown;
[0032] Figure 9 A schematic diagram of two vehicles before and after virtual marshalling in an operation process according to an embodiment of the present disclosure is shown;
[0033] Figure 10 A block diagram of an exemplary electronic device capable of implementing an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0034] In order to make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present disclosure.
[0035] In addition, the term "and / or" in this paper is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects.
[0036] In the present disclosure, by monitoring the remote network disk, the original log data uploaded to the remote network disk is automatically stored, and the original log data is automatically parsed into structured data, and the data is processed according to different business domains, which realizes the automatic acquisition, parsing and processing of virtual marshalling train log data, and improves the processing efficiency of virtual marshalling train log data.
[0037] Figure 1 A flowchart of a virtual marshalling train log data processing method 100 according to an embodiment of the present disclosure is shown.
[0038] Step S110, monitoring the remote network disk, and storing the original log data uploaded to the remote network disk to the distributed file system, wherein the original log data is a semi-structured text for recording virtual marshalling train running information, and the original log data includes train speed information;
[0039] Step S120, parsing the original log data into structured data and storing it into the target data table of the distributed file system;
[0040] Step S130, according to the business domain, obtaining and processing the structured data corresponding to the virtual marshalling and the business domain from the target data table to obtain the business data corresponding to the virtual marshalling and the business domain, wherein the corresponding relationship between the two trains in the virtual marshalling is determined according to the marshalling establishment condition and the train speed information.
[0041] The original log data includes ATO (Automatic Train Operation) and ATP (Automatic Train Protection) data of a VOBC (Vehicle On-Board Controller). For example, the ATO and ATP systems generate a period of semi-structured log data every 0.2 seconds, and a period of semi-structured log data includes more than 100 pieces of information, which is very large and heterogeneous. A period of semi-structured information needs to be parsed into structured text data to better analyze and use.
[0042] Optionally, in step S110, the original log data integrated from the remote network disk is stored in a specified directory of an HDFS (Hadoop Distributed File System), and data of different systems and different sources is integrated into a virtual marshalling vehicle log data processing platform through a data bus to provide the most original data source for subsequent data processing.
[0043] Optionally, in step S120, a task is started to parse semi-structured text into structured data and store the structured data in an original layer (ODS) of a HIVE (Hive, a data warehouse tool).
[0044] In some embodiments, in step S130, structured data corresponding to a virtual marshalling and a business domain is obtained and processed from a target data table to obtain business data corresponding to the virtual marshalling and the business domain, including:
[0045] The structured data is subjected to data deduplication, data cleaning, and dimension processing to obtain detailed data.
[0046] The detailed data corresponding to the virtual marshalling and the business domain is obtained, and the detailed data corresponding to the virtual marshalling and the business domain is analyzed and calculated according to a preset rule corresponding to the business domain to obtain business data corresponding to the virtual marshalling and the business domain.
[0047] Optionally, the detailed data is stored in a detailed data layer of a unified data warehouse layer, and the business data corresponding to the business domain is stored in a summary data layer of the unified data warehouse layer.
[0048] According to the embodiments of the present disclosure, the data processing flow is designed to obtain detailed data from structured data, and then obtain business data corresponding to a business domain by summarizing and analyzing the detailed data, so that the horizontal layered processing of data is realized, and the processing efficiency of virtual marshalling vehicle log data is improved.
[0049] In some embodiments, when the business domain is the running time of a train section, detailed data corresponding to the virtual train formation and the business domain are obtained, and the detailed data corresponding to the virtual train formation and the business domain are analyzed and calculated according to preset rules corresponding to the business domain to obtain the business data corresponding to the virtual train formation and the business domain, including:
[0050] Based on the train speed information in the detailed data, determine the stopping time of the train at the terminal station of the target section;
[0051] Two train sets whose stopping times at the terminal station of the target section are within a specified time difference will be considered as a two-car train, wherein the two-car train includes a leading car and a trailing car.
[0052] Based on the departure time of each train in the two-car train formation at the starting platform and the stopping time at the ending platform of the target section, the running time of the preceding train and the running time of the following train in the target section are determined.
[0053] The original log data of the virtual train formation contains train speed information, and the detailed data obtained from the original log data also contains train speed information. Based on the train speed information, the train departure and stop information can be obtained. Zero speed to non-zero speed is the departure mark, and non-zero speed to zero speed is the stop mark.
[0054] Find the departure time at the starting platform and the stopping time at the ending platform of the corresponding section for each train set. Then, subtract the departure time from the stopping time to obtain the running time of a single train set in a certain section. If the two trains in the virtual train set pass through the same section and the time difference between their arrival and stopping times is within a specific range, the running time of the two trains in the same section can be determined.
[0055] According to embodiments of this disclosure, by processing the original log data of the virtual train formation, and then analyzing, calculating and comparing the running times of the preceding and following trains in different sections, the running time of the train formation in the section can be monitored. When the running time of the train in the section exceeds a certain threshold, an alarm is triggered. Further analysis can be conducted on the circumstances under which the threshold will be exceeded during the running in the section, thereby finding factors to optimize the running of the train formation in the section and improving the efficiency of the train formation in the section.
[0056] In some embodiments, method 100 further includes:
[0057] We acquire business data based on application requirements and process the business data to provide data services corresponding to those application needs.
[0058] Optionally, data services include, but are not limited to: data analysis reports, curve analysis charts, indicator analysis charts, automated export of structured text data, and security performance analysis.
[0059] According to embodiments of this disclosure, data is obtained from a unified data warehouse layer to meet application requirements, and specific business data is processed to meet the specific needs of the business.
[0060] The method 100 of this disclosure embodiment will be described below using specific implementation examples:
[0061] With the increasing number of rail transit passengers and the continuous expansion of the network, passenger flow in urban transportation systems is dynamically changing. The uneven spatial and temporal distribution of urban rail transit passenger demand leads to a significant mismatch between the capacity provided and passenger demand in certain areas and at different times. In traditional fixed-formation trains, adjusting departure frequencies is typically used to increase capacity and achieve a better match between passenger and train flow. However, this system design has limitations; departure frequencies cannot be adjusted indefinitely. During peak hours, capacity needs to be further increased, while during off-peak and low-peak hours, passenger service needs to be further improved. Therefore, there is a strong demand for flexible train formations. Flexible train formations are divided into physical coupling and uncoupling technologies and virtual coupling and uncoupling technologies. Virtual coupling and uncoupling technologies have relatively lower requirements for coupler and coupling / uncoupling positions, offering extremely high flexibility and providing better solutions for train operation organization. Virtual train formation aims to achieve a better and more precise match between passenger and train flow. During peak hours, multiple train sets can operate in large formations to increase departure frequency, shorten train intervals, and quickly alleviate passenger flow pressure. During off-peak and low-peak hours, smaller train sets can operate to save resources, reduce operating costs, and improve passenger service.
[0062] Virtual train formations generate a large amount of operational data during operation. This data originates from various systems, including HAATP and ITP from ATP, and HAATO and ITO from ATO. This data is semi-structured text data. The ATO and ATP systems generate a cycle of semi-structured log data every 0.2 seconds, with over 100 entries per cycle. This data volume is substantial and highly heterogeneous. A cycle of semi-structured information needs to be parsed to form structured text data for better analysis and use. To perform big data analysis on the virtual train formation's onboard logs, it is necessary to collect, integrate, structure, and uniformly consolidate data from different systems. Based on this, a unified data model should be established, and an indicator system for the virtual train formation's onboard logs should be built, taking into account actual business needs. Currently, the analysis of virtual train formation onboard logs is still in its early stages. It primarily involves manually pulling data remotely, parsing it locally, analyzing the parsed data, and manually creating charts and tables. There is a lack of a systematic architectural solution to support the entire data processing process.
[0063] This disclosure presents a big data processing method for analyzing onboard logs of virtual train formations. It provides a complete solution encompassing data collection, parsing, modeling, storage, mining and analysis, data services, and front-end business presentation. Furthermore, it implements a comprehensive application solution for onboard logs of train formations. This can provide data support for improving the performance of virtual train formations, enhancing train operation service capabilities, and achieving better precise matching of train and passenger flows.
[0064] The method specifically includes the following steps: (1) First, integrate the virtual train vehicle-related log data, including VOBC ATO and ATO data, into the data platform through the data bus; (2) Parse the semi-structured text data integrated into the data platform, parse it into structured data, and uniformly put it into the raw data layer of the data warehouse; (3) Design the data model according to the business needs of the survey and the relevant integrated data. The data is horizontally layered according to the hierarchy of ODS (raw layer) -> DWD (unified data warehouse layer detailed data layer) -> DWS (unified data warehouse layer summary data layer) -> ADS (application layer). At the same time, the storage scheme is determined according to the data model; (4) Perform data calculation and processing based on the designed data model and the data integrated into the data warehouse to realize data analysis and support for business applications, realize the value of virtual train vehicle log data analysis and empower the business.
[0065] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0066] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.
[0067] Figure 2 A block diagram of a virtual formation vehicle log data processing platform according to an embodiment of the present disclosure is shown. Figure 2 As shown, the virtual train formation vehicle log data processing platform 200 includes:
[0068] Data bus 210 is used to monitor the remote cloud drive and store the raw log data uploaded to the remote cloud drive to the distributed file system. The raw log data is semi-structured text used to record the operation information of the virtual train formation, including train speed information.
[0069] The raw layer 220 is used to parse the raw log data into structured data and store it in the target data table of the distributed file system. The structured data includes multiple fields and the value corresponding to each field.
[0070] The unified data warehouse layer 230 is used to obtain and process the structured data corresponding to the virtual train and the business domain from the target data table according to the business domain, so as to obtain the business data corresponding to the virtual train and the business domain. The correspondence between the two cars in the virtual train is determined according to the train establishment conditions and the train speed information.
[0071] In some embodiments, the unified data warehouse layer 230 further includes: a detailed data layer, used to perform data deduplication, data cleaning and dimension processing on structured data to obtain detailed data; and a summary data layer, used to obtain detailed data corresponding to virtual groups and business domains, and to analyze and calculate the detailed data corresponding to virtual groups and business domains according to preset rules corresponding to business domains to obtain business data corresponding to virtual groups and business domains.
[0072] In some embodiments, the aggregated data layer is further configured to: determine the stopping time of the train set at the terminal station of the target section based on the train set speed information in the detailed data; classify two train sets whose stopping times at the terminal station of the target section differ within a specified time as a two-car train, wherein the two-car train includes a lead car and a follower car; and, when the business domain is the interval running time of the train set, determine the interval running time of the lead car and the interval running time of the follower car in the target section based on the departure time of each train set in the two-car train at the starting platform and the stopping time at the terminal platform of the target section.
[0073] In some embodiments, platform 200 further includes:
[0074] Application layer 240 is used to obtain business data from the unified data warehouse layer according to application requirements, and process the business data to provide data services corresponding to application requirements.
[0075] The platform 200 of this disclosure embodiment will be described below using specific implementation examples:
[0076] The technical roadmap for achieving big data analysis and processing of virtual trainset onboard log data and its integration with business applications includes the following aspects: integration of trainset onboard ATO and ATP log data, parsing of semi-structured text data from onboard logs into structured text data, data model design and storage solution construction, data flow process, data analysis and mining, and integration with actual operational scenarios. The specific process is as follows:
[0077] Step 1: Preparation of Virtual Train Onboard Log Data. Virtual train onboard log data includes ATO's HAATO and ITO logs, as well as ATP's HAATP and ITP logs. First, data from different systems and sources needs to be integrated and then aggregated into a data platform to provide the raw data source for subsequent data processing. The ATO and ATP log data integrated into the data platform is semi-structured text data. Semi-structured text data is not very user-friendly for analysis and mining. To facilitate unified calculation and processing later, the semi-structured text data needs to be parsed into structured text data.
[0078] Step 2: Determining the data model construction and storage scheme, as well as the entire data lifecycle process. Based on the actual business operation process of virtual train formation, the virtual train formation onboard log data integrated into the data platform, relevant data in the electronic map, actual operation business scenarios, and other factors, the data model construction is comprehensively considered to build a unified data model for virtual train formation analysis. At the same time, based on big data technologies such as Hadoop, Spark, HIVE, and CLICKHOUSE, the constructed data model is stored and processed.
[0079] Step 3: Based on actual business operation scenarios, conduct data analysis and mining according to the foundation laid in Steps 1 and 2. Through modeling, calculation, analysis, and mining, generate intelligent data services such as data dashboards, data analysis reports, curve analysis charts, indicator analysis charts, and automated export of structured text data.
[0080] The specific implementation technique of this disclosure is as follows:
[0081] (I) Preparation of Virtual Formation Vehicle Log Data
[0082] The preparation of virtual trainset vehicle log data includes two parts: virtual trainset vehicle log data integration and structured parsing of semi-structured vehicle log data. For example... Figure 3 A flowchart comparing traditional vehicle log integration and parsing schemes with the schemes of this disclosure.
[0083] 1. Traditional solutions for integrating raw virtual trainset vehicle log data involve manually retrieving compressed ATO and ATP files from different paths on a remote cloud drive, downloading them locally, decompressing them, and then parsing and processing them using code. This process is highly manual and inefficient, as local storage and computing capabilities cannot compare to those of a big data platform. This disclosure provides a timely processing function for ATO and ATP logs. By monitoring files on the remote cloud drive, a task can be initiated to promptly acquire and integrate the log data once data is available on the remote drive. The data is then stored on a big data platform, and automated parsing, calculation, analysis, and sharing are performed. This enables timely and effective data services, such as data front-end display and data analysis reports.
[0084] 2. The original data structure was semi-structured text data, which is very unfriendly to downstream data analysis and mining. It needs to be parsed into structured text data to facilitate downstream data processing using structured SQL. Therefore, after the data was integrated into the data platform, the data structuring and parsing task was immediately initiated. The original semi-structured text data is as follows: Figure 4 As shown, the data after structuring and parsing semi-structured text data is as follows: Figure 5 As shown.
[0085] The parsed structured data structure is described in Table 1.
[0086] Table 1:
[0087]
[0088]
[0089] (II) Design of virtual train-mounted log data model and establishment of storage scheme, as well as the data lifecycle process.
[0090] The vehicle log data model is divided into three layers according to the principle of horizontal layering: the ODS raw data layer, the DW unified data warehouse layer (which includes the DWD detailed data layer, the DWS summary data layer, and the DIM dimensional data layer), and the ADS data application layer. The ODS data layer mainly parses, aggregates, and integrates the raw vehicle log data into HIVE, providing a data source for the DW unified data warehouse layer. The DW unified data warehouse layer organizes data according to business domains and processes, defines consistency indicators and dimensions, and each business segment and business domain is independently constructed with unified specifications to form a unified and standardized standard business data system. The ADS data application layer is geared towards the needs of final business applications, obtaining data from the unified data warehouse layer and processing business-specific data to meet specific business needs. The detailed design of the model is shown in Table 2.
[0091] Table 2:
[0092]
[0093]
[0094]
[0095] Virtual trainset vehicle log data is stored on a data platform. The specific storage scheme is based on the vehicle log data analysis model; for details, please refer to [reference needed]. Figure 6 The diagram shows a detailed representation of the virtual trainset vehicle log data storage model and data flow process.
[0096] The raw semi-structured data integrated from the remote cloud drive is stored in a designated directory of HDFS. For the data model layer storage scheme, according to the horizontal layered architecture of the model, it is stored in the raw ODS layer, the unified data warehouse layer detailed data DWD, the unified data warehouse layer summary data DWS, and the application data layer ADS. The data in the model layer is calculated by SPARK and then stored in HIVE. For the application layer ADS data, considering the query performance of the data service, the ADS data in HIVE will be copied to CLICKHOUSE.
[0097] The specific calculation and flow process of the data is shown in the above-mentioned virtual trainset vehicle log data storage model and detailed data flow process diagram. After an operation is completed, the on-site staff will upload the virtual trainset vehicle log data to the remote network disk. The data acquisition system will monitor the designated remote network disk path. Once new data is detected, a task will be started to collect the data in the remote network disk and store it in HDFS on the data platform. Then, a task will be started to parse the semi-structured text file into a structured text file and store it in the ODS layer of HIVE. Next, the data processing and flow task in the data model will be started. First, the ODS data will be processed into DWD detailed data. Then, the DWD detailed data will be processed into DWS summary data. Then, the DWS summary data will be processed into the data required by the ADS application layer. The data processed in the data model is calculated and processed by the SPARK job and then stored in HIVE. After the data flow in the data model is completed, the data in the ADS layer will be copied to CLICKHOUSE to provide data services.
[0098] (III) Integration of in-vehicle data analysis with actual business operation scenarios
[0099] This disclosure provides various data usage solutions in actual business operation scenarios, including a central SMDU integrated dashboard, data analysis reports, curve analysis charts, indicator analysis charts, automated export of structured text data, safety performance analysis, and other data intelligence service scenarios. The following two business operation scenarios are used as examples for further explanation: a speed curve analysis chart of two train sets and a section running time analysis report of two train sets.
[0100] Analysis of speed curves of vehicles before and after virtual train formation: During operation, virtual train formation generates ATO (Autonomous Train Operation) log information. The ATO log information prints detailed speed information of the current train formation, including Protective Infrared (EBI) speed, Target Infrared (SBI) speed, and Actual Speed (SPD) speed. By considering the formation establishment conditions (changing from a non-formation train to a formation train) and the correspondence between the two trains (formation master control car and formation slave control car), the information of the two trains before and after the formation can be merged to obtain the corresponding EBI, SBI, and SPD speed details. This allows for the analysis and comparison of the speed of the two trains before and after the virtual formation during operation, and the generation of data such as... Figure 7 The velocity curve analysis diagram is shown.
[0101] Analysis Report on the Running Time of the Two Cars in the Virtual Train Formation: The goal is to monitor the running time of the train in different sections by processing the onboard log data of the virtual train formation and then analyzing, calculating and comparing the running time of the two cars in different sections. When the running time of the train in a section exceeds a certain threshold, an alarm is triggered. Further analysis can be conducted to identify the circumstances under which the threshold is exceeded during the running time in the section, thereby finding factors to optimize the running time of the train in the section and improving the efficiency of the train in the section.
[0102] A complete line operation consists of a train traveling from platform 2 down to platform 3, from platform 3 up to platform 2, from platform 2 up to platform 1, and from platform 1 down to platform 2. A complete section operation consists of a train departing from one platform and stopping at another. For example, a train departing from platform 2 and stopping at platform 3 constitutes a complete section operation. (Specific details follow...) Figure 8 As shown.
[0103] The virtual trainset's onboard log contains train departure and arrival information. Departure is indicated by speeds from zero to non-zero, and arrival is indicated by speeds from non-zero to zero. The departure time from the starting platform and the arrival time at the ending platform of the corresponding section are found. Subtracting the departure time from the arrival time gives the travel time of a single trainset within a given section. Furthermore, by ensuring that the two trains in the virtual trainset pass through the same section and that the difference in their arrival times is within a specific range, the travel time of the two trains in the same section can be determined. For example... Figure 9 The diagram shown is a schematic diagram of the running time analysis report between the two cars in a virtual train formation according to an embodiment of the present disclosure.
[0104] This disclosure addresses the issue of collecting, summarizing, and integrating semi-structured text data from virtual train formation logs into the TianShu data platform through data integration. It utilizes big data analytics to achieve fully automated pipeline operations, including automated integration of virtual train formation logs, automated parsing of semi-structured data, automated application analysis and calculation, and automated front-end data display. By incorporating events such as speed curve analysis and interval running time analysis reports into the data model for analysis and mining, it can guide the optimization and adjustment of relevant parameters in virtual train formation technology to improve the performance of virtual train formations. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the described modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0105] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0106] Figure 10A block diagram of an exemplary electronic device 1000 capable of implementing embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0107] Electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in ROM 1002 or a computer program loaded into RAM 1003 from storage unit 1008. RAM 1003 may also store various programs and data required for the operation of electronic device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. I / O interface 1005 is also connected to bus 1004.
[0108] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0109] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).
[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0115] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0116] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0117] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for processing virtual trainset onboard log data, characterized in that, include: Monitor the remote cloud drive and store the raw log data uploaded to the remote cloud drive in a distributed file system. The raw log data is semi-structured text used to record the operation information of virtual train formations, and the raw log data includes train speed information. The raw log data is parsed into structured data and stored in the target data table of the distributed file system; The structured data is deduplicated, cleaned, and dimensionality processed to obtain detailed data; Based on the train speed information in the detailed data, the stopping time of the train at the end station of the target section is determined; Two train sets whose stopping times at the terminal station of the target section are within a specified time difference are considered as a two-car train, wherein the two-car train includes a front car and a rear car. When the business domain is the running time of a train in a section, the running time of the preceding train and the running time of the following train in the target section are determined based on the departure time of each train group in the two trains in the section at the starting platform and the stopping time at the ending platform in the section. The correspondence between the two trains in the virtual train formation is determined based on the formation establishment conditions and the train speed information.
2. The method according to claim 1, characterized in that, The method further includes: Business data is acquired based on application requirements, and the business data is processed to provide data services corresponding to the application requirements.
3. A virtual train formation vehicle log data processing platform, characterized in that, include: A data bus is used to monitor a remote cloud drive and store the raw log data uploaded to the remote cloud drive in a distributed file system. The raw log data is semi-structured text used to record the operation information of virtual train formations, and the raw log data includes train speed information. The raw layer is used to parse the raw log data into structured data and store it in the target data table of the distributed file system; The detailed data layer is used to perform data deduplication, data cleaning, and dimensionality processing on the structured data to obtain detailed data; The summary data layer is used to determine the stopping time of the train at the terminal station of the target section based on the train speed information in the detailed data. Two train sets whose stopping times at the terminal station of the target section are within a specified time difference are considered as a two-car train, wherein the two-car train includes a front car and a rear car. When the business domain is the running time of a train in a section, the running time of the preceding train and the running time of the following train in the target section are determined based on the departure time of each train group in the two trains in the section at the starting platform and the stopping time at the ending platform in the section. The correspondence between the two trains in the virtual train formation is determined based on the formation establishment conditions and the train speed information.
4. The platform according to claim 3, characterized in that, The platform also includes: The application layer is used to obtain business data from the aggregated data layer according to application requirements, and process the business data to provide data services corresponding to the application requirements.
5. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 2.
6. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Train virtual marshaling method
CN113859326A
Train driving log data management method and system
CN115391302A