Train data processing classification method, device, computer equipment and storage medium

Through the distributed data processing architecture of Kafka message queue and SparkStreaming framework, the problem of large module coupling and insufficient real-time performance in real-time data processing in trains is solved, and flexible business module development and low-latency data interaction are realized.

CN115733896BActive Publication Date: 2025-07-18ZHUZHOU CSR TIMES ELECTRIC CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110991202.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-26
Publication Date
2025-07-18
Estimated Expiration
2041-08-26

AI Technical Summary

Technical Problem

In real-time train data processing, the traditional centralized streaming processing framework leads to high coupling between various service modules, increasing development and debugging complexity, and difficult to meet the real-time requirements of data processing.

Method used

Using a distributed data processing architecture based on Kafka message queue, using data flow as the core, the SparkStreaming processing framework is used to analyze and classify data, realize the decoupling of business modules, and provide the ability to subscribe to data on demand.

Benefits of technology

It improves the flexibility of system development and the stability of background processing, reduces the complexity of functional modules, and realizes low-latency data interaction in high concurrency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115733896B_ABST
    Figure CN115733896B_ABST
Patent Text Reader

Abstract

The present application relates to a train data processing and classification method, apparatus, computer device, and storage medium. The method includes: obtaining real-time data packets sent by train devices, encapsulating them into a binary data structure, and writing them into a Kafka message queue to form an original data stream; reading binary data from the original data stream, parsing the data into point information according to preset rules, and writing it into the Kafka message queue to form a parsed data stream; obtaining the data content from the parsed data stream, automatically filtering out the data points required by different business processing programs, and writing them into the Kafka message queue to form data streams of different business classifications. This method enables each business function module to be independent of each other, improves the flexibility of system development and product function assembly, and at the same time increases the stability of the background processing system. It provides the ability to subscribe to the required data on demand, significantly reducing the complexity of the function modules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a method, device, computer device, and storage medium for classifying train data processing. Background Art

[0002] In the context of the rapid development of the Internet of Things and big data, in order to better play the supporting role of data in the business field, the data collected from various sensors and module devices on the train will be transmitted back to the ground data center in real time. The real-time data needs to be received, parsed, and calculated before its true value can be obtained. In the current field of train real-time data applications, with the in-depth exploration of business, the derived business functions also increase accordingly. There will be operations such as correlation calculation and logical analysis based on the overall or partial categories of data, so as to achieve functions such as status monitoring, health diagnosis, and fault warning.

[0003] Due to the large amount of sensor and device status data of the train, in the case of high concurrency of multiple trains, the real-time data volume becomes very large. The traditional method of transmitting data through API interface calls cannot meet the real-time requirements of data processing. If a centralized streaming processing framework is adopted, all services need to be concentrated in a main program module, which leads to too high coupling between service modules. Every time a service module is added or modified, the entire main program module needs to be recompiled and released, which will increase the complexity of development and debugging; the main program module has redundant functions, and any problem in a function module may cause the main program to be abnormal. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, computer device, and storage medium for classifying train data processing in view of the above technical problems.

[0005] In a first aspect, an embodiment of the present invention provides a method for classifying train data processing, the method including:

[0006] Obtain real-time data packets sent by train devices, encapsulate them into a binary data structure, and write them into the kafka message queue to form an original data stream;

[0007] Read binary data from the original data stream, parse the data into point position information according to preset rules, and write it into the kafka message queue to form a parsed data stream;

[0008] Obtain the data content from the parsed data stream, automatically filter out the data point positions required by different service processing programs, and write them into the kafka message queue to form data streams of different service classifications.

[0009] Further, the steps of obtaining real-time data packets sent by train equipment, encapsulating them into a binary data structure, and writing them into the kafka message queue to form an original data stream include:

[0010] Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol;

[0011] After data verification, encapsulate them into a binary data structure of a specific protocol, and form the original data stream through writing into the kafka message queue;

[0012] While performing data stream transfer of the original data stream, perform data backup.

[0013] Further, the steps of reading binary data from the original data stream, parsing the data into point position information according to preset rules, and writing them into the kafka message queue to form a parsed data stream include:

[0014] Read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data;

[0015] According to preset rules, perform data length comparison, data parsing, and unit conversion on the original data stream, and form point position information;

[0016] After being processed by the data parsing engine, perform data packet encapsulation and stream writing to form a parsed data stream.

[0017] Further, the steps of obtaining data content from the parsed data stream, automatically screening out the data point positions required by different service processing programs, and writing them into the kafka message queue to form data streams of different service classifications include:

[0018] Build a subscription configuration table according to the data point positions, data stream names, and write corresponding kafka queue topic information required by different services;

[0019] Through the subscription configuration table, automatically screen out the data point positions required by different service processing programs from the parsed data stream according to the data content;

[0020] Automatically refresh the classification of data according to the refresh frequency configured according to project requirements, and dynamically adjust the required data point positions according to the service program.

[0021] On the other hand, an embodiment of the present invention further provides a train data processing and classification system, including:

[0022] The vehicle-ground data communication module is used to obtain real-time data packets sent by train equipment, encapsulate them into a binary data structure, and write them into the kafka message queue to form an original data stream;

[0023] The data parsing module is used to read binary data from the original data stream, parse the data into point position information according to preset rules, and write it into the kafka message queue to form a parsed data stream;

[0024] The data automatic classification module is used to obtain data content from the parsed data stream, automatically screen out data point positions required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications.

[0025] Further, the vehicle-ground data communication module includes a preprocessing unit, and the preprocessing unit is used for:

[0026] Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption and decompression according to the communication protocol;

[0027] After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream through the writing of the kafka message queue;

[0028] While performing data stream transmission on the original data stream, perform data backup.

[0029] Further, the data parsing module includes a distributed framework unit, and the distributed framework unit is used for:

[0030] Read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data;

[0031] According to preset rules, perform data length comparison, data parsing, unit conversion on the original data stream, and form point position information;

[0032] After being processed by the data parsing engine, perform data packet encapsulation and write stream to form a parsed data stream.

[0033] Further, the data automatic classification module includes a refresh classification unit, and the refresh classification unit is used for:

[0034] Construct a subscription configuration table according to the data point positions, data stream names required by different services, and write the corresponding kafka queue topic information;

[0035] Through the subscription configuration table, automatically screen out the data point positions required by different business processing programs from the parsed data stream;

[0036] Configure the classification of automatically refreshed data according to the project requirements, and dynamically adjust the required data points according to the business procedures.

[0037] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0038] Obtain the real-time data packet sent by the train equipment, encapsulate it into a binary data structure, and write it into the kafka message queue to form an original data stream;

[0039] Read the binary data from the original data stream, parse the data into point information according to the preset rules, and write it into the kafka message queue to form a parsed data stream;

[0040] Obtain the data content from the parsed data stream, automatically filter out the data points required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications.

[0041] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0042] Obtain the real-time data packet sent by the train equipment, encapsulate it into a binary data structure, and write it into the kafka message queue to form an original data stream;

[0043] Read the binary data from the original data stream, parse the data into point information according to the preset rules, and write it into the kafka message queue to form a parsed data stream;

[0044] Obtain the data content from the parsed data stream, automatically filter out the data points required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications.

[0045] The above train data processing and classification method, device, computer device, and storage medium. The method includes: obtaining real-time data packets sent by train devices, encapsulating them into a binary data structure, and writing them into a Kafka message queue to form an original data stream; reading binary data from the original data stream, parsing the data into point position information according to preset rules, and writing it into a Kafka message queue to form a parsed data stream; obtaining the data content from the parsed data stream, automatically screening out the data point positions required by different business processing programs, and writing them into a Kafka message queue to form data streams of different business classifications. Among them, a background processing architecture with the data stream as the core is adopted, making each business function module independent of each other, improving the flexibility of system development and product function assembly, and at the same time increasing the stability of the background processing system. Moreover, a high-performance big data message middleware is used as the data stream carrier to ensure low latency in data interaction under high data concurrency. At the same time, a data automatic classification module is adopted to provide the ability for each business processing module to subscribe to the required data on demand, significantly reducing the complexity of the function module. Description of the Drawings

[0046] Figure 1 It is a schematic flowchart of the train data processing and classification method in an embodiment;

[0047] Figure 2 It is a schematic flowchart of the preprocessing method in the vehicle-ground data communication process in an embodiment;

[0048] Figure 3 It is a schematic flowchart of the distributed data parsing in an embodiment;

[0049] Figure 4 It is a schematic flowchart of the data automatic classification step in an embodiment;

[0050] Figure 5 It is a structural block diagram of the train data processing and classification system in an embodiment;

[0051] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiment

[0052] In order to make the purpose, technical solution, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] In one embodiment, as Figure 1 shown, a train data processing and classification method is provided, including the following steps:

[0054] Step 101: Obtain the real-time data packets sent by the train equipment, encapsulate them into a binary data structure, and write them into the Kafka message queue to form the original data stream.

[0055] Step 102: Read the binary data from the original data stream, parse the data into point position information according to the preset rules, and write it into the Kafka message queue to form the parsed data stream.

[0056] Step 103: Obtain the data content from the parsed data stream, automatically screen out the data point positions required by different business processing programs, and write them into the Kafka message queue to form data streams of different business classifications.

[0057] Specifically, the data processing and classification method is a train real-time data processing and automatic classification method based on the data stream architecture. This architecture takes the data stream as the core, and each business function module can directly obtain the required data from the data stream to achieve the decoupling of business functions. This method uses a big data-based message middleware as the data stream carrier, supporting high-throughput and low-latency data interaction capabilities between different program modules. The data automatic classification module designed by this method can automatically configure and screen the required data point positions according to the needs of business processing programs, form new business data streams for professional program processing, and achieve the ability of data on-demand subscription and automatic distribution. This method provides the ability to independently develop and assemble business modules in the case of multiple vehicles and high data concurrency, increasing the flexibility of function module development and enhancing the stability of the background processing architecture.

[0058] Among them, the proposed data processing and classification method takes the data stream as the core. The data stream uses the big data message middleware Kafka as the carrier. One topic represents one real-time data stream. The upstream program module of the data stream corresponds to the producer of the Kafka topic, and the downstream program module of the data stream corresponds to the consumer of the Kafka topic. In this way, the producer writes data into the topic, and the consumer reads data from the topic, and the background data processing system forms a data-based streaming architecture. Kafka is an open-source log subscription system that can be used as a message middleware, providing producer and consumer modes.

[0059] In one embodiment, as Figure 2 shown, the preprocessing method in the vehicle-ground data communication process includes:

[0060] Step 201: Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol.

[0061] Step 202: After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream by writing through the kafka message queue;

[0062] Step 203: While performing data stream transfer on the original data stream, perform data backup.

[0063] Specifically, the vehicle-ground data communication program module receives the TCP / UDP real-time data packets sent by the train equipment through the network. After verification, decryption, and decompression according to the communication protocol, it is encapsulated into a binary data structure of a specific protocol and written into the kafka message queue to form the original data stream. Among them, in the original data stream, the data stream is the original binary data and cannot be used by relevant business systems such as analysis and display, but it can provide a data source for the original data backup module.

[0064] In one embodiment, as Figure 3 shown, the process of distributed data parsing includes:

[0065] Step 301: Read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data;

[0066] Step 302: According to the preset rules, perform data length comparison, data parsing, unit conversion on the original data stream, and form point position information;

[0067] Step 303: After being processed by the data parsing engine, perform data packet encapsulation and write stream to form a parsed data stream.

[0068] Specifically, the data parsing program reads the binary data from the original data stream, parses the data into point position information with specific meanings (including operations such as data length comparison, data parsing, and unit conversion) according to the corresponding rules, and writes it into the kafka message queue to form a parsed data stream. In order to improve data processing efficiency, the data parsing module adopts the SparkStreamming processing framework based on big data to implement fast data stream processing in a distributed task manner. Among them, SparkStreaming is a streaming distributed processing framework, which is implemented based on the micro-batch processing mode of Spark in-memory computing and has high-efficiency and scalable data processing characteristics. In the parsed data stream, the data stream is the parsed Key-Value data with specific meaning data points, such as voltage, current, etc. It can be directly used by data analysis, real-time display, and point position data storage modules.

[0069] In one embodiment, as Figure 4 shown, the process of data automatic classification includes:

[0070] Step 401: Construct a subscription configuration table based on the data points, data stream names, and Kafka queue topic information corresponding to different services that need to be written.

[0071] Step 402: Automatically screen out the data points required by different service processing programs from the parsed data stream through the subscription configuration table according to the data content.

[0072] Step 403: Automatically refresh the classification of data according to the configured refresh frequency according to project requirements, and dynamically adjust the required data points according to the service program. When the system operation instruction is the first run, obtain the system source code corresponding to the system where the engineering project to be detected is located.

[0073] Specifically, in the process of automatic classification, data is obtained from the parsed data stream. According to the service requirements subscription configuration table (the subscription configuration table contains the data points, data stream names, and Kafka queue topic information required by each service), the data points required by each service processing program are automatically screened out from the parsed data stream according to the data content and written into the corresponding Kafka message queue to form each service data stream. The automatic refresh function of the data classification rule of this module can configure the refresh frequency according to project requirements and can realize the dynamic adjustment of the required data points by the service program. The specific service data stream is a new data stream formed by part of the data points screened out from the parsed data stream according to the needs of the service processing program according to its predetermined data points. This data stream serves a specific service, and the amount of data points it contains changes according to service requirements.

[0074] It should be understood that although the steps in the above flow chart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flow chart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0075] In one embodiment, as Figure 5 shown, a train data processing and classification system is provided, including:

[0076] A vehicle-ground data communication module 501, configured to obtain real-time data packets sent by train equipment, encapsulate them into a binary data structure, and write them into a Kafka message queue to form an original data stream.

[0077] A data parsing module 502, configured to read binary data from the original data stream, parse the data into point location information according to a preset rule, and write it into a kafka message queue to form a parsed data stream;

[0078] A data automatic classification module 503, configured to obtain data content from the parsed data stream, automatically filter out data point locations required by different service processing programs, and write them into a kafka message queue to form data streams of different service classifications.

[0079] In one embodiment, as Figure 5 shown, the vehicle-ground data communication module 501 includes a preprocessing unit 5011, and the preprocessing unit 5011 is configured to:

[0080] Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol;

[0081] After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream through writing of the kafka message queue;

[0082] While performing data stream transfer on the original data stream, perform data backup.

[0083] In one embodiment, as Figure 5 shown, the data parsing module 502 includes a distributed framework unit 5021, and the distributed framework unit 5021 is configured to:

[0084] Read and process the original data stream in a distributed task manner through a SparkStreamming processing framework based on big data;

[0085] According to a preset rule, perform data length comparison, data parsing, and unit conversion on the original data stream, and form point location information;

[0086] After being processed by a data parsing engine, perform data packet encapsulation and stream writing to form a parsed data stream.

[0087] In one embodiment, as Figure 5 shown, the data automatic classification module 503 includes a refresh classification unit 5031, and the refresh classification unit 5031 is configured to:

[0088] Construct a subscription configuration table according to the data point locations, data stream names, and write corresponding kafka queue topic information required by different services;

[0089] According to the subscription configuration table, data points required by different service handlers are automatically screened out from the parsed data stream according to the data content;

[0090] Configure the refresh frequency according to project requirements to automatically refresh the classification of data, and dynamically adjust the required data points according to the service program.

[0091] For the specific limitations of the train data processing classification system, reference can be made to the limitations of the train data processing classification method in the above text, which will not be elaborated here. Each module in the above train data processing classification system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0092] Figure 6 The internal structure diagram of a computer device in an embodiment is shown. As Figure 6 shown, the computer device includes a processor, a memory, a network interface, an input device, and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the train data processing classification method. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the train data processing classification method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0093] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0094] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0095] Obtain the real-time data packet sent by the train device, encapsulate it into a binary data structure, and write it into the kafka message queue to form an original data stream;

[0096] Read binary data from the original data stream, parse the data into point information according to preset rules, and write it into the kafka message queue to form a parsed data stream;

[0097] Obtain the data content from the parsed data stream, automatically filter out the data points required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications.

[0098] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0099] Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol;

[0100] After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream through the writing of the kafka message queue;

[0101] While performing data stream transfer on the original data stream, perform data backup.

[0102] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0103] Read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data;

[0104] According to preset rules, perform data length comparison, data parsing, unit conversion on the original data stream, and form point information;

[0105] After being processed by the data parsing engine, perform data packet encapsulation and write stream to form a parsed data stream.

[0106] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0107] Construct a subscription configuration table according to the data points, data stream names required by different services, and write the corresponding kafka queue topic information;

[0108] Through the subscription configuration table, automatically filter out the data points required by different business processing programs from the parsed data stream according to the data content;

[0109] Configure the refresh frequency according to project requirements to automatically refresh the data classification, and dynamically adjust the data points required by the business program.

[0110] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0111] Obtain the real-time data packets sent by the train equipment, encapsulate them into a binary data structure, and write them into the kafka message queue to form an original data stream;

[0112] Read the binary data from the original data stream, parse the data into point position information according to the preset rules, and write it into the kafka message queue to form a parsed data stream;

[0113] Obtain the data content from the parsed data stream, automatically filter out the data point positions required by different service processing programs, and write them into the kafka message queue to form data streams of different service classifications.

[0114] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0115] Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption and decompression according to the communication protocol;

[0116] After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream through the writing of the kafka message queue;

[0117] While performing data stream transfer on the original data stream, perform data backup.

[0118] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0119] Read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data;

[0120] According to the preset rules, perform data length comparison, data parsing, unit conversion on the original data stream, and form point position information;

[0121] After being processed by the data parsing engine, perform data packet encapsulation and write stream to form a parsed data stream.

[0122] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0123] Build a subscription configuration table according to the data point positions, data stream names required by different services, and the kafka queue topic information to be written;

[0124] According to the subscription configuration table, data points required by different service processing programs are automatically filtered from the parsed data stream according to the data content;

[0125] Configure the classification of automatically refreshed data according to the refresh frequency required by the project, and dynamically adjust the required data points according to the service program.

[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0127] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0128] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A train data processing and classification method, characterized in that, It includes the following steps: Obtain the real-time data packets sent by the train equipment, encapsulate them into a binary data structure, and write them into the kafka message queue to form an original data stream; the real-time data packets are TCP / UDP real-time data packets; Read binary data from the original data stream, parse the data into point position information according to preset rules, and write it into the kafka message queue to form a parsed data stream; the process of reading binary data from the original data stream, parsing the data into point position information according to preset rules, and writing it into the kafka message queue to form a parsed data stream includes: reading and processing the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data; according to preset rules, comparing the data length, parsing the data, and performing unit conversion on the original data stream to form point position information; after being processed by the data parsing engine, performing data packet encapsulation and writing stream to form a parsed data stream; Obtain the data content from the parsed data stream, automatically screen out the data point positions required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications; the process of obtaining the data content from the parsed data stream, automatically screening out the data point positions required by different business processing programs, and writing them into the kafka message queue to form data streams of different business classifications includes: constructing a subscription configuration table according to the data point positions, data stream names, and writing corresponding kafka queue topic information required by different businesses; through the subscription configuration table, automatically screen out the data point positions required by different business processing programs from the parsed data stream according to the data content; configure the refresh frequency according to the project requirements to automatically refresh the data classification, and dynamically adjust the required data point positions according to the business program; the subscription configuration table contains the data point positions, data stream names, and writing corresponding kafka queue topic information required by each business.

2. The method according to claim 1, characterized in that, The process of obtaining the real-time data packets sent by the train equipment, encapsulating them into a binary data structure, and writing them into the kafka message queue to form an original data stream includes: Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol; After data verification, encapsulate it into a binary data structure of a specific protocol, and form the original data stream through the writing of the kafka message queue; While performing data stream transmission on the original data stream, perform data backup.

3. A train data processing and classification system, characterized in that, It includes: A vehicle-ground data communication module, which is used to obtain the real-time data packets sent by the train equipment, encapsulate them into a binary data structure, and write them into the kafka message queue to form an original data stream; the real-time data packets are TCP / UDP real-time data packets; A data parsing module, which is used to read binary data from the original data stream, parse the data into point position information according to preset rules, and write it into the kafka message queue to form a parsed data stream; the data parsing module includes a distributed framework unit, and the distributed framework unit is used to: read and process the original data stream in a distributed task manner through the SparkStreamming processing framework based on big data; compare the data length, parse the data, and perform unit conversion on the original data stream according to preset rules, and form point position information. After being processed by the data parsing engine, the data is packetized and written into a stream to form a parsed data stream. A data automatic classification module, which is used to obtain data content from the parsed data stream, automatically screen out the data point positions required by different business processing programs, and write them into the kafka message queue to form data streams of different business classifications; the process of obtaining data content from the parsed data stream, automatically screening out the data point positions required by different business processing programs, and writing them into the kafka message queue to form data streams of different business classifications includes: constructing a subscription configuration table according to the data point positions, data stream names required by different businesses, and the topic information of the corresponding kafka queue to be written. Through the subscription configuration table, automatically screen out the data point positions required by different business processing programs from the parsed data stream according to the data content; automatically refresh the data classification according to the refresh frequency configured according to project requirements, and dynamically adjust the data point positions required by the business program; the subscription configuration table contains the data point positions, data stream names required by each business, and the topic information of the corresponding kafka queue to be written.

4. The train data processing and classification system according to claim 3, wherein The vehicle-ground data communication module includes a preprocessing unit, and the preprocessing unit is used to: Perform network reception on the TCP / UDP real-time data packets sent by the train equipment, and perform data verification, decryption, and decompression according to the communication protocol. After data verification, it is encapsulated into a binary data structure of a specific protocol, and the original data stream is formed by writing through the kafka message queue. While transmitting the original data stream, data backup is performed.

5. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 2.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Log processing method and device, computer equipment and storage medium

    CN109726074A

  • Data processing method, system and device and storage medium

    CN110334070A

  • Integrated network management method and apparatus for rail traffic system, and system

    WO2020038447A1