Railway vehicle data management method and system

By creating data element standards and vehicle configurations, the problems of decentralized data storage and inconsistent standards for rail vehicles have been solved, automatic fusion and quality monitoring of multi-system data have been achieved, data accuracy and efficiency have been improved, and full life cycle management has been supported.

CN120687950APending Publication Date: 2025-09-23CRRC QINGDAO SIFANG ROLLING STOCK RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510972030.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The data storage of rail vehicles is scattered and the standards are not unified, which makes it difficult to associate and integrate data, and the fusion is difficult and the quality is poor, which affects operational decisions and inspection and maintenance.

Method used

By creating data element standards and vehicle configurations, defining field information and hierarchical relationships, using a tree structure for data fusion, and combining data quality monitoring, unified data services are provided.

Benefits of technology

It realizes the automatic integration of multi-system data, improves the accuracy and efficiency of data, ensures the reliability and security of data, and provides reliable data support for the full life cycle management of rail vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687950A_ABST
    Figure CN120687950A_ABST
Patent Text Reader

Abstract

The invention relates to a railway vehicle data management method and system, and the method comprises the steps: creating a data element standard, classifying the data of each system and equipment of a railway vehicle, and defining field information corresponding to the attribute information of each type of data; establishing a vehicle configuration, wherein the vehicle configuration presents a hierarchical relationship between each system and equipment of the railway vehicle in a tree structure; according to the corresponding relationship between the attribute information of each category of data in the data element standard and the defined field information, fusing the attribute information of each category of data of the system and the equipment with the corresponding nodes in the vehicle configuration; performing quality monitoring on the fused data; data meeting quality requirements are screened out through quality monitoring, and data services are provided for the outside. According to the embodiment provided by the invention, the whole process from data arrangement to data utilization can be optimized, and the problems of scattered data storage, non-uniform standards, high fusion difficulty, poor fusion quality and the like of the current railway vehicle are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data management, and in particular to a method and system for rail vehicle data management. Background Art

[0002] With the booming manufacturing industry, the intelligence of manufactured products continues to increase, and the amount of data generated during production and use is exploding. For example, rail vehicles, their data covers the entire life cycle, from design, manufacturing, testing, operation and maintenance to scrapping. Managing this data throughout its life cycle is crucial to the progress and sustainable development of the entire industry.

[0003] However, due to the lack of a unified data standards system, data from different systems lacks standardization in terms of format, interface, and semantics, making data correlation and integration difficult. For example, data formats vary significantly across devices provided by different vendors, and the definition and representation of the same concept vary across different systems. Furthermore, data storage is fragmented, with a variety of storage methods coexisting, including relational, time-series, and non-relational databases. For example, workshop data is stored in relational databases, Industrial IoT data in time-series databases, and field data in non-relational databases. This inconsistency in underlying storage complicates data correlation and hinders efficient fusion analysis.

[0004] Therefore, a method for managing the full life cycle data of rail vehicles is needed to solve many problems of current rail vehicle data storage, such as scattered storage, inconsistent standards, high integration difficulty, and poor integration quality.

[0005] Application Contents The present application solves at least one of the technical problems in the related art to a certain extent, and provides a rail vehicle data management method and system that can optimize the entire process from data collation to utilization.

[0006] To achieve the above-mentioned objectives, in a first aspect, the present application provides a rail vehicle data governance method, comprising: creating a data element standard, comprising: classifying the data of various systems and equipment of the rail vehicle, wherein the data of each system and equipment include attribute information describing various characteristics, states and parameters during its life cycle, and defining field information corresponding to the attribute information of each category of data; establishing a vehicle configuration, wherein the vehicle configuration is a tree structure, wherein the tree structure includes a vehicle node, each system node of the vehicle, and each device node in the system, and the tree structure presents the hierarchical relationship between the rail vehicle, each system, and each device through each node; fusing data with the vehicle configuration, comprising: fusing the attribute information of each category of data of the system and equipment with the corresponding node in the vehicle configuration according to the correspondence between the attribute information of each category of data in the data element standard and the defined field information; performing quality monitoring on the fused data, wherein the dimensions of the quality monitoring include one or a combination of the integrity, uniqueness, and timeliness of the data; screening out data that meets the quality requirements through the quality monitoring, and providing data services to the outside world.

[0007] This embodiment addresses the existing issues of fragmented rail vehicle data storage and inconsistent standards by establishing data element standards and vehicle configurations. This enables automatic fusion of multi-system data, reducing integration complexity. Integrating this with a data quality monitoring system ensures the quality of fused data and improves data reliability. Ultimately, unified data services are provided based on device dimensions, improving data utilization efficiency and reducing complexity, providing an effective solution for data governance throughout the rail vehicle lifecycle.

[0008] In some embodiments, the method of defining field information corresponding to attribute information of each category of data includes: defining the meaning of each field, defining the length of each field, and defining the information type of each field.

[0009] This implementation provides a unified and standardized description standard for the attribute information of various types of rail vehicle data by clearly defining the meaning, length, and information type of each field. This ensures consistency in format and semantics across data from different systems and sources, eliminating barriers to data exchange and laying the foundation for subsequent integration with vehicle configurations. This reduces the difficulty of data integration, improves the accuracy and efficiency of data fusion, provides a clear basis for data quality monitoring, and ensures standardization throughout the entire data governance process.

[0010] In some embodiments, the field information includes the operating status field of the system and equipment, and the method of defining the field information corresponding to the attribute information of each category of data also includes: defining the field of the operating status of key systems and equipment as not allowed to be empty; the key systems and equipment include systems and equipment that affect driving safety, and core functional systems and equipment, and the systems and equipment that affect driving safety include at least traction motors, braking systems, and signal systems, and the core functional systems and equipment include at least bogies and on-board network systems.

[0011] In this implementation, the operational status fields for systems and devices that impact train safety and core functions are defined as non-nullable. This ensures the integrity of this critical data and prevents safety hazards or misjudgments of functional failures due to missing critical information. This setting strengthens the rigor of data element standards, providing reliable foundational data support for subsequent data integration, quality monitoring, and safe operational decision-making, further enhancing the security and effectiveness of rail vehicle data governance.

[0012] In some embodiments, the method of fusing the data with the vehicle configuration includes: creating a data table for each system node and device node in the vehicle configuration according to the created data element standard, the data table including fields corresponding to all attribute information of the system and the device, the fields being used to store the corresponding attribute information; and filling the attribute information of each category of data into the corresponding fields in the data table.

[0013] In this implementation, data tables with corresponding fields are created for each system and device based on the data element standard. This ensures a precise association between attribute information and the standard, ensuring that each attribute has a standardized storage location. This lays the foundation for standardized data storage and management, enabling efficient interaction and integration of data from different sources based on a unified structure, reducing the complexity of data association and improving the efficiency and accuracy of subsequent data processing and integration.

[0014] In some embodiments, the method for fusing the attribute information of each category of data with the vehicle configuration includes: selecting a framework for fusing each category of data with the vehicle configuration based on the data volume and timeliness of each category of data; wherein, based on the data volume and timeliness, each category of data is divided into at least real-time data with a large data volume, historical data with a large data volume, and real-time data with a small data volume; the real-time data with a large data volume is fused with the vehicle configuration through the Flink framework, the historical data with a large data volume is fused with the vehicle configuration through the Spark framework, and the real-time data with a small data volume is fused with the vehicle configuration through a lightweight scheduling framework, and the lightweight scheduling framework includes at least the Quartz framework.

[0015] In this implementation, data types are divided according to data volume and timeliness, and fused using different frameworks. This allows for efficient adaptation of data with different characteristics to vehicle configurations. This approach enhances data fusion flexibility, rationally allocates system resources, reduces resource usage, ensures stable fusion tasks, and improves data fusion efficiency and accuracy, laying a solid foundation for subsequent data applications.

[0016] In some embodiments, the entire life cycle of a rail vehicle includes an operation period, a maintenance period, and a dormant period. The method for fusing the attribute information of the system and equipment with the vehicle configuration also includes: prioritizing the fusion of the real-time data with the vehicle configuration during the operation period, fusing the historical data with the vehicle configuration during the maintenance period, and archiving the fused data to a cold database during the dormant period.

[0017] In this embodiment, data fusion and processing are targeted based on the characteristics of the rail vehicle's operation, maintenance, and dormancy periods throughout its lifecycle, achieving dynamic adaptation of data fusion. During the operation phase, real-time data is prioritized for integration, ensuring timely operational monitoring and decision-making. During the maintenance phase, historical data is integrated to provide a comprehensive basis for fault analysis and maintenance optimization. During the dormancy phase, data is archived to a cold database, saving storage resources while preserving the complete data trajectory. This improves data governance efficiency and resource utilization overall, providing strong support for full lifecycle management.

[0018] In some embodiments, a method for monitoring the integrity of fused data includes: setting integrity rules for the fused data, and checking whether the fused data contains all fields and attribute information specified by the integrity rules according to the integrity rules; if not, determining that there is a problem with the fused data; a method for monitoring the uniqueness of the fused data includes: generating a unique data fingerprint or a unique identifier based on key information in the fused data, determining whether the fused data is duplicate data according to the data fingerprint or the unique identifier, and deleting or marking the duplicate data; a method for monitoring the timeliness of the fused data includes: setting time thresholds corresponding to different types of fused data, determining whether the time of data update exceeds the corresponding time threshold, and if so, determining that there is a problem with the fused data; wherein, the key information includes at least the device or system number and the timestamp of the fused data.

[0019] In this embodiment, integrity rules are set to check fields and attribute information to ensure that no key information is missing from the fused data. Unique identifiers are generated based on key information to identify and process duplicate data, ensuring data uniqueness. Time thresholds are set based on different data types to monitor timeliness and detect lagging data. These methods rigorously control the quality of fused data from multiple dimensions. The high-quality data selected can provide reliable support for operational decision-making, inspection and maintenance throughout the lifecycle of rail vehicles, enhancing the value of data applications.

[0020] In some embodiments, when it is determined that there is a problem with the fused data, a data quality report is generated, which includes relevant information about the problem data. The relevant information about the problem data includes at least one of the data source, problem type, and time when the problem occurred. The data source includes at least the category of the data.

[0021] In this embodiment, the quality of the fused data is monitored in multiple dimensions and a data quality report is generated to provide high-quality data support for operational decision-making and inspection and maintenance in the full life cycle management of rail vehicles, thereby enhancing the value of data application.

[0022] In some embodiments, each node in the tree structure of the vehicle configuration includes at least trackside detection data, section information data, vehicle network data, and basic information of the vehicle or system or equipment; the method for establishing the vehicle configuration also includes: assigning a unique identifier to each node in the tree structure, and the unique identifier is used to distinguish different devices or systems.

[0023] In this embodiment, a unique identifier is assigned to each node in the vehicle configuration tree structure, clearly distinguishing different devices or systems. This ensures that each node has a clear directionality during data association, fusion, and subsequent management, avoiding confusion. This provides a precise indexing basis for associating system and device attribute information with data element standards and integrating it with vehicle configurations, improving the accuracy and efficiency of data fusion and facilitating the full lifecycle tracking and management of the device or system corresponding to each node.

[0024] On the second aspect, the present application provides a rail vehicle data governance system, including a data element standard and vehicle configuration construction module, a data association and fusion module, a data quality monitoring module and a data service module; the data element standard and vehicle configuration construction module is configured to: classify the data of various systems and equipment of the rail vehicle, and define field information corresponding to the attribute information of each category of data; establish a vehicle configuration, and the vehicle configuration presents the hierarchical relationship between the rail vehicle, each system and each equipment in each node in a tree structure; the data association and fusion module is configured to: according to the correspondence between the attribute information of each category of data in the data element standard and the defined field information, the attribute information of each category of data of the system and equipment is integrated with each node of the vehicle configuration; the data quality monitoring module is configured to: perform quality monitoring on the integrated data, and the dimensions of the quality monitoring include one or a combination of the integrity, uniqueness and timeliness of the data; the data service module is configured to: provide data services to the outside world for the data that meets the quality requirements screened out by the quality monitoring.

[0025] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, which can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0027] Figure 1 is a flowchart of the main steps of the rail vehicle data management method according to an embodiment of the present application; Figure 2 It is a schematic diagram of the overall architecture of rail vehicle data management and post-management applications according to the implementation method of this application. DETAILED DESCRIPTION

[0028] In order to make the technical problems, technical solutions and beneficial effects to be solved by this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0029] In the embodiments of the present application, prefixes such as "first" and "second" are used only to distinguish different description objects and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a restriction on the described objects. For the statement of the described objects, please refer to the description in the context of the claims or embodiments, and the use of such prefixes should not constitute an unnecessary restriction. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "plurality" is two or more.

[0030] The following describes the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents "or." For example, A / B can represent A or B. "And / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, or B exists alone.

[0031] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection of some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0032] In this application, the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0033] With the booming manufacturing industry, the intelligence of manufactured products continues to increase, and the amount of data generated during production and use is exploding. For example, rail vehicles, their data covers the entire life cycle, from design, manufacturing, testing, operation and maintenance to scrapping. Managing this data throughout its life cycle is crucial to the progress and sustainable development of the entire industry.

[0034] However, due to the lack of a unified data standards system, data from different systems lacks standardization in terms of format, interface, and semantics, making data correlation and integration difficult. For example, data formats vary significantly across devices provided by different vendors, and the definition and representation of the same concept vary across different systems. Furthermore, data storage is fragmented, with a variety of storage methods coexisting, including relational, time-series, and non-relational databases. For example, workshop data is stored in relational databases, Industrial IoT data in time-series databases, and field data in non-relational databases. This inconsistency in underlying storage complicates data correlation and hinders efficient fusion analysis.

[0035] The present invention provides a method for the full life cycle management of rail vehicle data, which is used to solve many problems of current rail vehicle data such as decentralized storage, inconsistent standards, high fusion difficulty, and poor fusion quality.

[0036] In order to solve the above problems, this application proposes a rail vehicle data governance method. By optimizing the entire process from data collation to utilization, it effectively solves problems such as scattered data storage, inconsistent standards, high integration difficulty, and poor quality. It provides reliable data support for operational decision-making, inspection and maintenance in the full life cycle management of rail vehicles, and improves the value of data application and the efficiency of rail vehicle management.

[0037] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0038] As attached Figures 1 to 2 As shown, in an exemplary embodiment of the present application, the rail vehicle data governance method mainly includes the following steps: (1) construction of data element standards and vehicle configuration (SBOM); (2) fusion of data and vehicle configuration; (3) data quality monitoring and service.

[0039] In some embodiments, the step of constructing the data element standard and the vehicle configuration includes creating the data element standard and establishing the vehicle configuration.

[0040] In some embodiments, a method of creating a data element standard includes: S1: For different types of rail vehicles, with systems and equipment as the center, combined with various business systems related to systems and equipment, manually sort out the data of systems and equipment from multiple stages such as procurement, installation and commissioning, operation, maintenance and replacement.

[0041] Information at different stages is often stored in different business systems. For example, equipment supplier information at the procurement stage may come from the procurement management system, production parameters at the manufacturing stage may come from the production execution system, fault records at the operation and maintenance stage may come from the operation and maintenance management system, trackside inspection data may come from the trackside inspection system, and section yard informationization data may come from the section yard management system, etc. Therefore, combining multiple business systems can ensure the comprehensiveness of the sorted data.

[0042] S2: Classify the data of each system and device of the rail vehicle according to the data characteristics of each system and device of the vehicle. The classified data categories include at least trackside inspection data, section yard information data, vehicle network data and basic information of the rail vehicle.

[0043] S3: Based on the data characteristics of each vehicle system and device, define the field information corresponding to the attribute information of each category of data.

[0044] In some embodiments, the method of defining field information corresponding to attribute information of each category of data includes: defining the meaning of each field, defining the length of each field, and determining the information type of each field.

[0045] Attribute information for various types of system and device data refers to detailed information describing the characteristics, status, and parameters of systems and devices throughout their lifecycles. This covers the entire lifecycle of equipment, from procurement, installation and commissioning, operational replacement, to scrapping, including information from each stage of design, manufacturing, operation, and maintenance. Examples include: device model (string type, 20 bits in length), rated power (floating point), design life (integer), manufacturer, production date, quality inspection report number, real-time operating temperature (floating point), vibration frequency (floating point), energy consumption data (floating point), speed (km / h), and location (latitude and longitude / route section).

[0046] Defining the meaning of each field clarifies the specific attribute information it represents, ensuring consistent understanding of the same field across different systems and personnel. For example, in the traction motor equipment data for rail vehicles, the "Motor Model" field is defined as "the unique model identifier assigned by the manufacturer to the motor, used to distinguish motors of different specifications and performance." By defining clear meanings, data misunderstandings and association errors caused by ambiguous semantics can be avoided.

[0047] Define the length of each field to set reasonable length limits based on the actual needs of the data stored in the field to ensure standardized and efficient data storage. For example, for fixed-format fields such as "Device Number," set the length precisely according to encoding rules (e.g., 12 digits) to ensure consistent data format and facilitate subsequent retrieval and matching.

[0048] The information type of each field is defined based on the characteristics of the data it carries, ensuring data accuracy and operability. For example, fields that reflect numerical values, such as "rated power" and "operating temperature," are defined as floating-point types to accurately record decimal values. Textual information, such as "motor model" and "device status description," are defined as strings.

[0049] In some embodiments, the field information includes the operating status field of the system and equipment, and the method of defining the field information corresponding to the attribute information of each category of data also includes: defining the operating status field of the key system and equipment as not allowed to be empty to ensure the integrity of the data.

[0050] The key systems and equipment include those that affect driving safety, as well as core functional systems and equipment. The systems and equipment that affect driving safety include traction motors, braking systems, and signal systems, while the core functional systems and equipment include bogies and on-board network systems.

[0051] In some embodiments, a method for establishing a vehicle configuration includes establishing a complete vehicle configuration for different types of rail vehicles, wherein the vehicle configuration presents a hierarchical relationship between various systems and devices of the rail vehicle in a tree structure. The tree structure further includes nodes, each node including a vehicle node, nodes for various systems of the vehicle, and nodes for various devices in the system. The tree structure presents the hierarchical relationship between the rail vehicle, various systems, and various devices through the nodes.

[0052] Exemplarily, the hierarchical and node relationships of the tree structure are as follows: the first layer is the vehicle, and there is only one node "vehicle" in the first layer; the second layer is the system, and the nodes of the second layer are the systems, such as "traction system", "braking system", etc.; the third layer is the equipment, and the nodes of the third layer are the equipment units, such as "car body", "traction device", "bogie", "air conditioning", "lighting", etc.

[0053] In some embodiments, each node includes trackside detection data, section information data, vehicle network data, and basic information of the vehicle, system or equipment.

[0054] It should be understood that the data information included in each node is a data association designed based on the tree structure of the vehicle configuration. This establishes a preset correspondence between the nodes in the tree structure and various data sources—in other words, pre-defining which types of data should be associated with a particular device or system, providing a target and framework for subsequent data fusion. For example, the traction system needs to be associated with its basic parameters (model, power), real-time vehicle network operation data (temperature, speed), trackside inspection status data (wear and tear), and maintenance records at the yard.

[0055] In some embodiments, vehicle configurations are created by importing data from various business systems related to systems and devices. By importing the existing vehicle configuration tree structure, a clear index framework is provided for data integration and management, enabling accurate and rapid location and management of data for each system and device.

[0056] In some embodiments, the method for establishing a vehicle configuration further includes: assigning a unique identifier to each node in the tree structure, where the unique identifier is used to distinguish different devices or systems.

[0057] For example, the unique identifier can be a string code generated according to rules such as the level, type, and model of the vehicle equipment. For example, it can be composed in the format of "vehicle code-system number-equipment serial number", such as "CL01-QY02-DJ03", where "CL01" represents the vehicle, "QY02" represents the traction system, and "DJ03" represents the third motor equipment in the system. The code directly reflects the hierarchical association of the nodes.

[0058] The unique identifier ensures that each node is unique in the entire tree structure, thus avoiding confusion between different devices or systems.

[0059] Leveraging the tree structure of vehicle configurations and the standardized definition of data element standards, this system automatically identifies and matches data from different systems and devices, improving integration efficiency and accuracy. Compared to traditional methods, this eliminates the need for manual data association processing, significantly reducing labor costs and error rates. Furthermore, it enables deeper exploration of data connections between subsystems, providing more comprehensive decision support for rail transit operations.

[0060] In some embodiments, a data storage stage is included between the steps of constructing the data element standard and the vehicle configuration and the step of fusing the data with the vehicle configuration.

[0061] Data from various systems and devices is scattered across the underlying databases of different business systems, stored in diverse ways. This data lacks unified standards for format, interface, and semantics, making it difficult to connect and integrate the data. Therefore, during the data storage phase, it is necessary to optimize the unified storage of data from multiple business systems and devices, as well as from multiple databases.

[0062] Specifically, during the data storage stage, data compression technologies, such as LZ77 coding, Huffman coding and other compression algorithms, are used to compress the dispersedly stored data to reduce the storage space of the data while maintaining the integrity and recoverability of the data.

[0063] In some embodiments, during the data storage stage, column storage is adopted to centrally store all data of the same field (such as the operating temperature of all devices), and store data of different fields separately. This optimizes the data storage structure and improves data reading efficiency, which is particularly suitable for large-scale data processing scenarios.

[0064] Through data compression technology and column storage, the storage pressure problem caused by the scattered and large amount of data storage is solved, while the efficiency of data storage and reading is improved.

[0065] In some embodiments, in the step of fusing data with vehicle configuration, it includes associating the attribute information of each category of data with the data element standard, and according to the correspondence between the attribute information of each category of data and the defined field information in the data element standard, fusing the attribute information of each category of data of the system and equipment with the corresponding nodes in the vehicle configuration.

[0066] In some embodiments, the method of fusing data with vehicle configuration includes: creating a data table for each system node and device node in the vehicle configuration according to the created data element standard, the data table including fields corresponding to all attribute information of the system and the device, and the fields are used to store the corresponding attribute information; filling the attribute information of each category of data into the corresponding fields in the data table.

[0067] For example, according to the data element standard, attributes for traction motors include motor model (string type, 20 bits in length), rated power (floating point type), and operating temperature (floating point type). When creating a data table, corresponding fields are set to store these attribute data. This association and binding method achieves a one-to-one correspondence between system equipment data and standards, laying a solid foundation for standardized data storage and management.

[0068] In some embodiments, the method of fusing attribute information of each category of data with vehicle configuration includes: Selecting a framework for integrating each category of data with the vehicle configuration based on the volume and timeliness of each category of data; wherein, based on the volume and timeliness of each category of data, each category of data is divided into at least real-time data with a large volume of data, historical data with a large volume of data, and real-time data with a small volume of data; The Flink framework is used to integrate large amounts of real-time data with vehicle configurations, the Spark framework is used to integrate large amounts of historical data with vehicle configurations, and the lightweight scheduling framework is used to integrate small amounts of real-time data with vehicle configurations. The lightweight scheduling framework includes at least the Quartz framework.

[0069] IoV data is characterized by large volumes and high timeliness. Therefore, the Flink framework was chosen to integrate IoV data with vehicle configurations. Leveraging Flink's efficient stream processing capabilities, it can parse, clean, and integrate IoV data in real time, quickly identifying abnormal data and issuing timely warnings.

[0070] Trackside inspection data, while voluminous, has relatively low timeliness requirements. Most of this data consists of historically accumulated inspection data, such as track geometry and signal equipment status data. For this type of data, the Spark framework is used to fuse trackside inspection data with vehicle configurations. With its powerful distributed computing capabilities, Spark can batch process massive amounts of trackside inspection data, accelerating the data fusion process through cluster computing. For example, when analyzing long-term track deformation trends, Spark can quickly access large amounts of historical inspection data, perform data analysis, and build models, providing a scientific basis for track maintenance.

[0071] Depot information data, while small in volume, requires high timeliness. This primarily involves real-time operational data related to vehicle scheduling and maintenance within the depot. To address this, a lightweight scheduling framework is used to integrate depot information data with vehicle configurations. This framework is lightweight, flexible, and efficient, enabling rapid response to changes in depot data and precise scheduling and management of depot operations. For example, when a vehicle enters the depot for maintenance, the lightweight scheduling framework can quickly schedule maintenance tasks based on real-time maintenance resources and vehicle status, improving depot operational efficiency.

[0072] In some embodiments, the method of fusing the attribute information of each category of data with the vehicle configuration also includes: analyzing the data, intelligently and automatically matching the appropriate fusion framework based on the characteristics of each category of data analyzed, and triggering the fusion process based on preset fusion rules to automatically fuse the attribute information of the system and equipment with the vehicle configuration.

[0073] In some embodiments, a multi-source data synchronization and calibration mechanism is introduced during the data fusion process. Timestamp alignment and data calibration algorithms are used to ensure the temporal consistency of data from different systems and devices. The timestamp alignment method is to convert the timestamps of data from different systems and devices into a consistent format and accuracy with a unified time base (such as a standard clock signal) as a reference. In cases where there are still slight deviations in the timestamps or the data collection frequencies are different, corrections are made through algorithms. For example, for real-time operation data and historical monitoring data, an interpolation algorithm (supplementing reasonable data at missing time points) or a sliding window algorithm (matching data within a certain time range) is used to eliminate time misalignment caused by acquisition delays and frequency differences, so that data from different systems and devices accurately correspond in the time dimension.

[0074] The entire life cycle of a rail vehicle includes the operation period, maintenance period, and dormancy period. The operation period refers to the stage when the rail vehicle is in actual operation. During this period, the vehicle continuously generates real-time operating data, such as speed, location, and equipment status. The maintenance period refers to the stage when the vehicle enters the depot for inspection and maintenance. This period mainly involves historically accumulated test data and maintenance operation data. The real-time requirements are relatively low, but the data volume is large. The dormancy period refers to the stage when the vehicle is temporarily not involved in operation and does not require maintenance. During this period, the amount of data generated is greatly reduced, and the main focus is on the storage and management of historical data.

[0075] Therefore, in some embodiments, the method of fusing attribute information of each category of data with the vehicle configuration further includes: During the operational phase, IoV data and vehicle configuration are prioritized for integration to meet real-time requirements. For example, real-time calculation of average train speed and energy consumption indicators is performed to provide immediate support for operational decisions. During the maintenance phase, trackside inspection data and vehicle configuration are integrated to avoid excessive resource consumption during operational hours and provide data support for maintenance. During the dormant phase, the integrated data is archived to a cold database, freeing up computing resources and reducing storage costs. This ensures long-term data traceability and preserves a complete data foundation for subsequent data analysis or historical queries.

[0076] In some embodiments, the method of fusing attribute information of each category of data with vehicle configuration further includes: during the data fusion process, performing streaming processing optimization on data streams classified as real-time data (such as Internet of Vehicles data).

[0077] Specifically, the Flink framework enables efficient parsing, cleaning, and integration of real-time data streams, optimizes data stream processing logic, reduces data latency, and ensures the high timeliness of real-time data.

[0078] For example, for train operating status data, sliding window technology is used to calculate average speed and energy consumption indicators in real time, providing real-time support for operational decision-making. To meet the high security requirements for rail transit signal data, a data encryption module is embedded in the data fusion layer to ensure encryption of the train-to-ground communication link during the data transmission phase, preventing signal data from being stolen or tampered with. During the data storage phase, key technical parameters (such as train position and route information) are encrypted at the field level to meet the requirements of the "Information Security Technology - Network Security Technical Requirements for Rail Transit Signal Systems," filling a gap in rail transit data security technology. For cross-network data governance, a cross-operator data interaction protocol is designed to support the rapid alignment of data for vehicle configurations on different lines. Custom encoding rules for cross-network equipment are also defined, and data semantic conversion middleware is developed. This middleware, through preset conversion logic and mapping relationships, converts the proprietary data of different operators into a unified standard data format, ensuring consistency in data format across the network.

[0079] In some embodiments, the data quality monitoring and service steps include quality monitoring of the fused data and providing data services to external parties.

[0080] In some embodiments, the dimensions of quality monitoring of the fused data include one or more of the integrity, uniqueness, and timeliness of the data.

[0081] The method for monitoring the integrity of fused data includes: setting data integrity rules, and checking whether the fused data contains all fields and attribute information specified by the integrity rules according to the integrity rules. If so, the fused data is judged to be complete; if not, the fused data is judged to have problems. For example, for train operation status data, check whether key fields such as speed, position, and equipment status are included, check whether the vehicle operation speed (km / h), real-time position (latitude and longitude / line section), equipment status (traction / braking status, door switch status, on-board signal system status), energy consumption data (power / power consumption), car temperature and humidity, passenger flow, wheelset wear data, track geometry (gauge / height deviation) and other real-time monitoring indicators are included. If missing, the data is judged to be incomplete and the relevant information is recorded.

[0082] The method for monitoring the uniqueness of the fused data includes: generating a unique data fingerprint or a unique identifier based on key information in the fused data, determining whether the fused data is duplicate data based on the data fingerprint or the unique identifier, and deleting or marking the duplicate data.

[0083] Key information includes at least the device or system number and the timestamp of the fused data. For example, for device failure information, a unique data fingerprint is generated using key information such as the device number and failure time. If duplicate fingerprints are found, the data is considered duplicate.

[0084] The method for monitoring the timeliness of fused data includes setting time thresholds corresponding to different types of data and determining whether the updated time of the fused data exceeds the corresponding time threshold. If not, the fused data is considered updated in a timely manner. If so, it is determined that there is a problem with the fused data and that it is not updated in a timely manner. For example, for real-time train operation data, the time threshold is set to 5 seconds, requiring an update every 5 seconds. If it does not update for more than 5 seconds, it is determined that the data may be abnormal.

[0085] In some embodiments, when it is determined that there is a problem with the fused data, a data quality report is generated. The data quality report includes relevant information about the problem data. The relevant information about the problem data includes at least one of the data source, problem type, and time when the problem occurred. The data source includes at least the category of the data.

[0086] In some embodiments, the data governance system further has the function of backtracking dirty data. The method of backtracking dirty data includes: Each stage of data governance processing is versioned, recording the status of different versions for each piece of data. For example, if a piece of train operation status data experiences an anomaly due to a fault, the original unprocessed version and the cleaned, corrected version are retained, along with information such as the time of the fault, the time of resolution, and the operator for each version. Version management allows for the retrieval of every change to dirty data during processing.

[0087] At the same time, a full-process log of data governance is recorded, including the data source, processing process, and quality monitoring results. Specifically, the data source is recorded, including the original acquisition system (such as vehicle-to-vehicle sensors and trackside inspection equipment) of the dirty data, the acquisition time, the transmission path, etc., to locate the source of the data. The processing process is recorded, including the framework, conversion rules, and fusion algorithm used in the data fusion phase, to trace whether there are any problems with the processing logic. The quality monitoring results are recorded, including the time when the dirty data was identified, the quality rules triggered (integrity loss, uniqueness conflict, etc.), the basis for quality monitoring, etc., to clarify the cause of the anomaly.

[0088] When dirty data is discovered, the system can trace back to the original state and processing process of the problematic data through its version records and full-process logs, thereby confirming how the data problem occurred.

[0089] In some embodiments, after completing data quality monitoring, problematic data is marked, traced back, and processed, including correction or elimination, to filter out data that meets quality requirements.

[0090] In some embodiments, the data that meets the quality requirements is screened out and provided as data services to the outside world on a system or device basis. Specifically, through standardized interfaces, various types of rail transit systems or businesses can efficiently apply governed data to provide data support for various types of rail transit systems or businesses. For example, equipment operation status data is provided to the train intelligent operation and maintenance system to help operation and maintenance personnel promptly discover potential fault hazards; real-time position and speed data is provided to the train dispatching system to optimize train operation scheduling plans. In the industrial field, based on the results of equipment data governance, equipment production history and current status data are provided to the production management system to assist in production plan formulation; and product quality data throughout the entire process is provided to the quality inspection system for quality traceability and improvement.

[0091] In some embodiments, the rail vehicle data management method further includes a step of applying the data after management.

[0092] In some embodiments, the method for applying the governed data further includes: Data analysis supports rail transit operational decision-making. Operational reports are generated based on vehicle operation data, system and equipment health data, and energy consumption data. For example, by analyzing train operating efficiency and energy consumption data, train operation scheduling plans can be optimized; by analyzing equipment failure data, equipment procurement strategies can be optimized to improve operational efficiency and economic benefits.

[0093] Methods for optimizing train operation scheduling by analyzing train operation efficiency and energy consumption data include: An energy consumption optimization model is constructed based on historical and real-time monitoring data. First, multiple categories of data are collected from the treated systems and equipment, including vehicle operation data (operating speed, acceleration, braking frequency, traction power, mileage, wheel speed), equipment status data (traction inverter efficiency, air conditioning energy consumption ratio, door opening and closing times), environmental data (temperature, humidity, wind speed and direction, route slope, curve radius), and passenger flow data (real-time passenger density in carriages and passenger boarding and alighting at stations).

[0094] Machine learning algorithms (such as gradient boosting trees and long-short-term memory networks) and operations research optimization methods (dynamic programming and genetic algorithms) are used to preprocess the collected data, including cleaning and normalization, to extract key features. These key features include the correlation between speed and energy consumption at different slopes and the variation in air conditioning energy consumption at different passenger flow rates. The historical data is then divided into training and test sets. Through training, model parameters are optimized, enabling the energy optimization model to accurately learn the complex nonlinear relationships between various data types and energy consumption, and predict energy consumption trends under different operating conditions.

[0095] During the technical solution implementation phase, real-time data is transmitted to the dispatching center via the train-to-ground communication system. The energy consumption optimization model dynamically calculates this data and generates the optimal operating strategy. For example, on straight sections, the train operates at an economic speed calculated by the model, taking into account real-time wind speed and passenger flow. Before approaching a hilly section, power output is precisely adjusted in advance based on slope and curve information to avoid power waste. During braking, energy is recovered through optimized braking strategies. Simultaneously, digital twin technology is used to construct virtual train operation scenarios. The optimization strategy is simulated and verified before being issued to the train for execution. Energy consumption data after strategy execution is continuously collected and fed back into the energy consumption optimization model for iterative optimization. This forms a closed-loop management system of "data collection - model calculation - strategy execution - effect feedback - model optimization," enabling refined and dynamic management of rail transit vehicle energy consumption.

[0096] In some embodiments, the method for applying the governed data further includes: Develop a data visualization platform based on a rail network, combining rail transit vehicle operating data with geographic information. Maps display information such as vehicle location, operating status, and equipment health. For example, a graphical display of vehicle energy consumption distribution and a track diagram show vehicle operation paths provide intuitive decision support for operations management.

[0097] In some embodiments, the method for applying the governed data further includes: The full life cycle historical data of the managed systems and equipment (including operating parameters, fault logs, maintenance records, etc.) is used to train machine learning algorithms (such as random forests and LSTM neural networks) to build fault prediction models. Sensors (temperature, vibration, current sensors, etc.) deployed at key parts of the equipment collect and access operating data in real time (such as the bearing temperature of the traction motor, the stator winding vibration frequency, the load speed, etc.). After data cleaning and feature extraction, the data is input into the model for analysis. The fault risk probability is calculated in real time based on the preset threshold and algorithm output. When parameter abnormalities are detected (such as continuous exceeding of the effective value of vibration or exceeding of the temperature rise rate), early warning information is automatically generated and pushed to operation and maintenance personnel through multiple channels such as system pop-ups and SMS notifications, so that they can plan maintenance plans in advance, accurately locate potential fault points, effectively reduce the probability of sudden equipment shutdowns, and reduce operational interruption time and maintenance costs caused by faults.

[0098] In some embodiments, the method for applying the governed data further includes: Optimize maintenance plans based on system and equipment operating data and failure prediction models. By analyzing system and equipment operating time and failure probability, we develop personalized maintenance plans. For example, for equipment with high failure risk, we schedule preventive maintenance in advance to minimize the impact of sudden failures on operations.

[0099] In some embodiments, the method for applying the data after governance further includes data visualization: Develop a user-friendly data interaction interface to facilitate quick data query and analysis by operations and management personnel. Through a simple interface design, provide functions such as data query, report generation, and real-time monitoring. For example, dashboards display key indicators and charts show equipment operating trends, making data more convenient to use.

[0100] In some embodiments, the rail vehicle data governance method further includes protecting the data security of the governed data: Introduce data security and backup mechanisms to ensure data reliability and security. Use encryption technology to protect data during transmission and storage. Also, regularly back up data and employ off-site backup strategies to prevent data loss. For example, back up critical data to cloud storage to ensure data availability during disaster recovery.

[0101] In some embodiments, the present application also provides a rail vehicle data governance system for implementing the above-mentioned rail vehicle data governance method, which includes a data element standard and vehicle configuration construction module, a data association and fusion module, a data quality monitoring module and a data service module.

[0102] The data element standard and vehicle configuration construction module is configured to: classify the data of various systems and equipment of the rail vehicle, define field information corresponding to the attribute information of each category of data; establish a vehicle configuration, and the vehicle configuration presents the hierarchical relationship between the rail vehicle, various systems and various equipment as nodes in a tree structure; The data association and fusion module is configured to: fuse the attribute information of each category of data of the system and the device with each node of the vehicle configuration according to the correspondence between the attribute information of each category of data and the defined field information in the data element standard; The data quality monitoring module is configured to: perform quality monitoring on the fused data, wherein the dimensions of the quality monitoring include one or a combination of the integrity, uniqueness, and timeliness of the data; The data service module is configured to provide data services to external parties based on the data that meets the quality requirements and is screened out by the quality monitoring.

[0103] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be covered and fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A rail vehicle data management method, characterized in that: include: Creating a data element standard includes: classifying data of various systems and equipment of a rail vehicle, wherein the data of each system and equipment includes attribute information describing various characteristics, states, and parameters of the system and equipment during its life cycle, and defining field information corresponding to the attribute information of each category of data; Establishing a vehicle configuration, wherein the vehicle configuration is a tree structure, the tree structure including a vehicle node, each system node of the vehicle, and each device node in the system, the tree structure presenting a hierarchical relationship between the rail vehicle, each system, and each device through each node; The data and vehicle configuration fusion includes: according to the correspondence between the attribute information of each category of data and the defined field information in the data element standard, fusing the attribute information of each category of data of the system and the equipment with each node of the vehicle configuration; Performing quality monitoring on the fused data, wherein the dimensions of quality monitoring include one or a combination of data integrity, uniqueness, and timeliness; Through the quality monitoring, data that meets the quality requirements is screened out and provided as a data service to the outside world.

2. The rail vehicle data management method according to claim 1, characterized in that: The method for defining field information corresponding to attribute information of each category of data includes: Define the meaning of each field, define the length of each field, and define the information type of each field.

3. The rail vehicle data management method according to claim 2, characterized in that: The field information includes a system and device operating status field, and the method of defining field information corresponding to attribute information of each category of data further includes: defining the operating status field of key systems and devices as not allowed to be empty; The key systems and equipment include systems and equipment that affect driving safety, as well as core functional systems and equipment. The systems and equipment that affect driving safety include at least traction motors, braking systems, and signal systems. The core functional systems and equipment include at least bogies and on-board network systems.

4. The rail vehicle data management method according to claim 1, characterized in that: The method for fusing the data with the vehicle configuration includes: Creating a data table for each system node and device node in the vehicle configuration according to the created data element standard, wherein the data table includes fields corresponding to all attribute information of the system and the device, and the fields are used to store the corresponding attribute information; Fill in the attribute information of each category of data into the corresponding fields in the data table.

5. The rail vehicle data management method according to claim 1, characterized in that: The method for fusing the attribute information of each category of data with the vehicle configuration includes: Based on the volume and timeliness of each data category, a framework for integrating each data category with the vehicle configuration is selected; According to the data volume and timeliness, each category of data is divided into at least real-time data with a large data volume, historical data with a large data volume, and real-time data with a small data volume; The large-volume real-time data is fused with the vehicle configuration through the Flink framework, the large-volume historical data is fused with the vehicle configuration through the Spark framework, and the small-volume real-time data is fused with the vehicle configuration through a lightweight scheduling framework, which at least includes a Quartz framework.

6. The rail vehicle data management method according to claim 5, characterized in that: The entire life cycle of a rail vehicle includes an operation period, a maintenance period, and a dormant period. The method for integrating the attribute information of the system and equipment with the vehicle configuration further includes: During the operation period, the real-time data is preferentially integrated with the vehicle configuration; during the maintenance period, the historical data is integrated with the vehicle configuration; and during the dormant period, the integrated data is archived to a cold database.

7. The rail vehicle data management method according to any one of claims 1 to 6, characterized in that: The method for monitoring the integrity of the fused data includes: setting integrity rules for the fused data, and checking whether the fused data contains all fields and attribute information specified by the integrity rules according to the integrity rules, and if not, determining that there is a problem with the fused data; The method for monitoring the uniqueness of the fused data includes: generating a unique data fingerprint or a unique identifier based on key information in the fused data, determining whether the fused data is duplicate data based on the data fingerprint or the unique identifier, and deleting or marking the duplicate data; The method for monitoring the timeliness of the fused data includes: setting time thresholds corresponding to different types of fused data, determining whether the time of data update exceeds the corresponding time threshold, and if so, determining that there is a problem with the fused data; The key information at least includes the device or system number and the timestamp of the fused data.

8. The rail vehicle data management method according to claim 7, characterized in that: include: When it is determined that there is a problem with the fused data, a data quality report is generated, which includes relevant information of the problem data. The relevant information of the problem data includes at least one of the data source, problem type, and time when the problem occurred. The data source at least includes the category of the data.

9. The rail vehicle data management method according to claim 1, characterized in that: Each node in the tree structure of the vehicle configuration includes at least trackside detection data, section yard information data, vehicle network data, and basic information of the vehicle, system or equipment; The method for establishing a vehicle configuration further includes: assigning a unique identifier to each node in the tree structure, wherein the unique identifier is used to distinguish different devices or systems.

10. A rail vehicle data management system, characterized in that: include: The data element standard and vehicle configuration construction module is configured to: classify the data of various systems and equipment of the rail vehicle, define field information corresponding to the attribute information of each category of data; establish a vehicle configuration, and the vehicle configuration presents the hierarchical relationship between the rail vehicle, various systems and various equipment as nodes in a tree structure; a data association and fusion module configured to: fuse attribute information of each category of data of the system and the device with each node of the vehicle configuration according to a correspondence between attribute information of each category of data and defined field information in the data element standard; A data quality monitoring module is configured to: perform quality monitoring on the fused data, wherein the dimensions of the quality monitoring include one or a combination of the integrity, uniqueness, and timeliness of the data; The data service module is configured to provide data services to the outside world based on the data that meets the quality requirements and is screened out by the quality monitoring.