Rail transit data management method and device, equipment, storage medium and product
By employing data protocol conversion, knowledge graph matching, and encryption processing, the adaptation and security issues of rail transit data governance were resolved, achieving efficient and usable data governance results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, rail transit data governance relies on fixed templates, which are difficult to adapt to the unique professional terminology and business logic of various rail transit systems, resulting in poor availability, low efficiency, and difficulty in supporting operation and maintenance.
By acquiring rail transit data through data communication protocols, the data is converted into target data in a preset format. Then, it is matched with operation and maintenance knowledge entities in a pre-built rail transit knowledge graph to generate operation and maintenance semantic tags. The data is then classified and encrypted. Data cleaning and encryption are performed using data cleaning reinforcement learning models and conditional generative adversarial network models. Finally, the data is stored in a distributed file system and blockchain network.
It enables efficient and accurate management of rail transit data, is compatible with both new and old systems, improves data availability and security, and supports operation and maintenance decisions and scheduling control.
Smart Images

Figure CN121807976A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data governance technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for rail transit data governance. Background Technology
[0002] As rail transit networks become increasingly intelligent, the scale of data assets is also growing exponentially. The data assets of rail transit systems contain a wealth of multi-dimensional information on operations, equipment, and passengers, which is of great significance to rail transit operation and maintenance. Therefore, the demand for data governance of such data assets is becoming increasingly widespread.
[0003] Data governance is a data engineering process that transforms raw data into high-quality, secure, and available data. In related technologies, the governance of rail transit data relies on automated programs with fixed templates. However, the data structures of various rail transit systems are not uniform, and the rail transit field has unique terminology, business rules, and logical relationships. Such programs are difficult to adapt well, resulting in poor availability and low efficiency in data governance. Consequently, these data assets cannot better support the operation and maintenance of the rail transit network. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for rail transit data governance that can achieve high availability and high efficiency in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for rail transit data governance, including:
[0006] In response to a data governance request, the system acquires rail transit data from multiple rail transit systems and determines the data communication protocol used by each of the rail transit systems to transmit the rail transit data.
[0007] According to the data communication protocols described above, the rail transit data are converted into target data in a preset data format;
[0008] Match the operation and maintenance knowledge entities associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entities;
[0009] Based on the operation and maintenance semantic tags, the rail transit operation and maintenance events to which the target data belongs are classified to obtain the classification results;
[0010] Based on the classification results, the data governance results of the rail transit data are output.
[0011] In one embodiment, the operation and maintenance knowledge entity includes equipment failure knowledge entities, traffic line knowledge entities, and operation and maintenance measure knowledge entities associated with rail transit equipment. The step of classifying the rail transit operation and maintenance events to which the target data belongs, based on the operation and maintenance semantic tags, to obtain classification results includes:
[0012] Based on the operation and maintenance semantic tags, query the multiple entity association relationships of the equipment fault knowledge entity, the transportation line knowledge entity, and the operation and maintenance measures knowledge entity in the rail transit knowledge graph;
[0013] Based on the multiple entity associations, the target associations of the target data with the target rail transit equipment are determined; wherein, the target associations are used to characterize the association between the faults occurring in the target rail transit equipment and the operation and maintenance measures.
[0014] Based on the target association, the rail transit operation and maintenance events to which the target data belongs are classified to obtain the classification results.
[0015] In one embodiment, the step of outputting the data governance result of the rail transit data based on the classification result includes:
[0016] Obtain multiple user role attribute information and determine the data access permissions corresponding to each user role attribute information;
[0017] The classification results are encrypted using different data access permissions to obtain multiple encrypted classification results.
[0018] Based on the classification results obtained after encrypting multiple data sets, the data governance results of the rail transit data are output.
[0019] In one embodiment, the step of performing different data encryption processes on the classification results according to different data access permissions includes:
[0020] The data access permissions and the classification results are input into a pre-built conditional generative adversarial network model;
[0021] Using different data access permissions as the generation conditions of the conditional generative adversarial network model, the conditional generative adversarial network model determines the target fields in the classification results that correspond to different data access permissions and need to be encrypted, and generates replacement fields for each target field;
[0022] According to different data access permissions, the target field in the classification result is replaced with the corresponding replacement field to obtain multiple encrypted classification results.
[0023] In one embodiment, the method further includes:
[0024] The data governance results are stored in a distributed file system, and a data feature code of the data governance results is generated.
[0025] In response to data operations on the data governance results in the distributed file system, obtain the operation record of the data operation;
[0026] The operation record and the data feature code are stored in a blockchain network.
[0027] In one embodiment, after converting the rail transit data into target data according to a preset data format based on the respective data communication protocols, the method further includes:
[0028] The target data is input into a pre-trained data cleaning reinforcement learning model. The data features of the target data are extracted by the data cleaning reinforcement learning model. Data cleaning rules are generated based on the data features. The target data is then cleaned according to the data cleaning rules.
[0029] Secondly, this application also provides a rail transit data governance device 20, comprising:
[0030] The input module is used to respond to data governance requests, acquire rail transit data from multiple rail transit systems, and determine the data communication protocol used by each of the rail transit systems to transmit the rail transit data.
[0031] The conversion module is used to convert the rail transit data into target data according to the preset data format based on the data communication protocols.
[0032] The generation module is used to match the operation and maintenance knowledge entities associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entities.
[0033] The classification module is used to classify the rail transit operation and maintenance events to which the target data belongs based on the operation and maintenance semantic tags, and obtain the classification results;
[0034] The output module is used to output the data governance results of the rail transit data based on the classification results.
[0035] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-mentioned rail transit data governance method.
[0036] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described rail transit data governance method.
[0037] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described rail transit data governance method.
[0038] The aforementioned rail transit data governance methods, devices, computer equipment, computer-readable storage media, and computer program products convert various rail transit data into target data in a preset data format according to various data communication protocols, achieving differentiated adaptation for different rail transit systems and standardizing various types of rail transit data into target data in a preset data format, ensuring compatibility with both new and old rail transit systems. Through a pre-constructed rail transit knowledge graph, a rail transit knowledge network is built. Operation and maintenance knowledge entities associated with the target data are matched within the rail transit knowledge graph. Operation and maintenance semantic tags are generated for the target data based on these entities, mining potential rail transit operation and maintenance semantic information from the target data to achieve preliminary translation. Then, based on the operation and maintenance semantic tags, the rail transit operation and maintenance events to which the target data belongs are classified, automatically adapting to the unique professional terminology, business rules, and logical relationships in the rail transit field. Ultimately, the raw, underlying rail transit data is transformed into highly available and reliable rail transit operation and maintenance event classification results, thereby achieving precise and efficient governance of rail transit data. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 A schematic diagram illustrating the application environment of a rail transit data governance method provided in an embodiment of this application;
[0041] Figure 2 A flowchart illustrating the steps of a rail transit data governance method provided in an embodiment of this application;
[0042] Figure 3 A flowchart of sub-steps of a rail transit data governance method provided in another embodiment of this application;
[0043] Figure 4 This is a structural block diagram of a rail transit data management device 40 provided in an embodiment of this application;
[0044] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0046] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0047] The rail transit data governance method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0048] In one exemplary embodiment, such as Figure 2 As shown, a method for rail transit data governance is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 202 to 210. Wherein:
[0049] Step 202: In response to the data governance request, obtain rail transit data from multiple rail transit systems, and determine the data communication protocol used by each of the rail transit systems to transmit the rail transit data;
[0050] Rail transit systems can include control systems and station management systems for rail transit such as subways and high-speed railways; rail transit data can include equipment data, operation record data, sensor data, operation logs, video stream data, and third-party system data; data communication protocols can include MQTT (Message Queuing Telemetry Transport), Modbus, HTTP (Hypertext Transfer Protocol), and other data communication protocols.
[0051] In practical implementation, multiple data interfaces can be established to communicate with the rail transit system, receive and store this rail transit data, and parse the data communication protocols used by the rail transit data. Simultaneously, a data buffer queue can be configured to transmit data with these rail transit systems to handle high-concurrency scenarios. The data buffer queue can adopt a circular buffer design, prioritizing real-time data streams in high-concurrency scenarios to ensure low-latency transmission of critical data. The circular buffer allows the system to quickly integrate with various new and old equipment in the rail transit system, is compatible with existing industrial control systems, reduces deployment costs, and simultaneously increases data throughput.
[0052] Step 204: According to the data communication protocols, convert the rail transit data into target data in a preset data format;
[0053] In practical implementation, a pre-defined protocol parsing engine can be used to convert the data into a unified pre-defined data format according to the data communication protocols of various rail transit data, such as JSON (JavaScript Object Notation), XML (eXtensible Markup Language), CSV (Comma-Separated Values), etc., to obtain target data with a unified pre-defined data format. This allows for differentiated adaptation to different rail transit systems (e.g., adapting to new and old rail transit systems), facilitating subsequent data processing or data sharing.
[0054] In some embodiments, after converting the rail transit data into target data according to a preset data format based on the data communication protocols, the method further includes:
[0055] The target data is input into a pre-trained data cleaning reinforcement learning model. The data features of the target data are extracted by the data cleaning reinforcement learning model. Data cleaning rules are generated based on the data features. The target data is then cleaned according to the data cleaning rules.
[0056] In the specific implementation, data features of the target data are extracted through a data cleaning reinforcement learning model, such as data value distribution and data type. Based on the data cleaning rules, reinforcement learning is performed to dynamically generate corresponding data cleaning rules, including missing value imputation, outlier detection and redundant field removal. Temporal correlation analysis is also introduced to optimize the cleaning strategy and perform data cleaning on the target data.
[0057] The data cleaning process in related technologies relies on fixed cleaning templates. However, sensors in rail transit systems come from a wide range of sources and are of diverse types. The data generated by different sensors vary significantly in terms of format, accuracy, and noise characteristics. Fixed cleaning templates cannot automatically adjust and adapt to these differences, resulting in poor data cleaning performance and potentially missing important information or introducing erroneous data.
[0058] In this embodiment, a data cleaning reinforcement learning model is used to dynamically generate data cleaning rules based on the data characteristics of the target data, thereby dynamically adapting to various complex scenarios in the rail transit system and effectively improving the reliability of the target data.
[0059] Step 206: Match the operation and maintenance knowledge entities associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entities;
[0060] The rail transit knowledge graph stores relevant operation and maintenance knowledge of rail transit through a graph (Map) structure. The graph structure consists of multiple triples of "Operation and Maintenance Knowledge Entity - Relationship - Operation and Maintenance Knowledge Entity" to reflect the complex knowledge network within rail transit. Operation and maintenance knowledge entities can include line information (e.g., line distribution, line geographic information), station information (e.g., station name, station location), and equipment information (e.g., equipment type, equipment ID, equipment status). Operation and maintenance semantic tags are semantic tags that reflect relevant operation and maintenance information of rail transit operations, possessing characteristics such as interpretability and readability. For example, specific forms of operation and maintenance semantic tags can be: "Line Name: Line A", "Station Name: Station B", etc.
[0061] In practice, various fields can be extracted from the target data, and then each field can be matched with the operation and maintenance knowledge entities to obtain the associated operation and maintenance knowledge entities. The operation and maintenance knowledge entities can then be semantically processed according to a predefined data dictionary to generate operation and maintenance semantic tags for the target data.
[0062] In some examples, the target data includes operational data for a specific device, specifically: {device_id: 0x3A1F, status: 0x05}. Matching the device_id field to the device ID knowledge entity with device ID "0x3A1F" in the rail transit knowledge graph determines that the device type is an escalator. Matching the status field to the device status knowledge entity with status code "0x05" in the rail transit knowledge graph determines that the device is malfunctioning. Therefore, the generated operation and maintenance semantic tags for the target data can include: "Device Type: Escalator" and "Device Status: Malfunction".
[0063] Step 208: Based on the operation and maintenance semantic tags, classify the rail transit operation and maintenance events to which the target data belongs, and obtain the classification results;
[0064] Rail transit operation and maintenance events refer to event information that reflects rail transit operation and maintenance business, specifically including information such as the nature of the event, the cause of the event, and the impact of the event. For example, a safety-related rail transit operation and maintenance event can be represented as: Event nature: Train automatic operation system malfunction; Event cause: Positioning and stopping control failure; Event impact: Train delay.
[0065] In practical implementation, the operation and maintenance information reflected by the operation and maintenance semantic tags can be directly used to classify the rail transit operation and maintenance events to which the target data belongs according to the preset classification rules; alternatively, the operation and maintenance information can be further searched and matched based on the operation and maintenance semantic tags, and this operation and maintenance information can be input into the classification model (such as the BERT model) to classify the rail transit operation and maintenance events to which the target data belongs.
[0066] By classifying the target data into rail transit operation and maintenance events, the underlying, raw rail transit data is transformed into highly available rail transit operation and maintenance event classification results. These classification results not only facilitate operation and maintenance personnel in carrying out relevant operation and maintenance work, but also facilitate various rail transit systems in carrying out relevant scheduling control or operation and maintenance decisions, thereby achieving data governance of rail transit data.
[0067] In some embodiments, the operation and maintenance knowledge entity includes equipment fault knowledge entities, traffic line knowledge entities, and operation and maintenance measure knowledge entities associated with rail transit equipment, such as... Figure 3 As shown, the classification of rail transit operation and maintenance events to which the target data belongs, based on the operation and maintenance semantic tags, to obtain classification results includes:
[0068] Sub-step 2082: Based on the operation and maintenance semantic tags, query the multiple entity association relationships of the equipment fault knowledge entity, the traffic line knowledge entity, and the operation and maintenance measures knowledge entity in the rail transit knowledge graph;
[0069] Sub-step 2084: Based on the multiple entity association relationships, determine the target association relationship of the target rail transit equipment associated with the target data; wherein, the target association relationship is used to characterize the association relationship between the faults that occur in the target rail transit equipment and the operation and maintenance measures;
[0070] Sub-step 2086: Based on the target association relationship, classify the rail transit operation and maintenance events to which the target data belongs, and obtain the classification result.
[0071] In this embodiment, the equipment fault knowledge entity is a knowledge entity that reflects equipment fault information, such as equipment model, equipment fault code, and equipment fault cause; the transportation line knowledge entity is a knowledge entity that reflects the line information of rail transit lines, such as line name and line geographic information topology; and the operation and maintenance measures knowledge entity is a knowledge entity that reflects the information related to operation and maintenance measures, such as operation and maintenance procedure text, fault handling measure text, and maintenance handling measure text.
[0072] In practical implementation, based on the semantic tags of operation and maintenance (O&M), the system queries the rail transit knowledge graph to find multiple entity relationships among equipment fault knowledge entities, transportation line knowledge entities, and O&M measure knowledge entities. For example, one equipment fault knowledge entity may be associated with multiple transportation line knowledge entities, or with multiple O&M measure knowledge entities, or there may be mutual relationships among these entities. Based on these relationships, confidence analysis and correlation analysis are used to determine the target association between the actual faults of the target rail transit equipment and the O&M measures, i.e., to identify the association path with the highest correlation. Through this target association, an accurate equipment-fault-response measure mapping is established to accurately classify the rail transit O&M events corresponding to the target rail transit equipment.
[0073] In some examples, the target data is: {Equipment: "TR_MTR_01", Code: "LOC_ERR", Location: "SEC_T18"}, generating the operation and maintenance semantic tags as: "Equipment: Traction motor No. 1 of train G101", "Fault: Traction loss", "Location: T18 of Line 1 tunnel". Based on these operation and maintenance semantic tags, the following entity relationships are found in the rail transit knowledge graph: the "Traction loss" equipment fault knowledge entity is associated with other equipment fault knowledge entities such as "Motor overheat protection fault" and "Power interruption", and is also associated with the operation and maintenance measure knowledge entity "Check main circuit"; the "Traction motor No. 1" equipment fault knowledge entity is associated with other equipment fault knowledge entities such as "TM-2100 asynchronous motor" and "Cooling fan bearing is prone to jamming under continuous high load"; the "Line 1 tunnel section T18" traffic line knowledge entity is associated with other traffic line knowledge entities such as "Maximum gradient section of the line".
[0074] Correlation analysis of these entity relationships reveals the following: Train G101 is located on the section of the line with the steepest gradient. Traction motor No. 1 is a TM-2100 asynchronous motor. This type of motor is prone to jamming under continuous high load, leading to motor overheating protection failure. The maintenance measure for this type of failure is to check the main circuit. Therefore, the target relationship can be determined as follows: the target rail transit equipment, the "TM-2100 asynchronous motor," experiences a "motor overheating protection failure," and the corresponding maintenance measure is "check the main circuit."
[0075] Step 210: Based on the classification results, output the data governance results of the rail transit data.
[0076] Data governance results refer to rail transit data after data governance, which may include classification results, as well as metadata of other related data or rail transit data.
[0077] In practical implementation, classification results can be output as data governance results, or they can be combined with other relevant data, such as encapsulating target data as metadata as data governance results. Data governance results can be packaged into standardized API services, message queues, or binary file packages, and metadata description documents can be automatically generated to provide corresponding data services.
[0078] The step of outputting the data governance results of the rail transit data based on the classification results includes:
[0079] Obtain multiple user role attribute information and determine the data access permissions corresponding to each user role attribute information;
[0080] The classification results are encrypted using different data access permissions to obtain multiple encrypted classification results.
[0081] Based on the classification results obtained after encrypting multiple data sets, the data governance results of the rail transit data are output.
[0082] In practical implementation, different user role attribute information corresponds to different user roles, such as operations and maintenance personnel and administrators; different user roles have different data access permissions. Therefore, different data encryption processes can be applied to the classification results based on different data access permissions. For example, different levels and encryption grades can be used to desensitize private data in the classification results and improve data security. For instance, for ordinary operations and maintenance personnel, the device IDs in the classification results are obfuscated and encrypted, but information such as the device fault type is retained, allowing operations and maintenance personnel to perform operations and maintenance work based on the device fault type; for administrators, various passenger privacy data in the classification results are encrypted, but information such as statistical feature distribution is retained, facilitating related analysis tasks for administrators.
[0083] In some embodiments, the step of performing different data encryption processes on the classification results according to different data access permissions includes:
[0084] The data access permissions and the classification results are input into a pre-built conditional generative adversarial network model;
[0085] Using different data access permissions as the generation conditions of the conditional generative adversarial network model, the conditional generative adversarial network model determines the target fields in the classification results that correspond to different data access permissions and need to be encrypted, and generates replacement fields for each target field;
[0086] According to different data access permissions, the target field in the classification result is replaced with the corresponding replacement field to obtain multiple encrypted classification results.
[0087] In its implementation, the Conditional Generative Adversarial Network (CGAN) model identifies target fields in the classification results that require encryption, such as device ID and user ID. Using different data access permissions as generation conditions, it generates alternative fields for these target fields. For example, for ordinary maintenance personnel, CGAN generates fuzzy alternative fields for device-related target fields in the classification results, such as device ID, while retaining information such as device fault type, allowing maintenance personnel to perform maintenance work based on the fault type. For management personnel, CGAN generates corresponding alternative fields for various passenger privacy data target fields in the classification results, such as user ID and user identity information, while retaining information such as statistical feature distribution, facilitating related analysis tasks for management personnel.
[0088] In some embodiments, the method further includes:
[0089] The data governance results are stored in a distributed file system, and a data feature code of the data governance results is generated.
[0090] In response to data operations on the data governance results in the distributed file system, obtain the operation record of the data operation;
[0091] The operation record and the data feature code are stored in a blockchain network.
[0092] In this embodiment, the data feature code can be a character encoding such as a hash value or MD5 value, used to uniquely identify the data governance results.
[0093] In a practical implementation, the distributed file system can be IPFS (InterPlanetary File System). IPFS supports peer-to-peer transmission and stores complete data governance results. When participants perform data operations on the data governance results in the distributed file system, they obtain the operation record of the data operation, and store the data signature and operation record in the blockchain network to ensure that the data governance results are not illegally tampered with.
[0094] In practical applications, blockchain networks employ a lightweight node architecture, storing only the hash value of the data governance structure and operation records, and using smart contracts to automatically verify permissions. This lightweight node architecture ensures that each participant only stores the block header hash value and its own operation records. The smart contracts are configured with preset compliance rules (such as data retention periods and access permissions), and when a violation is detected (such as failure to delete data within the specified period), an alarm is automatically triggered and the account is frozen.
[0095] In this embodiment, storing complete data through a distributed file system and storing data operation records and data feature codes through a blockchain network is more effective in preventing post-event tampering compared to centralized storage methods. This meets the strong compliance requirements of safety audits in the rail transit industry and also reduces storage costs.
[0096] In some embodiments, a rail transit data asset governance model is also provided, which can implement the rail transit data governance method described above.
[0097] In some examples, the governance model for rail transit data assets includes:
[0098] The multi-source data access layer dynamically adapts rail transit data (such as sensor data, operation logs, video stream data, and third-party system data) through interface adapters, and uses a protocol parsing engine to uniformly convert heterogeneous data into a structured intermediate format to obtain the target data;
[0099] The interface adapter includes a protocol conversion submodule, supporting adaptive switching between MQTT, Modbus, and HTTP protocols, and features a built-in data cache queue to handle high-concurrency scenarios. Through a built-in protocol feature library and parsing rule engine, it automatically identifies the data source communication protocol and performs format conversion, avoiding efficiency bottlenecks caused by manual configuration. The data cache queue adopts a circular buffer design, prioritizing real-time data streams in high-concurrency scenarios to ensure low-latency transmission of critical data. This design allows the system to be quickly integrated with various new and old rail transit equipment, is compatible with existing industrial control systems, reduces deployment costs, and simultaneously increases data throughput.
[0100] The data preprocessing layer includes an intelligent cleaning module. The intelligent cleaning module dynamically generates data cleaning rules based on reinforcement learning, including missing value imputation, outlier detection and redundant field removal. It also introduces time-series correlation analysis to optimize the cleaning strategy and clean the target data.
[0101] The classification layer integrates the BERT model with the knowledge graph of the rail transit field to generate operation and maintenance semantic tags for the target data. It associates equipment codes, line topology and operation and maintenance knowledge through an attention mechanism and outputs multi-level classification results.
[0102] The domain knowledge graph comprises an equipment fault code library, line geographic information topology, and operation and maintenance procedure texts. Knowledge embedding and dynamic updates are achieved through graph neural networks. The knowledge graph specifically consists of three core knowledge categories: equipment fault code library, line geographic information topology, and operation and maintenance procedure texts. Graph neural networks transform discrete knowledge entities into vector representations, establishing a mapping between equipment, faults, and handling measures. During classification, attention weights dynamically fuse the knowledge graph with real-time data features, enabling the model to infer the root causes of faults. The knowledge graph uses a differential update method; when a new operation and maintenance case is added, only the weights of the associated edges are locally adjusted, reducing the resource consumption of full training and improving classification accuracy.
[0103] The dynamic governance engine adjusts data cleaning thresholds, classification confidence levels, and storage strategies in real time based on data quality assessment metrics and business needs feedback, and generates governance strategy iteration reports.
[0104] The dynamic governance engine can include a quality assessment submodule. This submodule calculates metrics covering data integrity, consistency, timeliness, and business relevance, and determines governance priorities based on fuzzy logic. Fuzzy logic algorithms are used to comprehensively calculate governance priorities; for example, when data latency exceeds a threshold and is associated with a high-priority work order, the processing level of that batch of data is automatically upgraded. Real-time visual dashboards provide feedback on the assessment results, driving dynamic adjustments to governance strategies and significantly improving the response speed of critical business data.
[0105] The security enhancement layer embeds a lightweight federated learning module, which can jointly train data across departments without exposing the original data. At the same time, it uses dynamic desensitization to conditionally blur sensitive fields.
[0106] In some embodiments, the rail transit data asset governance model can also be packaged into a microservice component, supporting dynamic scaling of Kubernetes (K8S) clusters, and integrating model version control and canary release functions. Combined with Kubernetes elastic scheduling, stability under high concurrency is ensured.
[0107] In practical applications, the CPU, memory, and network load metrics can be collected in real time to trigger automatic scaling strategies.
[0108] In some embodiments, a continuous optimization mechanism for the governance model of rail transit data assets can also be established. Based on an online learning framework, user behavior logs and system performance indicators are received, classification model weights and governance strategy library are updated regularly, and the long-term effectiveness of the model is maintained through online incremental training and drift detection.
[0109] In practical applications, online learning frameworks can employ incremental learning algorithms, retain historical model snapshots, and implement model drift detection mechanisms to trigger retraining. While preserving historical model parameters, weights are updated in batches using streaming data. Model drift detection calculates the difference between the current data distribution and the training set using KL divergence; when a threshold is exceeded, a full retraining is triggered. Historical snapshots are stored in an object storage system with timestamps, supporting rollback to any version as needed.
[0110] In some embodiments, the data flow path of data governance results can also be recorded based on blockchain technology, supporting full lifecycle auditing and compliance verification, establishing an immutable audit chain, and meeting industry compliance requirements.
[0111] In some embodiments, a visual encapsulation tool can also be built to provide a drag-and-drop data flow orchestration interface, which encapsulates the data governance results, i.e. the governed rail transit data assets, into standardized API services, message queues, or binary file packages, and automatically generates metadata description documents.
[0112] In practical applications, visualization and encapsulation tools can also provide data relationship graph display functions, supporting click-based queries of upstream and downstream dependencies and impact analysis.
[0113] By analyzing data relationship graphs, a cross-system data flow topology diagram is automatically generated, allowing users to view field-level change history by clicking on nodes. When data anomalies occur, the system can be traced back to the source system, and impact analysis algorithms can be used to mark affected business modules. This feature enables operations and maintenance personnel to quickly locate the root cause of data problems, significantly shortening troubleshooting time, while also providing visual support for data version management.
[0114] The embodiments of this application have the following advantages: By converting various rail transit data into target data in a preset data format according to various data communication protocols, differentiated adaptation to different rail transit systems is achieved, and various types of rail transit data are standardized into target data in a preset data format, making them compatible with both new and old rail transit systems; a rail transit knowledge network is constructed through a pre-built rail transit knowledge graph, and operation and maintenance knowledge entities associated with the target data are matched in the rail transit knowledge graph. Operation and maintenance semantic tags for the target data are generated based on the operation and maintenance knowledge entities, and potential rail transit operation and maintenance semantic information of the target data is mined, achieving preliminary translation of the target data. Then, based on the operation and maintenance semantic tags, the rail transit operation and maintenance events to which the target data belongs are classified, automatically adapting to the unique professional terms, business rules, and logical relationships in the rail transit field, and finally transforming the original, underlying rail transit data into highly available and highly reliable rail transit operation and maintenance event classification results, thereby achieving accurate and efficient governance of rail transit data.
[0115] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0116] Based on the same inventive concept, this application also provides a rail transit data governance device for implementing the rail transit data governance method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the rail transit data governance device provided below can be found in the limitations of the rail transit data governance method described above, and will not be repeated here.
[0117] In one exemplary embodiment, such as Figure 4 As shown, a rail transit data governance device 40 is provided, comprising:
[0118] Input module 402 is used to respond to a data governance request, acquire rail transit data from multiple rail transit systems, and determine the data communication protocol used by each of the rail transit systems to transmit the rail transit data;
[0119] The conversion module 404 is used to convert the rail transit data into target data according to a preset data format based on the data communication protocols.
[0120] The generation module is used to match the operation and maintenance knowledge entity 406 associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entity.
[0121] The classification module 408 is used to classify the rail transit operation and maintenance events to which the target data belongs based on the operation and maintenance semantic tags, and obtain the classification results;
[0122] The output module 410 is used to output the data governance results of the rail transit data based on the classification results.
[0123] In one embodiment, the operation and maintenance knowledge entity includes equipment fault knowledge entities, traffic line knowledge entities, and operation and maintenance measure knowledge entities associated with rail transit equipment, and the classification module 408 is further used for:
[0124] Based on the operation and maintenance semantic tags, query the multiple entity association relationships of the equipment fault knowledge entity, the transportation line knowledge entity, and the operation and maintenance measures knowledge entity in the rail transit knowledge graph;
[0125] Based on the multiple entity associations, the target associations of the target data with the target rail transit equipment are determined; wherein, the target associations are used to characterize the association between the faults occurring in the target rail transit equipment and the operation and maintenance measures.
[0126] Based on the target association, the rail transit operation and maintenance events to which the target data belongs are classified to obtain the classification results.
[0127] In one embodiment, the output module 410 is further configured to:
[0128] Obtain multiple user role attribute information and determine the data access permissions corresponding to each user role attribute information;
[0129] The classification results are encrypted using different data access permissions to obtain multiple encrypted classification results.
[0130] Based on the classification results obtained after encrypting multiple data sets, the data governance results of the rail transit data are output.
[0131] In one embodiment, the step of performing different data encryption processes on the classification results according to different data access permissions includes:
[0132] The data access permissions and the classification results are input into a pre-built conditional generative adversarial network model;
[0133] Using different data access permissions as the generation conditions of the conditional generative adversarial network model, the conditional generative adversarial network model determines the target fields in the classification results that correspond to different data access permissions and need to be encrypted, and generates replacement fields for each target field;
[0134] According to different data access permissions, the target field in the classification result is replaced with the corresponding replacement field to obtain multiple encrypted classification results.
[0135] In one embodiment, the rail transit data governance device 40 further includes:
[0136] The first storage module is used to store the data governance results in a distributed file system and generate data feature codes for the data governance results;
[0137] The operation record acquisition module is used to acquire the operation record of the data operation in response to the data operation of the data governance result in the distributed file system;
[0138] The second storage module is used to store the operation record and the data feature code in the blockchain network.
[0139] In one embodiment, the rail transit data governance device 40 further includes:
[0140] The data cleaning module is used to input the target data into a pre-trained data cleaning reinforcement learning model, extract data features of the target data through the data cleaning reinforcement learning model, generate data cleaning rules based on the data features, and perform data cleaning on the target data according to the data cleaning rules.
[0141] Each module in the aforementioned rail transit data management device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0142] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data including, but not limited to, rail transit data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a rail transit data management method.
[0143] Those skilled in the art will understand that Figure 3The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0144] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described rail transit data governance method.
[0145] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described rail transit data governance method.
[0146] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the rail transit data governance method described above.
[0147] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0148] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0149] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0150] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for governing rail transit data, characterized in that, The method includes: In response to a data governance request, the system acquires rail transit data from multiple rail transit systems and determines the data communication protocol used by each of the rail transit systems to transmit the rail transit data. According to the data communication protocols described above, the rail transit data are converted into target data in a preset data format; Match the operation and maintenance knowledge entities associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entities; Based on the operation and maintenance semantic tags, the rail transit operation and maintenance events to which the target data belongs are classified to obtain the classification results; Based on the classification results, the data governance results of the rail transit data are output.
2. The method according to claim 1, characterized in that, The operation and maintenance knowledge entities include equipment fault knowledge entities, transportation line knowledge entities, and operation and maintenance measure knowledge entities associated with rail transit equipment. The classification of rail transit operation and maintenance events to which the target data belongs, based on the operation and maintenance semantic tags, yields classification results, including: Based on the operation and maintenance semantic tags, query the multiple entity association relationships of the equipment fault knowledge entity, the transportation line knowledge entity, and the operation and maintenance measures knowledge entity in the rail transit knowledge graph; Based on the multiple entity associations, the target associations of the target data with the target rail transit equipment are determined; wherein, the target associations are used to characterize the association between the faults occurring in the target rail transit equipment and the operation and maintenance measures. Based on the target association, the rail transit operation and maintenance events to which the target data belongs are classified to obtain the classification results.
3. The method according to claim 1, characterized in that, The step of outputting the data governance results of the rail transit data based on the classification results includes: Obtain multiple user role attribute information and determine the data access permissions corresponding to each user role attribute information; The classification results are encrypted using different data access permissions to obtain multiple encrypted classification results. Based on the classification results obtained after encrypting multiple data sets, the data governance results of the rail transit data are output.
4. The method according to claim 3, characterized in that, The step of performing different data encryption processes on the classification results according to different data access permissions includes: The data access permissions and the classification results are input into a pre-built conditional generative adversarial network model; Using different data access permissions as the generation conditions of the conditional generative adversarial network model, the conditional generative adversarial network model determines the target fields in the classification results that correspond to different data access permissions and need to be encrypted, and generates replacement fields for each target field; According to different data access permissions, the target field in the classification result is replaced with the corresponding replacement field to obtain multiple encrypted classification results.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The data governance results are stored in a distributed file system, and a data feature code of the data governance results is generated. In response to data operations on the data governance results in the distributed file system, obtain the operation record of the data operation; The operation record and the data feature code are stored in a blockchain network.
6. The method according to any one of claims 1 to 4, characterized in that, After converting the rail transit data into target data according to a preset data format based on the respective data communication protocols, the method further includes: The target data is input into a pre-trained data cleaning reinforcement learning model. The data features of the target data are extracted by the data cleaning reinforcement learning model. Data cleaning rules are generated based on the data features. The target data is then cleaned according to the data cleaning rules.
7. A rail transit data management device, characterized in that, The device includes: The input module is used to respond to data governance requests, acquire rail transit data from multiple rail transit systems, and determine the data communication protocol used by each of the rail transit systems to transmit the rail transit data. The conversion module is used to convert the rail transit data into target data according to the preset data format based on the data communication protocols. The generation module is used to match the operation and maintenance knowledge entities associated with the target data in the pre-built rail transit knowledge graph, and generate operation and maintenance semantic tags for the target data based on the operation and maintenance knowledge entities. The classification module is used to classify the rail transit operation and maintenance events to which the target data belongs based on the operation and maintenance semantic tags, and obtain the classification results; The output module is used to output the data governance results of the rail transit data based on the classification results.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.