Railway station facility whole life cycle information query method based on digital twinning

By generating unified facility objects and constructing a full lifecycle causal chain model using digital twin technology, the problems of data silos and lack of logical explanation for fault causes in railway passenger station facility management have been solved, enabling accurate fault tracing and explanation.

CN121326953BActive Publication Date: 2026-03-31CHINA RAILWAY CONSTR ENG GRP FOURTH CONSTR CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies lack deep integration of cross-stage data logic relationships in railway passenger station facility management, resulting in a high false alarm rate in fault diagnosis, an inability to quantify the impact of early deviations on later faults, and rigid temporal model construction, making it difficult to accurately capture sudden changes in facility status.

Method used

By using a digital twin-based approach, a unified facility object is generated, a causal chain model of the facility's entire lifecycle is constructed, and the causal relationships between events are evaluated by combining engineering constraints and statistical correlations, enabling explanatory queries.

Benefits of technology

It effectively solves the problem of spurious correlation, quantifies the cumulative impact of early risks in the design and construction phases on later operation, and achieves accurate tracing and explanation of the causes of failures across stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121326953B_ABST
    Figure CN121326953B_ABST
Patent Text Reader

Abstract

The application discloses a kind of railway station facilities full life cycle information query method based on digital twinning, comprising: unified facility object with unique logical identification is generated based on the aggregation of multi-source heterogeneous data;Facility time version model containing time series node is constructed using event-driven rules;Combined with engineering constraint knowledge base and statistical characteristics, the physical feasibility between events and statistical correlation strength are evaluated, and a life cycle stage risk factor is introduced to construct a facility full life cycle causal chain model;Based on the model response explanatory query, a constrained path search is performed in the facility full life cycle causal chain model, and a life cycle query result set containing causal paths is generated.The present application can effectively solve the problem of pseudo-correlation generated by simple data mining, quantify the cumulative effect of early risk in the design and construction stages on the later operation, and realize accurate tracing and explanation of cross-stage fault causes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital twin technology, and in particular to a method for querying the entire lifecycle information of railway passenger station facilities based on digital twins. Background Technology

[0002] As ultra-large railway passenger stations develop towards intelligence and integration, the number of their facilities and equipment is growing exponentially. Full lifecycle management is crucial for ensuring the safe operation of these stations. Digital twin technology, as a digital mapping of physical entities, has become a key means to achieve real-time perception and historical traceability of facility status. By constructing a closed-loop interaction between virtual and physical spaces, the operational efficiency and emergency response capabilities of passenger stations can be significantly improved.

[0003] Currently, railway passenger station facility management largely adopts a phased and system-specific independent model. The design and construction phases primarily rely on Building Information Modeling (BIM) for static display, while the operation and maintenance phase depends on Asset Management System (EAM) or Building Automation System (BAS) to record operational logs. Existing query technologies mainly rely on keyword matching in relational databases or timestamp-based log retrieval, focusing on data extraction within a single system. Although some solutions attempt to aggregate multi-source data through data warehouses, this only achieves physical data centralization and simple report display, lacking deep integration of the logical relationships between data across different phases.

[0004] Existing technologies face several deep-seated problems when dealing with complex operation and maintenance attribution analysis: First, they lack a mechanism for integrating engineering constraints and data statistics. Relying solely on data mining algorithms can easily generate a large number of statistically relevant but physically meaningless pseudo-causal relationships in massive operation and maintenance data, leading to a high false alarm rate in fault diagnosis. Second, they lack a risk propagation quantification model based on lifecycle phase sequence, making it impossible to quantify how small deviations in the early stages (such as design selection and construction installation) are amplified step by step over a long lifecycle, leading to later failures and causing a gap in root cause tracing. Third, their temporal model construction is rigid. Traditional solutions often use fixed-frequency snapshots to record states, making it difficult to accurately capture and store discontinuous state changes triggered by critical events with low redundancy, resulting in a difficulty in balancing the granularity of historical state backtracking with storage costs. Summary of the Invention

[0005] The purpose of this invention is to provide a method for querying the entire lifecycle information of railway passenger station facilities based on digital twins, in order to solve at least one of the aforementioned problems existing in the prior art.

[0006] According to one aspect of this application, a method for querying the entire lifecycle information of railway passenger station facilities based on digital twins includes:

[0007] Based on multi-source heterogeneous basic data, a unified facility object with a unique facility logical identifier is generated through multi-dimensional attribute matching and aggregation.

[0008] By using facility logical identifiers to associate full lifecycle event data, a facility temporal version model containing time-series nodes is generated based on event-driven rules;

[0009] By combining a pre-stored engineering constraint knowledge base with statistical features extracted from the facility temporal version model, the physical feasibility and statistical correlation strength between events are evaluated, and a causal chain model of the entire facility life cycle connecting the design, construction and operation phases is constructed.

[0010] Parse the explanatory query request for the target facility, perform a constrained path search in the facility's full lifecycle causal chain model, and generate a lifecycle query result set containing causal paths.

[0011] According to another aspect of this application, a computer electronic device includes a processor and a memory, the memory storing a computer program, and when the processor executes the computer program, it implements specific steps of a method for querying the entire lifecycle information of railway passenger station facilities based on digital twins.

[0012] Beneficial effects: This invention can effectively solve the problem of spurious correlations caused by simple data mining, quantify the cumulative impact of early risks in the design and construction phases on later operation, and achieve accurate tracing and explanation of the causes of cross-phase failures. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the overall process for querying information on the entire lifecycle of railway passenger station facilities based on digital twins.

[0014] Figure 2 It is a schematic diagram of the process of constructing a causal chain model of the entire life cycle of a facility, connecting the design, construction and operation stages.

[0015] Figure 3 This is a schematic diagram of the process for obtaining a comprehensive causal score.

[0016] Figure 4 This is a schematic diagram of the causal chain model of the entire life cycle of the organization's generation facilities. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] According to one aspect of this application, such as Figure 1 As shown, this paper presents the overall process of a digital twin-based method for querying the full lifecycle information of railway passenger station facilities. It elaborates on the overall solution for establishing a unified identity for a single facility object within a railway passenger station from the data source, performing versioned modeling in the time dimension, and realizing explanatory queries through causal chain analysis. This solves the technical problems of data silos, difficulty in tracing historical status, and lack of logical explanation for the causes of failures in traditional operation and maintenance systems.

[0019] Step 101: Based on multi-source heterogeneous basic data, generate a unified facility object with a unique facility logical identifier through multi-dimensional attribute matching and aggregation.

[0020] In this step, multi-source heterogeneous basic data refers to a collection of original records from different business systems with varying formats. Specifically, this collection at least includes geometric model data from Building Information Modeling (BIM), asset ledger data from Enterprise Asset Management (EAM), equipment monitoring point table data from Building Automation System (BAS), and work order record data from Operation and Maintenance Management System. This data is typically stored in a scattered manner, lacking a unified association key. Through multi-dimensional attribute matching and aggregation, the system first identifies the scattered records as the same physical entity based on the overlap of spatial locations and the similarity of equipment attributes. For example, it can identify that the fan element located on the east side of the second-floor waiting hall in the BIM model is the same equipment as AHU-202 in the same location in the asset ledger. Based on this, a unified facility object is generated, and a globally unique facility logical identifier is assigned to it. This identifier can adopt the Universally Unique Identifier (UUID) format, serving as the unique primary key for connecting data at subsequent stages. This unified processing eliminates information barriers between systems, laying the data foundation for full lifecycle management.

[0021] In some alternative implementations, to improve the efficiency of attribute matching, the spatial coordinate range of the device can be used as the first-level index to filter out a candidate matching set, and a fuzzy matching algorithm for the device name string can be used for secondary confirmation. Furthermore, for newly built passenger stations, unified coding can be directly embedded in the BIM modeling stage to simplify the subsequent data aggregation process.

[0022] Step 102: Use facility logical identifiers to associate full lifecycle event data, and generate a facility temporal version model containing time series nodes based on event-driven rules.

[0023] This step constructs a temporal model that dynamically reflects the evolution of the facility's historical state. The full lifecycle event data covers all records affecting the facility's state from design and construction to operation and maintenance. Specifically, events in the design phase include freezing design parameters and issuing change orders; events in the construction phase include equipment installation and acceptance of concealed works; and events in the operation and maintenance phase include routine inspections, fault repairs, component replacements, and adjustments to operating strategies. Based on event-driven rules, the system does not mechanically record states at fixed time intervals, but only triggers a version generation operation when the aforementioned key events occur. For example, when the system receives a maintenance record regarding the replacement of a motor bearing, it extracts the latest geometric information, performance parameters, and operating status of the equipment at that moment, generating an independent version node. These version nodes are chained together in chronological order to form the facility's temporal version model. This model can accurately reconstruct the facility's configuration and state at any historical moment, providing reliable data support for accident backtracking.

[0024] Step 103: Combining the pre-stored engineering constraint knowledge base with the statistical features extracted from the facility temporal version model, assess the physical feasibility and statistical correlation strength between events, and construct a facility lifecycle causal chain model connecting the design, construction and operation phases.

[0025] The purpose of constructing a causal chain model of the entire facility lifecycle is to uncover the logical relationships hidden behind massive amounts of data. The engineering constraint knowledge base stores domain expert experience and physical laws, such as the rated airflow limit of a certain type of wind turbine and topological constraints of pipeline connections. Physical feasibility assessment uses this knowledge to verify whether there is a theoretical causal possibility between two events. For example, undersized components during the design phase may physically lead to overload failures during operation, while a simple change in appearance color cannot cause electrical failures. Statistical characteristics refer to the co-occurrence frequency or conditional probability of events calculated from historical data. Statistical correlation strength assessment quantifies the actual correlation between two events through data analysis. This step combines qualitative physical rules with quantitative statistical data, effectively eliminating spurious correlations and filtering out event chains with genuine causal relationships. The final model organically links design defects, construction deviations, and later operational failures, forming a cross-stage causal network.

[0026] Step 104: Parse the explanatory query request for the target facility, perform a constrained path search in the facility's full lifecycle causal chain model, and generate a lifecycle query result set containing causal paths.

[0027] This step enables intelligent user interaction and result output. Explanatory query requests typically include the name or location of the target facility, the type of question of interest, and a time range. For example, maintenance personnel might query the reasons for frequent malfunctions of an escalator over the past month. The system first parses the request, identifying the corresponding facility logical identifier and the target fault event node. In the facility's full lifecycle causal chain model, starting from this fault node, the system searches backward along the causal edges to find the preceding events that led to the fault. Constrained path search means that the search process is limited by time sequence, causal strength thresholds, and path length to ensure the accuracy of the results. The final generated lifecycle query result set not only contains the fault facts but also visually demonstrates the complete causal path, such as improper design selection leading to construction and installation difficulties that cause operational failures, helping managers quickly locate the root cause and develop improvement measures.

[0028] According to one aspect of this application, a further detailed description is provided of how to clean and aggregate high-quality unified facility objects from multi-source heterogeneous data, and the process of constructing a spatial location index is supplemented to support location-based fast queries.

[0029] Step 201: Read multi-source heterogeneous basic data, perform primary clustering based on the spatial proximity of the facility and the characteristics of the professional category, and generate candidate record clusters containing records from multiple sources.

[0030] In this step, multi-source heterogeneous basic data is loaded into the system's temporary processing buffer. The data typically contains different attribute fields; for example, BIM data includes a globally unique identifier (GUID) and geometric coordinates; ledger data includes asset codes and installation location descriptions; and monitoring point tables include point names and their respective systems. To improve processing efficiency, the system first performs preliminary clustering. Specifically, it extracts spatial location information from each record. For data containing explicit 3D coordinates, the coordinate values ​​are used directly; for data containing only semantic descriptions such as floor or room numbers, a pre-defined spatial dictionary maps the coordinates to an approximate range. Simultaneously, professional category features are used for auxiliary grouping, distinguishing equipment from different disciplines such as HVAC, plumbing, and electrical systems. Spatial proximity can be measured by calculating the Euclidean distance between the record center points. If the distance between two records is less than the preset clustering radius and they belong to the same professional category, they are grouped into the same candidate record cluster. Through this method, the originally scattered massive amount of data is divided into several small clusters, each containing multiple records that may point to the same physical entity, narrowing the scope for subsequent precise comparisons.

[0031] Step 202: Calculate the geometrical spatial overlap rate and key attribute consistency among the source records within the candidate record cluster to generate a matching score that quantifies the degree of identity.

[0032] This step performs a refined consistency check on records within the candidate record cluster. Geometric spatial overlap rate primarily targets records with geometric information, such as BIM model primitives and GIS spatial data. The system calculates the ratio of the intersection volume to the union volume of two geometric objects in three-dimensional space; the larger this ratio, the more spatially similar they are. Key attribute consistency targets text or numerical fields such as equipment model, rated parameters, and manufacturer. For text fields, an edit distance algorithm can be used to calculate string similarity; for numerical fields, it is determined whether the difference is within the allowable error range. The system calculates the above indicators for each pair of records and performs a weighted sum based on preset weighting coefficients to obtain the final matching score. This score is a value between 0 and 1, reflecting the likelihood that two records from different sources describe the same physical facility. For example, if two records have a very high spatial overlap rate and the equipment model is identical, their matching score will be close to 1.

[0033] Step 203: Records with matching scores exceeding the preset aggregation threshold are grouped into a unified facility object that is unique to the physical entity, and a globally unique facility logical identifier is assigned to it according to the preset coding rules, which serves as the primary key for subsequent data association.

[0034] In this step, the system sets a strict aggregation threshold, typically between 0.75 and 0.90. The optimal threshold can be determined through cross-validation based on historical labeled data, or an initial value, such as 0.85, can be set according to railway industry equipment management regulations. Only when the matching score between candidate records exceeds this threshold does the system determine that they belong to the same physical facility and perform a merge operation. The merge process includes integrating attribute information from various sources, eliminating conflicting fields, and forming a unified facility object containing comprehensive information. For example, this object retains the high-precision geometric data of the BIM model and integrates purchase price and warranty period information from the ledger system. Based on preset coding rules, such as using a combination of station code, professional code, and serial number, or generating a standard UUID, a facility logical identifier is assigned to this object. This identifier remains unchanged throughout the system's entire lifecycle; even if the equipment number in the original business system changes, this logical identifier can ensure data continuity and traceability.

[0035] Step 204: Construct a spatial location index based on an R-tree to map the geometric boundaries of unified facility objects to facility logical identifiers.

[0036] This embodiment also includes the step of constructing a spatial positioning index. To support users in querying facility information by clicking or selecting within a 3D scene, the system requires an efficient spatial retrieval mechanism. Specifically, the geometric bounding box of each uniform facility object is extracted, i.e., the smallest cuboid encompassing the object. Using the bounding box data, the system constructs an R-tree data structure (R-tree spatial index structure). An R-tree is a self-balancing tree data structure suitable for storing and querying high-dimensional spatial data. The leaf nodes of the R-tree store the geometric boundary information of the object and its corresponding facility logical identifier. When a user initiates a spatial query on the front-end interface, the system converts the query range into spatial coordinates and quickly locates all leaf nodes falling within that range by traversing the R-tree to obtain the relevant facility logical identifier. This indexing mechanism improves the response speed of location-based spatial queries, enabling real-time interaction in large-scale BIM scenarios.

[0037] According to one aspect of this application, the construction and optimization of a facility temporal version model are described in detail, addressing how to efficiently record and manage facility status data that changes over time, and optimizing storage space and query performance through a redundancy control mechanism.

[0038] Step 301: Define event triggering mode parameters that include features such as parameter freezing, entity replacement, and operation strategy adjustment, and perform time-series traversal of the full lifecycle event data.

[0039] This step first establishes an event triggering mechanism. Event triggering mode parameters serve as the basis for the system to identify critical state changes. Specifically, parameter freezing typically occurs at the end of the design phase or when major changes are confirmed, marking the formal establishment of the facility's technical specifications; physical replacement involves specific operational actions, such as replacing worn gears or burnt circuit boards, directly altering the facility's physical structure; and operational strategy adjustments refer to changing the equipment's control logic or setpoints, such as adjusting the air conditioning system's water supply temperature setpoint from 7 degrees Celsius to 9 degrees Celsius. The system scans the entire lifecycle of event data chronologically. During the scan, each event record is matched against a predefined triggering mode. Only events whose type and characteristics match the above mode are marked as valid triggering events, avoiding the misjudgment of numerous irrelevant daily logs as version change points.

[0040] Step 302: When an event that matches the event triggering mode parameters is detected, extract the geometric shape, key technical parameters and operating status of the unified facility object at that moment, and generate an attribute snapshot.

[0041] Once a triggering event is confirmed, a snapshot extraction operation is performed. An attribute snapshot is a complete record of the facility's overall appearance at a specific moment. The system queries all data sources associated with that moment to obtain the current geometric model version, the latest technical parameter configuration, and real-time operational monitoring data. For example, if a wind turbine motor replacement event is detected on a certain day, the system will record static parameters such as the new motor's model, power, and rated current, as well as dynamic operating indicators such as the wind turbine's speed and vibration amplitude at the moment the replacement was completed. This data is packaged into structured data blocks, i.e., the attribute snapshot. This snapshot not only contains the data itself but also implicitly includes the contextual relationships between the data, ensuring the interpretability of historical states.

[0042] Step 303: Encapsulate the attribute snapshot into version node data with an effective timestamp, and record the version sequence relationship between adjacent version node data according to the generated time order to form a facility temporal version model.

[0043] This step completes the basic construction of the temporal model. Each attribute snapshot is assigned an effective timestamp, corresponding to the time the triggering event occurred. The encapsulated data unit is called version node data. The system arranges all version nodes belonging to the same facility in chronological order and establishes pointer links between predecessors and successors, forming a version sequence relationship. The chained structure allows the system to easily trace the state evolution of the facility forward or backward along the timeline. For example, the system can quickly locate the previous version of the current version, compare the differences between the two, and analyze the specific content of the most recent change.

[0044] Step 304: Use the version sequence relationship to locate version node data with continuous time, calculate the degree of difference of key attributes between adjacent nodes, and generate version similarity.

[0045] To prevent a surge in the number of versions due to minor disturbances or frequent non-essential operations, the system introduces a similarity calculation mechanism. It iterates through the version sequence, selecting two adjacent version nodes as comparison objects. The comparison focuses primarily on key attributes, such as core geometric dimensions and major performance parameters. Version similarity can be calculated using algorithms such as weighted Euclidean distance or cosine similarity. For example, if two versions differ only in non-critical remarks fields, or if the changes in operating parameters are minimal, the calculated similarity value will be very high. This step quantifies the degree of similarity between two consecutive versions in terms of business substance.

[0046] Optionally, the generation of version similarity employs a hybrid calculation logic combining discrete attribute matching and numerical parameter differences. Specifically, the version similarity calculation formula is:

[0047] ;

[0048] Among them, the variability of numerical parameters The specific calculation formula is as follows:

[0049] ;

[0050] Among them, Sim(V t V t-1 The expression represents the overall similarity between the version node at time point t and the version node at time point t-1.

[0051] i represents the index of a discrete attribute, such as a device status code or a responsible person ID;

[0052] n represents the total number of discrete attributes;

[0053] w i This represents the preset weight of the i-th discrete attribute;

[0054] Attr i (t) represents the i-th discrete attribute value at time point t;

[0055] I is the indicator function, when Attr i (t)=Attr i When (t-1), the function value is 1; otherwise, it is 0.

[0056] λ is an adjustment coefficient used to balance the weights of discrete attributes and numerical parameters;

[0057] This represents the numerical parameter Euclidean distance between version t and version t-1;

[0058] j represents the index of a numerical parameter, such as the temperature setpoint or motor speed;

[0059] P j (t) represents the value of the j-th numerical parameter at time point t.

[0060] Judgment logic: If Sim(V) t V t-1 If the threshold is greater than Threshold, which is a preset merging threshold, then the two versions are determined to be similar, triggering a redundant merging operation.

[0061] Step 305: If the version similarity is higher than the preset merging threshold and the time interval is within the preset merging window, merge the change information of the subsequent node into the preceding node to generate updated version node data.

[0062] This step performs specific redundancy reduction operations. The system sets a merging threshold, such as 0.95, and a merging window, such as 24 hours. That is, only when two versions are extremely similar and the time interval between their occurrences is very short will they be considered redundant. In this case, the system considers the later version to have not brought about a substantial change in state, or to be a quick correction of the earlier version. The small amount of change information contained in the later version, such as corrected notes or fine-tuned parameters, is updated and integrated into the preceding node, forming new version node data containing the merged information. This processing method preserves the necessary change records while avoiding the generation of a large number of fragmented minor versions.

[0063] Step 306: Delete the merged nodes in the facility temporal version model and re-establish the updated version sequence relationship based on the updated version node data.

[0064] After merging the information, the system physically removes the original successor nodes from the model, freeing up storage space. Simultaneously, the system adjusts the pointer relationships in the version sequence, directing the successor pointers of previous nodes directly to the next node after the original successor node, maintaining the continuity of the chain. Through a series of optimizations, the final facility temporal version model is more concise and efficient, accurately reflecting the key historical evolution of the facility while reducing data storage pressure and computational overhead for querying and retrieval.

[0065] In some optional implementations, to support joint analysis across facilities, this embodiment can further configure a joint query template. This template defines multiple facility objects requiring correlation analysis and their time alignment rules. For example, for the linkage analysis of chillers and cooling towers, the template can specify that the operating version of the chiller is used as a baseline, mapping the version nodes of the cooling tower within the corresponding time period to the same time axis, constructing a combined view containing the status of multiple devices, providing data support for complex system-level fault diagnosis.

[0066] According to one aspect of this application, such as Figure 2 As shown, this paper further presents a causal assessment method based on the fusion of physical constraints and statistical data. It elaborates on the algorithm for deep mining of massive historical data from railway passenger station facilities, addressing how to identify logically related causal chains from seemingly chaotic event sequences within complex engineering systems. By integrating the hard constraints of the physical world with the statistical regularities of the data world, it can effectively overcome the spurious correlation problem caused by relying solely on data mining, and also discover hidden fault patterns that are difficult to detect using experience alone.

[0067] Step 401: Extract event pairs that satisfy the time sequence constraints from the facility temporal version model as candidate causal relationships.

[0068] In this step, the system first needs to construct an analysis sample set. The facility temporal version model stores a sequence of version nodes arranged along a timeline. Meeting the time sequence constraint means that the antecedent event must occur earlier than the consequent event, and the time interval between the two should be within a reasonable engineering impact window. Specifically, a sliding window algorithm is used to traverse the event sequences of all facilities. For example, a time window of 30 days is set. For each failure event in the sequence as a consequent event, the system traces back 30 days to all design changes, construction rectifications, or operational parameter adjustments, extracting each pair of events and marking them as candidate causal relationships. The extracted information includes the event type, occurrence time, facility ID, and the attribute status at that time.

[0069] In some alternative implementations, different time constraint strategies can be set for different types of causal assumptions. For example, for energy consumption changes caused by adjustments to operating parameters, the time window can be set to a short period of several hours to several days; while for equipment lifespan degradation caused by design selection, the time window may span several years. The system supports dynamically loading these time constraint parameters through configuration files to adapt to different analysis scenarios.

[0070] Step 402: Based on the design specifications and physical rules in the pre-stored engineering constraint knowledge base, verify the influence mechanism of the antecedent event on the consequent event in the candidate causal relationship, and generate a physical feasibility score that quantifies the possibility of the influence.

[0071] This step introduces domain knowledge as a filter and scorer. A large number of rule entries are digitally stored in the engineering constraint knowledge base. For example, a rule might be described as: if the cross-sectional area of ​​the duct is reduced by more than 10%, the probability of an increase in fan energy consumption is extremely high. The candidate causal relationships extracted in step 401 are matched against the knowledge base. The matching process is based on the semantic features of the events. If a candidate relationship violates common sense physics, such as attributing abnormal water pressure in a water supply and drainage system to voltage fluctuations in an electrical system, and there is no clear coupling between the two, the system assigns it a very low physical feasibility score, such as 0 points. Conversely, if a candidate relationship conforms to the failure modes described in the design specifications, a higher score is assigned based on the rule's credibility, such as 0.8 or 0.9 points.

[0072] Optionally, each rule may be pre-defined with an expert-annotated confidence value C. rule ∈[0,1]; when a candidate causal relationship matches a rule, the physical feasibility score S is calculated. phy =C rule If multiple rules are matched, the maximum value or weighted average is used; if no rule is matched but the general physical logic is met, S phy Use the default median value, such as 0.5; if it violates the laws of physics, S phy =0.

[0073] The physical feasibility score is a normalized numerical value used to measure the likelihood of a causal relationship at the mechanistic level.

[0074] Furthermore, for relationships not explicitly recorded in the knowledge base but conforming to general physical logic, a fuzzy reasoning mechanism can be used to assign intermediate scores. For example, the impact of a certain non-standard installation process during the construction phase on equipment vibration may not be explicitly recorded, but considering that it falls under the category of mechanical installation, it can be assigned a moderate physical feasibility score, such as 0.5 points, to be further verified by subsequent statistical data.

[0075] Optionally, the data structure of the knowledge base can adopt a triple format of antecedent-consequence-confidence or an ontology knowledge graph; the formal expression of rules can use production rules IF condition THEN consequence WITH confidence; the matching process between rules and events can be specifically based on exact matching of event type labels or approximate matching based on semantic vectors.

[0076] Step 403: Calculate the joint occurrence frequency and conditional probability of candidate causal relationships in the facility temporal version model, and generate a statistical strength score that characterizes the data association degree.

[0077] This step extracts correlation features from a pure data perspective. Global statistics are performed on each candidate causal relationship across the entire facility temporal version model. The main statistical indicators include lift and conditional probability. Specifically, the system calculates the probability of the consequential event occurring given the occurrence of the causal event, denoted as P(B|A); and the probability of the consequential event occurring given the absence of the causal event, denoted as P(B|not A). The statistical strength score can be constructed using the difference or ratio of these two probabilities. For example, if the probability of a certain type of wind turbine triggering a vibration alarm after frequency adjustment is significantly higher than the probability without frequency adjustment, the statistical strength score of this relationship will be close to 1. This score objectively reflects the correlation patterns presented in historical data.

[0078] In some optional implementations, a confidence check is introduced to eliminate statistical noise caused by an insufficient sample size. The statistical strength score is calculated only when the sample size exceeds a preset minimum support threshold; otherwise, the score is set to a default low value or marked as unreliable to prevent misjudgment due to chance.

[0079] Step 404: Perform weighted fusion of physical feasibility score and statistical intensity score to obtain comprehensive causal score, and establish candidate causal relationships with comprehensive causal scores exceeding a preset threshold as valid causal edges, and organize the generation of facility life cycle causal chain model.

[0080] In this step, the system uses a linear weighted model to calculate the comprehensive causal score. total The calculation formula can be: Score total (e i ,e j )=α*S phy (e i ,e j )+β*S stat (e i ,e j );

[0081] Among them, Score total (e i ,e j ) indicates the preceding event e i With consequence event e j Comprehensive causal score between them; e i Indicates the preceding event; e j Indicates a consequence event;

[0082] α represents the weighting coefficient for the physical feasibility score;

[0083] S phy (e i ,e j () represents the antecedent event e calculated based on the engineering constraint knowledge base. i For the consequence event e j The physical feasibility score ranges from 0 to 1.

[0084] β represents the weighting coefficient of the statistical intensity score, and satisfies α+β=1;

[0085] S stat (e i ,e j () represents the antecedent event e obtained based on historical data statistics. i With consequence event e j The statistical intensity score.

[0086] In addition, S stat (e i ,e j Specifically, it can be defined as the conditional probability P(e) j |e i ), that is, in the antecedent event e i The consequence event e under the conditions of occurrence j The probability of occurrence; or defined as the lift, i.e., P(e j |e i ) / P(e j ), used to measure e i The appearance of e jThe probability of occurrence is increased by a factor of 1. In this embodiment, the normalized increase is preferably used as the statistical strength score, and the specific calculation formula is: S stat =min(Lift / Lift max ,1), where Lift max This is a preset maximum lift limit used to normalize the score to the [0,1] interval for weighted fusion with the physical feasibility score. Conditional probability can also be retained as an alternative.

[0087] Optionally, in the initial stage of system operation, when data accumulation is insufficient, the physical weight coefficient can be set to 0.7 and the statistical weight coefficient to 0.3, relying more on expert experience. As the amount of operational data increases, the statistical weight coefficient can be gradually increased. After calculating the comprehensive causal score, the system compares it with a preset threshold. For example, the threshold can be set to 0.6. Candidate relationships with scores higher than 0.6 are identified as valid causal edges, retained, and stored in the graph database; relationships with scores lower than 0.6 are considered noise and removed. The final retained valid causal edges and their scores constitute the basic framework of the causal chain model.

[0088] According to one aspect of this application, the process of constructing a phased risk propagation and hierarchical causal graph is further described, such as... Figure 3 , Figure 4 As shown, this paper elaborates on how to introduce the concept of life cycle stage sequence, use stage risk factors to nonlinearly correct the causal score, and construct a hierarchical graph structure to calculate the multi-hop cumulative effect of risk on the time axis. This method can accurately identify early hidden dangers, such as design defects, that may have weak single-step effects but cause serious consequences after long-term accumulation.

[0089] Step 501: Identify the life cycle stage identifiers of the antecedent events and consequent events in the candidate causal relationships, and match the corresponding stage risk factors from the pre-set stage risk factor table based on the span characteristics of the transition from the early stage to the later stage.

[0090] This step begins by labeling events with their respective stage attributes. Based on the event's timestamp and business type, the system categorizes it into design, construction, commissioning, or operation and maintenance stages, assigning each stage a corresponding lifecycle stage identifier. It then analyzes the stage span of the causal relationship. For example, a design-stage event leading to an operational-stage failure is a long-span impact; an operational-stage operation causing an operational-stage failure is a short-span impact. Based on the engineering principle that upstream decisions have an exponentially amplified impact on downstream costs, the system pre-defines a stage risk factor table. For example, the impact factor of the design stage on the operational stage is set to 1.5, the construction stage's impact factor on the operational stage is set to 1.2, and the impact factor within the operational stage itself is set to 1.0. The system queries this table to match the corresponding stage risk factors based on the identified stage combinations.

[0091] Specifically, when both the causal event and the consequential event belong to the design phase, the phase risk factor is 1.0, indicating that causal transmission within the same phase is not amplified or corrected. When both the causal event and the consequential event belong to the construction phase, the phase risk factor is 1.2, reflecting a moderate amplification of the impact of design decisions on construction execution. When both the causal event and the consequential event belong to the operation phase, the phase risk factor is 1.5, reflecting a significant amplification of the impact of design defects on operational failures over a long period. When both the causal event and the consequential event belong to the design phase, the phase risk factor is 0, because events in the construction phase cannot temporally lead to events in the completed design phase; this zero value reflects the hard constraint of temporal causality. When both the causal event and the consequential event belong to the construction phase, the phase risk factor is 1.0. When both the causal event and the consequential event belong to the operation phase, the phase risk factor is 1.2. When both the causal event and the consequential event belong to the operation phase and the consequential event belong to the design or construction phase, the phase risk factor is 0, also reflecting the constraint of temporal causality. When both the current event and the consequence event are in the operational phase, the phase risk factor is set to 1.0.

[0092] In some optional implementations, the initial values ​​of the aforementioned stage risk factors can be calibrated based on the ten-fold rule in engineering economics. This rule indicates that the repair cost increases by approximately an order of magnitude for each stage a defect is propagated to in its lifecycle. Therefore, factors spanning two stages can be taken as the square root or linear sum of factors spanning one stage. Furthermore, the system supports calibrating and adjusting the aforementioned factors based on historical maintenance data from certain passenger stations; the calibration method can be found in the subsequent feedback optimization mechanism.

[0093] Step 502: Correct the weighted sum of the physical feasibility score and the statistical strength score using the stage risk factor to generate a risk propagation score that reflects the cumulative effect of risk across life cycle stages.

[0094] This step performs a revised risk score calculation. The system no longer relies solely on the basic weighted sum, but instead uses it as a base multiplied by the stage-specific risk factor matched in step 501. For example, a basic score of 0.6 might be for a design flaw leading to high energy consumption later on, but because it spans multiple stages from design to operation, the matched risk factor is 1.5, thus the final risk propagation score is revised to 0.9. This mechanism significantly amplifies early-stage problems with far-reaching potential impacts within the scoring system, making them easier to identify as key causes.

[0095] Optionally, for the correction of single-step causal relationships, the risk transmission score calculation formula is as follows:

[0096] Risk(e i ->e j )=Score total (e i ,e j )*r stage (Phase i Phase j );

[0097] Among them, Risk(e i ->e j () indicates that after considering the influence of the stage sequence, starting from the antecedent event e i To the consequences of the event e j Risk transmission score;

[0098] Score total (e i ,e j The ) represents the comprehensive causal score calculated in the preceding steps;

[0099] r stage (Phase i Phase j () indicates the phase to which the preceding event belongs. i Phase of the consequential event j Risk factors between phases. For example, when Phase i For the design phase and Phase j During the operational phase, the phase risk factor is set to a value greater than 1 to reflect the amplifying effect of early decisions on later risks.

[0100] Step 503: Use the risk propagation score as a comprehensive causal score to screen for effective causal edges.

[0101] In this step, the system uses a modified risk propagation score instead of the original score for threshold filtering. Under the same physical and statistical performance, deeper causes across stages are more likely to be selected as valid causal edges than superficial direct causes. The selected edges retain the modified score attributes, providing weighted connections for subsequent graph construction.

[0102] Step 504: For each unified facility object, connect the corresponding event nodes using effective causal edges, construct facility causal graph data, and establish stage layer nodes arranged in chronological order, mapping and associating event nodes with the corresponding stage layer nodes.

[0103] This step constructs a graph model with a specific topological structure. The system creates an empty directed graph container, using a unified facility object as the unit. The system then creates a backbone structure within the graph, consisting of stage-level nodes arranged chronologically, such as design-level nodes pointing to construction-level nodes, and construction-level nodes pointing to operation-level nodes. The causal and consequential events connected by the selected valid causal edges are instantiated as event nodes in the graph, and based on their temporal attributes, they are attached to the corresponding stage-level nodes. For example, design change event nodes are managed by design-level nodes, and maintenance record event nodes are managed by operation-level nodes. This forms a two-layer nested structure of stage layers and event layers, expressing both the macroscopic stage evolution and the microscopic event causality.

[0104] Step 505: In the facility causal graph data, perform a forward topological traversal based on the comprehensive causal score carried by the effective causal edges to calculate the node risk propagation data representing the multi-hop accumulation of risk impact from the early stage to the later stage.

[0105] This step calculates cumulative risk using a graph algorithm. The system starts from the earliest stage layer in the graph and performs a topological traversal along directed edges. For each node, its risk propagation data depends not only on its inherent risk but also on the product of the cumulative risks of all nodes pointing to it and the weights of the incoming edges. The specific calculation logic can employ a maximum path probability algorithm or a weighted accumulation algorithm. For example, when calculating the cumulative risk of a node experiencing operational failure, the system comprehensively considers the upstream design defect risk, construction hazard risk, and the transmission strength between them. If a path involves design defects leading to construction difficulties, construction difficulties leading to installation deviations, and installation deviations leading to operational failure, the system will perform a chain-like calculation of the scores along the path to ultimately obtain the comprehensive risk value of the failed node. This value quantifies the severity and root cause traceability of the failure from a full lifecycle perspective.

[0106] Optionally, for the accumulation of risk along multi-hop paths, the formula for calculating node risk propagation data is as follows:

[0107] Risk cum (e j )=Max (ei∈Pre(ej)) (Risk cum (e i )*Risk(e i ->e j ));

[0108] Among them, Risk cum (e j ) represents the consequence event node e j The cumulative risk value;

[0109] Pre(e j ) indicates pointing to node e j The set of all direct predecessor nodes;

[0110] Max indicates that when pointing to e j Take the maximum value among all the preceding paths;

[0111] e i ∈Pre(e j ), indicating e i It is e j Any node in the set of direct predecessor nodes;

[0112] Risk cum (e i ) represents the predecessor node e i The cumulative risk value.

[0113] This formula represents taking all possible outcomes, e. j In the causal path, the path with the largest risk product is used as the risk measure of that node, reflecting the worst-case estimate of the chain-like transmission of risk.

[0114] Step 506: Encapsulate the facility causal graph data, which includes stage layer nodes, event nodes, and node risk propagation data, into a facility lifecycle causal chain model.

[0115] The system serializes and encapsulates the constructed graph structure and all calculated risk attribute data. The encapsulated model not only contains graphical topological data but also detailed attribute tables for each node and edge, such as event description, occurrence time, responsible party, revised risk score, and cumulative risk value. The encapsulated object is the facility's entire lifecycle causal chain model, stored as a high-value knowledge asset in the graph database, awaiting retrieval by the query engine.

[0116] In some alternative implementations, graph compression can be further performed to optimize the storage efficiency and query performance of the graph model. Specifically, multiple consecutive intermediate event nodes in the graph with few branches and a risk propagation score change rate below a preset compression threshold can be merged into a single aggregation node. For example, multiple consecutive normal daily inspection records of the same type can be merged into a regular inspection cycle node, simplifying the graph structure and making the main branches of the causal path clearer.

[0117] According to one aspect of this application, the application closed-loop process of query search and feedback optimization is further described, elaborating on how the system responds to specific user query requests and how user feedback behavior is used to reverse-optimize model parameters to achieve adaptive evolution of the system.

[0118] Step 601: Semantically analyze the problem description and time conditions in the explanatory query request, map and locate the target event node in the facility lifecycle causal chain model, and generate search constraint data containing search direction and maximum hop count limit.

[0119] This step processes user input. When a user enters an inquiry, such as "Why is the elevator on the west side of the second-floor waiting hall frequently out of service this month?", the system first extracts keywords using natural language processing (NLP). Combining spatial indexing and unified identifiers, the elevator on the west side of the second-floor waiting hall is mapped to a specific facility logical identifier. Simultaneously, the system identifies the frequent outages as corresponding to the fault event type in the operational phase, and "this month" as the time-based filtering condition. Based on this information, the system locates a set of fault event nodes that meet the conditions in the causal chain model. The system generates search constraint data, which is a structured instruction package. Specifically, the search starting point is the located fault node, the search direction is reverse backtracking, and the maximum number of hops is set to, for example, 5 hops to prevent the search path from becoming too long and causing the results to diverge.

[0120] Optionally, the system has a pre-built query template library containing several standardized query patterns. The query template format is: query [Facility Name] within [Time Range] for [Event Type] [Query Intent]. The Facility Name slot is used to fill in the name or location description of the target facility; the Time Range slot is used to fill in the start and end times or relative time expressions; the Event Type slot is used to fill in the fault category or status change type; and the Query Intent slot is used to identify the type of information the user expects to obtain, such as cause tracing or status query.

[0121] When the system receives a natural language query request from a user, it first performs word segmentation, dividing the continuous text into independent word units. Then, it performs named entity recognition based on a pre-built entity dictionary. This entity dictionary contains at least three types of entries: a facility name dictionary, a time expression dictionary, and a fault type dictionary. The facility name dictionary stores the standard names and common aliases of all uniform facilities within the station; for example, the west elevator on the second-floor waiting hall and the west elevator both point to the same facility logical identifier. The time expression dictionary stores the mapping rules between relative and absolute time; for example, "this month" is mapped to the start and end dates of the current month, and "the last three days" is mapped to a time interval of three days prior to the current date. The fault type dictionary stores the standard terms and synonyms for various fault phenomena; for example, "elevator stoppage," "passenger entrapment," and "emergency stop" are all classified as elevator operation interruption faults.

[0122] The system fills the identified entities into the query template slots with the highest matching degree, generating structured query conditions. Specifically, it calculates the matching score between the user input and each template, selecting the template with the highest score as the parsing result. The matching score is calculated based on a weighted sum of slot filling completeness and keyword coverage. If the user input is ambiguous, such as a facility name matching multiple candidate objects, the system generates a clarification prompt and returns it to the user for confirmation, or selects the most likely candidate object based on the user's historical query records.

[0123] Step 602: Based on the search constraint data, perform path backtracking in the facility lifecycle causal chain model, starting from the target event node, to obtain candidate causal paths that satisfy logical connectivity.

[0124] The system performs a traversal of the graph model according to constraint instructions. Starting from the target fault node, it visits its predecessor nodes in reverse order along the incoming edges until the maximum hop count limit is reached or the source node without predecessors, such as the design node, is reached. In this process, all visited nodes and edges form an inverted tree structure. The system extracts each independent path from the root node (fault point) to the leaf node (root cause) to form candidate causal paths. Each path represents a possible explanation for the cause.

[0125] Step 603: Combine the node risk propagation data to sort the cumulative risk values ​​of candidate causal paths and select high-risk associated paths as the lifecycle query result set.

[0126] Since the number of candidate paths extracted may be large, the system needs to perform optimization. It reads the node risk propagation data for each node on each path and calculates the overall confidence score of the path. The calculation method can be the product of the risk propagation scores of all edges on the path, or the weighted sum of the accumulated risks of key nodes (such as cross-stage nodes) on the path. Based on the calculation results, all candidate paths are sorted in descending order, and the top-ranked paths are selected, for example, the top 3. These paths constitute the lifecycle query result set, representing the underlying causes that the system believes are most likely to lead to the current failure.

[0127] Step 604: Collect user confirmation or denial operations on the causal path in the lifecycle query results set, and generate user feedback data.

[0128] After the result set is displayed to the user, the system provides an interactive interface to collect feedback. Users can click to confirm whether a given explanation path is reasonable, unreasonable, or a false alarm. The system records the user's operation type, user ID, operation time, and corresponding path details, generating user feedback data. This is equivalent to providing manually labeled ground truth tags for the model output.

[0129] Step 605: Based on user feedback data, calculate the path accuracy under a certain lifecycle stage identifier combination, adjust the stage risk factors and weight coefficients in the weighted fusion calculation, and generate query optimization parameter data.

[0130] The system periodically or in real-time analyzes accumulated user feedback data. The feedback data is categorized and statistically analyzed according to the combination of stages involved in the path. For example, if the statistics show that paths involving the design to operation stages are frequently confirmed as reasonable by users, it indicates that the risk factor for that stage combination is set appropriately or too low; if a certain type of construction to operation path is frequently rejected, it indicates that its risk factor may be artificially high. Based on the statistical accuracy rate, the system dynamically fine-tunes parameters. For example, it increases the stage risk factor corresponding to high accuracy combinations and decreases the factor for low accuracy combinations; simultaneously, it can also adjust the weighting coefficients of physical scores and statistical scores when weighting and fusing them in the correct path.

[0131] Step 606: Write the query optimization parameter data back into the scoring model to generate updated stage risk factors and updated scoring strategies, so as to dynamically optimize the causal relationship filtering logic when responding to subsequent queries.

[0132] The system updates the calculated new parameter values ​​to the system's configuration database. The updated stage risk factors and scoring strategies will take effect immediately and be applied to subsequent assessments of newly generated causal relationships or reassessments of existing relationships. Through a closed-loop mechanism, the system can gradually learn causal judgment logic that better suits the actual situation of the station from feedback from maintenance personnel as the number of uses increases, continuously improving the accuracy and practical value of query results.

[0133] According to one aspect of the present application, the process of converting the calculated abstract causal chains and complex spatio-temporal data into intuitive visual forms is further elaborated in detail to assist maintenance personnel in performing efficient decision-making analysis. And how to enable users to intuitively understand the evolution of the facility status over multiple years and its failure causes in the mixed dimension of three-dimensional space and one-dimensional time axis.

[0134] The system reads the life cycle query result set. This result set mainly contains a number of high-risk association paths, and each path is composed of a series of event nodes connected by causal logic. To achieve intuitive display, a spatial mapping operation is first performed. The system analyzes the facility logical identifiers corresponding to each event node in the path, and uses the spatial positioning index to retrieve the corresponding three-dimensional geometric model data from the unified facility object set. For example, if the query result shows that a certain power outage accident is caused by overheating of the transformer, the system will automatically locate the specific three-dimensional coordinates and geometric contours of the transformer in the substation.

[0135] Furthermore, the system constructs a three-dimensional visualization scene. In the digital twin interface, highlighting or color coding techniques are used to prominently display the key facilities on the causal path. Specifically, the facility as the result of the failure can be rendered in eye-catching red, the pipelines or cables as the intermediate conduction links can be rendered in yellow, and the facility as the source cause can be rendered in dark orange. At the same time, dynamic streamlines with direction arrows are drawn between these facilities to visually express the risk propagation direction. The visual guidance mechanism can help users instantly identify the flow trajectory of the risk in the physical space, such as from the cooling tower on the top floor down along the riser to the heat exchanger on the bottom floor.

[0136] On this basis, the system synchronously constructs an interactive timeline view. All the event nodes involved in the causal path are mapped to a horizontal timeline according to their timestamps of occurrence. The key stage demarcation points of the facility life cycle, such as the design freeze date, the completion acceptance date, and the dates of previous major repairs, are clearly marked on the timeline. The system supports the联动 operation of the timeline and the three-dimensional scene. When the user drags the slider on the timeline to a certain historical moment, the facility models in the three-dimensional scene will automatically switch to the version state corresponding to that moment. For example, when the slider moves to the construction stage two years ago, the old model valves installed at that time and their spatial positions during construction will be displayed in the scene instead of the new valves that have been replaced currently. The historical state playback function enables users to review the source moment of the accident and witness the design form or construction environment at that time.

[0137] Furthermore, the system provides an evidence chain perspective function. When a user clicks on a causal connection in the visualization interface, a detailed information panel pops up. This panel displays the underlying data evidence supporting the causal relationship, including physical feasibility scores, statistical strength scores, and stage risk factors. The system translates abstract algorithm parameters into natural language descriptions, such as indicating to the user that the association is supported by physical rules and frequently appears in historical data. If the user believes that the inference does not conform to reality, they can directly click the false alarm button in the panel, and the system will then convert this action into user feedback data for subsequent model optimization.

[0138] According to another aspect of this application, a computer electronic device is provided for performing the above-described method for querying the entire lifecycle information of railway passenger station facilities based on digital twins.

[0139] A computer electronic device includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the method steps of any one of the present invention.

[0140] In this embodiment, the computer electronic device can specifically be a high-performance server cluster, an industrial-grade workstation, or a cloud-based computing platform deployed in a railway passenger station operation control center. The processor, as the computing core, can specifically include a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated AI acceleration chip. The CPU is responsible for executing logical control and routine data processing tasks, such as the aggregation and attribute matching of unified facility objects; the GPU is responsible for handling the 3D rendering and visualization mapping tasks of large-scale BIM models; and the AI ​​acceleration chip is used to accelerate causal graph traversal, matrix operations, and probabilistic statistical analysis to ensure millisecond-level query response when processing massive amounts of facility data.

[0141] The memory specifically includes high-speed random access memory (RAM) and non-volatile storage media. High-speed RAM is used to temporarily load running computer program modules and their intermediate data for real-time processing, such as candidate causal relationship sets and node risk propagation data. Non-volatile storage media, such as solid-state drive arrays, are used for persistent storage of massive amounts of multi-source heterogeneous basic data, facility temporal version model libraries, engineering constraint knowledge bases, and the finally generated facility lifecycle causal chain models.

[0142] In addition, the computer electronic device is also equipped with a high-speed network communication interface, which is used to connect the front-end data acquisition sensors, the database servers of various business subsystems, and the terminal display devices used by maintenance personnel through industrial Ethernet or fiber optic networks, so as to build a complete closed loop of data perception, processing and interaction.

[0143] According to one aspect of this application, an alternative implementation based on a distributed graph database is described as a preferred alternative technical architecture for the method of the present invention, applicable to the scenario of ultra-large railway hub passenger stations, solving the performance bottleneck problem of graph model construction and querying caused by excessive data volume in a single-machine environment.

[0144] In the above embodiments, the facility's entire lifecycle causal chain model can be described using a logical graph structure. In this alternative embodiment, the system directly uses a distributed graph database as its core storage engine. Specifically, the system directly maps each unified facility object and each event node to vertices in the graph database, and maps causal edges and version sequence relationships to edges in the graph database. Node attribute snapshot data and risk score data are then directly attached as attributes to the corresponding vertices or edges.

[0145] Based on this architecture, the constrained path search process can be transformed into execution using a native graph traversal query language. Leveraging the index-free adjacency feature of graph databases, the query engine's time complexity when performing reverse traversal from the faulty node to the root cause node is only related to the path length and is independent of the total data size. In other words, even when facing a massive passenger station model containing millions of component nodes, the system can still achieve real-time causal path retrieval.

[0146] Furthermore, for node risk propagation calculations, this embodiment employs a large-scale graph computing framework based on a batch synchronous parallel model. Within this framework, each event node is treated as an independent computing unit. In each round of overstep calculation, nodes receive risk values ​​from upstream nodes in parallel, perform correction calculations based on their own stage risk factors, and send the updated cumulative risk value to downstream nodes. This parallel computing mode shortens the time required for full-scale model risk assessment, enabling the system to support higher-frequency model updates and iterations, and more sensitively capture subtle trends in facility status changes.

[0147] According to one aspect of this application, the implementation details of the distributed graph database architecture are supplemented as follows:

[0148] Regarding the graph data sharding strategy, the system adopts a hash-based sharding method based on facility logical identifiers. Specifically, the system calculates a hash value for the facility logical identifier of each unified facility object, and then uses the hash value modulo the number of shards to determine which data shard the facility object and all its associated event nodes are stored in. This sharding strategy ensures that causal graph data for the same facility are physically stored within the same shard, allowing causal path queries for a single facility to be completed within a single shard, avoiding cross-shard communication overhead. The number of shards can be configured according to the scale of the passenger station, with selectable values ​​ranging from 8 to 64 shards.

[0149] Regarding cross-shard queries, when the query involves association analysis between multiple facility objects, the system adopts a two-stage processing mode of query routing and result aggregation. In the first stage, the query coordinating node parses the query request, identifies the set of logical identifiers of the involved facilities, and distributes the subqueries to the corresponding data shards for parallel execution. In the second stage, each shard returns local query results, and the coordinating node merges, deduplicates, and sorts the results to generate the final query response. For complex queries requiring cross-facility tracing, the system redundantly stores the target node shard location information of the cross-facility association edges in the event nodes of each shard to support multi-hop traversal in a distributed environment.

[0150] Regarding convergence control in the batch synchronous parallel model, the system employs an iterative convergence mechanism when performing node risk propagation calculations. In each round of overstep calculation, each node receives the risk increment value from upstream, updates its own cumulative risk value, and transmits the updated value to downstream nodes. The system sets two types of termination conditions. The first is a maximum iteration limit, for example, set to twice the maximum depth of the causal graph, to prevent infinite iterations caused by loops in the graph. The second is a convergence threshold judgment; when the change in risk value of all nodes in a certain overstep is less than a preset convergence threshold, such as 0.001, the system determines that the calculation has converged and terminates the iteration early. Finally, the cumulative risk value saved by each node is the node risk propagation data.

[0151] According to one aspect of this application, in the query and interpretation of facility lifecycle information based on causal chains, the assessment of causal strength based on physical constraints and data statistics specifically includes:

[0152] The candidate causal relationship data is filtered and scored to obtain a set of filtered causal edge data, and physical feasibility score data, statistical strength score data and comprehensive causal score data are generated in the process.

[0153] First, candidate causal relationships that are obviously impossible are eliminated from the perspective of physical and engineering rules.

[0154] Acquire causal candidate relationship data, where each candidate relationship includes antecedent events, consequent events, the time interval between the two, and information on related facilities. At the same time, read predefined design constraint knowledge base data, construction deviation knowledge base data, and operation strategy knowledge base data. The knowledge base records the physical constraints and engineering experience rules of various facilities in design, construction, and operation.

[0155] For each candidate causal relationship, based on facility type, event type, and time interval, a matching rule entry is searched in the knowledge base. For example, whether a certain construction deviation will affect a certain operational indicator within a specified time window. If the matched rule explicitly prohibits such an impact, the candidate relationship is marked as physically infeasible. If the rule allows it and there is an impact mechanism, a physical feasibility score is calculated, the score is stored in the physical feasibility scoring data, and candidate relationships with a physical feasibility score of zero or below the threshold are marked as low-confidence relationships.

[0156] Furthermore, the strength of the correlation between cause and effect is quantified from a statistical perspective.

[0157] Acquire causal candidate relationship data and facility version feature data set with physical feasibility tags. For each pair of antecedent event-consequence event combinations, count the frequency of the consequence event when the antecedent event occurs and the frequency of the consequence event when the antecedent event does not occur in the facility version feature data set. Calculate statistical indicators such as conditional probability and lift based on this, and organize these indicators into records in the statistical intensity score data.

[0158] For candidate relations with a physical feasibility score of zero or below the threshold, their statistical strength score is set to a default low value or not calculated, in order to avoid statistical noise interfering with the results and to ensure that physically infeasible relations are not incorrectly retained due to accidental statistical correlation during subsequent comprehensive scoring.

[0159] Based on this, a comprehensive causal score is obtained for each candidate relationship by combining physical and statistical scores.

[0160] Obtain physical feasibility score data and statistical strength score data, and for each candidate relationship, perform a comprehensive score using a pre-defined weighting strategy, such as defining a comprehensive causal score. total for:

[0161] Score total =α*S phy +β*S stat ;where S phy For physical feasibility scoring, S stat For statistical intensity scoring, α and β are configurable weight coefficients that satisfy α+β=1. Different weight combinations are set according to the problem type (such as abnormal energy consumption or frequent failures). The comprehensive score result of each relationship is recorded as comprehensive causal score data.

[0162] By combining the feedback from historical queries in the feedback evaluation dataset, the α and β weights of different types of candidate relationships are adjusted. For example, if users repeatedly report that a certain type of physically reasonable relationship with a small number of statistical samples is an effective relationship, the weight of the physical score is appropriately increased to achieve adaptive optimization of the comprehensive scoring strategy.

[0163] Furthermore, the selected causal edges are determined based on the comprehensive score and feedback information, and the threshold is dynamically adjusted.

[0164] Acquire comprehensive causal score data and historical feedback evaluation data set. Based on the target recall rate and fault tolerance requirements, set comprehensive score thresholds and backup threshold ranges for different problem types. Candidate relationships with comprehensive causal scores below the threshold are marked as non-selected relationships. Relationships with comprehensive causal scores above the threshold are selected into the filtered causal edge data set, which includes information such as facility logical identifiers, antecedent events, consequent events, and comprehensive causal scores.

[0165] Based on the evaluation of the effectiveness of the query results from the previous rounds by the operations and maintenance personnel in the feedback evaluation data set, the proportion of relationships that were confirmed or denied within different scoring intervals is statistically analyzed. If it is found that most relationships are denied within a certain scoring interval, the threshold is appropriately increased; otherwise, the threshold is decreased. The updated threshold parameters are written into the query optimization parameter data and used in the next round of comprehensive causal scoring calculation and weight adjustment, as well as causal threshold setting and dynamic learning, to achieve dynamic learning and continuous optimization of the threshold strategy.

[0166] According to one aspect of this application, an alternative to causal strength assessment based on physical constraints and data statistics is provided, specifically as follows:

[0167] First, lifecycle stages are labeled for the events involved in the candidate causal relationships, and a physical feasibility score is given based on the stage sequence and engineering rules.

[0168] Acquire causal candidate relationship data, facility version feature data set, and predefined design constraint knowledge base data, construction deviation knowledge base data, and operation strategy knowledge base data. Based on the time information and stage identifier in the facility version record, map each event to a specific life cycle stage category to obtain event stage label data including design stage, construction stage, commissioning stage, operation stage, modification stage, and failure stage.

[0169] Based on event phase label data, a phase number is assigned to each lifecycle phase, such as design phase 1, construction phase 2, operation phase 3, renovation phase 4, and failure phase 5. An allowed phase sequence relationship table is constructed. Relationships where the antecedent event phase number is less than or equal to the consequence event phase number are marked as reasonable phase sequence, and candidate relationships that violate this order are marked as unreasonable phase sequence.

[0170] For candidate relationships with reasonable phase sequence, the physical constraint rules recorded in the design constraint knowledge base, construction deviation knowledge base, and operation strategy knowledge base are combined to check whether the design changes, construction deviations, or operation strategy adjustments corresponding to the antecedent events are physically capable of affecting the performance indicators or failure types corresponding to the consequence events. For relationships that are physically impossible to have an impact, their physical feasibility score is set to 0. For relationships that may have a physical impact, a basic physical feasibility score is given based on the strength of the impact mechanism in the knowledge base, thus obtaining physical feasibility score data corresponding to each candidate relationship.

[0171] Candidate relationships with unreasonable stage order or a physical feasibility score of 0 are marked as physically infeasible in the causal candidate relationship data, in preparation for the subsequent steps to screen out these relationships.

[0172] Furthermore, by combining event stage tags, the candidate relationships that passed the initial screening of physical feasibility are grouped and statistically analyzed to calculate the statistical correlation.

[0173] We acquire causal candidate relationship data with stage sequence labels and physical feasibility scores, facility version feature data sets, and event stage label data. First, we filter out candidate relationships marked as physically infeasible and retain only candidate relationships with physical feasibility scores greater than 0 as the objects of statistical analysis.

[0174] All retained candidate relationships are grouped according to the combination of cause event stage - consequence event stage. For example, design stage -> failure stage, construction stage -> operation stage, operation stage -> failure stage, etc. are treated as different groups. For the candidate relationships in each group, the following information is statistically analyzed in the facility version feature data set: the frequency of consequence events when the cause event occurs, the frequency of consequence events when the cause event does not occur, and the frequency of consequence events within different time lag windows after the cause event occurs.

[0175] Based on the above statistical results, at least one statistical strength indicator is calculated for each candidate relationship, such as conditional probability, lift, or frequency-based correlation coefficient. The statistical indicators are organized into statistical strength score data corresponding to each candidate relationship, while retaining the stage combination label to which it belongs, so as to provide a basis for distinguishing different stage combinations in subsequent comprehensive scoring.

[0176] Furthermore, an inter-stage risk propagation factor is introduced, which integrates physical feasibility, statistical strength, and stage characteristics into a final causal score.

[0177] Acquire physical feasibility score data, statistical strength score data, and causal candidate relationship data with stage combination labels. For each candidate relationship, first calculate the distance between stages based on its antecedent event stage and consequence event stage (e.g., subtract the antecedent stage number from the consequence stage number). Then, according to the preset stage risk propagation rules, assign a stage risk factor r to different stage combinations. stage For example, the design phase -> failure phase r stage It can be set to a higher value, r during the running phase -> fault phase. stage Set to a medium value, the r value for the relationship between running events and running events within the same stage. stage Set to a lower value to reflect the importance of different phase combinations to the final risk propagation.

[0178] For each candidate relation, the physical score in the physical feasibility scoring data is denoted as S. phy The statistical score in the statistical intensity score data is denoted as S. stat Define the comprehensive causal score. total =α*S phy +β*S stat , where α and β are weight coefficients that satisfy α+β=1, and the corresponding weight configuration is selected from the pre-stored query optimization parameter data according to the problem type (such as abnormal energy consumption, frequent failures) and facility category.

[0179] After obtaining the Score total Then, the stage risk factor r is introduced. stage The risk propagation score is obtained by weighting the comprehensive basic score and recording the risk propagation score in the comprehensive causal score data, so that each candidate relationship has two fields: basic comprehensive score and risk propagation score.

[0180] Based on the historical feedback evaluation data set recording users' confirmation or denial of the combination relationship at different stages, the α, β, and r values ​​are evaluated. stage The values ​​are fine-tuned, and the new weights and stage risk configurations are updated in the query optimization parameter data for use in the next round of scoring.

[0181] Furthermore, based on risk propagation scores and dynamic threshold settings for stage combinations, a set of filtered causal edges is obtained.

[0182] Acquire comprehensive causal scoring data and historical feedback evaluation data sets. Based on different problem types and stage combinations, such as design->failure, construction->failure, operation->failure, operation->operation, etc., define an initial risk scoring threshold for each stage combination. Combine the confirmation and rejection rates of various relationships in the feedback evaluation data set to adjust these thresholds, so that the risk scoring threshold gradually increases in stage combinations that are rejected multiple times, and appropriately decreases in stage combinations that are confirmed multiple times.

[0183] For each candidate relationship, the corresponding risk score threshold is invoked according to its stage combination. The risk propagation score in the comprehensive causal score data is compared with the threshold. Relationships with risk propagation scores higher than the threshold are selected into the filtered causal edge data set, while relationships with risk propagation scores lower than the threshold are marked as not selected. For selected relationships, their facility logical identifier, antecedent event identifier, consequence event identifier, stage combination information, and risk propagation score value are retained in the filtered causal edge data set for subsequent use in building the causal chain model.

[0184] Based on the screening results and recent user feedback in the feedback evaluation dataset, update the threshold configurations corresponding to each stage combination, and adjust the updated thresholds, α, β, and r. stage These parameters form new query optimization parameter data, providing adaptive configuration for subsequent comprehensive scoring and filtering processes.

[0185] According to one aspect of this application, in the query and interpretation of facility lifecycle information based on causal chains, a facility lifecycle causal chain model is constructed, specifically as follows:

[0186] The filtered causal edge data set and the causal modeling candidate facility set data are used to construct a facility life cycle causal chain model, and causal chain index structure data is generated.

[0187] Obtain the filtered causal edge data set and the causal modeling candidate facility set data. For each facility logical identifier, filter out all causal edges related to that facility. Use the antecedent events and consequence events as nodes in the graph, and the causal edges as directed edges between nodes to construct facility causal graph data with facility logical identifiers as units.

[0188] Each causal edge is appended with its comprehensive causal score, event type information, and lifecycle stage information, and these attributes are saved in the facility causal graph data as the basis for subsequent path search and visualization interpretation.

[0189] Furthermore, the path structure of the facility causal graph is optimized to avoid excessively long or dense causal chains.

[0190] Acquire facility causal graph data, perform graph analysis algorithms on the causal graph of each facility, and identify highly dense local structures and excessively long causal path segments, such as highly repetitive intermediate nodes within the same stage or chains composed of multiple consecutive weak causal edges.

[0191] For highly repetitive and redundant intermediate nodes, multiple semantically similar intermediate events are aggregated into an aggregate node by merging nodes and converging edges. At the same time, the comprehensive causal score of related causal edges is updated (e.g., by taking the weighted average or the maximum value). While retaining the main influencing path information, the graph structure is compressed to obtain facility causal graph data with a simplified structure, making the causal path in subsequent queries simpler and easier to interpret.

[0192] Furthermore, an index structure is built for the causal graph based on common query patterns.

[0193] Obtain the optimized facility cause-effect graph data, and classify the nodes according to the event type (design, construction, operation, failure), life cycle stage (design stage, construction stage, operation stage, etc.) and problem type (abnormal energy consumption, frequent failures, etc.) based on the common query patterns in railway passenger station operation and maintenance scenarios. Generate a corresponding node index for each type of node set, and organize the index into a multi-dimensional node index table.

[0194] Using facility logical identifier, problem type, and time range as the primary key combination, edge indexes are constructed for directed edges in the causal graph. For example, an inverted index is created from the target event node to all possible antecedent event nodes, forming a causal chain index structure data. Finally, the facility causal graph data and the causal chain index structure data together constitute a facility full life cycle causal chain model.

[0195] This application employs a linear weighted fusion method of physical feasibility and statistical strength. By introducing an engineering constraint knowledge base for physical rule verification and combining it with joint frequency statistics of historical data, it effectively eliminates statistically relevant but physically infeasible noisy connections, improves the confidence of causal relationships, and solves the problem of spurious causality caused by the lack of fusion of engineering constraints and data statistics.

[0196] This application adopts a method of stage risk factor correction and multi-hop risk accumulation calculation. By defining the stages of design, construction, and operation, it uses nonlinear formulas to weight and correct cross-stage impacts, and calculates the multi-hop accumulation value of risk along the time axis based on a hierarchical graph structure. This achieves accurate quantitative tracing of the source from later failures to earlier design / construction defects, and solves the problem of lacking quantitative risk propagation based on life cycle stage sequence.

[0197] This application adopts an event-driven version generation and redundancy control method, abandoning fixed-frequency snapshots and instead monitoring key events (such as parameter freezing and component replacement) to trigger version generation. It also combines attribute similarity to merge redundant nodes. Under the premise of ensuring zero loss of key states, it achieves low-cost and high-precision historical state backtracking and solves the problem of rigid temporal model construction.

[0198] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A digital twin-based railway station facility full-life-cycle information query method, characterized in that, The method comprises the following steps: Based on multi-source heterogeneous basic data, unified facility objects with unique facility logical identifiers are generated through multi-dimensional attribute matching aggregation; Facility temporal version models containing time series nodes are generated based on event-driven rules by associating full life cycle event data with facility logical identifiers; Combining pre-stored engineering constraint knowledge base and statistical features extracted from facility temporal version models, the physical feasibility and statistical correlation strength between events are evaluated, and facility full life cycle causal chain models connecting design, construction and operation stages are constructed; Interpretive query requests for target facilities are parsed, and constrained path search is performed in facility full life cycle causal chain models to generate life cycle query result sets containing causal paths; Wherein, the facility temporal version models containing time series nodes are generated based on event-driven rules, which comprises the following steps: Define event trigger mode parameters containing parameter freezing, entity replacement and operation strategy adjustment features, and perform time sequence traversal on full life cycle event data; When an event meeting the event trigger mode parameters is monitored, the geometric shape, key technical parameters and running state of the unified facility object at that moment are extracted to generate an attribute snapshot; The attribute snapshot is packaged as version node data with an effective timestamp, and the version sequence relationship between adjacent version node data is recorded according to the generated time sequence to form a facility temporal version model; Redundancy control is performed on the facility temporal version model: Using version sequence relationship to locate time-continuous version node data, the difference degree of key attributes between adjacent nodes is calculated to generate version similarity; If the version similarity is higher than the preset merging threshold and the time interval is within the preset merging window, the change information of the later node is merged into the former node to generate updated version node data; The merged node is deleted in the facility temporal version model, and the updated version sequence relationship is re-established according to the updated version node data.

2. The method of claim 1, wherein, The facility full life cycle causal chain models connecting design, construction and operation stages are constructed, which comprises the following steps: Extract event pairs meeting time sequence constraints from the facility temporal version model as candidate causal relationships; According to the design specifications and physical rules in the pre-stored engineering constraint knowledge base, the influence mechanism of the cause event pair on the effect event in the candidate causal relationship is verified to generate a physical feasibility score of quantitative influence possibility; The joint occurrence frequency and conditional probability of the candidate causal relationship in the facility temporal version model are counted to generate a statistical strength score representing the degree of data association; The physical feasibility score and the statistical strength score are fused by weighting to obtain a comprehensive causal score, and the candidate causal relationship with a comprehensive causal score exceeding a preset threshold is established as an effective causal edge to organize and generate a facility full life cycle causal chain model.

3. The method of claim 2, wherein, The comprehensive causal score is obtained by: Identifying the life cycle stage identifiers of the cause event and the effect event in the candidate causal relationship, and matching the corresponding stage risk factors from the pre-stored stage risk factor table according to the span characteristics from early stage to later stage. The risk propagation score reflecting the cumulative effect of risk among life cycle stages is generated by modifying the weighted sum of the physical feasibility score and the statistical intensity score with the stage risk factor; The risk propagation score is used as the comprehensive causal score for screening effective causal edges.

4. The method of claim 3, wherein, The facility life cycle causal chain model generation facility includes: For each unified facility object, the facility causal graph data is constructed by connecting the corresponding event nodes with effective causal edges, and the stage layer nodes are arranged in time sequence, and the event nodes are mapped and associated to the corresponding stage layer nodes; In the facility causal graph data, the node risk propagation data representing the multi-hop accumulation of risk influence from early stages to later stages is calculated by performing forward topology traversal according to the comprehensive causal score carried by the effective causal edge; The facility life cycle causal chain model is encapsulated by the facility causal graph data including stage layer nodes, event nodes and node risk propagation data.

5. The method of claim 4, wherein, The explanatory query request for the target facility is parsed, and the constrained path search is performed in the facility life cycle causal chain model to generate a life cycle query result set including causal paths, including: The problem description and time condition in the explanatory query request are semantically analyzed to locate the target event node in the facility life cycle causal chain model, and search constraint data including search direction and maximum hop limit are generated; According to the search constraint data, the path backtracking is performed in the facility life cycle causal chain model with the target event node as the starting point to obtain candidate causal paths that meet the logical connectivity; The cumulative risk values of the candidate causal paths are sorted in combination with the node risk propagation data, and the high-risk associated paths are selected as the life cycle query result set.

6. The method of claim 5, wherein, Further including updating the model parameters by using feedback information: User feedback data is generated by collecting user confirmation or denial operations on the causal paths in the life cycle query result set; Based on the user feedback data, the path accuracy rate under the combination of the life cycle stage identifier is calculated, the stage risk factor and the weight coefficient in the weighted fusion are adjusted, and query optimization parameter data is generated; The query optimization parameter data is written back to the scoring model to generate updated stage risk factors and updated scoring strategies to dynamically optimize the selection logic of causal relationships when responding to subsequent queries.

7. The method of claim 1, wherein, The unified facility object with a unique facility logical identifier is generated by multi-dimensional attribute matching aggregation, including: Read multi-source heterogeneous basic data, perform primary clustering based on the spatial location proximity and professional category characteristics of the facility, and generate candidate record clusters containing multiple source records; Calculate the geometric spatial overlap rate and key attribute consistency between the source records in the candidate record cluster to generate a matching score quantifying the degree of identity; The records with a matching score exceeding a preset aggregation threshold are merged into a unified facility object with a unique physical entity, and a globally unique facility logical identifier is assigned to it according to a preset coding rule.

8. A computer electronic device, comprising: The processor and the memory are included, and the memory stores a computer program, when the processor executes the computer program, the method steps of any one of claims 1 to 7 are realized. The processor and the memory are included, and the memory stores a computer program, when the processor executes the computer program, the method steps of any one of claims 1 to 7 are realized.

Citation Information

Patent Citations

  • Road engineering low-carbon data credible interaction method and system

    CN120597253A

  • File storage system and method based on digital object architecture

    CN121051805A