A virtual-real fusion digital twin enhanced framework
Patent Information
- Application Number
- CN202610821759.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-11
AI Technical Summary
然而,仿真数据难以准确还原真实运营中的复杂分布特征,与实际数据之间存在系统性偏差
[0056] This invention provides a virtual-real fusion digital twin enhancement framework. The system utilizes a dual-diffusion model module, including a first diffusion model and a second diffusion model. Collected real-world rail transit data is input into the first diffusion model for data enhancement to generate reliable real-world data. Virtual data generated by a physical simulator module is input into the second diffusion model for distribution alignment enhancement to obtain aligned simulation data. The physical simulator module generates virtual data based on the real-world rail transit data and preset rail transit operation constraints, and outputs corresponding physical constraint residuals. A loss calculation module extracts trend deviations based on the temporal variation characteristics of the reliable real-world data and the aligned simulation data within the same time window, and generates loss information using a differentiated loss calculation method. The system uses a unified quantization approach based on the differences between reliable real data and aligned simulation data, physical constraint residuals, and trend deviations to generate a consistent virtual-real evolution potential. A threshold determination module performs virtual-real alignment determination, reliability verification determination, data filtering determination, and feedback optimization determination based on the loss signal and the consistent virtual-real evolution potential, generating a determination result. The virtual-real fusion digital twin module, based on real rail transit data, virtual data, reliable real data, aligned simulation data, and determination results, establishes a virtual-real mapping relationship through unified spatiotemporal indexing, rail transit object identification binding, and line topology matching. It also transmits feedback signals to the dual-diffusion model module and the physical simulator module based on the determination results. The resulting benefits include:
Smart Images

Figure CN122735441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital twins and data augmentation for rail transit, specifically a virtual-real fusion digital twin augmentation framework. Background Technology
[0002] Rail transit operation and management heavily rely on the accurate perception of multi-source spatiotemporal data, such as passenger flow, train operation, and station status. However, the raw data collected by automatic fare collection systems, train operation monitoring systems, and station detection equipment often contains noise, missing values, and abnormal offsets due to factors such as sensor errors and communication interruptions, which pose obstacles to intelligent decision-making tasks such as passenger flow prediction and capacity scheduling.
[0003] To address these issues, the industry has primarily explored two technological paths in recent years. The first is a data-driven approach, utilizing techniques such as statistical imputation, generative adversarial networks, or diffusion models to uncover spatiotemporal correlations within the data for completion and repair. While this method effectively restores the data's form, it lacks consideration of the physical operating mechanisms of rail transit, easily generating results that exceed station capacity or violate operational rules. The second approach is based on physical simulation, constructing traffic flow simulators to generate compliant simulation data under constraints such as passenger flow conservation, station capacity, and timetables. However, simulation data struggles to accurately reproduce the complex distribution characteristics of real operations, exhibiting systematic deviations from actual data.
[0004] Furthermore, existing systems typically treat data augmentation and simulation verification as independent processes, lacking a unified closed-loop feedback mechanism between modules. This makes it difficult to simultaneously determine the credibility of data during the augmentation process and also prevents the upstream modules from being optimized based on the verification results. Summary of the Invention
[0005] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a virtual-real fusion digital twin enhancement framework to solve the aforementioned technical problems.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a virtual-real fusion digital twin enhancement framework, comprising: a dual diffusion model module, a physical simulator module, a loss calculation module, a threshold determination module, and a virtual-real fusion digital twin module;
[0007] The dual diffusion model module includes a first diffusion model and a second diffusion model. The collected real rail transit data is input into the first diffusion model for data augmentation to generate reliable real data, and the virtual data generated by the physical simulator module is input into the second diffusion model for distribution alignment enhancement to obtain aligned simulation data.
[0008] The physics simulator module generates virtual data based on real rail transit data and preset rail transit operation constraints, and outputs the corresponding physical constraint residuals.
[0009] The loss calculation module extracts the trend deviation based on the temporal change characteristics of credible real data and aligned simulation data within the same time window, generates a loss signal using a differentiated loss calculation method, and performs unified quantization based on the difference information between credible real data and aligned simulation data, physical constraint residuals and trend deviations to generate a consistent virtual-real evolution potential.
[0010] The threshold determination module performs virtual-real alignment determination, credibility verification determination, data filtering determination, and feedback optimization determination based on the loss signal and the virtual-real consistency evolution potential to generate a determination result.
[0011] The virtual-real fusion digital twin module is based on real rail transit data, virtual data, credible real data, aligned simulation data and judgment results. It establishes a virtual-real mapping relationship through unified spatiotemporal index, rail transit object identification binding and line topology matching, and transmits feedback signals to the dual diffusion model module and physical simulator module according to the judgment results.
[0012] The present invention is further configured such that the collection and processing of the real rail transit data includes:
[0013] Raw data from rail transit is collected through multi-source data access methods, including data acquisition from automatic fare collection systems, train operation monitoring systems, station passenger flow detection equipment, transfer channel detection equipment, line dispatching systems, and operation management systems.
[0014] The raw rail transit data is processed by time alignment, station alignment, missing data marking, anomaly marking, normalization, and spatiotemporal coding to obtain standardized real rail transit data. The real rail transit data includes: rail transit object identifiers, passenger flow data, station data, line data, train operation data, transfer data, and operation status data.
[0015] The present invention is further configured such that the data augmentation process of the first diffusion model includes:
[0016] Data missing detection is performed on real rail transit data to determine the missing locations, and a missing location mask is generated based on the missing locations;
[0017] The first diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The attention mechanism is used to extract the spatiotemporal correlation features of real rail transit data between adjacent time windows, associated stations and associated lines.
[0018] Using real rail transit data and missing location masks as inputs to the first diffusion model, the distribution of real data is learned through forward noise addition and reverse denoising processes. In the reverse denoising process, missing regions are gradually filled in by combining missing location masks and spatiotemporal correlation features, and noise suppression and distribution calibration are performed on non-missing regions to generate reliable real data.
[0019] The present invention is further configured such that the distribution alignment enhancement process of the second diffusion model includes:
[0020] A second diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The second diffusion model adopts the same network structure as the first diffusion model.
[0021] The virtual data generated by the physical simulator module is used as the input of the second diffusion model, and the kernel density is estimated based on the real rail transit data to obtain the real side distribution density characteristics. The real side distribution density characteristics are used as the real side distribution constraint characteristics.
[0022] The distribution mapping relationship between virtual data and real-side distribution constraint features is learned through forward noise addition and reverse noise reduction processes. In the reverse noise reduction process, the distribution constraint features of the real side and the physical constraint residuals are combined to perform distribution alignment of the virtual data.
[0023] Based on the distributed and aligned virtual data, and in accordance with the preset distribution alignment constraints and feature consistency constraints, aligned simulation data is generated that retains the physical rationality of the virtual data while being consistent with the distribution of real rail transit data.
[0024] The present invention is further configured such that the virtual data generation process of the physics simulator module includes:
[0025] Virtual simulations are performed based on real rail transit data, preset rail transit operation constraints, and adjustable operating parameters to generate virtual data.
[0026] The constraints on rail transit operation include: passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints.
[0027] The adjustable operating parameters include: passenger flow loading coefficient, train speed, departure interval, station dwell time, transfer ratio, station entry and exit service capacity, and vehicle capacity.
[0028] Based on the degree of deviation of the virtual data from the constraints of passenger flow conservation, station capacity, vehicle capacity, train timetable, transfer route, and cross-sectional flow, physical constraint residuals are generated.
[0029] The present invention is further configured such that the trend deviation extraction includes:
[0030] Based on credible real data and aligned simulation data, time-series segments are divided into credible real data and aligned simulation data according to the same time window.
[0031] Match the corresponding reliable real data and aligned simulation data within the same time window;
[0032] By using sliding time window statistics, peak and valley period marking, and abrupt state identification, the temporal change features of the matched reliable real data and the aligned simulation data are extracted respectively. The temporal change features include: peak period passenger flow change trend, valley period passenger flow change trend, and sudden passenger flow change trend.
[0033] A trend bias is generated based on the deviation between the temporal variation characteristics of reliable real data and aligned simulation data.
[0034] The present invention is further configured such that the loss signal calculation includes:
[0035] For the data augmentation process of the first diffusion model, a first loss signal is generated based on the data reconstruction deviation between credible real data and real rail transit data, as well as the distribution smoothing deviation of credible real data relative to real rail transit data.
[0036] For the distribution alignment enhancement process of the second diffusion model, a second loss signal is generated based on the distribution alignment deviation of the aligned simulation data relative to the real rail transit data, and the consistency deviation between the aligned simulation data and the real rail transit data in the feature space.
[0037] For the virtual-real alignment determination process, a virtual-real alignment loss signal is generated based on the distribution difference between reliable real data and aligned simulation data;
[0038] For the trustworthy verification determination process, a trustworthy verification loss signal is generated based on the distribution difference, physical constraint residual, and trend deviation between trustworthy real data and aligned simulation data.
[0039] The first loss signal, the second loss signal, the virtual-real alignment loss signal, and the trusted verification loss signal are output as loss signals.
[0040] The present invention is further configured such that the generation of the virtual-real consistent evolution potential includes:
[0041] The differences between reliable real data and aligned simulation data, physical constraint residuals and trend deviations are normalized to obtain standardized virtual-real difference representations, physical deviation representations and trend deviation representations.
[0042] The contribution relationship of virtual-real difference representation, physical deviation representation and trend deviation representation is determined by a weighting method based on the loss signal. The weighting method includes determining the proportion of virtual-real difference representation, physical deviation representation and trend deviation representation in virtual-real consistency evaluation based on the virtual-real alignment loss signal, the credible verification loss signal, the deviation state of physical constraint residuals and the deviation state of trend deviation.
[0043] Based on the contribution relationship, the virtual-real difference representation, physical deviation representation, and trend deviation representation are uniformly quantified to generate a virtual-real consistent evolution potential.
[0044] The present invention is further configured such that the generation of the determination result includes:
[0045] The virtual-real alignment loss signal is compared with the preset virtual-real alignment threshold. Combined with the virtual-real consistency evolution potential, a joint judgment method combining threshold comparison and evolution potential level verification is used to generate the virtual-real alignment judgment result.
[0046] The trusted verification loss signal is compared with the preset trusted verification threshold. Combined with the virtual-real consistent evolution potential, a trusted verification judgment result is generated by a joint judgment method that combines threshold comparison and evolution potential level verification.
[0047] Based on the credibility verification judgment results, data filtering and judgment are performed on credible real data, aligned simulation data, and fused data generated from credible real data and aligned simulation data, and data output judgment results are generated.
[0048] When the virtual-real alignment judgment result or the credibility verification judgment result does not meet the preset judgment conditions, the feedback optimization object is determined based on the proportion relationship of virtual-real difference representation, physical deviation representation and trend deviation representation in the virtual-real consistency evolution potential, and the feedback optimization judgment result is generated.
[0049] The results of the virtual-real alignment judgment, the trusted verification judgment, the data output judgment, and the feedback optimization judgment are output as the judgment results.
[0050] The present invention is further configured such that the virtual-real fusion digital twin module includes a virtual-real mapping unit, an object binding unit, a topology matching unit, a judgment and display unit, and a feedback transmission unit, wherein:
[0051] The virtual-real mapping unit maps real rail transit data, virtual data, credible real data, aligned simulation data and judgment results to the same time base based on a preset unified spatiotemporal index.
[0052] The object binding unit binds real rail transit data, virtual data, trusted real data, and aligned simulation data corresponding to the same station, line, train, or transfer node based on the rail transit object identifier.
[0053] The topology matching unit maps the data that has been bound to the object to the corresponding rail transit line topology location based on the preset line topology relationship, and establishes a virtual-real mapping relationship.
[0054] The judgment display unit is used to display the judgment results and the virtual-real consistency evolution potential;
[0055] Based on the judgment result, the feedback transmission unit transmits feedback signals to the first diffusion model, the second diffusion model, and the physics simulator module.
[0056] This invention provides a virtual-real fusion digital twin enhancement framework. The system utilizes a dual-diffusion model module, including a first diffusion model and a second diffusion model. Collected real-world rail transit data is input into the first diffusion model for data enhancement to generate reliable real-world data. Virtual data generated by a physical simulator module is input into the second diffusion model for distribution alignment enhancement to obtain aligned simulation data. The physical simulator module generates virtual data based on the real-world rail transit data and preset rail transit operation constraints, and outputs corresponding physical constraint residuals. A loss calculation module extracts trend deviations based on the temporal variation characteristics of the reliable real-world data and the aligned simulation data within the same time window, and generates loss information using a differentiated loss calculation method. The system uses a unified quantization approach based on the differences between reliable real data and aligned simulation data, physical constraint residuals, and trend deviations to generate a consistent virtual-real evolution potential. A threshold determination module performs virtual-real alignment determination, reliability verification determination, data filtering determination, and feedback optimization determination based on the loss signal and the consistent virtual-real evolution potential, generating a determination result. The virtual-real fusion digital twin module, based on real rail transit data, virtual data, reliable real data, aligned simulation data, and determination results, establishes a virtual-real mapping relationship through unified spatiotemporal indexing, rail transit object identification binding, and line topology matching. It also transmits feedback signals to the dual-diffusion model module and the physical simulator module based on the determination results. The resulting benefits include:
[0057] Enhanced quality: The dual diffusion model works in conjunction with the physics simulator. The first diffusion model denoises and completes the real data while preserving its spatiotemporal features, while the second diffusion model aligns the simulation data to the real distribution and maintains physical plausibility. The two models are optimized in conjunction with the virtual-real alignment loss, which helps to alleviate the problem of physical infeasibility or distribution distortion in the enhanced results.
[0058] Trustworthy verification: The independent loss calculation module uses a differentiated loss calculation method to output multiple types of loss signals. The threshold determination module combines the virtual and real consistent evolution potential with the preset threshold to make a joint determination, thereby realizing the hierarchical evaluation of data trustworthiness and providing a quantifiable reference for data quality evaluation.
[0059] System performance: The five modules work together in a closed loop of virtual-real fusion alignment and threshold iteration verification, organically linking the real side, virtual side and model side, and driving the upstream module parameter optimization through feedback signals, thereby improving the system's adaptability to input data of different quality.
[0060] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0062] Figure 1 A schematic diagram of a virtual-real fusion digital twin enhancement framework is shown as an exemplary embodiment of the present invention;
[0063] Figure 2 This is a method architecture diagram illustrating a virtual-real fusion digital twin enhancement framework, which is an exemplary embodiment of the present invention. Detailed Implementation
[0064] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0065] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0066] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0067] Example:
[0068] A virtual-real integrated digital twin enhancement framework, such as Figure 1 As shown, it includes: a dual diffusion model module, a physics simulator module, a loss calculation module, a threshold determination module, and a virtual-real fusion digital twin module;
[0069] The dual diffusion model module includes a first diffusion model and a second diffusion model. The collected real rail transit data is input into the first diffusion model for data augmentation to generate reliable real data, and the virtual data generated by the physical simulator module is input into the second diffusion model for distribution alignment enhancement to obtain aligned simulation data.
[0070] The physics simulator module generates virtual data based on real rail transit data and preset rail transit operation constraints, and outputs the corresponding physical constraint residuals.
[0071] The loss calculation module extracts the trend deviation based on the temporal change characteristics of credible real data and aligned simulation data within the same time window, generates a loss signal using a differentiated loss calculation method, and performs unified quantization based on the difference information between credible real data and aligned simulation data, physical constraint residuals and trend deviations to generate a consistent virtual-real evolution potential.
[0072] The threshold determination module performs virtual-real alignment determination, credibility verification determination, data filtering determination, and feedback optimization determination based on the loss signal and the virtual-real consistency evolution potential to generate a determination result.
[0073] The virtual-real fusion digital twin module is based on real rail transit data, virtual data, credible real data, aligned simulation data and judgment results. It establishes a virtual-real mapping relationship through unified spatiotemporal index, rail transit object identification binding and line topology matching, and transmits feedback signals to the dual diffusion model module and physical simulator module according to the judgment results.
[0074] The present invention is further configured such that the collection and processing of the real rail transit data includes:
[0075] Raw data from rail transit is collected through multi-source data access methods, including data acquisition from automatic fare collection systems, train operation monitoring systems, station passenger flow detection equipment, transfer channel detection equipment, line dispatching systems, and operation management systems.
[0076] The raw rail transit data undergoes time alignment, station alignment, missing data marking, anomaly marking, normalization, and spatiotemporal coding to obtain standardized real rail transit data. This real rail transit data includes: rail transit object identifiers, passenger flow data, station data, line data, train operation data, transfer data, and operational status data. Specifically, raw rail transit data is collected through multi-source data access methods. These methods include: obtaining passenger flow data from the automatic fare collection system (ACS) for entering and exiting stations, gate throughput, and card swiping time; obtaining train numbers, arrival times, departure times, operating speeds, and operating sections from the train operation monitoring system; obtaining passenger flow data from station concourses, platforms, entrances / exits, and passageways from station passenger flow detection equipment; obtaining transfer node numbers, transfer directions, transfer passengers, and transfer times from transfer passage detection equipment; obtaining line numbers, timetables, departure intervals, capacity deployment, and temporary dispatch information from the line dispatching system; and obtaining operational status, equipment status, passenger flow control information, and abnormal event information from the operation management system. Data access can be achieved through interface collection, file import, or message subscription. Interface collection includes REST interfaces or database views; file import includes CSV, Excel, or JSON files; and message subscription includes MQTT, Kafka, or WebSocket. The collected raw rail transit data is first time-aligned. Time alignment uses a unified time window as the benchmark, with a default time window set to 15 minutes. For real-time monitoring scenarios, this can be set to 5 to 10 minutes, and for offline training scenarios, it can be set to 15 minutes. Data from the automatic fare collection system is assigned to the corresponding time window based on card swipe time; data from the train operation monitoring system is assigned to the corresponding time window based on arrival or departure time; data from passenger flow detection equipment is assigned to the corresponding time window based on collection time; and data from the line dispatching system and operation management system is assigned to the corresponding time window based on effective time. Passenger flow data is aggregated using a window summation method, operating speed data is aggregated using a window averaging method, and operational status data is supplemented using a previous value preservation method. After time alignment, the raw rail transit data is then station-aligned. Station alignment is performed based on a unified station coding table, which includes standard station numbers, station names, line numbers, transfer node numbers, station sequence numbers, and line topology location numbers. The gate numbers in the automatic fare collection system, the arrival and departure station codes in the train operation monitoring system, the equipment deployment locations in station passenger flow detection equipment, and the transfer node numbers in transfer passage detection equipment are all mapped to their corresponding numbers in the unified station coding table, ensuring consistent spatial identification for data from different sources. After station alignment, the raw rail transit data is marked for missing and anomalies. Missing data marking includes null value detection, sampling interruption detection, time window missing item detection, and object missing item detection; if data for the same rail transit object is not obtained in two or more consecutive time windows, it is marked as a sampling interruption missing item.Anomaly markers include one or more of the following: sliding window statistics, three-standard-deviation detection, median absolute deviation detection, and box plot detection. Sliding window statistics use three consecutive time windows as the detection window by default, and anomaly types include sudden increases, sudden decreases, continuous zero values, overcapacity, and operational status conflicts. Missing and anomaly markers do not directly delete the corresponding data; instead, they serve as auxiliary information for subsequent first diffusion modeling for missing data completion, noise suppression, and distribution calibration. Subsequently, the raw rail transit data undergoes normalization and spatiotemporal coding. Passenger flow data, operating speed, departure intervals, station dwell time, transfer ratio, station entry / exit service capacity, and vehicle capacity are converted to a zero-to-one range using the maximum-minimum normalization method. Station numbers, line numbers, train numbers, and transfer node numbers are processed using number mapping or embedded coding. Operational status data uses one-hot coding to represent normal operation, flow-limited operation, fault operation, delayed operation, and shutdown status. Spatiotemporal coding includes time coding and spatial coding. Time coding includes sampling time, time window number, weekday markers, holiday markers, peak period markers, off-peak period markers, and emergency event period markers. Peak periods by default include 7:00 to 9:00 and 17:00 to 19:00, while off-peak periods by default include 10:00 to 16:00 and 20:00 to 23:00. Emergency event periods are determined based on abnormal event information in the operation management system or detection results showing a passenger flow change rate exceeding 30% in adjacent time windows. Spatial coding includes station number, line number, train number, transfer node number, line direction, station sequence number, and line topology location number. Rail transit object identifiers are generated based on rail transit data types. Station-level data uses a combination of station number and line number to generate rail transit object identifiers; train-level data uses a combination of train number and line number; and transfer-level data uses a combination of transfer node number and line number. Rail transit object identifiers are used for subsequent object binding between reliable real data, virtual data, aligned simulation data, and judgment results. After the above processing, standardized real-world rail transit data is obtained. This standardized real-world rail transit data is organized using time window numbers and rail transit object identifiers as primary keys. Each data record includes the rail transit object identifier, passenger flow data, station data, line data, train operation data, transfer data, operational status data, missing data markers, anomaly markers, normalized values, time encoding, and spatial encoding. This standardized real-world rail transit data serves as the data augmentation input for the first diffusion model, the virtual-side simulation input for the physical simulator module, the foundational data for the second diffusion model to extract real-side distribution constraint features, and the data basis for the virtual-real fusion digital twin module to establish a unified spatiotemporal index and object binding relationships.
[0077] The present invention is further configured such that the data augmentation process of the first diffusion model includes:
[0078] Data missing detection is performed on real rail transit data to determine the missing locations, and a missing location mask is generated based on the missing locations;
[0079] The first diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The attention mechanism is used to extract the spatiotemporal correlation features of real rail transit data between adjacent time windows, associated stations and associated lines.
[0080] The first diffusion model uses real rail transit data and missing location masks as inputs. It learns the real data distribution through forward noise addition and reverse denoising. During reverse denoising, it progressively completes missing regions by combining missing location masks and spatiotemporal correlation features. For non-missing regions, it performs noise suppression and distribution calibration to generate reliable real data. Specifically, the first diffusion model enhances the reliability of standardized real rail transit data, targeting missing data, noisy data, anomalous offset data, and locally distorted data. The real rail transit data originates from previous data acquisition and processing steps. It is organized using time window numbers and rail transit object identifiers as primary keys. These identifiers correspond to stations, lines, trains, or transfer nodes. Data fields include passenger flow data, station data, line data, train operation data, transfer data, operational status data, missing data markers, anomaly markers, time encoding, and spatial encoding. The first diffusion model receives real rail transit data and missing location masks, and extracts spatiotemporal correlation features by combining time encoding, spatial encoding, and line topology relationships, ultimately outputting reliable real data. When performing missing data detection on real rail transit data, the missing detection results come from the preceding missing data markers. Alternatively, null value detection, time window missing data detection, and object missing data detection can be performed again before inputting the first diffusion model. Null value detection identifies data with empty fields, invalid fields, or fields not reported. Time window missing data detection identifies data lacking necessary passenger flow fields within the same time window. Object missing data detection identifies situations where a station, line, train, or transfer node lacks corresponding data within the same time window. After the missing location is determined, a missing location mask is generated according to the data dimensions of the real rail transit data. The missing location mask has the same time window dimension, rail transit object dimension, and data field dimension as the real rail transit data. Missing locations are marked as 0 by default, and non-missing locations are marked as 1 by default. The purpose of the missing location mask is to distinguish between data regions that need to be completed and data regions that need to maintain real observation constraints, thereby preventing the first diffusion model from destroying effective observation information in non-missing regions when completing missing data. The first diffusion model is constructed based on a U-Net backbone network with an embedded attention mechanism. The U-Net backbone network comprises an encoding section, a bottleneck section, and a decoding section. The encoding section extracts temporal variation features and spatial correlation features at different scales from real rail transit data. The decoding section progressively reconstructs the enhanced data representation. Skip connections between the encoding and decoding sections preserve details of local passenger flow changes. An attention mechanism is embedded in the encoding, bottleneck, or decoding sections to capture spatiotemporal correlation features between adjacent time windows, associated stations, and associated lines.Adjacent time windows are obtained based on their time window numbers, such as one or more time windows before and after the current time window; associated stations are obtained based on line topology relationships, such as upstream stations, downstream stations, and stations connected by the same transfer node on the same line; associated lines are obtained based on their line numbers and transfer node numbers, such as different lines connected by a transfer station. Through these methods, the first diffusion model can utilize temporal continuity, station adjacency, and line topology correlation during missing data completion, rather than relying solely on isolated data from a single time point or a single station. The first diffusion model is implemented using a conditional diffusion model. The forward noise addition process is used to gradually add noise to the real rail transit data during the training phase, enabling the first diffusion model to learn the real data distribution under different noise intensities. Gaussian noise can be used, and the noise scheduling strategy can be cosine scheduling or linear scheduling, preferably cosine scheduling to reduce the slow convergence problem caused by fixed noise scheduling. The default number of diffusion steps is set to 1000 steps, the default training learning rate is set to 5×10^-4, and the default training batch size is set to 32. The above parameters can be adjusted according to the training data scale. For example, when the data scale is small, the diffusion steps can be set to 500 steps, and when the data scale is large, it can be maintained at 1000 steps or increased to 1500 steps. The learning rate can be selected in the range of 1×10^-4 to 1×10^-3 through the validation set loss. The reverse denoising process is used to gradually recover reliable real data from noisy data. In the reverse denoising process, the first diffusion model simultaneously receives real rail transit data, missing location masks, time codes, spatial codes, and spatiotemporal correlation features. For missing regions, the first diffusion model gradually generates complete values based on passenger flow changes in adjacent time windows, passenger flow changes in associated stations, passenger flow changes in associated lines, and operational status information. For non-missing regions, the first diffusion model uses the original valid observation data as a preservation constraint, and only suppresses and corrects noise interference, abnormal spikes, and abnormal offsets, thereby avoiding excessive modification of the real observation data. For example, when a station has missing passenger flow within a 15-minute time window, the first diffusion model can combine passenger flow from previous and subsequent time windows, passenger flow from upstream and downstream stations on the same line, passenger flow from similar historical periods at the same station, and operational status codes to generate a complete result that conforms to the passenger flow change pattern. When an abnormal peak occurs in a non-missing time window, the first diffusion model can combine passenger flow from adjacent stations and related stations within the sliding window to suppress noise, preserve the true peak trend, and reduce abrupt changes caused by abnormal acquisition errors. During the training of the first diffusion model, loss constraints include data reconstruction bias and distribution smoothing bias. Data reconstruction bias measures the difference between reliable real data and real rail transit data in non-missing or effective observation areas, and can be implemented using the mean squared error algorithm. Distribution smoothing bias measures whether reliable real data deviates from the original distribution of real rail transit data, and can be implemented using the kernel density estimation algorithm.The distribution smoothing bias has a default weight of 0.3, used to balance reconstruction effectiveness and distribution consistency. Data reconstruction bias constrains credible real data to closely approximate actual observations, while distribution smoothing bias constrains credible real data to fall within a reasonable distribution range of actual rail transit data. Through these two types of constraints, the first diffusion model can both fill in missing regions and avoid generating abnormal data that does not conform to actual passenger flow patterns. Credible real data, as the output of the first diffusion model, maintains the same data organization structure as actual rail transit data. That is, credible real data is also organized using time window numbers and rail transit object identifiers as primary keys, and includes enhanced passenger flow data, station data, line data, train operation data, transfer data, and operational status data. Missing regions in the credible real data are filled in by the first diffusion model, while non-missing regions retain actual observation patterns after noise suppression and distribution calibration. The credible real data subsequently serves as the real-side input for the loss calculation module to extract trend bias and calculate virtual-real alignment loss, and as the data basis for the virtual-real fusion digital twin module to perform real-side mapping.
[0081] The present invention is further configured such that the distribution alignment enhancement process of the second diffusion model includes:
[0082] A second diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The second diffusion model adopts the same network structure as the first diffusion model.
[0083] The virtual data generated by the physical simulator module is used as the input of the second diffusion model, and the kernel density is estimated based on the real rail transit data to obtain the real side distribution density characteristics. The real side distribution density characteristics are used as the real side distribution constraint characteristics.
[0084] The distribution mapping relationship between virtual data and real-side distribution constraint features is learned through forward noise addition and reverse noise reduction processes. In the reverse noise reduction process, the distribution constraint features of the real side and the physical constraint residuals are combined to perform distribution alignment of virtual data.
[0085] Based on the distributed aligned virtual data, and according to preset distribution alignment constraints and feature consistency constraints, aligned simulation data is generated that maintains the physical rationality of the virtual data while being consistent with the distribution of real rail transit data. Specifically, the second diffusion model is used to enhance the distribution alignment of the virtual data generated by the physical simulator module, generating aligned simulation data. The virtual data is generated by the physical simulator module based on real rail transit data, rail transit operation constraints, and adjustable operating parameters. The virtual data includes virtual passenger flow entering the station, virtual passenger flow exiting the station, virtual transfer passenger flow, virtual train passenger capacity, virtual cross-sectional flow, and virtual operating status. The real rail transit data is obtained from the preceding data acquisition and processing flow, and is organized using time window number and rail transit object identifier as the primary key. The physical constraint residuals are output by the physical simulator module and are used to characterize the degree of deviation of the virtual data from passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints. The second diffusion model uses virtual data as the object to be aligned, takes the real-side distribution constraint features formed by real rail transit data as the alignment target, and uses physical constraint residuals as the basis for maintaining physical rationality. This ensures that the aligned simulation data closely approximates the distribution of real rail transit data while retaining the physical constraint features corresponding to the virtual data. The second diffusion model is built on a U-Net backbone network with an embedded attention mechanism and adopts the same network structure as the first diffusion model. The U-Net backbone network includes an encoding part, a bottleneck part, and a decoding part. The encoding part extracts multi-scale temporal variation features and spatial correlation features from the virtual data, while the decoding part restores the data representation after distribution alignment. The encoding and decoding parts are connected via skip connections to retain details of local passenger flow changes. The attention mechanism extracts the spatiotemporal correlation features of virtual data between adjacent time windows, associated stations, and associated lines, and guides the virtual data towards the real-side distribution constraint features. The second diffusion model uses the same network structure as the first diffusion model, ensuring consistency in feature extraction scale, time window processing methods, and line topology representation between the real-side enhancement links and the virtual-side alignment links. The first diffusion model enhances real rail transit data, while the second diffusion model performs distribution alignment on virtual data. The true side distribution constraint features are obtained from real rail transit data using the kernel density estimation method. The kernel density estimation method is used to estimate the probability density distribution of real rail transit data across different time windows, stations, lines, train operating states, and operational states. In practice, the real rail transit data is grouped according to time window numbers and rail transit object identifiers, and kernel density estimation is performed on passenger flow, transfer volume, cross-sectional flow, train passenger capacity, and operational state codes within each group to obtain the true side distribution density features.The kernel function can be a Gaussian kernel function. The kernel density estimation bandwidth can be determined using the Silverman empirical rule, the Scott empirical rule, or validation set cross-validation. The Silverman empirical rule is used as the initial bandwidth by default. The real-side distribution density features are input into the second diffusion model as real-side distribution constraint features to constrain the virtual data to closely approximate the real data distribution during distribution alignment. The second diffusion model is implemented using a conditional diffusion model. During the forward noise addition process, Gaussian noise is gradually added to the virtual data during the training phase, allowing the second diffusion model to learn the data distribution state of the virtual data under different noise intensities. The noise scheduling strategy can be cosine scheduling or linear scheduling, with cosine scheduling being preferred. The basic training parameters of the second diffusion model can be consistent with those of the first diffusion model. The default diffusion step count is set to 1000 steps, the default training learning rate is set to 5×10^-4, and the default training batch size is set to 32. The diffusion step count can be adjusted between 500 and 1500 steps depending on the virtual data size, and the training learning rate can be selected from 1×10^-4 to 1×10^-3 using the validation set loss. The second diffusion model's input includes virtual data, time encoding, spatial encoding, real-side distribution constraint features, and physical constraint residuals. Time encoding represents time windows and operating periods, spatial encoding represents stations, lines, trains, or transfer nodes, real-side distribution constraint features provide a reference for the true distribution, and physical constraint residuals provide a reference for deviations from the physical constraints. The reverse denoising process is used to gradually recover aligned simulation data from the noise-perturbed virtual data. During reverse denoising, the second diffusion model learns the distribution mapping relationship between virtual data and real-side distribution constraint features, and corrects areas in the virtual data that deviate significantly from the true distribution by incorporating the real-side distribution constraint features. For example, if the virtual data underestimates the passenger flow at a station during peak hours, the second diffusion model increases the virtual passenger flow distribution based on the corresponding time window and the real-side distribution density features of the corresponding station; if the virtual data shows a flow rate significantly higher than the real-side distribution density range at a certain interval, the second diffusion model compresses the flow rate distribution at the corresponding interval. Simultaneously, the reverse denoising process incorporates physical constraint residuals to prevent virtual data from exceeding rail transit operation constraints such as vehicle capacity, station capacity, train timetables, and transfer routes during distribution alignment. The second diffusion model training process employs distribution alignment constraints and feature consistency constraints. Distribution alignment constraints ensure consistency in probability distribution between aligned simulation data and real rail transit data. These constraints can be implemented using the Koolbek-Leibler divergence algorithm. The probability density in the Koolbek-Leibler divergence can be obtained using kernel density estimation methods. Real rail transit data and aligned simulation data share the same kernel density estimation framework, differing only in their input data sources, ensuring a unified evaluation standard for virtual and real data distributions and comparable results. The default weight for distribution alignment constraints is 0.5.Feature consistency constraints are used to ensure that the aligned simulation data and the real rail transit data remain consistent in the feature space. These constraints can be implemented using an adversarial consistency algorithm. This algorithm can employ a discriminator to distinguish between features in the real rail transit data and those in the aligned simulation data. The second diffusion model reduces the discriminator's ability to differentiate between the two types of features, gradually bringing the aligned simulation data closer to the real rail transit data in the feature space. The discriminator can be implemented using a multilayer perceptron, a 1D convolutional network, or a lightweight Transformer encoder; the default is a multilayer perceptron. Physical rationality preservation is achieved through the participation of physical constraint residuals in the reverse denoising process. After normalization, the physical constraint residuals are input into the second diffusion model along with the virtual data and the real-side distribution constraint features. Higher physical constraint residual values indicate a more significant deviation from the physical constraints for the corresponding time window or rail transit object. During the reverse denoising process, the second diffusion model adds physical preservation constraints to regions with high physical constraint residuals, ensuring that the aligned simulation data does not disrupt traffic flow patterns simply by closely approximating the real distribution. Physical preservation constraints can be implemented through residual gating, constraint masks, or physical deviation penalty terms. In engineering implementation, residual gating is preferred, with a default threshold of 0.1. When the physical constraint residual exceeds 0.1, the distribution adjustment amplitude for the corresponding region is reduced. The second diffusion model outputs aligned simulation data. The aligned simulation data maintains the same data organization structure as the virtual data; that is, it is also organized using time window numbers and rail transit object identifiers as primary keys, and includes aligned virtual passenger flow entering the station, virtual passenger flow exiting the station, virtual transfer passenger flow, virtual train passenger capacity, virtual cross-sectional flow, and virtual operating status. The aligned simulation data retains the physical rationality provided by the physical simulator module while approximating the distribution of real rail transit data through real-side distribution constraint features, distribution alignment constraints, and feature consistency constraints. The aligned simulation data is subsequently used as the virtual-side input for the loss calculation module to calculate virtual-real alignment loss, reliable verification loss, and trend deviation, and as the data basis for virtual-side mapping in the virtual-real fusion digital twin module.
[0086] The present invention is further configured such that the virtual data generation process of the physics simulator module includes:
[0087] Virtual simulations are performed based on real rail transit data, preset rail transit operation constraints, and adjustable operating parameters to generate virtual data.
[0088] The constraints on rail transit operation include: passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints.
[0089] The adjustable operating parameters include: passenger flow loading coefficient, train speed, departure interval, station dwell time, transfer ratio, station entry and exit service capacity, and vehicle capacity.
[0090] Based on the deviations of virtual data from constraints such as passenger flow conservation, station capacity, vehicle capacity, train timetable, transfer route, and cross-sectional flow, physical constraint residuals are generated. Specifically, the physical simulator module is used to construct virtual-side benchmark data for the digital twin of rail transit. The module uses standardized real rail transit data as driving input, combines rail transit operation constraints and adjustable operating parameters to perform virtual-side simulation, generate virtual data, and output physical constraint residuals. The virtual data serves as input to the second diffusion model for subsequent distribution alignment enhancement; the physical constraint residuals serve as physical rationality constraints, prompting the second diffusion model to retain rail transit operation patterns during distribution alignment. Real rail transit data comes from the preceding data acquisition and processing flow. This real data is organized using time window numbers and rail transit object identifiers as primary keys, and includes passenger flow data, station data, line data, train operation data, transfer data, and operational status data. Before simulation begins, the physical simulator module constructs a rail transit line topology model based on line numbers, station numbers, station sequence numbers, transfer node numbers, and line topology location numbers. The rail transit line topology model is represented using a node-edge structure, with stations, trains, and transfer nodes as topology nodes, and sections of track and transfer paths as topology edges. Using this topology model, the physical simulator module can determine the flow direction of passenger traffic between stations, trains, track sections, and transfer nodes. The physical simulator module can be implemented using discrete event simulation, queuing network models, passenger flow allocation models based on line topology, or passenger flow simulation algorithms based on individual agents. Discrete event simulation describes events such as train arrival, passenger entry, passenger boarding, passenger alighting, passenger transfer, and passenger exit; queuing network models describe the queuing and passage processes in turnstiles, escalators, platforms, and transfer passages; passenger flow allocation models based on line topology describe the distribution relationship of passenger traffic between track sections and transfer nodes; and passenger flow simulation algorithms based on individual agents simulate the movement behavior of individual passengers or passenger groups in refined scenarios. In engineering implementation, a combination of discrete event simulation and queuing network models is preferred to balance computational efficiency and physical interpretability. Rail transit operation constraints are used to define the physical boundaries of the virtual data. Passenger flow conservation constraints are used to maintain a balance between passenger flow entering, leaving, remaining at, and entering trains within the same time window. Station capacity constraints limit passenger flow in station halls, platforms, entrances / exits, and transfer passages from exceeding the station's spatial carrying capacity. Vehicle capacity constraints limit train passenger capacity from exceeding the rated or safe passenger capacity corresponding to the train type. Train timetable constraints ensure that train arrival times, departure times, operating sections, and departure intervals conform to the line operation schedule. Transfer path constraints limit the flow of transfer passengers along preset transfer paths between corresponding transfer nodes.Cross-sectional flow constraints are used to limit passenger flow between adjacent stations to match train frequency, train capacity, boarding / alighting passenger flow, and loitering passenger flow. Adjustable operating parameters are used to control the virtual simulation process. The passenger flow loading coefficient is used to convert the passenger flow in the real rail transit data into the input passenger flow for the virtual simulation. The default value of the passenger flow loading coefficient is 1.0, and the value range can be set from 0.8 to 1.2. The passenger flow loading coefficient can be adjusted based on historical passenger flow calibration results, holiday markings, emergency event markings, or feedback optimization judgment results. Train speed is used to control the train's running time within the line section. Train speed is preferentially obtained from the train operation monitoring system. When the train operation monitoring system data is missing, the historical average running speed of the same line and the same operating section is used. Departure interval is used to control the departure time difference between adjacent trains. The departure interval is preferentially obtained from the timetable in the line dispatching system. When the timetable data is missing, the historical average departure interval of the same line is used. The default range is 120 seconds to 600 seconds. Station dwell time controls the train's stop time at stations. Station dwell time is preferentially obtained from the train operation monitoring system or line dispatching system; when data is missing, the historical average dwell time is used. The default range is 20 to 60 seconds. Transfer ratio controls the distribution of passenger flow between different lines at the same transfer node. The transfer ratio can be determined by historical path inference results from the automatic fare collection system, statistical results from transfer channel detection equipment, or the historical average transfer ratio. Station entry / exit service capacity characterizes the passenger flow that turnstiles, escalators, passageways, and station halls can handle per unit time. Station entry / exit service capacity can be obtained from the operation management system, equipment parameter tables, or historical capacity statistics. Vehicle capacity characterizes the number of passengers a train can carry. Vehicle capacity is preferentially obtained from the train model configuration table or line dispatching system; when vehicle capacity is missing, the rated passenger capacity corresponding to the train formation on the same line is used. The virtual simulation process is executed sequentially according to time windows. First, the physical simulator module generates the simulation input passenger flow based on passenger flow data and passenger flow loading coefficients from real rail transit data, and loads the simulation input passenger flow onto the corresponding stations, lines, trains, or transfer nodes. Secondly, the physical simulator module simulates the processes of entering, exiting, and waiting within the station based on station capacity constraints and station service capabilities, generating virtual passenger flow, virtual passenger flow, and virtual passenger flow remaining within the station. Thirdly, the physical simulator module simulates train arrival, boarding / alighting, departure, and interval operation based on train speed, departure intervals, station dwell time, vehicle capacity, and train timetable constraints, generating virtual train passenger capacity and virtual train operation status. Subsequently, the physical simulator module simulates the transfer process between different lines based on transfer route constraints and transfer ratios, generating virtual transfer passenger flow. Finally, the physical simulator module generates virtual cross-sectional flow based on passenger flow between adjacent stations, passenger flow remaining within the station, train frequency, and vehicle capacity.Following the simulation process described above, the physical simulator module generates virtual data, including virtual passenger flow entering the station, virtual passenger flow exiting the station, virtual transfer passenger flow, virtual train passenger capacity, virtual cross-sectional flow, and virtual operational status. Physical constraint residuals are used to quantify the degree of deviation of the virtual data from the constraints of rail transit operation. The physical simulator module calculates the degree of deviation of the virtual data from passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints, and normalizes the different types of deviation to a range of 0 to 1. Passenger flow conservation residuals are used to represent the degree of imbalance between inflow, outflow, and retention within the same time window; station capacity residuals are used to represent the degree to which passenger flow in station halls, platforms, entrances / exits, or transfer passages exceeds the spatial carrying capacity; vehicle capacity residuals are used to represent the degree to which train passenger capacity exceeds vehicle capacity; train timetable residuals are used to represent the degree to which virtual train arrival time, departure time, or departure interval deviates from the timetable; transfer path residuals are used to represent the degree to which transfer passenger flow deviates from the preset transfer path or transfer ratio; and cross-sectional flow residuals are used to represent the degree of mismatch between passenger flow in the line section and train frequency, vehicle capacity, and boarding / alighting passenger flow. The aforementioned residuals can be aggregated according to preset weights to form physical constraint residuals. The default weights can be set as follows: passenger flow conservation residual 0.2, station capacity residual 0.15, vehicle capacity residual 0.2, train timetable residual 0.15, transfer path residual 0.15, and cross-sectional flow residual 0.15. During engineering deployment, the weights of vehicle capacity residuals or station capacity residuals can be increased according to the line's operational safety requirements. The underlying logic of physical constraint residuals is that virtual data not only needs to closely approximate the distribution of real rail transit data but also needs to conform to the basic operational rules of the rail transit system. For example, if the virtual passenger flow at a station increases significantly within a certain time window, but the corresponding train schedules and vehicle capacity cannot accommodate the increased passenger flow, then the vehicle capacity residual and cross-sectional flow residual will increase; if the virtual transfer passenger flow at a transfer node is assigned to a non-existent transfer path, then the transfer path residual will increase; if the virtual platform passenger flow at a station exceeds the platform's carrying capacity, then the station capacity residual will increase. A higher physical constraint residual indicates that the virtual data does not conform to physical operational constraints within the corresponding time window or for the corresponding rail transit object. After generating virtual data and physical constraint residuals, the physical simulator module outputs the virtual data to the second diffusion model and simultaneously outputs the physical constraint residuals to both the second diffusion model and the loss calculation module. The second diffusion model utilizes the physical constraint residual constraint distribution alignment process to avoid the aligned simulation data from violating physical rules such as vehicle capacity, station capacity, train timetables, and transfer routes by simply approximating the distribution of real rail transit data. The loss calculation module uses the physical constraint residuals to participate in the generation of the reliable verification loss signal and the virtual-real consistent evolution potential.In this way, the physics simulator module provides an interpretable virtual-side physical benchmark for the virtual-real fusion digital twin enhancement framework, and provides a physical constraint basis for subsequent virtual-real alignment and trusted verification.
[0091] The present invention is further configured such that the trend deviation extraction includes:
[0092] Based on credible real data and aligned simulation data, time-series segments are divided into credible real data and aligned simulation data according to the same time window.
[0093] Match the corresponding reliable real data and aligned simulation data within the same time window;
[0094] By using sliding time window statistics, peak and valley period marking, and abrupt state identification, the temporal change features of the matched reliable real data and the aligned simulation data are extracted respectively. The temporal change features include: peak period passenger flow change trend, valley period passenger flow change trend, and sudden passenger flow change trend.
[0095] A trend bias is generated based on the deviation between the temporal variation characteristics of the reliable real data and the aligned simulation data. Specifically, the trend bias measures the degree of inconsistency between the reliable real data and the aligned simulation data in terms of their temporal variation patterns. The trend bias serves as the input for the loss calculation module to generate the reliable verification loss signal and the virtual-real consistency evolution potential. The reliable real data is output by the first diffusion model, and the aligned simulation data is output by the second diffusion model. Both reliable real data and aligned simulation data are organized using time window numbers and rail transit object identifiers as primary keys. The rail transit object identifier corresponds to a station, line, train, or transfer node. Therefore, reliable real data and aligned simulation data can be matched one-to-one within the same time window and for the same rail transit object. Before extracting the trend bias, the reliable real data and aligned simulation data are first divided into temporal segments according to the same time window. The time window length is consistent with the rail transit real data acquisition and processing stage, with a default setting of 15 minutes. For real-time monitoring scenarios, this can be set to 5 to 10 minutes, and for offline training scenarios, it can be set to 15 minutes. When dividing time-series segments, a trend analysis segment is formed by multiple consecutive time windows. The default trend analysis segment contains three consecutive time windows, but this can be extended to five or seven consecutive time windows based on passenger flow fluctuation cycles. Time window numbers are used to determine the start and end positions of the time-series segment, and rail transit operation segment markers are used to distinguish between peak, off-peak, and normal periods. After completing the time-series segment division, reliable real data and aligned simulation data corresponding to the same rail transit object within the same time window are matched. The matching process is based on rail transit object identifiers: station-level data is matched according to station number and line number, train-level data is matched according to train number and line number, and transfer-level data is matched according to transfer node number and line number. The matched reliable real data and aligned simulation data have the same time window, the same station or line object, and the same data fields, such as the same station's inbound passenger flow, outbound passenger flow, or transfer passenger flow within the same 15-minute time window. Time-series change characteristics are extracted through sliding time window statistics, peak and off-peak period markers, and abrupt change state identification. Sliding time window statistics are used to extract the direction, magnitude, and duration of passenger flow changes within consecutive time segments. Passenger flow changes can be categorized as upward, downward, or stable. The magnitude of these changes is determined by the difference or rate of change in passenger flow between adjacent time windows. The duration of a change is determined by the number of consecutive upward, downward, or stable windows. Peak and off-peak periods are used to differentiate trend characteristics across different operating hours. Peak hours by default include 7:00 to 9:00 and 17:00 to 19:00, while off-peak periods by default include 10:00 to 16:00 and 20:00 to 23:00. Peak and off-peak periods can also be automatically determined based on the time periods in historical passenger flow distribution where passenger volume falls within the top 20% or bottom 20%.Abrupt change identification is used to identify sudden changes in passenger flow within a short period. The default trigger condition is a passenger flow change rate exceeding 30% in adjacent time windows; alternatively, the historical mean plus twice the standard deviation can be used. When extracting temporal variation features from reliable real-world data, passenger flow data from the reliable real-world data is used as a basis, combined with time coding, rail transit object identification, and operational status data to obtain peak-hour passenger flow trends, off-peak-hour passenger flow trends, and sudden passenger flow trends. When extracting temporal variation features from aligned simulation data, virtual inbound passenger flow, virtual outbound passenger flow, virtual transfer passenger flow, virtual train passenger capacity, and virtual cross-sectional flow rate from the aligned simulation data are used as a basis. The same sliding time window, the same peak-valley time period markings, and the same abrupt change identification rules are used to obtain the corresponding peak-hour passenger flow trends, off-peak-hour passenger flow trends, and sudden passenger flow trends in the aligned simulation data. The use of the same feature extraction rules for reliable real-world data and aligned simulation data ensures consistency in trend deviation evaluation standards. Trend deviation is generated based on the deviation between the temporal variation features of the reliable real-world data and the aligned simulation data. Deviation states include trend direction deviation, trend magnitude deviation, peak / valley position deviation, and sudden state deviation. Trend direction deviation indicates that the credible real data shows an upward trend while the aligned simulation data shows a downward or stable trend, or vice versa. Trend magnitude deviation indicates that the difference in passenger flow change magnitude between the credible real data and the aligned simulation data within the same time window exceeds a preset range, which is set to 0.1 times the normalized passenger flow change magnitude by default. Peak / valley position deviation indicates that the peak or trough occurrence times in the credible real data are inconsistent with those in the aligned simulation data; the default allowable offset range is one time window. Sudden state deviation indicates that a sudden passenger flow change is identified in the credible real data, but not in the aligned simulation data, or that a sudden change not present in the credible real data appears in the aligned simulation data. When generating trend deviations, peak-hour passenger flow trend deviations, trough-hour passenger flow trend deviations, and sudden passenger flow trend deviations are normalized to a range of 0 to 1. The closer the value is to 0, the more consistent the reliable real data and the aligned simulation data are in the corresponding trend type; the closer the value is to 1, the more significant the deviation in the corresponding trend type. By default, the weight for deviation of passenger flow trend during peak hours is set to 0.4, the weight for deviation during off-peak hours is set to 0.2, and the weight for deviation during sudden passenger flow trends is set to 0.4. Peak hours and sudden scenarios have a greater impact on the safety of rail transit operations; therefore, higher weights are used for deviations in passenger flow trends during peak hours and sudden passenger flow trends. During engineering deployment, the weights can be adjusted according to the line operation strategy. For example, commuter lines can have a higher weight for peak hours, while tourist lines or lines serving event venues can have a higher weight for sudden passenger flow.Through the above processing, the loss calculation module obtains the trend bias. The trend bias reflects the difference in temporal evolution between the reliable real data and the aligned simulation data, supplementing dynamic change information that distribution differences cannot reflect. When only distribution differences are used, the two datasets may be similar in overall distribution, but there may be differences in peak occurrence time, sudden passenger flow response, or trough decline speed. Introducing the trend bias allows the reliability verification process to simultaneously consider distribution consistency and temporal evolution consistency, thereby improving the accuracy of the enhanced data reliability judgment. The trend bias is subsequently input into the reliability verification loss signal calculation process and participates in the generation of the virtual-real consistency evolution potential.
[0096] The present invention is further configured such that the loss signal calculation includes:
[0097] For the data augmentation process of the first diffusion model, a first loss signal is generated based on the data reconstruction deviation between credible real data and real rail transit data, as well as the distribution smoothing deviation of credible real data relative to real rail transit data.
[0098] For the distribution alignment enhancement process of the second diffusion model, a second loss signal is generated based on the distribution alignment deviation of the aligned simulation data relative to the real rail transit data, and the consistency deviation between the aligned simulation data and the real rail transit data in the feature space.
[0099] For the virtual-real alignment determination process, a virtual-real alignment loss signal is generated based on the distribution difference between reliable real data and aligned simulation data;
[0100] For the trustworthy verification determination process, a trustworthy verification loss signal is generated based on the distribution difference, physical constraint residual, and trend deviation between trustworthy real data and aligned simulation data.
[0101] The first loss signal, the second loss signal, the virtual-real alignment loss signal, and the reliable verification loss signal are output as loss signals. Specifically, the loss calculation module is set up independently of the dual diffusion model module, the physical simulator module, and the threshold determination module. The loss calculation module is only used for loss signal calculation and quantization result output, and does not directly participate in the generation of reliable real data, virtual data generation, generation of aligned simulation data, model parameter optimization, or judgment decision. By setting the loss calculation module independently, the first diffusion model, the second diffusion model, the physical simulator module, and the threshold determination module can use a unified quantization standard, avoiding inconsistencies in evaluation criteria caused by different modules calculating losses separately. The data received by the loss calculation module includes real rail transit data, reliable real data, aligned simulation data, physical constraint residuals, and trend deviations. Among them, the real rail transit data comes from the acquisition and processing of real rail transit data, the reliable real data comes from the first diffusion model, the aligned simulation data comes from the second diffusion model, the physical constraint residuals come from the physical simulator module, and the trend deviation comes from the deviation results of the temporal change characteristics between the reliable real data and the aligned simulation data within the same time window. For the data augmentation process of the first diffusion model, the loss calculation module generates the first loss signal. The first loss signal measures whether the reliable real data generated by the first diffusion model both closely matches the effective observation information in the real rail transit data and maintains the original distribution pattern of the real rail transit data. The first loss signal includes data reconstruction bias and distribution smoothing bias. Data reconstruction bias measures the numerical difference between the reliable real data and the real rail transit data in the non-missing region or effective observation region. Data reconstruction bias can be implemented using the mean squared error algorithm. The number of samples in the mean squared error algorithm is determined by the number of data samples participating in training or evaluation within the same batch. The number of samples can be obtained by expanding the data according to the time window number, rail transit object identifier, and data fields. The effective observation region in the real rail transit data is determined by the non-missing positions in the missing position mask. The corresponding region in the reliable real data is obtained by matching the output results of the first diffusion model according to the same time window number and rail transit object identifier. Distribution smoothing bias measures whether the reliable real data deviates from the real data distribution formed by the real rail transit data. Distribution smoothing bias can be implemented using the kernel density estimation algorithm. The kernel density estimation algorithm uses real rail transit data as distribution reference data to calculate the degree to which reliable real data sample points fall within the probability density region of real rail transit data. The preferred kernel function is a Gaussian kernel function, and the kernel density estimation bandwidth is obtained by default using the Silverman empirical rule, but can also be determined through cross-validation on a validation set. The first loss signal is generated jointly by data reconstruction bias and distribution smoothing bias, with a default weight of 0.3 for distribution smoothing bias and 1.0 for data reconstruction bias. Using the first loss signal, the first diffusion model can fill in missing data and suppress noise while avoiding generating data that deviates from the actual passenger flow patterns.For the distribution alignment enhancement process of the second diffusion model, the loss calculation module generates a second loss signal. This second loss signal measures whether the aligned simulation data generated by the second diffusion model is consistent with the real rail transit data in terms of distribution and feature space. The second loss signal includes distribution alignment bias and feature consistency bias. Distribution alignment bias measures the difference in probability distribution between the aligned simulation data and the real rail transit data. This bias can be implemented using the Koolbek-Leibler divergence algorithm. The probability density in the Koolbek-Leibler divergence is obtained through a kernel density estimation algorithm. The real rail transit data and the aligned simulation data share the same kernel density estimation framework, differing only in the input data source, thus ensuring a unified evaluation standard for the distribution of real and virtual data and comparable calculation results. The default weight of distribution alignment bias is 0.5, used to enhance the adaptation effect of virtual data to the real distribution. Feature consistency bias measures the difference between the aligned simulation data and the real rail transit data in the feature space. This bias can be implemented using an adversarial consistency algorithm. The adversarial consensus algorithm includes a discriminator that takes as input features of real rail transit data and aligned simulation data features, and outputs a feature source discrimination result. The second diffusion model reduces the discriminator's ability to distinguish between features of real rail transit data and aligned simulation data, allowing the aligned simulation data to gradually approach real rail transit data in the feature space. The discriminator can be implemented using a multilayer perceptron, a 1D convolutional network, or a lightweight Transformer encoder; a multilayer perceptron is used by default to reduce computational complexity. Through a second loss signal, the second diffusion model can reduce the real-world distribution deviation between virtual data and real rail transit data while preserving the physical plausibility of the virtual data. For the virtual-real alignment determination process, the loss calculation module generates a virtual-real alignment loss signal. This signal measures the distribution difference between reliable real data and aligned simulation data. The virtual-real alignment loss signal can be implemented using a first-order Wasserstein distance algorithm. The first-order Wasserstein distance measures the minimum transmission cost required to transform one distribution into another, and is suitable for handling potential asymmetric distribution differences between reliable real data and aligned simulation data. In engineering implementation, to reduce computational complexity, the Sinkhorn iterative algorithm can be used to approximate the first-order Wasserstein distance. Before calculation, the reliable real data and aligned simulation data are aligned according to the same time window number, rail transit object identifier, and data fields to ensure that distribution differences originate from the same time reference and the same rail transit object. The smaller the virtual-real alignment loss signal, the more consistent the reliable real data and aligned simulation data are in terms of distribution; the larger the virtual-real alignment loss signal, the larger the deviation between the real-side enhancement result and the virtual-side alignment result. For the reliability verification judgment process, the loss calculation module generates a reliability verification loss signal.The trusted verification loss signal is used to determine whether the enhancement results meet the final trusted output conditions. The trusted verification loss signal not only considers the distribution differences between trusted real data and aligned simulation data, but also further introduces physical constraint residuals and trend deviations. Distribution differences represent the overall distribution consistency between the real-side enhancement results and the virtual-side alignment results. Physical constraint residuals represent the degree of physical deviation of the virtual data or aligned simulation data from constraints such as passenger flow conservation, station capacity, vehicle capacity, train timetable, transfer route, and cross-sectional flow. Trend deviations represent the degree of dynamic deviation between trusted real data and aligned simulation data in terms of peak-hour passenger flow trends, off-peak-hour passenger flow trends, and sudden passenger flow trends. The trusted verification loss signal can be generated using a weighted fusion method. The default weights can be set to 0.6 for distribution differences, 0.2 for physical constraint residuals, and 0.2 for trend deviations. During engineering deployment, the weights can be adjusted according to the line operation objectives. For example, lines with high safety constraints should increase the weight of physical constraint residuals, commuter lines should increase the weight of peak-hour trend deviations, and lines with frequent sudden passenger flow should increase the weight of trend deviations. The smaller the trustworthy verification loss signal, the more suitable the augmentation result is to be directly used as augmented trustworthy data output. The first loss signal, the second loss signal, the virtual-real alignment loss signal, and the trustworthy verification loss signal together constitute the loss signal output by the loss calculation module. The first loss signal mainly serves the data augmentation quality assessment of the first diffusion model, the second loss signal mainly serves the virtual data alignment quality assessment of the second diffusion model, the virtual-real alignment loss signal mainly serves the virtual-real alignment determination, and the trustworthy verification loss signal mainly serves the trustworthy verification determination and data filtering determination. The first loss signal and the second loss signal correspond to the real data augmentation branch and the virtual data alignment augmentation branch, respectively. The virtual-real alignment loss signal compares the outputs of the two branches in a unified manner, and the trustworthy verification loss signal further combines physical constraint residuals and trend deviations to evaluate the trustworthiness of the augmentation result. Through the above differentiated loss calculation method, the loss calculation module can provide a unified, traceable, and quantifiable basis for the virtual-real fusion alignment closed loop and the threshold iteration verification closed loop.
[0102] The present invention is further configured such that the generation of the virtual-real consistent evolution potential includes:
[0103] The differences between reliable real data and aligned simulation data, physical constraint residuals and trend deviations are normalized to obtain standardized virtual-real difference representations, physical deviation representations and trend deviation representations.
[0104] The contribution relationship of virtual-real difference representation, physical deviation representation and trend deviation representation is determined by a weighting method based on the loss signal. The weighting method includes determining the proportion of virtual-real difference representation, physical deviation representation and trend deviation representation in virtual-real consistency evaluation based on the virtual-real alignment loss signal, the credible verification loss signal, the deviation state of physical constraint residuals and the deviation state of trend deviation.
[0105] The virtual-real difference representation, physical deviation representation, and trend deviation representation are uniformly quantified according to their contribution relationships to generate a virtual-real consistency evolution potential. Specifically, the virtual-real consistency evolution potential is used to uniformly represent the comprehensive consistency deviation between credible real data and aligned simulation data, and serves as the quantitative basis for the threshold judgment module to perform virtual-real alignment judgment, credibility verification judgment, and feedback optimization judgment. The virtual-real consistency evolution potential is not a single loss value, but is jointly generated based on the difference information between credible real data and aligned simulation data, physical constraint residuals, and trend deviations, and is used to simultaneously reflect the deviation status at the distribution level, physical level, and temporal evolution level. When generating the virtual-real consistency evolution potential, the difference information between credible real data and aligned simulation data is first normalized to obtain the virtual-real difference representation. The difference information can be represented by the virtual-real alignment loss signal or the first-order Wasserstein distance. The credible real data and aligned simulation data are matched according to the same time window number, the same rail transit object identifier, and the same data fields before participating in the calculation. The virtual-to-real difference representation is normalized to a range of 0 to 1 by default. The closer the value is to 0, the more consistent the distribution between the reliable real data and the aligned simulation data; the closer the value is to 1, the more significant the distribution deviation between the augmentation results on the real side and the alignment results on the virtual side. Simultaneously, the physical constraint residuals are normalized to obtain the physical deviation representation. The physical constraint residuals are generated by the physical simulator module and are used to represent the degree of deviation of the virtual data from passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints. The physical deviation representation is normalized to a range of 0 to 1 by default. The closer the value is to 0, the more the virtual data or aligned simulation data conforms to the rail transit operation constraints; the closer the value is to 1, the more significant the deviation from the physical constraints. The default weights for each type of residual in the physical constraint residuals can be set as follows: passenger flow conservation residual 0.2, station capacity residual 0.15, vehicle capacity residual 0.2, train timetable residual 0.15, transfer route residual 0.15, and cross-sectional flow residual 0.15. Simultaneously, the trend deviation is normalized to obtain a trend deviation characterization. The trend deviation, generated during the trend deviation extraction process, characterizes the deviation between reliable real data and aligned simulation data in terms of peak-hour passenger flow trends, off-peak-hour passenger flow trends, and sudden passenger flow trends. The trend deviation characterization is normalized to a range of 0 to 1 by default; the closer the value is to 0, the more consistent the temporal evolution of the reliable real data and aligned simulation data; the closer the value is to 1, the more significant the temporal evolution deviation. The default weights for various trend deviations can be set as follows: peak-hour passenger flow trend deviation 0.4, off-peak-hour passenger flow trend deviation 0.2, and sudden passenger flow trend deviation 0.4. After obtaining the virtual-real difference characterization, physical deviation characterization, and trend deviation characterization, a weighting method is used based on the loss signal to determine the contribution relationship of the three characterizations.The initial contribution ratios are set by default as follows: virtual-real difference representation 0.4, physical deviation representation 0.3, and trend deviation representation 0.3. If the virtual-real alignment loss signal is higher than the preset virtual-real alignment threshold, or if the virtual-real difference representation is dominant among the three types of representations, the proportion of the virtual-real difference representation is increased; if the deviation state of the physical constraint residual is dominant among the three types of representations, the proportion of the physical deviation representation is increased; if the deviation state of the trend deviation is dominant among the three types of representations, the proportion of the trend deviation representation is increased. The step size for adjusting the proportions is set by default to 0.1, and the proportion of each item is set by default to be no less than 0.2 and no more than 0.6. After adjustment, renormalization is performed so that the sum of the proportions of the virtual-real difference representation, physical deviation representation, and trend deviation representation remains at 1. Subsequently, the virtual-real difference representation, physical deviation representation, and trend deviation representation are uniformly quantified according to the contribution relationship to generate a virtual-real consistent evolution potential. The virtual-real consistency evolution potential is normalized to the range of 0 to 1. The closer the value is to 0, the better the distribution consistency, physical rationality, and trend consistency between the reliable real data and the aligned simulation data; the closer the value is to 1, the higher the risk of deviation in the virtual-real fusion state. In engineering implementation, the virtual-real consistency evolution potential can be divided into low deviation level, medium deviation level, and high deviation level according to 0 to 0.1, 0.1 to 0.3, and greater than 0.3. The low deviation level is used to support the completion of virtual-real alignment or the determination of reliable output, the medium deviation level is used to trigger local optimization, and the high deviation level is used to trigger feedback optimization. Through the above processing, the virtual-real consistency evolution potential can unify the differences in virtual-real distribution, deviations in physical constraints, and deviations in trend changes into the same quantitative index, so that the threshold determination module can not only determine whether the reliable real data and the aligned simulation data meet the alignment requirements, but also identify the source of deviation. When the proportion of virtual-real difference representation is high, the feedback optimization target is prioritized for the second diffusion model; when the proportion of physical deviation representation is high, the feedback optimization target is prioritized for the physical simulator module; when the proportion of trend deviation representation is high, the feedback optimization target can be one or more of the first diffusion model, the second diffusion model, and the physical simulator module. Therefore, the virtual-real consistent evolution potential serves as a unified feedback quantization parameter between the virtual-real fusion alignment closed loop and the threshold iteration verification closed loop.
[0106] The present invention is further configured such that the generation of the determination result includes:
[0107] The virtual-real alignment loss signal is compared with the preset virtual-real alignment threshold. Combined with the virtual-real consistency evolution potential, a joint judgment method combining threshold comparison and evolution potential level verification is used to generate the virtual-real alignment judgment result.
[0108] The trusted verification loss signal is compared with the preset trusted verification threshold. Combined with the virtual-real consistent evolution potential, a trusted verification judgment result is generated by a joint judgment method that combines threshold comparison and evolution potential level verification.
[0109] Based on the credibility verification judgment results, data filtering and judgment are performed on credible real data, aligned simulation data, and fused data generated from credible real data and aligned simulation data, and data output judgment results are generated.
[0110] When the virtual-real alignment judgment result or the credibility verification judgment result does not meet the preset judgment conditions, the feedback optimization object is determined based on the proportion relationship of virtual-real difference representation, physical deviation representation and trend deviation representation in the virtual-real consistency evolution potential, and the feedback optimization judgment result is generated.
[0111] The threshold determination module outputs the virtual-real alignment judgment result, the credibility verification judgment result, the data output judgment result, and the feedback optimization judgment result. Specifically, the threshold determination module receives the virtual-real alignment loss signal, the credibility verification loss signal, and the virtual-real consistency evolution potential output by the loss calculation module, and generates the virtual-real alignment judgment result, the credibility verification judgment result, the data output judgment result, and the feedback optimization judgment result. The virtual-real alignment loss signal is obtained from the distribution difference between the credible real data and the aligned simulation data. The credibility verification loss signal is obtained from the distribution difference between the credible real data and the aligned simulation data, the physical constraint residual, and the trend deviation. The virtual-real consistency evolution potential is obtained by uniformly quantifying the virtual-real difference characterization, the physical deviation characterization, and the trend deviation characterization. By simultaneously introducing the loss signal and the virtual-real consistency evolution potential, the threshold determination module can simultaneously judge the degree of distribution alignment, the degree of physical constraint satisfaction, and the degree of temporal trend consistency. The virtual-real alignment judgment result is generated by a joint judgment method combining threshold comparison and evolution potential level verification. Threshold comparison is used to determine whether the virtual-real alignment loss signal meets the virtual-real distribution alignment requirements, and evolution potential level verification is used to determine whether the virtual-real consistency evolution potential is within the allowable deviation level. The default virtual-real alignment threshold is set to 0.08. During deployment, the optimal value can be searched within the range of 0 to 0.2 using 5-fold cross-validation with a step size of 0.01. The default virtual-real consistency evolution potential threshold is set to 0.10. During deployment, it can be determined by adding one standard deviation to the mean of the virtual-real consistency evolution potential in the training set, or by 5-fold cross-validation. The virtual-real consistency evolution potential levels are categorized by default as low deviation, medium deviation, and high deviation, with 0 to 0.1 representing low deviation, 0.1 to 0.3 representing medium deviation, and greater than 0.3 representing high deviation. When the virtual-real alignment loss signal is lower than the virtual-real alignment threshold and the virtual-real consistency evolution potential is at a low deviation level, a virtual-real alignment completion result is generated; when the virtual-real alignment loss signal is not lower than the virtual-real alignment threshold, or the virtual-real consistency evolution potential is not at a low deviation level, a virtual-real alignment failure result is generated. The trusted verification result is generated jointly by the trusted verification loss signal, the trusted verification threshold, and the virtual-real consistency evolution potential. The default trusted verification threshold is set to 0.12. During deployment, the optimal value can be searched within the range of 0 to 0.2 using 5-fold cross-validation with a step size of 0.01. A trusted verification pass result is generated when the trusted verification loss signal is below the trusted verification threshold and the virtual-real consistency evolution potential is at a low deviation level; a trusted verification fail result is generated when the trusted verification loss signal is not below the trusted verification threshold, or the virtual-real consistency evolution potential is not at a low deviation level. A trusted verification pass result indicates that the trusted real data and the aligned simulation data meet the output requirements in terms of distribution consistency, physical rationality, and trend consistency. The data output judgment result is generated based on the trusted verification judgment result. When trusted verification passes, the threshold judgment module determines the trusted real data as the enhanced trusted data output.When the trust verification fails and the number of feedback optimization attempts has not reached the upper limit, the threshold judgment module generates a feedback optimization judgment result and re-enters the virtual-real fusion alignment closed loop. The upper limit for the number of feedback optimization attempts is set to 10 by default, and can be set to 5 to 15 times depending on real-time requirements during engineering deployment. When the trust verification fails and the number of feedback optimization attempts reaches the upper limit, the threshold judgment module outputs fused data generated by weighting trusted real data and aligned simulation data. The fusion coefficient is set to 0.6 by default, indicating that the proportion of trusted real data is 0.6 and the proportion of aligned simulation data is 0.4; the fusion coefficient can also be set to 0.4 to 0.8 with a grid search in 0.1 steps, and the selection criterion is based on minimizing passenger flow prediction error, anomaly identification accuracy, or operational status assessment error. The feedback optimization judgment result is generated based on the proportional relationship between virtual-real difference representation, physical deviation representation, and trend deviation representation in the virtual-real consistency evolution potential. When the proportion of virtual-real difference representation is the highest, it indicates that the distribution difference between the credible real data and the aligned simulation data is the main source of deviation. The feedback optimization object is primarily determined to be the second diffusion model, and the feedback optimization content includes adjusting the participation strength of the real-side distribution constraint features, the weight of the distribution alignment constraint, and the weight of the feature consistency constraint. When the proportion of physical deviation representation is the highest, it indicates that the virtual data or aligned simulation data does not adequately meet the constraints of rail transit operation. The feedback optimization object is primarily determined to be the physical simulator module, and the feedback optimization content includes adjusting the passenger flow loading coefficient, train speed, departure interval, station dwell time, transfer ratio, station entry and exit service capacity, and vehicle capacity. When the proportion of trend deviation representation is the highest, it indicates that there is a deviation between the credible real data and the aligned simulation data in the trend of peak, trough, or sudden passenger flow changes. The feedback optimization object is determined to be one or more of the first diffusion model, the second diffusion model, and the physical simulator module, and the feedback optimization content includes strengthening the extraction of spatiotemporal correlation features, enhancing trend consistency constraints, and adjusting the passenger flow loading parameters during peak or sudden periods. The threshold determination module outputs the virtual-real alignment determination result, the trust verification determination result, the data output determination result, and the feedback optimization determination result to the virtual-real fusion digital twin module. The virtual-real fusion digital twin module then visualizes the determination results and transmits feedback signals. Through the above determination result generation process, the framework forms a closed loop of "generation—alignment—verification—filtering—feedback," enabling the simultaneous completion of rail transit data enhancement and trust verification.
[0112] The present invention is further configured such that the virtual-real fusion digital twin module includes a virtual-real mapping unit, an object binding unit, a topology matching unit, a judgment and display unit, and a feedback transmission unit, wherein:
[0113] The virtual-real mapping unit maps real rail transit data, virtual data, credible real data, aligned simulation data and judgment results to the same time base based on a preset unified spatiotemporal index.
[0114] The object binding unit binds real rail transit data, virtual data, trusted real data, and aligned simulation data corresponding to the same station, line, train, or transfer node based on the rail transit object identifier.
[0115] The topology matching unit maps the data that has been bound to the object to the corresponding rail transit line topology location based on the preset line topology relationship, and establishes a virtual-real mapping relationship.
[0116] The judgment display unit is used to display the judgment results and the virtual-real consistency evolution potential;
[0117] The feedback transmission unit transmits feedback signals to the first diffusion model, the second diffusion model, and the physical simulator module based on the judgment result. Specifically, the virtual-real fusion digital twin module is used to receive real data, virtual data, credible real data, aligned simulation data, judgment results, and virtual-real consistent evolution potential of the rail transit system, and establishes a virtual-real mapping relationship under a unified time base, a unified object base, and a unified line topology base. The virtual-real fusion digital twin module can be built based on Unity3D, Unreal Engine, or WebGL 3D visualization platforms, preferably based on the Unity3D platform to build the rail transit digital twin scene. Data communication can be implemented using the MQTT protocol, WebSocket protocol, or message queue method, preferably using the MQTT protocol, with the data transmission latency controlled within 100ms by default. The virtual-real mapping unit maps the real data, virtual data, credible real data, aligned simulation data, and judgment results of the rail transit system to the same time base based on a preset unified spatiotemporal index. The unified spatiotemporal index is generated by the sampling time, time window number, rail transit running segment marker, and rail transit object identifier. The time window number originates from the real data collection and processing phase of rail transit. The default time window length is 15 minutes, but it can be set to 5 to 10 minutes for real-time monitoring scenarios. Rail transit operating period markers include peak period markers, off-peak period markers, normal period markers, and emergency event period markers. Peak periods by default include 7:00 to 9:00 and 17:00 to 19:00, while off-peak periods by default include 10:00 to 16:00 and 20:00 to 23:00. Emergency event periods are determined based on abnormal event information in the operation management system or detection results showing a passenger flow change rate exceeding 30% in adjacent time windows. Through a unified spatiotemporal index, real-side data, virtual-side data, model output data, and judgment results within the same time window can enter the same digital twin refresh cycle. The object binding unit, based on rail transit object identifiers, binds real rail transit data, virtual data, reliable real data, and aligned simulation data corresponding to the same station, line, train, or transfer node. The identification of rail transit objects originates from the actual data collection and processing phase of rail transit. Station-level data uses a combination of station number and line number to generate the identification; train-level data uses a combination of train number and line number; and transfer-level data uses a combination of transfer node number and line number. After object binding is completed, a one-to-one correspondence is established between the data of the same rail transit object on the real side, virtual side, and model side, for subsequent difference display, judgment display, and feedback positioning. The topology matching unit, based on preset line topology relationships, maps the data of the completed object binding to the corresponding rail transit line topology location, establishing a virtual-real mapping relationship.The route topology can be represented using a node-edge structure, with stations, trains, and transfer nodes as topology nodes, and sections, train routes, and transfer routes as topology edges. Data sources for the route topology include route numbers, station sequence numbers, transfer node numbers, train operating sections, and route topology location numbers. Through topology matching, station passenger flow, train passenger capacity, transfer passenger flow, and cross-sectional flow can be mapped to corresponding station nodes, train nodes, transfer nodes, or section edges, forming a spatial topological correspondence between the real and virtual sides of the rail transit system. The judgment and display unit is used to display the judgment results and the virtual-real consistency evolution potential. Judgment results include virtual-real alignment judgment results, credibility verification judgment results, data output judgment results, and feedback optimization judgment results. The virtual-real consistency evolution potential includes a comprehensive quantitative result of virtual-real difference representation, physical deviation representation, and trend deviation representation. The judgment and display unit can use graphs, heatmaps, tables, color markers, and 3D scene markers to display the changing trends, deviation degrees, and credibility levels of credible real data and aligned simulation data. The virtual-real consistent evolution potential is defaulted to a low deviation level of 0 to 0.1, a medium deviation level of 0.1 to 0.3, and a high deviation level greater than 0.3. Low deviation levels are displayed as normal, medium deviation levels as a warning, and high deviation levels as an alarm. The feedback transmission unit transmits feedback signals to the first diffusion model, the second diffusion model, and the physical simulator module based on the judgment results. The generation of feedback signals is based on the feedback optimization judgment results, the proportion of virtual-real difference representation, physical deviation representation, and trend deviation representation in the virtual-real consistent evolution potential, as well as the corresponding time window, rail transit object identifier, and line topology location. When the proportion of virtual-real difference representation is the highest, the feedback transmission unit transmits instructions to the second diffusion model to adjust the distribution constraint features, distribution alignment constraint weights, and feature consistency constraint weights on the real side. When the proportion of physical deviation representation is the highest, the feedback transmission unit transmits instructions to the physical simulator module to adjust passenger flow loading coefficients, train speeds, departure intervals, station dwell times, transfer ratios, station entry / exit service capacity, and vehicle capacity. When the proportion of trend deviation representation is the highest, the feedback transmission unit transmits instructions to one or more of the first diffusion model, second diffusion model, and physical simulator module to strengthen spatiotemporal correlation features, adjust trend consistency constraints, and adjust passenger flow loading parameters during peak or sudden periods. Through the above settings, the virtual-real fusion digital twin module solves the time correspondence problem based on a unified spatiotemporal index, the object correspondence problem based on rail transit object identification, and the spatial correspondence problem based on line topology relationships. Furthermore, the deviation state is transformed into a model optimization direction through the judgment display unit and the feedback transmission unit.The virtual-real fusion digital twin module is not just a simple display module, but a hub that links the dual diffusion model module, the physical simulator module, the loss calculation module, and the threshold determination module, supporting a closed-loop process of "virtual-real mapping - determination and display - feedback optimization - regeneration".
[0118] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A virtual-real fusion digital twin enhancement framework, characterized in that, include: The module includes a dual-diffusion model module, a physics simulator module, a loss calculation module, a threshold determination module, and a virtual-real fusion digital twin module, among which: The dual diffusion model module includes a first diffusion model and a second diffusion model. The collected real rail transit data is input into the first diffusion model for data augmentation to generate reliable real data, and the virtual data generated by the physical simulator module is input into the second diffusion model for distribution alignment enhancement to obtain aligned simulation data. The physics simulator module generates virtual data based on real rail transit data and preset rail transit operation constraints, and outputs the corresponding physical constraint residuals. The loss calculation module extracts the trend deviation based on the temporal change characteristics of credible real data and aligned simulation data within the same time window, generates a loss signal using a differentiated loss calculation method, and performs unified quantization based on the difference information between credible real data and aligned simulation data, physical constraint residuals and trend deviations to generate a consistent virtual-real evolution potential. The threshold determination module performs virtual-real alignment determination, credibility verification determination, data filtering determination, and feedback optimization determination based on the loss signal and the virtual-real consistency evolution potential to generate a determination result. The virtual-real fusion digital twin module is based on real rail transit data, virtual data, credible real data, aligned simulation data and judgment results. It establishes a virtual-real mapping relationship through unified spatiotemporal index, rail transit object identification binding and line topology matching, and transmits feedback signals to the dual diffusion model module and physical simulator module according to the judgment results.
2. The virtual-real fusion digital twin enhancement framework according to claim 1, characterized in that, The collection and processing of real rail transit data includes: Raw data from rail transit is collected through multi-source data access methods, including data acquisition from automatic fare collection systems, train operation monitoring systems, station passenger flow detection equipment, transfer channel detection equipment, line dispatching systems, and operation management systems. The raw rail transit data is processed by time alignment, station alignment, missing data marking, anomaly marking, normalization, and spatiotemporal coding to obtain standardized real rail transit data. The real rail transit data includes: rail transit object identifiers, passenger flow data, station data, line data, train operation data, transfer data, and operation status data.
3. The virtual-real fusion digital twin enhancement framework according to claim 1, characterized in that, The data augmentation process for the first diffusion model includes: Data missing detection is performed on real rail transit data to determine the missing locations, and a missing location mask is generated based on the missing locations; The first diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The attention mechanism is used to extract the spatiotemporal correlation features of real rail transit data between adjacent time windows, associated stations and associated lines. Using real rail transit data and missing location masks as inputs to the first diffusion model, the distribution of real data is learned through forward noise addition and reverse denoising processes. In the reverse denoising process, missing regions are gradually filled in by combining missing location masks and spatiotemporal correlation features, and noise suppression and distribution calibration are performed on non-missing regions to generate reliable real data.
4. The virtual-real fusion digital twin enhancement framework according to claim 1, characterized in that, The distribution alignment enhancement process of the second diffusion model includes: A second diffusion model is constructed based on the U-Net backbone network with embedded attention mechanism. The second diffusion model adopts the same network structure as the first diffusion model. The virtual data generated by the physical simulator module is used as the input of the second diffusion model, and the kernel density is estimated based on the real rail transit data to obtain the real side distribution density characteristics. The real side distribution density characteristics are used as the real side distribution constraint characteristics. The distribution mapping relationship between virtual data and real-side distribution constraint features is learned through forward noise addition and reverse noise reduction processes. In the reverse noise reduction process, the distribution constraint features of the real side and the physical constraint residuals are combined to perform distribution alignment of virtual data. Based on the distributed and aligned virtual data, and in accordance with the preset distribution alignment constraints and feature consistency constraints, aligned simulation data is generated that retains the physical rationality of the virtual data while being consistent with the distribution of real rail transit data.
5. The virtual-real fusion digital twin enhancement framework according to claim 1, characterized in that, The virtual data generation process of the physics simulator module includes: Virtual simulation is performed based on real rail transit data, preset rail transit operation constraints, and adjustable operating parameters to generate virtual data. The constraints on rail transit operation include: passenger flow conservation constraints, station capacity constraints, vehicle capacity constraints, train timetable constraints, transfer route constraints, and cross-sectional flow constraints. The adjustable operating parameters include: passenger flow loading coefficient, train speed, departure interval, station dwell time, transfer ratio, station entry and exit service capacity, and vehicle capacity. Based on the degree of deviation of the virtual data from the constraints of passenger flow conservation, station capacity, vehicle capacity, train timetable, transfer route, and cross-sectional flow, physical constraint residuals are generated.
6. The virtual-real fusion digital twin enhancement framework according to claim 1, characterized in that, The trend deviation extraction includes: Based on credible real data and aligned simulation data, time-series segments are divided into credible real data and aligned simulation data according to the same time window. Match the corresponding reliable real data and aligned simulation data within the same time window; By using sliding time window statistics, peak and valley period marking, and abrupt state identification, the temporal change features of the matched reliable real data and the aligned simulation data are extracted respectively. The temporal change features include: peak period passenger flow change trend, valley period passenger flow change trend, and sudden passenger flow change trend. A trend bias is generated based on the deviation between the temporal variation characteristics of reliable real data and aligned simulation data.
7. The virtual-real fusion digital twin enhancement framework according to claim 6, characterized in that, The loss signal calculation includes: For the data augmentation process of the first diffusion model, a first loss signal is generated based on the data reconstruction deviation between credible real data and real rail transit data, as well as the distribution smoothing deviation of credible real data relative to real rail transit data. For the distribution alignment enhancement process of the second diffusion model, a second loss signal is generated based on the distribution alignment deviation of the aligned simulation data relative to the real rail transit data, and the consistency deviation between the aligned simulation data and the real rail transit data in the feature space. For the virtual-real alignment determination process, a virtual-real alignment loss signal is generated based on the distribution difference between reliable real data and aligned simulation data; For the trustworthy verification determination process, a trustworthy verification loss signal is generated based on the distribution difference, physical constraint residual, and trend deviation between trustworthy real data and aligned simulation data. The first loss signal, the second loss signal, the virtual-real alignment loss signal, and the trusted verification loss signal are output as loss signals.
8. The virtual-real fusion digital twin enhancement framework according to claim 7, characterized in that, The generation of the virtual-real consistent evolution potential includes: The differences between reliable real data and aligned simulation data, physical constraint residuals and trend deviations are normalized to obtain standardized virtual-real difference representations, physical deviation representations and trend deviation representations. The contribution relationship of virtual-real difference representation, physical deviation representation and trend deviation representation is determined by a weighting method based on the loss signal. The weighting method includes determining the proportion of virtual-real difference representation, physical deviation representation and trend deviation representation in virtual-real consistency evaluation based on the virtual-real alignment loss signal, the credible verification loss signal, the deviation state of physical constraint residuals and the deviation state of trend deviation. Based on the contribution relationship, the virtual-real difference representation, physical deviation representation, and trend deviation representation are uniformly quantified to generate a virtual-real consistent evolution potential.
9. The virtual-real fusion digital twin enhancement framework according to claim 8, characterized in that, The generation of the determination result includes: The virtual-real alignment loss signal is compared with the preset virtual-real alignment threshold. Combined with the virtual-real consistency evolution potential, a joint judgment method combining threshold comparison and evolution potential level verification is used to generate the virtual-real alignment judgment result. The trusted verification loss signal is compared with the preset trusted verification threshold. Combined with the virtual-real consistent evolution potential, a trusted verification judgment result is generated by a joint judgment method that combines threshold comparison and evolution potential level verification. Based on the credibility verification judgment results, data filtering and judgment are performed on credible real data, aligned simulation data, and fused data generated from credible real data and aligned simulation data, and data output judgment results are generated. When the virtual-real alignment judgment result or the credibility verification judgment result does not meet the preset judgment conditions, the feedback optimization object is determined based on the proportion relationship of virtual-real difference representation, physical deviation representation and trend deviation representation in the virtual-real consistency evolution potential, and the feedback optimization judgment result is generated. The results of the virtual-real alignment judgment, the trusted verification judgment, the data output judgment, and the feedback optimization judgment are output as the judgment results.
10. The virtual-real fusion digital twin enhancement framework according to claim 2, characterized in that, The virtual-real fusion digital twin module includes a virtual-real mapping unit, an object binding unit, a topology matching unit, a judgment and display unit, and a feedback transmission unit, among which: The virtual-real mapping unit maps real rail transit data, virtual data, credible real data, aligned simulation data and judgment results to the same time base based on a preset unified spatiotemporal index. The object binding unit binds real rail transit data, virtual data, trusted real data, and aligned simulation data corresponding to the same station, line, train, or transfer node based on the rail transit object identifier. Based on a preset line topology relationship, the topology matching unit maps the data that has been bound to the object to the corresponding rail transit line topology location, establishing a virtual-real mapping relationship; The judgment display unit is used to display the judgment results and the virtual-real consistency evolution potential; Based on the judgment result, the feedback transmission unit transmits feedback signals to the first diffusion model, the second diffusion model, and the physics simulator module.