Digital twin entity construction method and system based on spatiotemporal multi-dimensional data fusion

By constructing a weighted undirected graph and screening key stations, the problem of insufficient integration of spatiotemporal multidimensional factors in digital twin entities in urban rail transit was solved. This enabled efficient monitoring of passenger flow operation status and accurate synchronization of the virtual model, improving the effectiveness of operation and maintenance management and scheduling decisions.

CN122634880APending Publication Date: 2026-08-25LONGTEK (SUZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610753767.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing digital twin entities in urban rail transit have failed to fully integrate spatiotemporal multidimensional factors, resulting in insufficient reliability in monitoring passenger flow operation status.

Method used

By acquiring passenger flow data from each station in the urban rail transit network, a weighted undirected graph is constructed to obtain the importance coefficient of the comprehensive hub, key stations are screened and the monitoring frequency is increased. High-frequency monitoring data is then synchronized to the digital twin entity in real time to ensure the accuracy of the virtual-real synchronization of key stations.

Benefits of technology

It improves the reliability of digital twin entities in monitoring passenger flow operation status, enhances the accuracy of state synchronization between virtual models and physical entities at key stations, and optimizes operation and maintenance management and scheduling decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122634880A_ABST
    Figure CN122634880A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of digital twin, in particular to a digital twin entity construction method and system based on space-time multi-dimensional data fusion, which comprises the following steps: obtaining passenger flow data of each station in a city rail transit network, and then generating passenger flow of each station at a preset time granularity; obtaining passenger flow characteristic values of each station for any day; classifying all stations through the spatial positions of all stations and the passenger flow characteristic values of historical days before the any day; constructing a weighted undirected graph representing the spatial relationship of the city rail transit network; obtaining a comprehensive hub importance coefficient of each node in the city rail transit network; selecting key stations from all stations, and increasing the passenger flow data monitoring frequency of the key stations to ensure the virtual-real synchronization accuracy of the key stations. The application aims to improve the reliability of the digital twin entity in passenger flow operation state monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital twin technology, specifically to a method and system for constructing digital twin entities based on spatiotemporal multidimensional data fusion. Background Technology

[0002] Digital twin technology is an advanced digital technology that integrates digital twins and artificial intelligence. By constructing a high-fidelity dynamic mirror model of a physical entity in virtual space, and leveraging big data, the Internet of Things, and advanced algorithms, it achieves real-time mapping, intelligent analysis, and predictive optimization of the entity's entire lifecycle. In the field of urban rail transit, digital twin technology demonstrates enormous potential in operation and maintenance management, safety monitoring, scheduling decisions, and new line development and planning through real-time mapping of facilities, equipment, and operational status.

[0003] Rail transit operation is essentially a dynamic process with strong spatiotemporal coupling, exhibiting significant spatiotemporal complexity. This includes factors such as peak-hour passenger flow concentration, regular fluctuations between weekdays and weekends, and the propagation and diffusion effects of passenger flow between stations. Therefore, accurate analysis of passenger flow status is crucial. However, in practical applications, station passenger flow is influenced by the complex spatial network relationships of rail transit. Conventional digital twin entities fail to fully integrate the multidimensional spatiotemporal factors within the transportation system, resulting in insufficient reliability of digital twin entities in monitoring passenger flow operation status. Summary of the Invention

[0004] In light of the above, it is necessary to provide a method and system for constructing digital twin entities based on spatiotemporal multidimensional data fusion. Compared with traditional digital twin entity construction methods, this improves the reliability of digital twin entities in monitoring passenger flow operation status. In a first aspect, embodiments of this application provide a method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion, the method comprising the following steps: Obtain passenger flow data from each station in the urban rail transit network, and then generate passenger flow data for each station at a preset time granularity; For any given day, passenger flow characteristics are obtained for each station based on peak passenger flow concentration, frequency of passenger flow changes, and regular fluctuations in passenger flow. All stations are then categorized based on their spatial location and historical passenger flow characteristics from previous days, ensuring that stations within the same category are spatially adjacent and exhibit similar passenger flow patterns. Each category is treated as a node, and edges are constructed using the rail line connections between stations within each node. Edge weights are also constructed using the travel records between stations within each node, resulting in a weighted undirected graph representing the spatial relationships of the urban rail transit network. Finally, the comprehensive hub importance coefficient of each node in the urban rail transit network is obtained using the weights of the edges connected to each node and the shortest path hop count from each node to other nodes. By analyzing the distribution of the comprehensive hub importance coefficients of all nodes, key stations are selected from all stations. The frequency of passenger flow data monitoring for key stations is increased within a preset time interval after any given day. The high-frequency monitoring data is then synchronized in real time to the digital twin entity based on rail transit to ensure the accuracy of virtual-real synchronization of key stations.

[0005] In one embodiment, the process of obtaining the passenger flow characteristic value is as follows: By analyzing the fluctuation patterns of passenger flow at individual stations, we can obtain the degree of fluctuation in passenger flow patterns at each station. Obtain the peaks of passenger flow at a single station in any given day in time sequence, and sort all peaks in descending order. Calculate the sum of passenger flow within the peak width range of the peaks containing a preset number of peaks, and calculate the percentage of the sum in the total passenger flow of a single station in any given day. Calculate the dispersion of passenger flow at a single station within any given day; The passenger flow characteristic value of a single station is positively correlated with the numerical proportion and the passenger flow pattern fluctuation, and negatively correlated with the dispersion.

[0006] In one embodiment, the process of obtaining the passenger flow pattern fluctuation is as follows: If any day is a workday, a predetermined number of neighboring workdays preceding that day are used as reference days for that day. If any day is a rest day, a predetermined number of neighboring rest days preceding that day are used as reference days for that day. Calculate the temporal correlation coefficient of passenger flow at a single station between that day and its reference days. The fluctuation of passenger flow patterns at a single station is obtained by the correlation coefficient of that single station between any given day and all its control days.

[0007] In one embodiment, the calculation process for the passenger flow characteristic value is as follows: The fluctuation of passenger flow patterns at a single station is shifted to a non-negative number; the product of the numerical percentage and the non-negative number is recorded as the comprehensive product. The passenger flow characteristic value of a single station is directly proportional to the comprehensive product and inversely proportional to the dispersion.

[0008] In one embodiment, the process of classifying all stations is as follows: Calculate the arithmetic mean of the passenger flow characteristic values ​​of each station within a preset time range prior to the stated day; All stations are classified by comparing their arithmetic mean and geospatial coordinates.

[0009] In one embodiment, the construction process of the weighted undirected graph is as follows: If any two nodes contain stations that are adjacent on the track, connect the two nodes. For a single passage record, if any node contains a station as the entry station and any other node contains a station as the exit station, increment the weight of each edge along the shortest path between any node and any other node by 1. Obtain the edge weights from the passage records within a preset time period prior to the given day.

[0010] In one embodiment, the process of obtaining the comprehensive hub importance coefficient is as follows: Normalize the weights of all edges and calculate the mean of the normalized weights of all edges connected to each node. Count the shortest path hop count from each node to every other node, and sum the shortest path hop counts from each node to all other nodes. The importance coefficient of the integrated hub is directly proportional to the mean and inversely proportional to the summation result.

[0011] In one embodiment, the overall hub importance coefficient is the product of the reciprocal of the summation result and the mean.

[0012] In one embodiment, the selection process for the key station is as follows: A threshold for the segmentation of the comprehensive hub importance coefficient of all nodes is preset. Nodes whose comprehensive hub importance coefficient is greater than the segmentation threshold are designated as key nodes, and the stations contained in the key nodes are designated as key stations.

[0013] Secondly, embodiments of this application also provide a digital twin entity construction system based on spatiotemporal multidimensional data fusion, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any of the above-described digital twin entity construction methods based on spatiotemporal multidimensional data fusion.

[0014] This application has at least the following beneficial effects: This application comprehensively quantifies station passenger flow patterns from multiple perspectives, including peak passenger flow concentration, frequent passenger flow changes, and regular passenger flow fluctuations, based on the temporal distribution, volatility, and periodic stability of passenger flow. This helps identify differentiated passenger flow patterns at different stations due to variations in surrounding land use. Furthermore, by integrating the spatial location information and passenger flow characteristics of all stations for joint classification, stations within the same category are both geographically adjacent and highly similar in passenger flow patterns. The classification results provide a reasonable basis for node division in subsequent construction of a graphical model representing the spatial relationships of the urban rail transit network, reducing the complexity of the graphical model while retaining key spatial interaction information. By constructing a graphical model using track line connections and travel records, the application can comprehensively represent the spatial proximity, line connection relationships, and passenger flow interaction intensity of the urban rail transit network. Furthermore, by integrating passenger flow interaction characteristics and traffic location centrality, the hub status of each node in the urban rail transit network can be evaluated. This allows for the simultaneous capture of the node's busyness at the passenger flow interaction level and its core location characteristics at the spatial network level, improving the accuracy of key node identification. Consequently, key stations can be screened, and the monitoring frequency of these stations can be increased. Under limited resource constraints, high-frequency monitoring can be focused on stations that have a greater impact on the overall operation of the urban rail transit network, improving the utilization efficiency of monitoring resources. Real-time synchronization of high-frequency monitoring data to the digital twin entity can effectively reduce the state deviation between the virtual model and the physical entity of key stations, ensuring the accuracy of virtual-physical synchronization of key stations and significantly improving the reliability of the digital twin entity in passenger flow operation status monitoring. Attached Figure Description

[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 A flowchart illustrating the steps of a method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion, as provided in one embodiment of this application; Figure 2 A schematic diagram illustrating the process of obtaining the comprehensive hub importance coefficient; Figure 3 This is a schematic diagram illustrating the calculation process for the comprehensive hub importance coefficient. Detailed Implementation

[0017] In the description of the embodiments in this application, the words "exemplary," "or," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary," "or," and "for example" is intended to present the relevant concepts in a specific manner.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. It should be understood that, unless otherwise stated, " / " in this application means "or".

[0019] It should also be noted that the terms "first" and "second" in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0020] The following description, in conjunction with the accompanying drawings, details the specific scheme of the digital twin entity construction method and system based on spatiotemporal multidimensional data fusion provided in this application.

[0021] Please see Figure 1 The diagram illustrates a flowchart of a method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion according to an embodiment of this application. The method includes the following steps: Step 1: Obtain passenger flow data for each station in the urban rail transit network, and then generate passenger flow data for each station at a preset time granularity.

[0022] This application takes urban rail transit as an example. In urban rail transit applications, operation and maintenance management, safety monitoring, dispatching decisions, and new line development and planning are crucial aspects. This application employs digital twin technology to monitor the operational status of rail transit. When constructing the digital twin entity based on rail transit, it uses passenger flow data from each station in the urban rail transit network as a foundation. This passenger flow data includes passenger entry and exit times, as well as station location information. Based on this data, the passenger flow status of each station is analyzed to generate passenger flow data for each station at a preset time granularity.

[0023] In this embodiment, the time granularity is in minutes, that is, the passenger flow of each station is generated in minutes. Implementers can set the time granularity according to the actual passenger flow. This application does not impose any special restrictions.

[0024] In this embodiment, before generating passenger flow data, the collected passenger flow data is cleaned, including: removing passenger flow data that enters and exits the same station and deleting passenger flow data with incorrect dates.

[0025] Step 2: For any given day, obtain the passenger flow characteristic values ​​of each station; classify all stations based on their spatial location and the passenger flow characteristic values ​​of the previous several days; construct a weighted undirected graph representing the spatial relationships of the urban rail transit network; and obtain the comprehensive hub importance coefficient of each node in the urban rail transit network.

[0026] The factors influencing urban rail transit passenger flow are numerous and complex, among which the built environment is particularly important, directly determining the scale and direction of rail transit travel demand. For example, building type, holiday effects, and urban layout adjustments can all cause shifts in passenger flow patterns. These factors are intertwined, making it necessary to simultaneously address both regular patterns and uncertainties in accurately predicting passenger flow.

[0027] Step 2.1: For any given day, obtain the passenger flow characteristic values ​​of each station by analyzing the peak passenger flow concentration, the frequency of passenger flow changes, and the regular fluctuations in passenger flow.

[0028] The different travel purposes, time requirements, and concentration levels of people within different types of building areas lead to differentiated passenger flow characteristics at stations in urban rail transit networks, exhibiting significant differences in both time and quantity dimensions. For example, stations near residential and office areas typically show clear, regular fluctuations in passenger flow: during weekdays, passenger flow is highly concentrated during morning and evening peak hours, with large overall volume and a fast pace, mainly driven by rigid travel demands such as commuting and going to school; passenger flow is relatively stable at other times. On weekends, passenger flow at stations near residential and office areas is generally flat, with weaker peak characteristics. Stations near tourist attractions and transportation hubs, however, are affected by factors such as train schedules, flights, and buses, resulting in fluctuating passenger flow. While exhibiting some peak-tidal characteristics, passenger flow fluctuates multiple times throughout the day, with more frequent and unpredictable changes.

[0029] Based on the above analysis, for any given day, passenger flow characteristic values ​​for each station are obtained by analyzing peak passenger flow concentration, frequency of passenger flow changes, and regular fluctuations in passenger flow. The specific process is as follows: By analyzing the fluctuation patterns of passenger flow at individual stations, we can obtain the degree of fluctuation in passenger flow patterns at each station. Obtain the peaks of passenger flow at a single station in any given day in time sequence, and sort all peaks in descending order. Calculate the sum of passenger flow within the peak width range of the peaks containing a preset number of peaks. Use the percentage of the sum in the total passenger flow of a single station in any given day as the peak passenger flow concentration of a single station. The dispersion of passenger flow at a single station on any given day is taken as the frequency of passenger flow fluctuation at a single station. The fluctuation of passenger flow patterns at a single station is shifted to a non-negative number; the product of the peak passenger flow concentration and the non-negative number is denoted as the comprehensive product. The passenger flow characteristic value of a single station is directly proportional to the comprehensive product and inversely proportional to the frequency fluctuation of passenger flow.

[0030] The process for obtaining the passenger flow pattern fluctuation of a single station is as follows: Considering the difference in passenger flow between weekdays and rest days, if any day is a weekday, a preset number of neighboring weekdays before that day are used as reference days for that day; if any day is a rest day, a preset number of neighboring rest days before that day are used as reference days for that day; calculate the temporal correlation coefficient of passenger flow at a single station between that day and its reference days. The fluctuation of passenger flow patterns at a single station is obtained by the correlation coefficient of that single station between any given day and all its control days.

[0031] In this embodiment, the least squares method is used to perform nonlinear fitting on the time series of passenger flow of a single station on any given day, and the maximum value in the fitted curve is taken as the peak. The least squares method is a well-known technique and will not be described in detail in this application. As other implementation methods, based on the ability to obtain the peak of passenger flow of a single station on any given day in the time series, the implementer may use other existing feasible techniques, such as peak and trough detection algorithms, etc. This application does not impose any special restrictions.

[0032] In this embodiment, the method for obtaining the peak width range is as follows: extend the time axis from the peak to both sides until the value drops to the preset proportion of the peak value or encounters the cutoff interval of the adjacent trough. The condition reached first on the time axis shall prevail. The preset proportion is 20%. The preset proportion is preset by human intervention. The implementer can set it according to the actual situation. This application does not impose any special restrictions.

[0033] In this embodiment, since there are morning and evening peak hours during weekdays, the preset quantity is set to 2.

[0034] In this embodiment, the dispersion is specifically the coefficient of variation. The calculation of the coefficient of variation is a well-known technique and will not be described in detail here. As other implementation methods, based on the ability to measure the unevenness of passenger flow distribution, implementers may adopt other existing feasible techniques, such as variance, standard deviation, etc. This application does not impose any special restrictions.

[0035] In this embodiment, the preset number is 25. The preset number is preset by the user and the implementer can set it according to the actual situation. This application does not impose any special restrictions.

[0036] In this embodiment, the process of calculating the temporal correlation coefficient of passenger flow at a single station between any given day and each of the control days is as follows: the passenger flow of a single station within each day is arranged in chronological order to obtain the passenger flow sequence of a single station within each day, and the Spearman correlation coefficient of the passenger flow sequence of a single station between any given day and each of the control days is calculated. The calculation of the Spearman correlation coefficient is a well-known technique and will not be described in detail in this application. As other implementation methods, based on the ability to measure the degree of correlation between passenger flow sequences, the implementer may adopt other existing feasible techniques, such as the Pearson correlation coefficient, and this application does not impose any special restrictions.

[0037] In this embodiment, the average of the correlation coefficients of a single station between any given day and all its control days is taken as the passenger flow pattern fluctuation of a single station.

[0038] In this embodiment, the expression for the passenger flow characteristic value of a single station is: In the formula, A represents the passenger flow characteristic value of a single station; R represents the peak passenger flow concentration of a single station; P represents the passenger flow pattern fluctuation of a single station; and V represents the normalized value of the passenger flow frequency fluctuation of a single station. Specifically, by calculating P+1, the passenger flow pattern fluctuation is shifted to a non-negative number; This is denoted as the composite product.

[0039] In this embodiment, the normalized value of passenger flow frequency fluctuation is obtained by the maximum value normalization method. When normalizing the passenger flow frequency fluctuation, the maximum value refers to the maximum value of the passenger flow frequency fluctuation of all stations in any given day. The maximum value normalization method is a well-known technology and will not be described in detail in this application.

[0040] It should be noted that: the greater the calculated peak passenger flow concentration, the more significant the peak passenger flow concentration characteristic of a single station; the greater the calculated passenger flow frequency fluctuation, the more frequent the changes in passenger flow at a single station on any given day; the greater the calculated passenger flow pattern fluctuation, the more significant the regular fluctuation characteristics of passenger flow at a single station; and the larger the calculated passenger flow characteristic value, the closer the passenger flow pattern of a single station is to a commuter-dominated characteristic, and the closer it is to residential or office areas. By calculating the passenger flow characteristic value, the passenger flow change characteristics of each station can be accurately quantified, which helps to identify and distinguish station types with different land use properties.

[0041] Step 2.2: Classify all stations by their spatial location and the passenger flow characteristics of the previous days, so that stations in the same category are spatially close and have similar passenger flow patterns.

[0042] Passenger flow at stations near residential and office areas exhibits significant peak concentrations, relatively smooth overall changes, and strong regular fluctuations, resulting in relatively high passenger flow characteristic values. Conversely, passenger flow at stations near tourist attractions and transportation hubs fluctuates more frequently, with less pronounced peak concentrations and regular fluctuations, leading to relatively low passenger flow characteristic values. Because different stations have different passenger flow patterns, the required rail transit operation management and planning schemes also differ. Therefore, all stations are classified based on their spatial location and historical passenger flow characteristic values ​​from previous days, ensuring that stations within the same category are spatially adjacent and have similar passenger flow patterns. The specific process is as follows: Calculate the arithmetic mean of the passenger flow characteristic values ​​of each station within a preset time range prior to the stated day; All stations are classified by comparing their arithmetic mean and geospatial coordinates.

[0043] In this embodiment, the preset time range is 30 days. The preset time range is preset by the user and the implementer can set it according to the actual situation. This application does not impose any special restrictions.

[0044] In this embodiment, the normalized value of the arithmetic mean of each station and the normalized geospatial coordinates are used to form a feature vector. The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is used to classify all stations. The distance metric used in the classification process is the Euclidean distance between the feature vectors. The minimum number of points in the neighborhood, MinPts, is set to 4. The neighborhood radius is automatically obtained by determining the slope inflection point using a K-distance curve. The arithmetic mean and geospatial coordinates are normalized using the Min-Max normalization method. The DBSCAN algorithm, K-distance curve, and Min-Max normalization method are all well-known technologies and will not be described in detail in this application. As other implementation methods, based on the ability to classify all stations according to the arithmetic mean and geospatial coordinates of the stations, the implementer may use other existing feasible technologies. This application does not impose any special restrictions.

[0045] Step 2.3: Treat each type as a node, construct edges through the track line connection relationship between the stations contained in each node, and construct edge weights through the passage records between the stations contained in each node, to obtain a weighted undirected graph representing the spatial relationship of the urban rail transit network.

[0046] Furthermore, station passenger flow is also affected by the complex spatial network relationships of rail transit, resulting in complex diffusion effects. If there are frequent passage records between stations in one category and stations in another category, it indicates that the commuting demand and passenger flow interaction intensity between these two stations is high. Combining the transportation network relationships between stations can help extract key stations connecting multiple different stations, which helps to discover key station hubs and potential passenger flow intersections, thereby providing support for the development and planning of new lines.

[0047] Based on the above analysis, each type is treated as a node. Edges are constructed using the rail line connections between stations contained in each node, and edge weights are constructed using the travel records between stations contained in each node. This results in a weighted undirected graph representing the spatial relationships of the urban rail transit network, specifically: If any two nodes contain stations that are adjacent on the track, connect the two nodes. For a single passage record, if the station contained in any one node is the entry station and the station contained in any other node is the exit station, increment the weight of each edge along the shortest path from any one node to any other node by 1. Obtain the edge weights from the passage records within a preset time period prior to the given day. Note that any one node and any other node are not the same node.

[0048] In this embodiment, the length of the preset time period is 1 week. The length of the preset time period is preset by the user and the implementer can set it according to the actual situation. This application does not impose any special restrictions.

[0049] Step 2.4: Obtain the comprehensive hub importance coefficient of each node in the urban rail transit network by using the weight of the edges connected to each node and the number of shortest path hops from each node to the other nodes.

[0050] By using the weights of the edges connected to each node and the number of hops of the shortest path from each node to the other nodes, the comprehensive hub importance coefficient of each node in the urban rail transit network is obtained. The specific process is as follows: Normalize the weights of all edges and calculate the mean of the normalized weights of all edges connected to each node. Even if the weight of the edges directly connected to a node is not high, a node that is closer to the other nodes may still play an important role in the connectivity efficiency and passenger accessibility of the transportation network. Therefore, the number of shortest path hops from each node to every other node is counted, and the reciprocal of the sum of the number of shortest path hops from each node to all other nodes is used as the importance of the connectivity efficiency of each node. The product of the connection efficiency importance and the mean value is used as the comprehensive hub importance coefficient of each node in the urban rail transit network. A schematic diagram of the process for obtaining the comprehensive hub importance coefficient is shown below. Figure 2As shown in the diagram. The calculation process for the comprehensive hub importance coefficient is illustrated below. Figure 3 As shown.

[0051] In this embodiment, the Min-Max normalization method is used to normalize the weights of all edges.

[0052] It should be noted that: the higher the calculated mean, the more crucial the position of the stations contained in each node within the entire urban rail transit network; the higher the calculated importance of connection efficiency, the higher the accessibility of each node to other nodes; and the higher the calculated comprehensive hub importance coefficient, the more significant the passenger flow interaction intensity and spatial accessibility characteristics of the stations contained in each node. By integrating the passenger flow interaction characteristics and traffic location centrality of nodes, key hubs that are busy and centrally located can be identified more comprehensively.

[0053] Step 3: Select key stations from all stations by analyzing the distribution of the comprehensive hub importance coefficients of all nodes, increase the frequency of passenger flow data monitoring of key stations within the preset time interval after any given day, and synchronize the acquired high-frequency monitoring data to the digital twin entity based on rail transit in real time to ensure the accuracy of virtual-real synchronization of key stations.

[0054] The construction process of the digital twin entity based on rail transit is as follows: First, passenger flow data of each station in the urban rail transit network is collected from an open-source dataset. Then, using 3D modeling and Geographic Information System (GIS) technology, a virtual digital model that accurately maps to the physical entity is constructed. During the construction of the virtual digital model, key stations are selected from all stations based on the distribution of the comprehensive hub importance coefficient of all nodes. The frequency of passenger flow data monitoring at key stations is increased within a preset time interval after any given day. The acquired high-frequency monitoring data is synchronized in real time to the digital twin entity based on rail transit to ensure the accuracy of virtual-physical synchronization of key stations, providing closed-loop optimized digital support for dispatching, operation and maintenance management, and passenger services. The construction process of the digital twin entity is a well-known technology and will not be described in detail in this application.

[0055] The selection process for key stations is as follows: a threshold for the division of the comprehensive hub importance coefficient of all nodes is preset; nodes whose comprehensive hub importance coefficient is greater than the threshold are designated as key nodes; and stations contained in key nodes are designated as key stations.

[0056] In this embodiment, the third quartile of the comprehensive hub importance coefficient of all nodes is used as the segmentation threshold, which is calculated from experimental data.

[0057] In this embodiment, the length of the preset time interval is 1 day. The length of the preset time interval is preset by the user and the implementer can set it according to the actual situation. This application does not impose any special restrictions.

[0058] It should be added that: In this application, when calculating the ratio, if there is a case where the denominator is 0, the denominator is first mapped to a positive number before subsequent calculations are performed. There are many ways to map data to a positive number, and implementers can choose existing feasible methods according to the actual situation. In this embodiment, the purpose of mapping the data to a positive number is achieved by calculating the sum of the data and a preset value greater than 0. The value of the preset value greater than 0 is preset by humans, and implementers can set it according to the actual situation. This application does not impose any special restrictions. In this embodiment, the value of the preset value greater than 0 is 0.01.

[0059] Based on the same inventive concept as the above methods, this application also provides a digital twin entity construction system based on spatiotemporal multidimensional data fusion, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described methods for constructing digital twin entities based on spatiotemporal multidimensional data fusion.

[0060] In summary, this application, by integrating three dimensions—peak passenger flow concentration, frequent passenger flow changes, and regular fluctuations in passenger flow—can comprehensively quantify station passenger flow patterns from multiple perspectives, including the temporal distribution, volatility, and periodic stability of passenger flow. This helps identify differentiated passenger flow patterns at different stations due to variations in surrounding land use. Furthermore, by fusing spatial location information and passenger flow characteristics of all stations for joint classification, stations within the same category are both geographically adjacent and highly similar in passenger flow patterns. The classification results provide a reasonable basis for node partitioning in subsequent construction of a graphical model representing the spatial relationships of the urban rail transit network, reducing the complexity of the graphical model while retaining key spatial interaction information. By constructing a graphical model using track line connections and travel records, it can comprehensively represent the spatial proximity, line connection relationships, and passenger flow interaction intensity of the urban rail transit network. Furthermore, by integrating passenger flow interaction characteristics and traffic location centrality, the hub status of each node in the urban rail transit network can be evaluated. This allows for the simultaneous capture of the node's busyness at the passenger flow interaction level and its core location characteristics at the spatial network level, improving the accuracy of key node identification. Consequently, key stations can be screened, and the monitoring frequency of these stations can be increased. Under limited resource constraints, high-frequency monitoring can be focused on stations that have a greater impact on the overall operation of the urban rail transit network, improving the utilization efficiency of monitoring resources. Real-time synchronization of high-frequency monitoring data to the digital twin entity can effectively reduce the state deviation between the virtual model and the physical entity of key stations, ensuring the accuracy of virtual-physical synchronization of key stations and significantly improving the reliability of the digital twin entity in passenger flow operation status monitoring.

[0061] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0062] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from its essential characteristics. Therefore, the embodiments described above should be considered exemplary and non-limiting in all respects.

Claims

1. A method for constructing digital twin entities based on spatiotemporal multidimensional data fusion, characterized in that, The method includes the following steps: Obtain passenger flow data from each station in the urban rail transit network, and then generate passenger flow data for each station at a preset time granularity; For any given day, passenger flow characteristics are obtained for each station based on peak passenger flow concentration, frequency of passenger flow changes, and regular fluctuations in passenger flow. All stations are then categorized based on their spatial location and historical passenger flow characteristics from previous days, ensuring that stations within the same category are spatially adjacent and exhibit similar passenger flow patterns. Each category is treated as a node, and edges are constructed using the rail line connections between stations within each node. Edge weights are also constructed using the travel records between stations within each node, resulting in a weighted undirected graph representing the spatial relationships of the urban rail transit network. Finally, the comprehensive hub importance coefficient of each node in the urban rail transit network is obtained using the weights of the edges connected to each node and the shortest path hop count from each node to other nodes. By analyzing the distribution of the comprehensive hub importance coefficients of all nodes, key stations are selected from all stations. The frequency of passenger flow data monitoring for key stations is increased within a preset time interval after any given day. The high-frequency monitoring data is then synchronized in real time to the digital twin entity based on rail transit to ensure the accuracy of virtual-real synchronization of key stations.

2. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The process for obtaining the passenger flow characteristic value is as follows: By analyzing the fluctuation patterns of passenger flow at individual stations, we can obtain the degree of fluctuation in passenger flow patterns at each station. Obtain the peaks of passenger flow at a single station in any given day in the time sequence, and sort all peaks in descending order. Calculate the sum of passenger flow within the peak width range of the peaks containing a preset number of peaks, and calculate the percentage of the sum in the total passenger flow of a single station in any given day. Calculate the dispersion of passenger flow at a single station within any given day; The passenger flow characteristic value of a single station is positively correlated with the numerical proportion and the passenger flow pattern fluctuation, and negatively correlated with the dispersion.

3. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 2, characterized in that, The process for obtaining the fluctuation of the passenger flow pattern is as follows: If any day is a workday, a predetermined number of neighboring workdays preceding that day are used as reference days for that day. If any day is a rest day, a predetermined number of neighboring rest days preceding that day are used as reference days for that day. Calculate the temporal correlation coefficient of passenger flow at a single station between that day and its reference days. The fluctuation of passenger flow patterns at a single station is obtained by the correlation coefficient of that single station between any given day and all its control days.

4. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 2, characterized in that, The calculation process for the passenger flow characteristic value is as follows: The fluctuation of passenger flow patterns at a single station is shifted to a non-negative number; the product of the percentage of the value and the non-negative number is recorded as the comprehensive product. The passenger flow characteristic value of a single station is directly proportional to the comprehensive product and inversely proportional to the dispersion.

5. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The process of classifying all stations is as follows: Calculate the arithmetic mean of the passenger flow characteristic values ​​of each station within a preset time range prior to the stated day; All stations are classified by comparing their arithmetic mean and geospatial coordinates.

6. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The construction process of the weighted undirected graph is as follows: If any two nodes contain stations that are adjacent on the track, connect the two nodes. For a single passage record, if any node contains a station as the entry station and any other node contains a station as the exit station, increment the weight of each edge along the shortest path between any node and any other node by 1. Obtain the edge weights from the passage records within a preset time period prior to the given day.

7. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The process for obtaining the importance coefficient of the integrated hub is as follows: Normalize the weights of all edges and calculate the mean of the normalized weights of all edges connected to each node. Count the shortest path hop count from each node to every other node, and sum the shortest path hop counts from each node to all other nodes. The importance coefficient of the integrated hub is directly proportional to the mean and inversely proportional to the summation result.

8. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The importance coefficient of the integrated hub is the product of the reciprocal of the summation result and the mean.

9. The method for constructing a digital twin entity based on spatiotemporal multidimensional data fusion as described in claim 1, characterized in that, The selection process for the key stations is as follows: A threshold for the segmentation of the comprehensive hub importance coefficient of all nodes is preset. Nodes whose comprehensive hub importance coefficient is greater than the segmentation threshold are designated as key nodes, and the stations contained in the key nodes are designated as key stations.

10. A digital twin entity construction system based on spatiotemporal multidimensional data fusion, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the digital twin entity construction method based on spatiotemporal multidimensional data fusion as described in any one of claims 1-9.