A respiratory infectious disease cross-region transmission risk assessment method based on spatio-temporal signaling digital twinning
Patent Information
- Application Number
- CN202611064353.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-08-21
AI Technical Summary
[0003]现有方法往往难以有效整合具备高时空分辨率的微观时空轨迹数据与宏观交通网络数据,从而导致风险评估模型无法动态适应突发疫情下人群流动格局的快速变化,传统模型可能忽略疫情发生后人群跨区域流动的实时调整对病毒扩散的影响,或未能精细刻画不同来源地人群的输入风险差异、在重点场所的停留时长、跨区域连续移动轨迹及来源地风险叠加效应等因素,从而降低风险识别的空间精度和时间敏感性,同时多数系统过度依赖周期性批量数据处理,无法实现小时级或更短周期的风险动态更新,且评估维度单一,难以同时覆盖输入性风险与输出性风险的双向评估,例如,在疫情早期阶段,若模型未能根据实时信令数据识别重点传播通道和关键风险源区,可能导致防控响应滞后与资源配置失准
本申请提供的一种融合时空信令数字孪生的呼吸道传染病跨区域传播风险评估方法,方法包括如下步骤:采集多源时空数据,通过条件扩散模型实现轨迹超分辨率重建与隐私脱敏,构建个体移动数字孪生体,结合超图高阶聚集传播网络与多尺度自适应融合,输入时空图神经ODE传播模型并耦合计算流体动力学气溶胶传播建模,进行连续时间多尺度动力学推演,基于自组织临界性理论预测相变临界点,通过因果干预式反演识别关键传播路径,采用主动贝叶斯优化选择监测点并结合免疫景观演化输出风险评估结果,本申请有效整合微观个体行为与宏观传播机制,解决了现有技术中数据更新滞后、轨迹重建精度不足、高阶传播关系刻画不全面及风险评估时效性差等问题,能够实时动态评估能力、高精度轨迹重建、多尺度传播动力学建模、关键传播路径精准识别。
Smart Images

Figure CN122619408A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infectious disease transmission risk assessment technology, and in particular to a method for assessing the cross-regional transmission risk of respiratory infectious diseases by integrating spatiotemporal signaling digital twins. Background Technology
[0002] Existing technologies for assessing the risk of cross-regional transmission of respiratory infectious diseases mainly rely on historical epidemiological survey data, case report data, traditional traffic passenger flow statistics, and macro-level population flow information at the administrative division level. For example, some systems use the traditional SEIR transmission dynamics model, combined with population census data and static traffic network parameters for risk assessment. Some methods attempt to integrate railway and civil aviation passenger flow data or mobile communication data, but these are usually only used as auxiliary inputs and have not achieved in-depth collaborative analysis of macro and micro multi-source flow data.
[0003] Existing methods often struggle to effectively integrate high-spatiotemporal resolution micro-level spatiotemporal trajectory data with macro-level transportation network data. This results in risk assessment models being unable to dynamically adapt to the rapid changes in population movement patterns during sudden outbreaks. Traditional models may overlook the impact of real-time adjustments in cross-regional population movement on virus spread after an outbreak, or fail to accurately characterize factors such as differences in input risk from different origins, duration of stay in key locations, continuous cross-regional movement trajectories, and the cumulative effect of origin risks. This reduces the spatial accuracy and temporal sensitivity of risk identification. Furthermore, most systems rely excessively on periodic batch data processing, failing to achieve hourly or shorter-cycle dynamic risk updates. Moreover, their assessment dimensions are singular, making it difficult to simultaneously cover both input and output risks. For example, in the early stages of an outbreak, if the model fails to identify key transmission channels and critical risk sources based on real-time signaling data, it may lead to delayed prevention and control responses and inaccurate resource allocation.
[0004] Therefore, in response to the aforementioned problems of insufficient macro- and micro data fusion capabilities and lack of timeliness and precision in risk assessment, this invention proposes a method for assessing the cross-regional transmission risk of respiratory infectious diseases by integrating spatiotemporal signaling digital twins. Through the fusion calibration of operator spatiotemporal signaling trajectory data and macro-mobility network, coupled with dynamic risk quantification calculation of transmission parameters, a high-precision and timely cross-regional transmission risk assessment and key transmission path identification can be achieved. Summary of the Invention
[0005] The purpose of this application is to provide a method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins, which can perform real-time dynamic assessment, high-precision trajectory reconstruction, multi-scale transmission dynamics modeling, and accurate identification of key transmission paths.
[0006] In a first aspect, this application provides a method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins, the method comprising the following steps: Collect multi-source spatiotemporal data, including operator spatiotemporal signaling trajectory data, disease monitoring time series data, macro-population flow network data, building wind field data, and location POI data; The operator's spatiotemporal signaling trajectory data records the location and time information of the mobile terminal connecting to the base station at different time periods; The spatiotemporal signaling trajectory data is super-resolution reconstructed and privacy-desensitized using a conditional diffusion model to construct an individual mobile digital twin, identify stop points, and extract the microscopic OD matrix. The individual mobile digital twin simulates the spatiotemporal movement behavior pattern of each individual; A high-order hypergraph aggregation and propagation network is constructed. Based on the credibility-real-time performance-coverage joint weight mechanism, the micro-signaling flow network and the macro-population flow network are adaptively fused at multiple scales to generate a fused propagation hypernetwork. The hypergraph high-order aggregation propagation network uses regions as nodes and multi-group aggregation events as hyperedges to characterize ternary and higher-order propagation relationships. The fused propagation supernetwork is input into the spatiotemporal graph neural ODE propagation model, combined with computational fluid dynamics aerosol physics propagation modeling, to perform continuous-time multi-scale propagation dynamics deduction, and calculate the input propagation risk, output propagation risk and comprehensive propagation risk index of each region; Based on the theory of self-organized criticality, the phase transition critical point of the propagation system is predicted. The causal risk contribution of each cross-regional flow path is calculated through causal intervention-based inversion, and key propagation paths are identified. Active Bayesian optimization is used to select adaptive monitoring points, combined with an immune landscape spatiotemporal evolution model, to output risk confidence intervals, phase transition early warning signals, and dynamic risk maps.
[0007] Furthermore, the step of constructing an individual mobile digital twin by performing super-resolution reconstruction and privacy desensitization on the spatiotemporal signaling trajectory data using a conditional diffusion model includes: Through the formula: Calculate the same user The Article and No. Instantaneous movement speed between continuous trajectory records ;in, Indicates the first Trajectory Position With the Trajectory Position Spatial distance between them Indicates the first The timestamp of each trajectory record; When the instantaneous moving speed exceeds the preset maximum speed threshold, a conditional diffusion model is used to perform generative super-resolution completion of the missing trajectory segment; The conditional diffusion model uses known trajectory points as conditions to generate high-resolution intermediate trajectory points through stepwise denoising. The forward diffusion process in the conditional diffusion model is defined by the following formula: in, Indicates the first The trajectory after adding noise, Indicates the first noise variance of the step Indicates a normal distribution. Represents the identity matrix; The reverse generation process uses a neural network to predict noise and gradually remove it to reconstruct a high-resolution trajectory sequence. Through the formula: The user identifier is irreversibly de-identified using salted hashing to generate a de-identified user ID; where, Indicates the original user ID. Indicates the salting parameters. This represents an irreversible hash function; By mapping base station locations to standard spatial grid cells, an individual mobile digital twin is constructed, which includes the individual's spatiotemporal trajectory, stay preferences, travel patterns, and activity modes.
[0008] Furthermore, the construction of the hypergraph high-order aggregation propagation network and multi-scale adaptive fusion includes: Through the formula: Identify users Set of effective dwelling points ;in, Indicates the center location of the stop. Indicates the time of entry into the stop. Indicates the time of departure from the stop. Indicates the spatial radius threshold. Indicates the time threshold; Individual travel chains are extracted based on the continuous stop-point sequence, generating a micro-OD matrix with sample expansion weights, using the formula: Calculation time window Inner from the region Flow to area microfluidity ;in, Indicates an indicator function, Indicates user In the time window The origin and destination areas of the trip. Indicates user Sample expansion weights; Constructing a high-order aggregation and propagation network of a hypergraph ,in Represents a set of region nodes. Denotes the set of superedges. This represents the hyperedge weight matrix, which stores the weight value of each hyperedge. Each hyperedge connects three or more regional nodes that stay in the same location within the same time window, thus depicting higher-order propagation relationships. The hyperedge weight is determined by the formula: Calculate the first Superedge in the time window weight ;in, Indicates the first The set of users corresponding to each hyperedge Indicates user Infection pressure in the source area Indicates user The length of time spent at this location, This indicates the population density of the location; The edge weights of the micro-flow network and the macro-flow network are normalized respectively. The fusion weights are calculated through a joint weighting mechanism of credibility, real-time performance and coverage to generate a fusion propagation supernetwork.
[0009] Furthermore, the computational fluid dynamics aerosol physics propagation modeling includes: Based on building wind field data and topographic data, airflow is simulated using the Reynolds-averaged Navier-Stokes equations, as shown in the formula: Calculate the velocity distribution in the flow field; where, Indicates air density, express The velocity component in the direction, Indicates pressure, Represents the viscous stress tensor. Represents the components of gravitational acceleration; Based on computational fluid dynamics flow fields, the diffusion process of viral aerosols is simulated using the convection-diffusion equation: Calculate the viral concentration distribution; where, Indicates the concentration of viral aerosols. Indicates fluid velocity. Indicates the diffusion coefficient. Indicates the virus source item; The viral source is determined by the infected person's respiratory rate, viral load, and the correction factor for wearing a mask. Individual exposure dose is determined by the formula: Calculate users viral exposure dose ;in, Indicates user At the venue The duration of stay Indicates user In time Location, Indicates user respiratory rate; The probability of infection is calculated using the following dose-response model: ; in, This indicates the viral infectivity coefficient.
[0010] Furthermore, the spatiotemporal graph neural ODE propagation model is coupled with multi-scale dynamics, including: Construct a hypergraph neural network layer, extract high-order propagation features through hypergraph convolution, using the formula: Hypergraph convolution feature extraction is performed; among which, Represents the node at level l Hidden features, Indicates the presence of nodes The set of superedges Indicates the superedge The number of nodes, Represents a node The degree, Represents the l-th layer hyperedge The weight, Indicates the activation function; Using the initial features output by the hypergraph neural network as initial values, continuous-time evolution is performed through the neural network's ordinary differential equations, defined by the following formula: Define the time derivative of the hidden state; where, Indicates time The hidden state, Indicates by parameters Parameterized neural network functions; By adaptive step size The solver integrates from the initial time to the target time to obtain the hidden state at any time point; it realizes multi-scale coupling of the micro-individual layer, the meso-community layer and the macro-region layer, and the information exchange between the scales is carried out through upsampling and downsampling to form a hierarchical propagation dynamic system.
[0011] Furthermore, the prediction of the phase transition critical point of the propagation system based on the self-organizing criticality theory includes: The order parameters of the propagation system are calculated using the formula: Calculation time average infection rate As an order parameter; where... Indicates the total number of regions. Indicates the area The infection rate; The fluctuations of the order parameter are calculated using the formula: Calculate susceptibility ;in, This represents the variance of the infection rate in each region. This represents the average infection rate; When the system approaches the phase transition critical point, the susceptibility exhibits a divergent peak, as shown by the formula: Describe the critical divergence behavior; where, Indicates the propagation control parameters, Indicates the critical point. Indicates the critical index; By using finite-size scaling analysis and the renormalization group method, the phase transition critical point of the propagation system from the controllable state to the explosive state is predicted, and a phase transition early warning signal is output. The self-organized criticality is verified by avalanche size distribution analysis. When the avalanche size follows a power-law distribution, the system is determined to be in a critical state.
[0012] Furthermore, the spatiotemporal evolution of the immune landscape is bidirectionally coupled with the propagation dynamics, including: Construct a population immunity landscape, which describes the spatial distribution of population immunity levels in different regions and at different times; Regional immunity levels are expressed by the formula: Computational area In time immune level ;in, Indicates vaccination coverage rate. Indicates the proportion of the recovered population. This indicates the immune memory formed by historical infections. Indicates the weighting coefficient; Immune memory decays over time, as shown by the formula: Describing the temporal evolution of the immune landscape; among which, Indicates the rate of immune attenuation. This indicates new immunity, namely vaccination and natural infection; A two-way coupling between immune landscape and transmission dynamics: immune landscape influences transmission rate, and transmission dynamics alter the distribution of immune landscape; The proportion of susceptible individuals is determined by the immune landscape, as shown by the formula: Computational area Susceptibility ; Regional contact intensity is obtained by weighted summation of standardized population density, place activity, and public transport activity.
[0013] Furthermore, the causal intervention-based inversion calculation of the causal risk contribution of each cross-regional flow path includes: A propagation causal graph is constructed based on a structural causal model, defining causal variables and causal paths in the propagation process; Causal intervention is performed using the do operator, and the average causal effect of the path is calculated using the formula: Calculation path Average causal effect ;in, This indicates a causal intervention operation. Expressing expectations, Indicates the area The risk value, Indicates from arrive The flow rate; By controlling confounding factors through backdoor and frontdoor adjustments, the unbiasedness of causal effect estimation is ensured. Through the formula: Computation source region For the target area The causal risk contribution; among which, This represents a very small constant to prevent the denominator from being zero; When the causal risk contribution exceeds the contribution threshold, the path will be... Identified as a key input propagation path; Causal betweenness centrality measures the criticality of a node in the propagation causal chain and is used to identify key propagation hubs at the network level.
[0014] Furthermore, the adoption of active Bayesian optimization to select adaptive monitoring points includes: Gaussian processes are used to model the spatial distribution of risk, and the risk values of each region are modeled as samples of Gaussian processes. The prior distribution of a Gaussian process is defined by the mean function and the covariance function, and is defined by the following formula: ;in, Represents the mean function, This represents the covariance function, also known as the kernel function. Based on existing monitoring point data, a posterior estimate of the risk distribution is obtained through Bayesian inference. The next optimal monitoring point is selected by acquiring a function to maximize information gain. The acquisition function is calculated using the following formula: Calculate; where, Indicates position The posterior mean of risk, Indicates position The posterior standard deviation of risk, This indicates an exploration using equilibrium parameters; The upper confidence bound acquisition function is used to achieve a balance between regions of high uncertainty and high risk. By using Bayesian optimization and iterative selection of monitoring points, the accuracy of risk assessment can be maximized under the constraint of monitoring cost. Uncertainty is decomposed into two parts: cognitive uncertainty and accidental uncertainty, which correspond to model parameter uncertainty and inherent noise in the data, respectively.
[0015] Furthermore, the method also includes interpretable propagation reasoning and prevention strategy generation based on neural symbol fusion, specifically: A neural symbolic fusion reasoning framework is constructed, with a neural network perception layer at the bottom and a symbolic logic reasoning layer at the top; The neural network layer extracts high-dimensional propagation features, and the symbolic logic layer performs interpretable reasoning based on epidemiological knowledge. The continuous features of a neural network are mapped to discrete propositions of symbolic logic through a neural symbolic interface. Construct an interpretable rule set for propagating causal chains, with each rule containing preconditions and conclusions, supporting human experts' understanding and verification; The objective function for optimizing the prevention and control strategy is defined by the following formula: ;in, Indicates prevention and control strategies, Representing the policy space, Representation Strategy Lower region exist Risk value at any given moment Representation Strategy Implementation costs Representation Strategy The risks and uncertainties below Indicates the weighting coefficient; The optimal prevention and control strategy is solved using a deep reinforcement learning algorithm, and the strategy is verified and iteratively optimized in a digital twin. Output differentiated prevention and control recommendations, resource allocation plans, and reports explaining the causal chain of transmission; The model parameters were calibrated using a regularization optimization method, and the model's predictive performance was verified using historical epidemic data. The verification metrics included RMSE, Spearman correlation coefficient, AUC, and Top-K hit rate.
[0016] Compared with the prior art, this application has the following advantages: This application provides a method for risk assessment of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins. The method includes the following steps: collecting multi-source spatiotemporal data; achieving trajectory super-resolution reconstruction and privacy desensitization through a conditional diffusion model; constructing an individual mobile digital twin; combining a hypergraph high-order aggregation propagation network with multi-scale adaptive fusion; inputting a spatiotemporal graph neural ODE propagation model and coupling it with computational fluid dynamics aerosol propagation modeling; performing continuous-time multi-scale dynamic extrapolation; predicting phase transition critical points based on self-organized criticality theory; identifying key transmission paths through causal intervention-based inversion; and using active Bayesian optimization to select monitoring points and combining immune landscape evolution to output risk assessment results. This application effectively integrates microscopic individual behavior and macroscopic transmission mechanisms, solving problems such as data update lag, insufficient trajectory reconstruction accuracy, incomplete characterization of high-order transmission relationships, and poor timeliness of risk assessment in existing technologies. It has the ability to perform real-time dynamic assessment, high-precision trajectory reconstruction, multi-scale transmission dynamic modeling, and accurate identification of key transmission paths. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins, as provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0020] like Figure 1 As shown, this application provides a method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins.
[0021] S100 collects multi-source spatiotemporal data, including operator spatiotemporal signaling trajectory data, disease monitoring time series data, macro-population flow network data, building wind field data, and location POI data.
[0022] Operator spatiotemporal signaling trajectory data can be provided by operators in batch files on a regular basis; disease surveillance time series data can be manually extracted from statistical reports of public health departments; macro-population flow network data can be summarized based on annual statistical data of transportation departments; building wind field data can be estimated based on historical meteorological records or common building types; and location POI data can be downloaded in batches from public map services.
[0023] S200, The operator's spatiotemporal signaling trajectory data records the location and time information of the mobile terminal connecting to the base station at different time periods.
[0024] S300. The spatiotemporal signaling trajectory data is super-resolution reconstructed and privacy-desensitized using a conditional diffusion model to construct an individual mobile digital twin, identify stop points, and extract the microscopic OD matrix.
[0025] Super-resolution reconstruction of spatiotemporal signaling trajectory data can employ a simple linear interpolation method, which involves uniformly inserting intermediate points between two known base station connection points according to time intervals. Privacy desensitization can be achieved through anonymization, such as replacing user identifiers with randomly generated strings. Individual mobile digital twins can be constructed based on these raw or simply interpolated trajectory data by statistically analyzing the number of times and duration each user stays at different base stations, recording the most frequently occurring base station locations and approximate directions of movement. Stop point identification can be achieved by setting fixed spatial radii and time thresholds; for example, if a user stays near a base station for more than a preset time, it is considered a stop point. The microscopic OD matrix can be generated directly by statistically analyzing the number of movements from one stop point to another based on a continuous sequence of user stop points.
[0026] S400, the individual mobile digital twin simulates the spatiotemporal movement behavior pattern of each individual.
[0027] S500 constructs a high-order aggregation and propagation network of a hypergraph, and performs multi-scale adaptive fusion of the micro-signaling flow network and the macro-population flow network based on a joint weight mechanism of credibility, real-time performance and coverage to generate a fused propagation hypernetwork.
[0028] The construction of a high-order hypergraph aggregation propagation network can be simplified as follows: First, treat each region as a node. Then, identify multi-crowd aggregation events and construct hyperedges through manual analysis or based on simple rules (e.g., if individuals from three or more regions simultaneously appear at a specific location within the same time period, connect these regions to form a hyperedge). The weight of the hyperedge can be simply set as the number of regions participating in the aggregation. The fusion of the micro-signaling flow network and the macro-crowd flow network can employ a fixed-ratio weighted average method. For example, after standardizing the micro-flow data and macro-flow data respectively, they are linearly superimposed with preset fixed weights (e.g., each accounting for 50%) to generate a fused propagation hypernetwork.
[0029] S600 The hypergraph high-order aggregation propagation network uses regions as nodes and multi-group aggregation events as hyperedges to characterize ternary and higher-order propagation relationships.
[0030] S700. Input the fused propagation supernetwork into the spatiotemporal graph neural ODE propagation model, combine it with the aerosol physics propagation modeling of computational fluid dynamics, perform continuous time multi-scale propagation dynamics deduction, and calculate the input propagation risk, output propagation risk and comprehensive propagation risk index of each region.
[0031] The fusion propagation supernetwork can be input into the traditional SIR (Susceptible-Infected-Recovered) model, where the propagation probability between nodes is determined by the network connection strength. The spatiotemporal graph neural ODE propagation model can be replaced by a discrete-time-step graph convolutional network model, which updates the infection status of nodes at each time step. Computational fluid dynamics aerosol physics propagation modeling can be simplified to estimations based on empirical formulas or simplified diffusion models; for example, the viral concentration in the air can be directly estimated based on ventilation conditions and the number of infected individuals, without complex flow field simulations. The imported transmission risk, exported transmission risk, and overall transmission risk index for each region can be calculated based on simple infection rates and population density; for example, the imported risk can be defined as the number of infected individuals flowing in from other regions, and the exported risk can be defined as the number of infected individuals flowing out of the current region.
[0032] S800 predicts the phase transition critical point of a propagation system based on the theory of self-organized criticality, and identifies key propagation paths by calculating the causal risk contribution of each cross-regional flow path through causal intervention-based inversion.
[0033] Predicting the phase transition critical point of a transmission system can be based on statistical analysis of historical data, for example, by setting an empirical critical point by observing the infection rate thresholds during past outbreaks. Causal intervention-based inversion can be replaced by methods based on correlation analysis, such as calculating the Pearson correlation coefficient between different flow paths and the risk value of the target area, identifying paths with higher correlation coefficients as critical transmission paths. This method does not involve causal inference; it only reflects statistical associations.
[0034] The S900 employs active Bayesian optimization to select adaptive monitoring points, combined with an immune landscape spatiotemporal evolution model, to output risk confidence intervals, phase transition early warning signals, and dynamic risk maps.
[0035] The selection of adaptive monitoring points can employ a risk-ranking strategy, for example, selecting the N regions with the highest current risk index as monitoring points. The spatiotemporal evolution model of the immune landscape can be simplified to a simple linear model based on fixed immune decay rates and vaccination rates, without considering complex immune memory and cross-immunity. Risk confidence intervals can be simply estimated based on the standard deviation of historical data. Phase transition warning signals can be triggered based on preset infection rate thresholds. Dynamic risk maps can easily visualize the risk index of each region on a map using different colors.
[0036] This embodiment deeply integrates multi-source spatiotemporal information, including operator spatiotemporal signaling trajectory data, disease surveillance data, and macroscopic population movement networks, to construct individual mobile digital twins and a hypergraph high-order aggregation propagation network. It then combines a spatiotemporal graph neural ODE model with aerosol physical propagation modeling to perform multi-scale dynamic extrapolation. Therefore, this method overcomes the limitations of traditional assessment techniques, such as insufficient fusion of macro and micro data, poor model dynamic adaptability, and low accuracy and timeliness in risk identification. It achieves high-precision and high-timeliness assessment of the cross-regional transmission risk of respiratory infectious diseases, identifies key transmission paths, and outputs risk confidence intervals, phase transition warning signals, and dynamic risk maps, providing decision support for prevention and control response and resource allocation.
[0037] In some possible implementations, the step of performing super-resolution reconstruction and privacy desensitization on the spatiotemporal signaling trajectory data using a conditional diffusion model to construct an individual mobile digital twin includes: Through the formula: Calculate the same user The Article and No. Instantaneous movement speed between continuous trajectory records ;in, Indicates the first Trajectory Position With the Trajectory Position Spatial distance between them Indicates the first The timestamp of each trajectory record.
[0038] By analyzing the velocity between adjacent trajectory points, we can identify possible missing trajectory segments due to excessively large data sampling intervals or signal loss.
[0039] When the instantaneous moving speed exceeds the preset maximum speed threshold, a conditional diffusion model is used to perform generative super-resolution completion of the missing trajectory segment; The conditional diffusion model uses known trajectory points as conditions to generate high-resolution intermediate trajectory points through stepwise denoising. The forward diffusion process in the conditional diffusion model is defined by the following formula: in, Indicates the first The trajectory after adding noise, Indicates the first noise variance of the step Indicates a normal distribution. Represents the identity matrix.
[0040] This process simulates the gradual addition of noise to trajectory data. Subsequently, the reverse generation process predicts the noise and gradually removes it using a neural network, thereby reconstructing a high-resolution trajectory sequence. This effectively fills the sparsity of the original data and improves the continuity and precision of the trajectory.
[0041] The reverse generation process uses a neural network to predict noise and gradually remove it to reconstruct a high-resolution trajectory sequence. Through the formula: The user identifier is irreversibly de-identified using salted hashing to generate a de-identified user ID; where, Indicates the original user ID. Indicates the salting parameters. This indicates an irreversible hash function.
[0042] Salted hashing ensures the anonymity of user identifiers. Even if the original data is leaked, it is impossible to deduce the original user's identity, thus meeting the compliance requirements for data privacy protection.
[0043] By mapping base station locations to standard spatial grid cells, an individual mobile digital twin is constructed, which includes the individual's spatiotemporal trajectory, stay preferences, travel patterns, and activity modes.
[0044] By standardizing base station locations, the impact of differences in coverage and accuracy between different base stations can be eliminated, enabling individual mobile digital twins to more accurately reflect an individual's real mobile behavior patterns.
[0045] In some possible implementations, the construction of the hypergraph high-order aggregation propagation network and multi-scale adaptive fusion includes: Through the formula: Identify users Set of effective dwelling points ;in, Indicates the center location of the stop. Indicates the time of entry into the stop. Indicates the time of departure from the stop. Indicates the spatial radius threshold. Indicates the time threshold; Individual travel chains are extracted based on the continuous stop-point sequence, generating a micro-OD matrix with sample expansion weights, using the formula: Calculation time window Inner from the region Flow to area microfluidity ;in, Indicates an indicator function, Indicates user In the time window The origin and destination areas of the trip. Indicates user Sample expansion weights; Constructing a high-order aggregation and propagation network of a hypergraph ,in Represents a set of region nodes. Denotes the set of superedges. This represents the hyperedge weight matrix, which stores the weight value of each hyperedge. Each hyperedge connects three or more regional nodes that stay in the same location within the same time window, thus depicting higher-order propagation relationships. The hyperedge weight is determined by the formula: Calculate the first Superedge in the time window weight ;in, Indicates the first The set of users corresponding to each hyperedge Indicates user Infection pressure in the source area Indicates user The length of time spent at this location, This indicates the population density of the location; The edge weights of the micro-flow network and the macro-flow network are normalized respectively. The fusion weights are calculated through a joint weighting mechanism of credibility, real-time performance and coverage to generate a fusion propagation supernetwork.
[0046] As an implementation method, identifying the set of effective dwell points (Stayu) for user u is fundamental to understanding individual behavioral patterns and constructing micro-travel chains. Effective dwell points refer to locations where a user stays for more than a certain time threshold within a specific spatial range. By setting spatial radius thresholds ϵ and time thresholds τ, it is possible to accurately filter out a user's actual dwelling behavior at a particular location, rather than just a brief passage. This can be achieved through clustering analysis or sliding window analysis of the operator's spatiotemporal signaling trajectory data, grouping continuous, spatially close, and temporally sustained trajectory points into a single dwelling event.
[0047] Individual travel chains are extracted based on continuous stop sequence, and a micro-OD matrix with sample expansion weights is generated. This aims to quantify the flow of traffic generated by individual movement between different regions within a specific time window. Individual travel chains reflect users' travel patterns, while the micro-OD matrix provides refined flow data. The sample expansion weight wu is introduced to compensate for the bias in population coverage of signaling data. Statistical methods are used to infer sample data to the overall population, making the micro-OD matrix more representative. Travel chain extraction is accomplished by analyzing the temporal sequence of the effective stop set Stayu for user u, while the generation of the micro-OD matrix involves counting the number of trips each user makes from region i to region j within time window t and multiplying it by its corresponding sample expansion weight wu.
[0048] A high-order hypergraph clustering propagation network GH=VH,EH,WH is constructed, where VH represents the set of region nodes, EH represents the set of hyperedges, and WH represents the hyperedge weight matrix. This network is used to characterize the high-order propagation relationships formed when multiple people gather in the same place during the spread of respiratory infectious diseases. Hypergraphs can represent relationships between more than two nodes, unlike traditional graphs which can only represent binary relationships. Each hyperedge connects three or more region nodes that remain in the same place within the same time window. This more realistically reflects the transmission mechanism of viruses in crowd gathering events. For example, in scenarios such as meetings or dinners, the virus may simultaneously spread from one infected person to multiple susceptible individuals, or cross-infection may occur between people from multiple regions. This can be achieved by identifying events in which three or more individuals from different regions simultaneously remain in the same place within the same time window t, thus forming hyperedges.
[0049] The calculation of the hyperedge weight we(t) quantifies the importance or risk level of a specific high-order clustering event in its propagation. It comprehensively considers the set of users Ue participating in the clustering event, the infection pressure in the areas from which these users originate (Asrc(u)(t)), the duration of user stay at the location (τu,e(t)), and the population density (densitye(t)) of the location. This multi-factor weighting mechanism allows the hyperedge weight to more comprehensively reflect the potential propagation risk of high-order clustering events. For each hyperedge e, the infection pressure in the user's origin area can be obtained from disease surveillance data, the duration of user stay at the location can be obtained from individual mobile digital twins, and the population density of the location can be calculated by combining location POI data and real-time pedestrian flow data. These factors are then multiplied according to the formula.
[0050] The edge weights of the micro-flow network and the macro-flow network are normalized separately, and the fusion weights are calculated through a joint weighting mechanism of credibility, real-time performance, and coverage to generate a fused propagation supernetwork. Micro-flow networks (based on signaling data) and macro-flow networks (based on statistical data) each have their advantages and disadvantages. Normalization eliminates the dimensional differences, and the fusion is achieved through a joint weighting mechanism of credibility, real-time performance, and coverage. The aim is to combine the advantages of both to generate a more comprehensive and accurate fused propagation supernetwork that reflects both individual fine-grained behavior and macro-level trends. Credibility assesses the accuracy of the data source, real-time performance assesses the timeliness of the data, and coverage assesses the representativeness of the data to the population. Normalization can employ methods such as Min-Max normalization or Z-score normalization, while the joint weighting mechanism can be determined through expert scoring or quantitative methods based on data quality indicators.
[0051] In some possible implementations, the computational fluid dynamics aerosol physics propagation modeling includes: Based on building wind field data and topographic data, airflow is simulated using the Reynolds-averaged Navier-Stokes equations, as shown in the formula: Calculate the velocity distribution in the flow field; where, Indicates air density, express The velocity component in the direction, Indicates pressure, Represents the viscous stress tensor. Represents the components of gravitational acceleration; Based on computational fluid dynamics flow fields, the diffusion process of viral aerosols is simulated using the convection-diffusion equation: Calculate the viral concentration distribution; where, Indicates the concentration of viral aerosols. Indicates fluid velocity. Indicates the diffusion coefficient. Indicates the virus source item; The viral source is determined by the infected person's respiratory rate, viral load, and the correction factor for wearing a mask. Individual exposure dose is determined by the formula: Calculate users viral exposure dose ;in, Indicates user At the venue The duration of stay Indicates user In time Location, Indicates user respiratory rate; The probability of infection is calculated using the following dose-response model: ; in, This indicates the viral infectivity coefficient.
[0052] As one approach, this method uses building wind field data and topographic data to simulate airflow and calculate the velocity distribution of the flow field using the Reynolds-averaged Navier-Stokes equations. The aim is to accurately capture the motion of air—the medium for virus aerosol transmission. The Reynolds-averaged Navier-Stokes equations are a set of partial differential equations used to describe fluid motion. By performing time averaging or ensemble averaging on the instantaneous Navier-Stokes equations, turbulent flow can be effectively handled, thus capturing the macroscopic characteristics and average velocity field of airflow even with limited computational resources. In practical applications, this typically requires using specialized computational fluid dynamics (CFD) software to convert building structure and topographic information into a computational mesh, setting boundary conditions such as inlet, outlet, and walls, and then iteratively calculating a stable velocity distribution using a numerical solver.
[0053] Based on computational fluid dynamics flow fields, this study simulates the diffusion process of viral aerosols and calculates the viral concentration distribution using convection-diffusion equations. The aim is to track and predict the propagation trajectory and concentration changes of viral aerosol particles within a known airflow field. The convection-diffusion equations comprehensively consider the convection effect of aerosols with airflow (determined by fluid velocity u) and the diffusion effect due to concentration gradients (determined by the diffusion coefficient D). The viral source term S represents the rate at which an infected individual releases viral aerosols into the environment, and its accuracy directly affects the calculated concentration distribution.
[0054] The viral source term is jointly determined by the infected individual's respiratory rate, viral load, and a mask-wearing correction factor, representing a refined model of the viral release process. The respiratory rate reflects the volumetric flow rate of exhaled air; viral load quantifies the number of viral particles per unit volume of exhaled air; and the mask-wearing correction factor quantifies the mask's efficiency in blocking viral particles, typically a multiplicative factor between 0 and 1, representing the percentage reduction in viral release after wearing a mask. This comprehensive consideration of parameters allows the viral source term to more accurately reflect the viral contribution of different individuals to the environment under different conditions.
[0055] Individual exposure dose (Doseu) is calculated for user u using a formula. The core of this formula is to combine the viral concentration in the physical environment with individual behavior (duration of stay, location) and physiological characteristics (respiratory rate) to quantify the actual amount of virus inhaled by the individual. This integral process considers the viral concentration experienced by the user at different times and locations within location e, multiplied by their respiratory rate, to obtain the cumulative viral exposure.
[0056] The probability of infection is calculated using a dose-response model, which aims to translate a quantified viral exposure dose into the risk of infection in an individual. Dose-response models are commonly used tools in epidemiology and toxicology, describing the quantitative relationship between exposure dose and the organism's response (in this case, infection). The exponential model used in this application is a common simplified model, where the viral infectivity coefficient k is a key parameter reflecting viral virulence and host susceptibility, typically obtained through experimental data or epidemiological analysis.
[0057] In some possible implementations, the spatiotemporal graph neural ODE propagation model is coupled with multi-scale dynamics. The hypergraph neural network layer is a neural network structure capable of processing hypergraph data, its core being the aggregation of information through hypergraph convolution operations. A hypergraph is a generalized graph that allows an edge (called a hyperedge) to connect any number of nodes, unlike traditional graphs which only connect two nodes. In respiratory infectious disease propagation scenarios, this allows the model to capture more complex high-order propagation relationships. For example, when multiple individuals gather at the same time and place, and may be simultaneously exposed to pathogens, such multi-party propagation events cannot be effectively represented by traditional binary relation graphs. Hypergraph convolution, by aggregating information from all nodes within the hyperedge, can effectively extract the propagation features inherent in these high-order, multi-party interaction patterns, thus more accurately reflecting real-world propagation mechanisms, including: Construct a hypergraph neural network layer, extract high-order propagation features through hypergraph convolution, using the formula: Hypergraph convolution feature extraction is performed; among which, Represents the node at level l Hidden features, Indicates the presence of nodes The set of superedges Indicates the superedge The number of nodes, Represents a node The degree, Represents the l-th layer hyperedge The weight, Indicates the activation function; This mechanism allows each node's feature updates to depend not only on its direct neighbors but also on the higher-order group activities it participates in, thus enabling a more comprehensive capture of the complexity of propagation.
[0058] Using the initial features output by the hypergraph neural network as initial values, continuous-time evolution is achieved through the Neural Ordinary Differential Equation (NODE). The NODE is a deep learning model that treats the layers of a neural network as a continuous transformation process. It learns the continuous dynamics of data by solving an ordinary differential equation parameterized by the neural network. In infectious disease transmission modeling, the transmission process is essentially a continuous dynamic evolution process. However, traditional graph neural networks typically perform information propagation and feature updates at discrete time steps, which may lead to approximation errors or information loss regarding continuous dynamics. By using the discrete features output by the hypergraph neural network as the initial state of the NODE, this method can achieve continuous-time evolution of the transmission process, defined by the following formula: Define the time derivative of the hidden state; where, Indicates time The hidden state, Indicates by parameters Parameterized neural network functions.
[0059] It determines how the hidden state h(t) changes continuously with time t. By employing an ODE solver with an adaptive step size, the integration can proceed from the initial time to the target time, thereby obtaining the hidden state at any time point and achieving continuous, smooth, and accurate modeling of the propagation process.
[0060] By adaptive step size The solver integrates from the initial time to the target time to obtain the hidden state at any time point; it realizes multi-scale coupling of the micro-individual layer, the meso-community layer and the macro-region layer, and the information exchange between the scales is carried out through upsampling and downsampling to form a hierarchical propagation dynamic system.
[0061] This embodiment also achieves multi-scale coupling at the micro-individual, meso-community, and macro-regional levels. Information exchange between these scales occurs through upsampling and downsampling, forming a hierarchical transmission dynamics system. The spread of respiratory infectious diseases is a complex multi-scale phenomenon, involving individual contact behavior at the micro-level, population gatherings within communities or venues at the meso-level, and inter-regional population movement at the macro-level. Single-scale models often struggle to comprehensively capture these different granularity transmission mechanisms. This method constructs a hierarchical transmission dynamics system by achieving multi-scale coupling at the micro-individual, meso-community, and macro-regional levels. This means the model can simultaneously consider and integrate data and dynamic information at different granularities. Specifically, information exchange occurs between the scales through upsampling and downsampling mechanisms. For example, detailed movement trajectories, dwelling preferences, and contact events at the micro-individual level can be upsampled and aggregated to the meso-community level, forming population gathering events and potential transmission risks within communities or specific venues. The transmission trends and risk assessment results at the meso-community level can be further upsampled to the macro-regional level to assess inter-regional transmission risks and overall epidemic trends. Conversely, risk warnings, prevention and control strategies, or resource allocation plans at the macro-regional level can be refined and guided by downsampling mechanisms to inform behavioral interventions at the meso-community and micro-individual levels. This two-way, hierarchical information exchange ensures that the model can comprehensively understand and predict the spread of infectious diseases at different granularities, thereby providing a more comprehensive and refined risk assessment.
[0062] In some possible implementations, predicting the phase transition critical point of a propagation system based on self-organizing criticality theory first involves calculating the order parameters of the propagation system. Order parameters are key variables describing the macroscopic state of the system, reflecting the overall degree of "order" or activity level of the system, and include: The order parameters of the propagation system are calculated using the formula: Calculation time average infection rate As an order parameter; where... Indicates the total number of regions. Indicates the area The infection rate.
[0063] The average infection rate, as an order parameter, can intuitively reflect the activity level of the entire transmission system at a certain moment and is an important indicator for monitoring changes in the system's state.
[0064] Calculating the fluctuations of the order parameter, which represents the system's sensitivity to small disturbances, is a key indicator for determining whether the system is approaching a critical state. This can be achieved through the formula: Calculate susceptibility ;in, This represents the variance of the infection rate in each region. This indicates the average infection rate.
[0065] Susceptibility measures the strength of a system’s response to external disturbances or internal fluctuations. When a system approaches the critical point of a phase transition, its susceptibility increases significantly and may even exhibit divergent peaks.
[0066] When the system approaches the phase transition critical point, the susceptibility exhibits a divergent peak, as shown by the formula: Describe the critical divergence behavior; where, Indicates the propagation control parameters, Indicates the critical point. This represents the critical index.
[0067] This formula theoretically characterizes the susceptibility behavior of a system near a critical point, providing a mathematical basis for identifying critical points.
[0068] By using finite-size scaling analysis and the renormalization group method, the phase transition critical point of the propagation system from the controllable state to the explosive state is predicted, and a phase transition early warning signal is output.
[0069] Finite-size scaling analysis is a powerful tool for studying critical phenomena at finite system sizes, allowing the inference of critical behavior for infinitely large systems from finite data. Renormalization group methods provide a theoretical framework for understanding how system behavior evolves at different scales, revealing the universality of critical phenomena. The combination of these two methods enables accurate prediction of critical points in complex propagation systems, even under practically limited regions and data conditions.
[0070] The self-organized criticality is verified by avalanche size distribution analysis. When the avalanche size follows a power-law distribution, the system is determined to be in a critical state.
[0071] A key characteristic of self-organizing critical systems is that their "avalanche" events (e.g., the scale of an epidemic outbreak) follow a power-law distribution. When the avalanche size follows a power-law distribution, the system is considered to be in a critical state. Statistical analysis of actual or simulated propagation events verifies whether the avalanche size distribution conforms to a power law, thus providing empirical support for the applicability of the self-organizing criticality theory in this propagation system and further enhancing the reliability of the prediction results.
[0072] In some possible implementations, the spatiotemporal evolution of the immune landscape is bidirectionally coupled with transmission dynamics. Specifically, the method first constructs a population immune landscape, which aims to describe the spatial distribution of population immunity levels in different regions and at different times. The population immune landscape is a dynamic spatiotemporal data structure designed to quantify and visualize the overall immune status of a population in a specific geographic area against a respiratory infectious disease at different points in time. Its core function is to provide a fine and real-time perspective to understand the overall susceptibility distribution of the population, which is crucial for accurately assessing the disease's transmission potential. In practice, this can be constructed by combining a Geographic Information System (GIS) with a time-series database. Specifically, each predefined geographic area (e.g., urban administrative division, community unit, or grid area) is associated with a comprehensive immune level value at each discrete time step (e.g., daily, weekly). These values can come from sources including, but not limited to, vaccination records provided by public health departments, antibody positivity data obtained through serological surveys, and historical infection case data.
[0073] Construct a population immunity landscape, which describes the spatial distribution of population immunity levels in different regions and at different times; Regional immunity levels are expressed by the formula: Computational area In time immune level ;in, Indicates vaccination coverage rate. Indicates the proportion of the recovered population. This indicates the immune memory formed by historical infections. The formula represents the weighting coefficients; it provides a mathematical model for quantifying regional immunity levels, aiming to comprehensively consider the contributions of multiple immune sources to overall population immunity, thus making the calculation of immunity levels more comprehensive and accurate. Specifically, Vi(t) (vaccination coverage) can be obtained from the vaccination databases of public health institutions or disease control centers at all levels, statistically representing the proportion of the population in a specific region who have completed full vaccination or booster vaccination at a specific time point. Ri(t) (the proportion of recovered individuals) can be dynamically estimated based on the daily or weekly confirmed cases and discharged patients reported by the disease surveillance system, combined with a model of the average recovery period for the disease. InfHisti(t) (immune memory formed by historical infection) can be estimated through analysis of past epidemic data, serological survey results, and combined with an immune decay model to reflect the continued contribution of past infections to current population immunity. α, β, and γ (weighting coefficients) are parameters used to balance the relative importance of different immune sources. These coefficients can be set through expert experience, fitted by regression analysis based on historical epidemic data, or determined through machine learning optimization algorithms (such as genetic algorithms and particle swarm optimization) to ensure that the model accurately reflects the relative contributions of different immune components to the overall immunity level.
[0074] Immune memory decays over time, as shown by the formula: Describing the temporal evolution of the immune landscape; among which, Indicates the rate of immune attenuation. This indicates new immunity, namely vaccination and natural infection.
[0075] This formula aims to describe the dynamic changes in population immunity levels, considering not only the natural decline of immunity over time but also new immune acquisition (including new vaccinations and new natural recoveries). This dynamic evolution mechanism allows the immune landscape to reflect the continuous evolution of population immune status in real time and accurately, rather than a static, fixed value. δ (immune decay rate) is a key epidemiological parameter that can be estimated using data from medical studies, long-term cohort studies, or epidemiological surveys (e.g., by monitoring antibody titer decay curves over time). For example, for a specific respiratory virus, data on its antibody half-life or period of immune protection can be used to derive the corresponding immune decay rate. The ΔImmi(t) (new immunity) component is dynamically updated and is calculated by the contribution of new vaccinations and new natural recoveries to the immune level within a specific time period Δt. For example, the number of new vaccinations can be multiplied by the vaccine efficacy rate, and the number of new recoveries multiplied by the immune efficacy rate generated by natural infection; then, the two can be summed as the contribution of new immunity.
[0076] As an implementation approach, the immune landscape and transmission dynamics are bidirectionally coupled: the immune landscape influences the transmission rate, and transmission dynamics alter the distribution of the immune landscape. This is one of the core mechanisms of this method, designed to ensure a realistic interaction and feedback between population immune status and disease transmission. This bidirectional coupling mechanism allows the model to more accurately simulate real-world transmission processes. Specifically, changes in the immune landscape directly affect the number of susceptible individuals in the population, thus influencing the effective transmission rate of the disease; conversely, disease transmission (i.e., the occurrence of new infections) alters the distribution of the immune landscape in real time. At the implementation level, when the immune landscape shows a high level of immunity in a certain area, it means that the proportion of susceptible individuals in that area is reduced. This will be reflected in the transmission dynamics model (e.g., a variant based on the SIR model) as a decrease in the effective contact rate or basic reproduction number (R0) in that area, thereby slowing the rate of disease transmission in that area. Simultaneously, the number of infected and recovered individuals simulated and predicted by the transmission dynamics model will be fed back in real time as part of ΔImmi(t) and used to update the immune landscape. For example, when a model predicts a large-scale infection in a certain area, the proportion of recovered people in that area will increase significantly, thereby improving its overall immunity level and forming a closed-loop feedback mechanism.
[0077] A two-way coupling between immune landscape and transmission dynamics: immune landscape influences transmission rate, and transmission dynamics alter the distribution of immune landscape; The proportion of susceptible individuals is determined by the immune landscape, as shown by the formula: Computational area Susceptibility ; Regional contact intensity is obtained by weighted summation of standardized population density, place activity, and public transport activity.
[0078] The function of this formula is to directly convert the previously calculated regional immunity level Immi(t) into the proportion of susceptible individuals in that region, si(t). The proportion of susceptible individuals is a key input parameter in the transmission dynamics model, directly determining the potential for further disease spread within the population. This direct conversion ensures that the transmission dynamics model accurately reflects the impact of the current population's immune status on disease transmission. In practical applications, simply substituting the calculated regional immunity level Immi(t) from the preceding steps into this formula yields the corresponding susceptibility si(t). Furthermore, regional contact intensity is obtained by weighted summation of standardized population density, place activity, and public transport activity. Regional contact intensity is another crucial factor influencing the rate of disease transmission; it aims to quantify the probability and frequency of effective contact between populations. By comprehensively considering these three dimensions—standardized population density, place activity, and public transport activity—it can more comprehensively and accurately reflect the contact patterns within and between regions. In the implementation process, standardized population density can be obtained from census data of various regions through Geographic Information Systems (GIS) and standardized (e.g., by dividing by the maximum population density or using Z-score standardization to make the values within a comparable range). Venue activity can be calculated using various data sources such as operator spatiotemporal signaling trajectory data, venue POI data, and social media check-in data to statistically analyze the foot traffic, dwell time, and crowd density of various venues (such as shopping malls, restaurants, parks, office buildings, etc.) within a specific area, and then aggregated using weighted averages. Public transportation activity can be calculated by obtaining passenger flow data, card swipe data, and vehicle trajectory data from public transportation operators to statistically analyze the frequency and intensity of public transportation use between or within different regions. Finally, these three indicators are assigned different weights (these weights can be optimized through expert experience, fitting historical epidemic data, or machine learning methods), and then linearly weighted summed to obtain the final regional contact intensity.
[0079] In some possible implementations, the causal intervention-based inversion calculates the causal risk contribution of each cross-regional flow path, including: A propagation causal graph is constructed based on a structural causal model, defining causal variables and causal paths in the propagation process; Causal intervention is performed using the do operator, and the average causal effect of the path is calculated using the formula: Calculation path Average causal effect ;in, This indicates a causal intervention operation. Expressing expectations, Indicates the area The risk value, Indicates from arrive The flow rate; By controlling confounding factors through backdoor and frontdoor adjustments, the unbiasedness of causal effect estimation is ensured. Through the formula: Computation source region For the target area The causal risk contribution; among which, This represents a very small constant to prevent the denominator from being zero; When the causal risk contribution exceeds the contribution threshold, the path will be... Identified as a key input propagation path; Causal betweenness centrality measures the criticality of a node in the propagation causal chain and is used to identify key propagation hubs at the network level.
[0080] As an approach, structural causal models are constructed, defining causal variables and causal paths in the transmission process. The aim is to establish explicit causal relationships between variables, rather than merely statistical correlations. Structural causal models typically consist of a set of structural equations and a causal graph. Causal variables can include inter-regional mobility, regional infection rates, immunity levels, and various prevention and control measures, while causal paths are represented by directed edges indicating how these variables interact. This model can be constructed using domain expert knowledge, epidemiological theory, or data-driven causal discovery algorithms (e.g., PC algorithm, FCI algorithm) to ensure the rationality and accuracy of causal relationships.
[0081] The core operation in causal inference is calculating the average causal effect ACEi→j(t) of a path through causal intervention using the do operator. This simulates the effect of external intervention on a specific variable. For example, do(Flow_ij=0) means forcibly setting the flow from region i to region j to zero, and then observing the risk change in target region j. The average causal effect ACEi→j(t) quantifies the average impact of this intervention on the outcome variable (such as the risk value of the target region). Its calculation requires constructing a causal model capable of simulating different flow scenarios and is achieved through simulation based on structural causal models or causal inference algorithms (such as G-computation or inverse probability weighting).
[0082] Controlling confounding factors through backdoor and frontdoor adjustments aims to eliminate biases in causal effect estimation caused by confounding variables (i.e., variables that simultaneously affect both cause and effect). Backdoor adjustment blocks confounding effects by identifying and controlling a set of variables along a "backdoor path," such as controlling for factors like population density and immunity levels when calculating the impact of mobility on risk. Frontdoor adjustment identifies causal effects by recognizing a mediating variable, even in the presence of unobserved confounding factors. These adjustment methods effectively handle confounding factors during causal inference using statistical techniques (such as regression analysis and stratified analysis), thereby ensuring the unbiasedness of causal effect estimation.
[0083] The formula Contribi→jcausal(t) is used to calculate the causal risk contribution of source region i to target region j, aiming to quantify the relative importance of a specific flow path to the total risk of the target region. This contribution provides a normalized indicator by comparing the average causal effect of a single path with the sum of the average causal effects of all paths affecting the target region. The calculation involves iterating through all source regions p that may affect target region j, calculating their respective average causal effects ACEp→j(t), and then comparing the ACE of a specific path i→j with its sum.
[0084] When the causal risk contribution exceeds a contribution threshold, path i→j is identified as a key input propagation path. This threshold can be determined based on expert experience, historical data analysis, or sensitivity analysis, for example, set as a certain percentage of the total contribution. Once the causal risk contribution of all paths is calculated, the system will automatically mark the paths that exceed this threshold; these paths are the focus of prevention and control interventions.
[0085] Causal betweenness centrality measures the criticality of a node in a propagation causal chain, used to identify key propagation hubs at the network level. Causal betweenness centrality is a metric in network analysis that measures the importance of a node as a "bridge" or "mediator" in a causal path. A higher betweenness centrality means more causal paths pass through that node, thus indicating a more crucial role in the propagation causal chain. On the constructed propagation causal graph, the causal betweenness centrality of each regional node can be calculated. Regional nodes with high betweenness centrality are identified as key propagation hubs, and intervention in them can have a significant impact on the entire propagation network.
[0086] In some possible implementations, the adoption of active Bayesian optimization to select adaptive monitoring points includes: Gaussian processes are used to model the spatial distribution of risk, and the risk values of each region are modeled as samples of Gaussian processes. The prior distribution of a Gaussian process is defined by the mean function and the covariance function, and is defined by the following formula: ;in, Represents the mean function, This represents the covariance function, also known as the kernel function. A Gaussian process is a nonparametric Bayesian model capable of modeling and predicting unknown functions. In this case, it is used to infer the risk distribution across an entire spatial region, including the mean and uncertainty of risks in unmonitored areas, thereby capturing the spatial correlation of risks. The prior distribution of a Gaussian process is defined by the mean function m(x) and the covariance function k(x,x′), where m(x) represents the expected average behavior of risk values and can be a constant or a more complex function, while k(x,x′) defines the correlation of risk values between any two regions x and x′. For example, a squared exponential kernel or a Matérn kernel can be used to quantify this spatial correlation.
[0087] Based on existing monitoring point data, a posterior estimate of the risk distribution is obtained through Bayesian inference. Bayesian inference combines the prior distribution of the Gaussian process with actual observation data to generate a posterior distribution. This posterior distribution provides not only the predicted mean of risk values for unmonitored areas but also the uncertainty (variance) of these predictions. The next optimal monitoring point is selected using an acquisition function to maximize information gain.
[0088] The next optimal monitoring point is selected by acquiring a function to maximize information gain. The acquisition function is calculated using the following formula: Calculate; where, Indicates position The posterior mean of risk, Indicates position The posterior standard deviation of risk, This indicates an exploration using equilibrium parameters.
[0089] This acquisition function is a form of Upper Confidence Bound (UCB), which balances "exploration" (selecting areas with high uncertainty) and "exploitation" (selecting areas with high mean risk). Using the Upper Confidence Bound acquisition function achieves a balance between areas of high uncertainty and high risk, allowing monitoring resources to be directed to areas with both high risk and high uncertainty.
[0090] The upper confidence bound function is used to achieve a balance between regions of high uncertainty and regions of high risk.
[0091] By using Bayesian optimization to iteratively select monitoring points, the accuracy of risk assessment can be maximized under the constraint of monitoring cost.
[0092] Bayesian optimization is a sequential decision-making process that iteratively optimizes the objective function by selecting monitoring points. Each time a new monitoring point is selected, the Gaussian process model is updated, and the acquisition function is recalculated to guide the next selection. Uncertainty is decomposed into cognitive uncertainty and random uncertainty, corresponding to model parameter uncertainty and inherent data noise, respectively. Cognitive uncertainty stems from imperfections in the model itself or insufficient parameter estimation, which can be reduced by collecting more data; random uncertainty arises from the inherent randomness of the data or measurement noise. This decomposition helps to more comprehensively understand the reliability of risk assessment and guide subsequent decisions, such as prioritizing areas with high cognitive uncertainty to obtain maximum information gain.
[0093] Uncertainty is decomposed into two parts: cognitive uncertainty and accidental uncertainty, which correspond to model parameter uncertainty and inherent noise in the data, respectively.
[0094] In some possible implementations, the method further includes interpretable propagation reasoning and prevention strategy generation based on neural symbol fusion, specifically: A neural symbolic fusion reasoning framework is constructed, with the bottom layer being a neural network perception layer and the top layer being a symbolic logic reasoning layer.
[0095] The neural network layer is responsible for extracting high-dimensional propagation features from multi-source spatiotemporal data, such as the potential propagation intensity between regions and the risk factors of crowd gathering events. The symbolic logic layer, based on predefined epidemiological knowledge, expert experience, and causal rules, performs logical reasoning on the features extracted by the neural network layer to form an interpretable causal chain of propagation and risk assessment. This hierarchical design ensures that the system can handle complex data while providing a clear reasoning process.
[0096] The neural network layer extracts high-dimensional propagation features, while the symbolic logic layer performs interpretable reasoning based on epidemiological knowledge.
[0097] The continuous features of a neural network are mapped to discrete propositions of symbolic logic through a neural symbolic interface.
[0098] The risk index of a region can be converted into discrete propositions such as "high risk," "medium risk," and "low risk" through threshold judgment, or the causal contribution of a flow path can be converted into a "critical path" or "non-critical path." This mapping mechanism ensures the effective transmission and transformation of information between different levels, enabling the symbolic logic layer to reason based on the perception results of the neural network. Based on this, this method constructs an interpretable rule set for the propagation causal chain. Each rule includes preconditions and conclusions, supporting understanding and verification by human experts. These rules are presented in a human-readable form, allowing epidemiologists and policymakers to intuitively understand the model's reasoning process, verify its rationality, and adjust or supplement it according to the actual situation, thereby enhancing the model's transparency and credibility.
[0099] Construct an interpretable rule set for propagating causal chains, with each rule containing preconditions and conclusions, supporting human experts' understanding and verification; The objective function for optimizing the prevention and control strategy is defined by the following formula: ;in, Indicates prevention and control strategies, Representing the policy space, Representation Strategy Lower region exist Risk value at any given moment Representation Strategy Implementation costs Representation Strategy The risks and uncertainties below The weighting coefficients represent the weighting factors. This function aims to quantify the advantages and disadvantages of different prevention and control strategies a, comprehensively considering the risk value of each region j in the future time t+Δt after the implementation of strategy a, the implementation cost of strategy a, and the risk uncertainty under strategy a. By minimizing this objective function, the system can achieve a balance between reducing transmission risk, controlling implementation costs, and reducing uncertainty. The weighting coefficients λ1, λ2, and λ3 allow for adjustment of the relative importance of each optimization objective according to actual needs.
[0100] A deep reinforcement learning algorithm is employed to solve for the optimal prevention and control strategy, which is then validated and iteratively optimized within a digital twin. The deep reinforcement learning agent interacts with the digital twin environment to learn which prevention and control actions, under different epidemic conditions, will maximize long-term benefits (i.e., minimize the aforementioned objective function). The digital twin, as a high-fidelity simulation environment, can reflect the real-time impact of different strategies on transmission risk, population movement, and the immune landscape, thus providing feedback to the reinforcement learning agent. Through extensive strategy validation and iterative optimization within the digital twin, optimal or suboptimal prevention and control strategies that might be difficult to test directly in the real world can be discovered and refined, effectively avoiding the trial-and-error costs and risks inherent in the real world.
[0101] The system outputs differentiated prevention and control recommendations, resource allocation plans, and a report explaining the causal chain of transmission. Optimized prevention and control strategies will be output as differentiated recommendations, such as different isolation, testing, vaccination, or traffic control measures for different regions or populations. Simultaneously, the system will also output corresponding resource allocation plans to guide the rational allocation of medical supplies, human resources, and testing capabilities. Furthermore, to enhance the transparency and credibility of decision-making, the system will generate a report explaining the causal chain of transmission, detailing how the model derives risk assessment results and prevention and control recommendations from raw data, including key transmission paths, influencing factors, and reasoning logic, thereby helping decision-makers fully understand and adopt the recommendations. To ensure the model's accuracy and generalization ability, regularization optimization methods are used for model parameter calibration. Regularization helps prevent the model from overfitting to historical data and improves its predictive ability for future epidemics. The calibrated model undergoes rigorous predictive performance validation using historical epidemic data, employing a series of quantitative indicators to evaluate its performance, including RMSE, Spearman correlation coefficient, AUC, and Top-K hit rate. These validation indicators collectively ensure the model's reliability and practicality.
[0102] The model parameters were calibrated using a regularization optimization method, and the model's predictive performance was verified using historical epidemic data. The verification metrics included RMSE, Spearman correlation coefficient, AUC, and Top-K hit rate.
[0103] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins, characterized in that, The method includes the following steps: Collect multi-source spatiotemporal data, including operator spatiotemporal signaling trajectory data, disease monitoring time series data, macro-population flow network data, building wind field data, and location POI data; The operator's spatiotemporal signaling trajectory data records the location and time information of the mobile terminal connecting to the base station at different time periods; The spatiotemporal signaling trajectory data is super-resolution reconstructed and privacy-desensitized using a conditional diffusion model to construct an individual mobile digital twin, identify stop points, and extract the microscopic OD matrix. The individual mobile digital twin simulates the spatiotemporal movement behavior pattern of each individual; A high-order hypergraph aggregation and propagation network is constructed. Based on the credibility-real-time performance-coverage joint weight mechanism, the micro-signaling flow network and the macro-population flow network are adaptively fused at multiple scales to generate a fused propagation hypernetwork. The hypergraph high-order aggregation propagation network uses regions as nodes and multi-group aggregation events as hyperedges to characterize ternary and higher-order propagation relationships. The fused propagation supernetwork is input into the spatiotemporal graph neural ODE propagation model, combined with computational fluid dynamics aerosol physics propagation modeling, to perform continuous-time multi-scale propagation dynamics deduction, and calculate the input propagation risk, output propagation risk and comprehensive propagation risk index of each region; Based on the theory of self-organized criticality, the phase transition critical point of the propagation system is predicted. The causal risk contribution of each cross-regional flow path is calculated through causal intervention-based inversion, and key propagation paths are identified. Active Bayesian optimization is used to select adaptive monitoring points, combined with an immune landscape spatiotemporal evolution model, to output risk confidence intervals, phase transition early warning signals, and dynamic risk maps.
2. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The step of constructing an individual mobile digital twin by performing super-resolution reconstruction and privacy desensitization on the spatiotemporal signaling trajectory data using a conditional diffusion model includes: Through the formula: Calculate the same user The Article and No. Instantaneous movement speed between continuous trajectory records ;in, Indicates the first Track position With the Track position Spatial distance between them Indicates the first The timestamp of each trajectory record; When the instantaneous moving speed exceeds the preset maximum speed threshold, a conditional diffusion model is used to perform generative super-resolution completion of the missing trajectory segment; The conditional diffusion model uses known trajectory points as conditions to generate high-resolution intermediate trajectory points through stepwise denoising. The forward diffusion process in the conditional diffusion model is defined by the following formula: in, Indicates the first The trajectory after adding noise, Indicates the first noise variance of the step Indicates a normal distribution. Represents the identity matrix; The reverse generation process uses a neural network to predict noise and gradually remove it to reconstruct a high-resolution trajectory sequence. Through the formula: The user identifier is irreversibly de-identified using salted hashing to generate a de-identified user ID; where, Indicates the original user ID. This indicates the salting parameters. This represents an irreversible hash function; By mapping base station locations to standard spatial grid cells, an individual mobile digital twin is constructed, which includes the individual's spatiotemporal trajectory, stay preferences, travel patterns, and activity modes.
3. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins as described in claim 1, characterized in that, The construction of the hypergraph high-order aggregation propagation network and multi-scale adaptive fusion includes: Through the formula: Identify users Set of effective dwelling points ;in, Indicates the center location of the stop. Indicates the time of entry into the stop. Indicates the time of departure from the stop. Indicates the spatial radius threshold. Indicates the time threshold; Individual travel chains are extracted based on the continuous stop-point sequence, generating a micro-OD matrix with sample expansion weights, using the formula: Calculation time window Inner from the region Flow to area microfluidity ;in, Indicates an indicator function, Indicates user In the time window The origin and destination areas of the trip. Indicates user Sample expansion weights; Constructing a high-order aggregation and propagation network of a hypergraph ,in Represents a set of region nodes. Denotes the set of superedges. This represents the hyperedge weight matrix, which stores the weight value of each hyperedge. Each hyperedge connects three or more regional nodes that stay in the same location within the same time window, thus depicting higher-order propagation relationships. The hyperedge weight is determined by the formula: Calculate the first Superedge in the time window weight ;in, Indicates the first The set of users corresponding to each hyperedge Indicates user Infection pressure in the source area Indicates user The length of time spent at this location, This indicates the population density of the location; The edge weights of the micro-flow network and the macro-flow network are normalized respectively. The fusion weights are calculated through a joint weighting mechanism of credibility, real-time performance and coverage to generate a fusion propagation supernetwork.
4. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The computational fluid dynamics aerosol physics propagation modeling includes: Based on building wind field data and topographic data, airflow is simulated using the Reynolds-averaged Navier-Stokes equations, as shown in the formula: Calculate the velocity distribution in the flow field; where, Indicates air density, express The velocity component in the direction, Indicates pressure, Represents the viscous stress tensor. Represents the components of gravitational acceleration; Based on computational fluid dynamics flow fields, the diffusion process of viral aerosols is simulated using the convection-diffusion equation: Calculate the viral concentration distribution; where, Indicates the concentration of viral aerosols. Indicates fluid velocity. Indicates the diffusion coefficient. Indicates the virus source item; The viral source is determined by the infected person's respiratory rate, viral load, and the correction factor for wearing a mask. Individual exposure dose is determined by the formula: Calculate users viral exposure dose ;in, Indicates user At the venue The duration of stay Indicates user In time Location, Indicates user respiratory rate; The probability of infection is calculated using the following dose-response model: ; in, This indicates the viral infectivity coefficient.
5. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The spatiotemporal graph neural ODE propagation model is coupled with multi-scale dynamics, including: Construct a hypergraph neural network layer, extract high-order propagation features through hypergraph convolution, using the formula: Hypergraph convolution feature extraction is performed; among which, Represents the node at level l Hidden features, Indicates containing nodes The set of superedges Indicates the superedge The number of nodes, Represents a node The degree, Represents the l-th layer hyperedge The weight, Indicates the activation function; Using the initial features output by the hypergraph neural network as initial values, continuous-time evolution is performed through the neural network's ordinary differential equations, defined by the following formula: Define the time derivative of the hidden state; where, Indicates time The hidden state, Indicates by parameters Parameterized neural network functions; By adaptive step size The solver integrates from the initial time to the target time to obtain the hidden state at any time point; it realizes multi-scale coupling of the micro-individual layer, the meso-community layer and the macro-region layer, and the information exchange between the scales is carried out through upsampling and downsampling to form a hierarchical propagation dynamic system.
6. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 5, characterized in that, The prediction of phase transition critical points in a propagation system based on the self-organizing criticality theory includes: The order parameters of the propagation system are calculated using the formula: Calculation time average infection rate As an order parameter; where... Indicates the total number of regions. Indicates the region The infection rate; The fluctuations of the order parameter are calculated using the formula: Calculate susceptibility ;in, This represents the variance of the infection rate in each region. This represents the average infection rate; When the system approaches the phase transition critical point, the susceptibility exhibits a divergent peak, as shown by the formula: Describe the critical divergence behavior; where, Indicates the propagation control parameters, Indicates the critical point. Indicates the critical index; By using finite-size scaling analysis and the renormalization group method, the phase transition critical point of the propagation system from the controllable state to the explosive state is predicted, and a phase transition early warning signal is output. The self-organized criticality is verified by avalanche size distribution analysis. When the avalanche size follows a power-law distribution, the system is determined to be in a critical state.
7. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The spatiotemporal evolution and propagation dynamics of the immune landscape are bidirectionally coupled, including: Construct a population immunity landscape, which describes the spatial distribution of population immunity levels in different regions and at different times; Regional immunity levels are expressed by the formula: Computational area In time immune level ;in, Indicates vaccination coverage rate. Indicates the proportion of the recovered population. This indicates the immune memory formed by historical infections. Indicates the weighting coefficient; Immune memory decays over time, as shown by the formula: Describing the temporal evolution of the immune landscape; among which, Indicates the rate of immune attenuation. This indicates new immunity, namely vaccination and natural infection; A two-way coupling between immune landscape and transmission dynamics: immune landscape influences transmission rate, and transmission dynamics alter the distribution of immune landscape; The proportion of susceptible individuals is determined by the immune landscape, as shown by the formula: Computational area Susceptibility ; Regional contact intensity is obtained by weighted summation of standardized population density, place activity, and public transport activity.
8. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The causal intervention-based inversion calculation of the causal risk contribution of each cross-regional flow path includes: A propagation causal graph is constructed based on a structural causal model, defining causal variables and causal paths in the propagation process; Causal intervention is performed using the do operator, and the average causal effect of the path is calculated using the formula: Calculation path Average causal effect ;in, This indicates a causal intervention operation. Expressing expectations, Indicates the region The risk value, Indicates from arrive The flow rate; By controlling confounding factors through backdoor and frontdoor adjustments, the unbiasedness of causal effect estimation is ensured. Through the formula: Computation source region For the target area The causal risk contribution; among which, This represents a very small constant to prevent the denominator from being zero; When the causal risk contribution exceeds the contribution threshold, the path will be... Identified as a key input propagation path; Causal betweenness centrality measures the criticality of a node in the propagation causal chain and is used to identify key propagation hubs at the network level.
9. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The method of using active Bayesian optimization to select adaptive monitoring points includes: Gaussian processes are used to model the spatial distribution of risk, and the risk values of each region are modeled as samples of Gaussian processes. The prior distribution of a Gaussian process is defined by the mean function and the covariance function, and is defined by the following formula: ;in, Represents the mean function, This represents the covariance function, also known as the kernel function. Based on existing monitoring point data, a posterior estimate of the risk distribution is obtained through Bayesian inference. The next optimal monitoring point is selected by acquiring a function to maximize information gain. The acquisition function is calculated using the following formula: Calculate; where, Indicates position The posterior mean of risk, Indicates position The posterior standard deviation of risk, This indicates an exploration using equilibrium parameters; An upper confidence bound acquisition function is used to achieve a balance between regions of high uncertainty and high risk. By using Bayesian optimization and iterative selection of monitoring points, the accuracy of risk assessment can be maximized under the constraint of monitoring cost. Uncertainty is decomposed into two parts: cognitive uncertainty and accidental uncertainty, which correspond to model parameter uncertainty and inherent noise in the data, respectively.
10. The method for assessing the risk of cross-regional transmission of respiratory infectious diseases by integrating spatiotemporal signaling digital twins according to claim 1, characterized in that, The method also includes interpretable propagation reasoning and prevention strategy generation via neural symbol fusion, specifically: A neural symbolic fusion reasoning framework is constructed, with a neural network perception layer at the bottom and a symbolic logic reasoning layer at the top; The neural network layer extracts high-dimensional propagation features, and the symbolic logic layer performs interpretable reasoning based on epidemiological knowledge. The continuous features of a neural network are mapped to discrete propositions of symbolic logic through a neural symbolic interface. Construct an interpretable rule set for propagating causal chains, with each rule containing preconditions and conclusions, supporting human experts' understanding and verification; The objective function for optimizing the prevention and control strategy is defined by the following formula: ;in, Indicates prevention and control strategies, Representing the policy space, Representation Strategy Lower region exist Risk value at any given moment Representation Strategy Implementation costs Representation Strategy The risks and uncertainties below Indicates the weighting coefficient; The optimal prevention and control strategy is solved using a deep reinforcement learning algorithm, and the strategy is verified and iteratively optimized in a digital twin. Output differentiated prevention and control recommendations, resource allocation plans, and reports explaining the causal chain of transmission; The model parameters were calibrated using a regularization optimization method. The model's predictive performance was validated using historical epidemic data. Validation metrics included RMSE, Spearman correlation coefficient, AUC, and Top-K hit rate.