Urban infectious disease dynamic intervention strategy optimization method based on reinforcement learning

By constructing a dynamic intervention strategy optimization method for urban infectious diseases, and combining reinforcement learning and cost-benefit evaluation, this method solves the problems of unrealistic environmental characterization and static fixed strategies in infectious disease prevention and control. It achieves a dynamic balance between health benefits and intervention costs during the evolution of the epidemic and generates an adaptive optimal intervention strategy.

CN122369998APending Publication Date: 2026-07-10FUDAN UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUDAN UNIVERSITY
Filing Date
2026-03-27
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing research on urban infectious disease prevention and control suffers from unrealistic depictions of the transmission environment, inconsistent cost-benefit objectives, static and fixed intervention strategies, and insufficient utilization of regional spatial dependence, making it difficult to achieve a dynamic balance between health benefits and implementation costs at different stages and in different regions.

Method used

A reinforcement learning-based method for optimizing dynamic intervention strategies for urban infectious diseases is proposed. This method constructs a simulation environment for urban spatial unit-level transmission and intervention, reconstructs individual spatiotemporal trajectories and temporal contact networks, establishes an event-driven transmission model, and builds a cost-benefit evaluation index system for incremental health benefits and intervention costs. By combining regional node state characteristics and graph state representation, the optimal dynamic intervention strategy is generated.

Benefits of technology

It has enabled regional-level dynamic intervention strategies that can be adaptively adjusted during the evolution of the epidemic, balancing health benefits and intervention costs, and generating better comprehensive prevention and control results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369998A_ABST
    Figure CN122369998A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of complex network propagation control technology, specifically a method for optimizing dynamic intervention strategies for urban infectious diseases based on reinforcement learning. The invention includes: reconstructing individual spatiotemporal trajectories and time-series contact networks based on urban signaling data; establishing an event-driven SEIRHD propagation simulation environment embedding non-pharmaceutical intervention mechanisms; constructing incremental intervention cost, incremental health benefit, and net health gain indicators, and using these to establish a reinforcement learning reward function and long-term optimization objective; defining a regional-level dynamic intervention action space, and constructing a graph state representation by combining regional state characteristics and population flow relationships; and generating dynamic intervention strategies throughout the entire propagation time domain through interactive training between the reinforcement learning agent and the propagation environment. This invention can achieve dynamic optimization of urban infectious disease intervention strategies while balancing health benefits and implementation costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of complex network propagation control technology, specifically relating to a method for optimizing dynamic intervention strategies for urban infectious diseases based on reinforcement learning. Background Technology

[0002] The spread of infectious diseases in cities is influenced by factors such as high-density population gathering, frequent inter-regional movement, and continuous changes in contact structure, and has significant spatiotemporal heterogeneity and dynamic evolution characteristics [1]. In particular, in large cities, there are significant differences in population size, activity intensity, and movement patterns between different regions, which makes the spread of the epidemic not uniform, but rather a complex spatial diffusion process along the population flow network. Therefore, research on the prevention and control of infectious diseases in cities cannot simply adopt the assumption of uniform mixing, but needs to combine real population flow data with spatial structure to construct a more refined transmission model.

[0003] In situations where drug reserves are insufficient, vaccines are not yet widely available, or in the early stages of an emerging infectious disease outbreak, non-pharmaceutical interventions such as movement restrictions, testing and isolation, and personal protective equipment remain crucial control measures. Existing research typically evaluates different intervention strategies based on epidemiological indicators such as the number of infections, peak load, and transmission rate. However, relying solely on the effectiveness of transmission control as the basis for decision-making often fails to simultaneously reflect the trade-off between health benefits and intervention implementation costs, thus providing insufficient and comparable policy support for public health management departments.

[0004] Therefore, cost-benefit analysis methods from health economics have been gradually introduced into the assessment of infectious disease interventions[2]. By quantifying the changes in health losses and the costs of intervention implementation corresponding to intervention strategies, the health benefits and economic costs of different intervention strategies can be measured within a unified framework. However, most existing related studies still remain at the level of comparing pre-set static intervention strategies, that is, assuming that the intensity, scope, and duration of intervention remain fixed throughout the entire transmission period. Although such strategies are convenient for overall evaluation, they are difficult to adapt to the ever-changing transmission risks of urban epidemics at different stages and in different regions.

[0005] For the problem of optimizing dynamic intervention for urban infectious diseases, traditional rule-based decision-making methods usually rely on local observation information to determine intervention actions. Due to the high-dimensional state space, strong nonlinear propagation dynamics, and long-term decision chains in the urban transmission environment, the above methods often fail to obtain globally optimal dynamic intervention schemes. Although reinforcement learning provides a new technical path for such problems [3], existing research on infectious disease intervention based on reinforcement learning still generally suffers from problems such as oversimplification of the transmission environment, lack of health economics explanation for the optimization objective, and insufficient utilization of inter-regional spatial dependencies. Therefore, there is an urgent need for a method that can combine cost-benefit evaluation and utilize regional spatial structure information to optimize dynamic intervention strategies in a real urban transmission environment. Summary of the Invention

[0006] The purpose of this invention is to address the problems existing in current urban infectious disease intervention research, such as unrealistic depiction of the transmission environment, inconsistent cost-benefit objectives, static and fixed intervention strategies, and insufficient utilization of regional spatial dependence. It proposes a method for optimizing dynamic intervention strategies for urban infectious diseases based on reinforcement learning, so as to achieve a dynamic balance between health benefits and implementation costs in urban infectious disease prevention and control strategies.

[0007] The proposed method for optimizing dynamic intervention strategies for urban infectious diseases based on reinforcement learning, as described in this invention, has the following overall process: Figure 1 As shown, the process includes: constructing a city spatial unit-level propagation and intervention simulation environment based on city signaling data; reconstructing individual spatiotemporal trajectories and temporal contact networks; and establishing an event-driven propagation model embedding non-pharmaceutical intervention mechanisms. It also involves constructing a cost-benefit evaluation index system centered on incremental intervention costs, incremental health benefits, and net health gains, and further forming an immediate reward function and long-term optimization objective for reinforcement learning. Furthermore, it involves constructing a dynamic action space based on the target intervention strategy type, extracting regional node state features, and forming a graph state representation. Finally, it involves inputting the graph state representation into a reinforcement learning agent, and through environmental interaction, experience storage, and network iterative updates, generating the optimal dynamic intervention strategy for the entire propagation time domain. The specific steps are as follows:

[0008] Step 1: Construction of a Simulation Environment for Urban Infectious Disease Transmission and Intervention. This includes preprocessing urban signaling data, dividing the city into spatial units based on the study area, reconstructing individual spatiotemporal trajectories and temporal contact networks, establishing an event-driven transmission model on these contact networks, and embedding lockdown strategies, detection and isolation strategies, and individual protection strategies to form a transmission and intervention simulation environment for evaluating the effectiveness of dynamic intervention strategies; specifically:

[0009] Step 1-1: De-identify, remove missing values, clean up abnormal records and align time for signaling data, and divide the city into multiple spatial units according to the spatial range of the study area to obtain the basic regional nodes for propagation simulation and intervention decision-making;

[0010] Steps 1-2: Reconstruct the spatiotemporal activity trajectory of individuals based on the preprocessed signaling data, and construct a temporal contact network based on the spatiotemporal co-existence relationship of individuals in the same spatial unit and the same time window;

[0011] Steps 1-3: Build an event-driven SEIRHD propagation model based on the temporal contact network, classifying individual health status into susceptibility groups. Exposure ,Infect Rehabilitation Hospitalization and death The state is determined and propagated through a sequence of timestamped events, driving the state transition process.

[0012] Steps 1-4: Embed lockdown strategies, detection and isolation strategies, and individual protection strategies into the transmission model, and characterize the impact of different non-pharmaceutical intervention mechanisms on the transmission process through methods such as trajectory adjustment, contact removal, and transmission probability scaling.

[0013] Step 2: Construction of Cost-Benefit Evaluation Indicators and Optimization Objectives. This includes constructing incremental intervention cost and incremental health benefit indicators using a no-intervention scenario as the baseline scenario, and further unifying them into a net health benefit indicator. Based on this, the net health benefit is embedded into the reinforcement learning training process to construct an immediate reward function and a long-term cumulative optimization objective, in order to achieve cost-benefit optimization of the dynamic intervention problem of urban infectious diseases. Specifically:

[0014] Step 2-1: Using the no-intervention scenario as the baseline scenario, construct the incremental intervention cost index corresponding to the target intervention strategy. The incremental intervention cost includes the direct costs incurred during the intervention process and the indirect costs caused by the intervention;

[0015] Step 2-2: Using the no-intervention scenario as the baseline scenario, construct an incremental health benefit indicator expressed in quality-adjusted life years (QALY). The incremental health benefits are obtained by comparing the differences in health loss between the intervention scenario and the no-intervention scenario;

[0016] Steps 2-3: Based on the incremental intervention cost and incremental health benefits We construct net health benefit (NHB) as a unified cost-benefit evaluation indicator, expressed as:

[0017]

[0018] in, This represents the willingness-to-pay (WTP) threshold corresponding to a unit of health benefit.

[0019] Steps 2-4: Embed the phased net health gains into the reinforcement learning training process and construct the instant reward function at each decision moment as a feedback signal for strategy learning;

[0020] Steps 2-5: Define the cumulative discount return over multiple decision cycles, and construct a global optimization objective function for reinforcement learning with the goal of maximizing the long-term cumulative net health benefit.

[0021] Step 3: Construction of Dynamic Intervention Decision Space and Graph State Representation. This includes determining the corresponding adjustable intervention parameters based on the target intervention strategy type and constructing a reinforcement learning action space, mapping dynamic intervention actions to effective intervention parameters in the propagation simulation environment; simultaneously, extracting regional node state features, combining them with inter-regional population flow relationships to construct a spatial graph structure, and forming a graph state representation for reinforcement learning decision-making; specifically:

[0022] Step 3-1: Set the target intervention strategy type as any one of the following: lockdown strategy, detection and isolation strategy, and individual protection strategy, and determine the corresponding adjustable intervention parameters according to the set target intervention strategy type to construct the action space of reinforcement learning;

[0023] Step 3-2: For the division of cities Each spatial region defines a dynamic intervention action vector at each decision-making moment:

[0024]

[0025] in, Indicates the region At the moment of decision The target intervention intensity is used to characterize the proportion of inter-regional flows retained in each region or the corresponding intervention coverage.

[0026] Step 3-3: Map the action vectors to effective intervention parameters in the propagation simulation environment to update inter-regional population flow relationships or individual intervention states, and further influence contact structures and propagation processes;

[0027] Steps 3-4: Extract the state features of each regional node at the decision-making time. The state features include at least the proportion of people in each health state, historical intervention intensity, and regional population size information.

[0028] Steps 3-5: Construct a spatial graph structure based on the inter-regional population flow relationship, and encode it using a graph attention network by combining node features and edge features to form a graph state representation for reinforcement learning decision-making.

[0029] Step 4: Reinforcement Learning Algorithm Training and Intervention Strategy Generation. This includes initializing the reinforcement learning agent for dynamic intervention optimization, inputting graph state representations into the agent to generate continuous intervention actions at the region node level, and interacting with the propagation simulation environment; based on the environment interaction, model training is completed through experience replay and iterative updates of network parameters, ultimately outputting the optimal dynamic intervention action sequence throughout the entire propagation time domain; specifically:

[0030] Step 4-1: Initialize the reinforcement learning agent for dynamic intervention optimization, including the Actor network, Critic network, target Critic network, and experience replay pool;

[0031] Step 4-2: Input the graph state representation at the current decision moment into the Actor network to generate intervention actions for each region node. ;

[0032] Step 4-3: Generate the intervention actions The input propagation and intervention simulation environment is used to interact with the environment, obtain the next time-map state and immediate reward, and store the obtained transfer samples into the experience replay pool;

[0033] Step 4-4: Sample a small batch of samples from the experience replay pool, construct the target Q value using the target network, and update the Critic network parameters;

[0034] Steps 4-5: Based on the evaluation results of the state-action value by the Critic network, update the Actor network parameters to improve the long-term cumulative return of the strategy and maintain the action exploration capability;

[0035] Steps 4-6: Perform soft updates on the target Critic network and iteratively execute the graph state encoding, action generation, environment interaction, experience storage, and network update processes until the preset number of training rounds is reached;

[0036] Steps 4-7: After training is completed, the current urban propagation environment state is input into the trained reinforcement learning model. The Actor network outputs the dynamic intervention action sequence of each regional node at the current and subsequent decision times, thereby obtaining the optimal dynamic intervention strategy in the entire propagation time domain.

[0037] This invention uses a spatiotemporal contact network constructed from urban signaling data and an event-driven propagation simulation model as its environmental foundation. It combines incremental intervention costs, incremental health benefits, and net health gains to construct a unified cost-benefit optimization objective. Furthermore, it integrates inter-regional population flow relationships, regional epidemic status characteristics, and graph structure representations to establish a reinforcement learning dynamic decision-making framework for heterogeneous urban spatial propagation processes. This framework generates regional-level dynamic intervention strategies that can adaptively adjust as the epidemic evolves, thereby achieving a better balance between health benefits and intervention costs. Attached Figure Description

[0038] Figure 1 This is a flowchart of the algorithm framework of the present invention.

[0039] Figure 2 The diagram shows the effect verification of the method of the present invention in the embodiments, where a is the curve of change of dynamic intervention intensity, and b is the curve of change of net health benefit of the method of the present invention and each comparison strategy. Detailed Implementation

[0040] To make the above-mentioned objectives and innovations of the present invention easier to understand, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Step 1: Construction of a simulation environment for urban infectious disease transmission and intervention; the specific process is as follows:

[0042] Step 1-1: Signaling Data Preprocessing and Urban Spatial Unit Division. Input signaling data, including anonymous user identifiers, timestamps, and location identifiers. The raw signaling data undergoes de-identification, missing value removal, outlier cleaning, and time alignment to obtain a sequence of individual location records arranged chronologically. Based on the spatial extent of the study area, the city is divided into... Each spatial unit serves as a basic regional node for subsequent propagation simulation and dynamic intervention decision-making. This spatial unit can be determined using methods such as regular grid division, administrative region division, or functional zone division. This embodiment uses Shanghai, China as the study area, and divides the study area into grids to obtain urban spatial units for subsequent propagation modeling.

[0043] Steps 1-2: Individual Spatiotemporal Trajectory Reconstruction and Temporal Contact Network Construction. Based on the preprocessed signaling data, the location records of each anonymous user on consecutive time slices are encoded into a spatial location sequence arranged at hourly resolution, thereby reconstructing the individual spatiotemporal activity trajectory. Based on the reconstructed grid-level trajectory, potential contact relationships are identified using a spatiotemporal coexistence approach: when two individuals are in the same spatial cell and within the same hourly window, they are considered to have formed a potential contact pair; correspondingly, all potential contact pairs on all time slices are constructed into a temporal contact network.

[0044] Steps 1-3: Constructing an event-driven SEIRHD propagation model. Based on the temporal contact network, a SEIRHD propagation model is established. Individual health status includes susceptibility status. Exposure status Infection status Recovery status Hospitalization status and death status When a susceptible individual has effective contact with an infected individual, the probability of transmission per unit of contact is... Infection occurs, and immediately by Transfer After the incubation period, exposed individuals are transferred to... Infected individuals then transition to or The hospitalized individual was further transferred to or The propagation process is implemented using an event-driven framework, which represents infection events and various state transition events as timestamped event sequences, and executes them one by one on the global timeline in chronological order of occurrence to update the state of the corresponding individuals;

[0045] Steps 1-4: Embedding Non-Pharmaceutical Intervention Strategies. Three types of intervention strategies are embedded into the above transmission model: lockdown strategies, detection and isolation strategies, and individual protection strategies. Lockdown strategies reduce inter-regional movement by modifying individual movement trajectories; the reduced inter-regional movement is redirected within the residential grid. Detection and isolation strategies identify individuals in an exposed state. or infected status Individuals who test positive are identified and placed under isolation management after a predetermined time lag following detection. Once isolated, individuals are removed from the contact network and no longer participate in transmission until they recover or progress to hospitalization. Individual protection strategies for individuals taking protective measures aim to reduce the probability of transmission per unit of contact. The scale is adjusted, but the remaining individuals still participate in the propagation according to the original propagation probability.

[0046] Step 2: Construction of cost-benefit evaluation indicators and optimization objectives; the specific process is as follows:

[0047] Step 2-1: Construction of Incremental Intervention Cost Indicators. Using the no-intervention scenario as the baseline scenario, construct incremental cost indicators corresponding to non-pharmacological intervention strategies. Incremental costs are estimated from a socioeconomic perspective, including both direct costs incurred during the intervention's implementation and indirect costs resulting from it. Direct costs include testing costs, isolation costs, personal protective equipment costs, and hospitalization costs; indirect costs include productivity losses due to lockdowns or isolation. For a given intervention strategy... Let its total cost be... The total cost under the baseline scenario without intervention is The incremental intervention cost is then expressed as:

[0048]

[0049] Step 2-2: Construction of Health Benefit Indicators. Using the no-intervention scenario as the baseline scenario, incremental health benefit indicators expressed in quality-adjusted life years (QALY) are constructed. Health benefits are obtained by comparing the differences in health losses between the intervention scenario and the no-intervention scenario, which include years lived with disability (YLD) and years of life lost (YLL). For a given intervention strategy... Let their corresponding YLD and YLL be respectively and The corresponding health losses in the no-intervention scenario are as follows: and The incremental health benefits are expressed as:

[0050]

[0051] Steps 2-3: Constructing Net Health Benefit. To further transform the cost-benefit evaluation indicators into a unified objective function that can be used for dynamic optimization, incremental health benefits and incremental costs are combined into net health benefit (NHB):

[0052]

[0053] in, This represents the willingness-to-pay (WTP) threshold, representing the cost-to-health benefit per unit of health benefit, used to convert economic costs into a scale comparable to health benefits. When... The larger the value, the better the intervention strategy is in balancing health benefits and intervention costs;

[0054] Steps 2-4: Reward Function Construction. To embed net health gains into the reinforcement learning training process, decision points are constructed based on the phased net health gains. The immediate reward function. Recording the decision moment. The environmental state is The target intervention action parameters are The reward function is then defined as:

[0055]

[0056] in, This indicates the target intervention strategy at the decision time relative to the baseline scenario without intervention. The incremental health benefits brought about within the corresponding period. This represents the incremental intervention cost generated by the target intervention strategy within the same decision-making cycle. The reward function is used to characterize the immediate health and economic benefits generated by the target intervention strategy within a single decision-making cycle and serves as a direct feedback signal for subsequent strategy learning.

[0057] Steps 2-5: Optimize Objective Construction. The reinforcement learning agent optimizes by comprehensively considering the long-term cumulative effects over multiple decision cycles. Let the policy function be... The discount factor is , The duration of the propagation simulation is indicated by the decision-making moment. The initial cumulative return on discounts is defined as:

[0058]

[0059] Substituting further into the reward function expression, we obtain the optimization objective as:

[0060]

[0061] in, This represents the optimal dynamic intervention strategy. By maximizing the expected discount cumulative return, the cumulative net health benefit over the entire intervention time domain is optimized, thereby ensuring that the strategy learning objective of reinforcement learning is consistent with the cost-effectiveness optimization objective of the dynamic intervention problem for infectious diseases.

[0062] Step 3: Construction of dynamic intervention decision space and graph state representation; the specific process is as follows:

[0063] Step 3-1: Setting the Target Intervention Strategy Type. The target intervention strategy type is any one of the lockdown strategy, detection and isolation strategy, and personal protective strategy described in Steps 1-4. Based on the set target intervention type, determine the corresponding adjustable intervention parameters in the transmission simulation environment, and construct the action space for reinforcement learning accordingly.

[0064] The following section uses lockdown strategies as the target intervention type to illustrate how to construct the action space and state representation. For detection and isolation strategies and personal protection strategies, the same framework can be used to define the coverage of detection and isolation or personal protection as the corresponding action parameters.

[0065] Step 3-2: Define the space for dynamic intervention actions. Assume the city is divided into... A spatial region, at the moment of decision-making Define the baseline urban area flow matrix as follows: ,in, Indicates time From the region Flow to the region The population proportion. To characterize the moderating effect of dynamic lockdown strategies on inter-regional movement, a decision time is defined. The action vector is

[0066]

[0067] in Indicates the region Residents at all times The permitted percentage of cross-regional flows. When When, it indicates that residents of that area are not allowed to move across areas; when When this occurs, it indicates that the residents of that area maintain their original level of inter-regional mobility. The action vector... According to fixed decision intervals The update is performed and remains unchanged between two adjacent decision points. This embodiment sets... sky;

[0068] Step 3-3: Mapping motion to simulation parameters. Given a motion vector... Then, it is mapped into an effective urban mobility matrix after dynamic intervention. The elements of the effective flow matrix are defined as follows:

[0069]

[0070] in, For the Kronecker delta function, when The value is 1 if the condition is met, and 0 otherwise. The above mapping represents: for residents of region... The proportion of individuals who originally retained inter-regional mobility was The portion of transregional flow that was suppressed They are redirected to stay within this area. Thus, dynamic lockdown measures can regulate cross-regional population flow at the regional node granularity, and further influence the contact structure and propagation process in the propagation simulation environment;

[0071] Steps 3-4: Extraction of regional node state features. At the decision-making time... The simulation environment propagates and returns the region-level aggregate state to the reinforcement learning agent. For any region node... Construct its node feature vector for:

[0072]

[0073] in, Representing regions The number of people in a vulnerable, exposed, infected, hospitalized, recovering, and deceased state. Indicates the region Total population This represents the total population of the city. The first six items are calculated by dividing by... Normalization is performed to characterize the relative severity of the epidemic within a region; Indicates the effect of the previous decision on the region. The intensity of intervention is used to provide historical control information to the agent; Used to explicitly characterize population size heterogeneity between regions;

[0074] Steps 3-5: Graph State Representation Construction. To represent the spatial propagation dependencies caused by inter-regional population flows, a spatial graph structure is constructed using regional nodes, and baseline regional flow relationships are used as edge features. For flows from regions... Pointing to area A directed edge, whose edge characteristics are defined as follows: Based on this, a graph attention network is used to encode the region state. First, a linear mapping is performed on the node feature vectors:

[0075]

[0076] in, The weight matrix is ​​learnable. Then, a graph attention mechanism incorporating edge features is introduced to focus on the nodes. and its neighboring nodes Calculate the attention coefficient:

[0077]

[0078] in, and For learnable parameters, This represents vector concatenation. Represents a node The set of neighboring nodes. Attention coefficient. Reflecting neighboring nodes For nodes The relative importance of nodes is determined, and the flow intensity between regions is explicitly considered. Furthermore, the neighborhood node information is weighted and aggregated to obtain the node... Graph embedding representation:

[0079]

[0080] in, This represents a non-linear activation function. All nodes are embedded. Together they constitute the entire system at the decision-making moment. The graph state representation is used as input to the subsequent reinforcement learning agent for dynamic intervention strategy optimization.

[0081] Step 4: Training the reinforcement learning algorithm and generating intervention strategies; the specific process is as follows:

[0082] Step 4-1: Initialization of reinforcement learning agent. Construct a reinforcement learning agent for dynamic intervention strategy optimization, including Actor network, Critic network, target Critic network and experience replay pool [4]. Actor network is used to generate continuous action vectors based on the current graph state embedding, Critic network is used to estimate the value of state-action pairs, target Critic network is used to stabilize the training process, and experience replay pool is used to store environmental interaction samples;

[0083] Step 4-2: Generation of regional node-level intervention actions. At the decision-making moment... The graph state representation is composed of all embedded region nodes. The input Actor network learns a random policy to generate continuous intervention actions.

[0084]

[0085] in, The parameters of the Actor network are represented. Indicates by parameters The determined random policy function;

[0086] Step 4-3: Environmental Interaction and Experience Sample Storage. The action vectors output from Step 4-2 are... In the propagation and intervention simulation environment constructed in step 1, the environment updates the effective region flow matrix and propagation process according to the action mapping relationship, and then returns the graph state at the next decision time. and current rewards Let the transfer sample obtained from a single environmental interaction be denoted as:

[0087]

[0088] The transferred samples are stored in the experience replay pool for subsequent network parameter updates;

[0089] Step 4-4: Critic Network Update. Randomly sample mini-batch transfer samples from the empirical replay pool and construct the target Q-value using the target network. For any sample... Its target value is defined as:

[0090] in, As a discount factor, The entropy temperature coefficient Indicates the relationship with the first A target Critic network. Based on this, the first... The loss function of a Critic network is defined as:

[0091]

[0092] in, This represents the sample distribution in the experience replay pool. The loss function is then used to... Perform gradient descent on the parameters of the two Critic networks. Update them separately;

[0093] Steps 4-5: Actor Network Update. After completing the Critic network update, update the Actor network based on the current Critic network's value assessment results, maximizing its long-term cumulative returns while maintaining its exploration capabilities. The optimization objective of the Actor network is defined as:

[0094]

[0095] The first term corresponds to the policy entropy term, which encourages action exploration, while the second term represents the value estimate of the action generated in the current state. This is achieved by minimizing the loss function. Update network parameters This allows the Actor network to gradually learn node-level dynamic intervention strategies that can bring higher net health benefits;

[0096] Steps 4-6: Target Network Soft Update and Iterative Training. To improve training stability, a soft update is performed on the target Critic network after each round of parameter updates. The soft update rule is as follows:

[0097]

[0098] in, These are soft update coefficients. The process involves executing a cyclical process of "graph state encoding—action generation—environment interaction—experience storage—network update" until the preset number of training rounds is reached.

[0099] Steps 4-7: Generation of the optimal dynamic intervention strategy. After training, the current urban propagation environment state is input into the trained reinforcement learning model, and the Actor network outputs the continuous action vectors of each regional node at the current decision time. This process is repeated at each subsequent decision point to obtain a dynamic sequence of intervention actions throughout the simulation time domain. .

[0100] To verify the effectiveness of the dynamic intervention strategy optimization method proposed in this invention, the following comparative intervention strategies were constructed:

[0101] (1) Set a no-intervention strategy as the baseline scenario, that is, do not impose movement restrictions on any area during the entire propagation process;

[0102] (2) Fixed-High Intervention Strategy: This involves imposing a fixed, high level of inter-regional flow restriction on all regions throughout the entire decision-making time domain. );

[0103] (3) Fixed-Low Intervention Strategy, which means applying a fixed low intensity of inter-regional flow restrictions to all regions throughout the entire decision-making time domain. );

[0104] (4) Threshold-triggered strategy: This strategy triggers intervention based on whether the current number of infected people in each region exceeds a preset threshold. At any decision-making moment... Record the area The number of individuals in the endoinfected state is Let the infection threshold be... The control action corresponding to this strategy is defined as follows:

[0105]

[0106] That is, when the area The number of infections exceeded the threshold If necessary, the area will be completely sealed off; otherwise, it will remain open.

[0107] (5) Risk ranking strategy: This strategy ranks regions based on their spread risk at each decision point and prioritizes intervention in high-risk areas. (Note: The last part, "region," appears to be a typo and can be left as is.) At the moment of decision Risk score Its definition is:

[0108]

[0109] in, Indicates the region Current number of infections Indicates time From the region Flow to the region The proportion of the population. For each decision moment. ,according to Sort all regions in descending order and select those with the highest risk scores. The region constitutes a high-risk area set And define the corresponding control action as:

[0110] In one embodiment, take .

[0111] Based on the above comparison methods, the method of this invention (denoted as RL) is compared with each comparison strategy under the same propagation environment. In this embodiment, the willingness-to-pay threshold WTP = 320,000 yuan / QALY is selected as the cost-benefit trade-off scenario. For each strategy, simulations are repeatedly performed under multiple sets of mutually independent propagation random seeds, and the results are averaged to improve the stability and reliability of the evaluation results.

[0112] The comparison results are as follows Figure 2 As shown, Figure 2 (a) is the dynamic intervention intensity change curve generated by the present invention, wherein the intervention intensity This represents the average intervention intensity of all regional nodes at the corresponding time. Figure 2 (b) Net health benefits for different strategies The curve represents the cumulative net health benefit from the initial time to the current time point. This result indicates that the method of this invention can adaptively and dynamically adjust the intervention intensity according to the epidemic's spread; under the same cost-benefit evaluation conditions, the net health benefit corresponding to the method of this invention is... The curve was higher than other comparison strategies, and the gap widened significantly in the middle of the propagation and remained at a better level in subsequent simulation periods.

[0113] Incremental intervention costs for each strategy Incremental health benefits and net health benefits The summary results are shown in Table 1. Table 1 further shows that the method of this invention achieved the highest net health benefit at the end of the simulation period. Its incremental health benefits It is significantly superior to other comparison strategies, while the cost of incremental intervention is lower. The method employs a fixed high-intensity strategy and a threshold-triggered strategy. In summary, this invention balances health benefits and intervention costs at different stages of transmission, achieving superior overall performance and validating its effectiveness in optimizing dynamic interventions for urban infectious diseases.

[0114] Table 1

[0115] .

[0116] References:

[0117] [1] Nouvellet P, Bhatia S, Cori A, et al. Reduction in mobility and COVID-19 transmission[J]. Nature communications, 2021, 12(1): 1090.

[0118] [2] Du Z, Pandey A, Bai Y, et al. Comparative cost-effectiveness ofSARS-CoV-2 testing strategies in the USA: a modelling study[J]. The LancetPublic Health, 2021, 6(3): e184-e191.

[0119] [3] Yao Y, Zhou H, Cao Z, et al. Optimal adaptive nonpharmaceuticalinterventions to mitigate the outbreak of respiratory infections followingthe COVID-19 pandemic: a deep reinforcement learning study in Hong Kong,China[J]. Journal of the American Medical Informatics Association, 2023, 30(9): 1543-1551.

[0120] [4] Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor[C]. International conference on machine learning, 2018: 1861-1870。

Claims

1. A method for optimizing dynamic intervention strategies for urban infectious diseases based on reinforcement learning, characterized in that, include: Based on urban signaling data, individual spatiotemporal trajectories and time-series contact networks are reconstructed, and an event-driven SEIRHD propagation simulation environment embedding non-pharmaceutical intervention mechanisms is established. Incremental intervention costs, incremental health benefits, and net health gains indicators are constructed, and a reinforcement learning reward function and long-term optimization objective are established accordingly. A regional-level dynamic intervention action space is defined, and a graph state representation is constructed by combining regional state characteristics with population flow relationships. Through interactive training between reinforcement learning agents and the propagation environment, dynamic intervention strategies are generated throughout the entire propagation time domain. The specific steps are as follows: Step 1: Construct a simulation environment for the spread and intervention of urban infectious diseases, including preprocessing urban signaling data, dividing urban spatial units according to the study area, reconstructing individual spatiotemporal trajectories and temporal contact networks, establishing an event-driven propagation model on the contact network, and embedding lockdown strategies, detection and isolation strategies, and individual protection strategies, thereby forming a simulation environment for evaluating the effectiveness of dynamic intervention strategies. Step 2: Construct cost-benefit evaluation indicators and optimization objectives, including using the no-intervention scenario as the baseline scenario, constructing incremental intervention cost and incremental health benefit indicators, and further unifying the two into a net health benefit indicator; on this basis, embed the net health benefit into the reinforcement learning training process, construct an immediate reward function and a long-term cumulative optimization objective, so as to achieve cost-benefit optimization of the dynamic intervention problem of urban infectious diseases. Step 3: Construct a dynamic intervention decision space and graph state representation, including determining the corresponding adjustable intervention parameters according to the target intervention strategy type and constructing a reinforcement learning action space, mapping dynamic intervention actions to effective intervention parameters in the propagation simulation environment; at the same time, extracting regional node state features, combining the inter-regional population flow relationship to construct a spatial graph structure, and forming a graph state representation for reinforcement learning decision-making. Step 4: Generate reinforcement learning algorithm training and intervention strategies, including initializing reinforcement learning agents for dynamic intervention optimization, inputting graph state representations into the agents to generate continuous intervention actions at the region node level, and interacting with the propagation simulation environment; based on the interaction with the environment, complete model training through experience playback and iterative updates of network parameters, and finally output the optimal dynamic intervention action sequence in the entire propagation time domain.

2. The method for optimizing dynamic intervention strategies for urban infectious diseases according to claim 1, characterized in that, Step 1 describes the construction of a simulation environment for the spread and intervention of urban infectious diseases. The specific process is as follows: Step 1-1: Signaling data preprocessing and urban spatial unit division; the signaling data includes anonymous user identifiers, timestamps, and location identifiers; the original signaling data undergoes de-identification, missing value removal, abnormal record cleaning, and time alignment processing to obtain a sequence of individual location records arranged in chronological order; based on the spatial extent of the study area, the city is divided into... Each spatial unit serves as a basic regional node for subsequent propagation simulation and dynamic intervention decision-making; the spatial unit is determined by regular grid division, administrative region division, or functional area division. Steps 1-2: Construct individual spatiotemporal trajectory reconstruction and temporal contact network; Based on the preprocessed signaling data, encode the location records of each anonymous user on continuous time slices into a spatial location sequence arranged in hourly resolution, thereby reconstructing the individual spatiotemporal activity trajectory; Based on the reconstructed grid-level trajectory, identify potential contact relationships using a spatiotemporal coexistence approach: When two individuals are in the same spatial unit and within the same hourly window, they are considered to form a potential contact pair; Correspondingly, construct a temporal contact network for all potential contact pairs on all time slices. Steps 1-3: Construct an event-driven SEIRHD propagation model; build a SEIRHD propagation model based on the temporal contact network; individual health status includes susceptibility status. Exposure status Infection status Recovery status Hospitalization status and death status When a susceptible individual has effective contact with an infected individual, the probability of transmission per unit of contact is... Infection occurs, and immediately by Transfer After the incubation period, exposed individuals are transferred to... Infected individuals then transition to or The hospitalized individual was further transferred to or ; The propagation process is implemented using an event-driven framework, which represents infection events and various state transition events as a sequence of timestamped events, and executes them one by one on the global timeline in the order of their occurrence to update the state of the corresponding individuals. Steps 1-4: Embedding non-pharmaceutical intervention strategy mechanisms; embedding three types of intervention strategies into the above transmission model: lockdown strategy, detection and isolation strategy, and individual protection strategy; the lockdown strategy reduces inter-regional flow by modifying individual movement trajectories, and the reduced inter-regional flow is redirected to the residential grid; The detection and isolation strategy identifies individuals in an exposed state. or infected status Individuals who test positive are identified and placed under isolation management after a predetermined time lag following detection; isolated individuals are removed from the contact network and no longer participate in transmission until they recover or progress to hospitalization; individual protection strategies for individuals taking protective measures reduce their probability of transmission per unit of contact. The scale is adjusted, but the remaining individuals still participate in the propagation according to the original propagation probability.

3. The method for optimizing dynamic intervention strategies for urban infectious diseases according to claim 1, characterized in that, Step 2, which involves constructing cost-benefit evaluation indicators and optimization objectives, follows this process: Step 2-1: Construct incremental intervention cost indicators; using the no-intervention scenario as the baseline scenario, construct incremental cost indicators corresponding to non-pharmacological intervention strategies. ; Incremental costs are estimated from a socioeconomic perspective, including direct costs incurred during the intervention implementation process and indirect costs caused by the intervention; direct costs include testing costs, isolation costs, personal protective equipment costs, and hospitalization costs; indirect costs include productivity losses caused by lockdowns or isolation; for a given intervention strategy Let its total cost be... The total cost under the baseline scenario without intervention is The incremental intervention cost is then expressed as: Step 2-2: Constructing Health Benefit Indicators; Using the no-intervention scenario as the baseline scenario, construct incremental health benefit indicators expressed in quality-adjusted life years (QALY). Health benefits are obtained by comparing the differences in health losses between the intervention scenario and the no-intervention scenario, whereby health losses include two components: lost disability years of survival (YLD) and lost life years (YLL); for a given intervention strategy Let their corresponding YLD and YLL be respectively and The corresponding health losses in the no-intervention scenario are as follows: and The incremental health benefits are expressed as: Steps 2-3: Constructing Net Health Benefit; To further transform the cost-benefit evaluation indicators into a unified objective function that can be used for dynamic optimization, incremental health benefits and incremental costs are combined into Net Health Benefit (NHB): in, This represents the willingness-to-pay threshold (WTP) for a unit of health benefit, used to convert economic costs into a scale comparable to health benefits; when The larger the value, the better the intervention strategy is in balancing health benefits and intervention costs; Steps 2-4: Construct the reward function; to embed net health gains into the reinforcement learning training process, construct decision points based on the phased net health gains. The instant reward function; recording the decision moment. The environmental state is The target intervention action parameters are The reward function is then defined as: in, This indicates the target intervention strategy at the decision time relative to the baseline scenario without intervention. The incremental health benefits brought about within the corresponding period. This represents the incremental intervention cost generated by the target intervention strategy within the same decision-making cycle; the reward function is used to characterize the immediate health and economic benefits generated by the target intervention strategy within a single decision-making cycle, and serves as a direct feedback signal for subsequent strategy learning. Steps 2-5: Construct the optimization objective; the reinforcement learning agent comprehensively considers the long-term cumulative effect over multiple decision cycles for optimization; let the policy function be... The discount factor is , The duration of the propagation simulation is indicated by the decision-making moment. The initial cumulative return on discounts is defined as: Substituting further into the reward function expression, we obtain the optimization objective as: in, This represents the optimal dynamic intervention strategy; by maximizing the expected discount cumulative return, it optimizes the cumulative net health benefits over the entire intervention time domain, thereby ensuring that the strategy learning objective of reinforcement learning is consistent with the cost-effectiveness optimization objective of the dynamic intervention problem of infectious diseases.

4. The method for optimizing dynamic intervention strategies for urban infectious diseases according to claim 1, characterized in that, Step 3, which involves constructing the dynamic intervention decision space and graph state representation, follows this process: Step 3-1: Set the target intervention strategy type; the target intervention strategy type is any one of the lockdown strategy, detection and isolation strategy and personal protection strategy mentioned in Step 1-4; according to the set target intervention type, determine the corresponding adjustable intervention parameters in the transmission simulation environment, and construct the action space of reinforcement learning accordingly; Step 3-2: Define the dynamic intervention action space; for the lockdown strategy, assume the city is divided into... A spatial region, at the moment of decision-making Define the baseline urban area flow matrix as follows: ,in Indicates time From the region Flow to the region The population proportion; to characterize the regulatory effect of dynamic lockdown strategies on inter-regional movement, the decision time is defined. The action vector is: in, Indicates the region Residents at all times Permissible percentage of cross-regional flows; when When, it indicates that residents of that area are not allowed to move across areas; when When, it indicates that the residents of the area maintain the original level of inter-regional mobility; the action vector According to fixed decision intervals Updated and remains unchanged between two adjacent decision points; Step 3-3: Mapping motion to simulation parameters; given a motion vector Then, it is mapped into an effective urban mobility matrix after dynamic intervention. The elements of the effective flow matrix are defined as follows: in, For the Kronecker delta function, when The value is 1 if the condition is met, and 0 otherwise; the above mapping means: for residents of the region The proportion of individuals who originally retained inter-regional mobility was The portion of transregional flow that was suppressed They are redirected to stay within this area; thus, dynamic lockdown actions can adjust cross-regional population flow at the regional node granularity, and further affect the contact structure and propagation process in the propagation simulation environment; Steps 3-4: Extract the state of regional nodes; at the decision-making moment The simulation environment propagates and returns the region-level aggregate state to the reinforcement learning agent; for any region node... Construct its node feature vector for: in, Representing regions The number of people in a vulnerable, exposed, infected, hospitalized, recovering, and deceased state. Indicates the region Total population This represents the total population of the city; the first six items are calculated by dividing by... Normalization is performed to characterize the relative severity of the epidemic within a region; Indicates the effect of the previous decision on the region. The intensity of intervention is used to provide historical control information to the agent; Used to explicitly characterize population size heterogeneity between regions; Steps 3-5: Constructing a graph state representation; To represent the spatial propagation dependencies caused by inter-regional population flows, a spatial graph structure is constructed using regional nodes, and baseline regional flow relationships are used as edge features; For flows from regions Pointing to area A directed edge, whose edge characteristics are defined as follows: Based on this, a graph attention network is used to encode the region state; first, a linear mapping is performed on the node feature vectors: in, The weight matrix is ​​learnable; then, a graph attention mechanism incorporating edge features is introduced to focus on the nodes. and its neighboring nodes Calculate the attention coefficient: in, and For learnable parameters, This represents vector concatenation. Represents a node The set of neighboring nodes; attention coefficient Reflecting neighboring nodes For nodes The relative importance of nodes is determined, and the inter-regional flow intensity is explicitly considered; furthermore, the neighborhood node information is weighted and aggregated to obtain the node... Graph embedding representation: in, Represents a non-linear activation function; all nodes are embedded. Together they constitute the entire system at the decision-making moment. The graph state representation is used as input to the subsequent reinforcement learning agent for dynamic intervention strategy optimization. For both testing and isolation strategies and personal protective equipment (PPE) strategies, the same framework is used to define the coverage of testing and isolation or PPE as the corresponding action parameter.

5. The method for optimizing dynamic intervention strategies for urban infectious diseases according to claim 4, characterized in that, Step 4, which involves generating reinforcement learning algorithm training and intervention strategies, follows this process: Step 4-1: Initialize the reinforcement learning agent; construct a reinforcement learning agent for dynamic intervention strategy optimization, including an Actor network, a Critic network, a target Critic network, and an experience replay pool; the Actor network is used to generate continuous action vectors based on the current graph state embedding, the Critic network is used to estimate the value of state-action pairs, the target Critic network is used to stabilize the training process, and the experience replay pool is used to store environmental interaction samples. Step 4-2: Generate regional node-level intervention actions; at the decision-making moment. The graph state representation is composed of all embedded region nodes. The Actor network learns a random policy to generate continuous intervention actions. in, The parameters of the Actor network are represented. Indicates by parameters The determined random policy function; Step 4-3: Store the interaction and experience samples in the storage environment; convert the action vector output from Step 4-2 into a storage environment. In the propagation and intervention simulation environment constructed in step 1, the environment updates the effective region flow matrix and propagation process according to the action mapping relationship, and then returns the graph state at the next decision time. and current rewards Let the transfer sample obtained from a single environmental interaction be denoted as: The transferred samples are stored in the experience replay pool for subsequent network parameter updates; Step 4-4: Update the Critic network; randomly sample a small batch of transfer samples from the empirical replay pool, and construct the target Q-value using the target network; for any sample Its target value is defined as: in, As a discount factor, The entropy temperature coefficient Indicates the relationship with the first The first target Critic network; based on this, the second... The loss function of a Critic network is defined as: in, This represents the sample distribution in the experience replay pool; through the loss function Perform gradient descent on the parameters of the two Critic networks. Update them separately; Steps 4-5: Update the Actor Network; After completing the Critic Network update, update the Actor Network based on the current Critic Network's value assessment results, maximizing its long-term cumulative returns while maintaining its exploration capabilities; the optimization objective of the Actor Network is defined as: The first term corresponds to the policy entropy term, which encourages action exploration; the second term represents the value estimate of the action generated in the current state; this is achieved by minimizing the loss function. Update network parameters This allows the Actor network to gradually learn node-level dynamic intervention strategies that can bring higher net health benefits; Steps 4-6: Target Network Soft Update and Iterative Training; To improve training stability, a soft update is performed on the target Critic network after each round of parameter updates; the soft update rule is as follows: in, The soft update coefficients are used; the process of "graph state encoding - action generation - environment interaction - experience storage - network update" is executed until the preset number of training rounds is reached. Steps 4-7: Generate the optimal dynamic intervention strategy; after training, input the current urban propagation environment state into the trained reinforcement learning model, and the Actor network outputs the continuous action vectors of each regional node at the current decision time. This process is repeated at each subsequent decision point to obtain a dynamic sequence of intervention actions throughout the simulation time domain. .