Infectious disease prevention and control strategy determination method and device, electronic equipment and storage medium
By building a warehouse model and Markov decision-making model, combining deep reinforcement learning to train decision-making agents, the infectious disease prevention and control strategies are optimized, and the problem of poor effectiveness of prevention and control strategies in uncertain scenarios in uncertain scenarios is solved, and more effective infectious disease control is achieved.
Patent Information
- Application Number
- CN202510219347.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, prevention and control strategies under certain scenarios have poor prevention and control effects on the transmission of infectious diseases in real uncertain scenarios, especially due to the incompleteness of observation information, the effectiveness of prevention and control strategies is poor.
A method for determining strategies for infectious disease prevention and control is proposed. By constructing a warehouse model and Markov decision-making model for infectious disease transmission, combining deep reinforcement learning methods to train decision-making agents, and optimizing prevention and control strategies to adapt to uncertain scenarios.
This method can more effectively control the spread of infectious diseases in uncertain scenarios, reduce infection costs, detection costs and isolation costs, and improve the overall effectiveness of prevention and control strategies.
Smart Images

Figure CN120148894A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of public health, and particularly to a method, device, electronic device, and computer-readable storage medium for determining an infectious disease prevention and control strategy. Background Art
[0002] In order to control the spread of infectious diseases, non-pharmaceutical prevention and control strategies such as restricting travel, tracing contacts, and isolating infected persons are usually adopted. Although these measures can effectively slow down the spread of the virus and reduce the infection peak, they also have a significant negative impact on social and economic activities, such as suppressing economic activities and reducing income. In order to balance the effect of controlling the spread of infectious diseases and the cost generated by the negative impact on social and economic activities, it is necessary to optimize the infectious disease prevention and control strategy.
[0003] Current optimization schemes for prevention and control strategies usually optimize the preset prevention and control strategies based on preset deterministic scenarios. However, in real epidemic control, there are often large uncertainties, especially the incompleteness of observation information. In this case, the numbers of susceptible persons, exposed persons, infected persons, isolated persons, and recovered persons in the control area are constantly changing, and the scenarios are complex and changeable. Therefore, many prevention and control strategies optimized under deterministic scenarios have poor prevention and control effects on the spread of infectious diseases in real uncertain scenarios.
[0004] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of the present application is to provide a method, device, electronic device, and computer-readable storage medium for determining an infectious disease prevention and control strategy, aiming to solve the technical problem that the prevention and control strategy under a deterministic scenario has a poor prevention and control effect on the spread of infectious diseases in a real uncertain scenario.
[0006] To achieve the above purpose, the present application proposes a method for determining an infectious disease prevention and control strategy, and the method for determining an infectious disease prevention and control strategy includes:
[0007] Based on a preset variety of population types and the change amounts of the numbers of each population type, construct a compartment model for the spread of infectious diseases, where the population types at least include susceptible persons, exposed persons, infected persons, isolated persons, and recovered persons;
[0008] According to the compartment model, construct a corresponding Markov decision model, where the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values, and the reward values are inversely proportional to the infection costs, detection costs, and isolation costs corresponding to the prevention and control strategies respectively;
[0009] Establish a decision-making agent corresponding to the prevention and control strategy in the Markov decision model, and iteratively train the decision-making agent based on the Markov decision model until the reward value of the decision-making agent reaches a preset training goal to obtain a target agent;
[0010] Input the current system state data into the target agent to obtain a target prevention and control strategy corresponding to the current system state data, where the target prevention and control strategy is used to control the spread of infectious diseases.
[0011] In one embodiment, the exposed persons include detected exposed persons and undetected exposed persons, the infected persons include undetected infected persons, detected infected persons, and reported infected persons, the isolated persons include isolated exposed persons and isolated infected persons, and the change amount of the quantity of each population type at least includes the natural change amount;
[0012] Before the step of constructing a compartment model for the spread of infectious diseases based on a preset variety of population types and the change amount of the quantity of each population type, the method further includes:
[0013] Determine the natural change amount of susceptible persons according to the exposure rate of the control area and the current number of susceptible persons, where the exposure rate is determined by the sensing rate of the control area, the flow matrix, and the total number of infected persons;
[0014] Determine the natural change amount of undetected exposed persons according to the exposure rate of the control area, the current number of susceptible persons, the infection conversion rate, and the current number of undetected exposed persons;
[0015] Determine the natural change amount of detected exposed persons according to the infection conversion rate and the current number of detected exposed persons;
[0016] Determine the natural change amount of undetected infected persons according to the infection conversion rate, the current number of undetected exposed persons, the recovery rate, and the current number of undetected infected persons;
[0017] Determine the natural change amount of detected infected persons according to the infection conversion rate, the current number of detected exposed persons, the recovery rate, and the current number of detected infected persons;
[0018] Determine the natural change amount of reported infected persons according to the recovery rate and the current number of reported infected persons;
[0019] Determine the natural change amount of isolated exposed persons according to the infection conversion rate and the current number of isolated exposed persons;
[0020] Determine the natural change amount of isolated infected persons according to the infection conversion rate, the current number of isolated exposed persons, the recovery rate, and the current number of isolated infected persons;
[0021] Determine the natural change amount of recovered individuals based on the recovery rate, the current number of undetected infected individuals, the current number of detected infected individuals, the current reported number of infected individuals, and the current number of quarantined infected individuals.
[0022] In one embodiment, the change amount of each of the population types further includes a reported change amount, a detected change amount, and a quarantined change amount;
[0023] Before the step of constructing a compartment model for infectious disease transmission based on a preset variety of population types and the change amounts of the quantities of each of the population types, the method further includes:
[0024] Determine the reported change amounts corresponding to the undetected infected individuals and the detected infected individuals respectively according to the active reporting rate of the control area and the current number of undetected infections;
[0025] Determine the detected change amounts corresponding to the undetected exposed individuals and the detected exposed individuals respectively according to the detection rate of the control area, the effectiveness rate of the detection reagent for exposed individuals, and the current number of undetected exposed individuals;
[0026] Determine the detected change amounts corresponding to the undetected infected individuals and the detected infected individuals respectively according to the detection rate of the control area, the effectiveness rate of the detection reagent for infected individuals, and the current number of undetected infected individuals;
[0027] Determine the quarantined change amount of the detected exposed individuals and the quarantined change amount of the quarantined exposed individuals according to the quarantine rate of the control area and the current number of detected exposed individuals;
[0028] Determine the quarantined change amount of the detected infected individuals according to the quarantine rate of the control area and the current number of detected infected individuals;
[0029] Determine the quarantined change amount of the reported infected individuals according to the quarantine rate of the control area and the current reported number of infections;
[0030] Determine the quarantined change amount of the quarantined infected individuals according to the quarantine rate of the control area, the current number of detected infected individuals, and the current reported number of infections.
[0031] In one embodiment, the prevention and control strategy includes a detection ratio and a quarantine ratio;
[0032] The step of constructing a corresponding Markov decision model according to the compartment model includes:
[0033] Compose a system state according to the number of infected individuals in the control area, the change amount of exposed individuals in the compartment model, and the number of infected individuals in the global area where the control area is located, wherein the system state includes a local state and a global state;
[0034] Initialize the prevention and control strategy, where the prevention and control strategy includes a detection ratio and a quarantine ratio;
[0035] Based on the quantities and changes of the population types in the compartment model, predict the original state transition probabilities corresponding to the current system state transitioning to various predicted system states;
[0036] Based on the quantities and changes of the population types in the compartment model, predict the action state transition probabilities corresponding to the current system state transitioning to various predicted system states after implementing the prevention and control strategy;
[0037] According to the sum of the number of exposed and infected individuals in each sub-region of the global region, the proportion of the number of exposed and infected individuals from the control region in each sub-region to the sum of the number of exposed and infected individuals in each sub-region, and the newly added exposed individuals in each sub-region, determine the infection cost of the control region;
[0038] Based on the detection ratio and quarantine ratio of the control region, calculate the corresponding detection cost and quarantine cost;
[0039] Perform weighted calculation on the infection cost, the detection cost, and the quarantine cost of the control region to obtain the reward value of the Markov decision model.
[0040] In one embodiment, the decision-making agent includes a policy network and a value network, and the policy network includes a policy function;
[0041] The steps of establishing a decision-making agent corresponding to the prevention and control strategy in the Markov decision model include:
[0042] Initialize the policy function in the policy network, where the policy function is used to determine the corresponding prevention and control strategy according to the input system state data of the control region;
[0043] Initialize the value function in the value network, where the value function is used to determine the corresponding state value according to the input local state data and global state data of the control region.
[0044] In one embodiment, the training objective is to maximize and converge the cumulative reward value. The steps of iteratively training the decision-making agent based on the Markov decision model until the reward value of the decision-making agent reaches the preset training objective to obtain the target agent include:
[0045] Collect multiple sets of system state data with continuous time;
[0046] Input each of the system state data into the policy function to obtain the corresponding prevention and control strategy;
[0047] Input each of the system state data and the corresponding prevention and control strategies into the Markov decision model to obtain the corresponding reward value and transition state;
[0048] Train the policy network and the value network based on each of the system state data, control measurements, reward values, and transition states;
[0049] When the cumulative reward value of the decision-making agent is maximized and converges within a preset time period, stop the training and set the decision-making agent as the target agent.
[0050] In one embodiment, before the step of inputting the current system state data into the target agent to obtain the target prevention and control strategy corresponding to the current system state data, the method further includes:
[0051] Obtain the number of exposed persons and the number of infected persons in the current system state data;
[0052] Based on the number of exposed persons and the number of infected persons, the total number of people in the control area, and the flow matrix, reconstruct the reported number of infected persons in the control area to obtain the updated reported number of infected persons;
[0053] Update the number of exposed persons and the number of infected persons in the current system state data according to the updated reported number of infected persons to obtain the updated system state data, where the number of exposed persons includes the number of detected exposed persons and the number of undetected exposed persons, and the number of infected persons includes the number of undetected infected persons, the number of detected infected persons, and the reported number of infected persons;
[0054] Input the updated system state data into the target agent to obtain the target prevention and control strategy corresponding to the updated system state data.
[0055] In addition, to achieve the above object, the present application also proposes an infectious disease prevention and control strategy determination device, and the infectious disease prevention and control strategy determination device includes:
[0056] A compartment model establishment module, configured to construct a compartment model for infectious disease transmission based on a preset variety of population types and the change amounts of the quantities of each of the population types, where the population types at least include susceptibles, exposed persons, infected persons, quarantined persons, and recovered persons;
[0057] A decision model establishment module, configured to construct a corresponding Markov decision model according to the compartment model, where the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values, and the reward values are inversely proportional to the infection costs, detection costs, and quarantine costs corresponding to the prevention and control strategies respectively;
[0058] An agent training module, configured to establish a decision-making agent corresponding to a prevention and control strategy in the Markov decision model, and iteratively train the decision-making agent based on the Markov decision model until the reward value of the decision-making agent reaches a preset training target, thereby obtaining a target agent;
[0059] A strategy output module, configured to input current system state data into the target agent to obtain a target prevention and control strategy corresponding to the current system state data, wherein the target prevention and control strategy is used to control the spread of infectious diseases.
[0060] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the infectious disease prevention and control strategy determination method as described above.
[0061] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the infectious disease prevention and control strategy determination method as described above.
[0062] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the infectious disease prevention and control strategy determination method as described above.
[0063] The present application proposes a method for determining an infectious disease prevention and control strategy. In the method for determining an infectious disease prevention and control strategy, first, a compartment model for infectious disease transmission is constructed based on a preset variety of population types and the change amounts of the numbers of each of the population types. Among them, the population types at least include susceptibles, exposed individuals, infected individuals, quarantined individuals, and recovered individuals. This realizes the construction of a simulation model for the numbers of various population types during the transmission of infectious diseases in an incomplete scenario, facilitating the analysis of the transmission situation of infectious diseases from the perspective of population types. Additionally, according to the compartment model, a corresponding Markov decision model is constructed. Among them, the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values. The reward value is inversely proportional to the infection cost, detection cost, and quarantine cost corresponding to the prevention and control strategy. The Markov decision model belongs to a decision model and can be used for the simulation and modeling of the decision-making process. It combines the prevention and control strategy of infectious diseases with the conversion of system states, and quantifies the value of the prevention and control strategy through the reward value. Then, a decision agent corresponding to the prevention and control strategy in the Markov decision model is established, and the decision agent is iteratively trained based on the Markov decision model until the reward value of the decision agent reaches a preset training target, obtaining a target agent. The decision agent and the target agent in the present application are agents for making prevention and control strategies according to the input system state data. During the process of training the decision agent, the decision-making process in the previously established Markov decision model is combined, and the reward value is used as the training target to obtain a target agent that meets the requirements. Finally, the current system state data is input into the target agent to obtain the target prevention and control strategy corresponding to the current system state data. Among them, the target prevention and control strategy is used to control the transmission of infectious diseases. The technical solution of the present application constructs a compartment model that conforms to the actual prevention and control scenario, thereby simulating the process of infectious disease transmission in combination with the prevention and control strategy, can adapt to complex scenarios with incomplete observations and continuously changing observation information, and combines the Markov decision model and the reasonable design of system states and reward values therein to establish a decision agent for outputting prevention and control strategies, thereby training a target agent that can make the reward value of the prevention and control strategy reach the training target. Finally, the trained target agent is used to analyze based on the current system state data and output a relatively optimal target prevention and control strategy to effectively control the process of infectious disease transmission in the control area, while ensuring a certain effect of controlling the transmission of infectious diseases and avoiding excessive social and economic activity costs brought by the prevention and control strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0065] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1 It is a schematic flowchart of the first embodiment of the method for determining the infectious disease prevention and control strategy in the embodiment of the present application;
[0067] Figure 2 It is a schematic diagram of a compartment model based on SEIRQ in the method for determining the infectious disease prevention and control strategy of the present application;
[0068] Figure 3 It is a statistical chart of the infectious disease prevention and control effect of the target prevention and control strategy determined by applying the method for determining the infectious disease prevention and control strategy of the embodiment of the present application in City A and under the condition of full observation and control;
[0069] Figure 4 It is a statistical chart of the infectious disease prevention and control effect of the target prevention and control strategy determined by applying the method for determining the infectious disease prevention and control strategy of the embodiment of the present application in City A and under the condition of incomplete observation and control;
[0070] Figure 5 It is a statistical chart of the infectious disease prevention and control effect of the target prevention and control strategy determined by applying the method for determining the infectious disease prevention and control strategy of the embodiment of the present application in City A and under the condition of incomplete observation and control after information reconstruction;
[0071] Figure 6 It is a schematic diagram of the structural composition of the device for determining the infectious disease prevention and control strategy in the embodiment of the present application;
[0072] Figure 7 It is a schematic diagram of the device structure of the hardware operating environment involved in the method for determining the infectious disease prevention and control strategy in the embodiment of the present application.
[0073] The realization of the purpose, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the drawings. Specific Embodiments
[0074] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0075] To better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings in the specification and the specific embodiments.
[0076] The execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, a server, etc., or an electronic device, a control device, etc. that can implement the above functions. Hereinafter, a computer is taken as an example of the execution subject to illustrate this embodiment and the following embodiments.
[0077] Currently, when optimizing the prevention and control strategies for infectious diseases, the prevention and control strategies are usually optimized based on deterministic scenarios. However, in the actual epidemic control, there are often great uncertainties, especially the incompleteness of observational information. In this case, it is difficult to fully grasp the information of the infected, resulting in poor prevention and control effects of the prevention and control strategies. Many prevention and control strategies optimized under deterministic scenarios often fail to maintain good prevention and control effects when facing incomplete observations. In order to break through the limitations of existing research in uncertain scenarios, especially in the case of incomplete observations, the embodiments of this application propose a method for infectious disease control strategies. First, a transmission compartment model that fits the reality is constructed, and the incomplete observation situation in the real scenario is simulated, so as to more realistically reproduce the shortage of observational information in the actual control process of infectious disease transmission. Secondly, the prevention and control strategies are learned through the deep reinforcement learning method, and the impact of incomplete observations on the prevention and control effects is analyzed. The idea of information reconstruction can also be introduced to alleviate to a certain extent the problem of poor prevention and control strategy effects caused by insufficient observational information.
[0078] The embodiments of this application provide a method for determining an infectious disease prevention and control strategy, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the method for determining an infectious disease prevention and control strategy of this application. The method for determining an infectious disease prevention and control strategy includes:
[0079] Step S10, based on a preset variety of population types and the change amounts of the numbers of each population type, construct a compartment model for the transmission of infectious diseases, where the population types at least include susceptibles, exposed individuals, infected individuals, quarantined individuals, and recovered individuals;
[0080] In the embodiments of this application, first, based on the actual prevention and control scenario, five population types of susceptibles (Susceptible), exposed individuals (Exposed), infected individuals (Infected), quarantined individuals (Quarantined), and recovered individuals (Recovered) are considered. It should be noted that there may be cross relationships between various different population types. For example, a person is both an infected individual and a quarantined individual. After that, in order to simulate the change situations of the numbers of these different population types, a corresponding compartment model for the transmission of infectious diseases is constructed.
[0081] Furthermore, the exposed can be further divided into undetected exposed and detected exposed, the infected can be divided into undetected infected, detected infected, and reported (voluntarily reported) infected, and the quarantined can be divided into quarantined exposed and quarantined infected.
[0082] Exemplarily, the compartment model can be represented as the SEIQR model (this abbreviation represents the above various population types). In the implementation process, the real-world transmission scenario can be abstracted based on the ordinary differential equation (ODE) method. On the one hand, this compartment model strives to accurately reflect the actual transmission process and be close to the real situation; on the other hand, it avoids over-refinement and abstracts general laws to ensure good applicability and universality.
[0083] Specifically, considering the general situation of epidemic prevention and control in reality and extracting the common control actions: detection and quarantine, a metapopulation ODE compartment model is designed. The population types can include:
[0084] Susceptibles represent the population that has not been infected and can be infected by the exposed and the infected. When coming into contact with the exposed or the infected, there is a possibility of being infected and turning into an exposed person.
[0085] The exposed represent the population that has been infected but is in the early stage of infection and does not yet have full infectivity. The exposed usually have an incubation period and may not show obvious symptoms. Furthermore, the exposed can be subdivided into the following two categories: undetected (Eundetected, Eun), detected (Edetected, Ede), representing undetected and detected exposed.
[0086] The infected represent the population that already has infectivity and can transmit the pathogen to others. The infected can be further subdivided according to their detection status: undetected (Iundetected, Iun), detected (Idetected, Ide), voluntarily reported (Ireported, Ire), representing undetected, detected, and voluntarily reported infected.
[0087] The quarantined refer to the population that has been infected and needs to be quarantined and no longer has infectivity. The quarantined can be divided into: quarantined exposed (QE), quarantined infected (QI).
[0088] The recovered refer to the population that has recovered after infection, no longer participates in the transmission, and has immunity themselves.
[0089] Combining the above content, referring to Figure 2The shown compartment division is considered, and each area is regarded as a metapopulation. Considering the mutual conversion among different population types, a corresponding (metapopulation) compartment model is constructed. Among them, S can be converted into Eun, Eun can be converted into Ede or Iun, Ede can be converted into Ide or QE, Iun can be converted into Ire or Ide or QI, QE can be converted into QI, I (including Ire, Ide, and Iun) can be converted into R, and QI can be converted into R.
[0090] In another feasible embodiment, without changing the compartment design, the ODE model can also be replaced by an AMB (Agent-Based Model).
[0091] Step S20: According to the compartment model, a corresponding Markov decision model is constructed. The Markov decision model includes system state, prevention and control strategies, original state transition probability, action state transition probability, and reward value. The reward value is inversely proportional to the infection cost, detection cost, and isolation cost corresponding to the prevention and control strategies respectively.
[0092] Based on the aforementioned metapopulation ODE model with control actions such as detection and isolation, a corresponding Markov decision (MDP, Markov Decision Process) model is further constructed to determine and optimize the prevention and control strategies. Among them, system state, prevention and control strategies, original state transition probability, action state transition probability, and reward value are all essential components of the Markov decision model.
[0093] Step S30: Establish a decision agent corresponding to the prevention and control strategy in the Markov decision model, and iteratively train the decision agent based on the Markov decision model until the reward value of the decision agent reaches the preset training target to obtain the target agent.
[0094] In order to generate more scientific and reasonable prevention and control strategies, in the embodiments of the present application, a decision agent is used to generate the prevention and control strategy for control area i, where the prevention and control strategy includes the detection level and the isolation level.
[0095] Exemplarily, the PPO (Proximal Policy Optimization) algorithm can be used to train the reinforcement learning agent to optimize the prevention and control strategy. PPO is an algorithm based on policy optimization in reinforcement learning and is an improved version of the policy gradient method, combining efficiency and stability. It ensures that while optimizing the policy, it avoids destroying the previously learned policy by restricting the update amplitude during policy update, taking into account both performance improvement and training stability.
[0096] The decision-making agent is a deep learning network model, which includes a network for generating prevention and control strategies. This network includes various model parameters. During the training process, the reward value corresponding to the generated prevention and control strategy is calculated, and the reward value needs to be determined in combination with the aforementioned Markov decision model. The Markov decision model can reflect the transfer of the system state after the implementation of the prevention and control strategy. Finally, the corresponding reward value is calculated based on the infection cost, detection cost, isolation cost, etc. corresponding to the transferred system state.
[0097] Step S40: Input the current system state data into the target agent to obtain the target prevention and control strategy corresponding to the current system state data, where the target prevention and control strategy is used to control the spread of infectious diseases.
[0098] The trained target agent can generate the corresponding target prevention and control strategy according to the input system state data. This target prevention and control strategy is a prevention and control strategy with the relatively highest reward value under the condition of this system state data. After obtaining the target prevention and control strategy, the prevention and control work of infectious disease spread in the real scenario can be guided based on the detection level and isolation level in the target prevention and control strategy, so that the spread of infectious diseases can be controlled to a certain extent and the detection cost, isolation cost, etc. are ensured not to be too high, taking into account both the prevention and control effect of infectious diseases and the prevention and control cost of infectious diseases.
[0099] An embodiment of the present application proposes a method for determining an infectious disease prevention and control strategy. In the method for determining an infectious disease prevention and control strategy, first, based on a preset variety of population types and the change amounts of the quantities of each of the population types, a compartment model for infectious disease transmission is constructed, wherein the population types at least include susceptibles, exposed individuals, infected individuals, quarantined individuals, and recovered individuals. This realizes the construction of a simulation model for the number of people of various population types during the infectious disease transmission process in an incomplete scenario, facilitating the analysis of the infectious disease transmission situation from the perspective of population types. Additionally, according to the compartment model, a corresponding Markov decision model is constructed, wherein the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values. The reward values are inversely proportional to the infection costs, detection costs, and quarantine costs corresponding to the prevention and control strategies respectively. The Markov decision model belongs to a decision model and can be used for simulating and modeling the decision-making process, combining the prevention and control strategies of infectious diseases with the conversion of system states, and quantifying the prevention and control strategies through reward values for subsequent optimization of the prevention and control strategies. Then, a decision agent corresponding to the prevention and control strategy in the Markov decision model is established, and the decision agent is iteratively trained based on the Markov decision model until the reward value of the decision agent reaches a preset training target to obtain a target agent. The decision agent and the target agent in the embodiment of the present application are agents for making prevention and control strategies according to the input system state data. During the process of training the decision agent, the decision-making process in the previously established Markov decision model is combined, and the reward value is used as the training target to obtain a target agent that meets the requirements. Finally, the current system state data is input into the target agent to obtain the target prevention and control strategy corresponding to the current system state data, wherein the target prevention and control strategy is used to control the transmission of infectious diseases. The technical solution of the embodiment of the present application constructs a compartment model that conforms to the actual prevention and control scenario, thereby simulating the process of infectious disease transmission in combination with the prevention and control strategy, can adapt to complex scenarios with incomplete observations and continuously changing observation information, and combines the Markov decision model and the reasonable design of system states and reward values therein to establish a decision agent for outputting prevention and control strategies, thereby training a target agent that can make the reward value of the prevention and control strategy reach the training target. Finally, the trained target agent is used to analyze based on the current system state data and output a relatively optimal target prevention and control strategy to effectively control the process of infectious disease transmission in the control area, while ensuring a certain effect of controlling the transmission of infectious diseases and avoiding excessive social and economic activity costs brought by the prevention and control strategy.
[0100] Further, considering disease transmission dynamics, during the transmission of an infectious disease, the number of people of various population types will change, which is called the natural change amount.
[0101] In a feasible implementation, the exposed persons include detected exposed persons and undetected exposed persons, the infected persons include undetected infected persons, detected infected persons, and reported infected persons, and the quarantined persons include quarantined exposed persons and quarantined infected persons. The change amount of the number of each population type at least includes a natural change amount;
[0102] Therefore, in order to simulate the natural change amounts corresponding to susceptible persons, undetected exposed persons, detected exposed persons, undetected infected persons, detected infected persons, reported infected persons, quarantined exposed persons, quarantined infected persons, recovered persons, etc., before the step of constructing a compartment model for infectious disease transmission based on a preset variety of population types and the change amounts of the numbers of each population type, the method may further include:
[0103] Step A10, determining the natural change amount of susceptible persons according to the exposure rate of the control area and the current number of susceptible persons, where the exposure rate is determined by the sensing rate of the control area, the flow matrix, and the total number of infected persons;
[0104] Among them, the number of susceptible persons is affected by the exposure rate of the control area. The higher the exposure rate, the faster the susceptible persons are converted into exposed persons or infected persons. The calculation expression for the natural change amount of susceptible persons is:
[0105]
[0106] Among them, λ i is the exposure rate of control area i, and S i (t) is the current number of susceptible persons at time t.
[0107] In addition, the exposure rate of the control area is affected by the sensing rate of the control area, the flow matrix, and the number of existing total infected persons, and can be expressed as:
[0108]
[0109] Among them, λ i is the exposure rate of control area i, m ij is the flow matrix between area i and area j, β is the infection rate, N l is the total number of people in area l, r E represents the moving proportion of exposed persons E compared to infected persons I, and P M represents the moving proportion of infected persons.
[0110] Step A20, determining the natural change amount of undetected exposed persons according to the exposure rate of the control area, the current number of susceptible persons, the infection conversion rate, and the current number of undetected exposed persons;
[0111] Among them, the increase in the natural change of undetected exposed individuals is equal to the decrease in the number of susceptible individuals. The decrease in the natural change of undetected exposed individuals is affected by the infection conversion rate. The infection conversion rate σ refers to the conversion rate from exposed individuals to infected individuals.
[0112] Exemplarily, the calculation expression for the natural change of undetected exposed individuals is:
[0113]
[0114] Among them, E un,i (t) is the number of current undetected exposed individuals at time t, and σ is the infection conversion rate.
[0115] Step A30: Determine the natural change of detected exposed individuals according to the infection conversion rate and the current number of detected exposed individuals;
[0116] Among them, the natural change of detected exposed individuals is affected by the infection conversion rate. The calculation expression for the natural change of detected exposed individuals is:
[0117]
[0118] Among them, E de,i (t) is the number of current detected exposed individuals at time t, and σ is the infection conversion rate.
[0119] Step A40: Determine the natural change of undetected infected individuals according to the infection conversion rate, the current number of undetected exposed individuals, the recovery rate, and the current number of undetected infected individuals;
[0120] Among them, the natural change of undetected infected individuals is affected by the infection conversion rate (influence increment) and the recovery rate (influence decrement). The recovery rate refers to the probability of converting from infected individuals to recovered individuals. Specifically, the calculation expression for the natural change of undetected infected individuals is:
[0121]
[0122] Among them, I un,i (t) is the number of current undetected infected individuals at time t, E un,i (t) is the number of current undetected infected individuals at time t, σ is the infection conversion rate, and γ is the recovery rate.
[0123] Step A50: Determine the natural change of detected infected individuals according to the infection conversion rate, the current number of detected exposed individuals, the recovery rate, and the current number of detected infected individuals;
[0124] Among them, the natural change of detected infected individuals is affected by the infection conversion rate (influence increment) and the recovery rate (influence decrement). Specifically, the calculation expression for the natural change of detected infected individuals is:
[0125]
[0126] Among them, I de,i (t) is the number of currently detected infected persons at time t, and E de,i (t) is the number of currently detected infected persons at time t, σ is the infection conversion rate, and γ is the recovery rate.
[0127] Step A60: Determine the natural change amount of the reported infected persons according to the recovery rate and the current number of reported infected persons;
[0128] Among them, the natural change amount of the reported infected persons is mainly affected by the recovery rate (affecting the reduction amount). It should be noted that the increase amount of the reported infected persons is affected by the reporting action and is not considered within the natural change amount.
[0129] Exemplarily, the calculation expression for the natural change amount of the reported infected persons is:
[0130]
[0131] Among them, I re,i (t) is the current number of reported infected persons at time t, and γ is the recovery rate.
[0132] Step A70: Determine the natural change amount of the quarantined exposed persons according to the infection conversion rate and the current number of quarantined exposed persons;
[0133] Among them, the natural change amount of the quarantined exposed persons is mainly affected by the infection conversion rate (affecting the reduction amount) and is converted into the corresponding exposed infected persons. It should be noted that the increase amount of the quarantined exposed persons is affected by the quarantine action and is not considered within the natural change amount.
[0134] Exemplarily, the calculation expression for the natural change amount of the quarantined exposed persons is:
[0135]
[0136] Among them, Q E (t) is the current number of quarantined exposed persons at time t, and σ is the infection conversion rate.
[0137] Step A80: Determine the natural change amount of the quarantined infected persons according to the infection conversion rate, the current number of quarantined exposed persons, the recovery rate, and the current number of quarantined infected persons;
[0138] Among them, the natural change amount of the quarantined infected persons is mainly affected by the infection conversion rate (affecting the increase amount) and the recovery rate (affecting the reduction amount), and the quarantined exposed persons are converted into quarantined infected persons based on the infection conversion rate.
[0139] Exemplarily, the calculation expression for the natural change in the number of isolated infected individuals is:
[0140]
[0141] where Q E (t) is the current number of isolated exposed individuals at time t, Q I (t) is the current number of isolated infected individuals at time t, σ is the infection conversion rate, and γ is the recovery rate.
[0142] Step A90: Determine the natural change in the number of recovered individuals based on the recovery rate, the current number of undetected infected individuals, the current number of detected infected individuals, the current reported number of infected individuals, and the current number of isolated infected individuals.
[0143] Among them, the natural change in the number of recovered individuals is mainly affected by the recovery rate (incremental effect), and the source population types for conversion include current undetected infected individuals, current detected infected individuals, current reported infected individuals, and current isolated infected individuals.
[0144] Exemplarily, the calculation expression for the natural change in the number of recovered individuals is:
[0145]
[0146] where R i (t) is the current number of recovered individuals at time t, Q I (t) is the current number of isolated infected individuals at time t, I re,i (t) is the current reported number of infected individuals at time t, I de,i (t) is the current detected number of infected individuals at time t, I un,i (t) is the current undetected number of infected individuals at time t, and γ is the recovery rate.
[0147] Furthermore, in a feasible embodiment, considering the impact of control actions such as reporting, detection, and isolation on the changes in the number of each population type, the change in each of the population types further includes a reporting change, a detection change, and an isolation change;
[0148] Before the step of constructing a compartment model for infectious disease transmission based on a preset variety of population types and the changes in the number of each population type, the method further includes:
[0149] Step B10: Determine the reporting changes corresponding to undetected infected individuals and detected infected individuals respectively based on the active reporting rate of the control area and the current number of undetected infections;
[0150] It is understandable that once some of the undetected infected individuals report themselves proactively, they will be converted into detected infected individuals. Therefore, the absolute value of the reporting change in the number of undetected infected individuals is equal to the absolute value of the reporting change in the number of detected infected individuals.
[0151] Exemplarily, the calculation expression for the reporting change corresponding to the undetected infected individuals is:
[0152]
[0153] In addition, the calculation expression for the reporting change corresponding to the detected infected individuals is:
[0154]
[0155] where P re,i represents the proactive reporting rate of region i, that is, the probability that an infected individual reports their infection status proactively, and I un,i (t) is the number of current undetected infected individuals at time t.
[0156] Step B20: Determine the detection changes corresponding to the undetected exposed individuals and the detected exposed individuals respectively according to the detection rate of the control area, the effectiveness rate of the detection reagent for the exposed individuals, and the number of current undetected exposed individuals.
[0157] It is understandable that once some of the undetected exposed individuals are detected, they will be converted into detected exposed individuals. Therefore, the absolute value of the detection change in the number of undetected exposed individuals is equal to the absolute value of the detection change in the number of detected exposed individuals. In addition, when considering the reduction in the number of undetected exposed individuals, in addition to considering the detection rate, the effectiveness rate of the detection reagent for the exposed individuals also needs to be considered.
[0158] Exemplarily, the calculation expression for the detection change of the undetected exposed individuals is:
[0159]
[0160] where E un,i (t) represents the number of current undetected exposed individuals at time t, and P test,i is the detection rate of control area i, and P E,eff represents the effectiveness rate of the detection reagent for the exposed individuals.
[0161] In addition, the calculation expression for the detection change of the detected exposed individuals is:
[0162]
[0163] where E de,i (t) represents the number of current detected exposed individuals at time t, and P test,i is the detection rate of control area i, and P E,effIndicates the efficiency of the detection reagent for the exposed
[0164] Step B30: Determine the detection change amounts corresponding to the undetected infected individuals and the detected infected individuals respectively according to the detection rate in the control area, the efficiency of the detection reagent for the infected individuals, and the current number of undetected infected individuals.
[0165] It can be understood that when some of the undetected infected individuals are detected, they are converted into detected infected individuals. Therefore, the absolute value of the detection change amount of the undetected infected individuals is equal to the absolute value of the detection change amount of the detected infected individuals. In addition, when considering the reduction of the undetected infected individuals, in addition to combining the detection rate, the efficiency of the detection reagent for the infected individuals also needs to be considered.
[0166] Exemplarily, the calculation expression for the detection change amount of the undetected infected individuals is:
[0167]
[0168] where I un,i (t) represents the number of current undetected infected individuals at time t, and I un,i (t) is the number of current undetected infected individuals at time t, and P test,i is the detection rate of control area i, and P I,eff represents the efficiency of the detection reagent for the infected individuals.
[0169] Exemplarily, the calculation expression for the detection change amount of the detected infected individuals is:
[0170]
[0171] where I de,i (t) represents the number of current detected infected individuals at time t, and I un,i (t) is the number of current undetected infected individuals at time t, and P test,i is the detection rate of control area i, and P I,eff represents the efficiency of the detection reagent for the infected individuals.
[0172] Step B40: Determine the isolation change amount of the detected exposed individuals and the isolation change amount of the isolated exposed individuals according to the isolation rate in the control area and the current number of detected exposed individuals.
[0173] It can be understood that after the detected exposed individuals are isolated, they are converted into isolated exposed individuals. Therefore, the absolute value of the isolation change amount of the detected exposed individuals is equal to the absolute value of the isolation change amount of the isolated exposed individuals.
[0174] Exemplarily, the calculation expression for the isolation change amount of the detected exposed individuals is:
[0175]
[0176] Among them, E de,i (t) represents the number of currently detected exposed individuals at time t, and P quara,i represents the isolation rate of area i.
[0177] In addition, the calculation expression for the change in the number of isolated exposed individuals is:
[0178]
[0179] Among them, QE i (t) represents the number of isolated exposed individuals at time t, and E de,i (t) represents the number of currently detected exposed individuals at time t, and P quara,i represents the isolation rate of area i.
[0180] Step B50: Determine the change in the number of isolated detected infected individuals based on the isolation rate of the control area and the number of currently detected infected individuals;
[0181] It can be understood that the change in the number of isolated detected infected individuals is related to the isolation rate. Exemplarily, the calculation expression for the change in the number of isolated detected infected individuals is:
[0182]
[0183] Among them, I de,i (t) represents the number of detected infected individuals at time t, and P quara,i represents the isolation rate of area i.
[0184] Step B60: Determine the change in the number of isolated reported infected individuals based on the isolation rate of the control area and the currently reported number of infections;
[0185] Among them, the change in the number of isolated reported infected individuals is related to the isolation rate. Exemplarily, the calculation expression for the change in the number of isolated reported infected individuals is:
[0186]
[0187] Among them, I re,i (t) represents the number of reported infected individuals at time t, and P quara,i represents the isolation rate of area i.
[0188] Step B70: Determine the change in the number of isolated infected individuals based on the isolation rate of the control area, the currently detected number of infected individuals, and the currently reported number of infections.
[0189] Among them, the change in the isolation of quarantined infected individuals is related to the change in the isolation of detected infected individuals and the change in the isolation of reported infected individuals. It can be understood that the sum of the absolute values of the change in the isolation of detected infected individuals and the change in the isolation of reported infected individuals is equal to the change in the isolation of quarantined infected individuals (increment).
[0190] Exemplarily, the calculation expression for the change in the isolation of quarantined infected individuals is:
[0191]
[0192] Among them, QI i (t) represents the number of quarantined infected individuals at time t, P quara,i represents the isolation rate of area i, I re,i (t) represents the number of reported infected individuals at time t, I de,i (t) represents the number of detected infected individuals at time t.
[0193] In the embodiments of the present application, the mutual conversion and quantity changes of various population types due to prevention and control actions such as detection and isolation are simulated. Then, combined with the natural change amounts of each population type determined by the disease transmission dynamics, the spread of infectious diseases can be better simulated in an uncertain environment with prevention and control actions, making the reliability of the compartment model higher and providing a data basis for making better prevention and control strategies.
[0194] In a feasible embodiment, the prevention and control strategy includes a detection ratio and an isolation ratio; the step of constructing a corresponding Markov decision model according to the compartment model may include:
[0195] Step S21: Compose a system state based on the number of infected individuals in the control area, the change in the number of exposed individuals in the compartment model, and the number of infected individuals in the global area where the control area is located. Among them, the system state includes a local state and a global state;
[0196] Step S22: Initialize the prevention and control strategy, where the prevention and control strategy includes a detection ratio and an isolation ratio;
[0197] Step S23: Based on the number and change amount of each population type in the compartment model, predict the original state transition probabilities corresponding to the current system state transferring to various predicted system states;
[0198] Step S24: Based on the number and change amount of each population type in the compartment model, predict the action state transition probabilities corresponding to the current system state transferring to various predicted system states after taking the prevention and control strategy;
[0199] Step S25: Determine the infection cost of the control area based on the sum of the number of exposed and infected individuals in each sub-region of the global area, the proportion of the number of exposed and infected individuals from the control area in each sub-region to the sum of the number of exposed and infected individuals in each sub-region, and the newly added exposed individuals in each sub-region.
[0200] Step S26: Calculate the corresponding detection cost and isolation cost respectively based on the detection ratio and isolation ratio of the control area.
[0201] Step S27: Perform weighted calculation on the infection cost, detection cost, and isolation cost of the control area to obtain the reward value of the Markov decision model.
[0202] In the embodiments of the present application, a method for constructing and determining various parameters in a Markov decision (MDP, Markov Decision Process) model is also disclosed.
[0203] Among them, the Markov decision process can be defined by (S, A, P, R). Among them, S represents the state space, that is, the system state; A represents the action space where actions can be taken, that is, the prevention and control strategy; P represents the state transition probability that the system state predicted by the model transfers to other system states without intervention through control actions, that is, the original state transition probability; P(s′|s,a) represents the probability that the system state predicted by the model transfers to state s′ after taking action a in state s, and these are all determined by the ODE model; R represents the reward value. After taking action a in state s and the system transfers to state s′, the corresponding reward value. The reward value depends on the infection cost, detection cost, and isolation cost corresponding to the prevention and control strategy. The lower the total cost, the higher the reward value.
[0204] Exemplarily, the definitions of the system state S, control action A, and reward value R are as follows:
[0205] State: The state of control area i consists of two parts, the local state S local,i and the global state S global . The local state S local,i = [I combined , ΔE detected,i , where I combined represents the existing infection scale of control area i, and ΔE detected,i represents the newly detected exposed individuals. These two parameters can reflect the existing scale and new trend of infected individuals. The global state S global = ∑ i S local, representing the overall infection situation of the entire area. When determining the prevention and control strategy for control area i, it is necessary to combine the local state of this control area and the overall global state to formulate a prevention and control strategy suitable for this area.
[0206] Action: Based on the previously constructed metapopulation ODE model, the control action of control area i can be expressed as: is used to represent the detection level and isolation level of control area i. The detection level is the detection ratio of this area, and the isolation level is the isolation ratio. Among them, the detection ratio refers to the ratio of the number of people detected to the total number of people in this control area, and the isolation ratio refers to the ratio of the number of people isolated to the total number of people in this control area.
[0207] Furthermore, the concept of detection rate is proposed to reflect the probability that an infected person is detected. The higher the detection rate, the greater the probability that an infected person is detected. The corresponding relationship between the detection ratio and the detection rate is:
[0208]
[0209] Among them, P test,i represents the detection rate, represents the total number of infected people in area i (including all exposed people and infected people), and the function represents that at the infection rate of level, the number of test person - times required to detect a single infected person. This function can be taken as f(x)=x -0.6 , and the higher the infection rate, the smaller the required unit number of test person - times.
[0210] Furthermore, the concept of isolation rate is proposed to reflect the probability that a known infected person is isolated. The higher the isolation rate, the greater the probability that an infected person is isolated. The corresponding relationship between the isolation level and the isolation rate is:
[0211]
[0212] Among them, P quara,i represents the isolation rate, represents the known infected part of area i (including Ede, Ide, Ire). The detection and isolation levels of this area will affect the state transition of the ODE model, and thus achieve the prevention and control of infectious diseases. Different detection and isolation levels result in different prevention and control effects. In this way, subsequently, by selecting the optimal prevention and control strategy (detection level and isolation level), a relatively better prevention and control effect can be achieved.
[0213] Reward value (Reward): Reward is an evaluation of the control decision, assessing the situation where the current prevention and control strategy takes into account both the prevention and control effect and the control cost. Therefore, the reward function can be expressed as:
[0214] r i =-a1 ×C EI,i -a 2 ×C test,i -a 3 ×C quara,i ;
[0215] Among them, r i represents the reward value under the current prevention and control strategy for control area i, which is jointly determined by the infection cost C EI,i , the detection cost C test,i , and the isolation cost C quara,i .
[0216] Specifically, when calculating the infection cost of a certain control area, it is necessary to consider not only the changes in the number of infected people in this area, but also the changes in the number of infected people in other areas affected by this area. Therefore, for the infection cost C EI,i of control area i, the calculation expression can be:
[0217]
[0218] Among them, j is each local area (including i) in the global area where the control area is located, and can be obtained through I eff,j = r E E j + P m I j represents the effective number of infected people in area j. The calculation of C EI,i consists of two parts. The first bracket part represents the proportion of the infected people from area i to area j among the infected people arriving at area j, and the second bracket part represents the total newly added exposed people in area j, that is, ΔE un . Through the above calculation, the cost of the newly added exposed people can be shared among each area, which is beneficial to determining the clear corresponding relationship between the execution actions and the reward values. For the detection cost C test,i and the isolation cost C quara,i of area i, they are determined by the corresponding detection level (i.e., detection ratio) and isolation level (i.e., isolation ratio). The higher the detection level and isolation level, the higher the corresponding detection cost and isolation cost.
[0219] In a feasible embodiment, the decision-making agent includes a policy network and a value network, and the policy network includes a policy function; the steps of establishing the decision-making agent corresponding to the prevention and control strategy in the Markov decision model may include:
[0220] Step S31, initialize the policy function in the policy network, where the policy function is used to determine the corresponding prevention and control strategy according to the input system state data of the control area;
[0221] Step S32, initialize the value function in the value network, where the value function is used to determine the corresponding state value according to the input local state data and global state data of the control area.
[0222] After constructing the meta-population ODE model and formalizing the prevention and control process of infectious disease transmission as a Markov decision process (MDP), the PPO (Proximal Policy Optimization) algorithm can be used to train the reinforcement learning agent to optimize the prevention and control strategy. PPO is an algorithm based on policy optimization in reinforcement learning and is an improved version of the policy gradient method, combining efficiency and stability. It ensures that while optimizing the policy, it avoids destroying the previously learned policy by restricting the update amplitude during policy updates, taking into account both performance improvement and training stability.
[0223] In the process of constructing the decision-making agent, it is first necessary to clarify the input and output of the decision-making agent. For the PPO algorithm, there are two sets of networks: Actor (policy) and Critic (value). In the embodiment of the present application, a multi-agent architecture is adopted, that is, each control area has a corresponding decision-making agent. For area i, the policy network (Actor) inputs the local system state S local,i , and the policy function π(a|s) can output the prevention and control strategy (including the detection level and isolation level) in this state. The value network (Critic) inputs the concatenation of the local system state and the global state [S local,i , S global , and through the value function V π (s), evaluates the value in the current state and reflects the corresponding reward value.
[0224] In another feasible embodiment, deep reinforcement learning algorithms such as DQN (Deep Q-Network) and SAC (Soft Actor-Critic, an algorithm based on the maximum entropy reinforcement learning framework) can also be used to replace the PPO algorithm.
[0225] Furthermore, in the process of training the decision-making agent, the training objective is to maximize and converge the cumulative reward value. The step of iteratively training the decision-making agent based on the Markov decision model until the reward value of the decision-making agent reaches the preset training objective to obtain the target agent may include:
[0226] Step S31, collect multiple sets of system state data with continuous time;
[0227] Before formally training the decision-making agent, a certain amount of training data needs to be collected. Specifically, reinforcement learning learns by interacting with the environment. First, it is necessary to collect the system state s in the control area.
[0228] Step S32: Input each system status data into the policy function respectively to obtain the corresponding prevention and control strategies.
[0229] Step S33: Input each system status data and the corresponding prevention and control strategies into the Markov decision model to obtain the corresponding reward values and transition states.
[0230] During the training of the decision-making agent, the system status data s can be input into the policy function in chronological order to obtain the corresponding prevention and control strategies a, and the corresponding reward values r and transition states s′ can be calculated in combination with the calculation method of the reward value in the corresponding Markov decision model.
[0231] Specifically, for the PPO algorithm, action control can be performed through the policy function π(a|s). By the above method, multiple sets of system status data (s, a, r, s′) that are continuous in time distribution are collected, where s represents the system status, a represents the prevention and control strategy, r represents the corresponding reward value, and s′ represents the transition state after taking the prevention and control strategy. After collecting a certain amount of data, the training steps are executed.
[0232] Step S34: Train the policy network and the value network based on each system status data, control measurements, reward values, and transition states.
[0233] During the training of the agent, supervised learning training can be adopted. The collected system status data, control measurements, reward values, and transition states are regarded as real data or label data to iteratively update and train the model parameters of the policy network and the value network respectively. During the training process, the loss function value of the network needs to be calculated, and the parameter updates of the policy network and the reward network in the decision-making agent are guided based on the change of the loss function value.
[0234] Exemplarily, the training of the policy network can be based on the policy gradient theorem. The PPO algorithm simultaneously adopts the advantage estimation and the clip clipping method to ensure the stability of the policy gradient. The loss function expression of the policy network is as follows:
[0235] L(θ) = E[min(r t (θ)·A t , clip(r t (θ), 1 - ε, 1 + ε)·A t )];
[0236] where L(θ) is the loss function value of the policy network, is used to represent the ratio between the current policy and the pre-updated policy, A t is the advantage estimation function, and ε is the clip threshold, which is used to control the amplitude of the policy update.
[0237] On the other hand, the training of the value network is based on the idea of Temporal Difference (TD). It mainly uses the TD target of the next step as the target for iterative training. Correspondingly, the expression of the loss function of the value network is as follows:
[0238]
[0239] Among them, L V (θ) refers to the loss function value of the value network, and γ is used to control the update amplitude of the value network.
[0240] Step S35: When the cumulative reward value of the decision-making agent is maximized and converges within the preset time period, stop the training and set the decision-making agent as the target agent.
[0241] Repeat the aforementioned steps S31 to S34, repeat sampling data and optimizing network hyperparameters, and improve the strategy through multiple rounds of iteration. When the cumulative reward value corresponding to the prevention and control strategy output by the decision-making agent reaches the maximum and converges within the preset time period (that is, a higher reward value cannot be obtained through training), it can be considered that the training of the decision-making agent is completed, and the training can be stopped to obtain the target agent for determining the infectious disease prevention and control strategy.
[0242] It can be understood that as the infectious disease transmission process progresses, the matching degree between the target agent and the environment may decrease. At this time, the process of repeating steps S31 to S35 can be continued to update the target agent and improve the performance of the target agent to output a prevention and control strategy that is more matched to the current environment.
[0243] In a feasible embodiment, before the step of inputting the current system state data into the target agent to obtain the target prevention and control strategy corresponding to the current system state data, the method may further include:
[0244] Step S41: Obtain the number of exposed persons and the number of infected persons in the current system state data;
[0245] Step S42: Based on the number of exposed persons, the number of infected persons, the total number of people in the control area, and the flow matrix, reconstruct the reported number of infected persons in the control area to obtain the updated reported number of infected persons;
[0246] Step S43: Update the number of exposed persons and the number of infected persons in the current system state data according to the updated reported number of infected persons to obtain the updated system state data, where the number of exposed persons includes the number of detected exposed persons and the number of undetected exposed persons, and the number of infected persons includes the number of undetected infected persons, the number of detected infected persons, and the reported number of infected persons;
[0247] Step S44: Input the updated system status data into the target agent to obtain the target prevention and control strategy corresponding to the updated system status data.
[0248] In the embodiments of the present application, expressing the situation of incomplete observation in reality using the ODE model is actually in an ideal scenario where infected individuals will actively report at a rate of 100%. However, in reality, the active reporting rate in each region may randomly be 0, making it difficult to accurately estimate the scale of infection. Based on this, the incomplete observation structure in the embodiments of the present application can be expressed as:
[0249] Pre,i
[0250] I un,i →I re,i ;
[0251] where I un,i is the number of undetected infected individuals in region i, I re,i is the number of reported infected individuals in region i, and P re,i is the active reporting rate in region i, with a value ranging from 0 to 1. The smaller the value, the more deviated the observation of infected individuals in region i is from the truth, that is, the observation is incomplete. This will result in a poor prevention and control effect of the prevention and control strategy output by the decision-making agent in the scenario of incomplete observation.
[0252] To overcome the above defects, the embodiments of the present application propose an information reconstruction method to improve the prevention and control effect of the prevention and control strategy of the decision-making agent.
[0253] Specifically, first obtain the observation vector I obs at time t. The observation vector I obs includes the number of exposed individuals and the number of infected individuals in the current system status data corresponding to time t. For untested regions, it is considered that the observation is inaccurate and can be represented by O mask , where being true means untested, that is, inaccurate.
[0254] Then, perform information reconstruction on I pre based on the following reconstruction formula. The calculation formula is expressed as:
[0255]
[0256] where I pre represents the number of reported infected individuals in the region, N is the total number of people in the region, and M is the flow matrix.
[0257] After the reconstruction is completed, backfill the original current system status data based on the updated number of reported infected individuals I pre . The backfill formula is expressed as: I obs [O mask = I pre[O mask ], in order to realize the observation vector I obs The reconstruction backfill of the initially obtained I obs Update and observe the vector I obs It is part of the current system status data and updates I obs That is, the system status data is updated. It can be understood that the reconstructed and updated system status data is more consistent with the actual status of each control area. The updated system status data is input into the target intelligent agent to obtain the corresponding target prevention and control strategy. Compared with directly determining the target prevention and control strategy based on the system status data without information reconstruction, information reconstruction helps the target intelligent agent to make accurate target prevention and control strategies, improve the decision-making ability of the target intelligent agent, and more effectively control the spread of infectious diseases.
[0258] The technical solution of the embodiment of the present application is equivalent to constructing an incomplete observation model on the basis of the SEIQR model established in step S10 to describe the incomplete observation phenomenon existing in the real infectious disease prevention and control scenario. In addition, through the aforementioned information reconstruction technology, the control performance of the decision-making agent under incomplete observation is effectively improved. Specifically, in the modeling process of incomplete observation, the complexity of the real scene should be fully considered. Among them, the active reporting rate of the infected person is unknown and random, that is, the conversion rate from Iun to Ire is a variable with a random value. This randomness makes the acquisition and utilization of infectious disease transmission data challenging, and increases the difficulty of accurately judging the situation of infectious disease prevention and control. The information reconstruction method adopted is based on the spatial correlation between various regions (such as communities) brought by the flow network. By analyzing the information presented by the detected area, the relevant information of the undetected area can be reasonably inferred, and then under the condition of incomplete observation, the relatively accurate acquisition of information can be achieved, which provides strong support for the subsequent precise control. In the technical solution of the embodiment of the present application, by using the aforementioned information reconstruction method, the control performance under incomplete observation can be effectively improved.
[0259] The infectious disease prevention and control strategy determination method of the embodiment of the present application has been field tested and statistically analyzed. Taking City A in my country as the research area, the research area is subdivided into communities, and the anonymous mobile phone positioning data from a certain operator is used to establish a spatiotemporal propagation network. For example, the epidemic simulation parameter settings include: infection rate β = 0.8, infection conversion rate σ = 1 / 3, recovery rate γ = 1 / 7, and the mobile ratio P of the infected person. m =0.4. At the beginning of the simulation, 100 initial infected people were distributed to the 100 communities with the largest population, and the total simulation duration was 120 days. Among them, the parameters of the Markov decision process, the weight a of the reward function 1 =400, a 2 =0.0125, a 3 =1,a1 and a 2 and a 3 The corresponding infection cost, detection cost, and quarantine cost respectively. The detection level actions can be selected from 5 levels: [0, 0.05, 0.2, 0.5, 1.0]×0.1, and the quarantine level actions can be selected from 2 levels: [0, 1.0]×0.01.
[0260] Based on the foregoing control scenarios, in the case of fully observable control, the statistical chart of the infectious disease prevention and control effect is as shown in Figure 3 ; in the case of partially observable control, the statistical chart of the infectious disease prevention and control effect is as shown in Figure 4 ; in the case of partially observable control after information reconstruction, the statistical chart of the infectious disease prevention and control effect is as shown in Figure 5 . In each of the above figures, the abscissa is time, and the ordinate is the value, including the observed total number of infected people, the actual total number of infected people, as well as the detection level and quarantine level.
[0261] As shown in Figure 3 , in the fully observable scenario, the prevention and control strategy output by the decision-making agent in the embodiment of the present application can effectively control the spread of infectious diseases in a relatively short time. However, in the partially observable scenario ( Figure 4 ), the prevention and control effect significantly decreases, and it takes a certain period of time to end the control. This indicates that partial observability has a greater impact on control. If information reconstruction is performed, effective control can still be achieved ( Figure 5 ). This demonstrates that the information reconstruction technology can effectively alleviate the impact of partial observability and still achieve a relatively good prevention and control effect.
[0262] Furthermore, during the field test of infectious disease control in City A, the prevention and control effects of the decision-making agent on the spread of infectious diseases in various scenarios are as shown in the following table:
[0263]
[0264] As can be seen from the above table, in the partially observable scenario, after information reconstruction, with a relatively low daily average detection rate and the number of quarantined people, the total number of infected people significantly decreases compared to when information reconstruction is not performed, and the cumulative reward significantly increases, both approaching the prevention and control effect in the fully observable scenario. This shows that information reconstruction can effectively improve the effectiveness of the prevention and control strategy in the partially observable scenario.
[0265] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the method for determining the infectious disease prevention and control strategy of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0266] The present application also provides an apparatus for determining an infectious disease prevention and control strategy. Referring to Figure 6 , the apparatus for determining an infectious disease prevention and control strategy includes:
[0267] A compartment model establishment module 10, configured to construct a compartment model for infectious disease transmission based on a preset variety of population types and the change amounts of the quantities of each of the population types, where the population types at least include susceptibles, exposed individuals, infected individuals, isolated individuals, and recovered individuals;
[0268] A decision model establishment module 20, configured to construct a corresponding Markov decision model according to the compartment model, where the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values, and the reward values are inversely proportional to the infection costs, detection costs, and isolation costs corresponding to the prevention and control strategies respectively;
[0269] An agent training module 30, configured to establish a decision agent corresponding to the prevention and control strategy in the Markov decision model, and perform iterative training on the decision agent based on the Markov decision model until the reward value of the decision agent reaches a preset training target, thereby obtaining a target agent;
[0270] A strategy output module 40, configured to input current system state data into the target agent to obtain a target prevention and control strategy corresponding to the current system state data, where the target prevention and control strategy is used to control the spread of infectious diseases.
[0271] In an embodiment, the exposed individuals include detected exposed individuals and undetected exposed individuals, the infected individuals include undetected infected individuals, detected infected individuals, and reported infected individuals, the isolated individuals include isolated exposed individuals and isolated infected individuals, and the change amounts of the quantities of each of the population types at least include natural change amounts;
[0272] The compartment model establishment module 10 is further configured to:
[0273] Determine the natural change amount of susceptibles according to the exposure rate of the control area and the current number of susceptibles, where the exposure rate is determined by the sensing rate of the control area, the flow matrix, and the total number of infected individuals;
[0274] Determine the natural change amount of undetected exposed individuals according to the exposure rate of the control area, the current number of susceptibles, the infection conversion rate, and the current number of undetected exposed individuals;
[0275] Determine the natural change amount of detected exposed individuals according to the infection conversion rate and the current number of detected exposed individuals;
[0276] Determine the natural change amount of undetected infected persons based on the infection conversion rate, the current number of undetected exposed persons, the recovery rate, and the current number of undetected infected persons;
[0277] Determine the natural change amount of detected infected persons based on the infection conversion rate, the current number of detected exposed persons, the recovery rate, and the current number of detected infected persons;
[0278] Determine the natural change amount of reported infected persons based on the recovery rate and the current number of reported infected persons;
[0279] Determine the natural change amount of quarantined exposed persons based on the infection conversion rate and the current number of quarantined exposed persons;
[0280] Determine the natural change amount of quarantined infected persons based on the infection conversion rate, the current number of quarantined exposed persons, the recovery rate, and the current number of quarantined infected persons;
[0281] Determine the natural change amount of recovered persons based on the recovery rate, the current number of undetected infected persons, the current number of detected infected persons, the current number of reported infected persons, and the current number of quarantined infected persons.
[0282] In one embodiment, the change amount of each of the population types further includes a reported change amount, a detection change amount, and a quarantine change amount;
[0283] The compartment model establishment module 10 is further configured to:
[0284] Determine the reported change amounts corresponding to undetected infected persons and detected infected persons respectively based on the active reporting rate of the control area and the current number of undetected infections;
[0285] Determine the detection change amounts corresponding to undetected exposed persons and detected exposed persons respectively based on the detection rate of the control area, the effective rate of the detection reagent for exposed persons, and the current number of undetected exposed persons;
[0286] Determine the detection change amounts corresponding to undetected infected persons and detected infected persons respectively based on the detection rate of the control area, the effective rate of the detection reagent for infected persons, and the current number of undetected infected persons;
[0287] Determine the quarantine change amount of detected exposed persons and the quarantine change amount of quarantined exposed persons based on the quarantine rate of the control area and the current number of detected exposed persons;
[0288] Determine the quarantine change amount of detected infected persons based on the quarantine rate of the control area and the current number of detected infected persons;
[0289] Determine the quarantine change amount of reported infected persons based on the quarantine rate of the control area and the current number of reported infections;
[0290] Determine the change in the isolation of infected individuals based on the isolation rate of the control area, the current number of detected infected individuals, and the current reported number of infections.
[0291] In one embodiment, the prevention and control strategy includes a detection ratio and an isolation ratio;
[0292] The decision-making model establishment module 20 is further configured to:
[0293] Based on the number of infected individuals, the change in the number of exposed individuals in the control area in the chamber model, and the number of infected individuals in the global area where the control area is located, form a system state, where the system state includes a local state and a global state;
[0294] Initialize the prevention and control strategy, where the prevention and control strategy includes a detection ratio and an isolation ratio;
[0295] Based on the number and change amount of each population type in the chamber model, predict the original state transition probabilities corresponding to the current system state transitioning to various predicted system states;
[0296] Based on the number and change amount of each population type in the chamber model, predict the action state transition probabilities corresponding to the current system state transitioning to various predicted system states after adopting the prevention and control strategy;
[0297] Based on the sum of the number of exposed and infected individuals in each sub-region of the global area, the proportion of the sum of the number of exposed and infected individuals from the control area in each sub-region to the sum of the number of exposed and infected individuals in each sub-region, and the newly added exposed individuals in each sub-region, determine the infection cost of the control area;
[0298] Based on the detection ratio and isolation ratio of the control area, calculate the corresponding detection cost and isolation cost respectively;
[0299] Perform weighted calculation on the infection cost, the detection cost, and the isolation cost of the control area to obtain the reward value of the Markov decision-making model.
[0300] In one embodiment, the decision-making agent includes a policy network and a value network, and the policy network includes a policy function;
[0301] The agent training module 30 is further configured to:
[0302] Initialize the policy function in the policy network, where the policy function is used to determine the corresponding prevention and control strategy according to the input system state data of the control area;
[0303] Initialize the value function in the value network, where the value function is used to determine the corresponding state value according to the local state data and global state data of the input control area.
[0304] In one embodiment, the agent training module 30 is further configured to:
[0305] Collect multiple sets of system state data with continuous time;
[0306] Input each of the system state data into the policy function respectively to obtain the corresponding prevention and control policy;
[0307] Input each of the system state data and the corresponding prevention and control policy into the Markov decision model to obtain the corresponding reward value and transition state;
[0308] Train the policy network and the value network based on each of the system state data, control measurements, reward values, and transition states;
[0309] When the cumulative reward value of the decision-making agent is maximized and converges within a preset time period, stop training and set the decision-making agent as the target agent.
[0310] In one embodiment, the infectious disease prevention and control policy determination device further includes an information reconstruction module, and the information reconstruction module is configured to:
[0311] Obtain the number of exposed persons and the number of infected persons in the current system state data;
[0312] Based on the number of exposed persons and the number of infected persons, the total number of people in the control area, and the flow matrix, reconstruct the reported number of infected persons in the control area to obtain the updated reported number of infected persons;
[0313] Update the number of exposed persons and the number of infected persons in the current system state data according to the updated reported number of infected persons to obtain updated system state data, where the number of exposed persons includes the number of detected exposed persons and the number of undetected exposed persons, and the number of infected persons includes the number of undetected infected persons, the number of detected infected persons, and the reported number of infected persons;
[0314] Input the updated system state data into the target agent to obtain the target prevention and control policy corresponding to the updated system state data.
[0315] The infectious disease prevention and control strategy determination device provided by this application adopts the infectious disease prevention and control strategy determination method in the above-mentioned embodiment, and can solve the technical problem that the prevention and control effect of the prevention and control strategy in a deterministic scenario on the spread of infectious diseases is not good in the real uncertain scenario. Compared with the prior art, the beneficial effects of the infectious disease prevention and control strategy determination device provided by this application are the same as those of the infectious disease prevention and control strategy determination method provided by the above-mentioned embodiment, and other technical features in the infectious disease prevention and control strategy determination device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0316] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the infectious disease prevention and control strategy determination method in Embodiment 1 above.
[0317] The following refers to Figure 7 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any restrictions on the functions and usage scope of the embodiments of this application.
[0318] As Figure 7As shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0319] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0320] The electronic device provided by the present application adopts the method for determining the infectious disease prevention and control strategy in the above-mentioned embodiment, and can solve the technical problem that the prevention and control effect of the prevention and control strategy in a deterministic scenario on the spread of infectious diseases is poor in a real uncertain scenario. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the method for determining the infectious disease prevention and control strategy provided in the above-mentioned embodiment, and other technical features in this electronic device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0321] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0322] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0323] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the infectious disease prevention and control strategy determination method in the above embodiments.
[0324] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0325] The above computer-readable storage medium can be included in an electronic device; or it can exist separately without being assembled into the electronic device.
[0326] The above computer-readable storage medium stores one or more programs, which, when executed by an electronic device, cause the electronic device to perform the following operations: constructing a compartment model for infectious disease transmission based on a preset variety of population types and the change amounts of the quantities of the respective population types, where the population types at least include susceptibles, exposed individuals, infected individuals, quarantined individuals, and recovered individuals; constructing a corresponding Markov decision model according to the compartment model, where the Markov decision model includes system states, prevention and control strategies, original state transition probabilities, action state transition probabilities, and reward values, and the reward values are inversely proportional to the infection costs, detection costs, and quarantine costs corresponding to the prevention and control strategies respectively; establishing a decision agent corresponding to the prevention and control strategy in the Markov decision model, and iteratively training the decision agent based on the Markov decision model until the reward value of the decision agent reaches a preset training target to obtain a target agent; inputting current system state data into the target agent to obtain a target prevention and control strategy corresponding to the current system state data, where the target prevention and control strategy is used to control the spread of infectious diseases.
[0327] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0328] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0329] The modules described in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0330] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned method for determining the infectious disease prevention and control strategy, which can solve the technical problem that the prevention and control effect of the prevention and control strategy in a deterministic scenario on the spread of infectious diseases is poor in a real-world uncertain scenario. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the method for determining the infectious disease prevention and control strategy provided by the above embodiments, and will not be elaborated herein.
[0331] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method for determining the infectious disease prevention and control strategy as described above.
[0332] The computer program product provided by the present application can solve the technical problem that the prevention and control effect of the prevention and control strategy in a deterministic scenario on the spread of infectious diseases is poor in a real-world uncertain scenario. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the method for determining the infectious disease prevention and control strategy provided by the above embodiments, and will not be elaborated herein.
[0333] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for determining an infectious disease prevention and control strategy, characterized in that: The infectious disease prevention and control strategy determination method comprises: Based on the preset multiple population types and the change in the number of each population type, a compartment model of infectious disease transmission is constructed, wherein the population types at least include susceptible people, exposed people, infected people, isolated people and recovered people; According to the compartment model, a corresponding Markov decision model is constructed, wherein the Markov decision model includes system state, prevention and control strategy, original state transition probability, action state transition probability and reward value, and the reward value is inversely proportional to the infection cost, detection cost and isolation cost corresponding to the prevention and control strategy; Establishing a decision-making agent corresponding to the prevention and control strategy in the Markov decision model, and iteratively training the decision-making agent based on the Markov decision model until the reward value of the decision-making agent reaches a preset training target, thereby obtaining a target agent; The current system state data is input into the target intelligent agent to obtain a target prevention and control strategy corresponding to the current system state data, wherein the target prevention and control strategy is used to control the spread of infectious diseases.
2. The infectious disease prevention and control strategy determination method according to claim 1, characterized in that: The exposed persons include detected exposed persons and undetected exposed persons, the infected persons include undetected infected persons, detected infected persons and reported infected persons, the isolated persons include isolated exposed persons and isolated infected persons, and the change in the number of each type of population at least includes the natural change; Before the step of constructing a compartment model of infectious disease transmission based on the preset multiple types of people and the change in the number of each type of people, the method further includes: Determine the natural variation of susceptible persons according to the exposure rate of the control area and the current number of susceptible persons, wherein the exposure rate is determined by the sensing rate of the control area, the flow matrix and the total number of infected persons; Determine the natural change amount of undetected exposed persons according to the exposure rate, the current number of susceptible persons, the infection conversion rate and the current number of undetected exposed persons in the control area; Determine the natural change amount of detected exposed persons according to the infection conversion rate and the current number of detected exposed persons; Determining the natural change amount of undetected infected persons according to the infection conversion rate, the number of currently undetected exposed persons, the recovery rate, and the number of currently undetected infected persons; Determine the natural change amount of detected infected persons according to the infection conversion rate, the number of currently detected exposed persons, the recovery rate, and the number of currently detected infected persons; Determine the natural change in reported infected persons based on the recovery rate and the current number of reported infected persons; Determine the natural change amount of isolated exposed persons according to the infection conversion rate and the number of currently isolated exposed persons; Determine the natural change amount of isolated infected persons based on the infection conversion rate, the number of currently isolated exposed persons, the recovery rate, and the number of currently isolated infected persons; The natural change in the number of recovered persons is determined based on the recovery rate, the current number of undetected infected persons, the current number of detected infected persons, the current number of reported infected persons, and the current number of isolated infected persons.
3. The infectious disease prevention and control strategy determination method according to claim 2, characterized in that: The change amount of each type of population also includes the reported change amount, the detected change amount, and the isolated change amount; Before the step of constructing a compartment model of infectious disease transmission based on the preset multiple types of people and the change in the number of each type of people, the method further includes: Determine the reported changes corresponding to the undetected infected persons and the detected infected persons respectively according to the active reporting rate of the control area and the current number of undetected infected persons; Determine the detection changes corresponding to the undetected exposed persons and the detected exposed persons, respectively, according to the detection rate of the control area, the effectiveness of the detection reagent on the exposed persons, and the number of currently undetected exposed persons; Determine the detection changes corresponding to the undetected infected persons and the detected infected persons respectively according to the detection rate of the control area, the effectiveness of the detection reagent for the infected persons, and the number of currently undetected infected persons; Determining the isolation change amount of the detected exposed persons and the isolation change amount of the isolated exposed persons according to the isolation rate of the control area and the number of currently detected exposed persons; Determining the change in isolation of the detected infected persons according to the isolation rate of the control area and the current number of detected infected persons; Determining the change in isolation of reported infected persons based on the isolation rate of the control area and the current number of reported infections; The isolation change of the isolated infected persons is determined based on the isolation rate of the control area, the current number of detected infected persons and the current number of reported infections.
4. The infectious disease prevention and control strategy determination method according to claim 1, characterized in that: The prevention and control strategies include detection ratio and isolation ratio; The step of constructing a corresponding Markov decision model according to the compartment model comprises: According to the number of infected persons in the control area in the warehouse model, the change in the exposed persons, and the number of infected persons in the global area where the control area is located, a system state is formed, wherein the system state includes a local state and a global state; Initialize the prevention and control strategy, wherein the prevention and control strategy includes a detection ratio and an isolation ratio; Based on the number and change of each type of people in the compartment model, predict the probability of the current system state transferring to the original state corresponding to various predicted system states; Based on the number and change of each type of people in the compartment model, predict the probability of the current system state transferring to the action state corresponding to each of the predicted system states after the prevention and control strategy is adopted; Determine the infection cost of the control area according to the sum of the number of exposed persons and infected persons in each sub-region of the global area, the ratio of the sum of the number of exposed persons and infected persons from the control area in each sub-region to the sum of the number of exposed persons and infected persons in each sub-region, and the newly exposed persons in each sub-region; Based on the detection ratio and isolation ratio of the control area, calculating the corresponding detection cost and isolation cost respectively; The infection cost, the detection cost and the isolation cost of the control area are weightedly calculated to obtain a reward value of the Markov decision model.
5. The infectious disease prevention and control strategy determination method according to claim 1, characterized in that: The decision-making agent includes a strategy network and a value network, and the strategy network includes a strategy function; The step of establishing a decision-making agent corresponding to the prevention and control strategy in the Markov decision model includes: Initializing a policy function in the policy network, wherein the policy function is used to determine a corresponding prevention and control policy according to input system status data of the control area; Initialize a value function in the value network, wherein the value function is used to determine a corresponding state value according to local state data and global state data of an input control area.
6. The infectious disease prevention and control strategy determination method according to claim 5, characterized in that: The training goal is to maximize and converge the cumulative reward value. The step of iteratively training the decision agent based on the Markov decision model until the reward value of the decision agent reaches the preset training goal, and obtaining the target agent includes: Collect multiple sets of system status data in a continuous time period; Input each of the system status data into the strategy function to obtain the corresponding prevention and control strategy; Inputting each of the system state data and the corresponding prevention and control strategy into the Markov decision model to obtain the corresponding reward value and transfer state; Training the policy network and the value network based on the system state data, control measurements, reward values, and transition states; When the cumulative reward value of the decision-making agent within a preset time period is maximized and converged, the training is stopped and the decision-making agent is set as the target agent.
7. The method for determining infectious disease prevention and control strategies according to any one of claims 1 to 6, characterized in that: Before the step of inputting the current system status data into the target intelligent agent to obtain the target prevention and control strategy corresponding to the current system status data, the method further includes: Get the number of exposed persons and infected persons in the current system status data; Based on the number of exposed persons and infected persons, the total number of people in the control area, and the flow matrix, information on the number of reported infected persons in the control area is reconstructed to obtain an updated number of reported infected persons; The number of exposed persons and the number of infected persons in the current system status data are updated according to the updated number of reported infected persons to obtain updated system status data, wherein the number of exposed persons includes the number of detected exposed persons and the number of undetected exposed persons, and the number of infected persons includes the number of undetected infected persons, the number of detected infected persons, and the number of reported infected persons; The updated system status data is input into the target intelligent agent to obtain the target prevention and control strategy corresponding to the updated system status data.
8. A device for determining infectious disease prevention and control strategies, characterized in that: The infectious disease prevention and control strategy determination device comprises: A compartment model building module is used to build a compartment model of infectious disease transmission based on a plurality of preset population types and the change in the number of each population type, wherein the population types at least include susceptible persons, exposed persons, infected persons, isolated persons, and recovered persons; A decision model building module, used to build a corresponding Markov decision model according to the compartment model, wherein the Markov decision model includes system state, prevention and control strategy, original state transition probability, action state transition probability and reward value, and the reward value is inversely proportional to the infection cost, detection cost and isolation cost corresponding to the prevention and control strategy; An agent training module is used to establish a decision agent corresponding to the prevention and control strategy in the Markov decision model, and iteratively train the decision agent based on the Markov decision model until the reward value of the decision agent reaches a preset training target, thereby obtaining a target agent; The strategy output module is used to input the current system status data into the target intelligent agent to obtain the target prevention and control strategy corresponding to the current system status data, wherein the target prevention and control strategy is used to control the spread of infectious diseases.
9. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the infectious disease prevention and control strategy determination method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the infectious disease prevention and control strategy determination method according to any one of claims 1 to 7 are implemented.