A method and system for individual mobile intervention in infectious disease prevention and control

By constructing an infectious disease prevention and control model for individual mobile intervention, using graph neural network and deep reinforcement learning, accurately analyze the impact of implicitly infected people, formulate individual prevention and control measures, the impact of implicitly infected people on infectious disease prevention and control is solved, and the effect of reducing the number of infected people and reducing costs under low travel intervention is achieved.

CN113889282BActive Publication Date: 2025-08-12TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110858829.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-28
Publication Date
2025-08-12
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

The existing methods for infectious disease prevention and control fail to effectively consider the impact of hidden infections on infectious disease prevention and control, resulting in out-of-control infectious diseases and rising medical pressure, and the rough prevention and control measures affect travel and cause economic losses.

Method used

By obtaining the historical status information and relationship information of individual users, using graph neural networks, long-term and short-term neural networks and intelligent models, an individual mobile intervention infectious disease prevention and control model is constructed, the impact of hidden infections is accurately analyzed, and individualized prevention and control measures are formulated, combining deep reinforcement learning to optimize the number of infected people and travel intervention.

Benefits of technology

With low travel intervention, effectively reduce the number of infected people, reduce the social and economic costs of infectious diseases, and achieve precise infectious disease prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889282B_ABST
    Figure CN113889282B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for individual mobile intervention in infectious disease prevention and control. The method comprises: obtaining daily historical status information and individual relationship information of individual users in a target city within a preset time interval; inputting the historical status information and individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city. The trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-short-term neural network, and an intelligent agent based on a partially observable Markov decision process using the individual status information and the sample individual relationship information of sample users. The sample individual status information includes health status information of individuals who have transitioned from latently infected to overtly infected. The present invention can minimize the number of infections while maintaining low travel intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of individual infectious disease prevention and control, and in particular to a method and system for individual mobile intervention in infectious disease prevention and control. Background Art

[0002] Individual Mobility Intervention Epidemic Control refers to the use of individual historical trajectories and individual characteristics to implement travel intervention measures of varying intensity for some high-risk close contact groups, thereby achieving the effect of infectious disease prevention and control. However, in the actual prevention and control of infectious diseases, latently infected people in the incubation period cannot be discovered through non-drug methods, and these latently infected people are often highly contagious, which can easily lead to the loss of control of infectious diseases and a sharp increase in medical pressure, which poses a great challenge to the prevention and control of infectious diseases. Existing practical applications often only adopt some crude prevention and control measures. Although these infectious disease prevention and control measures can limit the spread of infectious diseases to a certain extent, they will seriously affect people's travel and transportation, thereby causing huge economic losses.

[0003] The current intervention methods for infectious disease prevention and control mainly include obtaining the user's physical parameter information and the user's identity information to determine whether the user's physical parameter information is abnormal. If the user's physical parameter information is found to be abnormal, the user's location information is determined based on the user's identity information, and then the user's movement trajectory is determined based on the user's location information, thereby facilitating infectious disease control; or, based on the positioning displacement and time difference between the user's real-time positioning information, the user's historical instantaneous velocity vector and acceleration are calculated, solved and plotted, and then the user's mobile tool and mobile behavior are determined, so as to more accurately analyze the infection site and infection mechanism during the infectious disease period; or, based on previously known epidemic symptoms, an infectious disease situation prediction model is constructed, and the model is used to infer the duration of the epidemic, providing reference data for disease prevention and control decision-making.

[0004] However, in the current infectious disease prevention and control, the impact of latent infections on infectious disease prevention and control is ignored, and there is a lack of an existing framework to model the impact of latent infections, resulting in the inability to accurately implement infectious disease prevention and control intervention measures. Summary of the Invention

[0005] In response to the problems existing in the prior art, the present invention provides a method and system for individual mobile intervention in infectious disease prevention and control.

[0006] The present invention provides a method for individual mobile intervention in infectious disease prevention and control, comprising:

[0007] Obtain daily historical status information and individual relationship information for individual users in the target city within a preset time interval. The historical status information includes the user's movement trajectory, health status, intervention status, infection probability, and number of acquaintance contacts. The individual relationship information includes relationships between individuals, relationships between individuals and regions, and relationships between regions.

[0008] Inputting the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city;

[0009] Among them, the trained individual mobile intervention infectious disease prevention and control model is obtained by training the graph neural network, long-term and short-term neural network and intelligent agent based on the individual status information and sample individual relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process; the individual status information of sample users includes the health status information of latently infected people converted to overtly infected people.

[0010] According to a method for individual mobile intervention in infectious disease prevention and control provided by the present invention, the trained individual mobile intervention in infectious disease prevention and control model is obtained by the following steps:

[0011] According to the sample individual relationship information, the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions are obtained to construct a first training sample set;

[0012] Inputting the first training sample set into a graph neural network for training, obtaining a graph neural network based on individual contact risk, as well as the individual regular commuting characteristics of each sample user on each day and the social relationships between the sample users;

[0013] Constructing a second training sample set based on the individual regular commuting characteristics of the sample users, the social relationships between the sample users, and the corresponding historical status sample information; the historical status sample information includes historical status sample information of uninfected persons, historical status sample information of latently infected persons, historical status sample information of overtly infected persons, and historical status sample information of recovered persons;

[0014] Inputting the second training sample set into a long-short-term neural network for training, thereby obtaining a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents a change trend of the individual state information of the sample user within a preset time period obtained through prediction;

[0015] Based on hierarchical reinforcement learning and the sample user individual state representation based on the time state sequence, the intelligent agent is trained to obtain a pre-trained intelligent agent; the hierarchical reinforcement learning model is constructed based on the FuN model;

[0016] According to the graph neural network based on individual contact risk, the time state sequence representation model and the pre-trained intelligent agent, a trained individual mobile intervention infectious disease prevention and control model is obtained.

[0017] According to the present invention, a method for individual mobile intervention in infectious disease prevention and control is provided. The intelligent agent is constructed with the goal of minimizing the number of infected people and the occurrence of the lowest intervention. The intelligent agent is used to observe the state representation of the sample user individuals based on the time state sequence, and to determine the prevention and control intervention measures for each user individual based on the observation results. The observations include the sample user individual's health status observation, intervention status observation, infection probability observation, and number of acquaintance contacts observation. The infection probability observation is obtained by estimating the movement trajectory of the sample user individual.

[0018] The agent's reward is expressed as:

[0019]

[0020] Among them, r t represents the reward on the tth day, L represents the total number of days of infectious disease prevention and control, ΔI represents the change in the number of infected people during the infectious disease prevention and control period, θ I represents the tolerance threshold of the medical system, ΔQ represents the change in travel intervention during the infectious disease prevention and control period, and θ Q Indicates the tolerance threshold of the economic system.

[0021] According to a method for individual mobile intervention in infectious disease prevention and control provided by the present invention, the graph neural network based on individual contact risk is constructed through a graph convolutional network and a GraphSAGE algorithm.

[0022] According to a method for individual mobile intervention in infectious disease prevention and control provided by the present invention, the time state sequence representation model is obtained by improving the long short-term memory network through a variational autoencoder.

[0023] According to an individual mobile intervention infectious disease prevention and control method provided by the present invention, the hierarchical reinforcement learning includes a Manager network and a Worker network, wherein the loss function formula of the Manager network is:

[0024]

[0025]

[0026]

[0027]

[0028]

[0029] Among them, La m Represents the behavioral loss function of the Manager network, represents the advantage function of the Manager network, represents the reward of the Manager network, γ represents the hyperparameter, Represents the state value of the Manager network's observation history on day t+1, Represents the state value of the observation history of the Manager network on day t, d w represents the similarity between the observation and the target, represents the observation of the Manager network on day t+c, represents observations based on incomplete information, represents the action probability distribution ratio of the new and old policy networks of the Manager network on day t, p represents the probability value of the policy output action, g t represents an abstract internal goal, represents the observation history of the Manager network on day t, p old represents the probability distribution of the old policy network, ε represents the hyperparameter, Represents the loss function of the time state sequence representation model, Lc m Represents the judgment loss function of the Manager network;

[0030] The loss function formula of the Worker network is:

[0031]

[0032]

[0033]

[0034]

[0035] Among them, La w represents the behavioral loss function of the Worker network, represents the advantage function of the Worker network, Represents the reward of the Worker network, Represents the state value of the observation history of the Worker network on day t+1, Represents the state value of the observation history of the Worker network on day t, It represents the ratio of the action probability distribution of the new and old policy networks of the Worker network on day t. represents the observation history of the tth day in the Worker network, Represents the loss function of the time state sequence representation model, Lc w Represents the evaluation loss function of the Worker network.

[0036] The present invention also provides an individual mobile intervention infectious disease prevention and control system, comprising:

[0037] An individual daily historical information acquisition module is used to obtain daily historical status information and individual relationship information of individual users in the target city within a preset time interval. The historical status information includes the user's movement trajectory, health status, intervention status, infection probability, and number of acquaintance contacts. The individual relationship information includes relationships between individuals, relationships between individuals and regions, and relationships between regions.

[0038] a prevention and control intervention strategy generation module, configured to input the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city;

[0039] Among them, the trained individual mobile intervention infectious disease prevention and control model is obtained by training the graph neural network, long-term and short-term neural network and intelligent agent based on the individual status information and sample individual relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process; the individual status information of sample users includes the health status information of latently infected people converted to overtly infected people.

[0040] According to the present invention, a personal mobile intervention infectious disease prevention and control system is provided, the system further comprising:

[0041] A first sample set construction module is used to obtain the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions based on the sample individual relationship information, and construct a first training sample set;

[0042] A first training module is configured to input the first training sample set into a graph neural network for training, thereby obtaining a graph neural network based on individual contact risk, as well as the individual regular commuting characteristics of each sample user on each day and the social relationships between the sample users;

[0043] A second sample set construction module is configured to construct a second training sample set based on the individual regular commuting characteristics of the sample users, the social relationships between the sample users, and the corresponding historical status sample information; the historical status sample information includes historical status sample information of uninfected persons, historical status sample information of latently infected persons, historical status sample information of overtly infected persons, and historical status sample information of recovered persons;

[0044] a second training module configured to input the second training sample set into a long-short-term neural network for training to obtain a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents a predicted trend of individual state information changes of the sample user within a preset time period;

[0045] A third training module is configured to train an agent based on hierarchical reinforcement learning and the sample user individual state representation based on the time state sequence to obtain a pre-trained agent; the agent is constructed based on the FuN model;

[0046] A model deployment module is used to obtain a trained individual mobile intervention infectious disease prevention and control model based on the graph neural network based on individual contact risk, the time state sequence representation model and the pre-trained intelligent agent.

[0047] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the steps of any of the above-mentioned individual mobile intervention infectious disease prevention and control methods are implemented.

[0048] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned individual mobile intervention infectious disease prevention and control methods.

[0049] The individual mobile intervention infectious disease prevention and control method and system provided by the present invention analyzes the historical trajectories of individual users in the target area and the characteristics of the individual users themselves, and models the impact of latently infected people on infectious diseases, thereby conducting precise travel interventions of varying degrees for some high-risk groups. The ultimate goal is to minimize the number of infections with lower travel interventions. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 A schematic diagram of the process of the individual mobile intervention infectious disease prevention and control method provided by the present invention;

[0052] Figure 2 A schematic diagram of the process of infectious disease prevention and control by intervening in individual travel provided by the present invention;

[0053] Figure 3 This is a schematic diagram of the principle of the workflow of individual mobile intervention in infectious disease prevention and control provided by the present invention;

[0054] Figure 4 A schematic diagram of the process of the feature representation method of the graph neural network based on individual exposure risk provided by the present invention;

[0055] Figure 5 A schematic diagram of the construction process of the graph neural network based on individual exposure risk provided by the present invention;

[0056] Figure 6 A schematic diagram of the structural principle of the VLSTM provided by the present invention;

[0057] Figure 7 A schematic diagram of the structural principle of hierarchical reinforcement learning provided by the present invention;

[0058] Figure 8 This is a schematic diagram of the structure of the individual mobile intervention infectious disease prevention and control system provided by the present invention;

[0059] Figure 9 This is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0060] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0061] Current approaches to individual infectious disease prevention and control primarily include graph-based approaches, infectious disease prediction-based approaches, heuristic algorithm-based approaches, and reinforcement learning-based approaches. Graph-based approaches often treat individuals as nodes, with edges representing their contacts and connections. These approaches leverage the graph's topological knowledge to identify individuals at high risk of infection and isolate them. However, these approaches struggle to effectively leverage individual characteristics for infectious disease prevention and control, hindering precise individual control. Infectious disease prediction-based approaches typically isolate potentially infected individuals based on predictions of individual health status using algorithms such as gradient boosting decision trees (GBDTs) or long short-term memory networks (LSTMs). However, these approaches fail to consider the long-term impact of current measures and often fall into local optima. While heuristic algorithm-based approaches can effectively leverage individual information for infectious disease prevention and control, they often overlook the impact of contacts and connections between individuals on infection control. Although current reinforcement learning-based methods can effectively utilize individual information for long-term prevention and control, and related work also takes into account the impact of inter-individual contact on infectious disease prevention and control, they often ignore the impact of latent infections on infectious disease prevention and control and lack dedicated modules to model the impact of latent infections. In addition, because the reinforcement learning reward function measures the final effect of infectious disease prevention and control, the corresponding reward must wait until the infectious disease prevention and control is completed before receiving the corresponding reward. Therefore, this is a sparse reward problem, which current reinforcement learning-based algorithms rarely take into account.

[0062] The present invention adopts the framework of deep reinforcement learning. It first encodes the real-time characteristic attributes of individuals and uses graph neural networks to model the social relationships and commuting characteristics of individuals. Then, it uses long short-term memory networks (preferably, the present invention uses Variational LSTM) to model the individual feature sequences, thereby strengthening the individual state representation. The dual goals of the number of infections and the occurrence of intervention are modeled as a reinforcement learning reward function, and the reinforcement learning method is used to explore an optimal infectious disease prevention and control strategy.

[0063] Figure 1 This is a flow chart of the individual mobile intervention infectious disease prevention and control method provided by the present invention, such as Figure 1 As shown, the present invention provides a method for individual mobile intervention in infectious disease prevention and control, comprising:

[0064] Step 101: Obtain daily historical status information and individual relationship information of individual users in the target city within a preset time interval. The historical status information includes the individual user's movement trajectory, health status, intervention status, infection probability, and number of contacts with acquaintances. The individual relationship information includes relationships between individuals, relationships between individuals and regions, and relationships between regions.

[0065] In the present invention, relevant information of an individual user can be collected through a mobile terminal used by the user. For example, through a mobile communication device such as a mobile phone, the user's daily travel trajectory, health status, intervention status, infection probability and number of acquaintance contacts within a preset time interval are collected, where the health status and intervention status can be obtained through information such as the user's health code; the infection probability can be judged by the user's appearance trajectory and estimated based on the infectious disease risk level of the area where the user is located (for example, if the user has been to a medium-risk infectious disease area, the infection probability is estimated to be 40% to 60%); the number of acquaintance contacts can be obtained by counting the number of times the user is in the same area within a certain period of time. If the number exceeds a preset threshold, it is determined that the two individuals are in an acquaintance relationship. For example, if a user individual is in the same office building area with five other user individuals between 9 am and 6 pm from Monday to Friday, it is determined that the user individual and the other five user individuals are in an acquaintance relationship, and the number of acquaintance contacts of the user individual per day is 5. Furthermore, when counting the number of acquaintances that an individual user has contact with, we can obtain the relationships between individuals (for example, acquaintance relationships, stranger relationships, etc.), the relationships between individuals and regions (for example, the user works in a certain area or the user goes to a certain hospital for medical treatment), and the relationships between regions (for example, the relationship between residential areas and work areas, the relationship between residential areas and entertainment venues, etc.).

[0066] Step 102: Input the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city;

[0067] Among them, the trained individual mobile intervention infectious disease prevention and control model is obtained by training the graph neural network, long-term and short-term neural network and intelligent agent based on the individual status information and sample individual relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process; the individual status information of sample users includes the health status information of latently infected people converted to overtly infected people.

[0068] In this invention, the trained individual mobile intervention infectious disease prevention and control model targets a target city with a population of M and N regions, taking into account the differences in characteristics of different individuals, enabling more accurate and effective infectious disease prevention and control, and reducing unnecessary travel interventions. During the model training phase, the relationships between individuals are considered and modeled based on commuting characteristics, which can estimate the risk of individual infection and the probability of transmission, thereby further improving the individual's status characteristics; considering and modeling the impact of latently infected individuals can reasonably estimate the infectious disease transmission process and provide support for infectious disease prevention and control; based on a deep reinforcement learning model, it can consider the long-term impact of current infectious disease prevention and control measures, and can comprehensively consider the dual-objective optimization of minimizing the number of infections and minimizing travel interventions.

[0069] Furthermore, in the present invention, the health status of each individual in the city is defined as: Susceptible (susceptible to infection, indicating that the individual has not been infected yet), Asymptomatic (latent infection), Symptomatic (overt infection) and Recovered (recovered). Specifically, individuals with health status of Asymptomatic and Symptomatic are infected individuals. People who go to the same place at the same time have a certain probability of contacting these infected individuals and even being infected. After infection, their health status will change from Susceptible to Asymptomatic. Individuals in the Asymptomatic state will become Symptomatic after the incubation period. Individuals in the Symptomatic state will be sent to the hospital immediately and will become Recovered after the recovery period. Policy makers (intelligent agents) cannot discover the difference between the Susceptible and Asymptomatic states of individuals through non-drug methods. Therefore, the individual mobile intervention infectious disease prevention and control model provided by the present invention aims to select corresponding prevention and control measures for individuals with a status of Susceptible or Asymptomatic, so that the number of infections and travel interventions can be minimized. According to different health statuses, the present invention defines four types of prevention and control measures: No Intervention (no isolation), Confine (no contact with people outside the residence, i.e., isolation between communities), Quarantine (no contact with strangers, i.e., home isolation), and Isolate (no contact with anyone). Figure 2 This is a schematic diagram of the process of preventing and controlling infectious diseases by intervening in individual travel provided by the present invention, such as Figure 2As shown in the figure, based on the movement trajectory of individual users and the health status of individual users in the area, corresponding infectious disease prevention and control strategies are formulated. For example, those who are obviously infected are sent to the hospital for treatment, and those who are at risk of infection are sent to the CDC (center for disease control and prevention) for corresponding isolation.

[0070] In the present invention, the goal of the individual mobility intervention infectious disease prevention and control model is to simultaneously reduce the number of infections and minimize travel interventions. However, these two optimization goals are contradictory, because the lower the number of infections, the more travel interventions are often required. Therefore, from the perspective of social cost, the present invention unifies the two optimization goals into infection-spread-cost (Icost) and mobility-intervention-cost (Mcost). Once the number of infections exceeds a certain threshold (defined in the present invention as the tolerance threshold of the medical system, for example, an increase in the number of hospitalizations), the medical system will be broken down, resulting in a rapid increase in social costs; on the other hand, when the restrictions on people's travel are higher than a certain threshold (defined in the present invention as the tolerance threshold of the economic system, for example, excessive travel interventions), the social production system will also be paralyzed, which will also lead to a rapid increase in social costs. Therefore, this embodiment uses exponential form to express the above costs:

[0071] Q=λ h *N h +λ i *N i +λ q *N q +λ c *N c ;

[0072]

[0073]

[0074] Among them, I represents the total number of infections during the period of infectious disease prevention and control, Q represents the total number of travel interventions during the period of infectious disease prevention and control, and N h ,N i ,N q ,N c They represent the total number of hospitalized (number of hospitalized patients), isolated (number of people who need to be isolated from everyone), quarantined (number of people who need to be isolated from strangers), and confined (number of people who are restricted) during the period of infectious disease prevention and control, θ I represents the tolerance threshold of the medical system, θ Q represents the tolerance threshold of the economic system, λ h ,λ i,λ q ,λ c Represents the correlation coefficient.

[0075] In real life, latently infected people exist objectively. However, the agent (strategist) can only obtain observations of the true health status of latently infected people in a non-drug-based way. t (o t = represents the observation on day t). Therefore, the present invention models the infectious disease prevention and control problem as a partially observable Markov decision process (POMDP). POMDP describes the state s where the agent (strategist) cannot obtain complete information. t , we can only get observations o with information loss t Therefore, the goal of individual-based infectious disease prevention and control is to observe the t , select corresponding travel intervention measures for individuals with Susceptible or Asymptomatic status to minimize social cost.

[0076] Figure 3 The schematic diagram of the principle of the workflow of individual mobile intervention in infectious disease prevention and control provided by the present invention is as follows: Figure 3 As shown, the present invention establishes a reinforcement learning framework to obtain optimal infectious disease prevention and control strategies. Specifically, the model uses days as the minimum time interval for strategy implementation. It first connects the collected individual states and the relationships between individuals and regions as input. After passing through an agent module, it outputs prevention and control measures for each individual. During the training process, the model aims to train a reinforcement learning network to output specific prevention and control measures for different individuals every day, ultimately minimizing the social cost of the entire infectious disease prevention and control period.

[0077] The individual mobile intervention infectious disease prevention and control method provided by the present invention analyzes the historical trajectories of individual users in the target area and the characteristics of the individual users themselves, and models the impact of latently infected people on infectious diseases, thereby conducting precise travel interventions of varying degrees for some high-risk groups. The ultimate goal is to minimize the number of infections with lower travel interventions.

[0078] Based on the above embodiment, the trained individual mobile intervention infectious disease prevention and control model is obtained by the following steps:

[0079] Step 201: Based on the sample individual relationship information, the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions are obtained to construct a first training sample set.

[0080] In step 202, the first training sample set is input into a graph neural network for training to obtain a graph neural network based on individual contact risk, as well as the individual regular commuting characteristics of each sample user on each day and the social relationships between individual sample users.

[0081] In the present invention, since the individual user trajectory information available in real life is often incomplete and coarse-grained, and infection after close contact between individuals is also a probabilistic event, it is very difficult to accurately measure contact between individuals and the spread of infectious diseases. In addition, tracking all contacts and infections caused by latently infected people is also extremely challenging. In order to solve this challenge, the present invention proposes a graph neural network based on individual contact risk (Individual Contact-Risk, abbreviated as GNN) to model individual regular commuting and social relationships between individuals, thereby estimating the individual's infection risk. The GNN regards individuals and urban areas as two types of nodes. In the GNN, there are edges between nodes of the same type and nodes of different categories. Figure 4 This is a process diagram of the feature representation method of the graph neural network based on individual exposure risk provided by the present invention, such as Figure 4 As shown in the figure, through GNN, it is possible to model the regular commuting and social relations between individuals through the relationships between individuals, individuals and regions, and regions and regions, so as to better estimate the contact infection risk between individuals.

[0082] Preferably, the graph neural network based on individual exposure risk is constructed by graph convolutional network and GraphSAGE (Graph Sample and aggregate) algorithm. Represents the regional node features of the k-th layer GNN network, Represents the characteristics of individual nodes in the k-th layer GNN network. The detailed calculation of the GNN network layer is as follows:

[0083]

[0084] f c =softmax(f m );

[0085]

[0086]

[0087]

[0088] Among them, f mrepresents the individual regular commuting calculated through the visit history between individuals and regions, f c represents the regular commuting distribution between individuals and regions, A d Indicates the acquaintance relationship between individuals, A r Indicates the commuting relationship between regions, The graph embedding of individual features at the k-1th time step is expressed as W k ,B k represents the learnable parameters.

[0089] Figure 5 This is a schematic diagram of the construction process of the graph neural network based on individual exposure risk provided by the present invention, which can be referred to Figure 5 As shown, in step 1, the relationship between individuals is used as an edge to calculate the characteristics of individual nodes. The characteristics are calculated by formula Calculated; in step 2, the individual characteristics of the commuting relationship with the region are aggregated to calculate the characteristics of the regional node, which is calculated by the formula Calculated; in step 3, the commuting relationship between regions is used as an edge to calculate the representation of regional nodes. This feature is calculated by formula Calculated; in step 4, by aggregating the regional characteristics of individual regular commuting, the final individual node feature representation is obtained, which is expressed by the formula get.

[0090] Step 203: construct a second training sample set based on the regular commuting characteristics of the sample users, the social relationships between the sample users, and the corresponding historical status sample information; the historical status sample information includes historical status sample information of uninfected persons, historical status sample information of latently infected persons, historical status sample information of overtly infected persons, and historical status sample information of recovered persons;

[0091] In step 204, the second training sample set is input into the long-short term neural network for training to obtain a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents the individual state information change trend of the sample user within a preset time period obtained by prediction.

[0092] Since the status of latently infected people is unobservable, we can directly use the current observation o t Making decisions will result in decision errors. In the present invention, the sample observation history h t = (o t ,o t-1,…o0), contains the information of latent infection (asymptomatic) transformed into obvious infection (symptomatic), so as to better extract the characteristics of individuals. Therefore, the intelligent agent of the present invention should be based on the observation history h t Instead of the current observation o t Long Short-Term Memory (LSTM) is often used to process state history information. However, the effect of LSTM is often affected by state noise caused by the dynamic environment. Therefore, based on the above embodiment, the time state sequence representation model is obtained by improving the long short-term memory network through the Variational Auto Encoder (VAE), namely Variational LSTM (VLSTM), which is a time recursive latent variable model that can be obtained through a random latent variable z t Learn to encode and predict complex sequences of observations x t , so that the model can predict the changing trend of individual status information over a period of time based on the daily regular commuting characteristics of individual sample users, the social relationships between individual sample users, and the corresponding historical status sample information.

[0093] Furthermore, the inference model is given x t and d t To estimate the latent variable z t , the inference model is represented by the parameter φ, as follows:

[0094]

[0095]

[0096] d t It can be obtained by the following formula:

[0097] d t =f LSTM (d t-1 ;z t ,x t );

[0098] In this invention, the generative model of VLSTM is used to help predict the distribution of the next state variable of VLSTM and calculate the loss of VLSTM. The generative model is represented by the parameter θ, which is as follows:

[0099]

[0100] Among them, φ encoder and θ priorRepresents parameter mapping, similar to the parameters of a neural network, d t is the state variable of LSTM. Figure 6 Schematic diagram of the structural principle of the VLSTM provided by the present invention.

[0101] It should be noted that the present invention uses VLSTM as the encoder of the entire individual mobile intervention infectious disease prevention and control model, which can be referred to Figure 4 As shown, the variational loss of VLSTM is as follows:

[0102]

[0103] Among them, D KL Indicates q φ (z t ) and p θ (z t )’s KL divergence, q φ represents the probability distribution of the inferred model, p θ Represents the probability distribution of the prior model.

[0104] Step 205 : Based on hierarchical reinforcement learning and the individual state representation of sample users based on the time state sequence, the intelligent agent is trained to obtain a pre-trained intelligent agent; the hierarchical reinforcement learning model is constructed based on the FuN model.

[0105] In this paper, the problem of dynamic infectious disease prevention and control strategies is modeled as a POMDP problem and solved using a reinforcement learning (RL) algorithm. Specifically, the present invention uses each day as the decision-making interval. The intelligent agent is constructed with the goal of minimizing the number of infected people and the occurrence of interventions as the goal. It is used to observe the state representations of individual sample users based on the time state sequence, and then determine the prevention and control intervention measures for each individual user based on the observation results. The observations include observations of the health status of the individual sample users, observations of the intervention status, observations of the infection probability, and observations of the number of acquaintances contacted. The infection probability observations are estimated by the movement trajectory of the individual sample users.

[0106] The design of the agent, observation, state, action, and reward function in the POMDP setting is as follows:

[0107] Agent: A global agent is designed to make decisions. That is, the agent's observations, states, and actions are specific to everyone.

[0108] Observation: The observation of an agent is a concatenation of each individual's characteristics. For each individual, observations include health status, intervention status, number of acquaintances, and infection probability. The observation design takes into account characteristics related to latently infected individuals. The number of acquaintances reflects the risk of latently infected individuals spreading the virus to others, as the probability of infection is higher through contact with acquaintances. The individual infection probability estimates inter-individual contact based on the individual's recent historical trajectory, thereby measuring the risk of infection from random socially driven contact. Due to the objective existence of latently infected individuals, the health status of the individuals in the observation is partially observable.

[0109] State: The state of an agent measures observations under complete information. The definition of state is the same as observation, but in the state, the true health status of latently infected individuals can be observed.

[0110] Action: For the agents in the individual mobile intervention infectious disease prevention and control model, each day's action determines the corresponding prevention and control measures for each individual. Actions include No Intervention (no isolation), Confine (no contact with people outside the residence, i.e., intra-community isolation), Quarantine (no contact with strangers, i.e., home isolation), and Isolate (no contact with anyone).

[0111] Reward: The optimization goal is to minimize the social cost during the entire epidemic prevention and control period, and this cost can only be obtained on the last day of epidemic prevention and control. Assuming that the number of days of epidemic prevention and control is L, and setting the reward to a negative social cost, the reward of the agent is expressed as:

[0112]

[0113] Among them, r t represents the reward on the tth day, L represents the total number of days of infectious disease prevention and control, ΔI represents the change in the number of infected people during the infectious disease prevention and control period, θ I represents the tolerance threshold of the medical system, ΔQ represents the change in travel intervention during the infectious disease prevention and control period, and θ Q Indicates the tolerance threshold of the economic system.

[0114] Furthermore, since a reward can only be obtained on the last day of infectious disease prevention and control, travel intervention in infectious disease prevention and control is a sparse reward problem. For this problem, the present invention sets a daily reward, which can achieve the problem of partially observable rewards, because the newly added latent infections in the number of newly infected people every day are unobservable. In order to solve the sparse and partially observable problems of rewards, the present invention introduces hierarchical reinforcement learning (HRL) into the individual mobile intervention infectious disease prevention and control model. HRL includes a Manager module (Manager network) and a Worker module (Worker network). Preferably, the hierarchical reinforcement learning model of the present invention is constructed based on the FuN (FeUdal Networks for Hierarchical Reinforcement Learning) model. On the basis of the FuN model, the present invention adds a graph neural network containing individual contact risk and an Agent module of VLSTM. Figure 7 The schematic diagram of the structural principle of the hierarchical reinforcement learning provided by the present invention can be referred to Figure 7 As shown, the principle structure of hierarchical reinforcement learning is specifically explained:

[0115] Manager module: The Manager module does not interact directly with the environment. It inputs the observation history and outputs the target goal to the Worker. This goal constitutes the intrinsic reward that guides the Worker's learning. The external reward that measures the overall effectiveness of infectious disease prevention and control is used to train the Manager. Because most latently infected people will become overtly infected during the long-term prevention and control process, this overall reward can help solve some of the observability problems of the reward. The specific settings of the Manager module are as follows: m : On day t, the Manager's observation is the same as the definition of Observation in the above embodiment, including health status, intervention status, number of acquaintances, and infection probability.

[0116] Goal: In the setting of hierarchical reinforcement learning, the Manager’s action is to generate an abstract internal goal g for the Worker. t , abbreviate Agent Module as AM, then:

[0117]

[0118] Reward m :Manager's Reward The rewards are the same as those defined in the above embodiment, namely:

[0119]

[0120] Worker module: The Worker is the module that directly interacts with the environment. It inputs the observation history and outputs actions to the environment. The intrinsic rewards provided by the Manager are used to train the Manager and are available daily, thus solving the sparse reward problem. The specific settings of the Worker are as follows:

[0121] Observation w : On day t, the worker's observation is the same as the definition of Observation in the above embodiment, including health status, intervention status, number of acquaintances, and infection probability.

[0122] Action w :Worker action The definition of Action is the same as in the above embodiment. One of the four actions is output for each individual, namely No Intervention (no isolation), Confine (no contact with people outside the residence, i.e., isolation within the community), Quarantine (no contact with strangers, i.e., home isolation), and Isolate (no contact with anyone). Figure 7 As shown, this action is caused by and g t jointly decided,

[0123]

[0124]

[0125] in, Represents the observation history after passing through the agent module of the Worker network.

[0126] Reward w :Worker through internal rewards To train, this reward is to guide the Worker to achieve the goal given by the Manager. The specific formula is:

[0127]

[0128] Among them, d cos (α,β)=α T β(|α|·|β|), represents the cosine similarity of two phasors, and c is a hyperparameter.

[0129] Furthermore, the training process of reinforcement learning is as follows:

[0130] Action Space Exploration Strategy: To improve the efficiency of the agent's exploration of the action space, the exploration of the action space is constrained based on the individual infection probability mentioned in the above embodiment (i.e., estimating the contact between individuals based on their historical trajectories over the past few days, thereby measuring the risk of infection from socially driven random movement). Assuming that individuals with a higher probability of infection should be subject to more stringent prevention and control measures, this ensures that individuals with a high infection rate are also considered high-risk, thereby reducing unnecessary exploration.

[0131] The individual mobile intervention infectious disease prevention and control model of the present invention uses the proximal policy optimization algorithm (PPO) for both the Manager module and the Worker module. The model loss includes the loss of the RL model and the variational loss of the VLSTM. m ,Lc m ,La w ,Lc w and h t To represent the Manager's actor loss (behavior loss function), the Manager's critic loss (critic loss function), the Worker's actor loss, the Worker's critic loss and the observation history, respectively. Based on the above embodiment, the hierarchical reinforcement learning includes a Manager network and a Worker network, wherein the loss function formula of the Manager network is:

[0132]

[0133]

[0134]

[0135]

[0136]

[0137] Among them, La m Represents the behavioral loss function of the Manager network, represents the advantage function of the Manager network, represents the reward of the Manager network, γ represents the hyperparameter, Represents the state value of the Manager network's observation history on day t+1, Represents the state value of the observation history of the Manager network on day t, d wrepresents the similarity between the observation and the target, represents the observation of the Manager network on day t+c, represents observations based on incomplete information, represents the action probability distribution ratio of the new and old policy networks of the Manager network on day t, p represents the probability value of the policy output action, g t represents an abstract internal goal, represents the observation history of the Manager network on day t, p old represents the probability distribution of the old policy network, ε represents the hyperparameter, Represents the loss function of the time state sequence representation model, Lc m Represents the judgment loss function of the Manager network;

[0138] The loss function formula of the Worker network is:

[0139]

[0140]

[0141]

[0142]

[0143] Among them, La w represents the behavioral loss function of the Worker network, represents the advantage function of the Worker network, Represents the reward of the Worker network, Represents the state value of the observation history of the Worker network on day t+1, Represents the state value of the observation history of the Worker network on day t, It represents the ratio of the action probability distribution of the new and old policy networks of the Worker network on day t. represents the observation history of the tth day in the Worker network, Represents the loss function of the time state sequence representation model, Lc w Represents the evaluation loss function of the Worker network.

[0144] During the training process, based on the above formula, by minimizing and Update critic; by maximizing and Update the actor to obtain the updated strategy, then interact the updated strategy with the environment, collect new data after the interaction, and use it for the next round of strategy updates.

[0145] Step 206: Obtain a trained individual mobile intervention infectious disease prevention and control model based on the graph neural network based on individual contact risk, the time state sequence representation model, and the pre-trained intelligent agent.

[0146] In the present invention, in the model training part, it is necessary to first collect the travel trajectories of users in the city to be deployed for a period of time through mobile communication devices such as mobile phones, and divide the city to be deployed into blocks. Through the individual travel trajectory, the regional visit history, the closeness relationship between individuals, the commuting relationship between individuals and regions, and the commuting relationship between regions can be modeled; and through the parameters of real infectious diseases, the corresponding SEIR (Susceptible Exposed Infectious Recovered) model is constructed, and the infectious disease simulator is constructed in combination with the individual mobility modeling, so as to train the individual mobile intervention infectious disease prevention and control model on the infectious disease simulator, and deploy the trained individual mobile intervention infectious disease prevention and control model on the central server. In the model deployment and use part, it is necessary to collect the historical movement trajectory, health status and intervention status of individual users to the central server, and input them into the RL model deployed on the central server to obtain the corresponding prevention and control measures for each user, and send them to each user through the mobile terminal to achieve the prevention and control effect.

[0147] Preferably, in one embodiment, the individual mobile intervention infectious disease prevention and control model can be applied to individual prevention and control of different infectious diseases at different stages within a city. Specifically, during the model training phase, based on the individual's regional visit history and inter-individual relationships obtained based on the mobility data of urban users, different scenarios can be adapted by setting the number of infected people in the simulator for different infectious diseases. For example, in the early stages of an infectious disease, there are many cases of infection caused by external contact with the population, so the simulator can be set to have random external contact infections for a certain period of time to adapt to this scenario; while in the middle stages of an infectious disease, the model is based on the infectious disease prevention and control based on a certain number of infected people in the city, so a certain number of people can be set in the initial infected population setting of the simulator to meet the needs of this scenario. For different infectious diseases, the parameters of the simulator can be modified based on the settings of the SEIR model of the infectious disease in existing studies and reports to adapt to different infectious disease settings.

[0148] Figure 8 This is a structural diagram of the individual mobile intervention infectious disease prevention and control system provided by the present invention, such as Figure 8As shown, the present invention provides an individual mobility intervention epidemic prevention and control system, which is an individual mobility intervention epidemic prevention and control system via deep reinforcement learning, including an individual daily historical information acquisition module 801 and a prevention and control intervention strategy generation module 802, wherein the individual daily historical information acquisition module 801 is used to obtain the daily historical status information and individual relationship information of individual users in the target city within a preset time interval, the historical status information including the user's movement trajectory, health status, intervention status, infection probability and number of acquaintance contacts, and the individual relationship information including the relationship between individuals, the relationship between individuals and regions, and the relationship between regions; the prevention and control intervention strategy generation module 802 is used to input the historical status information and the individual relationship information into a trained individual mobility intervention epidemic prevention and control model to obtain the prevention and control intervention measures for each individual user in the target city;

[0149] Among them, the trained individual mobile intervention infectious disease prevention and control model is obtained by training the graph neural network, long-term and short-term neural network and intelligent agent based on the individual status information and sample individual relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process; the individual status information of sample users includes the health status information of latently infected people converted to overtly infected people.

[0150] The individual mobile intervention infectious disease prevention and control system provided by the present invention analyzes the historical trajectories of individual users in the target area and the characteristics of the individual users themselves, and models the impact of latently infected people on infectious diseases, thereby conducting precise travel interventions of varying degrees for some high-risk groups. The ultimate goal is to minimize the number of infections with lower travel interventions.

[0151] On the basis of the above embodiment, the system also includes a first sample set construction module, a first training module, a second sample set construction module, a second training module, a third training module and a model deployment module, wherein the first sample set construction module is used to obtain the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions based on the sample individual relationship information, and construct a first training sample set; the first training module is used to input the first training sample set into the graph neural network for training to obtain a graph neural network based on individual contact risk, as well as the regular commuting characteristics of each sample user individual and the social relations between sample users on each day; the second sample set construction module is used to construct a second training sample set based on the regular commuting characteristics of the sample users, the social relations between the sample users and the corresponding historical status sample information; the historical status sample information includes the historical status of uninfected persons. Sample information, historical status sample information of latently infected persons, historical status sample information of explicitly infected persons and historical status sample information of recovered persons; the second training module is used to input the second training sample set into the long-term and short-term neural network for training to obtain a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents the individual state information change trend of the sample user within a preset time period obtained by prediction; the third training module is used to train the intelligent agent based on hierarchical reinforcement learning and the sample user individual state representation based on the time state sequence to obtain a pre-trained intelligent agent; the intelligent agent is constructed based on the FuN model; the model deployment module is used to obtain a trained individual mobile intervention infectious disease prevention and control model based on the graph neural network based on individual contact risk, the time state sequence representation model and the pre-trained intelligent agent.

[0152] The system provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.

[0153] Figure 9 A schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 9As shown, the electronic device may include: a processor (processor) 901, a communication interface (Communications Interface) 902, a memory (memory) 903 and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904. The processor 901 can call the logic instructions in the memory 903 to execute the individual mobile intervention infectious disease prevention and control method, which includes: obtaining the historical status information and individual relationship information of individual users in the target city every day within a preset time interval, the historical status information including the individual user's movement trajectory, health status, intervention status, infection probability and number of acquaintance contacts, and the individual relationship information including the relationship between individuals, the relationship between individuals and regions, and the relationship between regions; inputting the historical status information and the individual relationship information into the trained individual mobile intervention infectious disease prevention and control model to obtain the prevention and control intervention measures for each individual user in the target city; wherein the trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-short-term neural network and an intelligent agent based on sample user individual status information and sample individual relationship information, and the intelligent agent is constructed based on a partially observable Markov decision process; the sample user individual status information includes health status information of latently infected people converted to overtly infected people.

[0154] In addition, the logic instructions in the above-mentioned memory 903 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0155] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the individual mobile intervention infectious disease prevention and control method provided by the above methods, the method including: obtaining historical status information and individual relationship information of individual users in a target city every day within a preset time interval, the historical status information including the user's movement trajectory, health status, intervention status, infection probability and number of acquaintance contacts, and the individual relationship information including inter-individual relationships, relationships between individuals and regions, and relationships between regions; inputting the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city; wherein the trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-short-term neural network and an intelligent agent based on sample user individual status information and sample individual relationship information, and the intelligent agent is constructed based on a partially observable Markov decision process; the sample user individual status information includes health status information of latently infected people converted to overtly infected people.

[0156] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the individual mobile intervention infectious disease prevention and control method provided in the above-mentioned embodiments, the method comprising: obtaining historical status information and individual relationship information of individual users in a target city every day within a preset time interval, the historical status information including the individual user's movement trajectory, health status, intervention status, infection probability and number of contacts with acquaintances, and the individual relationship information including relationships between individuals, relationships between individuals and regions, and relationships between regions; inputting the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city; wherein the trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-short-term neural network and an intelligent agent based on sample user individual status information and sample individual relationship information, and the intelligent agent is constructed based on a partially observable Markov decision process; the sample user individual status information includes health status information of a latently infected person converted to an overtly infected person.

[0157] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for individual mobile intervention in infectious disease prevention and control, characterized in that: include: Obtain daily historical status information and individual relationship information for individual users in the target city within a preset time interval. The historical status information includes the user's movement trajectory, health status, intervention status, infection probability, and number of acquaintance contacts. The individual relationship information includes relationships between individuals, relationships between individuals and regions, and relationships between regions. Inputting the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city; The trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-term and short-term neural network, and an intelligent agent based on the individual status information and relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process. The individual status information of sample users includes the health status information of latently infected people who have transformed into overtly infected people. The trained individual mobile intervention infectious disease prevention and control model is obtained by the following steps: According to the sample individual relationship information, the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions are obtained to construct a first training sample set; Inputting the first training sample set into a graph neural network for training, obtaining a graph neural network based on individual contact risk, as well as the individual regular commuting characteristics of each sample user on each day and the social relationships between the sample users; Constructing a second training sample set based on the individual regular commuting characteristics of the sample users, the social relationships between the sample users, and the corresponding historical status sample information; the historical status sample information includes historical status sample information of uninfected persons, historical status sample information of latently infected persons, historical status sample information of overtly infected persons, and historical status sample information of recovered persons; Inputting the second training sample set into a long-short-term neural network for training, thereby obtaining a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents a change trend of the individual state information of the sample user within a preset time period obtained through prediction; Based on hierarchical reinforcement learning and the sample user individual state representation based on the time state sequence, the intelligent agent is trained to obtain a pre-trained intelligent agent; the hierarchical reinforcement learning model is constructed based on the FuN model; Obtaining a trained individual mobile intervention infectious disease prevention and control model based on the graph neural network based on individual contact risk, the time state sequence representation model, and the pre-trained intelligent agent; The agent is constructed with the goal of minimizing the number of infections and the occurrence of minimal interventions, and is used to observe the state representations of the sample user individuals based on the time state sequence, so as to determine the prevention and control intervention measures for each user individual based on the observation results. The observations include the sample user individual's health status observation, intervention status observation, infection probability observation, and number of acquaintance contacts observation. The infection probability observation is obtained by estimating the movement trajectory of the sample user individual; The agent's reward is expressed as: Among them, r t represents the reward on the tth day, L represents the total number of days of infectious disease prevention and control, ΔI represents the change in the number of infected people during the infectious disease prevention and control period, θ I represents the tolerance threshold of the medical system, ΔQ represents the change in travel intervention during the infectious disease prevention and control period, and θ Q Indicates the tolerance threshold of the economic system.

2. The method for individual mobile intervention in infectious disease prevention and control according to claim 1, characterized in that: The graph neural network based on individual exposure risk is constructed through a graph convolutional network and a GraphSAGE algorithm.

3. The method for individual mobile intervention in infectious disease prevention and control according to claim 1, characterized in that: The time state sequence representation model is obtained by improving the long short-term memory network through a variational autoencoder.

4. The method for individual mobile intervention in infectious disease prevention and control according to claim 3, characterized in that: The hierarchical reinforcement learning includes a Manager network and a Worker network, wherein the loss function formula of the Manager network is: Among them, La m Represents the behavioral loss function of the Manager network, represents the advantage function of the Manager network, represents the reward of the Manager network, γ represents the hyperparameter, Represents the state value of the Manager network's observation history on day t+1, Represents the state value of the observation history of the Manager network on day t, d w represents the similarity between the observation and the target, represents the observation of the Manager network on day t+c, represents observations based on incomplete information, represents the action probability distribution ratio of the new and old policy networks of the Manager network on day t, p represents the probability value of the policy output action, g t represents an abstract internal goal, represents the observation history of the Manager network on day t, p old represents the probability distribution of the old policy network, ε represents the hyperparameter, Represents the loss function of the time state sequence representation model, Lc m Represents the judgment loss function of the Manager network; The loss function formula of the Worker network is: Among them, La w represents the behavioral loss function of the Worker network, represents the advantage function of the Worker network, Represents the reward of the Worker network, Represents the state value of the observation history of the Worker network on day t+1, Represents the state value of the observation history of the Worker network on day t, It represents the ratio of the action probability distribution of the new and old policy networks of the Worker network on day t. represents the observation history of the tth day in the Worker network, Represents the loss function of the time state sequence representation model, Lc w Represents the evaluation loss function of the Worker network.

5. An individual mobile intervention infectious disease prevention and control system, characterized in that: include: An individual daily historical information acquisition module is used to obtain daily historical status information and individual relationship information of individual users in the target city within a preset time interval. The historical status information includes the user's movement trajectory, health status, intervention status, infection probability, and number of acquaintance contacts. The individual relationship information includes relationships between individuals, relationships between individuals and regions, and relationships between regions. a prevention and control intervention strategy generation module, configured to input the historical status information and the individual relationship information into a trained individual mobile intervention infectious disease prevention and control model to obtain prevention and control intervention measures for each individual user in the target city; The trained individual mobile intervention infectious disease prevention and control model is obtained by training a graph neural network, a long-term and short-term neural network, and an intelligent agent based on the individual status information and relationship information of sample users. The intelligent agent is constructed based on a partially observable Markov decision process. The individual status information of sample users includes the health status information of latently infected people who have transformed into overtly infected people. The system further comprises: A first sample set construction module is used to obtain the relationship between sample users, the relationship between sample users and regions, and the relationship between sample regions based on the sample individual relationship information, and construct a first training sample set; A first training module is configured to input the first training sample set into a graph neural network for training, thereby obtaining a graph neural network based on individual contact risk, as well as the individual regular commuting characteristics of each sample user on each day and the social relationships between the sample users; A second sample set construction module is configured to construct a second training sample set based on the individual regular commuting characteristics of the sample users, the social relationships between the sample users, and the corresponding historical status sample information; the historical status sample information includes historical status sample information of uninfected persons, historical status sample information of latently infected persons, historical status sample information of overtly infected persons, and historical status sample information of recovered persons; a second training module configured to input the second training sample set into a long-short-term neural network for training to obtain a time state sequence representation model and a sample user individual state representation based on the time state sequence; the sample user individual state representation based on the time state sequence represents a predicted trend of individual state information changes of the sample user within a preset time period; A third training module is configured to train an agent based on hierarchical reinforcement learning and the sample user individual state representation based on the time state sequence to obtain a pre-trained agent; the agent is constructed based on the FuN model; A model deployment module, configured to obtain a trained individual mobile intervention infectious disease prevention and control model based on the graph neural network based on individual contact risk, the time state sequence representation model, and the pre-trained intelligent agent; The agent is constructed with the goal of minimizing the number of infections and the occurrence of minimal interventions, and is used to observe the state representations of the sample user individuals based on the time state sequence, so as to determine the prevention and control intervention measures for each user individual based on the observation results. The observations include the sample user individual's health status observation, intervention status observation, infection probability observation, and number of acquaintance contacts observation. The infection probability observation is obtained by estimating the movement trajectory of the sample user individual; The agent's reward is expressed as: Among them, r t represents the reward on the tth day, L represents the total number of days of infectious disease prevention and control, ΔI represents the change in the number of infected people during the infectious disease prevention and control period, θ I represents the tolerance threshold of the medical system, ΔQ represents the change in travel intervention during the infectious disease prevention and control period, and θ Q Indicates the tolerance threshold of the economic system.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the individual mobile intervention infectious disease prevention and control method as described in any one of claims 1 to 4 are implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for individual mobile intervention in infectious disease prevention and control as claimed in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Major infectious disease propagation risk early warning and prevention and control analysis system for COVID-19

    CN111863271A

  • Conversation method and device based on hierarchical reinforcement learning network, and storage medium

    CN112860869A