Virus transmission dynamic prediction method based on multi-strain multi-region population model and related equipment
By constructing a multi-strain and multi-regional population model, integrating multi-source data and dividing the population into subgroups, the problems of data integration and real-time updating in existing technologies have been solved, high-precision prediction and early warning of virus transmission dynamics have been achieved, and the response capabilities of the public health system have been enhanced.
Patent Information
- Application Number
- CN202510189704.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology of dynamic prediction of virus transmission, it is difficult to effectively integrate multi-source data, update in real time and improve prediction accuracy, especially in complex situations with multiple variants and multiple regions.
A multi-strain and multi-regional population model is adopted. By constructing the spatiotemporal multi-regional population model framework SpatPOMP, epidemic data, viral gene sequence data and population mobility data are integrated to divide the population into susceptible, immune, exposed, infected and recovered sub-populations, and dynamic predictions are made using the transmission dynamics model.
It has achieved comprehensive pathogen monitoring and accurate early warning, and can update and respond to rapid pathogen mutations and cross-regional spread in real time, improving the real-time and accuracy of predictions, and enhancing the practicality and effectiveness of epidemic early warnings.
Smart Images

Figure CN120748769A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of dynamic prediction of epidemic spread, and in particular relates to a method and related equipment for dynamic prediction of virus spread based on a multi-strain and multi-region population model. Background Art
[0002] Globally, respiratory pathogen surveillance and early warning technologies have made considerable progress. In particular, the application of bioinformatics and the development of high-throughput sequencing technologies have greatly improved the ability to rapidly measure and analyze pathogen genomes. These technologies have become essential tools for accurately identifying and tracking pathogens. However, despite these technological advances, current systems still face several key challenges when processing large-scale epidemic data:
[0003] 1. Data integration problem: Existing technologies often lack effective methods to integrate data of different types and sources (such as genomic data, environmental monitoring data, and population mobility data), which limits the comprehensiveness of analysis and the accuracy of predictions.
[0004] 2. Real-time issues: Many existing models cannot be updated in real time or quickly adapt to new data inputs, resulting in delayed responses to rapidly changing epidemics.
[0005] 3. Prediction accuracy issues: Despite the use of technologies such as high-throughput sequencing, existing models often fail to fully utilize these data to effectively predict transmission dynamics, especially in complex situations with multiple variants and multiple regions. Summary of the Invention
[0006] The purpose of this application is to provide a method and related equipment for dynamic prediction of virus transmission based on a multi-strain and multi-regional population model, aiming to solve the problems existing in the existing technology. By developing a population model that can update and integrate multi-source data in real time, it not only improves the comprehensiveness of pathogen monitoring and the accuracy of early warning, but also effectively responds to the challenges of rapid mutation and cross-regional spread of pathogens.
[0007] In a first aspect, the present application provides a method for predicting the dynamics of virus transmission based on a multi-strain, multi-regional population model, the method comprising:
[0008] Obtain historical epidemic data related to the target area; the target area includes multiple regions; historical epidemic data includes epidemic data related to the target virus, viral gene sequence data, population mobility data, and epidemic management policy data in each region within a specified past time period; the target virus has multiple strains;
[0009] The population mobility data is organized into an asymmetric population mobility matrix between different regions every day;
[0010] Determine the epidemic management policies adopted by different regions at different time periods based on epidemic management policy data, and sort out the epidemic prevention levels corresponding to different epidemic management policies based on the epidemic management policies adopted by different regions at different time periods;
[0011] Using the spatiotemporal multi-regional population model framework SpatPOMP as the model construction framework, and using epidemic data, viral gene sequence data, population mobility matrix, and epidemic prevention levels corresponding to different epidemic management policies as basic data to construct a multi-strain multi-regional population model; the multi-strain multi-regional population model is used to determine the potential dynamics of each region at each moment within a specified time period;
[0012] Generate dynamic prediction information on virus transmission based on a multi-strain and multi-region population model.
[0013] In some embodiments, the multi-strain multi-region population model includes multiple independent underlying dynamic process models corresponding to different regions;
[0014] Each latent dynamic process model consists of a transmission dynamics group model, and different transmission dynamics group models are connected through the covariates of the population flow matrix that actually changes every day;
[0015] The transmission dynamics group model associated with each region divides the population in the region into susceptible, immune, exposed, infected, recovered, and deceased populations;
[0016] The exposed population in each region includes one or more exposed subpopulations subdivided according to the strains affecting the region; the infected population in each region includes one or more infected subpopulations subdivided according to the strains affecting the region.
[0017] In some embodiments, the underlying dynamics are described by the following formula:
[0018] X(t)=(S u (t),V u (t),E uk (t),I uk (t),R u (t),D u (t),u∈0:U,k∈1:K)
[0019] X(t) represents the potential dynamics of each region in time period t; U represents the number of regions, K represents the number of strains; S u (t) represents the dynamics of susceptible population in the uth region during time period t; V u (t) represents the dynamics of the immune population in the uth region during the t period; E uk (t) represents the dynamics of the exposed population associated with the kth strain in the uth region during the tth period; Iuk (t) represents the dynamics of the infected population associated with the kth strain in the uth region during the tth period; R u (t) represents the dynamics of the rehabilitation population in the uth region during the t-time period; D u (t) represents the death population dynamics in the u-th region during time period t.
[0020] In some embodiments, each exposed subpopulation includes an imported exposed population and a community exposed population, and the imported exposed population includes one or more source-subdivided populations subdivided according to the source of the case; each infected subpopulation includes an imported infected population and a community infected population, and the imported infected population includes one or more source-subdivided populations subdivided according to the source of the case; the population in the target area includes a traveler population composed of all individuals traveling between regions except transit passengers.
[0021] In some embodiments, the dynamics of the susceptible population in the u-th region is determined by the following formula:
[0022]
[0023] dS u represents the dynamics of susceptible population in the u-th region, represents the conversion increment of susceptible population into immune population in the u-th region within infinitesimal time, It represents the conversion increment of susceptible population in the u-th region into community exposed population associated with the k-th strain in infinitesimal time;
[0024] The dynamics of the immune population in region u is determined by the following formula:
[0025]
[0026] dV u represents the dynamics of the immune population in the u-th region, It represents the conversion increment of the immune population in the u-th region into the community exposed population associated with the k-th strain in infinitesimal time;
[0027] The dynamics of the community exposure population associated with the kth strain in the uth region is determined by the following formula:
[0028]
[0029] dE u,k,c represents the dynamics of the community exposed population associated with the kth strain in the uth region, represents the conversion increment of the community exposure population associated with the k-th strain in the u-th region into the traveler population within infinitesimal time, It represents the conversion increment of the community exposed population associated with the k-th strain in the u-th region into the community infected population associated with the k-th strain within infinitesimal time;
[0030] The dynamics of the community infection population associated with the kth strain in the uth region is determined by the following formula:
[0031] dI u,k,c represents the dynamics of community infection population associated with the kth strain in the uth region, represents the conversion increment of the community infection population related to the k-th strain in the u-th region into the passenger population within infinitesimal time, represents the conversion increment of the community infected population related to the kth strain in the uth region into the recovered population within infinitesimal time, It represents the conversion increment of the number of community infected people related to the k-th strain in the u-th region into the number of deaths within infinitesimal time;
[0032] The dynamics of the population exposed to the input associated with the kth strain in the uth region is determined by the following formula:
[0033]
[0034] dE u,k,i represents the dynamics of the input-exposed population associated with the kth strain in the uth region, represents the conversion increment of the passenger population into the input exposure population associated with the kth strain in the uth region within infinitesimal time, It represents the conversion increment of the input exposed population related to the k-th strain in the u-th region into the input infected population related to the k-th strain within infinitesimal time;
[0035] The dynamics of the imported infected population associated with the kth strain in the uth region is determined by the following formula:
[0036]
[0037] dI u,k,i represents the dynamics of the imported infected population associated with the kth strain in the uth region, represents the conversion increment of the passenger population into the imported infected population associated with the kth strain in the uth region within infinitesimal time, represents the conversion increment of the input infected population related to the k-th strain in the u-th region into the recovered population within infinitesimal time, It represents the conversion increment of the imported infected population related to the k-th strain in the u-th region into the dead population within infinitesimal time;
[0038] The dynamics of the rehabilitation population in the uth region is determined by the following formula:
[0039]
[0040] dR u represents the dynamics of the recovered population in the u-th region;
[0041] The death population dynamics in the u-th region is determined by the following formula:
[0042]
[0043] dD u Represents the dynamics of the death population in the u-th region.
[0044] In some embodiments, The conversion rate is defined as:
[0045] The conversion rate is defined as:
[0046] and The conversion rate is defined as:
[0047] and The conversion rate is defined as:
[0048] and The conversion rate is defined as:
[0049] in,
[0050] λ uk (t) represents the infectious force; ω uk (t) represents the infection rate; β 0,u represents the basic infection rate in each region; represents the total population effectively infected;
[0051] λ k represents the relative reduction in infection rate due to effective vaccination against the kth strain;
[0052] η k Relative values indicating the increase or decrease in transmissibility due to different strains;
[0053] ι u,n(t) represents the relative reduction in the virus's transmissibility due to the epidemic management policy reaching epidemic prevention level n;
[0054] ε k represents the inverse of the incubation period of the kth strain;
[0055] h u (t) represents the pre-estimated ratio of fatal infection cases to infection cases in the u-th region within time t;
[0056] γ represents the rate at which infected individuals recover or die;
[0057] dΓ / dt represents non-negative multiplicative gamma white noise with variance parameter σ.
[0058] In some embodiments, Calculated by the following formula:
[0059]
[0060] It represents the conversion increment of the passenger population into the imported infected population in the u-th region related to the k-th strain and originating from the o-th region within infinitesimal time;
[0061] M ou (t) represents the time-varying transmission matrix representing the number of individuals moving from the oth region to the uth region during the time period t;
[0062] I o,k,c represents the community infection population associated with the kth strain in the oth region;
[0063] N o represents the population in the oth region;
[0064] Calculated by the following formula:
[0065]
[0066] It represents the conversion increment of the passenger population into the imported exposed population in the u-th region related to the k-th strain and originating from the o-th region within infinitesimal time;
[0067] E o,k,c represents the community exposed population associated with the kth strain in the oth region;
[0068] Calculated by the following formula:
[0069]
[0070] M uj (t) is the time-varying transmission matrix representing the number of individuals moving from the uth region to the jth region in time period t;
[0071] I uk,c represents the community infection population associated with the kth strain in the uth region;
[0072] N u represents the population of the u-th region;
[0073] Calculated by the following formula:
[0074]
[0075] E uk,c represents the community exposed population associated with the kth strain in the uth region.
[0076] In some embodiments, the process of constructing a multi-strain, multi-regional population model includes:
[0077] Build multiple basic potential dynamic process models based on basic data;
[0078] Construct corresponding observation process models based on each underlying potential dynamic process model;
[0079] Each basic latent dynamic process model was calibrated based on the observation process model corresponding to each basic latent dynamic process model to obtain a multi-strain and multi-region population model.
[0080] In a second aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0081] In a third aspect, the present application provides a computer device comprising: one or more processors; a memory; and one or more computer programs, the processor and the memory being connected via a bus, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the processor executes the computer program, the method provided in the above embodiment is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 This is a flowchart of a method for dynamically predicting virus transmission based on a multi-strain, multi-regional population model provided by an embodiment of the present application;
[0083] Figure 2 This is a flow chart of the SVEIRD model provided in one embodiment of the present application;
[0084] Figure 3This is a specific structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0085] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0086] The current public health system faces a series of challenges in responding to respiratory disease outbreaks, especially the COVID-19 outbreak caused by rapidly mutating pathogens such as SARS-CoV-2. Through interdisciplinary collaboration and integration of research methods in fields such as bioinformatics, epidemiology, and environmental science, this application provides a method for predicting the dynamics of virus transmission based on a multi-strain and multi-regional population model. This method will provide the ability to detect, respond quickly, and effectively control respiratory disease pathogens early, significantly improving the comprehensiveness, accuracy, and geographical precision of pathogen monitoring. In particular, the multi-strain and multi-regional population model is constructed based on multi-source data such as population mobility data, epidemic prevention policy data, and viral genome evolution data, which can improve the ability to accurately predict the dynamics of pathogen transmission and enhance the practicality and effectiveness of epidemic warnings. Through the above-mentioned innovative approach, the method can enhance the response capacity of the global public health system, reduce the risk of transmission and outbreak of respiratory diseases, and at the same time promote the integration and innovation of interdisciplinary research methods, and contribute to the development of fields such as bioinformatics, public health, and environmental science.
[0087] See also Figure 1 , is a flowchart of a method for dynamic prediction of virus transmission based on a multi-strain and multi-regional population model provided in one embodiment of the present application. This embodiment mainly uses the application of this method to a computer device as an example to illustrate.
[0088] S110: Obtain historical epidemic-related data in the target area.
[0089] The target area includes multiple regions. Historical epidemic data includes epidemic data related to the target virus, viral gene sequence data, population mobility data, and epidemic management policy data for each region within a specified past time period; the target virus has multiple strains.
[0090] In some embodiments, the target region may be the Guangdong-Hong Kong-Macao Greater Bay Area, and the regions included in the target region may be multiple major cities in the Guangdong-Hong Kong-Macao Greater Bay Area (e.g., Hong Kong, Macao, Guangzhou, Shenzhen, Zhuhai, Foshan, Huizhou, Dongguan, Zhongshan, Jiangmen, Zhaoqing). The target virus and its strain may be multiple strains of the new coronavirus (e.g., Alpha, Delta, Omicron, Others).
[0091] Accordingly, regarding data collection, we can focus on collecting key epidemic data and corresponding gene sequence information to ensure the accuracy and timeliness of the monitoring and early warning system. The specific collection scope includes:
[0092] 1) Confirmed cases, deaths, and viral gene sequence data:
[0093] Hong Kong epidemic data and genetic sequences from 2021 to 2022: COVID-19 epidemic report data in Hong Kong during this period will be collected, including the number of confirmed cases, the number of deaths, and mutation information of viral genetic sequences, as basic data for analyzing the epidemic spread trend and virus evolution.
[0094] Epidemic data and genetic sequences of various cities in Guangdong Province from 2022 to 2023: Focus on the epidemic data and viral genetic sequences of various cities in the Greater Bay Area during the same period. Considering the active economy and frequent population mobility in the Greater Bay Area, these data are crucial for understanding the study of virus transmission patterns between different cities.
[0095] Macao epidemic data and genetic sequences of the outbreak on June 18, 2022: Collect relevant data and viral genetic sequences of the Macao epidemic since that date to assess the spread and mutation of the virus.
[0096] 2) Population mobility data:
[0097] Population mobility data between the 11 major cities in the Greater Bay Area can be collected between 2021 and 2023. This data will be organized into a daily asymmetric population flow matrix between different cities to support the dynamic population mobility process in the model. In addition, if transit passengers (defined as passengers passing through a city briefly) are to be considered, the proportion of transit passengers to total passengers in each city will also need to be collected.
[0098] 3) Policy data:
[0099] Based on the specific epidemic management policies implemented by different cities at different time periods, corresponding epidemic prevention levels can be compiled. Each level corresponds to an independent parameter in the model, which can be used to assess the impact of different epidemic prevention measures on the spread of the epidemic. Each level corresponds to an independent parameter in the model that affects the ability of imported cases to spread local cases.
[0100] S120: Organize population mobility data into an asymmetric population mobility matrix between different regions every day.
[0101] S130: Determine the epidemic management policies adopted by different regions in different time periods based on the epidemic management policy data, and sort out the epidemic prevention levels corresponding to the different epidemic management policies based on the epidemic management policies adopted by different regions in different time periods.
[0102] S140: Using the spatiotemporal multi-regional population model framework SpatPOMP as the model construction framework, and using epidemic data, viral gene sequence data, population mobility matrix and epidemic prevention levels corresponding to different epidemic management policies as basic data to construct a multi-strain and multi-regional population model; the multi-strain and multi-regional population model is used to determine the potential dynamics of each region at each moment within a specified time period.
[0103] This application aims to construct a multi-strain (variant) multi-regional population model, and simulate and calibrate the transmission information contained in the viral genome sequence in the model. However, the existing models cannot meet the requirements, so it is necessary to build a dedicated model from scratch. This is a highly personalized model. This application uses a flexible and advanced scalable framework, namely the spatiotemporal multi-regional population model framework SpatPOMP as the model construction framework, and uses epidemic data, viral gene sequence data, population flow matrix, and epidemic prevention levels corresponding to different epidemic management policies and other multi-source data to construct the model. SpatPOMP was proposed by a team from the School of Statistics at the University of Michigan in 2021. It provides a very flexible Monte Carlo statistical research framework for multi-space nonlinear non-Gaussian mathematical models.
[0104] In some embodiments, the process of constructing a multi-strain and multi-regional population model can be to construct multiple basic potential dynamic process models based on basic data, construct corresponding observation process models based on each basic potential dynamic process model, and calibrate each basic potential dynamic process model based on the observation process model corresponding to each basic potential dynamic process model to obtain a multi-strain and multi-regional population model.
[0105] In some embodiments, the multi-strain and multi-region population model includes multiple independent potential dynamic process models corresponding to different regions; each potential dynamic process model is composed of a transmission dynamics population model, and different transmission dynamics population models are connected through the covariates of the population flow matrix that actually changes every day; the transmission dynamics population model related to each region divides the population in the region into susceptible population, immune population, exposed population, infected population, recovered population and dead population; the exposed population in each region includes one or more exposed sub-populations subdivided according to the strains affected by the region; the infected population in each region includes one or more infected sub-populations subdivided according to the strains affected by the region.
[0106] Specifically, SpatPOMP, a framework for constructing multi-strain, multi-regional population models, is an advanced spatiotemporal multi-regional population modeling framework developed by a team from the Department of Statistics at the University of Michigan, building on its predecessor, POMP. The POMP family of models, known as partially observed Markov process models (also known as hidden Markov models or state-space models), is a common tool for time series analysis. This suite of methods provides a highly flexible Monte Carlo statistical framework for studying nonlinear, non-Gaussian POMP models. Many modern statistical methods for POMP models have been implemented within this framework, including sequential Monte Carlo, iterative filtering, particle Markov chain Monte Carlo, approximate Bayesian computation, maximum composite likelihood estimation, nonlinear prediction, and trajectory matching. SpatPOMP is a partially observed Markov process (POMP) model with spatial structure. A SpatPOMP model consists of two main processes: a latent process model (i.e., a continuous-time Markov chain) and an observation model (observation model, which obtains noisy and / or incomplete observations from the latent process at different time points).
[0107] The underlying dynamic process model consists of a SVEIRD population model. The SVEIRD population model is a transmission dynamics population model, specifically a mathematical model used to describe the spread of infectious diseases, particularly when considering vaccination and different stages of infection. The model divides the population into six compartments: susceptible (S), vaccinated (V), exposed (E), infected (infected and contagious) (I), recovered (R), and deceased (D).
[0108] The underlying dynamics are described by the following formula:
[0109] X(t)=(S u (t),V u (t),E uk (t),I uk (t),R u (t),D u (t),u∈0:U,k∈1:K)
[0110] X(t) represents the potential dynamics of each region in time period t; U represents the number of regions in the target area, K represents the number of strains of the target virus; S u (t) represents the dynamics of susceptible population in the uth region during time period t; V u (t) represents the dynamics of the immune population in the uth region during the t period; E uk(t) represents the dynamics of the exposed population associated with the kth strain in the uth region during the tth period; I uk (t) represents the dynamics of the infected population associated with the kth strain in the uth region during the tth period; R u (t) represents the dynamics of the rehabilitation population in the uth region during the t-time period; D u (t) represents the death population dynamics in the u-th region during time period t.
[0111] For the imported exposed (infected) population in any region, individuals in it can come from other regions, while individuals in the community exposed (or infected) population in the region do not come from other regions, and they can only flow out to other regions. Each exposed subpopulation can be further subdivided into the imported exposed population and the community exposed population, and each infected subpopulation can be further subdivided into the imported infected population and the community infected population.
[0112] For the imported exposed (or infected) population in any region, it can be further subdivided into populations corresponding to different source areas according to the different source areas. Therefore, each imported exposed population can include one or more source-based subdivided populations subdivided according to the source area of the case, and each imported infected population can include one or more source-based subdivided populations subdivided according to the source area of the case.
[0113] The population in the target area includes all individuals traveling between regions except transit passengers, and the passenger population is added to the above population for the convenience of notation.
[0114] For any two groups of people in any region, N WM (t) is used to represent the direction change from W group to M group, and dN WM To express the conversion increment in infinitesimal time, dN WM =N WM (t+dt)-N WM (t).
[0115] The following describes how to determine the dynamics of each population in each region.
[0116] The dynamics of the susceptible population in the u-th region is determined by the following formula:
[0117]
[0118] dS u represents the dynamics of susceptible population in the u-th region, represents the conversion increment of susceptible population into immune population in the u-th region within infinitesimal time, It represents the conversion increment of susceptible population in the u-th region into community exposed population associated with the k-th strain in infinitesimal time.
[0119] The dynamics of the immune population in region u is determined by the following formula:
[0120]
[0121] dV u represents the dynamics of the immune population in the u-th region, It represents the conversion increment of the immune population in the u-th region into the community exposed population associated with the k-th strain in infinitesimal time.
[0122] The dynamics of the community exposure population associated with the kth strain in the uth region is determined by the following formula:
[0123]
[0124] dE u,k,c represents the dynamics of the community exposed population associated with the kth strain in the uth region, represents the conversion increment of the community exposure population related to the kth strain in the uth region into the traveler population within infinitesimal time, It represents the conversion increment of the community exposed population associated with the k-th strain in the u-th region into the community infected population associated with the k-th strain within infinitesimal time.
[0125] The dynamics of the community infection population associated with the kth strain in the uth region is determined by the following formula:
[0126] dI u,k,c represents the dynamics of community infection population associated with the kth strain in the uth region, represents the conversion increment of the community infection population related to the k-th strain in the u-th region into the passenger population within infinitesimal time, represents the conversion increment of the community infected population related to the kth strain in the uth region into the recovered population within infinitesimal time, It represents the conversion increment of the number of community infected people related to the k-th strain in the u-th region into the number of deaths within infinitesimal time.
[0127] The dynamics of the population exposed to the input associated with the kth strain in the uth region is determined by the following formula:
[0128]
[0129] dE u,k,irepresents the dynamics of the input-exposed population associated with the kth strain in the uth region, represents the conversion increment of the passenger population into the input exposure population associated with the kth strain in the uth region within infinitesimal time, It represents the conversion increment of the imported exposed population associated with the k-th strain in the u-th region into the imported infected population associated with the k-th strain within infinitesimal time.
[0130] The dynamics of the imported infected population associated with the kth strain in the uth region is determined by the following formula:
[0131] dI u,k,i represents the dynamics of the imported infected population associated with the kth strain in the uth region, represents the conversion increment of the passenger population into the imported infected population associated with the kth strain in the uth region within infinitesimal time, represents the conversion increment of the input infected population related to the k-th strain in the u-th region into the recovered population within infinitesimal time, It represents the conversion increment of the imported infected population related to the k-th strain in the u-th region into the death population within infinitesimal time.
[0132] The dynamics of the rehabilitation population in the uth region is determined by the following formula:
[0133]
[0134] dR u Represents the dynamics of the recovered population in the u-th region.
[0135] The death population dynamics in the u-th region is determined by the following formula:
[0136]
[0137] dD u Represents the dynamics of the death population in the u-th region.
[0138] In the model, each individual only travels between regions when they become exposed or infected. In addition, it is important to note that in practice it is not necessary to match each individual entering and leaving a T compartment (i.e., a traveling population), so there may be small random variations in the total population of each region.
[0139] Each transition between groups has a corresponding transition rate, which may depend on the state of other groups, or on covariate processes, or on specific parameters, or on time itself. For groups not related to travel, it is convenient to specify the rate as a per capita ratio.
[0140] In some embodiments, The conversion rate is defined as:
[0141] The conversion rate is defined as:
[0142] and The conversion rate is defined as:
[0143] and The conversion rate is defined as:
[0144] and The conversion rate is defined as:
[0145] in,
[0146] λ uk (t) represents the infectious force; ω uk (t) represents the infection rate; β 0,u represents the basic infection rate in each region; represents the total population effectively infected;
[0147] λ k represents the relative reduction in infection rate due to effective vaccination against the kth strain;
[0148] η k Relative values indicating the increase or decrease in transmissibility due to different strains;
[0149] l u,n (t) represents the relative reduction in the virus's transmissibility due to the epidemic management policy reaching epidemic prevention level n;
[0150] ε k represents the inverse of the incubation period of the kth strain;
[0151] h u (t) represents the pre-estimated ratio of fatal infection cases to infection cases in the u-th region within time t;
[0152] γ represents the rate at which infected individuals recover or die;
[0153] dΓ / dt represents a non-negative multiplicative gamma white noise with variance parameter σ, i.e., Γ(t), Γ(t) is a gamma process with independent increments such that Γ(t) − Γ(s) has mean ts and has a gamma distribution with variance σ(ts).
[0154] For the transitions in and out of the passenger crowd, the number can be based on the time-varying transfer matrix M ju (t) is specified, and the matrix represents the number of individuals moving from region j to region u at time t.
[0155] In some embodiments, Calculated by the following formula:
[0156]
[0157] It represents the conversion increment of the passenger population into the imported infected population in the u-th region related to the k-th strain and originating from the o-th region within infinitesimal time;
[0158] M ou (t) represents the time-varying transmission matrix representing the number of individuals moving from the oth region to the uth region during the time period t;
[0159] I o,k,c represents the community infection population associated with the kth strain in the oth region;
[0160] N o represents the population in the oth region;
[0161] Calculated by the following formula:
[0162]
[0163] It represents the conversion increment of the passenger population into the imported exposed population in the u-th region related to the k-th strain and originating from the o-th region within infinitesimal time;
[0164] E o,k,c represents the community exposed population associated with the kth strain in the oth region;
[0165] Calculated by the following formula:
[0166]
[0167] M uj (t) is the time-varying transmission matrix representing the number of individuals moving from the uth region to the jth region in time period t;
[0168] I uk,c represents the community infection population associated with the kth strain in the uth region;
[0169] N u represents the population of the u-th region;
[0170] Calculated by the following formula:
[0171]
[0172] E uk,c represents the community exposed population associated with the kth strain in the uth region.
[0173] Furthermore, the observation process model will be constructed based on the above potential dynamic process model. n In the above, the observation data for each region include 2 official reports (i.e., daily reported cases and deaths) and the total number of sequenced cases for each strain, as well as the total number of sequenced cases of input events from different sources of infection with different strains (which can be inferred from the evolutionary tree model). Daily reported cases and deaths can be modeled as random variables and The number of strains sequenced is reported. The simulation is part of the reported case. Specifically,
[0174]
[0175] in, and They are used to measure the specific number of infected individuals who have entered the recovered population and the death population from the infected population. Here, NB(ΔN,θ) represents a negative binomial distribution with mean ΔN and variance θ. u (t) and ρ u (t) are vector-valued covariate processes specifying the infection detection ratio (IDR) and detection sequencing ratio (DSR), respectively.
[0176] In the process of constructing the multi-strain, multi-regional population model, there are at least the following innovations:
[0177] (1) Multi-region model construction: The model treats different regions as discrete spatial units, and different spatial units are connected through the covariate of the travel movement matrix, which actually changes every day. Each spatial unit has its own independent transmission dynamics group model, namely the SVEIRD model, which is composed of six groups: susceptible, immune, exposed, infected, recovered, and deceased. In addition, to simplify the model calculation, we assume that cross-regional transmission events can occur in exposed and infected populations in different spatial units, while ignoring the impact of population movement in other populations.
[0178] (2) Multi-strain model construction: The model further subdivides the exposed and infected populations in each spatial unit into sub-compartments of their respective strains according to the strains they are affected by. Specifically, assuming that a certain area is affected by four strains, the exposed population can be subdivided into four exposed sub-populations, and the infected population can also be subdivided into four infected sub-populations.
[0179] (3) Recording case import events from multiple sources: There are two possible sources for confirmed cases in a region, namely community cases and imported cases. For imported cases, in theory, they may come from any other spatial unit in the system. It is important to record the origins (places of origin) of different imported cases in the multi-strain and multi-region population model, so that the origin information of each imported case can be inferred by using the viral genome sequence, and this information will be used to calibrate the model. Accordingly, in order to record the origin information of each individual, the exposed subpopulation and infected subpopulation in each spatial unit can be further subdivided. For example, assuming that there are 11 regions in the target area, the exposed subpopulation and infected subpopulation can be subdivided into multiple origin subpopulations according to the 11 possible origins (Origin), for example, I u,k,o It can represent the imported infected population in the area numbered u who are infected with the kth strain and come from the area numbered o.
[0180] (4) Multi-level observation model: The multi-strain and multi-region population model distinguishes between process models and observation models. The process model is the above-mentioned SVEIRD differential equation model, which is used to simulate the spread, infection and recovery events of the virus within the population. Considering that in the real world, the confirmed cases observed are only a part of the infected cases, and a considerable part of the infected population may not be monitored, the embodiment of the present application introduces a realistic observation model to simulate and calibrate this process. The observation model includes two main steps: 1. Infection to diagnosis; 2. Diagnosis to sequencing. Each of these two steps is determined by two pre-set covariates, namely the infection detection ratio (IDR) and the detection sequencing ratio (DSR).
[0181] (5) Virus importation events derived from phylogenetic models: Combining points 3 and 4 above, the observation model can be used to observe the number of imported cases from different sources monitored in each region every day. However, how to calibrate this observed number is a difficult problem. To address this problem, the embodiment of the present application derives the source information of each virus importation event based on the viral genome sequence and phylogenetic model. The importation events derived through the above operation are based on multi-temporal and spatiotemporal genetic information and do not rely on epidemiological surveys. Therefore, they have the advantages of saving manpower, high sensitivity and accuracy.
[0182] (6) Simultaneously calibrate virus input events and daily case counts: Common transmission dynamics models in the field of epidemiology only model and calibrate case counts. The model provided in this application will simultaneously calibrate multiple observation indicators, such as the total number of confirmed cases in each region on each day, the total number of deaths, the total number of sequenced cases of each infected strain, and the total number of sequenced cases of input events from different sources of each infected strain. Specifically, the maximum likelihood estimation method can be used to add the likelihood values of each observation indicator, and then use an optimization algorithm suitable for high-dimensional models, such as an iterated block particle filter (IBPF) to optimize the search space, so as to find the model parameters that best match the actual observation values.
[0183] (7) Introducing multiple innovative key parameters: In addition to the aforementioned innovations, the model can also introduce vaccination-related data to simulate the immune-mediated impact on infectiousness. At the same time, each region has its own independent basic virus transmission parameters to reflect regional differences; different strains have independent infectiousness parameters; under different levels of epidemic prevention policies (such as when quarantine policies are implemented), the transmission capacity of imported cases will vary to varying degrees; most of the other parameters not mentioned above can be set as shared parameters between different spatial units, thereby reducing the amount of calculation and avoiding overfitting.
[0184] (8) Random process model: Considering that the real world is full of various randomness, this application uses a preset statistical model to randomly select values in each propagation step of the model, and deliberately adds random variables in some processes (such as gamma white noise in the basic propagation parameters), which can improve the model's adaptability and fitting ability to real-world data.
[0185] S150: Generate dynamic prediction information on virus transmission based on a multi-strain and multi-region population model.
[0186] After obtaining a calibrated multi-strain and multi-region population model, the model can be used to predict the transmission dynamics of the target virus in the target area, which helps predict the outbreak of new virus variants and provide early warnings to different cities in the region.
[0187] This application can effectively improve the efficiency of pathogen monitoring and early warning, building a more powerful support system for global public health security, especially in rapidly responding to the spread of widespread and rapidly mutating pathogens. It can at least bring the following beneficial effects:
[0188] 1. Improve the real-time and accuracy of predictions: This application adopts a real-time data update mechanism to quickly respond to new epidemic information and changes, ensuring that the model output is closest to the current condition, thereby enhancing the timeliness and accuracy of early warnings; by optimizing the data analysis process through machine learning algorithms, this application can more accurately predict the transmission path and speed of pathogens, helping public health decision makers to develop more effective prevention and control strategies.
[0189] 2. Enhanced data integration and multi-source information fusion capabilities: This application effectively integrates multiple data sources, from genomic data to environmental monitoring and population mobility, to provide a comprehensive epidemic monitoring perspective that helps explain and predict the complex dynamics of pathogen transmission. The deep integration of multi-source data enhances the comprehensiveness of the analysis, enabling the predictive model to support decision-making on epidemic management and response measures at multiple levels.
[0190] 3. Optimizing resource allocation and emergency response: The model provided in this application can guide the optimal allocation of public health resources, such as the distribution of testing equipment, medical resources, and vaccines, to ensure that these resources can be effectively utilized in areas where the epidemic is most urgent. By predicting epidemic development trends and high-risk areas through the model, emergency response measures can be deployed more quickly to reduce the social and economic impact of the epidemic.
[0191] 4. Promote the advancement of public health and scientific research: The interdisciplinary research approach of this application has promoted the development of fields such as bioinformatics, public health, and environmental science, and provided new scientific tools and concepts for the research and monitoring of other infectious diseases. The development and application of the model will enhance the understanding of the mechanisms of epidemic transmission and promote the scientific and precise formulation of public health policies.
[0192] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the virus transmission dynamic prediction method based on a multi-strain and multi-regional population model as provided in an embodiment of the present application.
[0193] Figure 3 A specific structural block diagram of a computer device provided in an embodiment of the present application is shown. The computer device 100 includes: one or more processors 101, a memory 102, and one or more computer programs, wherein the processor 101 and the memory 102 are connected via a bus, the one or more computer programs are stored in the memory 102, and are configured to be executed by the one or more processors 101. When the processor 101 executes the computer program, it implements the virus transmission dynamic prediction method based on the multi-strain and multi-regional population model provided in the present application.
[0194] It should be understood that each step in each embodiment of the present application is not necessarily performed in sequence according to the order indicated by the step numbers. Unless clearly stated herein, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, in each embodiment, at least a portion of steps may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or the sub-steps or stages of other steps.
[0195] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink), DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0196] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0197] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A method for predicting the dynamic spread of viruses based on a multi-strain, multi-regional population model, characterized in that: The method comprises: Obtaining historical epidemic data related to a target region; the target region includes multiple regions; the historical epidemic data includes epidemic data, viral gene sequence data, population mobility data, and epidemic management policy data related to the target virus in each of the regions within a specified past time period; the target virus has multiple strains; Arrange the population flow data into an asymmetric population flow matrix between different regions every day; Determine the epidemic management policies adopted by different regions at different time periods based on the epidemic management policy data, and sort out the epidemic prevention levels corresponding to the different epidemic management policies based on the epidemic management policies adopted by different regions at different time periods; A multi-strain, multi-regional population model is constructed using the spatiotemporal multi-regional population model framework SpatPOMP as a model construction framework, and a multi-strain, multi-regional population model is constructed based on the epidemic data, viral gene sequence data, the population mobility matrix, and the epidemic prevention levels corresponding to different epidemic management policies; the multi-strain, multi-regional population model is used to determine the potential dynamics of each region at each moment within a specified time period; Virus transmission dynamic prediction information is generated based on the multi-strain and multi-region population model.
2. The method according to claim 1, wherein The multi-strain multi-region population model includes multiple independent potential dynamic process models corresponding to different regions; Each of the potential dynamic process models is composed of a transmission dynamics group model, and different transmission dynamics group models are connected through the covariates of the population flow matrix that actually changes every day; The transmission dynamics population model associated with each of the regions divides the population in the region into susceptible, immune, exposed, infected, recovered, and deceased; The exposed population in each of the regions includes one or more exposed subpopulations subdivided according to the strains that affect the region; the infected population in each of the regions includes one or more infected subpopulations subdivided according to the strains that affect the region.
3. The method according to claim 2, wherein The underlying dynamics are described by the following formula: X(t)=(S u (t),V u (t),E uk (t),I uk (t),R u (t),D u (t),u∈0:U,k∈1:K) X(t) represents the potential dynamics of each region in time period t; U represents the number of regions, K represents the number of strains; S u (t) represents the dynamics of susceptible population in the u-th region in time period t; V u (t) represents the dynamics of the immune population in the u-th region in the t-time period; E uk (t) represents the dynamics of the exposed population in the u-th region during the t-th period of time associated with the k-th strain; I uk (t) represents the dynamics of the infected population associated with the kth strain in the uth region during the tth period; R u (t) represents the dynamics of the rehabilitation population in the u-th region during the t-time period; D u (t) represents the death population dynamics in the u-th region during time period t.
4. The method according to claim 3, wherein Each of the exposed subpopulations includes the imported exposed population and the community exposed population, and the imported exposed population includes one or more source-subdivided populations subdivided according to the source of the case; each of the infected subpopulations includes the imported infected population and the community infected population, and the imported infected population includes one or more source-subdivided populations subdivided according to the source of the case; the population in the target area includes the passenger population composed of all individuals traveling between regions except transit passengers.
5. The method according to claim 4, wherein The dynamics of the susceptible population in the u-th region is determined by the following formula: dS u represents the dynamics of susceptible population in the u-th region, It represents the conversion increment of susceptible population into immune population in the u-th region in infinitesimal time, It represents the conversion increment of susceptible population in the u-th region into community exposed population associated with the k-th strain in infinitesimal time; The dynamics of the immune population in the u-th region is determined by the following formula: dV u represents the dynamics of the immune population in the u-th region, It represents the conversion increment of the immune population in the u-th region into the community exposed population associated with the k-th strain within infinitesimal time; The dynamics of the community exposure population associated with the kth strain in the uth region is determined by the following formula: dE u,k,c represents the dynamics of the community exposed population associated with the kth strain in the uth region, It represents the conversion increment of the community exposure population associated with the k-th strain in the u-th region into the traveler population within infinitesimal time, It represents the conversion increment of the community exposed population associated with the k-th strain in the u-th region into the community infected population associated with the k-th strain within infinitesimal time; The dynamics of the community infection population associated with the kth strain in the uth region is determined by the following formula: dI u,k,c represents the dynamics of the community infection population associated with the kth strain in the uth region, It represents the conversion increment of the community infection population associated with the k-th strain in the u-th region into the passenger population within infinitesimal time, It represents the conversion increment of the community infection population related to the k-th strain in the u-th region into the recovered population within infinitesimal time, It represents the conversion increment of the number of community infected people related to the k-th strain in the u-th region into the number of deaths within infinitesimal time; The dynamics of the population exposed to the input associated with the kth strain in the uth region is determined by the following formula: dE u,k,i represents the dynamics of the population exposed to the kth strain in the uth region, represents the conversion increment of the passenger population into the input exposure population associated with the k-th strain in the u-th region within infinitesimal time, It represents the conversion increment of the imported exposed population associated with the k-th strain in the u-th region into the imported infected population associated with the k-th strain within infinitesimal time; The dynamics of the imported infected population associated with the kth strain in the uth region is determined by the following formula: dI u,k,i represents the dynamics of the imported infected population associated with the kth strain in the uth region, represents the conversion increment of the passenger population into the imported infected population associated with the k-th strain in the u-th region within infinitesimal time, represents the conversion increment of the imported infected population related to the k-th strain in the u-th region into the recovered population within infinitesimal time, It represents the conversion increment of the number of imported infected people related to the k-th strain in the u-th region into the number of deaths within infinitesimal time; The dynamics of the recovered population in the u-th region is determined by the following formula: dD u represents the dynamics of the recovered population in the u-th region; The dynamics of the death population in the u-th region is determined by the following formula: dD u Represents the dynamics of the death population in the u-th region.
6. The method according to claim 5, wherein The conversion rate is defined as: The conversion rate is defined as: and The conversion rate is defined as: and The conversion rate is defined as: and The conversion rate is defined as: in, λ uk (t) represents the infectious force; ω uk (t) represents the infection rate; β o,u represents the basic infection rate in each of the said areas; represents the total population effectively infected; λ k represents the relative reduction in infection rate due to effective vaccination against the kth strain; η k Relative values indicating the increase or decrease in transmissibility due to different strains; l u,n (t) represents the relative reduction in the virus's transmissibility due to the epidemic management policy reaching epidemic prevention level n; ε k represents the inverse of the incubation period of the kth strain; h u (t) represents the estimated ratio of fatal infections to infections in the uth region within time t; γ represents the rate at which infected individuals recover or die; dΓ / dt represents non-negative multiplicative gamma white noise with variance parameter σ.
7. The method according to claim 6, wherein described Calculated by the following formula: It represents the conversion increment of the passenger population into the imported infected population related to the kth strain and originating from the oth region in the uth region within infinitesimal time; M ou (t) represents the time-varying transmission matrix representing the number of individuals moving from the oth region to the uth region during the time period t; I o,k,c represents the community infection population associated with the kth strain in the oth region; N o represents the population of the oth region; described Calculated by the following formula: It represents the conversion increment of the passenger population into the imported exposed population in the u-th region related to the k-th strain and originating from the o-th region within an infinitesimal time; E o,k,c represents the community exposed population associated with the kth strain in the oth region; described Calculated by the following formula: M uj (t) is a time-varying transmission matrix representing the number of individuals moving from the uth region to the jth region in time period t; I uk,c represents the community infection population associated with the kth strain in the uth region; N u represents the population of the u-th region; described Calculated by the following formula: E uk,c Represents the community exposed population associated with the kth strain in the uth region.
8. The method according to claim 3, wherein The construction process of the multi-strain multi-region population model includes: constructing a plurality of basic potential dynamic process models based on the basic data; Constructing a corresponding observation process model based on each of the basic potential dynamic process models; Each of the basic latent dynamic process models is calibrated based on the observation process model corresponding to each of the basic latent dynamic process models to obtain the multi-strain multi-region population model.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer device comprising one or more processors, a memory, and one or more computer programs; the processors and the memory are connected via a bus, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Agricultural enzyme adaptive generation and optimization method based on multi-source data fusion
CN121744207A
Agricultural enzyme adaptive generation and optimization method based on multi-source data fusion
CN121744207B