A method for quantifying early epidemic infection risk based on regional mobility
By constructing a dynamic mobile flow perception map and a spatial interaction function model, and combining historical risk values, the risk of infection within and outside the region is quantified. This solves the problems of adaptability and insufficient data in the early epidemic risk prediction of existing technologies, and achieves rapid and accurate risk identification and prediction.
Patent Information
- Application Number
- CN202411250102.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-09-06
AI Technical Summary
Existing methods for predicting the risk of early epidemic infection are ill-suited to rapidly changing emerging diseases and are unable to accurately identify high-risk areas in the early stages when training data is lacking.
By employing a regional flow-based approach, a dynamic mobile flow perception map and spatial interaction function model are constructed. Combined with historical risk values, the risk of infection within and outside the region is quantified. Wasserstein distance is used to measure flow differences, calculate infection risk values, and perform risk superposition to achieve rapid and accurate risk quantification.
In the early stages of an epidemic, it can quickly and accurately identify high-risk areas and predict future disease transmission trends without requiring in-depth understanding of the infectious disease or the virus transmission mechanism, and is applicable to areas of different sizes.
Smart Images

Figure BDA0005032011620000031 
Figure BDA0005032011620000032 
Figure BDA0005032011620000034
Abstract
Description
Technical Field
[0001] This invention belongs to the field of epidemic infection risk assessment technology, and in particular relates to a method for quantifying early epidemic infection risk based on regional mobility. Background Technology
[0002] Infectious diseases, such as those caused by COVID-19, pose a serious threat to global public health and the economy. These viruses can spread rapidly in cities through daily human activities and social interactions. Accurate and effective assessment of disease risk levels in different regions is crucial in controlling and managing such early-stage epidemics.
[0003] Existing methods for predicting and assessing the risk of early epidemic infections include using data on confirmed cases and deaths, and employing proxy room models to model the transmission process of infectious diseases. These models can predict the viral concentration in a specific time and region to assess the risk level; however, they usually require a deep understanding of the disease and are therefore not suitable for newly emerging infections. Another approach inputs data on infection cases, deaths, and recoveries into deep learning models to predict epidemic trends, population infection rates, and medical resource needs. Considering that mechanistic data such as contact type, frequency, and transmission probability are often difficult to obtain, recent studies have begun to use human movement data. For example, some researchers have used human movement data to identify potential areas of Zika transmission in Singapore and conducted a feasibility analysis (Rajarethinam J, Ong J, Lim SH, et al. Using human movement data to identify potential areas of Zika transmission: case study of the largest Zika cluster in Singapore[J]. International journal of environmental research and public health, 2019, 16(5):808.). Other researchers have found that the inherent spatial structure of urban mobility has influenced the progression of epidemics (Aguilar J, Bassolas A, Ghoshal G, et al. Impact of urban structure on infectious diseases spreading[J]. Scientific reports,2022,12(1):3816.). In Namibia, researchers used spatially defined mobility risk networks to calculate the annual risk of HIV infection and identify high-risk areas for HIV transmission. However, such methods have two problems. First, these models are often relatively static because they focus primarily on modeling and predicting the behavior of a single viral type, so they may not adapt well when an epidemic shifts to a new or different stage of transmission. Second, these models rely on large amounts of training data, which is often difficult to obtain in the early stages of an epidemic.
[0004] Therefore, how to effectively and accurately identify high-risk areas in the early stages of an epidemic remains an important and unresolved challenge. Summary of the Invention
[0005] In light of this, this invention proposes a novel method for quantifying early epidemic infection risk based on regional mobility, specifically targeting early outbreaks of epidemics. The method focuses on two core factors: internal regional infection factors and cross-regional infection factors. The former stems from the observation that infectious diseases are primarily transmitted through interpersonal contact; therefore, the infection risk in a region is closely related to the rate of interpersonal interaction. The latter is inspired by research findings that inter-regional population mobility can reveal infection risk levels, with regions exhibiting higher mobility tending to be more susceptible to infectious diseases. To effectively adapt to rapid viral changes and intervene effectively in the early stages of an epidemic, the method allows for real-time adjustment of infection parameters in the model to accommodate viral variations. Furthermore, regardless of regional size, the method effectively identifies high-risk areas, demonstrating its broad applicability. Particularly in the early outbreak stage, when case data is limited and viral transmission mechanisms are unclear, the method can rapidly and accurately quantify regional infection risk and predict future disease transmission trends.
[0006] This invention provides a method for quantifying the early epidemic infection risk based on regional mobility, characterized by comprising:
[0007] Step 1: Construct a dynamic mobility flow perception map using human mobility data to obtain a weight matrix W that records the correlation between the flow patterns of each region constructed by the dynamic flow and other regions. DFAG The human mobility data mentioned therein is the origin-destination matrix (OD matrix) of human mobility.
[0008] Step 2: Calculate the epidemic infection risk value caused by intra-county movement based on the Spatial Interaction Function (SIF) model. The infection risk value generated by intra-county movement within time period t is calculated using the following formula:
[0009]
[0010] Among them, Z ii Z represents the number of individuals remaining in county i during time period t, while Z represents the number of individuals remaining in county i during time period t. ba and Z ab These represent the number of individuals who migrated from region a to region b and from region b to region a during time period t, respectively.
[0011] Step 3: Quantify the impact of inter-county flows based on the inflow and outflow volumes of counties and districts, and generate a county-level dynamic mobility flow perception map to dynamically quantify the risk of cross-county transmission. The cross-county flow risk value of county i in time period t is obtained by the following formula:
[0012]
[0013] Among them, WDFAG (i,j) is a weight matrix representing the correlation between the flow pattern of county i and county j, constructed from dynamic flows. ji Z represents the number of people who migrated from county j to county i during time period t, while Z... ij This represents the number of people who migrated from county i to county j during time period t;
[0014] Step 4: By integrating inter-county risk quantification with intra-county risk quantification, the infection risk value of the studied county / district is calculated. Each risk value is normalized by dividing it by the maximum value across all counties / districts and time periods. Then, the inter-county and intra-district risk values are weighted according to their respective weights to obtain the risk value of county / district i. As shown in the formula below:
[0015] Furthermore, in step 1, assuming there are N regions, for each region i, its daily inflow data Z to all other regions is... d Consider it as the vector of day d Right now Where d∈[1,D], the daily flow pattern information of each region is obtained by calculating the magnitude of the flow vector, that is, the daily inflow data pattern from region i to all other regions j in the time interval D:
[0016]
[0017] In this equation, P n Represents the probability distribution obtained from the vector sequence, which indicates the proportion of people flowing from region i to region j on day d;
[0018] Then, the dot product similarity between pedestrian flow vectors is used as the cost function.
[0019]
[0020] Incorporating the above formula into the calculation of the Wasserstein distance yields a metric called dynamic flow-sensing distance, used to measure flow differences between regions; the flow-sensing distance between region i and region j on day d is calculated as follows:
[0021]
[0022] The Wasserstein distance is calculated by applying the formula to the OD matrix of each small-scale census area, resulting in the weight matrix W. DFAG W DFAG [i,j]=d DFAG The weight matrix ∈[0,1] records the degree of correlation between the flow pattern of each region constructed by the dynamic flow and other regions.
[0023] Furthermore, in step 2, SIF is used. t The SIF of time t is expressed mathematically as follows:
[0024]
[0025] Among them, D ab λ represents the basic connectivity between or within regions; λ represents the scaling parameter; γ is a parameter used to determine the weights at the tail of the function; P t This indicates the proportion of people moving within a small-scale census area during time period t, relative to the total number of people moving in and out of the area.
[0026] Furthermore, the early epidemic infection risk quantification method of the present invention also includes a risk overlay module for quantifying the impact of historical risk values on regional risk values. Specifically, when calculating time t, risk values from historical time 0 to time t-1 are included in the calculation, as shown in the following formula:
[0027]
[0028] in, β represents the sum of risk values for county / district i from historical time 0 to t-1, and β represents the probability of viral infection transmission, which is adjusted according to the diverse and dynamic viral situation.
[0029] Compared with the prior art, the present invention has the following advantages:
[0030] 1. This invention proposes a novel method that can effectively quantify the infection risk level of a given geographical area without requiring in-depth knowledge of the details of the infectious disease or the spread of the virus.
[0031] 2. The method of the present invention integrates analysis at multiple geographical granularities, comprehensively considering intra-regional and cross-regional mobility to accurately quantify risk levels.
[0032] 3. The inventors evaluated the method of this invention based on actual COVID-19 data from four US states. The results showed that the method of this invention performed excellently across short-, medium-, and long-term risk quantification scenarios. The risk quantification method of this invention demonstrated high performance across all time categories and overall accuracy. Detailed Implementation
[0033] The present invention will be further described below with reference to specific embodiments, but the present invention is not limited thereto.
[0034] To describe the invention more clearly, the following definitions are provided.
[0035] Definition of Census Communities and Counties: The areas involved in this invention are mainly divided into two types: small-scale census areas and county-level census areas. The difference between small-scale census areas and counties lies in their different levels of geographical granularity. Small-scale census areas are smaller regions with a population of approximately 3,000 to 5,000 people, while counties are larger regions based on county boundaries. In this invention, "small-scale census area" is used as a geographical unit one level smaller than "county".
[0036] Define a two-region location map: model these regions as a directed graph. Here, V represents the set of small-scale census tracts, E represents the set of edges, and W represents the edge weights of the graph. The weight of each edge is not calculated solely based on the geographical distance between different regions. Instead, it is determined by considering both geographical distance and the frequency of interactions between regions. This will be discussed in the section on model structure and algorithm design.
[0037] Definition 3: Areas at risk of pandemic infection: During a certain period [t] f ,t p Within [the context], the risk of infection (R) for a certain infectious disease. d =[t p ,t f [N] is defined as the number of new cases in region d. d The sum of (t), where the calculation method is as follows: If at some future moment t f The number of new cases in region d is relatively high, and d can be designated as the region in period [t]. p ,t f The high-risk counties within [ ]. The specific methods for determining high-risk counties will be explained in detail below. It should be noted here that N d Quantification and period of (t) [t] p ,t f The settings may vary depending on various factors such as actual circumstances and pandemic forecasts, which will be discussed below.
[0038] The objective of this invention is to extract historical mobility data from various counties. Starting point, input into the infection risk quantification model. In order to predict the infection risk value of each county at a given point in time, that is, the infectious disease vulnerability of each county, the spatial and temporal risk area is quantified.
[0039] Specifically, suppose we want to predict the pandemic risk value of a certain disease in county n at time T+t, denoted as R. n Given (T+t), we can obtain the following expression:
[0040]
[0041] Unlike methods that rely on historical infection data, the early epidemic infection risk quantification method based on regional mobility of this invention can quickly and accurately quantify regional viral infection risk and predict future disease transmission trends. The entire method consists of four main modules: (1) Dynamic mobility flow perception map: It uses a data-driven approach to discover differences in flow patterns between different regions and measure the spatial heterogeneity of risk transmission. (2) Intra-regional factor modeling: It measures the risk of epidemic transmission caused by intra-county mobility based on social interaction models and uses this information to estimate the probability of virus transmission. (3) Inter-regional factor modeling: It considers the interaction between counties and measures the degree to which epidemic transmission is affected by distance. (4) Intra-regional-inter-regional integrated risk quantification: By integrating intra-regional factor modeling and inter-regional factor modeling, the infection risk value of all small-scale census areas in the study county is calculated.
[0042] Next, the method for quantifying early epidemic infection risk based on regional mobility according to the present invention will be described in detail.
[0043] Construction of Dynamic Mobile Traffic Awareness Map
[0044] In this section, the inventors will discuss how to use human mobility data to construct epidemic risk weight maps. They propose a data-driven approach called Dynamic Flow-Aware Mapping (DFAG), which utilizes the origin-destination (OD) matrix of human mobility to determine spatial correlations between regions. Epidemic risk weight maps are constructed from mobility data to capture trends in epidemic spread and identify high-risk areas.
[0045] Specifically, by analyzing daily origin-destination (OD) data flows between regions, Wasserstein distance is used to capture the similarity of movement probability distributions among small census areas, representing the dynamic spatial dependencies between them. This method of measuring the similarity of population movement patterns can provide a better understanding of virus transmission trends. For example, if both regions have large populations moving to high-risk areas, this movement pattern may increase their likelihood of contracting the virus.
[0046] Suppose there are N regions. For each region i, we can obtain its daily inflow data Z to all other regions. d Consider it as the vector of day d Right now Where d∈[1,D]. The daily flow pattern information of each region is obtained by calculating the magnitude of the flow vector, that is, the daily inflow data pattern from region i to all other regions j in the time interval D.
[0047]
[0048] In this equation, P n Represents the probability distribution obtained from the vector sequence, which indicates the proportion of people flowing from region i to region j on day d.
[0049] Then we need to obtain the transformation cost for each probability quality. We use the dot product similarity between pedestrian flow vectors as the cost function.
[0050]
[0051] Incorporating the above formula into the calculation of the Wasserstein distance yields a metric called dynamic flow-aware distance, used to measure flow differences between regions. The flow-aware distance between region i and region j on day d can then be calculated.
[0052]
[0053] By applying the formula to the OD matrix of each small-scale census area to calculate the Wasserstein distance, a weight matrix W can be obtained. DFAG W DFAG [i,j]=d DFAG ∈[0,1]. This matrix records the degree of correlation between the flow patterns of each region constructed by dynamic flows and other regions, that is, the degree of correlation between residents with the same flow patterns and other regions.
[0054] County-level intra-district modeling based on social interaction model
[0055] Intra-regional mobility plays a crucial role in the spread of infectious diseases. Particularly within counties comprised of multiple small census districts, the interactions between these districts can vividly depict intra-regional mobility, thus revealing one of the main drivers of virus transmission. To quantify this impact, a Dynamic Mobility Flow Awareness Map (DFAG) was generated at the small census district level using the aforementioned formula. However, it is important to note that infectious disease outbreaks may occur at a more granular geographical scale. Obtaining such fine-grained data can be challenging due to privacy and other concerns. Therefore, alternative measurement methods have been explored to assess risk at a more detailed geographical level.
[0056] In this invention, the Spatial Interaction Function (SIF) model is employed. This model, based on small-scale census tracts, describes the probability of any two small-scale census tracts within a county forming a connection (i.e., making contact) at a given distance. Using SIF... t Let SIF represent time t, and its mathematical expression is shown below.
[0057]
[0058] Among them, D ab Representing the fundamental connectivity between or within regions, it can be interpreted as the probability of face-to-face contact between or within two regions. λ represents a scaling parameter used to define the phenomenological distance unit of probability decay, while γ is a parameter used to determine the weight of the tail of the function (higher values indicate fewer long-range connections, etc.).
[0059] Using small-scale census tracts as the smallest spatial analysis unit, and referencing this spatial interaction function model, the risk of infection associated with movement within small-scale census tracts is assessed. Here, SIF represents the intra-regional infection model, and P... t This represents the proportion of people moving within a small-scale census area during time period t, relative to the total inflow and outflow of the area. The model parameters are fixed at λ² = 0.035 and γ² = 6. Finally, the infection risk value generated by intra-county movement within time period t is calculated using the following formula:
[0060]
[0061] Among them, Z ii Z represents the number of individuals remaining in county i during time period t, while Z represents the number of individuals remaining in county i during time period t. ba and Z ab These represent the number of individuals who migrated from region a to region b and from region b to region a during time period t, respectively.
[0062] Inter-county interaction modeling
[0063] The spread of infectious diseases is influenced by a complex interplay of factors, including lifestyle, economic conditions, and transportation, and this influence extends across different geographical regions. Therefore, in addition to the potential for infection due to intra-regional movement, inter-county population movement patterns constitute a key mechanism for disease transmission. The objective here is to effectively capture these movement patterns to reveal the underlying mechanisms of disease transmission and provide a basis for prevention and control strategies.
[0064] The cross-county infection risk quantification map focuses on the infection risk caused by inter-county migration at the county level. This infection risk is mainly related to the risk transmission correlation of inflows and outflows between counties. The impact of inter-county movement is quantified using the inflow and outflow volumes of each county, and a county-level dynamic mobility flow perception map (DFAG) is generated using the previous formula to dynamically quantify the risk transmission correlation.
[0065] In this figure, the inter-county mobility risk of county i in time period t can be obtained by the following formula:
[0066]
[0067] Among them, W DFAG (i,j) is a weight matrix representing the correlation between the flow pattern of county i and county j, constructed from dynamic flows. ji Z represents the number of people who migrated from county j to county i during time period t, while Z... ij This represents the number of people who migrated from county i to county j within the time period t.
[0068] Integration of inter-county risk quantification and intra-county risk quantification
[0069] As described above, a quantitative model is used to assess the risk values from inter-county and intra-county flows to measure the risk value of a county / district. Then, assuming that the risks generated by these two flows are equivalent in disease transmission, they are weighted and summed. This is consistent with the mechanism of disease transmission, i.e., the probability of disease transmission does not change due to different modes of movement and scales. First, each risk value is normalized by dividing it by the maximum value across all counties / districts and time periods. Then, the risk values between counties and within counties / districts are weighted according to their respective weights to obtain the risk value of county / district i at time t. Finally, by summing the two weighted risk values, the risk value of county i is obtained, as shown in the following formula:
[0070]
[0071] Furthermore, the inventors referenced the transmission mechanisms of diseases and discovered that most diseases have an incubation period during which patients are asymptomatic but infectious. Therefore, risks generated in a region in the past may persist in future risk values. This impact is influenced by the virus itself. To simulate actual epidemic situations involving transmission through movement, the inventors designed a risk overlay module to quantify the impact of historical risk values on regional risk values. Specifically, when calculating time t, the risk values from historical time 0 to time t-1 also need to be included in the calculation, as shown in the following formula:
[0072]
[0073] in, β represents the sum of risk values for county / district i from historical time 0 to t-1, and β represents the probability of viral infection transmission, which can be adjusted according to the diverse and dynamic viral situation.
[0074] The quantitative method of this invention yielded a continuous and dynamic epidemic transmission risk value R within the study area. i .
[0075] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for quantifying the risk of early epidemic infection based on regional flow, characterized in that, Comprising: Step 1: Constructing a dynamic mobile flow-aware graph using human mobility data to obtain a weight matrix W that records the degree of association between the flow patterns of each region built by dynamic flow and other regions DFAG wherein the human mobility data is a start-end matrix of human mobility; Step 2: Calculate the epidemic infection risk value caused by the flow within the county based on the spatial interaction function model SIF, wherein the infection risk value generated by the flow within the county i in the time period t is calculated by the following formula: where Z ii represents the number of individuals remaining in county i over time period t, while Z ba and Z ab represent the number of individuals migrating from area a to area b and from area b to area a, respectively, over time period t. Step 3: Quantify the influence of inter-county flow based on the inflow and outflow quantity of the county, generate a county-level dynamic mobile flow perception map, and dynamically quantify the cross-county transmission risk, wherein the cross-county flow risk value of county i in the time period t is obtained by the following formula: wherein W DFAG (i,j) is a weight matrix of the degree of association between the flow pattern of county i constructed from dynamic flows and county j, Z ji represents the number of people who migrated from county j to county i in the time period t, and Z ij represents the number of people who migrated from county i to county j in the time period t; Step 4: Calculate the infection risk value of the county by integrating the inter-county risk quantification and the intra-county risk quantification, wherein each risk value is divided by the maximum value of all counties and time periods to be normalized, and then the inter-county and intra-county risk values are weighted by respective weights to obtain the risk value of the county i As shown in the following formula:
2. The method of quantifying the risk of infection of an early epidemic according to claim 1, characterized in that, In said step 1, assuming there are N regions, for each region i, its daily inflow data Z d Vector considered as day d That is where d e [1, D], the daily flow pattern information for each region is obtained by calculating the size of the flow vector, i.e., the daily inflow data pattern from region i to all other regions j in time interval D: In this equation, P n represents the probability distribution obtained from the vector sequence, which represents the proportion of human flow from region i to region j on day d; Then use the dot product similarity between human flow vectors as the cost function, The above formula is incorporated into the calculation of Wasserstein distance to obtain a measure called dynamic flow perception distance, which is used to measure the flow difference between regions; Calculate the flow perception distance between region i and region j on the dth day: The Wasserstein distance is calculated for each origin-destination matrix of the small-scale census zones by applying the formula to the origin-destination matrix DFAG where W DFAG [i,j] = d DFAG ∈ [0,1], the weight matrix records the degree of association between the flow pattern of each zone built from dynamic flows and other zones.
3. The method of quantifying the risk of infection of an early epidemic according to claim 1, characterized in that, In the step 2, SIF is used t SIF representing time t is mathematically expressed as follows: where D ab represents the basic connectivity between regions or within a region; λ represents the scaling parameter; γ is a parameter that determines the weight of the tail of the function; P t represents the proportion of the small-scale census population within the region that flows within the region during the time period t.
4. The method of quantifying the risk of infection of an early epidemic according to claim 1, characterized in that, Also includes a risk superposition module for quantifying the influence of historical risk values on regional risk values, wherein when calculating time t, the risk values from historical time 0 to time t-1 are included in the calculation, as shown in the following formula: where, represents the risk value superposition of county i at historical time 0 to t-1, β represents the transmission probability of virus infection, which is adjusted according to the diversification and dynamics of the virus.
Citation Information
Patent Citations
Early risk situation analysis method for epidemic situation of infectious disease based on input-diffusion function
CN111063451A
Epidemic situation risk grade assessment method based on people flow density
CN111128399A