Infectious disease early warning prediction method and system based on regional network

By constructing regional network maps and monitoring the changes in the minimum spanning tree weight, the problem of seasonal infectious disease warning is solved, early warning and resource optimization are achieved, and the accuracy and efficiency of public health decisions are improved.

CN120260968APending Publication Date: 2025-07-04THE CENT FOR DISEASE CONTROL & PREVENTION OF XINJIANG UYGUR AUTONOMOUS REGION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334245.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively warn of the outbreak of seasonal infectious diseases, resulting in tight medical resources and loss of productivity.

Method used

The infectious disease warning prediction method based on regional networks is constructed, and edge weights are calculated using the network entropy method. The Prim algorithm and Logistic model are combined to monitor the weight changes of the minimum spanning tree, and the CUSUM control chart is used for early warning.

Benefits of technology

It has achieved early warnings for seasonal infectious diseases, improved the accuracy and efficiency of public health decisions, and reduced medical resource shortage and productivity losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260968A_ABST
    Figure CN120260968A_ABST
Patent Text Reader

Abstract

The invention discloses an infectious disease early warning prediction method and system based on a regional network, and belongs to the field of infectious disease monitoring, and the prediction method comprises the following steps: constructing a dynamically evolved regional network diagram by using real-time population flow data according to the adjacent information of a geographic position; sequentially processing the infectious disease sequence data until the whole sequence is processed; calculating the weight of each edge ek = (vi, vj) of the regional network graph by using a network entropy method to obtain a weighted undirected graph; giving a minimum spanning tree and a corresponding minimum weight Lt according to a Prim algorithm; and observing whether the Lt is greater than a preset threshold value h1, monitoring whether the accumulated deviation of the Lt time sequence is greater than a preset threshold value h2, if the Lt or the accumulated deviation is greater than the preset threshold value, determining that the current moment t is an early warning time point, otherwise, returning to monitor the next moment. According to the scheme, the occurrence condition of infectious diseases in a region can be well monitored, and early warning is carried out when early warning is needed, so that people and related organizations and units can carry out early warning and prevention in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of disease early warning, and relates to an infectious disease early warning prediction method and system based on a regional network. Background Art

[0002] Infectious diseases generally refer to a class of diseases caused by various pathogens that are transmitted between humans or animals, while seasonal infectious diseases refer to infectious diseases that tend to occur more frequently in different seasons due to changes in environmental and climatic factors. Common seasonal infectious diseases include influenza (seasonal flu), norovirus, chickenpox, mumps, insect-borne infectious diseases, and hand, foot and mouth disease. Seasonal infectious diseases can lead to large-scale epidemics and pose a great threat to people's health. For example, influenza causes millions of severe illnesses and 290,000 to 650,000 deaths worldwide each year; and there may be sequelae and complications: for example, chickenpox may latently cause shingles, and severe dengue fever may cause shock, etc. Seasonal outbreaks of infectious diseases (such as influenza) will also increase the burden on hospitals, affect the treatment of other diseases, cause serious crowding out of medical resources, and cause shutdowns, resulting in serious productivity losses.

[0003] Therefore, there is an urgent need for a method that can provide early warning and prediction for seasonal infectious diseases that are about to occur on a large scale. Summary of the invention

[0004] The purpose of the present invention is to provide an infectious disease early warning prediction method and system based on a regional network, which is based on dynamic network markers, combined with minimum spanning tree, Logistic model, CUSUM control chart method, etc., to propose an infectious disease early warning prediction method and system based on a regional network. Early warning prediction can be carried out for seasonal and periodic infectious diseases, and the people or relevant departments can be reminded in advance to take protective measures.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for early warning and prediction of infectious diseases based on a regional network comprises the following steps performed in sequence:

[0007] Step S1: Based on the adjacency information of the geographical location, a dynamically evolving regional network diagram is constructed using real-time population mobility data such as mobile phone signaling and traffic flow;

[0008] Step S2: Process the infectious disease sequence data in sequence, starting from time t=1, until the entire sequence is processed;

[0009] Step S3: Using the network entropy method, calculate each edge e of the regional network graph k =(v i ,v j )’s weights, and after calculation, a weighted undirected graph is obtained;

[0010] Step S4: According to Prim's algorithm, a minimum spanning tree and the corresponding minimum weight corresponding to the weighted undirected graph at time t are given. The minimum weight is recorded as L t ;

[0011] Step S5: Observe L t Is it greater than the preset threshold h1? At the same time, monitor L t Whether the cumulative deviation of the time series is greater than the preset threshold h2, if L t Or if the accumulated deviation is greater than the preset threshold, the current time t is the warning time point, otherwise the process returns to step S2.

[0012] In order to better implement this solution, further, in step S3, the network entropy method is used to calculate each edge e of the regional network graph. k =(v i ,v j ) The specific calculation formula is:

[0013]

[0014] in, Represents edge e k , SD is the variance of the corresponding node, PCC is the Pearson coefficient between the two variables, δ is the absolute value of the variance change of the node at time t and the previous time, LNE is the local network entropy of the corresponding node, the network entropy is constructed according to the DNM principle, and L is the number of neighbors of the regional node. Among them, represents the time series vector of cases in region v at time t, Represents the case time series vector set of the neighboring areas of area v, edge weight Means the quantized area v i With v j risk of transmission between them.

[0015] In the dynamic change formula of variance δ, SD t (v i ,v j ) represents the region v i and adjacent area v j The joint variance of the two nodes is calculated and averaged, which represents the joint change amplitude of the fluctuation of the two regions. The variance δ is used to amplify the weight of LNE. When the fluctuation of the two regions is drastic, the edge weight Significantly increased. Edge weight In the formula, the absolute value of LNE of the two nodes is averaged and then multiplied by δ to ensure that the edge weight is non-negative. If the LNE of the two nodes is negative and the absolute value is large (i.e., the fluctuation is violent and the neighbor correlation is concentrated), and δ is large (fluctuation mutation), then the edge weight Significantly increased, indicating edge e k has a high transmission risk.

[0016] In the formula of LNE, the entropy term is the Shannon entropy, which is used to measure the uncertainty of neighbor correlation. The larger the entropy, the more dispersed the influence of neighbors on v i , that is, there is no dominant propagation direction; the smaller the entropy, the more concentrated the propagation of a few neighbors, that is, centralized propagation. is the coefficient of variation of variance, is the variance of the time series of its own cases in region v i , reflecting the intensity of fluctuations. The numerator part in the coefficient of variation of variance is used to capture the sudden change of the self-fluctuation in region v i , such as a sudden increase in cases. The denominator logL is used to standardize the entropy value range, playing a role similar to normalization. LNE combines the self-fluctuation change and the uncertainty of neighbor correlation distribution. Due to the existence of the negative sign and absolute value, LNE t (v i ) ≤ 0. The closer LNE is to 0, the more stable the propagation mode of region v i ; the more negative LNE is, the more intense the fluctuations and the more concentrated the neighbor correlation, that is, the more likely it is to be a propagation hot spot.

[0017] p i (t) represents the probability distribution. p i (t) normalizes the absolute value of the Pearson correlation coefficient (PCC) of the case sequences in region v i and its neighbor region v j into a probability distribution. The larger p i (t) is, the more highly correlated the case changes in region v i are with its neighbors, and it may be a key node on the propagation path. The dynamic variance change δ is incorporated into the probability distribution to capture the suddenness of case fluctuations between two regions, and the local network entropy LNE is also incorporated to measure the uncertainty of the region's own case pattern and its influence on neighbors. This method has three advantages:

[0018] (1) Dynamics: By the variance change terms in δ and LNE, capture the suddenness of the evolution of infectious diseases and the transfer of propagation paths.

[0019] (2) Spatial correlation: Use PCC to construct a probability distribution of neighbor correlation, reflecting the statistical dependence of case changes between regions;

[0020] (3) Resistance to noise dependence: The calculation of entropy reduces the sensitivity to a single outlier and pays more attention to the overall distribution pattern.

[0021] To better implement this solution, further, in step S5, observe L tThe method of determining whether the minimum weight L is greater than the preset threshold h1 is as follows: using the Logistic algorithm, the S-shaped function is used to convert all the minimum weight L t Mapped to between 0 and 1, the specific formula is:

[0022]

[0023] Among them, ω is the weight obtained through data training;

[0024] Observe P(y t =1|L t ; ω) is greater than the preset threshold h1.

[0025] In order to better implement this solution, further, the preset threshold h1 is set to 0.5.

[0026] In order to better implement this solution, further, in step S5, monitoring L t The method for determining whether the cumulative deviation of the time series is greater than the preset threshold h2 is as follows:

[0027] Based on the historical period without warning, L t Calculate the mean μ0 of the sequence and calculate the history L t The standard deviation σ of the sequence, set the sensitivity parameter σ0 = Nσ, N is the set minimum amplitude constant;

[0028] Step S501: Minimum weight L at each time t t , calculate the standardized deviation S t : This step is used to eliminate the dimensional effect and facilitate unified monitoring;

[0029] Step S502: Calculate the minimum weight L t The positive cumulative sum Here we use It is used to adjust the sensitivity to small changes and avoid noise interference;

[0030] Step S503: If Then the time t is determined to be a mutation point. If the time t is a mutation point, reset the forward cumulative sum Avoid historical bias from affecting subsequent detection.

[0031] In addition, in this step, the baseline can be adjusted dynamically according to seasonal factors. That is, if the network structure changes over a long period of time, such as seasonal factors, μ0 and σ can be re-estimated using a sliding window, for example, once every 7 days. Furthermore, h2 can also be adjusted dynamically according to the distribution of recent cumulative sums, for example, h = 3 × recent The rolling standard deviation of .

[0032] To better implement this solution, further, in step S1, each node represents a region, and the connecting edges between nodes represent the adjacency relationship between regions.

[0033] To better implement this solution, further, step S2 is specifically as follows: Process the infectious disease sequence data in turn in the manner of a sliding window. Each moment t corresponds to a sliding window. Starting from moment t = 1, continue until the entire sequence is processed.

[0034] An infectious disease early warning and prediction system based on a regional network, based on the foregoing early warning and prediction method, includes:

[0035] Regional Linking Module: Construct a corresponding regional network diagram according to the adjacency information of geographical locations;

[0036] Processing and Generation Module: Process the infectious disease sequence data in turn, starting from moment t = 1 until the entire sequence is processed; Use the network entropy method to calculate the weight of each edge e k =(v i , v j ) of the regional network diagram, and obtain a weighted undirected graph after calculation; According to the Prim algorithm, for the weighted undirected graph corresponding to moment t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ; Use the Prim algorithm, for the weighted undirected graph corresponding to moment t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ;

[0037] Early Warning and Identification Module: Use the Logistic algorithm to map all the minimum weights L t to between 0 and 1 through the S-shaped function. The specific formula is:

[0038]

[0039] where ω is the weight value obtained through data training;

[0040] In addition, the Early Warning and Identification Module also uses the CUSUM control chart monitoring method to monitor the mutation points of the L t sequence. Calculate the mean μ0 based on the L t sequence of the historical non-early warning period, calculate the standard deviation σ of the historical L t sequence, and set the sensitivity parameter σ0 = Nσ, where N is the set minimum amplitude constant, usually 1;

[0041] For the minimum weight L t at each moment t, calculate the standardized deviation S t :

[0042] Calculate the minimum weight Lt Forward cumulative sum

[0043] If is greater than the preset threshold h2 or P(y t = 1|L t ; ω) is greater than the preset threshold h1, then it is determined that the current time t is the early warning time point, otherwise the current time t is not the early warning time point.

[0044] To better implement this solution, further, in the processing and generation module, the network entropy method is used to calculate the weight of each edge e of the regional network diagram k =(v i ,v j ) The specific calculation formula of the weight is:

[0045]

[0046] Among them, represents the weight of edge e k , SD is the variance of the corresponding node, PCC represents the Pearson coefficient between two variables, δ represents the absolute value of the variance change of the node at time t and the previous time, LNE represents the local network entropy of the corresponding node, the network entropy is constructed based on the DNM principle, and L represents the number of neighbors of the regional node.

[0047] This solution has outstanding advantages in spatial correlation modeling, computational efficiency, interpretability, and scalability. It is especially suitable for large-scale infectious disease monitoring scenarios with limited resources but requiring rapid response, providing reliable technical support for public health decision-making.

[0048] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0049] 1. The method for early warning and prediction of infectious diseases based on a regional network described in the present invention has the ability to capture the spatial spread dynamics. By constructing a network through regional adjacency information, it directly reflects the physical law of the spread of infectious diseases along geographical paths, avoiding the defect of pure statistical models ignoring spatial correlation, and dynamically calculating the edge weights to capture the information transmission intensity of case fluctuations between regions, such as the potential impact of a sudden increase in cases in a certain region on adjacent regions, which is more sensitive than traditional correlation coefficients;

[0050] 2. The method for early warning and prediction of infectious diseases based on a regional network described in the present invention uses the key path extraction and dimensionality reduction analysis method to simplify the complex network into key propagation paths, focusing on core risk links such as inter-provincial transportation arteries, reducing the noise interference of high-dimensional data, and at the same time identifying network anomalies through the sudden increase of the total weight L of the MST t , which is more efficient than monitoring all edge weights and can reflect global structural changes;

[0051] 3. The infectious disease early warning and prediction method based on regional network described in the present invention has interpretability and decision support. The network edge corresponds to the actual geographical connection. The MST result can be directly mapped to the actual transmission path, such as the city chain along the Beijing-Guangzhou line, which is convenient for locating high-risk areas. In addition, the preset threshold can be set in combination with public health experience, such as L in the historical outbreak period. t Statistical quantiles make it easy for decision makers to understand and adjust strategies;

[0052] 4. The method for early warning and prediction of infectious diseases based on regional networks described in the present invention has low computational complexity of the system constructed by the method: the complexity of the Prim algorithm is O(ElogV), which is suitable for real-time updates of medium-sized regional networks, such as 300+ city-level nodes across the country, and can perform incremental processing and process data streams moment by moment, without the need to roll back the full amount of historical data, and is suitable for online monitoring scenarios;

[0053] 5. The infectious disease early warning and prediction method based on regional network described in the present invention is compatible with multi-source data. The network construction in step S1 can integrate dynamic data such as population mobility and traffic flow to improve the model precision, and the early warning mechanism is scalable. The CUSUM control chart method is used to balance the sensitivity and false alarm rate, and other graph features such as node centrality can be combined to construct a comprehensive index;

[0054] 6. The infectious disease early warning prediction method based on regional network described in the present invention has high practical application value, and early warning is achieved through L t Sudden increases capture structural changes in the transmission network. For example, if a certain place becomes a new hub, it may trigger an early warning several days earlier than the obvious increase in the number of cases. It can also optimize resources, such as outputting the key edges of the MST during early warning to guide precise blockades or material delivery, such as focusing on monitoring traffic nodes connecting two outbreak areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative work, among which:

[0056] Figure 1 It is a flow chart of the infectious disease early warning and prediction method based on regional network described in the present invention. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0058] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0059] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0060] The features and performance of the present invention will be further described in detail below in conjunction with the embodiments.

[0061] Embodiment 1

[0062] A method for early warning and prediction of infectious diseases based on a regional network provided by a preferred embodiment of the present invention, as Figure 1 shown, includes the following steps carried out in sequence:

[0063] Step S1: According to the adjacency information of geographical locations, use real-time population flow data such as mobile phone signaling, traffic flow, etc. to construct a dynamically evolving regional network diagram;

[0064] Step S2: Process the infectious disease sequence data in sequence, starting from time t = 1 until the entire sequence is processed;

[0065] Step S3: Use the network entropy method to calculate each edge e of the regional network diagram k =(v i , v j)'s weight, and after calculation, a weighted undirected graph is obtained;

[0066] Step S4: According to Prim's algorithm, for the weighted undirected graph corresponding to time t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ;

[0067] Step S5: Observe L t Whether it is greater than the preset threshold h1, and at the same time monitor whether the cumulative deviation of the time series of L t is greater than the preset threshold h2. If L t or the cumulative deviation is greater than the preset threshold, then the current time t is the early warning time point, otherwise return to execute Step S2.

[0068] In Step S1, the use of real-time population mobility data can more accurately reflect the movement law of the population, so as to construct a dynamically evolving regional network and enhance the accuracy of the infectious disease transmission model. Its data sources generally come from:

[0069] (1) Mobile signaling data: Obtain the user's base station handover records through the operator to identify cross-regional movement trajectories (high precision, but need to be desensitized);

[0070] (2) Transportation ticket data: Departure-arrival records of railways, airlines, and highways (such as 12306 train ticket data, flight steward API);

[0071] (3) Internet platform data: Real-time traffic conditions, navigation requests, and location service (LBS) heat maps of Baidu Maps / Amap;

[0072] (4) Smart device data: Origin-destination (OD) data of shared bicycles and online car-hailing.

[0073] These data first need to be preprocessed. The preprocessing link generally includes cleaning and denoising links to eliminate abnormal trajectories and impute missing values. Then, spatial aggregation is required to map the original longitude and latitude to administrative regions. Finally, a time alignment link is required to summarize the flow volume according to a fixed time window (such as 1 hour / day) to generate a time series. After preprocessing, we need to define network nodes and edges. The nodes are generally administrative regions such as cities and districts, and each node represents a geographical unit; the edge is the population flow connection between regions, and its weight changes dynamically with time.

[0074] Embodiment 2

[0075] In the said Step S3, using the network entropy method, the specific calculation formula for calculating the weight of each edge e k =(v i , v j ) of the regional network graph is:

[0076]

[0077] in, Represents edge e k , SD is the variance of the corresponding node, PCC is the Pearson coefficient between the two variables, δ is the absolute value of the variance change of the node at time t and the previous time, LNE is the local network entropy of the corresponding node, the network entropy is constructed according to the DNM principle, and L is the number of neighbors of the regional node. Among them, represents the time series vector of cases in region v at time t, Represents the case time series vector set of the neighboring areas of area v, edge weight Means the quantized area v i With v j risk of transmission between them.

[0078] In the dynamic change formula of variance δ, SD t (v i ,v j ) represents the region v i and adjacent area v j The joint variance of the two nodes is calculated and averaged, which represents the joint change amplitude of the fluctuation of the two regions. The variance δ is used to amplify the weight of LNE. When the fluctuation of the two regions is drastic, the edge weight Significantly increased. Edge weight In the formula, the absolute value of LNE of the two nodes is averaged and then multiplied by δ to ensure that the edge weight is non-negative. If the LNE of the two nodes is negative and the absolute value is large (i.e., the fluctuation is violent and the neighbor correlation is concentrated), and δ is large (fluctuation mutation), then the edge weight Significantly increased, indicating that edge e k The risk of transmission is high.

[0079] In the LNE formula, the entropy term is Shannon entropy, which is used to measure the uncertainty of neighbor correlation. The larger the entropy, the more likely the neighbor is to have a close relationship with v. i The more dispersed the influence is, the more there is no dominant propagation direction; the smaller the entropy is, it means that a few neighbors dominate the propagation, that is, centralized propagation. is the coefficient of variation of variance, is the region v i The variance of the case time series reflects the intensity of fluctuations. The numerator of the variance variation coefficient is used to capture the region v i The denominator logL is used to standardize the entropy range, which plays a similar role to normalization. LNE combines the uncertainty of its own fluctuations and the distribution of neighbor correlations. Due to the existence of negative signs and absolute values, LNE t (v i) ≤ 0, the closer LNE is to 0, the more stable the propagation mode of region v i ; the more negative LNE is, the more intense the fluctuations and the more concentrated the neighbor correlation, that is, it is more likely to be a propagation hotspot.

[0080] p i (t) represents the probability distribution, and p i (t) normalizes the absolute value of the Pearson correlation coefficient (PCC) of the case sequences in region v i and its neighbor region v j into a probability distribution. The larger p i (t) is, the more highly correlated the case changes in region v i are with its neighbors, and it may be a key node on the propagation path. The dynamic variance change δ is incorporated into the probability distribution to capture the suddenness of case fluctuations between two regions, and the local network entropy LNE is also incorporated to measure the uncertainty of the case pattern in the region itself and its impact on neighbors. This method has three advantages:

[0081] (1) Dynamics: By the variance change terms in δ and LNE, capture the suddenness of the evolution of infectious diseases and the transfer of propagation paths.

[0082] (2) Spatial correlation: Use PCC to construct a probability distribution of neighbor correlation, reflecting the statistical dependence of case changes between regions;

[0083] (3) Resistance to noise dependence: The calculation of entropy reduces the sensitivity to a single outlier and pays more attention to the overall distribution pattern.

[0084] The following gives an example demonstration:

[0085] Suppose at a certain moment t, region vi has 3 neighbors, that is, L = 3, and the PCC values of the 3 neighbors are [0.8, 0.2, 0.1], then:

[0086]

[0087] We choose the natural logarithm e as the logarithm base of entropy, so:

[0088] Entropy = -(0.727ln0.727 + 0.182ln0.182 + 0.091ln0.091) ≈ 0.760 nats

[0089] If then:

[0090]

[0091] If the LNE of the neighbor region v j is t (v j ) = -1.5 and δ = 1.2, then:

[0092]

[0093] Example 3

[0094] Based on Example 1, in step S5, the method for observing whether L t is greater than the preset threshold h1 is as follows: Using the Logistic algorithm, all the minimum weights L t are mapped to between 0 and 1 through the S-shaped function. The specific formula is:

[0095]

[0096] where ω is the weight value obtained through data training;

[0097] Observe whether P(y t =1|L t ; ω) is greater than the preset threshold h1. Here, the preset threshold h1 can be set to 0.5.

[0098] In step S5, the method for monitoring whether the cumulative deviation of the L t time series is greater than the preset threshold h2 is as follows:

[0099] Based on the L t sequence in the historical no-warning period, calculate the mean μ0, calculate the standard deviation σ of the historical L t sequence, and set the sensitivity parameter σ0 = Nσ, where N is the set minimum amplitude constant;

[0100] Step S501: For the minimum weight L t at each moment t, calculate the standardized deviation S t : This step is used to eliminate the influence of dimensions for convenient unified monitoring;

[0101] Step S502: Calculate the forward cumulative sum t of the minimum weight L Here, is used to adjust the sensitivity to small changes and avoid noise interference;

[0102] Step S503: If then determine that the moment t is a mutation point. If the moment t is a mutation point, reset the forward cumulative sum to avoid the influence of historical deviations on subsequent detections.

[0103] In addition, in this step, dynamic baseline adjustment can be performed according to seasonal factors, that is, if the network structure changes over a long period, such as seasonal factors, etc., a sliding window is used to re-estimate μ0 and σ. For example, it can be updated every 7 days. Further, h2 can also be dynamically adjusted according to the distribution of the recent cumulative sum. For example, take h = 3 × recent rolling standard deviation.

[0104] Here we give an example. Suppose the historical mean μ0 = 10 and the standard deviation σ = 2 of a certain sequence L in a place, the minimum amplitude constant N = 1 is set, and the preset threshold h2 = 3.5. Then the final early warning result in the following data is t It can be seen that on the fourth day

[0105]

[0106] triggers an early warning.

[0107] Example 4

[0108] This example is a further supplementary explanation of Example 1. In step S1, each node represents a region, and the connecting edges between the nodes represent the adjacency relationship between regions.

[0109] Step S2 is specifically as follows: The infectious disease sequence data is processed sequentially in the way of a sliding window. Each moment t corresponds to a sliding window, starting from the moment t = 1 until the entire sequence is processed.

[0110] Example 5

[0111] Based on the early warning prediction method described in any one of Examples 1 to 4, this example constructs an infectious disease early warning prediction system based on a regional network, including:

[0112] Regional link module: Construct a corresponding regional network diagram according to the adjacency information of geographical locations;

[0113] Processing and generating module: Process the infectious disease sequence data sequentially, starting from the moment t = 1 until the entire sequence is processed; Use the network entropy method to calculate the weight of each edge e k =(v i , v j ) of the regional network diagram, and obtain a weighted undirected graph after calculation; According to the Prim algorithm, for the weighted undirected graph corresponding to the moment t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ; Use the Prim algorithm, for the weighted undirected graph corresponding to the moment t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ;

[0114] Early warning recognition module: Using the Logistic algorithm, all the minimum weights L are mapped to between 0 and 1 through the S-shaped function. The specific formula is: t Specifically,

[0115]

[0116] where ω is the weight obtained through data training;

[0117] In addition, the early warning recognition module also uses the CUSUM control chart monitoring method to monitor the mutation points of the L sequence. Based on the L sequence in the historical non-warning period, the mean μ0 is calculated, the standard deviation σ of the historical L sequence is calculated, and the sensitivity parameter σ0 = Nσ is set, where N is the set minimum amplitude constant, generally 1; t For the minimum weight L at each moment t, t calculate the standardized deviation S: t Specifically,

[0118] For the minimum weight L at each moment t, t calculate the standardized deviation S: t Specifically,

[0119] Calculate the forward cumulative sum of the minimum weight L: t Specifically,

[0120] If is greater than the preset threshold h2 or P(y t = 1|L t ; ω) is greater than the preset threshold h1, then it is determined that the current moment t is the early warning time point, otherwise the current moment t is not the early warning time point.

[0121] Example 6

[0122] Based on Example 5, in this example, in the processing and generation module, the network entropy method is used. The specific calculation formula for the weight of each edge e k =(v i ,v j ) of the regional network diagram is:

[0123]

[0124] where represents the weight of edge e k , SD is the variance of the corresponding node, PCC represents the Pearson coefficient between two variables, δ represents the absolute value of the variance change of the node at time t compared with the previous moment, LNE represents the local network entropy of the corresponding node, the network entropy is constructed based on the DNM principle, and L represents the number of neighbors of the regional node.

[0125] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made by those skilled in the art within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An infectious disease early warning and prediction method based on a regional network, characterized in that It includes the following steps carried out sequentially: Step S1: Construct a dynamically evolving regional network diagram using real-time population flow data based on the adjacency information of geographical locations; Step S2: Process the infectious disease sequence data sequentially, starting from time t = 1 until the entire sequence is processed; Step S3: Use the network entropy method to calculate the weight of each edge e k =(v i , v j ) of the regional network graph. After calculation, a weighted undirected graph is obtained; Step S4: According to Prim's algorithm, for the weighted undirected graph corresponding to time t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ; Step S5: Observe L t to see if it is greater than a preset threshold h1, and at the same time monitor the cumulative deviation of the L t time series to see if it is greater than a preset threshold h2. If L t or the cumulative deviation is greater than the preset threshold, then the current moment t is the early warning time point; otherwise, return to execute Step S2.

2. The method for infectious disease early warning and prediction based on a regional network according to claim 1, wherein: In the step S3, the network entropy method is used to calculate the weight of each edge e k =(v i , v j ) of the regional network diagram, and the specific calculation formula is as follows: Among them, represents the weight of edge e k The weight of, SD is the variance of the corresponding node, PCC represents the Pearson coefficient between two variables, δ represents the absolute value of the variance change of the regional node at time t compared with the previous time, LNE represents the local network entropy of the corresponding node, and L represents the number of neighbors of the regional node.

3. A method for early warning and prediction of infectious diseases based on a regional network according to claim 1, characterized in that: In the step S5, observe L t The method for determining whether it is greater than the preset threshold h1 is as follows. Using the Logistic algorithm, all the minimum weights L are mapped to the range between 0 and 1 through the S-shaped function t The specific formula is as follows: where ω is the weight obtained through data training; Observe whether P(y t = 1|L t ; ω) is greater than a preset threshold h1.

4. The method for early warning and prediction of infectious diseases based on a regional network according to claim 3, wherein: The preset threshold h1 is set to 0.

5.

5. A method for early warning and prediction of infectious diseases based on a regional network according to claim 1, characterized in that, In step S5, monitor L t The method for determining whether the cumulative deviation of the time series is greater than a preset threshold h2 is specifically as follows: L based on the historical non-warning period t Calculate the mean μ0 of the sequence, and calculate the standard deviation σ of the historical L t sequence, and set the sensitivity parameter σ0 = Nσ, where N is the set minimum amplitude constant; Step S501: Calculate the normalized deviation S for the minimum weight L at each moment t t , where S t is defined as follows: Step S502: Calculate the minimum weight L t of the forward cumulative sum Step S503: If then it is determined that the time t is a mutation point. If the time t is a mutation point, reset the forward cumulative sum 6. The method for infectious disease early warning and prediction based on a regional network according to claim 1, characterized in that: In the said Step S1, each node represents a region, and the connecting edges between the nodes represent the adjacency relationship between regions.

7. A method for infectious disease early warning and prediction based on a regional network according to claim 1, characterized in that: The said Step S2 is specifically: Process the infectious disease sequence data sequentially in the way of a sliding window. Each time t corresponds to a sliding window, starting from time t = 1 until the entire sequence is processed.

8. An infectious disease early warning and prediction system based on a regional network, based on the early warning and prediction method described in any one of claims 1-7, characterized in that, It includes: Regional Link Module: Construct a corresponding regional network diagram according to the adjacency information of geographical locations; Processing generation module: sequentially process the infectious disease sequence data, starting from time t = 1 until the entire sequence is processed; use the network entropy method to calculate the weight of each edge e k =(v i ,v j ) of the regional network graph, and obtain a weighted undirected graph after calculation; according to the Prim algorithm, for the weighted undirected graph corresponding to time t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ; use the Prim algorithm, for the weighted undirected graph corresponding to time t, give a minimum spanning tree and the corresponding minimum weight, and the minimum weight is denoted as L t ; Early warning recognition module: Using the Logistic algorithm, all the minimum weights L are mapped between 0 and 1 through the S-shaped function. The specific formula is as follows: t The mapping is between 0 and 1, and the specific formula is: where ω is the weight obtained through data training; In addition, the early warning recognition module also uses the CUSUM control chart monitoring method to monitor the mutation points of the L t sequence, calculates the mean value μ0 based on the L t sequence during the historical non-early warning period, calculates the standard deviation σ of the historical L t sequence, and sets the sensitivity parameter σ0 = Nσ, where N is the set minimum amplitude constant The minimum weight L at each moment t t Calculate the standardized deviation S t : Calculate the minimum weight L t Forward cumulative sum of If is greater than the preset threshold h2 or P(y t = 1|L t ; ω) is greater than the preset threshold h1, then it is determined that the current time t is the warning time point; otherwise, the current time t is not the warning time point.

9. The early warning and prediction system for infectious diseases based on a regional network according to claim 8, characterized in that: In the processing and generating module, the network entropy method is used to calculate the weight of each edge e of the regional network graph k =(v i , v j ) and the specific calculation formula of the weight is as follows: Among them, represents the weight of edge e k SD is the variance of the corresponding node, PCC represents the Pearson coefficient between two variables, δ represents the absolute value of the variance change of the node at time t compared with the previous time, LNE represents the local network entropy of the corresponding node, and L represents the number of neighbors of the regional node.