Intelligent epidemic disease prediction system based on deep learning
By using deep learning technology to integrate multi-source data to analyze the dynamics of epidemic transmission, identify high-risk areas and optimize the allocation of medical resources, the problems of single data and lagging resources in traditional methods are solved, and efficient and accurate prediction and resource utilization are achieved in epidemic prevention and control.
Patent Information
- Application Number
- CN202511262227.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Traditional epidemic prediction methods rely on a single data source and find it difficult to capture the spatiotemporal dynamic characteristics of epidemic transmission, resulting in insufficient accuracy and timeliness of prediction results. In addition, the allocation of medical resources lacks systematization and intelligence, which makes it easy to waste resources and delays.
An intelligent epidemic prediction system based on deep learning is used to obtain environmental meteorological, population mobility and medical resource distribution data through a multivariate data acquisition module. The spatiotemporal graph neural network and long-short-term memory network are used to analyze the dynamic characteristics of regional transmission, identify high-risk areas, calculate the coverage of medical resources, and generate prevention and control strategies.
It has achieved accurate prediction of epidemic transmission trends and efficient resource allocation, improved the targetedness and efficiency of prevention and control work, avoided waste of resources, and enhanced public health safety.
Smart Images

Figure CN120767004A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of epidemic prediction, in particular to an intelligent epidemic prediction system based on deep learning. BACKGROUND
[0002] In the global public health security system, the outbreak and spread of epidemics will have a great impact on social order, economic development and public health. Timely and accurate prediction of epidemic transmission trends, identification of high-risk areas and rational allocation of medical resources are important links in prevention and control. Traditional epidemic prediction methods rely on epidemiological survey data and simple statistical models, which have obvious limitations. From the perspective of data utilization, traditional methods often only focus on historical epidemic transmission data, ignoring the influence of environmental meteorological data, population flow data and medical resource distribution data on epidemic transmission. In fact, environmental meteorological factors such as temperature, humidity and precipitation directly affect the survival, reproduction and activity of the transmission medium of the pathogen. For example, some viruses are more likely to survive and spread in low-temperature and humid environments. Population flow data reflects the cross-regional movement of personnel and is a key driving factor for the spread of epidemics across regions. In particular, during holidays and other periods of large-scale population flow, the risk of epidemic transmission will increase significantly. Medical resource distribution data is closely related to epidemic prevention and control capabilities. The number, type and distribution density of medical resources in a region will affect the speed and effectiveness of epidemic control. In terms of prediction models, traditional statistical models are difficult to capture the spatio-temporal dynamic characteristics of epidemic transmission. Epidemic transmission has obvious spatio-temporal correlation, and the transmission situation of different regions at different time points is different, and the transmission between regions is mutually influenced. Traditional models cannot effectively integrate multi-source heterogeneous data, and it is also difficult to accurately describe this complex spatio-temporal dynamic relationship, resulting in insufficient accuracy and timeliness of the prediction results, which often cannot provide timely and effective reference for prevention and control decisions. In terms of risk area identification and resource allocation, traditional methods often use manual analysis and experience-based judgment, lacking systematic and intelligent technical support. Manual analysis is not only inefficient, but also difficult to dynamically adjust the risk level division according to real-time changing epidemic data, which may lead to lag in identifying high-risk areas. At the same time, in the process of matching medical resources, it is often difficult to quickly and accurately calculate the resource coverage range of each medical institution, resulting in over-concentration of medical resources in some areas and shortage of resources in high-risk areas, affecting the orderly development of epidemic prevention and control work. With the frequent occurrence of epidemics around the world in recent years, traditional methods have been difficult to meet the actual needs of epidemic prevention and control under current complex situations, and there is an urgent need for a technical solution that can integrate multi-source data, have accurate prediction ability, intelligently identify risk areas and reasonably match resources. SUMMARY
[0003] The purpose of the present invention is to provide an intelligent epidemic prediction system based on deep learning to solve the problems raised in the above background technology.
[0004] To achieve the above objectives, the present invention provides an epidemic intelligent prediction system based on deep learning, the system comprising: Multivariate data collection module, used to obtain real-time epidemic-related environmental and meteorological data, population mobility data, medical resource distribution data, and historical epidemic spread data; A deep learning prediction module is used to analyze the dynamic characteristics of regional transmission through the collaborative architecture of spatiotemporal graph neural network and long short-term memory network based on the environmental meteorological data, population mobility data, medical resource distribution data and historical epidemic spread data; The risk area identification module is used to construct a regional risk level heat map based on the dynamic characteristics of regional transmission and identify high-risk areas with risk levels exceeding a preset threshold. The risk area identification module constructs a regional risk level heat map based on the output embedding vector of the deep learning prediction module. The process includes: inputting the embedding vector into a classification model, which uses a support vector machine to divide the risk level into three categories: low, medium, and high; the heat map is drawn through a geographic information system, and each grid cell corresponds to a risk probability; the preset threshold is set to be above the median value as high risk, and the system automatically identifies high-risk areas with risk levels exceeding the threshold and outputs them as raster data; A resource matching analysis module is used to calculate the resource coverage of each medical institution based on the medical resource distribution data and the spatial location of the high-risk area; When the resource matching analysis module calculates the resource coverage of each medical institution, it specifically performs the following operations: The first step is to obtain the real-time resource capacity parameters of each medical institution, including the number of medical staff on duty, the number of available beds, and the amount of emergency equipment in stock, and standardize these parameters into a resource adequacy index in the range of 0-1; The second step is to combine population density data and transportation accessibility data of high-risk areas to establish a basic coverage radius model; The third step is to dynamically adjust the coverage radius based on the resource adequacy index: when the resource adequacy index is ≥ 0.8, the coverage radius is expanded by 20% based on the basic model results; when 0.5 ≤ resource adequacy index < 0.8, the coverage radius of the basic model is maintained; when the resource adequacy index is < 0.5, the coverage radius is reduced by 30% based on the basic model results; The fourth step is to perform spatial boundary verification on the adjusted coverage radius, exclude natural geographical obstacles such as rivers and mountains and traffic control areas, and finally determine the actual resource coverage of each medical institution, and output it as polygonal vector data with geographic coordinates.
[0005] Preferably, the resource matching analysis module includes: Perform spatial overlay analysis on the resource coverage of each medical institution and the high-risk area to generate the overlapping area and location of the resource coverage and high-risk area; Calculate the distance between the medical institution and the overlapping area based on the location data of the medical institution and the location of the overlapping area; Based on the overlapping area, interval distance and real-time load data of medical resources, the prevention and control potential coefficient of each medical institution is calculated.
[0006] Preferably, the deep learning prediction module performs: The temperature and humidity parameters in the environmental meteorological data, the cross-regional migration intensity in the population flow data, and the bed turnover rate in the medical resource distribution data are used as input features; The spatial correlation of multi-source data is processed through spatiotemporal graph neural networks, providing an accurate spatial feature foundation for long-term and short-term memory networks; Extracting spatiotemporal propagation feature vectors through long short-term memory networks; The propagation rate and mutation risk probability within a future specified time window are predicted based on the spatiotemporal propagation feature vector.
[0007] Preferably, the system further comprises: A prevention and control strategy generation module is used to integrate the prevention and control potential coefficient of each medical institution with the transmission rate and mutation risk probability; Output the prevention and control matching score of each area through the multi-layer perceptron network; A prevention and control resource scheduling priority list is generated based on the prevention and control matching score.
[0008] Preferably, the system further comprises: A dynamic decision-making module for scheduling prevention and control resources based on the priority list and the real-time epidemic severity index; When the real-time epidemic severity index exceeds the dynamically adjusted threshold, the graded response mechanism is activated; Based on the matching result of the control and prevention matching score and the preset response rule, a control and prevention parameter adjustment instruction is output.
[0009] Preferably, the control parameter adjustment instruction includes: Adjust the detection point density parameters according to the rate of change of the high-risk area; Adjust the isolation range parameters according to the ratio of transmission rate to medical resource load; Adjust vaccine allocation weight parameters according to the probability of mutation risk.
[0010] Preferably, the system further comprises: The feedback module is configured to collect actual transmission attenuation rate, resource usage deviation value and new case distribution data after implementation of the prevention and control measures; The actual transmission attenuation rate is compared with the predicted transmission rate to generate a first feedback coefficient; The resource usage deviation value is analyzed in correlation with the prevention and control matching degree score to generate a second feedback coefficient.
[0011] Preferably, the system further comprises: The model optimization module is configured to calculate spatial error between the new case distribution data and the predicted high-risk area; The first feedback coefficient and the second feedback coefficient are fused to construct a loss function; The weight parameters of the spatio-temporal graph neural network are dynamically updated through a back propagation algorithm.
[0012] Preferably, the system further comprises: The iterative early warning module is configured to regenerate regional transmission dynamic characteristics according to the updated deep learning prediction module; When the regenerated mutation risk probability exceeds a historical peak value, a cross-regional collaborative early warning protocol is triggered; A prevention and control resource scheduling priority list is updated based on the recalculated prevention and control matching degree score.
[0013] Preferably, according to the deep learning-based intelligent epidemic prediction system, the system is deployed on a distributed computing platform and comprises a spatio-temporal database configured to store output of the multi-element data collection module; The deep learning prediction module accelerates matrix operation of the spatio-temporal graph neural network through a graphics processing unit; Output instructions of the dynamic decision module are synchronized to a public health response terminal through an Internet of Things interface.
[0014] Compared with the prior art, the present application has the following advantages: The multivariate data acquisition module can acquire epidemic-related environmental and meteorological data, population mobility data, medical resource distribution data, and historical epidemic spread data in real time, breaking the limitation of traditional prediction methods that rely on a single data source. This module realizes the comprehensive collection and real-time updating of multi-source heterogeneous data, allowing the system to fully consider various key factors affecting the spread of epidemics when conducting predictive analysis. The introduction of environmental and meteorological data allows the system to accurately grasp the environmental conditions for the survival and spread of pathogens, thereby more accurately judging the transmission risks in different environments; the real-time acquisition of population mobility data can promptly reflect the trend of personnel mobility and provide a data basis for capturing the dynamics of cross-regional epidemic spread; the inclusion of medical resource distribution data provides data support for the system in the subsequent resource matching link, avoiding irrational resource allocation due to missing information; and historical epidemic spread data provides a basis for model training and summarizing the laws of transmission, helping to improve the reliability of predictions. The deep learning prediction module uses spatiotemporal graph neural networks to analyze the dynamic characteristics of regional transmission. Compared to traditional statistical models, it offers significant advantages in addressing the spatiotemporal correlations of epidemic spread. Spatiotemporal graph neural networks can effectively integrate multi-source data and, by constructing spatiotemporal correlation models, accurately depict the transmission dynamics of different regions at different points in time and the mutual influence between regions. This model architecture fully exploits the hidden spatiotemporal information in the data, capturing the dynamic patterns of epidemic spread, thereby improving the accuracy and timeliness of transmission trend predictions. Through this module's analysis, it is possible to predict in advance the scope, speed, and number of infections of an epidemic over a period of time. This provides scientific guidance for epidemic prevention and control departments to formulate prevention and control strategies and deploy prevention and control personnel in advance, and facilitates timely intervention measures in the early stages of an epidemic to delay or block its spread. The risk area identification module constructs a regional risk level heat map based on the dynamic characteristics of regional transmission and identifies high-risk areas, realizing intelligent and visual risk area identification. The heat map format can intuitively present the risk level distribution of different regions, making it easier for prevention and control personnel to quickly grasp the overall epidemic risk situation. At the same time, the module conducts risk level assessments based on real-time updated transmission dynamic characteristics, which can promptly identify areas with rising risk levels and avoid the problem of lagging identification of high-risk areas. By accurately identifying high-risk areas, prevention and control resources can be concentrated on the areas with the highest risks, improving the pertinence and efficiency of prevention and control work, reducing unnecessary waste of resources, and maximizing the effectiveness of limited prevention and control forces. The Resource Matching Analysis module calculates the resource coverage of each medical institution based on medical resource distribution data and the spatial location of high-risk areas, providing intelligent support for the rational allocation of medical resources. This module's resource coverage calculations are not simply based on geographic distance, but rather dynamically adjust multiple parameters. First, the module collects the real-time resource capacity of medical institutions (such as staffing, available beds, and emergency equipment) and converts it into a resource adequacy index. The module also considers the population density and accessibility of high-risk areas to determine a baseline coverage radius. This baseline radius is then adjusted based on the resource adequacy index. For example, resource-rich institutions can expand their service areas, while resource-constrained institutions can reduce theirs. Finally, the module eliminates natural obstacles and traffic control areas to ensure that the calculated coverage area aligns with actual service capacity and geographic accessibility, avoiding biased resource coverage assessments caused by a single geographic distance calculation. This module quickly and accurately analyzes the area each medical institution can cover and the availability of medical resources around different high-risk areas. Based on these analysis results, prevention and control departments can clearly understand which high-risk areas have sufficient medical resources and which areas have resource gaps. They can then formulate resource allocation plans based on this information, allocating medical resources from areas with sufficient resources to areas with resource gaps. This resource matching method avoids the subjectivity and lag of manual allocation, ensuring that medical resources can be quickly and accurately allocated to the areas most in need, ensuring that patients in high-risk areas can receive medical treatment in a timely manner. It also avoids the waste caused by excessive accumulation of medical resources in certain areas, improves the overall efficiency of medical resource utilization, and provides strong support for medical treatment work during epidemic prevention and control. The various modules of the entire system work together to form a complete technical chain from data collection, trend forecasting, risk identification, to resource matching. The multivariate data collection module provides comprehensive, real-time data support for subsequent modules. The analysis results of the deep learning prediction module provide a basis for risk area identification. The risk area identification results in turn guide the resource matching analysis module to carry out accurate resource coverage calculations. The mutual cooperation of these modules makes the entire system have efficient and accurate epidemic prevention and control assistance capabilities. In practical applications, this system can be widely used in disease control centers, health departments, and other institutions at all levels to provide comprehensive technical support for epidemic prevention and control decision-making, help improve the overall level of epidemic prevention and control, and better protect public health safety and public health. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a time sequence diagram of the deep learning-based intelligent epidemic prediction system of the present invention; Figure 2 A detailed workflow diagram for the resource matching analysis module; Figure 3 Generate a workflow diagram for the prevention and control strategy module; Figure 4 This is the workflow diagram of the dynamic decision-making module. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0017] See also Figure 1 The present invention provides an intelligent epidemic prediction system based on deep learning, which includes a multivariate data acquisition module, a deep learning prediction module, a risk area identification module and a resource matching analysis module.
[0018] The multivariate data acquisition module acquires real-time epidemic-related environmental and meteorological data, population mobility data, medical resource distribution data, and historical epidemic spread data. This data is integrated through heterogeneous data interfaces. Environmental and meteorological data is sourced from the National Meteorological Center's real-time sensor network, including parameters such as temperature, humidity, and wind speed. Population mobility data is collected through mobile operator signaling platforms and traffic monitoring systems, recording the intensity of cross-regional migration. Medical resource distribution data is extracted from health management department databases, covering hospital bed numbers, equipment distribution, and staffing. Historical epidemic spread data, including the distribution of previous cases and transmission cycles, is downloaded from the disease control agency's historical records. The multivariate data acquisition module cleans and standardizes data, employing timestamp alignment technology to ensure temporal consistency, and outputs structured datasets to the deep learning prediction module.
[0019] The deep learning prediction module receives the output of the multivariate data acquisition module and analyzes the dynamic characteristics of regional transmission using a spatiotemporal graph neural network. The spatiotemporal graph neural network architecture consists of an input layer, a graph convolutional layer, and an output layer. The graph convolutional layer constructs a spatial adjacency matrix based on geographic partitions to account for interactions between regions. Input features include multidimensional vectors of environmental and meteorological data, population mobility data, medical resource distribution data, and historical epidemic transmission data. The graph convolutional layer applies an attention mechanism to dynamically weight the influence of neighboring nodes, and the output layer generates an embedding vector representing the dynamic characteristics of regional transmission. The deep learning prediction module is deployed on a distributed cluster, achieving efficient computation through batch processing.
[0020] The risk area identification module constructs a regional risk level heat map based on the embedded vectors output by the deep learning prediction module. This process involves: inputting the embedded vectors into a classification model, which uses a support vector machine to classify risk levels into three categories: low, medium, and high. The heat map is then created using a geographic information system, with each grid cell corresponding to a risk probability. A preset threshold is set above the median value, indicating high risk. The system automatically identifies high-risk areas with risk levels exceeding this threshold and outputs them as raster data.
[0021] The resource matching analysis module uses medical resource distribution data and the spatial location of high-risk areas to calculate the resource coverage of each medical institution. The coverage area is defined as the maximum radius of the area that can be served within the medical institution's buffer zone. The buffer zone radius is dynamically adjusted based on the medical resource capacity formula, which includes bed density and staff density. This module executes a spatial analysis algorithm, combining regional population density and transportation network to calculate the coverage radius, and outputs polygonal boundary data. The following further details the extended implementation of the system through multiple examples.
[0022] Example 1: See Figure 2 The resource matching analysis module executes based on medical resource distribution data and the spatial location data of high-risk areas. Medical resource distribution data is stored in a vector polygon format, representing the resource coverage of each medical institution. Coverage is defined as the contiguous geographic area within a specific service radius where the institution can effectively provide medical support. High-risk area data is derived from the raster heat map output by the risk area identification module and converted into a polygon layer with clear boundaries after binarization. The system launches a spatial overlay analysis program, invoking the spatial operation interface of the GIS engine. This program loads the medical institution coverage layer and the high-risk area layer into the in-memory workspace and applies spatial indexing techniques to accelerate data processing. The overlay analysis performs a polygon intersection operation, determining the positional relationship between all geometric objects in the two layers. The operation iterates through the set of high-risk areas for each institution and calculates the geometric intersection area between the coverage polygon and the high-risk area polygon. This intersection area is calculated in real time by parsing the coordinates of the polygon vertices using a numerical integration algorithm. The geographic identifier of each high-risk area that results in the intersection is recorded, and the coordinates of the center point of its bounding rectangle are extracted. The calculation process dynamically generates a result dataset, which contains the medical institution ID, high-risk area ID, intersection area, and intersection area center point coordinate fields.
[0023] After obtaining the intersection area, the system calculates the distance between the medical institution and the corresponding high-risk area. The reference point location of the medical institution extracts the longitude and latitude coordinates from the medical resource distribution database, which is usually set as the location point of the institution's main building. The interval distance is calculated using the geodetic formula, and based on the spatial reference coordinate system conversion, the geodesic distance between the coordinate point of the medical institution and the center point of the intersection of the aforementioned high-risk area is calculated. The calculation engine automatically handles the projection transformation problem of the coordinate system to ensure the accuracy of the distance on the earth's curved surface. For traffic factor optimization, the system is connected to the real-time road network database. When special terrain or traffic control is detected between the high-risk area and the medical institution, the path planning algorithm is called to output the actual travel distance instead of the straight-line distance. The calculation result is appended to the result data set in the form of a distance value.
[0024] After spatial overlay analysis and distance calculation, the system integrates real-time medical resource load data. Real-time load data is updated every ten minutes via an application programming interface (API) and includes dynamic indicators such as remaining hospital beds, ventilator utilization, and the number of medical staff on duty. The load data processing module normalizes the raw indicators to generate load pressure values ranging from 0 to 1. The synthesis of the prevention and control potential coefficient utilizes a multidimensional decision-making model, integrating three core parameters: intersection area, separation distance, and load pressure value. The model architecture uses intersection area as a positive factor, separation distance as a negative factor, and load pressure value as a modulating factor. The data processing process first normalizes area values to a uniform dimension and performs a reciprocal conversion of distance values to conform to the law of distance attenuation. The three sets of parameters are combined using preset weights to calculate an intermediate value, with the weights dynamically calibrated and verified through historical events. The coefficient value is calculated using scalar arithmetic rules, retaining three decimal places of precision. The calculation process implements an outlier detection mechanism, triggering a data review process when a parameter exceeds a reasonable threshold. The final output data table, containing the institution ID, potential coefficient value, and a timestamp, is written to a central database for subsequent module access.
[0025] The technical implementations involved in the data flow process include an R-tree index structure used in the geometric operation engine to accelerate polygon queries; multi-threaded parallel processing for distance calculations; and real-time data access to message queues. Spatial overlay analysis is performed in blocks based on administrative divisions to avoid bottlenecks in processing extremely large datasets. Each medical institution's overlay analysis task is encapsulated as an independent computational unit, with compute nodes allocated by a distributed task scheduler. The interval distance calculation module has a built-in caching mechanism to locally store and reuse high-risk area locations that are repeatedly calculated. The entire process triggers a full computation every 60 minutes, but an incremental update is immediately initiated if a boundary change of more than 10% is detected in a high-risk area. Computational results are output in a standardized JSON format, preserving complete spatial topology and parameter metadata. Spatial fields are used in the database table structure to store geometric data, supporting dynamic geographic queries and visualization. Version information is recorded for all intermediate data during processing to meet audit and backtracking requirements. The system implements an automatic retry mechanism for computational failures, and repeated failures trigger manual intervention.
[0026] Example 2: See Figure 3 The data processing flow of the deep learning prediction module starts with the input feature extraction stage. The temperature and humidity parameters in the environmental meteorological data are input in the form of floating-point numbers, and the parameters are updated every sixty minutes. The cross-regional migration intensity index in the population flow data is derived from the location signaling of mobile devices. This indicator is generated by the base station switching frequency statistics, and the value represents the proportion of the migrating population per unit time. The bed turnover rate in the medical resource distribution data comes from the hospital information system and is calculated as the ratio of the number of discharged patients per day to the number of available beds. The three feature parameters are standardized after being received by the input interface. The processing rule uses the historical mean and standard deviation of each parameter for z-score conversion. The converted feature tensor forms a fixed-dimensional matrix structure, with the row dimension corresponding to the time step and the column dimension corresponding to the feature category. The time step is configured as a seven-day period.
[0027] This matrix is input into a long-short-term memory network for feature extraction. The network structure consists of two hidden layers, each containing 128 neurons. Network operations utilize a gating mechanism to handle temporal dependencies: a forget gate controls the proportion of historical information retained; an input gate selects important features at the current moment; and an output gate modulates the output strength of the feature vector. Connections between network units utilize a cyclic computational logic, where the hidden state from the previous moment and the current input jointly participate in the gating decision at the current moment. Through iterative computations over multiple time steps, the network captures the dynamic patterns of regional transmission from time series data. The hidden layer output is passed to a fully connected layer for dimensionality compression, generating a fixed-length spatiotemporal transmission feature vector. The feature vector's dimensionality matches the number of geographic regions, with each element encoding the dynamic state of a specific region.
[0028] Based on this spatiotemporal propagation feature vector, the module performs two parallel prediction tasks. The propagation rate prediction branch uses a linear regression layer to output a continuous value that quantifies the intensity of the epidemic spread per unit time within a seven-day window. The mutation risk probability prediction branch applies a sigmoid activation function to transform the output value, mapping the regression result into a probability estimate between 0 and 1. The time window is fixed to a seven-day period, and the prediction engine performs a full calculation every 24 hours. Prediction results are stored by region ID, and the output data structure contains the predicted value, timestamp, and confidence interval.
[0029] The input end of the prevention and control strategy generation module loads three data sources: the data table of transmission rate and mutation risk probability output by the deep learning prediction module, and the list of prevention and control potential coefficients generated by the resource matching analysis module. The data fusion process starts the feature alignment mechanism and matches the relevant parameters according to the regional ID. The aligned parameter group is normalized and scaled to eliminate the dimensional difference to form a joint feature vector. This vector is input into the multi-layer perceptron network for pattern analysis. The network structure contains three fully connected layers. The input layer receives a five-dimensional vector (regional transmission rate, mutation risk probability, and prevention and control potential coefficients associated with the three medical institutions); the first hidden layer is set with 64 neurons, and the activation function is ReLU; the second hidden layer is set with 32 neurons, and ReLU activation is also used; the output layer is designed with a single neuron structure to process the score mapping. The weight parameters of each layer are initialized through historical operation data. The prevention and control matching score calculation follows the following mapping relationship: ; Where: represents the control matching score, is the output layer weight parameter vector, is the joint eigenvector, Represents the sigmoid function. The output score is thresholded, and values outside the 0-1 range are forced to the boundary value. The scoring results are stored in a cache database.
[0030] The logic for generating a priority list for scheduling prevention and control resources is based on matching score sorting. The sorting algorithm uses a maximum heap structure to implement a priority queue, and the queue node contains the region ID and matching score value. During the heap sorting process, the score difference between the new input data and the top element of the heap is dynamically compared, and when the new score is higher than the top element of the heap, the structure adjustment is triggered. The final output list entries are sorted in descending order by score, and each entry is attached with a time validity flag. The list update cycle is synchronized with the prediction module for twenty-four hours, but it is immediately recalculated when a new alarm is received in a high-risk area. The queue structure adopts a persistent storage design, and each update records the operation log for reference. The data output format is compatible with the public health database, and the fields include area code, medical resource type weight allocation recommendation, and scheduling emergency level label.
[0031] Containerized deployment is used at the system implementation level. The long-short-term memory network model publishes an API interface through the TensorFlowServing framework; the multi-layer perceptron model is loaded into the in-memory database to accelerate real-time reasoning. Apache Kafka message queues are used for data transmission between modules, and an independent control channel is set up to transmit score update instructions. The cache mechanism uses the Redis database to temporarily store intermediate feature vectors, effectively reducing the frequency of repeated model calculations. The distributed task scheduler monitors the load of computing nodes, and abnormal timed-out tasks are automatically transferred to backup nodes. The log system fully records the comparison between the predicted value and the actual number of cases, but does not involve the conclusion of the effect evaluation. Model version control integrates Git repository management, and a snapshot of the network weight is submitted with each update. The hardware layer connects to the GPU accelerator through the PCIe channel to optimize the efficiency of matrix multiplication operations.
[0032] Example 3: See Figure 4 The operating mechanism of the dynamic decision-making module starts with dual data source input: the priority list for scheduling prevention and control resources and the real-time epidemic severity index. The priority list is stored in the form of a sorted queue, containing regional identifiers and corresponding prevention and control matching scores. The real-time epidemic severity index is obtained through the multivariate data acquisition module. The calculation logic integrates three dimensions: the current growth rate of confirmed cases, the proportion of severe cases, and the regional transmission coefficient. The index calculation adopts a weighted summation model, and the weight distribution is based on the importance of epidemiological parameters. The calculation results are normalized to floating-point numbers between 0 and 1, and the data snapshot is updated every 30 minutes.
[0033] The module's core logic includes a threshold detection mechanism. The system presets dynamically adjusted thresholds in the configuration library, which are set to different values based on historical epidemic stage classifications. The detection program continuously compares the real-time epidemic severity index with the currently effective threshold. When the index exceeds the threshold for two consecutive cycles, a state transition is triggered. This state transition activates a graded response mechanism, which defines three response modes: primary, intermediate, and advanced. Mode switching is determined by the magnitude of the index exceeding the threshold: primary mode is activated when the threshold exceeds 10%, intermediate mode is activated when the threshold exceeds 10% to 25%, and advanced mode is activated when the threshold exceeds 25%. The mode selection result is written to the system status register.
[0034] Once the response mode is determined, the module performs a match analysis between the control match score and the preset response rules. These rules are stored in a relational database rule table. Each rule contains four fields: the response mode field defines the scope of the rule; the score interval field defines the threshold for the match score; the response action field encodes the specific action instructions; and the priority field resolves rule conflicts. The rule engine loads the rule subset corresponding to the current response mode and scans each rule according to the match score interval. The matching process uses an interval tree data structure to accelerate queries and generate a rule hit mark for each area.
[0035] The generation of control parameter adjustment instructions is based on the rule matching results. The instruction generator parses the response action field of the matching rule and dynamically synthesizes executable control commands. The instruction set includes three types of core parameter adjustments: The density of detection points is adjusted based on the rate of change in the high-risk area. This rate of change is calculated by comparing the heat maps of the high-risk areas in the current and previous periods, using the grid difference statistics method to determine the percentage change. The adjustment logic follows a linear response model: ; Where: is the density of detection points after adjustment (unit: points / square kilometer), is the base density configuration value, is the percentage of area change, is the density adjustment coefficient (default value is 100). The calculation result is rounded to generate the density parameter update instruction.
[0036] Adjustment of the quarantine radius parameters depends on the ratio of the transmission rate to the healthcare resource load. The transmission rate is predicted by the deep learning prediction module, while the healthcare resource load is based on the real-time bed occupancy rate. The ratio is calculated as the transmission rate divided by the load factor, and the result is input into a piecewise function converter. The converter has preset ratio thresholds: below 1.0, the quarantine radius remains at the baseline; between 1.0 and 1.5, the radius increases linearly; above 1.5, the maximum quarantine radius is activated. The output command contains a geofence coordinate update packet.
[0037] Vaccine allocation weighting parameters are adjusted based on variant risk probability. Probability values are mapped to a weighting table, which defines the relationship between probability intervals and regional priorities. Probability values below 0.3 are assigned standard weights; those between 0.3 and 0.6 receive a 50% increase in weight; and those above 0.6 receive double the weight. The adjuster outputs a list of vaccine allocation weight coefficients for each region, with coefficient values retained to two decimal places.
[0038] The command transmission mechanism utilizes a lightweight message protocol encapsulation. Each command contains a four-tuple structure: timestamp, region code, parameter type, and new parameter value. The IoT interface establishes a dedicated message topic channel to transmit command packets by region. A message confirmation and retransmission mechanism is implemented during transmission, and the public health response terminal returns an execution status code upon receipt. The system implements a command lifecycle manager that retries unconfirmed commands three times. If no response is received after the timeout, the command is marked as abnormal.
[0039] The module's fault-tolerant design includes a rule conflict detection algorithm. When multiple rules match the same region, the algorithm selects the highest-level command based on the rule priority field; if priorities are equal, the most recently updated command is used. During the state machine monitoring module's operation, abnormal events trigger rollbacks: if command generation fails, the last valid configuration is restored; if a network interruption occurs, commands are cached in a local queue. Historical command versions are stored in a blockchain repository, supporting traceability and auditing of parameter adjustments. A real-time monitoring panel visualizes the currently effective parameters and regional coverage status, assisting in manual decision-making and intervention.
[0040] At the system integration level, a service bus connects upstream and downstream modules. Priority list changes trigger subscription updates, and epidemic index violations activate asynchronous processing threads. At the hardware level, FPGAs are used to accelerate rule matching operations, and hot reloading of rule tables supports dynamic policy adjustments. The logging system records the entire decision-making chain: every state change, from threshold detection and rule matching to instruction generation, generates a timestamped log entry.
[0041] Example 4: The operational process of executing the feedback module takes the administrative division of a certain city as an example. The module starts the data collection task and retrieves the three core indicators after the implementation of prevention and control measures from the public health monitoring system. The actual transmission attenuation rate field is derived from the daily case statistics report, and the calculation logic is the percentage change value of the new cases this week compared to last week. The resource usage deviation value field is extracted from the hospital resource management system, and the calculation method is the difference between the planned number of ventilators deployed and the actual number received. The new case distribution data is obtained from the spatiotemporal database of the municipal disease control center, which contains the latitude and longitude coordinates and the confirmed timestamp of each confirmed patient. The data collection cycle is fixed to batch processing at zero o'clock every day, covering all records of the previous 24 hours. The collection results are stored in the feedback data table. For typical data structure, see Table 1.
[0042] Table 1: Execution feedback module data collection table:
[0043] The data processing phase executes parallel branches: The first feedback coefficient generation branch loads actual transmission attenuation data and simultaneously extracts predicted transmission rate data for the same time window from the deep learning prediction module. After aligning the region identifier and time tag, the system calculates the difference between the two values. For example, if a region's attenuation rate is -15% (declining cases) and the predicted transmission rate is 0.2 (expanding trend), the difference is -0.35. This value is converted to the range [-1, 1] using a normalization function and output as the first feedback coefficient. The conversion function has built-in boundary control logic, triggering the data review process when the difference exceeds the historical maximum fluctuation range.
[0044] The second feedback coefficient generation branch obtains the resource use deviation value and the prevention and control matching degree score. The data matching is in units of regions, and the deviation value sequence and the score sequence are aligned on the time axis. The system applies a statistical correlation algorithm to calculate the correlation degree of the change of the resource use deviation value with the prevention and control matching degree score. The algorithm outputs the Pearson correlation coefficient, whose value range is [-1, 1], which is directly used as the output value of the second feedback coefficient. The calculation process excludes invalid samples, such as deviation data of periods when the matching degree score is not updated.
[0045] The model optimization module simultaneously starts spatial error analysis. The newly added case distribution data is processed by the geographic information engine: first, the case point data is converted into a density grid layer, with a resolution of 100-meter grid; then, the historical high-risk area prediction polygon generated by the risk area identification module is loaded. The system performs spatial overlay analysis, using a grid-by-grid comparison strategy: when a case high-density grid falls within the predicted high-risk area, it is recorded as a successful match, and when it falls outside the area, it is recorded as a prediction deviation. The error quantification calculation includes three indicators: prediction omission rate (the proportion of case grids outside the predicted high-risk area), prediction overcoverage rate (the proportion of the area of the predicted area without cases to the total area of the predicted high-risk area), and regional overlap accuracy (the overlapping area ratio of the predicted high-risk area and the case area). The calculation results generate a spatial error report file.
[0046] The model optimization module fuses the above results to construct a loss function. The function structure adopts a multi-source input design: the first feedback coefficient reflects the time prediction deviation weight, the second feedback coefficient embodies the resource prediction deviation weight, and the spatial error report contributes to the geographical prediction deviation weight. The loss value calculation process does not depend on mathematical formula expression, but is realized through the configuration of weight parameter table: the time prediction weight parameter is initially set to 0.5; the resource prediction weight parameter is initially set to 0.3; the spatial error weight parameter is initially set to 0.2; the module integrates the three inputs in proportion to the weights, and outputs a total loss value in the range of 0-1. This value is input into the neural network trainer, and the backpropagation algorithm is activated to update the weight parameters of the spatio-temporal graph neural network.
[0047] The training process implements a phased strategy: Freeze the input layer weight: fine-tune the spatial feature extraction layer according to the characteristics of the newly added case distribution data; Adjust the graph convolution layer: optimize the regional correlation parameters based on the spatial error report; Update the fully connected layer: calibrate the prediction output according to the time series feedback data; The training period is configured to be executed once every three days, and each iteration retains a weight checkpoint. The update mechanism implements an incremental training mode: only the latest feedback data is loaded instead of the full historical data, saving computing resources. The model weight after training is automatically deployed to the production environment, and the version number is incremented.
[0048] The exception handling mechanism includes data quality monitoring: automatically switching to a backup data source if the actual propagation attenuation rate is missing for three consecutive days; pausing coefficient calculations for that dimension if resource usage deviation exceeds a reasonable threshold; and triggering a coordinate system correction procedure when spatial analysis detects map coordinate drift. The operation log records the complete flow of each feedback collection to model update, with log entries including the original data hash value, loss weight configuration parameters, and the updated model performance baseline.
[0049] The hardware layer is configured with an independent feedback processing cluster: time series feedback processing is performed using a CPU cluster; spatial error calculations are scheduled using GPU nodes to accelerate raster operations; and model training is assigned to dedicated AI accelerator cards. The network topology adopts a split architecture: collection nodes are connected to the government-specific network, processing nodes are deployed on the internal computing network, and update nodes are connected to the model repository.
[0050] Example 5: The operation process of the iterative warning module starts with the triggering of a model update event. When the model optimization module completes the update of the weight parameters of the spatiotemporal graph neural network, the system automatically issues a model version upgrade notification. The iterative warning module receives the notification and loads the latest weight file from the model warehouse into the memory. The loading process implements a verification mechanism: the hash value of the weight file is calculated and compared with the warehouse record. If the verification fails, it falls back to the previous stable version. After successful loading, the module calls the operation interface of the deep learning prediction module and inputs the multivariate data set at the current moment to regenerate the dynamic characteristics of regional transmission. The input data includes real-time monitoring values of environmental meteorology, statistics of population mobility in the past 24 hours, snapshots of the real-time status of medical resources, and rolling window data of historical epidemic transmission. The feature generation process reuses the original network architecture, but uses the updated weight parameters to perform forward propagation calculations. The output feature vector dimension is consistent with the geographic grid division, and each vector element corresponds to the transmission situation code of a specific grid.
[0051] After regenerating the dynamic characteristics of regional transmission, the module performs threshold monitoring of the mutation risk probability. The system extracts the mutation risk probability component from the feature vector. The probability value indicates the possibility of a major mutation of the virus strain. At the same time, the historical database is queried to retrieve the probability peak records of the region over the past 365 days. The comparison logic uses a sliding window algorithm: taking seven days as the window unit, the maximum probability within the window is compared with the historical peak. When the regenerated mutation risk probability exceeds the historical peak for three consecutive window periods, the cross-regional collaborative early warning protocol is triggered. Triggering conditions include additional regional correlation verification: if there is a cross-administrative region population flow hotspot in the probability exceeding standard area, the warning range will be expanded to the three adjacent administrative regions.
[0052] The implementation of the cross-regional collaborative early warning protocol includes information encapsulation and routing distribution. The protocol defines a standardized early warning message structure: the message header contains the warning ID, release time and validity period; the message body records the list of areas exceeding the standard, the probability of exceeding the standard, and the flow intensity value of the associated area. After the message is generated, it is submitted to the distribution engine, and the engine selects the transmission path based on the pre-configured list of recipients. The provincial health department terminal receives directly through the government dedicated line; the municipal terminal relies on the health emergency communication network for transmission; and the cross-provincial collaborative area uses an encrypted Internet channel for transmission. The transmission protocol implements priority grading: real-time streaming transmission is used in the province, and batch compression transmission is enabled in the cross-provincial area. The receiving terminal returns a receipt confirmation code, and unconfirmed messages are resent every five minutes until timeout.
[0053] After the early warning is triggered, the update of the prevention and control resource scheduling priority list is started synchronously. The system inputs the regenerated regional propagation dynamic features into the prevention and control strategy generation module, which recalculates the prevention and control matching score based on the updated propagation dynamic features. The score calculation logic maintains the original multi-layer perceptron network structure, but replaces the input feature vector with the latest data. The recalculated score value is input into the sorting subsystem, which maintains the global priority queue. The update operation adopts an incremental refresh strategy: only the area where the mutation risk probability exceeds the standard and its associated areas are recalculated, and the non-affected areas retain the original score. The queue sorting algorithm applies the minimum heap adjustment technology to dynamically update the queue position of the area where the score changes. Finally, the new version of the prevention and control resource scheduling priority list is output. The list entries contain the area code, the latest score value, the version identifier and the medical resource scheduling recommendation level.
[0054] The system's overall deployment architecture is based on a distributed computing platform. The platform's physical layer comprises multiple clusters of computing nodes. The multivariate data acquisition module runs on edge computing nodes, close to the data source to reduce transmission latency. The deep learning prediction module is deployed on a GPU acceleration cluster equipped with high-speed NVLink interconnects. The risk area identification and resource matching analysis module runs on a dedicated geographic information server. The spatiotemporal database utilizes a distributed columnar storage architecture, with data sharded both by time partitions and spatial grids. The storage engine optimizes compression algorithms for time series data and establishes R-tree indexes for spatial data.
[0055] Matrix operations in the deep learning prediction module are accelerated using CUDA kernel functions. Adjacency matrix operations in the graph convolution layer are decomposed into sparse matrix multiplications, processed in parallel by the GPU's TensorCores. The time-step loop of the long short-term memory network is expanded into parallel thread blocks, using shared memory to cache intermediate states. Gradient calculations during training are optimized using the automatic differentiation engine, and backpropagation uses a mixed-precision strategy to reduce video memory usage.
[0056] The dynamic decision-making module's output command transmission relies on the IoT interface layer. This interface layer implements a protocol conversion gateway, encoding system-wide control parameter adjustment commands as MQTT protocol payloads. Topic naming adheres to regional tiering standards, such as publishing provincial-level commands to the topic / cn / province / command-type. Public health response terminals subscribe to the corresponding topics, decode the received commands, and perform parameter adjustments. Terminal execution status is transmitted back via telemetry channels, and the gateway monitors online status and redirects messages from offline terminals. The security mechanism implements two-way certificate authentication, and command payloads are encrypted using a national encryption algorithm.
[0057] Full-chain monitoring is implemented at the system operation and maintenance level. The log collector aggregates operational metrics from each module, including data collection latency, model inference time, and command transmission success rate. The monitoring dashboard displays real-time regional risk heat maps and resource scheduling paths, assisting operations personnel in identifying bottlenecks. Version releases utilize a blue-green deployment model, with traffic switching occurring after the new model version has been verified in the shadow environment.
[0058] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0059] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent epidemic prediction system based on deep learning, characterized by: Includes the following modules: Multivariate data collection module, used to obtain real-time epidemic-related environmental and meteorological data, population mobility data, medical resource distribution data, and historical epidemic spread data; A deep learning prediction module is used to analyze the dynamic characteristics of regional transmission through the collaborative architecture of spatiotemporal graph neural network and long short-term memory network based on the environmental meteorological data, population mobility data, medical resource distribution data and historical epidemic spread data; The risk area identification module is used to construct a regional risk level heat map based on the dynamic characteristics of regional transmission and identify high-risk areas with risk levels exceeding a preset threshold. The risk area identification module constructs a regional risk level heat map based on the output embedding vector of the deep learning prediction module. The process includes: inputting the embedding vector into a classification model, which uses a support vector machine to divide the risk level into three categories: low, medium, and high; the heat map is drawn through a geographic information system, and each grid cell corresponds to a risk probability; the preset threshold is set to be above the median value as high risk, and the system automatically identifies high-risk areas with risk levels exceeding the threshold and outputs them as raster data; A resource matching analysis module is used to calculate the resource coverage of each medical institution based on the medical resource distribution data and the spatial location of the high-risk area; The deep learning prediction module performs: The temperature and humidity parameters in the environmental meteorological data, the cross-regional migration intensity in the population flow data, and the bed turnover rate in the medical resource distribution data are used as input features; The spatial correlation of multi-source data is processed through spatiotemporal graph neural networks, providing an accurate spatial feature foundation for long-term and short-term memory networks; Extracting spatiotemporal propagation feature vectors through long short-term memory networks; Predicting the transmission rate and mutation risk probability within a specified future time window based on the spatiotemporal transmission feature vector; When the resource matching analysis module calculates the resource coverage of each medical institution, it specifically performs the following operations: The first step is to obtain the real-time resource capacity parameters of each medical institution, including the number of medical staff on duty, the number of available beds, and the amount of emergency equipment in stock, and standardize these parameters into a resource adequacy index in the range of 0-1; The second step is to combine population density data and transportation accessibility data of high-risk areas to establish a basic coverage radius model; The third step is to dynamically adjust the coverage radius based on the resource adequacy index: when the resource adequacy index is ≥ 0.8, the coverage radius is expanded by 20% based on the basic model results; when 0.5 ≤ resource adequacy index < 0.8, the coverage radius of the basic model is maintained; when the resource adequacy index is < 0.5, the coverage radius is reduced by 30% based on the basic model results; The fourth step is to perform spatial boundary verification on the adjusted coverage radius, exclude natural geographical obstacles such as rivers and mountains and traffic control areas, and finally determine the actual resource coverage of each medical institution, and output it as polygonal vector data with geographic coordinates.
2. The deep learning-based epidemic intelligent prediction system according to claim 1, characterized in that: The resource matching analysis module includes: Perform spatial overlay analysis on the resource coverage of each medical institution and the high-risk area to generate the overlapping area and location of the resource coverage and high-risk area; Calculate the distance between the medical institution and the overlapping area based on the location data of the medical institution and the location of the overlapping area; Based on the overlapping area, interval distance and real-time load data of medical resources, the prevention and control potential coefficient of each medical institution is calculated.
3. The deep learning-based epidemic intelligent prediction system according to claim 1, characterized in that: Also includes: A prevention and control strategy generation module is used to integrate the prevention and control potential coefficient of each medical institution with the transmission rate and mutation risk probability; Output the prevention and control matching score of each area through the multi-layer perceptron network; A prevention and control resource scheduling priority list is generated based on the prevention and control matching score.
4. The deep learning-based epidemic intelligent prediction system according to claim 3, characterized in that: Also includes: A dynamic decision-making module for scheduling prevention and control resources based on the priority list and the real-time epidemic severity index; When the real-time epidemic severity index exceeds the dynamically adjusted threshold, the graded response mechanism is activated; Based on the matching result of the control and prevention matching score and the preset response rule, a control and prevention parameter adjustment instruction is output.
5. The deep learning-based epidemic intelligent prediction system according to claim 4 is characterized in that: The control parameter adjustment instructions include: Adjust the detection point density parameters according to the rate of change of the high-risk area; Adjust the isolation range parameters according to the ratio of transmission rate to medical resource load; Adjust vaccine allocation weight parameters according to the probability of mutation risk.
6. The deep learning-based epidemic intelligent prediction system according to claim 5, characterized in that: Also includes: The execution feedback module is used to collect actual transmission attenuation rate, resource utilization deviation value and new case distribution data after the implementation of prevention and control measures; Comparing the actual propagation attenuation rate with the predicted propagation rate to generate a first feedback coefficient; The second feedback coefficient is generated by performing correlation analysis between the resource utilization deviation value and the prevention and control matching score.
7. The deep learning-based epidemic intelligent prediction system according to claim 6, characterized in that: Also includes: A model optimization module is used to calculate the spatial error between the new case distribution data and the predicted high-risk areas; The loss function is constructed by fusing the first feedback coefficient and the second feedback coefficient; The weight parameters of the spatiotemporal graph neural network are dynamically updated through the back-propagation algorithm.
8. The deep learning-based intelligent epidemic prediction system according to claim 7, characterized in that: Also includes: An iterative early warning module, used to regenerate regional transmission dynamic characteristics based on the updated deep learning prediction module; When the probability of regenerated mutation risk exceeds the historical peak, a cross-regional collaborative early warning protocol is triggered; Update the prevention and control resource scheduling priority list based on the recalculated prevention and control matching score.
9. The deep learning-based intelligent epidemic prediction system according to any one of claims 1 to 8, characterized in that: The system is deployed on a distributed computing platform and includes a spatiotemporal database for storing the output of the multivariate data acquisition module; The deep learning prediction module accelerates the matrix operations of the spatiotemporal graph neural network through a graphics processor; The output instructions of the dynamic decision-making module are synchronized to the public health response terminal through the Internet of Things interface.
Citation Information
Patent Citations
Multichannel LSTM neural network influenza epidemic situation prediction method based on attention mechanism
CN110085327A
Medical staff resource matching analysis system based on big data
CN111081358A
Early warning method and system for infectious diseases and readable storage medium
CN114141385A
Infectious disease trend prediction algorithm based on space-time hypergraph neural network
CN118609846A
Epidemic mixing data prediction method based on deep learning
CN119314696A
Cited By
Safety early warning system for underground gas pipe network
CN120997982A
Medical market division method and system based on regional behavior data
CN121707638A