Dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural network
By integrating multi-source data and constructing a multi-dimensional graph structure, a prediction and early warning system based on heterogeneous data fusion and graph neural networks is developed. This solves the problems of insufficient data fusion and inadequate early warning timeliness in existing technologies, and enables accurate and timely monitoring and control support for dengue fever outbreaks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SCIENCE & TECHNOLOGY RESEARCH CENTER OF CHINA CUSTOMS
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-23
AI Technical Summary
Existing dengue fever prediction systems are inadequate in terms of the breadth of data fusion, the depth of spatial relationships, and the timeliness of early warning. They are unable to effectively integrate multi-source heterogeneous data, construct multi-dimensional graph structures, and lack real-time early warning capabilities, resulting in insufficient support for epidemic monitoring and prevention.
A prediction and early warning system based on heterogeneous data fusion and graph neural networks is adopted. Through data acquisition module, data fusion module, graph neural network prediction module, early warning decision module and display module, it integrates case data, meteorological data, remote sensing vegetation index data, population flow data and social media sentiment data to construct multi-dimensional spatiotemporal feature data, construct dynamic multi-dimensional graph structure, perform spatiotemporal coupled prediction, and perform visualization display and early warning decision.
It has improved the precision and accuracy of dengue fever outbreak forecasting, enhanced the reliability of forecast results and the sensitivity and stability of early warning, and achieved full-process automation from data perception to decision support, thereby improving emergency response efficiency and the ability to allocate prevention and control resources.
Smart Images

Figure CN122266816A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infectious disease prediction and early warning technology, specifically a dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks. Background Technology
[0002] Currently, dengue fever prediction research mainly relies on traditional statistical models and machine learning methods, such as LASSO regression and support vector machines (SVM). These methods primarily focus on time-series prediction of the overall number of cases in a region, lacking the ability to model the complex relationships between micro-spatial units within a city, such as streets and towns. With the development of deep learning technologies such as graph neural networks, some studies have attempted to incorporate spatial structure information to improve prediction accuracy. For example, the authorized publication number "CN111554408A" describes a "dengue fever spatiotemporal prediction method, system, and electronic equipment within a city." This method improves the prediction performance of dengue fever cases at the town scale to some extent by constructing a graph structure reflecting the adjacency relationships of towns and using a graph convolutional neural network (GCN) model for joint prediction. However, this existing technology still has several limitations: First, its input features mainly rely on case data and meteorological data, namely temperature, rainfall, and static population data, failing to fully integrate the diverse and heterogeneous data widely present in the urban environment, such as social media sentiment, traffic flow, land use, and distribution of medical resources, which often contain potential signals of epidemic spread; second, the graph structure it constructs is based only on geographical adjacency relationships, without considering complex spatial dependencies such as population flow, transportation networks, and socio-economic connections, limiting the model's ability to capture the dynamics of cross-regional epidemic spread; in addition, this technology focuses on post-event prediction and is weak in real-time early warning and dynamic risk assessment, making it difficult to support prevention and control departments in rapid response and pre-deployment of resources.
[0003] Therefore, existing dengue fever prediction systems still have room for improvement in terms of the breadth of data fusion, the depth of spatial relationships, and the timeliness of early warnings. There is an urgent need for a dengue fever prediction and early warning system that can integrate multi-source heterogeneous data, construct multi-dimensional graph structures, and possess real-time early warning capabilities, in order to achieve more accurate and timely urban epidemic monitoring and prevention support. Summary of the Invention
[0004] The purpose of this invention is to provide a dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks, so as to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural network, including a data acquisition module, a data fusion module, a graph neural network prediction module, an early warning decision module, and a display module;
[0006] The data acquisition module is used to collect case data, meteorological data, remote sensing vegetation index data, population flow data, and social media sentiment data.
[0007] The data fusion module communicates with the data acquisition module and is used to clean, time-align, and process missing values of the acquired heterogeneous data to generate fused multidimensional spatiotemporal feature data.
[0008] The graph neural network prediction module communicates with the data fusion module to receive the multidimensional spatiotemporal feature data, construct a graph structure reflecting the relationship between regions with administrative division units as nodes, and input the multidimensional spatiotemporal feature data and the graph structure into the spatiotemporal coupling prediction model for training and prediction, and output the predicted value of dengue fever cases and the corresponding risk level of each node in the future specified time period.
[0009] The early warning decision module communicates with the graph neural network prediction module and is used to generate dynamic early warning thresholds based on historical case data analysis. Based on the case prediction values and real-time monitored multi-source indicators, it performs early warning judgment and level triggering through preset linkage rules.
[0010] The display module communicates with the graph neural network prediction module and the early warning decision module to visualize the predicted values, risk levels and early warning information of the cases.
[0011] Furthermore, the data fusion module includes a data assimilation unit and a quality verification unit;
[0012] The data assimilation unit is used to estimate and fill in missing values in the meteorological data and the remote sensing vegetation index data using the Kalman filter method to form a spatiotemporally continuous data sequence.
[0013] The quality verification unit is used to verify the lead-lag relationship between the social media sentiment data and the population flow data and historical case data through statistical causal analysis, and to screen out indicators with statistical leadership as effective features.
[0014] Furthermore, the graph neural network prediction module, when constructing the graph structure, is specifically used for:
[0015] Obtain multiple basic adjacency matrices that reflect the relationships between different regions, wherein the basic adjacency matrices include at least a geographical adjacency matrix and a population flow matrix;
[0016] The multiple basic adjacency matrices are weighted and combined to generate a comprehensive adjacency matrix, which serves as the weighting basis for the edges connecting nodes in the graph structure.
[0017] The population flow matrix is constructed based on the inter-regional population flow intensity data processed by the data fusion module.
[0018] Furthermore, the spatiotemporal coupling prediction model includes a spatial feature extraction unit, a temporal feature extraction unit, a cross-validation unit, and a result arbitration unit;
[0019] The spatial feature extraction unit employs a graph attention network to spatially aggregate node features based on the comprehensive adjacency matrix.
[0020] The time feature extraction unit uses a time convolutional network to model the time dependency relationship of the historical feature sequence of each node.
[0021] The cross-verification unit runs a prediction model based on a gradient boosting decision tree in parallel. The input of the gradient boosting decision tree model is the node feature data after feature engineering, which is used to generate independent prediction results.
[0022] The result arbitration unit is used to compare the master prediction result of the spatiotemporal coupled prediction model with the independent prediction result of the gradient boosting decision tree model. When the difference between the two exceeds a preset threshold, a secondary verification network is activated to arbitrate the prediction discrepancy in order to determine the final prediction value.
[0023] Furthermore, the graph attention network in the spatial feature extraction unit is used to learn dynamic attention weights between nodes, and the temporal convolutional network in the temporal feature extraction unit contains multiple dilated causal convolutional layers.
[0024] Furthermore, the early warning decision module includes a dynamic threshold generation unit and a linkage judgment unit;
[0025] The dynamic threshold generation unit uses a change point detection algorithm to analyze the time series of historical cases, identify the stages of the epidemic development, and calculate the quantile threshold of the number of cases based on a rolling time window in each stage.
[0026] The linkage judgment unit is pre-set with multi-level early warning rules. The early warning rules compare the predicted value of the case with the quantile threshold and make a joint judgment based on the real-time status of the mosquito vector monitoring index, social media sentiment index and Internet search popularity index. The triggering and cancellation of the early warning level must meet the condition of a continuous time window.
[0027] Furthermore, the system also includes a model optimization module, which communicates with the early warning decision module and the graph neural network prediction module;
[0028] The model optimization module is used to collect actual case data and corresponding system internal status data after the early warning is issued, form a feedback dataset, and use the feedback dataset to periodically fine-tune the parameters of the spatiotemporal coupling prediction model.
[0029] Furthermore, the display module provides an electronic map visualization interface for rendering the risk level on geographical partitions with different color levels, and provides a time trend chart of the predicted case values and a marker of the warning trigger time.
[0030] Furthermore, the data fusion module also includes a data version management unit, which is used to attach version identifiers to the fused multidimensional spatiotemporal feature data and intermediate processing data to achieve traceability of the data processing process.
[0031] Furthermore, the system also includes an automatic report generation module, which communicates with the early warning decision module and the display module to automatically integrate the current early warning level, risk area information, key indicator status and recommended measures to generate a structured early warning report document when an early warning is triggered.
[0032] This invention provides a dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks. It has the following beneficial effects:
[0033] This dengue fever prediction and early warning system, based on heterogeneous data fusion and graph neural networks, effectively improves the precision and accuracy of dengue fever outbreak prediction by integrating multi-source heterogeneous data and constructing a dynamic multi-dimensional graph structure for spatiotemporal coupled prediction. The system employs data assimilation and causal verification to ensure the quality of input information, utilizes graph attention networks and temporal convolutional networks to capture complex spatial dependencies and temporal dynamics, and combines parallel model cross-validation and arbitration mechanisms to enhance the reliability of prediction results. Dynamic early warning thresholds and multi-indicator linkage triggering rules further improve the sensitivity and stability of risk identification and reduce false alarms.
[0034] This dengue fever prediction and early warning system, based on heterogeneous data fusion and graph neural networks, achieves full-process automation and closed-loop optimization from data perception and intelligent analysis to decision support. A visual interactive interface provides intuitive risk assessment, version control ensures full traceability, and automatic report generation and publishing functions improve emergency response efficiency. This system provides timely, accurate, and actionable risk warning information for dengue fever prevention and control, helping to optimize the allocation of prevention and control resources and enhance the ability to respond to public health emergencies. Attached Figure Description
[0035] Figure 1 This is a data flow diagram of the dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks of the present invention;
[0036] Figure 2 This is a state diagram of the early warning decision-making logic of the dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural network of the present invention.
[0037] Figure 3 This is a graph neural network structure diagram of the dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural network of the present invention. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] Please see Figures 1 to 3 The present invention provides a technical solution: a dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural network, including a data acquisition module, a data fusion module, a graph neural network prediction module, an early warning decision module and a display module;
[0040] The data acquisition module is used to collect case data, meteorological data, remote sensing vegetation index data, population flow data, and social media sentiment data;
[0041] The data fusion module communicates with the data acquisition module and is used to clean, time-align, and handle missing values of the acquired heterogeneous data to generate fused multidimensional spatiotemporal feature data.
[0042] The graph neural network prediction module communicates with the data fusion module to receive multidimensional spatiotemporal feature data, construct a graph structure reflecting the relationship between regions with administrative division units as nodes, and input the multidimensional spatiotemporal feature data and graph structure into the spatiotemporal coupling prediction model for training and prediction, and output the predicted value of dengue fever cases and the corresponding risk level of each node in the future specified time period.
[0043] The early warning decision module communicates with the graph neural network prediction module to generate dynamic early warning thresholds based on historical case data analysis, and to make early warning judgments and trigger levels based on case prediction values and real-time monitored multi-source indicators through preset linkage rules.
[0044] The display module communicates with the graph neural network prediction module and the early warning decision module to visualize the predicted values, risk levels and early warning information of cases.
[0045] It should be further explained that the implementation of the dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks begins with the data acquisition module's synchronous acquisition of multi-source heterogeneous data, including case report data from disease control agencies, temperature, precipitation and humidity data provided by meteorological departments, vegetation index sequences retrieved from remote sensing satellites, population flow heat map data generated based on mobile signaling or location applications, and sentiment index data extracted from social media texts through natural language processing technology.
[0046] The data fusion module then performs spatiotemporal alignment and cleaning on the raw data. It uses the Kalman filter algorithm to assimilate the meteorological and remote sensing data sequences to generate a continuous and consistent gridded data field. In particular, it introduces statistical causal analysis to verify the quality of social media sentiment and population flow data. By calculating the transfer entropy with historical case sequences or performing Granger causality tests, only indicators with significant leading relationships are retained as effective features, thereby ensuring that the input information has both breadth and predictive relevance.
[0047] The graph neural network prediction module builds a prediction model based on the high-quality multidimensional feature data fused above. The prediction process of its dynamic multidimensional graph structure construction and built-in cross-validation includes: First, the system uses administrative divisions as nodes, not only generating a static adjacency matrix based on geographical boundaries, but also using population flow intensity data to construct a dynamic flow matrix. The two are then linearly combined through learnable weight parameters to form a comprehensive adjacency matrix that simultaneously encodes geographical proximity and population mobility. This serves as the basis for graph convolution operations, thereby more realistically simulating disease transmission paths.
[0048] The main body of the model adopts a coupled architecture of graph attention network and temporal convolutional network. The graph attention network is responsible for learning dynamic attention weights between nodes on the comprehensive adjacency matrix to aggregate spatial information, while the temporal convolutional network captures the long-term temporal dependence of each node's own feature sequence by expanding causal convolutional layers.
[0049] To further improve prediction reliability, the system sets up a parallel cross-validation mechanism: while the main model is running, an independent gradient boosting decision tree model makes predictions based on the same node features, where the node features do not contain graph structure; the system sets up a difference arbitration unit, when the prediction results of the two models deviate from the preset range, a lightweight bidirectional long short-term memory network is triggered as a secondary validator to arbitrate the discrepancies based on historical context information, and finally output the verified case prediction value and risk level.
[0050] The early warning decision module dynamically divides historical epidemic stages based on the Bayesian variable detection algorithm and calculates the quantile threshold of the number of cases in each stage. It integrates multiple sources of indicators such as the predicted number of cases, real-time mosquito vector monitoring index, social media sentiment index and online search popularity, and makes a comprehensive judgment based on preset multi-level linkage rules and continuous triggering conditions. For example, when the predicted number of cases exceeds a certain quantile threshold continuously and the relevant sentiment index rises synchronously, the corresponding level of early warning is triggered. The rise and fall of the early warning level must meet the continuous time window condition to avoid false alarms.
[0051] Finally, the display module renders the prediction results, risk levels, and early warning information in color-coded zones using an electronic map, and generates a time-series trend chart for visualization output. Meanwhile, the automatic report generation module integrates key information to generate structured early warning briefings, thus forming a full-chain, closed-loop prediction and early warning system that integrates multi-source information perception, spatiotemporal depth modeling, intelligent risk assessment, and intuitive decision support.
[0052] The data fusion module includes a data assimilation unit and a quality verification unit;
[0053] The data assimilation unit is used to estimate and fill in missing values in meteorological data and remote sensing vegetation index data using the Kalman filter method to form a spatiotemporally continuous data sequence.
[0054] The quality verification unit is used to verify the lead-lag relationship between social media sentiment data and population mobility data and historical case data through statistical causal analysis, and to screen out indicators with statistical leadership as effective features.
[0055] It should be further explained that the implementation method of the data assimilation unit in the data fusion module is as follows: For time-series observation data such as temperature, precipitation, and humidity in meteorological data, as well as remote sensing vegetation index data, an ensemble Kalman filter algorithm is used for processing. This algorithm defines the state variable of each grid point or administrative region at each time step as the observation element value of that location, and establishes a state equation describing its evolution over time. At the same time, it uses the data of all available observation stations in the same time period as observation values, and performs optimal estimation and smoothing of the original data with missing or noisy data by recursively executing prediction and update steps, thereby generating a spatiotemporally continuous and consistent meteorological field and vegetation index field.
[0056] The quality verification unit targets indirect indicators such as social media sentiment data and population mobility data. First, the unit performs time-series alignment and smoothing on the sentiment data and population mobility intensity data. The sentiment data is obtained through daily sentiment scores obtained from a sentiment analysis model. Next, using the Granger causality test, at multiple preset time lags, it examines whether these indicator sequences have a statistically significant causal impact on historical case number sequences. Specifically, it constructs a vector autoregression model and calculates the F-statistic for hypothesis testing. Only when the test p-value is below a set significance level is the indicator considered to have predictive leadership. For population mobility data, the transfer entropy between the data and the case sequence is also calculated to quantify the intensity and direction of information flow. Finally, only those sentiment and mobility indicators that pass statistical tests and are proven to have a significant leading indicator effect on case changes are retained as effective features input into the downstream prediction model. This design, by performing physical law-based assimilation repair and statistical causal correlation screening on heterogeneous data, expands the data dimensions while ensuring the quality and predictive value of the input features.
[0057] The graph neural network prediction module is specifically used for constructing the graph structure as follows:
[0058] Obtain multiple basic adjacency matrices that reflect the relationships between different regions. The basic adjacency matrices include at least a geographical adjacency matrix and a population flow matrix.
[0059] Multiple basic adjacency matrices are weighted and combined to generate a comprehensive adjacency matrix, which serves as the basis for the weights of edges connecting nodes in the graph structure.
[0060] The population flow matrix is constructed based on the inter-regional population flow intensity data processed by the data fusion module.
[0061] It should be further explained that the specific implementation method of graph structure construction is as follows: The system first generates a geographic adjacency matrix based on the vector boundary file of administrative divisions through spatial relationships. If the geographic boundaries of two regions are adjacent, the corresponding matrix element value is 1, otherwise it is 0. At the same time, the system uses the time series data of inter-regional population flow intensity obtained by the data fusion module after preprocessing. This data comes from the anonymization processing of mobile signaling or the aggregation statistics of location application. By calculating the mean of the number of people flowing from region i to region j within a specified time window and normalizing it, a population flow matrix is constructed.
[0062] The core steps in generating the comprehensive adjacency matrix are as follows: The geographic adjacency matrix and the population flow matrix are normalized separately, and then linearly weighted and fused using preset or trainable weight coefficients. That is, the comprehensive adjacency matrix = α × geographic adjacency matrix + β × population flow matrix, where α and β are weight parameters greater than 0. Their specific values can be determined through grid search based on the performance of the model validation set, or designed as parameters that can be trained along with the model. This comprehensive adjacency matrix ultimately serves as the input to a graph convolutional or graph attention network layer, defining the strength and direction of message propagation between nodes. This implementation method, by fusing static geographic proximity and dynamic population mobility, constructs a graph structure that better reflects the cross-regional transmission patterns of diseases, providing a crucial foundation for subsequent spatial feature extraction.
[0063] The spatiotemporal coupling prediction model includes a spatial feature extraction unit, a temporal feature extraction unit, a cross-validation unit, and a result arbitration unit;
[0064] The spatial feature extraction unit employs a graph attention network to spatially aggregate node features based on a comprehensive adjacency matrix.
[0065] The temporal feature extraction unit uses a temporal convolutional network to model the temporal dependencies of the historical feature sequences of each node.
[0066] The cross-validation unit runs a prediction model based on gradient boosting decision tree in parallel. The input of the gradient boosting decision tree model is the node feature data after feature engineering, which is used to generate independent prediction results.
[0067] The result arbitration unit is used to compare the master prediction result of the spatiotemporal coupled prediction model with the independent prediction result of the gradient boosting decision tree model. When the difference between the two exceeds a preset threshold, the secondary verification network is activated to arbitrate the prediction discrepancy in order to determine the final prediction value.
[0068] Further explanation is needed regarding the specific implementation of the spatiotemporal coupling prediction model as follows: The spatial feature extraction unit takes a multi-dimensional comprehensive adjacency matrix and a node feature matrix as input, and adopts a multi-layer graph attention network architecture. Each layer of the graph attention network first calculates the attention coefficient for each node pair in the graph. This coefficient is obtained by concatenating the feature vectors of the target node and its neighboring nodes, inputting them into a single-layer feedforward neural network, and applying the LeakyReLU activation function. Subsequently, the attention coefficient is normalized, and finally, the normalized attention coefficient is used as a weight to perform a weighted summation of the features of the neighboring nodes, thereby completing a spatial information aggregation. This process can learn the dynamic and heterogeneous spatial dependencies between nodes. Among them, the neighboring nodes are determined based on the comprehensive adjacency matrix.
[0069] The temporal feature extraction unit processes the historical feature sequence of each node using a temporal convolutional network. This network consists of multiple stacked residual blocks, each containing a causal convolutional layer with an inflation factor, a weight normalization layer, and a gated activation unit. The inflation factor increases exponentially with the network depth to expand the receptive field. Causal convolution ensures that the modeling of historical information does not use future data. Finally, this unit outputs the high-level abstract features of each node in the temporal dimension.
[0070] The cross-validation unit operates in parallel with the aforementioned spatiotemporal feature extraction process. It receives node feature data after feature engineering, which excludes graph structure information. For example, it concatenates the number of cases, meteorological factors, and verified flow and sentiment indicators of each region over the past few weeks into a feature vector, and uses this vector to train a gradient boosting decision tree model. This model independently outputs the predicted value of cases for each region.
[0071] The resulting arbitration unit is a lightweight bidirectional long short-term memory network. It is activated when the absolute difference between the predictions of the main model GAT-TCN and the gradient boosting decision tree model exceeds a preset tolerance threshold. The preset threshold for warning of divergence is set at 15%-25%, meaning that when the relative difference between the main prediction of the spatiotemporally coupled prediction model and the independent predictions of the gradient boosting decision tree model falls within this range, the secondary verification network is activated. It takes the current feature sequence of the divergent node, the state of its neighboring nodes, and the two disputed predictions as input. This bidirectional long short-term memory network, after training, can output an adjusted weight or directly select a more reliable prediction, thus generating a final prediction result that has undergone double verification and arbitration. The lightweight nature of the secondary verification network... The Bi-Short Memory Long Short-Term Memory (BiLSTM) network is specifically configured with a 2-3 layer structure, with 32-64 neurons in each hidden layer. The input dimension of the network is the sum of the node feature dimension and 2, where 2 corresponds to the master prediction value of the spatiotemporally coupled prediction model and the independent prediction value of the gradient boosting decision tree model. The output dimension is 1, used to output the final case prediction value after arbitration. The network uses the tanh activation function, and a dropout layer with a dropout rate of 0.1-0.2 is added after each hidden layer to avoid overfitting. The batch size during network training is 8-32, and the number of training iterations is 50-100. Training stops when the prediction error on the validation set does not decrease for 10 consecutive iterations. If the difference does not exceed a threshold, the weighted average of the two prediction values is directly taken as the output. This implementation effectively improves the robustness and accuracy in a single prediction scenario by constructing a closed loop of parallel prediction and intelligent arbitration of heterogeneous models.
[0072] The graph attention network in the spatial feature extraction unit is used to learn dynamic attention weights between nodes, while the temporal convolutional network in the temporal feature extraction unit contains multiple dilated causal convolutional layers.
[0073] Further explanation is needed regarding the specific implementation details of the spatial feature extraction unit and the temporal feature extraction unit, as follows: The graph attention network adopts a multi-head attention mechanism. The calculation process of each attention head is as follows: For the center node i and any of its neighboring nodes j, j is determined according to the comprehensive adjacency matrix. First, the feature vector h_i of node i and the feature vector h_j of node j are concatenated to obtain the concatenated vector [h_i||h_j]. Then, the concatenated vector is multiplied by a trainable weight vector a, and nonlinearity is introduced through the LeakyReLU activation function to obtain the original attention score e_{ij}=LeakyReLU(a^T·[h_i||h_j]). Next, the Softmax function is used to normalize the original scores of all neighboring nodes of the center node i, including its own, to obtain the normalized attention coefficient α_{ij}. Finally, the features of the neighboring nodes are weighted and summed using these attention coefficients to obtain the updated features of node i under this attention head. The outputs of multiple attention heads will be concatenated or averaged to form the final spatial aggregated features of the nodes.
[0074] In temporal convolutional networks, dilated causal convolutional layers operate on the input time series using a one-dimensional convolutional kernel. The dilation factor d controls the interval skipped by the kernel when processing the sequence. Specifically, in a network with l layers, the dilation factor is typically set to d = 2^{l-1}, which allows higher layers to capture long-range temporal dependencies with exponentially expanded receptive fields. Simultaneously, causal convolution ensures that the output at any given time t depends only on the input at time t and earlier by performing convolution operations only in the forward temporal direction of the sequence, thus avoiding the leakage of future information. Multiple such dilated causal convolutional layers, combined with residual connections, weight normalization, and gated linear units, constitute the main structure for temporal feature extraction.
[0075] The early warning decision module includes a dynamic threshold generation unit and a linkage judgment unit;
[0076] The dynamic threshold generation unit uses a change point detection algorithm to analyze the time series of historical cases, identify the stages of the epidemic development, and calculate the quantile threshold of the number of cases based on a rolling time window in each stage;
[0077] The linkage judgment unit has multiple pre-set early warning rules. The early warning rules compare the predicted value of cases with the quantile threshold and make a joint judgment based on the real-time status of mosquito vector monitoring index, social media sentiment index and Internet search popularity index. The triggering and cancellation of the early warning level must meet the condition of a continuous time window.
[0078] Further explanation is needed regarding the specific implementation details of the early warning decision-making module: The dynamic threshold generation unit first analyzes the historical dengue fever case time series using a change point detection algorithm. Specifically, it employs the PELT algorithm based on penalized likelihood, which automatically identifies multiple time points where the number of cases changes significantly by minimizing a cost function that includes a penalty term for the fitting error and the number of change points. This divides the entire epidemic process into several relatively stable development stages. Within each identified stage, the system uses a fixed-length rolling time window, for example, a window length of the past four weeks. The 50th, 75th, and 90th percentiles of the number of cases within this window are calculated and used as the dynamic thresholds for the current stage's yellow, orange, and red alerts, respectively. These thresholds are updated as the time window rolls to reflect the recent epidemic intensity. The length of the rolling time window used in the dynamic threshold calculation can be selected from a range of 2-6 weeks, depending on the temporal resolution of the epidemic data and the regional epidemic fluctuation characteristics. For high-incidence areas or rapid spread stages of dengue fever, a short window of 2-3 weeks is preferred, while a long window of 4-6 weeks can be used for low-incidence areas or stable stages.
[0079] The linkage judgment unit implements multi-level early warning logic based on a well-defined state machine. This state machine receives case prediction values from the graph neural network prediction module, mosquito vector density indices from the monitoring network (such as the Breteau index), verified social media sentiment indices, and internet search popularity indices for specific keywords. The early warning rules are specifically programmed as follows: a yellow warning is triggered when the number of predicted cases in a certain area exceeds the current dynamic yellow threshold for two consecutive time periods (two weeks); an orange warning is triggered when the number of predicted cases exceeds the dynamic orange threshold for two consecutive periods, and the social media sentiment index or search popularity index also exceeds its respective historical baseline threshold during the same period; a red warning is triggered when the number of predicted cases exceeds the dynamic red threshold for two consecutive periods, the mosquito vector density index exceeds its high-risk threshold, and both the sentiment index and search popularity are at a high level; a downgrade of the warning level requires all relevant indicators to fall below the current level threshold for several consecutive periods. This system aims to improve the accuracy and stability of early warning signals and reduce false alarms caused by single data fluctuations through this multi-indicator collaborative and continuously triggered judgment mechanism.
[0080] The system also includes a model optimization module, which communicates with the early warning decision module and the graph neural network prediction module;
[0081] The model optimization module is used to collect actual case data and corresponding system internal status data after the early warning is issued, forming a feedback dataset, and to periodically fine-tune the parameters of the spatiotemporal coupling prediction model using the feedback dataset.
[0082] Further explanation is needed regarding the specific implementation of the model optimization module, which establishes a closed-loop feedback learning mechanism. When the early warning decision module issues an early warning at a certain level, the system automatically records the timestamp of the early warning event, the administrative regions involved, the early warning level, and all input feature states used to trigger the early warning, including predicted case values, various dynamic thresholds, and all linked real-time indicator data. After a fixed time delay, the system automatically obtains the actual confirmed dengue fever case data for the corresponding region from the CDC database through the data acquisition module after the early warning period.
[0083] The function of the model optimization module is to associate the system's internal state snapshots at these warning times with the actual case data obtained afterward to construct structured feedback training samples. Each sample contains input features, the model's initial prediction output, and the actual case observations. The system is configured with a periodically executed model fine-tuning task, for example, starting every four weeks. During fine-tuning, the system extracts all recent samples from the stored feedback dataset and uses this data to update the parameters of the spatiotemporal coupled prediction model. This update process adopts the idea of transfer learning, which does not retrain the entire model, but fixes most of the low-level parameters in the graph attention network and temporal convolutional network, and only optimizes the parameters of the fully connected regression layer at the end of the network or a small number of high-level network parameters with a small learning rate through supervised gradient descent, so that the model's predicted output moves closer to the actual observed values. The learning rate used in the model fine-tuning process ranges from 1e-5 to 5e-4. In the early stages of fine-tuning, a larger learning rate of 1e-4 to 5e-4 can be selected. As the feedback dataset accumulates and the model's prediction accuracy improves, the learning rate is gradually adjusted to a smaller learning rate of 1e-5 to 1e-4, and the learning rate remains fixed in each fine-tuning process. At the same time, to prevent overfitting to the limited feedback data, an L2 regularization penalty term for the original model parameters is added to the fine-tuning loss function.
[0084] This module continuously integrates the latest actual development data of the epidemic to adjust the model, enabling the entire prediction and early warning system to adapt to the evolution of the epidemic, thereby maintaining and improving its prediction accuracy and early warning reliability over time.
[0085] The display module provides an electronic map visualization interface, which renders risk levels on geographic zones with different color levels, and provides time trend charts of case prediction values and markers of warning trigger times.
[0086] Further explanation is needed regarding the specific implementation of the demonstration module: This module constructs a web-based interactive visualization platform. Its electronic map visualization interface is developed using open-source geographic information system libraries, such as Leaflet or Mapbox, and the underlying layer loads vector boundary data of administrative divisions. The system receives risk level data and case prediction values for each region from the graph neural network prediction module. Risk levels are divided into multiple levels. Based on preset color mapping rules, the platform assigns a specific color from a gradient color system from light to dark to each risk level. For example, low risk corresponds to light green, medium risk to yellow, high risk to orange, and extremely high risk to red. The map rendering engine then fills each administrative division with the color corresponding to its risk level, thereby generating a heat map of risk level distribution. At the same time, when the user hovers the mouse over a certain area, the interface dynamically displays detailed prediction data for that area in the form of a pop-up information box, including the specific predicted number of cases, risk level, and a summary of the main influencing factors.
[0087] In addition, this module provides a time-series trend chart display function. For one or more regions selected by the user, the system retrieves the historical case observation sequence, model prediction sequence, and dynamic threshold sequence calculated by the early warning decision module from the backend database for that region, and uses the front-end chart library to draw a time-series line chart. Different sequences are distinguished by different line styles and colors, and the specific time points of early warning triggering and lifting are clearly marked on the time axis. This trend chart supports zooming and panning interaction, making it easy for users to analyze the details of the epidemic evolution over a specific time period. All data in the visualization views are synchronized in real time with the backend prediction and early warning engine through the application programming interface to ensure the timeliness of the displayed information.
[0088] The data fusion module also includes a data version management unit, which is used to attach version identifiers to the fused multidimensional spatiotemporal feature data and intermediate processing data to achieve traceability of the data processing process.
[0089] It should be further explained that the specific implementation of the data version management unit is as follows: This unit provides systematic version control and traceability capabilities for the entire data fusion and processing chain. It automatically generates and attaches structured version identifiers for each batch of raw data, intermediate feature data after data assimilation and quality verification, and the final output multidimensional spatiotemporal feature data matrix.
[0090] The generation rules for version identifiers combine time elements, data source fingerprints, and processing configurations: First, the system assigns a globally unique transaction ID when each data acquisition task is started; second, it calculates the content hash value, such as SHA-256, as a data fingerprint for each type of raw data file, such as meteorological observation files and social media data stream snapshots; finally, it serializes the key configuration parameters of the current data processing pipeline, such as homogenization algorithm type, significance level threshold for causal testing, and time window size.
[0091] The version identifier is composed of these three parts concatenated in a fixed format, such as "V_Transaction ID_Data Fingerprint Summary_Configuration Summary". All version identifiers and their associated metadata, such as data time range, geographical coverage, and processing timestamps, are persistently stored in a dedicated version metadata database. When a subsequent graph neural network prediction module calls feature data of a certain version for training or prediction, the version identifier will be automatically recorded in the model training log or prediction result metadata. When the early warning decision module or display module needs to trace back the basis for a certain early warning or visualization result, it can query this metadata database to fully trace back to the original data batch, specific processing steps, and parameter configurations on which the result depends.
[0092] This mechanism ensures end-to-end traceability from the final warning signal to the original data source, providing a reliable foundation for the verification, auditing, and performance comparison analysis between different versions of the model results.
[0093] The system also includes an automatic report generation module, which communicates with the early warning decision-making module and the display module to automatically integrate the current early warning level, risk area information, key indicator status and recommended measures to generate a structured early warning report document when an early warning is triggered.
[0094] It should be further explained that the specific implementation of the automatic report generation module is as follows: When the early warning decision module triggers an early warning event of any level, the event acts as a signal to automatically activate the automatic report generation module. The module then initiates a structured information aggregation and document generation process. First, it obtains the core metadata of this early warning from the early warning decision module, including the early warning level, trigger time, list of geographical areas involved, and specific triggering rules. Next, the module queries the graph neural network prediction module for the case prediction data and risk level change trends of the relevant area during the current early warning period and a certain time window in the future. At the same time, it retrieves the risk rendering snapshot of the corresponding area on the electronic map and the time series chart data of key indicators from the background database of the display module.
[0095] Based on these multi-source inputs, the module's core report generation engine populates the content according to a preset template: the report begins by generating an execution summary, outlining the warning level, core risk areas, and main basis; then the main body elaborates in detail in chapters, including a comparative analysis of the specific predicted case numbers and thresholds for risk areas, the real-time status of related mosquito vector monitoring indices and social media sentiment indices, an assessment of the epidemic development trend based on historical data, and a brief description of the population and socioeconomic characteristics of the affected areas; the report also automatically includes relevant visualization charts as illustrations.
[0096] After the content is populated, the module converts the report into a standardized document format according to the warning level and preset release rules. The report automatically includes version identifiers for the data and models used in the analysis, provided by the data version management unit, ensuring traceability. The final complete warning report can be automatically sent to a pre-configured email list, internal collaboration platform, or public warning information release system via an integrated communication interface, achieving an automated closed loop from internal system judgment to structured, traceable external release of warning information.
[0097] This system effectively improves the precision and accuracy of dengue fever outbreak prediction by integrating multi-source heterogeneous data and constructing a dynamic multi-dimensional graph structure for spatiotemporal coupled prediction. The system employs data assimilation and causal verification to ensure the quality of input information, utilizes graph attention networks and temporal convolutional networks to capture complex spatial dependencies and temporal dynamics, and combines parallel model cross-validation and arbitration mechanisms to enhance the reliability of prediction results. Dynamic early warning thresholds and multi-indicator linkage triggering rules further improve the sensitivity and stability of risk identification and reduce false alarms.
[0098] The system achieves end-to-end automation and closed-loop optimization from data perception and intelligent analysis to decision support. A visual interactive interface provides intuitive risk assessment, version control ensures full traceability, and automatic report generation and publishing enhances emergency response efficiency. This system provides timely, accurate, and actionable risk warning information for dengue fever prevention and control, helping to optimize the allocation of prevention and control resources and improve the ability to respond to public health emergencies.
[0099] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0100] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks, characterized in that, It includes a data acquisition module, a data fusion module, a graph neural network prediction module, an early warning and decision-making module, and a display module; The data acquisition module is used to collect case data, meteorological data, remote sensing vegetation index data, population flow data, and social media sentiment data. The data fusion module communicates with the data acquisition module and is used to clean, time-align, and process missing values of the acquired heterogeneous data to generate fused multidimensional spatiotemporal feature data. The graph neural network prediction module communicates with the data fusion module to receive the multidimensional spatiotemporal feature data, construct a graph structure reflecting the relationship between regions with administrative division units as nodes, and input the multidimensional spatiotemporal feature data and the graph structure into the spatiotemporal coupling prediction model for training and prediction, and output the predicted value of dengue fever cases and the corresponding risk level of each node in the future specified time period. The early warning decision module communicates with the graph neural network prediction module and is used to generate dynamic early warning thresholds based on historical case data analysis. Based on the case prediction values and real-time monitored multi-source indicators, it performs early warning judgment and level triggering through preset linkage rules. The display module communicates with the graph neural network prediction module and the early warning decision module to visualize the predicted values, risk levels and early warning information of the cases.
2. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 1, characterized in that: The data fusion module includes a data assimilation unit and a quality verification unit; The data assimilation unit is used to estimate and fill in missing values in the meteorological data and the remote sensing vegetation index data using the Kalman filter method to form a spatiotemporally continuous data sequence. The quality verification unit is used to verify the lead-lag relationship between the social media sentiment data and the population flow data and historical case data through statistical causal analysis, and to screen out indicators with statistical leadership as effective features.
3. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 2, characterized in that: The graph neural network prediction module, when constructing the graph structure, is specifically used for: Obtain multiple basic adjacency matrices that reflect the relationships between different regions, wherein the basic adjacency matrices include at least a geographical adjacency matrix and a population flow matrix; The multiple basic adjacency matrices are weighted and combined to generate a comprehensive adjacency matrix, which serves as the weighting basis for the edges connecting nodes in the graph structure. The population flow matrix is constructed based on the inter-regional population flow intensity data processed by the data fusion module.
4. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 3, characterized in that: The spatiotemporal coupling prediction model includes a spatial feature extraction unit, a temporal feature extraction unit, a cross-validation unit, and a result arbitration unit; The spatial feature extraction unit employs a graph attention network to spatially aggregate node features based on the comprehensive adjacency matrix. The time feature extraction unit uses a time convolutional network to model the time dependency relationship of the historical feature sequence of each node. The cross-verification unit runs a prediction model based on a gradient boosting decision tree in parallel. The input of the gradient boosting decision tree model is the node feature data after feature engineering, which is used to generate independent prediction results. The result arbitration unit is used to compare the master prediction result of the spatiotemporal coupled prediction model with the independent prediction result of the gradient boosting decision tree model. When the difference between the two exceeds a preset threshold, a secondary verification network is activated to arbitrate the prediction discrepancy in order to determine the final prediction value.
5. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 4, characterized in that: The graph attention network in the spatial feature extraction unit is used to learn dynamic attention weights between nodes, and the temporal convolutional network in the temporal feature extraction unit contains multiple dilated causal convolutional layers.
6. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 1, characterized in that: The early warning decision module includes a dynamic threshold generation unit and a linkage judgment unit; The dynamic threshold generation unit uses a change point detection algorithm to analyze the time series of historical cases, identify the stages of the epidemic development, and calculate the quantile threshold of the number of cases based on a rolling time window in each stage. The linkage judgment unit is pre-set with multi-level early warning rules. The early warning rules compare the predicted value of the case with the quantile threshold and make a joint judgment based on the real-time status of the mosquito vector monitoring index, social media sentiment index and Internet search popularity index. The triggering and cancellation of the early warning level must meet the condition of a continuous time window.
7. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 6, characterized in that: The system also includes a model optimization module, which communicates with the early warning decision module and the graph neural network prediction module. The model optimization module is used to collect actual case data and corresponding system internal status data after the early warning is issued, form a feedback dataset, and use the feedback dataset to periodically fine-tune the parameters of the spatiotemporal coupling prediction model.
8. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 1, characterized in that: The display module provides an electronic map visualization interface, which is used to render the risk level on the geographic partition with different color levels, and provides a time trend chart of the predicted case values and a mark of the warning trigger time.
9. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 1, characterized in that: The data fusion module also includes a data version management unit, which is used to attach version identifiers to the fused multidimensional spatiotemporal feature data and intermediate processing data to achieve traceability of the data processing process.
10. The dengue fever prediction and early warning system based on heterogeneous data fusion and graph neural networks according to claim 1, characterized in that: The system also includes an automatic report generation module, which communicates with the early warning decision module and the display module. When an early warning is triggered, the module automatically integrates the current early warning level, risk area information, key indicator status, and recommended measures to generate a structured early warning report document.
Citation Information
Patent Citations
CN111554408A