A method and system for dynamic monitoring and early warning of sfts based on multi-modal data
By using multimodal data fusion and graph neural network analysis, the problem of multi-source data correlation and dynamic changes in SFTS epidemic monitoring was solved, enabling accurate risk prediction and rapid prevention and control response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DITAN HOSPITAL CAPITAL MEDICAL UNIVERSTY
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-12
AI Technical Summary
Existing monitoring methods are insufficient to perform semantic-level correlation and real-time capture of dynamic changes in multi-source heterogeneous data, resulting in a lack of accuracy and timeliness in SFTS epidemic risk prediction, and making it difficult to implement targeted and rapid prevention and control measures.
A dynamic monitoring system is constructed by multimodal data fusion, graph neural network correlation analysis, and interpretable risk assessment. Attention mechanism is used to screen key features, graph neural network is used to model semantic associations, conditional independence test and gradient boosting tree are combined to identify driving factors, temporal convolutional network is used to predict risks, and interpretable early warning signals are generated by SHAP value and dynamic weight adjustment.
It has achieved comprehensive correlation and dynamic capture of multi-source data, improved the comprehensiveness and accuracy of epidemic risk assessment, the timeliness of early warning and the pertinence of prevention and control measures, and realized the transformation from passive response to proactive prediction.
Smart Images

Figure CN122201815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring and early warning technology, specifically to a dynamic monitoring and early warning method and system for SFTS based on multimodal data. Background Technology
[0002] Severe fever with thrombocytopenia syndrome (SFTS) is an acute infectious disease transmitted by ticks. It has a high mortality rate and its transmission risk dynamically changes with the environment and population, posing a serious threat to public health security. Effective monitoring and early warning of regional SFTS risks are crucial for epidemic control, as they are related to timely interruption of transmission chains and reduction of socioeconomic losses. However, existing monitoring methods have significant shortcomings in dealing with complex and dynamic environments and the integration of multi-source information, making it difficult to meet the needs of precise prevention and control. Traditional SFTS monitoring methods mainly rely on data collection from single scenarios, such as case reports from medical institutions or laboratory virus testing, lacking comprehensive analysis of multi-dimensional information on the environment, virus, and population. This method struggles to capture the correlations between different factors, such as how changes in tick habitats affect virus transmission or how population movement exacerbates regional risks. The problem of information silos prevents the monitoring system from fully reflecting the dynamics of the epidemic, often resulting in delayed early warnings and missed opportunities for optimal prevention and control.
[0003] The core technical challenge lies in achieving deep correlation and real-time capture of dynamic changes from multi-source data. Data from different sources, such as environmental temperature and humidity, clinical symptoms, viral gene sequences, and population movement trajectories, have different formats and time scales, making it difficult to integrate them into meaningful relationships. For example, tick density in a region may surge due to increased seasonal rainfall, but current technology struggles to correlate this environmental change with the decline in platelet counts in cases and viral mutation characteristics in real time. This semantic integration challenge across different data domains results in a lack of accuracy and timeliness in risk prediction. Capturing dynamic changes is another key technical bottleneck. Because SFTS transmission is influenced by multiple factors such as season, geography, and population behavior, the temporal correlations between data are complex. For instance, a township may report suspected cases continuously within a week, while simultaneously monitoring increased tick activity. However, existing systems struggle to determine the risk level in real time based on these dynamic signals, and cannot identify which factors contribute most to the risk, making it difficult for control departments to quickly formulate targeted measures. Therefore, how to semantically correlate multi-source heterogeneous data and capture its dynamic trends in real time to generate accurate regional risk predictions and interpretable early warning information has become a key issue in building a dynamic monitoring and early warning system for SFTS. Summary of the Invention
[0004] The purpose of this invention is to provide a dynamic monitoring and early warning method and system for SFTS based on multimodal data. By using multimodal data fusion, graph neural network correlation analysis and interpretable risk assessment, a dynamic monitoring system capable of accurate early warning of SFTS outbreaks is constructed, effectively improving the initiative and accuracy of epidemic prevention and control.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] This application provides a method for dynamic monitoring and early warning of SFTS based on multimodal data, including the following steps:
[0007] S1. Standardize environmental monitoring data, clinical symptom reports, viral gene sequences and population movement trajectories through sensor networks and database interfaces. Employ a feature weighting method based on attention mechanism to strengthen key features strongly correlated with SFTS transmission and establish an enhanced dataset with spatiotemporal consistency.
[0008] S2. A graph neural network is used to perform semantic integration processing on the augmented dataset, and different modal data are modeled as multi-type nodes. Through the temporal graph attention mechanism, the correlation strength between nodes is dynamically learned, especially the coupling relationship between environmental changes and virus transmission paths, to generate a multi-scale correlation graph with time dimension.
[0009] S3. Identify causal paths between environmental factors, viral mutations, and case reports from multi-dimensional correlation maps, and use a combination of conditional independence tests and gradient boosting trees to determine key driving factors and their mechanisms of action.
[0010] S4. Based on the causal analysis results, a prediction model integrating temporal convolutional networks and attention mechanisms is established to capture both short-term fluctuations and long-term trends. Monte Carlo dropout technology is used to quantify prediction uncertainty and generate a risk evolution vector containing confidence intervals.
[0011] S5. Quantitatively analyze the contribution of each component in the risk evolution vector through SHAP value. When specific components such as population flow are identified as the dominant factors, automatically start the dynamic weight adjustment algorithm to generate regional risk heat maps that mark the main sources of risk.
[0012] S6. Based on the dominant risk sources marked by risk heatmaps, generate an interpretable early warning signal sequence containing risk level, spatial location and dominant cause from the intelligent template library;
[0013] S7. Based on the early warning signal sequence, if the signal sequence indicates an increase in risk level, a real-time notification mechanism is triggered to push the correlation details to the prevention and control system, thereby obtaining the final regional risk prediction output.
[0014] This invention provides a dynamic monitoring and early warning system for SFTS based on multimodal data, applied to a dynamic monitoring and early warning method for SFTS based on multimodal data, comprising:
[0015] The data augmentation module is used to collect environmental, clinical, viral, and population trajectory data from sensors and databases. After standardization, it uses an attention mechanism to filter key features and combines K-means clustering and trajectory matching to construct an augmented dataset.
[0016] The graph construction module divides the augmented dataset into multiple types of nodes according to modality, models semantics through graph neural networks, calculates the association strength using temporal attention if time series information is included, analyzes the coupling relationship between the environment and virus transmission, and generates a multi-scale association graph with time dimension.
[0017] The causal identification module, based on the association graph, extracts the initial causal path using the conditional independence test, and then determines the key driving factors by calculating the path weights through gradient boosting tree; if the environmental or viral mutation weights exceed the limit, specific variables are extracted and a mechanism model is constructed to generate a causal path graph.
[0018] The risk prediction module combines causal analysis results with a temporal convolutional network to highlight long-term trends using short-term fluctuation characteristics and an attention mechanism, and integrates these to construct a prediction model. If the output deviation exceeds the standard, the uncertainty is quantified through Monte Carlo sampling and dropout techniques to generate a risk evolution vector with confidence intervals.
[0019] The heatmap generation module performs SHAP analysis on the risk evolution vector. If the contribution of population movement exceeds the standard, it is identified as a core risk. The parameters are updated using a dynamic weight adjustment algorithm to generate a preliminary heatmap. After smoothing optimization and overlaying with geographic information, a final heatmap with labeled risk sources is obtained.
[0020] The early warning push module extracts risk sources from heat maps, uses algorithms to determine risk levels and causes, and generates early warning sequences by matching template libraries. If the risk increases, a notification is triggered, pushing relevant details to the prevention and control system. Logistic regression is then used to assess regional risks and output accurate prediction results.
[0021] The beneficial effects of this invention are as follows:
[0022] By integrating multimodal data and aligning it with spatiotemporal data, the problem of information silos was solved. By employing feature weighting with attention mechanisms and graph neural network modeling, heterogeneous data from multiple sources, such as environmental monitoring, clinical symptoms, viral genes, and population trajectories, were semantically integrated to construct an enhanced dataset and multi-scale association map with spatiotemporal consistency. This breakthrough overcame the limitations of traditional single data sources and enabled the comprehensive capture of the complex relationships between tick density, viral mutations, and case reports. It also achieved a shift from fragmented monitoring to systematic risk assessment, significantly improving the comprehensiveness and accuracy of dynamic epidemic perception.
[0023] By using dynamic causal analysis and uncertainty quantification, the problem of delayed early warning is solved. By combining conditional independence test and gradient boosting tree to identify key driving factors, and using temporal convolutional network and Monte Carlo dropout technology to build a prediction model, it can capture short-term fluctuations and long-term trends and quantify uncertainty. It can analyze the dynamic impact of environmental changes, virus mutations and other factors on the risk of transmission in real time and generate risk evolution vectors with confidence intervals. This realizes the transformation from passive response to active prediction and greatly improves the timeliness and reliability of early warning.
[0024] By generating and pushing interpretable early warnings in real time, the problem of insufficient targeting in prevention and control has been solved. The SHAP value analysis is used to quantify the risk contribution, and a heat map with marked risk sources is generated by dynamic weight adjustment. An interpretable early warning signal containing risk level, location and cause is automatically generated through an intelligent template library. When the risk escalates, the relevant details are pushed to the prevention and control system in real time, enabling the prevention and control department to quickly identify the dominant risk factors and key areas. This realizes the transformation from generalized early warning to precise intervention, effectively improving the targeting and response efficiency of prevention and control measures. Attached Figure Description
[0025] To better understand and implement this application, the technical solution is described in detail below with reference to the accompanying drawings.
[0026] Figure 1 This is a flowchart illustrating a dynamic monitoring and early warning method for SFTS based on multimodal data, provided in Embodiment 1 of this application.
[0027] Figure 2 This is a flowchart illustrating step S2 in an SFTS dynamic monitoring and early warning method based on multimodal data provided in Embodiment 1 of this application.
[0028] Figure 3 This is a flowchart illustrating step S4 in a method for dynamic monitoring and early warning of SFTS based on multimodal data provided in Embodiment 1 of this application.
[0029] Figure 4 This is a schematic diagram of the structure of an SFTS dynamic monitoring and early warning system based on multimodal data, provided in Embodiment 2 of this application. Detailed Implementation
[0030] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.
[0031] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0032] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.
[0033] Example 1
[0034] Please see Figures 1-3 This embodiment provides a method for dynamic monitoring and early warning of SFTS based on multimodal data, including the following steps:
[0035] S1. Standardize environmental monitoring data (temperature and humidity, tick density), clinical symptom reports (fever level, platelet count), viral gene sequences (mutation sites, evolutionary branches), and population movement trajectories through sensor networks and database interfaces. Employ a feature weighting method based on attention mechanism to strengthen key features strongly correlated with SFTS transmission and establish an enhanced dataset with spatiotemporal consistency.
[0036] Further, step S1 specifically includes:
[0037] Acquire temperature and humidity data, tick density, fever level, platelet count, mutation sites, evolutionary branches, and movement trajectories from sensor network and database interfaces. Normalize the data using standardized processing methods to obtain an initial dataset in a unified format.
[0038] The temperature and humidity data, tick density, fever level, platelet count, mutation sites, evolutionary branches and movement trajectories in the initial dataset are weighted using an attention mechanism to determine the weights of key features related to SFTS transmission.
[0039] Specifically, an attention mechanism was used to apply targeted weighting to temperature and humidity data, tick density, fever level, platelet count, viral mutation sites, viral evolutionary branches, and population movement trajectories. By combining the biological characteristics and epidemiological patterns of SFTS transmission, such as ticks being the main transmission vector and their density and temperature and humidity directly affecting viral host activity; fever and platelet count being key indicators for the clinical diagnosis of SFTS; viral mutation sites and evolutionary branches being associated with changes in transmissibility; and population movement trajectories determining the spatial spread range of the virus, the correlation strength between various features and SFTS transmission cases, as well as the contribution gradient of features in transmission events, was calculated to dynamically assign weights to different features: features with a more significant impact on transmission were assigned higher weights, such as tick density in high-risk areas, platelet count thresholds related to severe illness, and viral mutation sites that promote transmission; features with a weaker correlation to transmission were assigned lower weights, such as routine temperature and humidity fluctuations in non-critical areas. Ultimately, the correlation weight values between each type of feature and SFTS transmission were clearly quantified and determined, providing a core basis for subsequent screening of key features and construction of high-quality augmented datasets.
[0040] If the weight of a key feature is higher than a preset threshold, the feature is retained and its spatiotemporal information is extracted to obtain a feature set containing spatiotemporal labels. The K-means algorithm is used to cluster the feature set to obtain data groups with spatiotemporal consistency. Based on the data groups, temperature and humidity data, tick density, fever level, platelet count, mutation sites and evolutionary branches are fused to construct an enhanced dataset.
[0041] By matching the movement trajectory with the spatiotemporal consistency of the augmented dataset, the correlation between the trajectory and SFTS propagation is determined, resulting in an augmented dataset with spatiotemporal consistency.
[0042] Among them, matching the spatiotemporal consistency of the movement trajectory with the enhanced dataset includes: integrating data such as temperature and humidity, tick density, clinical symptoms, and viral gene sequences, and each data point is accompanied by a clear spatiotemporal label, such as collection timestamp, geographic coordinates or administrative region code. At the same time, the movement trajectory data needs to be preprocessed into a standardized format that includes the timestamps and geographic coordinates of each trajectory point to ensure that the units and precision of the spatiotemporal dimensions of the two are consistent.
[0043] Next, spatiotemporal consistency matching is initiated: In the time dimension, based on the time window of a single record in the augmented dataset, movement trajectory segments whose timestamps fall within that window are selected; in the spatial dimension, using geospatial indexes (such as grid indexes or R-tree indexes), the geographic coordinates of the trajectory points are matched to see if they overlap with the spatial range of the augmented dataset record (such as a township or a 500-meter grid). If the trajectory segment is perfectly aligned with a certain augmented data record in both time and space, it is considered a preliminary spatiotemporal match. Subsequently, SFTS transmission correlation judgment is performed: Combining the key transmission characteristics of the spatiotemporal record in the augmented dataset, such as whether the tick density is higher than the transmission risk threshold, whether there are clinical case reports, and whether the virus mutation is a highly transmissible branch, the population corresponding to the matched trajectory is analyzed to see if there is a possibility of transmission exposure. If the trajectory segment stays in the spatiotemporal range for more than a preset threshold, or the trajectory path covers areas with high tick density or case activity areas, if the exposure conditions are met, it is determined that the trajectory is correlated with SFTS transmission. Finally, all the mobile trajectory data that were determined to be spatiotemporally matched and had a transmission association were merged with the corresponding augmented dataset records to supplement the associated information such as trajectory dwell time and movement path. At the same time, augmented data records and trajectory data that failed to match or had weak association were removed. The final result is a spatiotemporally consistent augmented dataset in which each record contains environmental, clinical, and viral data and associated mobile trajectory information, and the spatiotemporal dimensions are fully aligned and the transmission association is clear.
[0044] Specifically, by standardizing multi-source heterogeneous data and using attention-based feature weighting, the integration challenges caused by inconsistent formats and spatiotemporal scales of environmental, clinical, viral, and population movement data were effectively solved. Spatiotemporal consistency matching technology was used to accurately correlate movement trajectories with multimodal data, constructing an enhanced dataset that combines spatiotemporal alignment and transmission correlation. This enabled the transformation from fragmented information to systemic risk characteristics, laying a high-quality data foundation for subsequent accurate monitoring and early warning.
[0045] S2. A graph neural network is used to perform semantic integration processing on the augmented dataset, and different modal data are modeled as multi-type nodes. Through the temporal graph attention mechanism, the correlation strength between nodes is dynamically learned, especially the coupling relationship between environmental changes and virus transmission paths, to generate a multi-scale correlation graph with time dimension.
[0046] Further, step S2 specifically includes:
[0047] S21. Divide the augmented dataset into multiple types of nodes according to modality, and use graph neural networks to perform semantic modeling on the multiple types of nodes to generate initial node embedding representations. If the initial node embedding representations contain time series information, calculate the dynamic association strength between nodes through a temporal attention mechanism to obtain a weighted association matrix.
[0048] Based on the pre-defined multi-type nodes, such as environmental nodes carrying information on temperature, humidity, and tick density; clinical nodes recording fever levels and platelet counts; and virus nodes storing mutation sites and evolutionary branches, a graph neural network (GNN) is initiated for semantic modeling. First, each type of node is assigned an initial feature representation: environmental nodes are initialized with specific temperature and humidity values and tick density statistics; clinical nodes are initialized with fever level codes and platelet count interval vectors; and virus nodes are initialized with gene sequence feature matrices. Then, the GNN uses a message passing mechanism to allow different types of nodes to exchange semantic information. For example, environmental nodes pass the influence of temperature and humidity on virus survival features to virus nodes, and virus nodes pass the influence of mutations on pathogenicity features to clinical nodes. This allows each node to retain its own modal core information while integrating the semantic associations of related nodes. Finally, the GNN's aggregation layer integrates and maps the interacting node features, transforming high-dimensional, heterogeneous multimodal semantics into a low-dimensional, unified vector form, generating an initial node embedding representation containing its own features and cross-modal association semantics.
[0049] S22. Based on the weighted correlation matrix, analyze the interaction patterns between environmental change nodes and virus transmission nodes, generate a coupling relationship representation, and combine the coupling relationship representation with time dimension information through multi-scale graph construction technology to generate a multi-scale correlation graph.
[0050] This process involves extracting submatrices of two types of nodes from the weighted association matrix, focusing on the association strength data between environmental nodes and virus transmission nodes, including weight values at different time steps. Combined with time series analysis of interaction patterns, cross-correlation is calculated to identify the time lag between changes in environmental characteristics and fluctuations in virus transmission indicators (such as increased frequency of mutation sites) (e.g., increased virus transmission 3-7 days after environmental changes). Through rule mining, typical interaction patterns are extracted, such as when humidity > 60% and tick density > threshold, the association strength of virus evolutionary branch transmissibility increases by more than 20%. Finally, these patterns are transformed into a structured representation of coupling relationships: static associations are recorded using triples, including environmental node characteristics, association strength, and virus transmission node characteristics; dynamic trends are labeled using time series vectors; and coupling type labels are added, forming coupling relationship data that can be directly used for map construction.
[0051] By using graph neural networks and temporal attention mechanisms, the problem of semantic fragmentation and difficulty in capturing dynamic associations in multimodal data is effectively solved. Heterogeneous data such as environmental, clinical, and viral data are modeled as interactive nodes and dynamically weighted association graphs are generated. This enables deep semantic fusion and spatiotemporal association mining of cross-domain features, thereby accurately revealing the coupling mechanism between environmental changes and virus transmission, and providing reliable structured knowledge support for risk tracing and trend prediction.
[0052] S3. Identify causal paths between environmental factors, viral mutations, and case reports from multi-dimensional correlation maps, and use a combination of conditional independence tests and gradient boosting trees to determine key driving factors and their mechanisms of action.
[0053] Further, step S3 specifically includes:
[0054] Data from multi-dimensional correlation graphs were obtained, and conditional independence tests were used to extract significant data correlations from environmental factors, viral mutations, and case reports to obtain an initial set of causal paths. Then, gradient boosting trees were used to calculate the weights of each path and determine the key driving factors.
[0055] The initial set of causal paths obtained from the conditional independence test is used as input, including paths such as environmental factors → case reports, virus mutation → case reports, and environmental factors → virus mutation → case reports. First, the features of each path are transformed into numerical feature vectors that can be processed by the gradient boosting tree. Path features include the correlation strength of the variables involved in the path, the spatiotemporal range covered by the path, and the frequency of the path's occurrence in historical transmission cases. Then, the gradient boosting tree model is trained, iteratively generating multiple weak classifiers. Each classifier focuses on correcting the bias in the preceding model's prediction of path weights. Finally, by accumulating the prediction results of all weak classifiers, the weight value of each causal path is quantified and output. A higher weight indicates a more significant impact of the path on the occurrence of SFTS cases. Finally, a weight threshold is set, such as an impact contribution threshold calibrated based on historical SFTS outbreak data, to screen out paths with weights exceeding the threshold. The core influencing factors of these paths are identified as key driving factors.
[0056] If the path weight of environmental factors in the key driving factors is higher than the preset threshold, then temperature, humidity and pollution index are extracted from the environmental factors to obtain specific environmental driving variables. If the proportion of viral mutations in the key driving factors is higher than the preset threshold, then gene sequence changes and mutation rates are extracted from viral mutations to obtain specific mutation driving variables.
[0057] Based on key driving factors and specific driving variables, the mechanism of action of causal paths is constructed to obtain a mechanism description model. Then, the interaction between environmental driving variables and variation driving variables is analyzed to determine the interaction path weights, generate the dynamic influence path of key driving factors on case reports, and obtain the final causal path diagram.
[0058] The resulting mechanism description model includes: clarifying the correspondence between key driving factors and their corresponding specific driving variables. Key driving factors include environmental factors and viral mutations. Specific driving variables include temperature / humidity under environmental factors and gene sequence changes / mutation rates under viral mutations. Then, combining the biological mechanisms and epidemiological patterns of SFTS transmission, the logical relationships between variables are identified: for example, "increased temperature → enhanced tick activity → expanded viral host transmission range → increased number of reported cases"; "mutation at a certain site in the gene sequence → enhanced viral binding to host cells → enhanced viral transmissibility → increased number of reported cases." Next, through statistical validation and data fitting, such as calculating the correlation between case incidence and tick density in different temperature ranges and the differences in case growth trends under different mutation rates, these logical relationships are transformed into quantifiable rules or mathematical expressions. For example, when the temperature is between 25-30℃ and humidity > 70%, for every 10 ticks / hectare increase in tick density, the number of reported cases increases by an average of 15%; when the viral mutation rate is > 0.002 / site / year, the case transmission radius expands by 2... Finally, using the structure of key driving factors → specific driving variables → intermediate action links → case report impact on results, all quantitative rules and logical connections are integrated to form a mechanism description model. The mechanism details are optimized through model fit test, such as comparing the deviation between the case change trend predicted by the model and the actual case data, to ensure that the model can accurately reflect the logic of the role of key driving factors in SFTS case reporting.
[0059] Specifically, by combining conditional independence testing with gradient boosting tree analysis, the problem of identifying key driving paths in a multi-factor environment was effectively solved. It not only quantitatively assessed the causal influence strength among environmental factors, viral mutations, and case reports, but also constructed an interpretable causal path diagram with a clear mechanism of action. This enabled source analysis from surface correlations to deep driving factors, providing a scientific basis for precise intervention and strategy formulation.
[0060] S4. Based on the causal analysis results, a prediction model integrating temporal convolutional networks and attention mechanisms is established to capture both short-term fluctuations and long-term trends. Monte Carlo dropout technology is used to quantify the prediction uncertainty and generate a risk evolution vector containing confidence intervals.
[0061] Further, step S4 specifically includes:
[0062] S41. Obtain the causal analysis results, determine the causal impact of key variables on the time series, obtain the causal weight set, process the causal weight set through a temporal convolutional network, extract the short-term fluctuation characteristics in the time series, and obtain the short-term fluctuation vector.
[0063] The causal weight set obtained from causal analysis includes the causal influence weights of key variables such as temperature, humidity, and viral mutation on the spread of SFTS at different time steps, forming time-series data. This data is then input into a temporal convolutional network (TCN) for processing. The TCN uses multi-layer dilated convolution operations with convolution kernels of different sizes to focus on capturing short-term, local dynamic changes in the time series. For example, a sudden increase in temperature over a few days can lead to a sudden increase in its causal weight, or rapid fluctuations in the spread association weights after the emergence of viral mutation sites. These convolutional layers can effectively extract short-term fluctuation features such as high-frequency oscillations and sudden peaks / troughs in the weight sequence. Through feature integration and dimensionality compression, the final output is a short-term fluctuation vector containing key information about short-term dynamic changes.
[0064] S42. Employ an attention mechanism to weight the short-term fluctuation vector, highlighting the long-term trend information at key time steps, to obtain the long-term trend vector. Integrate the short-term fluctuation vector and the long-term trend vector to construct a prediction model and generate initial prediction results.
[0065] The construction of the prediction model includes: using the short-term fluctuation vector output by the temporal convolutional network, which contains short-term dynamic features related to SFTS transmission at each time step, such as sudden changes in daily temperature and humidity and features corresponding to short-term fluctuations in the number of cases; activating the attention mechanism to calculate the attention weight of each time step; and measuring the correlation between the features of this time step and the overall time series trend. For example, if the viral mutation weight suddenly increases at a certain time step and is affected by it in multiple subsequent time steps, the weight of this step will be amplified. Higher weights are assigned to the features of key time steps in the short-term fluctuation vector, such as the tick activity period and the time of occurrence of highly pathogenic mutations that drive the long-term transmission trend. Lower weights are assigned to the features of non-key time steps. After weighted integration, a long-term trend vector reflecting the overall trend of SFTS transmission is extracted. Subsequently, element-level fusion (or feature splicing) is used to combine the detailed information of the short-term fluctuation vector with the overall trend information of the long-term trend vector to construct an integrated prediction model of temporal convolutional network and attention mechanism. The fused feature vector is input into the output layer of the model and mapped to the predicted value of SFTS transmission risk for a specific future time period through an activation function, generating initial prediction results, such as the risk index sequence of each time step.
[0066] S43. If the output deviation of the prediction model exceeds the preset threshold, the model parameters are sampled multiple times using the Monte Carlo method to obtain the parameter distribution set. The dropout technique is then applied to randomly deactivate the neurons in the model to generate a prediction distribution containing uncertainty estimates, thus obtaining a confidence interval vector. Finally, the risk evolution trend is calculated to generate a risk evolution vector with confidence intervals.
[0067] The model parameter distribution set obtained based on the Monte Carlo method is used to apply dropout technology to the prediction model corresponding to each set of parameters, randomly deactivating some neurons during the inference stage. For example, some neurons in the convolutional layer and attention layer are temporarily turned off according to a preset probability to simulate the uncertainty of the model structure. Then, the same input data (such as time series data of key variables) is input into multiple sub-models after parameter sampling and neuron deactivation. After multiple runs, a set of SFTS propagation risk prediction results are obtained. The discrete distribution of these results forms the prediction distribution containing uncertainty estimates. Statistical analysis is performed on this prediction distribution to determine the upper and lower fluctuation range of the risk prediction value at each time step, resulting in a confidence interval vector. Then, the time dimension is combined to sort out the change law of risk over time, such as the short-term risk rising sharply, the medium-term risk stabilizing, and the long-term risk declining trend. The confidence interval vector and the risk evolution trend are deeply integrated to finally generate a risk evolution vector with confidence intervals. Each time node in the vector contains both the risk prediction value of that node and the corresponding uncertainty interval, clearly reflecting the risk change trend and the reliability of the prediction.
[0068] Specifically, by integrating a prediction architecture with temporal convolutional networks and attention mechanisms, and combining it with Monte Carlo dropout technology, the problem that traditional prediction models struggle to balance short-term fluctuations and long-term trends, and lack uncertainty quantification, is effectively solved. This enables multi-scale accurate prediction of SFTS propagation risk, generates risk evolution vectors containing confidence intervals, significantly improves the reliability and practicality of prediction results, and provides a scientific basis for phased and differentiated prevention and control decisions.
[0069] S5. Quantitatively analyze the contribution of each component in the risk evolution vector through SHAP value. When specific components such as population flow are identified as the dominant factors, automatically start the dynamic weight adjustment algorithm to generate regional risk heat maps that mark the main sources of risk.
[0070] Further, step S5 specifically includes:
[0071] Obtain risk evolution vector data, calculate the contribution of each component through SHAP value analysis, and obtain the contribution distribution results. If the contribution of the population flow component is higher than the preset threshold, extract the feature data of the component, determine the dominant risk factor, and update the weight allocation rule using a dynamic weight adjustment algorithm based on the dominant risk factor to obtain the adjusted weight parameters.
[0072] A regional risk heat map is generated by adjusting the weight parameters to obtain preliminary heat map data. A spatial smoothing algorithm is then applied to optimize the boundary area to obtain a smoothed heat map. Finally, by overlaying geographic information data, a regional risk heat map with the risk sources marked is generated.
[0073] After obtaining risk evolution vector data with confidence intervals, the contribution of each component is quantified using the SHAP (SHapley Additive ex Planations) value analysis method. This includes factors such as population movement, temperature and humidity, tick density, viral mutation, and clinical symptom association. The SHAP value is calculated based on the association between each component and the time step and regional risk value in the risk evolution vector. It calculates the specific contribution of each component to the risk value at different spatiotemporal nodes. Positive values indicate that the component promotes the increase of risk, while negative values indicate the inhibition of risk. The average contribution intensity, contribution ratio, and spatiotemporal distribution characteristics of each component are statistically integrated. For example, the contribution of a certain component is more significant in high-risk areas. Finally, a contribution distribution result that can intuitively reflect the degree of influence of each component on risk evolution is formed, such as a list of components sorted by contribution size and a thermal distribution of the contribution ratio of components in different regions.
[0074] Specifically, by quantifying the contribution of each component in the risk evolution vector through SHAP values, the problem of traditional risk assessment methods being unable to accurately identify the dominant risk factors and their spatial impact patterns is effectively solved. Combined with dynamic weight adjustment and spatial smoothing techniques, regional heat maps labeled with the main risk sources are generated, realizing the transformation from abstract risk prediction to spatial visualization of specific risk sources. This provides intuitive and reliable decision support for accurately locating high-risk areas and formulating targeted prevention and control measures.
[0075] S6. Based on the dominant risk sources marked on the risk heat map, generate an interpretable early warning signal sequence containing risk level, spatial location and dominant cause from the intelligent template library.
[0076] Further, step S6 specifically includes:
[0077] The dominant risk sources and their corresponding risk distributions are extracted from the risk heat map data. The dominant risk sources are analyzed through risk assessment algorithms to determine the risk level and dominant cause. A smart template library is used for template matching to generate a warning signal sequence containing the risk level, spatial location and dominant cause.
[0078] If the risk level exceeds the preset threshold, a location identifier is generated based on the spatial location to mark the risk distribution. Through the signal generation process, the template matching result is converted into an interpretable early warning signal sequence, generating a spatial risk mapping that includes the dominant cause.
[0079] The process of generating a spatial risk map that includes the dominant cause involves: identifying the existing structured data such as risk level, spatial location, and dominant cause in the template matching results; then initiating a signal conversion process to decompose the structured data into interpretable units consisting of spatiotemporal nodes, risk attributes, and cause descriptions; organizing these units according to temporal logic or risk diffusion order; and supplementing each sequence node with a simplified description to form a warning signal sequence, such as explaining the current risk level, core cause, and future short-term risk trend of a certain area. When generating the spatial risk map, the spatial location in the warning signal sequence is linked to a geographic information database, and risk areas are marked on the map with corresponding color blocks. A floating label is added to each color block, clearly indicating the dominant cause. Arrows are added to the map in conjunction with the risk diffusion trend, and the map legend explains the correspondence between risk level, dominant cause, and color / symbol. This ultimately forms a spatial risk map that combines geospatial visualization with explicit dominant cause identification, making the location, cause, and diffusion direction of the risk clearly identifiable.
[0080] Specifically, through intelligent template library matching and interpretable signal generation technology, the problem of high formatting and poor readability of traditional early warning information is effectively solved. The dominant risk sources in the risk heat map are transformed into early warning signal sequences that include risk level, spatial location and dominant cause, and a visual mapping that integrates geographic information and risk causes is generated. This realizes the leap from data-driven early warning to knowledge-guided decision-making, and significantly improves the operability of early warning information and the accuracy of prevention and control response.
[0081] S7. Based on the early warning signal sequence, if the signal sequence indicates an increase in risk level, a real-time notification mechanism is triggered to push the correlation details to the prevention and control system, thereby obtaining the final accurate regional risk prediction output.
[0082] Further, step S7 specifically includes:
[0083] The system acquires early warning signal sequences, analyzes these sequences to determine if the risk level has increased, and if so, triggers a real-time notification mechanism to generate notification data. It then extracts correlation details from the notification data and pushes them to the prevention and control system.
[0084] Based on the details of the push notifications, a logistic regression algorithm is used to assess regional risks and obtain a risk probability distribution. High-risk area features are extracted from the risk probability distribution to generate regional risk predictions. By combining the regional risk predictions with preset thresholds, the final accurate regional risk prediction is determined.
[0085] Specifically, by monitoring early warning signal sequences in real time and establishing an automatic triggering mechanism, the problem of delayed risk response and low information transmission efficiency in the traditional prevention and control system has been effectively solved. It can immediately push key related details to the prevention and control system when the risk level rises, and generate accurate regional risk predictions by combining logistic regression algorithms. This has realized the transformation from passive alarm reception to proactive intervention, and significantly improved the response speed and decision-making accuracy of epidemic prevention and control.
[0086] Example 2
[0087] Please see Figure 4 This embodiment provides a dynamic monitoring and early warning system for SFTS based on multimodal data, applied to a dynamic monitoring and early warning method for SFTS based on multimodal data, including:
[0088] The data augmentation module is used to collect environmental, clinical, viral, and population trajectory data from sensors and databases. After standardization, it uses an attention mechanism to filter key features and combines K-means clustering and trajectory matching to construct an augmented dataset.
[0089] The graph construction module divides the augmented dataset into multiple types of nodes according to modality, models semantics through graph neural networks, calculates the association strength using temporal attention if time series information is included, analyzes the coupling relationship between the environment and virus transmission, and generates a multi-scale association graph with time dimension.
[0090] The causal identification module, based on the association graph, extracts the initial causal path using the conditional independence test, and then determines the key driving factors by calculating the path weights through gradient boosting tree; if the environmental or viral mutation weights exceed the limit, specific variables are extracted and a mechanism model is constructed to generate a causal path graph.
[0091] The risk prediction module combines causal analysis results with a temporal convolutional network to highlight long-term trends using short-term fluctuation characteristics and an attention mechanism, and integrates these to construct a prediction model. If the output deviation exceeds the standard, the uncertainty is quantified through Monte Carlo sampling and dropout techniques to generate a risk evolution vector with confidence intervals.
[0092] The heatmap generation module performs SHAP analysis on the risk evolution vector. If the contribution of population movement exceeds the standard, it is identified as a core risk. The parameters are updated using a dynamic weight adjustment algorithm to generate a preliminary heatmap. After smoothing optimization and overlaying with geographic information, a final heatmap with labeled risk sources is obtained.
[0093] The early warning push module extracts risk sources from heat maps, uses algorithms to determine risk levels and causes, and generates early warning sequences by matching template libraries. If the risk increases, a notification is triggered, pushing relevant details to the prevention and control system. Logistic regression is then used to assess regional risks and output accurate prediction results.
[0094] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for dynamic monitoring and early warning of SFTS based on multimodal data, characterized in that: Includes the following steps: S1. Standardize environmental monitoring data, clinical symptom reports, viral gene sequences and population movement trajectories through sensor networks and database interfaces. Employ a feature weighting method based on attention mechanism to strengthen key features strongly correlated with SFTS transmission and establish an enhanced dataset with spatiotemporal consistency. S2. A graph neural network is used to perform semantic integration processing on the augmented dataset, and different modal data are modeled as multi-type nodes. Through the temporal graph attention mechanism, the correlation strength between nodes is dynamically learned, especially the coupling relationship between environmental changes and virus transmission paths, to generate a multi-scale correlation graph with time dimension. S3. Identify causal paths between environmental factors, viral mutations, and case reports from multi-dimensional correlation maps, and use a combination of conditional independence tests and gradient boosting trees to determine key driving factors and their mechanisms of action. S4. Based on the causal analysis results, a prediction model integrating temporal convolutional networks and attention mechanisms is established to capture both short-term fluctuations and long-term trends. Monte Carlo dropout technology is used to quantify prediction uncertainty and generate a risk evolution vector containing confidence intervals. S5. Quantitatively analyze the contribution of each component in the risk evolution vector through SHAP value. When a specific component of population flow is identified as the dominant factor, automatically start the dynamic weight adjustment algorithm to generate a regional risk heat map that marks the main sources of risk. S6. Based on the dominant risk sources marked by risk heatmaps, generate an interpretable early warning signal sequence containing risk level, spatial location and dominant cause from the intelligent template library; S7. Based on the early warning signal sequence, if the signal sequence indicates an increase in risk level, a real-time notification mechanism is triggered to push the correlation details to the prevention and control system, thereby obtaining the final regional risk prediction output.
2. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S1 specifically includes: Acquire temperature and humidity data, tick density, fever level, platelet count, mutation sites, evolutionary branches, and movement trajectories from sensor network and database interfaces. Normalize the data using standardized processing methods to obtain an initial dataset in a unified format. The temperature and humidity data, tick density, fever level, platelet count, mutation sites, evolutionary branches and movement trajectories in the initial dataset are weighted using an attention mechanism to determine the weights of key features related to SFTS transmission. If the weight of a key feature is higher than a preset threshold, the feature is retained and its spatiotemporal information is extracted to obtain a feature set containing spatiotemporal labels. The K-means algorithm is used to cluster the feature set to obtain data groups with spatiotemporal consistency. Based on the data groups, temperature and humidity data, tick density, fever level, platelet count, mutation sites and evolutionary branches are fused to construct an enhanced dataset. By matching the movement trajectory with the spatiotemporal consistency of the augmented dataset, the correlation between the trajectory and SFTS propagation is determined, resulting in an augmented dataset with spatiotemporal consistency.
3. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S2 specifically includes: S21. Divide the augmented dataset into multiple types of nodes according to modality, and use graph neural networks to perform semantic modeling on the multiple types of nodes to generate initial node embedding representations. If the initial node embedding representations contain time series information, calculate the dynamic association strength between nodes through a temporal attention mechanism to obtain a weighted association matrix. S22. Based on the weighted correlation matrix, analyze the interaction patterns between environmental change nodes and virus transmission nodes, generate a coupling relationship representation, and combine the coupling relationship representation with time dimension information through multi-scale graph construction technology to generate a multi-scale correlation graph.
4. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S3 specifically includes: Data from multi-dimensional correlation graphs were obtained, and conditional independence tests were used to extract significant data correlations from environmental factors, viral mutations, and case reports to obtain an initial set of causal paths. Then, gradient boosting trees were used to calculate the weights of each path and determine the key driving factors. If the path weight of environmental factors in the key driving factors is higher than the preset threshold, then temperature, humidity and pollution index are extracted from the environmental factors to obtain specific environmental driving variables. If the proportion of viral mutations in the key driving factors is higher than the preset threshold, then gene sequence changes and mutation rates are extracted from viral mutations to obtain specific mutation driving variables. Based on key driving factors and specific driving variables, the mechanism of action of causal paths is constructed to obtain a mechanism description model. Then, the interaction between environmental driving variables and variation driving variables is analyzed to determine the interaction path weights, generate the dynamic influence path of key driving factors on case reports, and obtain the final causal path diagram.
5. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S4 specifically includes: S41. Obtain the causal analysis results, determine the causal impact of key variables on the time series, obtain the causal weight set, process the causal weight set through a temporal convolutional network, extract the short-term fluctuation characteristics in the time series, and obtain the short-term fluctuation vector. S42. Employ an attention mechanism to weight the short-term fluctuation vector, highlighting the long-term trend information at key time steps, to obtain the long-term trend vector. Integrate the short-term fluctuation vector and the long-term trend vector to construct a prediction model and generate initial prediction results. S43. If the output deviation of the prediction model exceeds the preset threshold, the model parameters are sampled multiple times using the Monte Carlo method to obtain the parameter distribution set. The dropout technique is then applied to randomly deactivate the neurons in the model to generate a prediction distribution containing uncertainty estimates, thus obtaining a confidence interval vector. Finally, the risk evolution trend is calculated to generate a risk evolution vector with confidence intervals.
6. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S5 specifically includes: Obtain risk evolution vector data, calculate the contribution of each component through SHAP value analysis, and obtain the contribution distribution results. If the contribution of the population flow component is higher than the preset threshold, extract the feature data of the component, determine the dominant risk factor, and update the weight allocation rule using a dynamic weight adjustment algorithm based on the dominant risk factor to obtain the adjusted weight parameters. A regional risk heat map is generated by adjusting the weight parameters to obtain preliminary heat map data. A spatial smoothing algorithm is then applied to optimize the boundary area to obtain a smoothed heat map. Finally, by overlaying geographic information data, a regional risk heat map with the risk sources marked is generated.
7. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S6 specifically includes: The dominant risk sources and their corresponding risk distributions are extracted from the risk heat map data. The dominant risk sources are analyzed through risk assessment algorithms to determine the risk level and dominant cause. A smart template library is used for template matching to generate a warning signal sequence containing the risk level, spatial location and dominant cause. If the risk level exceeds the preset threshold, a location identifier is generated based on the spatial location to mark the risk distribution. Through the signal generation process, the template matching result is converted into an interpretable early warning signal sequence, generating a spatial risk mapping that includes the dominant cause.
8. The SFTS dynamic monitoring and early warning method based on multimodal data according to claim 1, characterized in that: Step S7 specifically includes: The system acquires early warning signal sequences, analyzes these sequences to determine if the risk level has increased, and if so, triggers a real-time notification mechanism to generate notification data. It then extracts correlation details from the notification data and pushes them to the prevention and control system. Based on the details of the push notifications, a logistic regression algorithm is used to assess regional risks and obtain a risk probability distribution. High-risk area features are extracted from the risk probability distribution to generate regional risk predictions. By combining the regional risk predictions with preset thresholds, the final accurate regional risk prediction is determined.
9. A dynamic monitoring and early warning system for SFTS based on multimodal data, applied to the dynamic monitoring and early warning method for SFTS based on multimodal data as described in any one of claims 1-8, characterized in that: include: The data augmentation module is used to collect environmental, clinical, viral, and population trajectory data from sensors and databases. After standardization, it uses an attention mechanism to filter key features and combines K-means clustering and trajectory matching to construct an augmented dataset. The graph construction module divides the augmented dataset into multiple types of nodes according to modality, models semantics through graph neural networks, calculates the association strength using temporal attention if time series information is included, analyzes the coupling relationship between the environment and virus transmission, and generates a multi-scale association graph with time dimension. The causal identification module, based on the association graph, extracts the initial causal path using the conditional independence test, and then determines the key driving factors by calculating the path weights through gradient boosting tree. If the environmental or viral mutation weights exceed the limit, extract specific variables and construct a mechanism model to generate a causal path diagram. The risk prediction module combines the results of causal analysis with temporal convolutional networks to highlight long-term trends by using short-term fluctuation characteristics and attention mechanisms, and integrates them to build a prediction model. If the output deviation exceeds the limit, the uncertainty is quantified by Monte Carlo sampling and dropout techniques to generate a risk evolution vector with confidence intervals; The heatmap generation module performs SHAP analysis on the risk evolution vector. If the contribution of population movement exceeds the standard, it is identified as a core risk. The parameters are updated using a dynamic weight adjustment algorithm to generate a preliminary heatmap. After smoothing optimization and overlaying with geographic information, a final heatmap with labeled risk sources is obtained. The early warning push module extracts risk sources from heat maps, uses algorithms to determine risk levels and causes, and generates early warning sequences by matching template libraries. If the risk increases, a notification is triggered, pushing relevant details to the prevention and control system. Logistic regression is then used to assess regional risks and output accurate prediction results.