An influenza early warning method and system based on crowd stratification multi-source data fusion
Patent Information
- Application Number
- CN202610848406.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-04
AI Technical Summary
针对现有技术的不足,本发明提供了一种基于人群分层多源数据融合的流感早期预警方法及系统,具备信号与人群精准匹配、抗年龄结构分布漂移能力强、多时域预测稳定且预警假阳率低等优点,解决了现有流感早期预警技术中人群聚合建模引入大量噪声、疫情后流感年龄结构漂移导致系统性低估、直接拟合绝对值预测性能衰减快且固定阈值预警造成公共卫生资源浪费的问题
1、该基于人群分层多源数据融合的流感早期预警方法及系统,先将流感相关搜索关键词划分为多个语义类别,针对不同年龄段人群的行为特征匹配对应的语义信号,特别为无法独立进行信息检索的低龄儿童匹配家长代理查询类信号,再分年龄层独立建模,仅使用对应年龄层自身的自回归特征,避免跨年龄层信息干扰,精准匹配不同人群的流感活动信号,有效提升整体预测精度。
Smart Images

Figure CN122696401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of influenza early warning technology, specifically to an influenza early warning method and system based on population-stratified multi-source data fusion. Background Technology
[0002] Seasonal influenza is a persistent and significant public health threat worldwide, causing approximately 290,000 to 650,000 respiratory-related deaths and millions of severe cases annually. Influenza-like illness (ILI) surveillance is a core component of the influenza prevention and control system. By identifying early trends in influenza activity, it can provide a scientific basis for vaccination, allocation of medical resources, and public health education. Traditional ILI surveillance mainly relies on weekly reports from sentinel hospitals in the national influenza surveillance network, which has an inherent time lag of 1 to 2 weeks, making it difficult to meet the emergency response needs of public health emergencies.
[0003] With the development of digital epidemiology, influenza prediction methods based on multi-source data fusion have gradually become a research hotspot. Google Flu Trends first demonstrated a strong correlation between search engine query data and influenza activity. Subsequently, the ARGO framework further improved prediction accuracy by integrating search engine, meteorological, and electronic health record data, and shortened the monitoring lag to some extent.
[0004] However, existing influenza early warning technologies still have the following problems that urgently need to be addressed: First, population aggregation modeling leads to a mismatch between signals and populations. The correlation between search behavior and influenza activity varies across different age groups. Hybrid modeling introduces a lot of noise and reduces prediction accuracy. In particular, children aged 0-14 cannot independently retrieve information, and existing technologies cannot effectively capture influenza activity signals in this population. Second, after the COVID-19 pandemic, the characteristics of influenza epidemics have changed significantly, showing an age structure shift dominated by school-aged children aged 5 to 14. In some regions, the proportion of ILI in this age group has increased from 16.3% before the pandemic to 54.7%. Under this distribution shift, existing population-blind models will systematically underestimate the intensity of influenza activity. Third, existing multi-time-domain prediction methods directly fit the absolute ILI% value, which is easily affected by long-term trend terms and shows significant performance degradation when the prediction is made more than 2 weeks in advance. Furthermore, they generally use fixed warning thresholds, resulting in a high false positive rate in different regions, different epidemic seasons, and different vaccination rates, which can easily lead to unnecessary mobilization of public health resources. Therefore, a method and system for early warning of influenza based on multi-source data fusion in a population-stratified manner is proposed. Summary of the Invention
[0005] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides an early influenza warning method and system based on population-stratified multi-source data fusion. It has advantages such as accurate matching of signals and populations, strong resistance to age structure distribution drift, stable multi-time domain prediction, and low false positive rate. It solves the problems of existing early influenza warning technologies, such as the introduction of a large amount of noise in population aggregation modeling, systematic underestimation caused by post-epidemic influenza age structure drift, rapid performance decay of direct fitting absolute value prediction, and waste of public health resources caused by fixed threshold warning.
[0006] (II) Technical Solution To achieve the goals of accurate matching of signals with populations, strong resistance to age-related structural distribution drift, stable multi-temporal prediction, and low false positive rate in early warning, this invention provides the following technical solution: an early warning method for influenza based on population-stratified multi-source data fusion, comprising the following steps: S1: Multi-source data collection and alignment with the ISO weekly calendar. Collect ILI monitoring data from sentinel hospitals by age group, meteorological data, internet search index data, and seasonal influenza vaccine coverage data. All data are uniformly aligned to the ISO standard weekly calendar. S2: Semantic classification and feature dimensionality reduction of search keywords. The flu-related keywords in the Internet search index are divided into multiple semantic categories. After normalizing and fusing the common keywords across platforms, the principal components of each semantic category are extracted as composite features. S3: Hierarchical feature selection based on parental agent behavior. For the 0-14 age group, parental agent query type, influenza specific type and symptom type signals are matched to represent children's influenza activity using parental agent retrieval behavior; for the ≥15 age group, influenza specific type, symptom type and treatment type signals are matched; the model for each age group only uses the ILI% lag term of the age group itself as an autoregressive feature. S4: Feature engineering, constructing multi-order lag versions of each composite feature, calculating the first-order difference between ILI% and the week-on-week ratio of the core composite feature, and using sine and cosine pairing to encode the seasonality of the week calendar. S5: Age-stratified Ridge regression ensemble modeling, fitting a Ridge regression sub-model for each age stratum, using L2 regularization to suppress overfitting caused by the sparsity of age-stratified data, and weighting and aggregating the predicted values of multiple sub-models according to the proportion of the resident population in the corresponding age stratum to obtain the real-time estimate of the population ILI%. S6: Multi-time domain prediction based on Δ target. The Δ target strategy is adopted for 1-4 week advance prediction. The weekly change Δ of ILI% is defined as the prediction target. The model is trained to predict the Δ value and reconstruct the absolute prediction value to eliminate the influence of the time series trend term and improve the stability of the 1-4 week advance prediction. S7: Threshold-based early warning decision, triggers corresponding level early warnings based on early warning thresholds, and simultaneously calculates and outputs the sensitivity, specificity, positive predictive value, and F1 score of the early warning.
[0007] Preferably, in step S5, the loss function of the regression sub-model is: ; In the formula Here is the regularization coefficient, in A 5-fold cross-validation grid search is performed on the set, with the optimization objective being to minimize the mean absolute error of the validation set.
[0008] Preferably, in step S6, the Δ-Ridge predicted value and the seasonal baseline predicted value based on the historical average for the same period are weighted according to adaptive weights. Perform weighted fusion: ; Weight The weighting is dynamically adjusted based on influenza activity: when influenza activity is stable, the weight of the seasonal baseline forecast is increased; when the Δ value changes for two consecutive weeks by more than 1.5 times the standard deviation for the same period in history, the weight of the Δ-Ridge forecast is automatically increased. The range of values is The initial values are determined by a grid search on the validation set, and the optimization objective is the validation set determination coefficient. maximum.
[0009] Preferably, in step S7, a dynamic early warning threshold adjustment mechanism is adopted, which dynamically calculates the early warning threshold based on historical ILI data for the same period and the current vaccination rate: ; In the formula This is the average of the ILI% for the same period over the past 5 years. This represents the standard deviation of ILI% over the same period over the past 5 years. This is the threshold coefficient (default value is 1.5). The vaccination rate adjustment coefficient is used, and a segmented adjustment method is adopted: when the vaccination rate is below 30%, Take 0.7; when the vaccination rate is higher than or equal to 30%, Take 0.5; The current influenza vaccination rate; when the predicted population ILI% exceeds An orange alert is triggered when the predicted value exceeds [a certain threshold]. A red alert was triggered at that time.
[0010] An early warning system for influenza based on population-stratified multi-source data fusion includes: Data acquisition module: Used to automatically collect outpatient data by age group from sentinel hospitals, meteorological data, internet search index data, and vaccination data through API interface, and perform ISO calendar alignment and missing value filling; Data preprocessing module: used to perform semantic classification of search keywords, cross-platform index normalization fusion, principal component dimensionality reduction, hierarchical feature selection and outlier cleaning operations; Real-time estimation module: used to perform age-stratified Ridge regression ensemble modeling and output real-time estimates of the population ILI%. Multi-time-domain prediction module: used to perform multi-time-domain prediction based on the Δ target, outputting the predicted population ILI% value for the next 1 to 4 weeks; Early warning decision module: used to trigger corresponding level early warnings based on dynamic early warning thresholds and calculate early warning evaluation indicators.
[0011] Preferably, the internet search index data collected by the data acquisition module includes search engine indexes and short video platform indexes; the specific method for cross-platform normalization fusion is as follows: For keywords shared by search engines and short video platforms, the weekly index values of the two platforms are min-max normalized to the [0,1] interval, and then the corresponding values are added together to obtain the fusion index value.
[0012] Preferably, it also includes a model update and monitoring module, which is used to incrementally fine-tune the model weekly using newly added real ILI data. When the model prediction error exceeds a preset threshold for four consecutive weeks, a full retraining process is triggered. At the same time, the data quality and model performance indicators of each data source are monitored in real time, and a model health report is generated.
[0013] Preferably, the early warning decision module is also used to generate a visual early warning report, displaying the real-time ILI% trend, the prevalence of different age groups, the prediction curve and early warning information, and supporting multi-dimensional filtering and querying by administrative region, age group and time range.
[0014] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements an early warning method for influenza based on population-stratified multi-source data fusion.
[0015] (III) Beneficial Effects Compared with existing technologies, this invention provides an early warning method and system for influenza based on population-stratified multi-source data fusion, which has the following beneficial effects: 1. This influenza early warning method and system based on population-stratified multi-source data fusion first divides influenza-related search keywords into multiple semantic categories, matches corresponding semantic signals to the behavioral characteristics of different age groups, and specifically matches parent proxy query signals for young children who cannot independently retrieve information. Then, it models independently by age group, using only the autoregressive features of the corresponding age group to avoid cross-age information interference, accurately matches influenza activity signals of different groups, and effectively improves the overall prediction accuracy.
[0016] 2. This influenza early warning method and system based on population stratification and multi-source data fusion fits regression models for each age group. Each model only learns the influenza transmission pattern of the corresponding population. Regularization is used to suppress overfitting caused by the sparsity of age-stratified data. The prediction results of each age group are then weighted and aggregated according to the proportion of the resident population of the corresponding age group to obtain the overall estimate. It can adapt to the changes in the epidemic intensity of different age groups and resist the systematic bias caused by age structure drift.
[0017] 3. This influenza early warning method and system based on population stratification and multi-source data fusion eliminates the influence of time series trend terms by predicting the weekly change in the percentage of influenza-like cases. Then, the predicted change is dynamically weighted and fused with the seasonal baseline prediction value according to the influenza activity status. At the same time, the warning threshold is dynamically adjusted by combining historical data and vaccination rate, which significantly improves the stability of multi-time domain prediction and effectively reduces the false positive rate of warnings. Attached Figure Description
[0018] Figure 1 This is an overall flowchart of the influenza early warning method of the present invention; Figure 2 This is a schematic diagram of the influenza early warning system of the present invention; Figure 3 This is a structural diagram of the age-stratified Ridge regression ensemble model of the present invention; Figure 4 This is a structural diagram of the multi-temporal prediction model based on the Δ target of the present invention. Detailed Implementation
[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1: This embodiment describes in detail the sources of multi-source data collection, field definitions, and standardization alignment methods, providing a high-quality data foundation for subsequent modeling.
[0021] The data acquisition module automatically acquires the following five types of data through a standardized API interface: 1. Sentinel Hospital ILI Surveillance Data: Data was collected from the HIS system of sentinel hospitals, and the number of ILI cases and total outpatient visits were recorded daily for five age groups: 0-4 years, 5-14 years, 15-24 years, 25-59 years, and ≥60 years. ILI was defined as a patient with a body temperature ≥38℃ and accompanied by one of the symptoms of cough or sore throat. 2. Meteorological data: Daily average temperature, daily average absolute humidity, and daily average relative humidity were obtained from the China Meteorological Data Network and summarized into weekly averages; 3. Search Engine Index: Weekly search indexes for 25 flu-related keywords obtained from search engine index platforms. Keywords include: flu, influenza A, influenza B, influenza A, influenza B, flu symptoms, fever, cough, sore throat, headache, muscle aches, oseltamivir, Tamiflu, antipyretics, cough suppressants, cold medicine, pneumonia, bronchitis, complications, fever in children, cold in babies, flu in children, flu in pregnant women, flu in the elderly, and flu vaccine. 4. Short Video Platform Index: Weekly search indexes for 19 flu-related keywords obtained from short video platform index platforms. Keywords include: flu, influenza A, influenza B, influenza A, influenza B, flu symptoms, fever, cough, sore throat, oseltamivir, antipyretics, cough medicine, cold medicine, pneumonia, what to do if a child has a fever, baby cold, childhood flu, elderly flu, flu vaccine; 5. Vaccine coverage data: Weekly first-dose influenza vaccination rates by age group were obtained from the immunization program information system.
[0022] All data is aligned to the ISO standard weekly calendar, which means the first week of each year includes January 4th, and each week runs from Monday to Sunday. For missing values, linear interpolation is used to fill them in. Data that is missing for more than 3 consecutive weeks is marked as abnormal data and triggers a manual review process.
[0023] Example 2: This embodiment details the semantic classification criteria for keywords, cross-platform fusion methods, and hierarchical feature selection rules to achieve accurate signal matching at the data layer.
[0024] First, the 44 flu-related keywords were divided into five categories based on their semantics: 1. Specific types of influenza: Influenza A, Influenza B, Type A Influenza, Type B Influenza in children, Influenza in pregnant women, and Influenza in the elderly; 2. Symptoms: Flu symptoms, fever, cough, sore throat, headache, muscle aches; 3. Treatment medications: Oseltamivir, Tamiflu, antipyretics, cough suppressants, and cold medicines; 4. Complications: Pneumonia, bronchitis, and other complications; 5. Parental assistance inquiry category: Child has a fever, baby has a cold, what to do if child has a fever.
[0025] Cross-platform normalization fusion was performed on 19 keywords shared by search engines and short video platforms: the weekly index values of the two platforms were min-max normalized to the [0,1] interval, and then the corresponding values were added to obtain the fusion index value; principal component analysis was performed on each semantic category with parameters set to center and standardization, retaining the first principal component, and requiring the cumulative variance contribution rate to be no less than 75%.
[0026] Stratified feature selection based on the behavioral characteristics of different age groups: Children aged 0-4 and 5-14 are unable to independently retrieve information, and their flu-related searches are 100% completed by their parents. In the post-COVID-19 era, parents' attention to children's health has significantly increased, and the number of searches related to children's fever has increased 3.2 times compared to before the pandemic. Therefore, matching parent-assisted query, flu-specific, and symptom-related signals to this age group can accurately capture early signs of flu activity in children. Adolescents and young adults aged 15-24 are primarily concerned with influenza itself and its symptoms, therefore matching influenza-specific and symptom-related signals; Adults aged 25-59 are more concerned about treatment drugs after experiencing flu symptoms, thus matching treatment signals; People aged 60 and older have weaker immune systems and are more concerned about potential complications and treatments from influenza. Therefore, signals related to complications and treatments are matched.
[0027] Each age group model uses only the ILI% lagged 1-week term of its own age group as an autoregressive feature, avoiding model overfitting caused by cross-age group information leakage, while eliminating noise caused by differences in the popularity rhythm of different age groups.
[0028] Example 3: This embodiment details the construction, training, and ensemble methods of age-stratified Ridge regression models, and verifies the model's prediction accuracy and resistance to distribution drift through comparative experiments.
[0029] The input features for each age group include: The 0-week and 1-week lag values of the first principal component of the matched semantic category; The ILI% value of this age group lagged by 1 week; Weekly average temperature, weekly average absolute humidity, and weekly average relative humidity; Features of the sine and cosine coding of the weekly calendar: ; ; The weekly vaccination rate for this age group lagged by 2 and 4 weeks, taking into account that it takes 2-4 weeks for antibodies to develop after vaccination; ILI% and the first-order difference of the circumferential ratio of the three strongest composite features.
[0030] The week-on-week first-order difference feature can effectively capture the "acceleration" feature of influenza transmission, that is, the rate of increase or decrease of influenza activity. When combined with the Ridge regression model, this feature can significantly improve the technical problem of "trend prediction lag" in traditional methods. Compared with the year-on-year feature, the week-on-week difference is more sensitive to changes in short-term epidemic trends. Compared with the moving average feature, the week-on-week difference can retain more instantaneous change information.
[0031] The mathematical expression for Ridge regression is: ; The loss function is: ; In the formula This is the regularization coefficient, used to control model complexity and prevent overfitting.
[0032] Age-group modeling reduces the training sample size for each sub-model, making it prone to overfitting. Ridge regression, by introducing an L2 regularization term to penalize larger regression coefficients, can effectively suppress overfitting caused by data sparsity, enabling the model to maintain good generalization ability even with small sample sizes. This is the core reason why this invention chooses Ridge regression instead of ordinary linear regression or other complex models.
[0033] Regularization coefficient exist A 5-fold cross-validation grid search is performed on the dataset, with the optimization objective being to minimize the mean absolute error of the validation set. Data from 2018 to 2023 is used as the training set, data from 2024 to 2025 is used as the validation set, and data from 2025 to 2026 is used as the test set.
[0034] After the five age-group sub-models were trained, their predicted values were weighted and summed according to the proportion of the permanent residents of City A by age group in the corresponding year as published by the National Bureau of Statistics, to obtain the real-time estimate of the group ILI%. ; In the formula For the first The percentage of permanent residents in each age group For the first Predicted ILI% values for each age group.
[0035] Experiment 1: To verify the prediction accuracy of the model of this invention, it was compared with mainstream machine learning models and baseline models that did not perform signal-to-population matching. The experimental results are shown in the table below: ; As can be seen from the above experiments, the model of this invention... It achieved the highest signal-to-population (MAE) and the lowest MAE, demonstrating significantly better prediction accuracy than other mainstream models. Compared to the Ridge model, which did not perform signal-population matching... The performance was improved by 12.4%, and the MAE was reduced by 68.3%, proving that the hierarchical feature selection mechanism based on parent agent behavior can effectively improve model performance.
[0036] Experiment 2: To verify the stability of the model when influenza epidemic characteristics change, the performance of the model of this invention and the blind baseline model of the population were compared before and after age structure drift. The experimental results are shown in the table below: ; The above experiments show that under the condition of drastic distributional drift in the 5-14 year old age group, where the proportion of ILI increased from 16.3% to 54.7%, the performance of the model of this invention only slightly decreased, while the performance of the blind baseline model of the population was greatly degraded; this proves that the age-stratified modeling strategy can effectively resist the influence of age structure drift and has stronger robustness.
[0037] Experiment 3: To verify the effectiveness of cross-platform data fusion, the performance of models based on a single data source and dual-platform fusion was compared. The experimental results are shown in the table below: ; The above experiments show that after integrating data from both search engines and short video platforms, the model... Compared to using only search engines, the index increased by 8.5%, and the MAE decreased by 59.6%, proving that cross-platform data fusion can more comprehensively capture changes in public health behavior and significantly improve prediction accuracy.
[0038] Example 4: This embodiment describes in detail the principle of the Δ target strategy, the adaptive weighted fusion method, and the segmented dynamic threshold warning rule, and verifies the multi-time domain prediction performance and warning accuracy through experiments.
[0039] Traditional direct prediction methods directly fit The relationship with historical data is easily affected by long-term trends, leading to a significant degradation in multi-time-domain prediction performance. This invention employs a Δ-target strategy to improve prediction stability by predicting the weekly variation of ILI%. 1. For advance Zhou's prediction, definition ,in For the first Zhou's actual group ILI% 2. Train the Ridge regression model for prediction The input features are the same as those of the instantaneous estimation model; 3. Reconstructing absolute predictions: .
[0040] The Δ-target strategy transforms the non-stationary absolute ILI% time series into a stationary series of changes, eliminating the influence of the trend term on the forecast and significantly improving the stability of multi-time-domain forecasts. To further enhance forecast robustness, the Δ-Ridge forecast value and the seasonal baseline forecast value based on the historical average of the same period are adaptively weighted. Perform weighted fusion: ; Weight The weighting is dynamically adjusted based on influenza activity: when influenza activity is stable, the weight of the seasonal baseline forecast is increased; when the Δ value changes for two consecutive weeks by more than 1.5 times the standard deviation for the same period in history, the weight of the Δ-Ridge forecast is automatically increased. The range of values is Initial values are determined through a validation set grid search, and the optimization objective is the validation set. Maximum; Experimental results show that the initial optimal weight for prediction is 0.7 for week 1, 0.6 for week 2, 0.5 for week 3, and 0.4 for week 4.
[0041] This invention employs a segmented dynamic early warning threshold adjustment mechanism to replace the traditional fixed threshold, enabling it to adapt to the differences in epidemic characteristics across different regions, years, and influenza subtypes. Specifically, the early warning threshold is linked to the current influenza vaccination rate, using a segmented adjustment coefficient: when the vaccination rate is below 30%, the herd immunity barrier is weak, and influenza outbreaks are more likely, so the system uses a larger adjustment coefficient of 0.7 to tighten the early warning threshold; when the vaccination rate is above or equal to 30%, the system uses a smaller adjustment coefficient of 0.5 to appropriately relax the early warning threshold, reducing unnecessary early warnings. ; In the formula This is the average of the ILI% for the same period over the past 5 years. This represents the standard deviation of ILI% over the same period over the past 5 years. This is the threshold coefficient (default value is 1.5). This is the adjustment factor for vaccination rate. This represents the current influenza vaccination rate.
[0042] When the predicted population ILI% exceeds An orange alert is triggered when influenza activity enters its epidemic season; when the predicted value exceeds [a certain threshold], [the alert is triggered]. A red alert was triggered, indicating that influenza activity was entering its peak season.
[0043] Experiment 4: To verify the performance improvement of the Δ-target strategy on multi-time-domain prediction, the direct prediction method and the method of this invention were compared at different lead times. The experimental results are shown in the table below: ; As can be seen from the above experiments, the method of the present invention achieves the desired effect in all lead times. All were significantly higher than direct prediction methods, and the improvement gradually increased with the increase of the lead time; 4-week lead time prediction... The improvement of 18.3% proves that the Δ target strategy can effectively eliminate the influence of the trend term and significantly improve the stability of multi-time domain forecasts.
[0044] Experiment 5: To verify the effectiveness of the dynamic early warning threshold mechanism, the early warning performance of a fixed threshold and the dynamic threshold of this invention were compared. The experimental results are shown in the table below: ; The above experiments show that the dynamic threshold mechanism of this invention significantly improves the specificity of early warning while maintaining high sensitivity; the specificity of early warning at 1 week advance warning is increased by 12.0%, and the specificity of early warning at 4 weeks advance warning is increased by 22.4%, effectively reducing the false positive rate and reducing unnecessary mobilization of public health resources.
[0045] Example 5: This embodiment describes in detail the methods for detecting and cleaning outliers in multi-source data, as well as the methods for online model updates and performance monitoring, to ensure that the system can operate stably for a long time.
[0046] Internet search indexes are easily affected by non-disease factors, such as media reports, trending events, celebrity illnesses, and platform algorithm changes, which can cause sudden spikes or drops in search indexes that are unrelated to actual flu activity. This invention uses the sliding window standard deviation method for outlier detection: for each keyword's time series, the mean and standard deviation are calculated within a sliding window of 4 weeks; if the index value in a certain week exceeds the mean plus 3 times the standard deviation, it is marked as a spike outlier; if the index value in a certain week suddenly drops to zero but the corresponding keyword index on another platform is normal, it is marked as a zero-point outlier.
[0047] For confirmed sudden increases in outliers, the average value of the sample over the preceding and following two weeks is used for replacement; for confirmed zero outliers, the normalized index value of the corresponding keyword from another platform is used for filling; as a preferred solution, the accuracy of outlier detection can be further improved by combining the isolated forest algorithm. This part is disclosed in the specification and is not included in the scope of protection of the claims.
[0048] This invention establishes a two-level model update mechanism of "incremental fine-tuning + full retraining": 1. Incremental Fine-tuning: The system performs incremental fine-tuning of the model every Monday using the real ILI data added the previous week; during the fine-tuning process, most of the model's parameters are fixed, and only the regression coefficients of the last layer are updated, with the learning rate set to 0.01; incremental fine-tuning can quickly absorb the latest popular information and maintain the model's timeliness. 2. Full Retraining: When the average absolute error of the model exceeds the preset threshold (0.5% by default) for 4 consecutive weeks, the full retraining process is triggered. Full retraining uses all historical data to retrain the model and update the regularization coefficients and feature weights. Full retraining can solve the performance degradation problem after the model has been running for a long time.
[0049] Meanwhile, the system monitors the following performance metrics in real time: Prediction accuracy metrics: MAE, RMSE; Early warning performance indicators: sensitivity, specificity, and positive predictive value; Data quality metrics: missing rate, outlier rate.
[0050] The system generates a model health report weekly and automatically issues an alert when an indicator shows abnormalities, prompting administrators to intervene manually.
[0051] Example 6: This embodiment describes in detail the deployment architecture and operation process of the early warning system, and verifies the actual application effect of the system.
[0052] The system adopts a B / S architecture, which is divided into four layers: data layer, preprocessing layer, modeling layer, and application layer. 1. Data Layer: Stores raw data from multiple sources and preprocessed feature data, managed using a PostgreSQL relational database; 2. Preprocessing layer: Performs data cleaning, alignment, semantic classification, principal component dimensionality reduction, hierarchical feature selection, and outlier cleaning operations. It runs automatically every day at midnight. 3. Modeling layer: Loads pre-trained age-stratified Ridge and Δ-Ridge models, automatically performs incremental model fine-tuning every Monday, and generates prediction results for the next 1 to 4 weeks; 4. Application Layer: This includes an early warning and decision-making module, a model update and monitoring module, and a visualization module, providing decision support for staff at the municipal CDC.
[0053] The system is deployed on a Linux server configured with an Intel Xeon E5-2680 v4 CPU and 32GB of memory. The operating environment is Python 3.8, and the dependent libraries include scikit-learn, pandas, numpy, flask, etc. No GPU support is required. The system automatically collects data daily and generates the latest early warning report every Monday at 9:00 AM, which is pushed to the staff of the municipal CDC and the CDCs of various districts and counties via web interface and email.
[0054] In summary, this influenza early warning method and system based on population-stratified multi-source data fusion first divides influenza-related search keywords into multiple semantic categories, matches corresponding semantic signals to the behavioral characteristics of different age groups, and specifically matches parent proxy query signals for young children who cannot independently retrieve information. Then, it models independently for each age group, using only the autoregressive features of the corresponding age group to avoid cross-age information interference, accurately matches influenza activity signals of different populations, and effectively improves the overall prediction accuracy.
[0055] Furthermore, this influenza early warning method and system based on population stratification and multi-source data fusion fits regression models for each age group separately. Each model only learns the influenza transmission patterns of the corresponding population. Regularization is used to suppress overfitting caused by the sparsity of data in different age groups. The prediction results of each age group are then weighted and aggregated according to the proportion of the resident population in the corresponding age group to obtain the overall estimate. This can adapt to the changes in the epidemic intensity of different age groups and resist the systematic bias caused by age structure drift.
[0056] Furthermore, this influenza early warning method and system based on population-stratified multi-source data fusion eliminates the influence of time series trend terms by predicting the weekly changes in the percentage of influenza-like cases. It then dynamically weights and fuses the predicted changes with the seasonal baseline predictions according to the influenza activity status. Simultaneously, it dynamically adjusts the warning threshold by combining historical data from the same period and vaccination rates, significantly improving the stability of multi-time-domain predictions and effectively reducing the false positive rate. This solves the problems of existing influenza early warning technologies, such as the introduction of a large amount of noise in population aggregation modeling, systematic underestimation due to the drift of influenza age structure after the epidemic, rapid performance decay of direct fitting of absolute value predictions, and waste of public health resources caused by fixed threshold warnings.
[0057] The relevant modules involved in this system are all hardware system modules or functional modules that combine computer software programs or protocols with hardware in the prior art. The computer software programs or protocols involved in these functional modules are technologies known to those skilled in the art and are not improvements to this system. The improvement of this system lies in the interaction or connection between the modules, that is, in improving the overall structure of the system to solve the corresponding technical problems that this system aims to address.
[0058] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for early influenza warning based on population-stratified multi-source data fusion, characterized in that, Includes the following steps: S1: Multi-source data collection and alignment with the ISO weekly calendar. Collect ILI monitoring data from sentinel hospitals by age group, meteorological data, internet search index data, and seasonal influenza vaccine coverage data. All data are uniformly aligned to the ISO standard weekly calendar. S2: Semantic classification and feature dimensionality reduction of search keywords. The flu-related keywords in the Internet search index are divided into multiple semantic categories. After normalizing and fusing the common keywords across platforms, the principal components of each semantic category are extracted as composite features. S3: Hierarchical feature selection based on parental agent behavior. For the 0-14 age group, parental agent query type, influenza specific type and symptom type signals are matched to represent children's influenza activity using parental agent retrieval behavior; for the ≥15 age group, influenza specific type, symptom type and treatment type signals are matched; the model for each age group only uses the ILI% lag term of the age group itself as an autoregressive feature. S4: Feature engineering, constructing multi-order lag versions of each composite feature, calculating the first-order difference between ILI% and the week-on-week ratio of the core composite feature, and using sine and cosine pairing to encode the seasonality of the week calendar. S5: Age-stratified Ridge regression ensemble modeling, fitting a Ridge regression sub-model for each age stratum, using L2 regularization to suppress overfitting caused by the sparsity of age-stratified data, and weighting and aggregating the predicted values of multiple sub-models according to the proportion of the resident population in the corresponding age stratum to obtain the real-time estimate of the population ILI%. S6: Multi-time domain prediction based on Δ target. The Δ target strategy is adopted for 1-4 week advance prediction. The weekly change Δ of ILI% is defined as the prediction target. The model is trained to predict the Δ value and reconstruct the absolute prediction value to eliminate the influence of the time series trend term and improve the stability of the 1-4 week advance prediction. S7: Threshold-based early warning decision, triggers corresponding level early warnings based on early warning thresholds, and simultaneously calculates and outputs the sensitivity, specificity, positive predictive value, and F1 score of the early warning.
2. The influenza early warning method based on population-stratified multi-source data fusion according to claim 1, characterized in that, In step S5, the loss function of the regression sub-model is: ; In the formula Here is the regularization coefficient, in A 5-fold cross-validation grid search is performed on the set, with the optimization objective being to minimize the mean absolute error of the validation set.
3. The influenza early warning method based on population-stratified multi-source data fusion according to claim 1, characterized in that, In step S6, the Δ-Ridge predicted value and the seasonal baseline predicted value based on the historical average of the same period are adjusted by adaptive weighting. Perform weighted fusion: ; Weight The system dynamically adjusts the weighting of the seasonal baseline forecast based on influenza activity: when influenza activity is stable, the weighting of the seasonal baseline forecast is increased; when the change in the Δ value exceeds 1.5 times the standard deviation of the historical period for two consecutive weeks, the weighting of the Δ-Ridge forecast is automatically increased. Weight The range of values is The initial values are determined by a grid search on the validation set, and the optimization objective is the validation set determination coefficient. maximum.
4. The influenza early warning method based on population-stratified multi-source data fusion according to claim 1, characterized in that, In step S7, a dynamic early warning threshold adjustment mechanism is adopted, which dynamically calculates the early warning threshold based on historical ILI data for the same period and the current vaccination rate. ; In the formula This is the average of the ILI% for the same period over the past 5 years. This represents the standard deviation of ILI% over the same period over the past 5 years. This is the threshold coefficient (default value is 1.5). The vaccination rate adjustment coefficient is used, and a segmented adjustment method is adopted: when the vaccination rate is below 30%, Take 0.7; when the vaccination rate is higher than or equal to 30%, Take 0.5; The current influenza vaccination rate; when the predicted population ILI% exceeds An orange alert is triggered when the predicted value exceeds [a certain threshold]. A red alert was triggered at that time.
5. An early warning system for influenza based on population-stratified multi-source data fusion, characterized in that, include: Data acquisition module: Used to automatically collect outpatient data by age group from sentinel hospitals, meteorological data, internet search index data, and vaccination data through API interface, and perform ISO calendar alignment and missing value filling; Data preprocessing module: used to perform semantic classification of search keywords, cross-platform index normalization fusion, principal component dimensionality reduction, hierarchical feature selection and outlier cleaning operations; Real-time estimation module: used to perform age-stratified Ridge regression ensemble modeling and output real-time estimates of the population ILI%. Multi-time-domain prediction module: used to perform multi-time-domain prediction based on the Δ target, outputting the predicted population ILI% value for the next 1 to 4 weeks; Early warning decision module: used to trigger corresponding level early warnings based on dynamic early warning thresholds and calculate early warning evaluation indicators.
6. The influenza early warning system based on population-stratified multi-source data fusion according to claim 5, characterized in that, The internet search index data collected by the data acquisition module includes search engine indexes and short video platform indexes; the specific method for cross-platform normalization fusion is as follows: For keywords shared by search engines and short video platforms, the weekly index values of the two platforms are min-max normalized to the [0,1] interval, and then the corresponding values are added together to obtain the fusion index value.
7. An early warning system for influenza based on population-stratified multi-source data fusion as described in claim 5, characterized in that, It also includes a model update and monitoring module, which is used to incrementally fine-tune the model weekly using newly added real ILI data. When the model prediction error exceeds a preset threshold for four consecutive weeks, a full retraining process is triggered. Simultaneously, it monitors the data quality and model performance metrics of each data source in real time and generates a model health report.
8. An early warning system for influenza based on population-stratified multi-source data fusion as described in claim 5, characterized in that, The early warning decision module is also used to generate a visual early warning report, which displays the real-time ILI% trend, the prevalence of the disease by age group, the prediction curve and the early warning information, and supports multi-dimensional filtering and querying by administrative region, age group and time range.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the influenza early warning method based on population stratification and multi-source data fusion as described in any one of claims 1-4.