Dynamic disease prediction system based on real-time search trend and public data

By building a dynamic disease prediction system based on real-time search trends and public data, the difficulty and lag problems of data acquisition in the existing technology are solved, and dynamic prediction of disease search index is realized, and the accuracy and timeliness of prediction are improved.

CN119993551APending Publication Date: 2025-05-13BEIJING CHINESE MEDICINE HOSPITAL AFFILIATED CAPITAL MEDICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510039955.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing disease prediction methods, data acquisition difficulties, data lag, and dynamic changes between data are not considered, resulting in bias in the analysis results.

Method used

A dynamic disease prediction system based on real-time search trends and public data is adopted, and a multi-dimensional public data set is obtained through the data acquisition module. The dynamic disease analysis module constructs a multi-dimensional dynamic analysis model, identify and quantify the data that affects the disease search index and the degree of impact, and finally predicts it by the disease prediction module.

Benefits of technology

It realizes dynamic timing prediction of disease search index, improves the accuracy and timeliness of disease prediction, and can be effectively applied to the prediction and management of infectious diseases and chronic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993551A_ABST
    Figure CN119993551A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic disease prediction system based on a real-time search trend and public data, belongs to the technical field of disease prediction, and solves the problems that during existing disease prediction, data acquisition is difficult, hysteresis exists, and a dynamic change relation between data is not considered, so that a prediction result is deviated. The system comprises a data acquisition module which is used for acquiring a public data set of continuous time with a city as a unit and performing preprocessing; the public data set comprises disease search indexes, meteorological data and social economic data; the dynamic disease analysis module is used for constructing and applying a multi-dimensional dynamic analysis model on the basis of the multi-dimensional public data set, and identifying and quantifying the data influencing the search index of the disease and the influence degree; wherein the multi-dimensional dynamic analysis model is used for capturing the dynamic relationship of the data and obtaining the time sequence relationship among the data; and the disease prediction module is used for predicting the disease development trend based on the data influencing the search index of the disease and the influence degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease prediction, and in particular to a dynamic disease prediction system based on real-time search trends and public data. Background Art

[0002] With the rapid development of information technology, disease prediction technology has gradually expanded from relying on the analysis of single medical data to a comprehensive analysis method that integrates multi-dimensional data, including real-time search data, environmental indicators, economic indicators, etc. However, existing disease prediction methods mainly rely on traditional medical data, such as emergency records, medical insurance data, etc. Although these data play an important role in analyzing the disease burden, it is difficult to achieve the timeliness and accuracy of disease prediction due to the difficulty of obtaining and the lag.

[0003] In recent years, the application of real-time search data in infectious disease prediction has gradually gained attention. However, for the prediction and management of chronic diseases, existing technologies usually use static correlation analysis, which fails to effectively capture the dynamic changes between data, resulting in large deviations in analysis results and affecting the accuracy of predictions. In addition, environmental factors (such as climate conditions and air pollution levels) are closely related to the incidence of chronic diseases, and socioeconomic status also significantly affects the health status of individuals. Existing technologies fail to fully combine real-time data and dynamic models, making it difficult to accurately reflect the time-varying characteristics of disease influencing factors. Summary of the invention

[0004] In view of the above analysis, an embodiment of the present invention aims to provide a dynamic disease prediction system based on real-time search trends and public data, so as to solve the problems in existing disease prediction methods, such as difficulty in data acquisition, data lag, and failure to consider the dynamic change relationship between data, which leads to deviations in analysis results.

[0005] The purpose of the present invention is mainly achieved through the following technical solutions:

[0006] The present invention provides a dynamic disease prediction system based on real-time search trends and public data, comprising:

[0007] A data collection module, used to obtain and pre-process a multi-dimensional public data set of continuous time in cities; wherein the multi-dimensional public data set includes a disease search index, meteorological data, and socio-economic data;

[0008] A dynamic disease analysis module, used to construct and apply a multidimensional dynamic analysis model based on the preprocessed multidimensional public data set, to identify and quantify the data and degree of influence on the search index of the disease; wherein the multidimensional dynamic analysis model is used to capture the dynamic relationship between the data in the multidimensional public data set and obtain the time series relationship between the data;

[0009] The disease prediction module is used to predict the development trend of the disease based on the data that affects the search index of the disease and the degree of influence.

[0010] Furthermore, the dynamic disease analysis module includes a multi-dimensional public data set preliminary analysis module, a multi-dimensional dynamic modeling and analysis module, and a causal analysis module;

[0011] The multi-dimensional public data set preliminary analysis module is used to analyze the data related to the search index of the disease based on the pre-processed multi-dimensional public data set to obtain a preliminary linear relationship between the search index of the disease and the related data;

[0012] The multidimensional dynamic modeling and analysis module is used to perform further causal analysis based on the preliminary linear relationship between the search index of the disease and the relevant data, construct the multidimensional dynamic analysis model and obtain the dynamic relationship between the search index of the disease and each relevant data;

[0013] The causal analysis module is used to analyze the impact of each relevant data on the search index of the disease based on the dynamic relationship between the search index of the disease and each relevant data.

[0014] Furthermore, the multi-dimensional public data set preliminary analysis module includes: a correlation data selection module, an optimal lag order selection module and a correlation preliminary analysis module;

[0015] The correlation data selection module uses a correlation analysis method to obtain meteorological data and socioeconomic data related to the search index of the disease, and constructs an initial dependent variable vector and a control variable matrix; wherein the initial dependent variable vector includes the search index of the disease and the related meteorological data; and the control variable matrix includes the socioeconomic data related to the search index of the disease;

[0016] The lag order selection module uses the information criterion method to obtain the optimal lag order of the data in the multidimensional public data set;

[0017] The preliminary correlation analysis module constructs a preliminary analysis model based on the initial dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the preliminary analysis model to obtain the preliminary linear relationship between the disease search index and related data.

[0018] Furthermore, the multi-dimensional dynamic modeling and analysis module includes a key data selection module, a multi-dimensional dynamic analysis model construction and a dynamic relationship analysis module;

[0019] The key data selection module, based on the preliminary linear relationship between the search index of the disease and the relevant data obtained by the preliminary analysis model, uses the Granger causality test method to perform a causal test on the initial dependent variable vector, removes the meteorological data with low correlation in the initial dependent variable vector, and obtains a screened dependent variable vector;

[0020] The multidimensional dynamic analysis model construction and dynamic relationship analysis module constructs a multidimensional dynamic analysis model based on the screened dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the multidimensional dynamic analysis model to obtain the dynamic relationship between the variables in the multidimensional dynamic analysis model.

[0021] Furthermore, the preliminary analysis model is expressed as:

[0022]

[0023] in, represents the initial dependent variable vector at time t; n represents the optimal lag order; represents the coefficient of the i-th lag order; X t represents the control variable matrix at time t; B pre represents the influence coefficient of the initial control variable matrix; represents the error term at time t;

[0024] The multi-dimensional dynamic analysis model is expressed as:

[0025]

[0026] Among them, Y t represents the dependent variable vector after screening at time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening; ε t represents the error term at time t.

[0027] Further, the cause-effect analysis module includes: an impact response analysis module and an impact rate ranking module;

[0028] The impact response analysis module performs impact response analysis on the search index of the disease and each related data in the multidimensional dynamic analysis model to obtain the dynamic impact of the unit standard deviation change of each related data on the disease search index;

[0029] The influence rate ranking module uses the variance decomposition method to obtain the influence rate of each relevant data on the search index fluctuation of the disease in the multidimensional dynamic analysis model and each relevant data, and ranks the influence rate.

[0030] Furthermore, the disease prediction module includes a trend prediction module, a disease early warning module and a report generation module;

[0031] The trend prediction module predicts the changing trend of the disease search index in the next few days based on the coefficients of each lag order and the influence coefficients of each relevant data of the multidimensional dynamic analysis model;

[0032] The disease warning module, when the disease search index prediction value obtained by the trend prediction module is greater than a preset threshold, performs a disease warning based on the influence rate ranking result of each relevant data;

[0033] The report generation module generates a disease search index prediction curve, an influence rate ranking result and an impact response curve respectively based on the disease search index prediction value obtained by the trend prediction module, the influence rate ranking result of each relevant data and the impact response analysis result.

[0034] Furthermore, based on the coefficients of each lag order of the multidimensional dynamic analysis model, the influence coefficients of each relevant data, the historical data of the screened dependent variable vector and the control variable matrix, the following formula is used iteratively to predict the change trend of the disease search index of the next day until the change trend of the disease search index in the next few days is obtained:

[0035]

[0036] Among them, Y t+1 represents the vector of the dependent variable after screening on the day after time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the predicted control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening.

[0037] Furthermore, the multi-dimensional public data set is preprocessed, including: removing outliers of each data in the multi-dimensional public data set and aligning them in time series, and normalizing each data.

[0038] Furthermore, the meteorological data includes daily temperature, humidity, wind speed, precipitation and air pollutant concentrations of each city;

[0039] The socio-economic data include GDP, unemployment rate and CPI of each city every quarter.

[0040] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0041] 1. The present invention integrates real-time search trends with multi-dimensional public data including meteorological data and socioeconomic data to perform dynamic time series prediction of the disease search index. The constructed multi-dimensional dynamic analysis model belongs to the enhanced panel vector autoregression (EPVAR) model, which can accurately capture the dynamic relationship between disease-related variables, predict the changing trend of the disease search index, and improve the accuracy of disease prediction.

[0042] 2. The present invention utilizes the real-time retrieval data of the search engine to timely reflect the changing trend of the disease, solves the problem of lagging traditional medical data, and makes disease prediction more timely.

[0043] 3. The present invention utilizes the real-time retrieval data of search engines and various types of easily available public data, which is not only suitable for the prediction of infectious diseases, but can also be effectively applied to the management and prediction of chronic diseases, expanding the application scope of real-time search data and facilitating the analysis of influencing factors and long-term management of chronic diseases.

[0044] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like components throughout the drawings.

[0046] Figure 1 A schematic diagram of the structure of a dynamic disease prediction system based on real-time search trends and public data in an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the structure of a dynamic disease analysis module in an embodiment of the present invention;

[0048] Figure 3 Schematic diagram of the structure of a multi-dimensional public data set preliminary analysis module in an embodiment of the present invention;

[0049] Figure 4 A schematic diagram of the structure of a multi-dimensional dynamic modeling and analysis module in an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of the structure of a cause-effect analysis module in an embodiment of the present invention;

[0051] Figure 6 Schematic diagram of the structure of the disease prediction module in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0053] A specific embodiment of the present invention discloses a dynamic disease prediction system based on real-time search trends and public data, such as Figure 1 As shown, it includes: a data acquisition module, a dynamic disease analysis module and a disease prediction module; wherein,

[0054] The data acquisition module is used to obtain and preprocess a multi-dimensional public data set of continuous time in cities; wherein the multi-dimensional public data set includes a disease search index, meteorological data, and socio-economic data.

[0055] Specifically, the search index of the disease uses the search engine index as a proxy variable for disease attention. For example, Baidu Index or Google Trends can be used. The search index of the disease can reflect the public's attention and search behavior to a certain disease in real time. It is usually closely related to the epidemic trend of the disease and the public's health awareness. When a certain disease begins to become popular, people will search for relevant information more frequently, resulting in an increase in the search index. For some new or rare diseases, since traditional medical data is difficult to obtain in a timely manner or there is a lag, the search index can be used as an indicator to predict the epidemic trend and spread of the disease, and provide a reference for public health monitoring. It should be noted that the search index of the disease is obtained on a daily basis in cities.

[0056] Furthermore, the meteorological data includes daily temperature, humidity, wind speed, precipitation and air pollutant concentration of each city.

[0057] Specifically, the meteorological data include meteorological change indicators and air quality indicators. The occurrence and spread of many diseases are closely related to meteorological conditions. For example, respiratory infectious diseases such as influenza are more easily spread in cold and dry winters, and some mosquito-borne diseases (such as dengue fever) are related to hot and rainy climatic conditions. Therefore, by analyzing meteorological data, it is possible to understand the impact mechanism of environmental factors on diseases. At the same time, meteorological data can assist in identifying the seasonal variation of diseases, thereby providing a basis for the prevention and control of seasonal diseases. For example, based on historical meteorological data and disease incidence data, the start and end time of the flu season can be predicted, and prevention and control preparations can be made in advance. The meteorological data in this embodiment is obtained from the China Meteorological Data Network.

[0058] Furthermore, the socio-economic data include GDP, unemployment rate and CPI of each city every quarter.

[0059] Specifically, among the socio-economic data, the gross domestic product (GDP) is used to measure the overall level of economic activity; the unemployment rate is used to reflect the dynamic situation of the social labor market; the consumer price index (CPI) is used to reflect the price level and inflation. Socio-economic data is one of the important factors affecting the health of individuals. People in high socio-economic areas usually have better medical resources, health awareness and lifestyles, which reduces the risk of certain diseases. For example, higher gross domestic product (GDP) levels are usually associated with better public health infrastructure and medical services, while higher unemployment rates may cause people to face greater life pressure and economic difficulties, thereby affecting their health. The socio-economic data in this embodiment is obtained from the quarterly data of the National Bureau of Statistics of China.

[0060] Furthermore, the multi-dimensional public data set is preprocessed, including: removing outliers of each data in the multi-dimensional public data set and aligning them in time series, and normalizing each data.

[0061] Specifically, the actual data in the obtained multi-dimensional public data set usually contains noise and outliers, which may mislead the analysis results. By removing these outliers to improve the accuracy and reliability of the data, the real characteristics and laws of the data can be better captured. Exemplarily, the method for removing outliers can use the box plot IQR method based on the interquartile range. Specifically, data points below the first quartile minus 1.5 times the IQR or above the third quartile plus 1.5 times the IQR are considered outliers.

[0062] In order to ensure that the timestamps of each data in the multidimensional public dataset are consistent, each data is aligned in time series; wherein, since the disease search index and meteorological data in the multidimensional public dataset are data obtained daily, and the socioeconomic data are data obtained quarterly, the quarterly data in the socioeconomic data are evenly distributed to each day in the quarter using the average distribution method, and the split socioeconomic data are aligned with the daily disease search index and meteorological data.

[0063] For example, the current quarter of a city is 90 days, and the gross domestic product (GDP) of the quarter is 100 billion yuan, and the average GDP per day is 1000÷90≈1.111 billion yuan; the unemployment rate of the quarter is 5%, and the average unemployment rate per day is 5÷90≈0.056%; the consumer price index (CPI) of the quarter is 102%, and the average consumer price index (CPI) per day is 102÷90≈1.133%.

[0064] It should be noted that when aligning the multi-dimensional public data set in time series, if there is missing data in the data, it is necessary to fill the missing data through interpolation or mean filling method to maintain the integrity of the data.

[0065] More specifically, the normalization of each data includes: performing logarithmic transformation on non-normally distributed variables in the multidimensional public data set with aligned timestamps, normalizing all data to the interval [0,1], and fusing the normalized data into a structured data set in chronological order.

[0066] It should be noted that according to the central limit theorem, even if the original data is not normally distributed, as long as the sample size is large enough, the distribution of the sample mean will be close to the normal distribution. The normal distribution is the basic assumption of many statistical analysis methods. Therefore, in this embodiment, for the data in the multi-dimensional public data set, if it is non-normally distributed data, it is logarithmically transformed to convert the skewed distributed data into data closer to the normal distribution. The nonlinear relationship can be converted into a linear relationship, so that the distribution of the data is more symmetrical, so as to improve the effect of statistical analysis and modeling.

[0067] More specifically, for the logarithmically transformed data in the multi-dimensional public data set, the minimum-maximum normalization method is used to normalize the variables to the interval [0,1] to eliminate the dimensional effects between different variables.

[0068] The normalized data are fused in chronological order to ensure that the data at each time point contains all relevant data, so as to obtain a structured data set, which provides a basis for subsequent model building and causal analysis; wherein, the structured data set is a data set that is stored and presented in an organized and formatted manner. In a structured data set, the data is organized into rows and columns, each row represents a record, and each column represents a feature. The structured data set makes the data easy to store, retrieve and analyze.

[0069] Furthermore, the dynamic disease analysis module is used to construct and apply a multidimensional dynamic analysis model based on the preprocessed multidimensional public data set to identify and quantify the data and degree of influence on the search index of the disease; wherein, the multidimensional dynamic analysis model is used to capture the dynamic relationship between each data in the multidimensional public data set and obtain the time series relationship between each data.

[0070] Specifically, the dynamic disease analysis module converts the complex dynamic relationships of data in multidimensional public datasets into quantifiable information, reveals how different factors affect the changing trends of disease search indexes over time, and provides accurate data support for public health decision-making.

[0071] Further, such as Figure 2 As shown, the dynamic disease analysis module includes a multi-dimensional public data set preliminary analysis module, a multi-dimensional dynamic modeling and analysis module, and a causal analysis module.

[0072] The multi-dimensional public data set preliminary analysis module is used to analyze the data related to the search index of the disease based on the pre-processed multi-dimensional public data set, and obtain the preliminary linear relationship between the search index of the disease and the related data.

[0073] Specifically, the preliminary analysis model is used to preliminarily predict the changing trend of the disease search index, providing a basis for subsequent in-depth analysis and model optimization.

[0074] Further, such as Figure 3 As shown, the multi-dimensional public data set preliminary analysis module includes: a correlation data selection module, an optimal lag order selection module and a correlation preliminary analysis module.

[0075] The correlation data selection module uses a correlation analysis method to obtain meteorological data and socioeconomic data related to the search index of the disease, and constructs an initial dependent variable vector and a control variable matrix; wherein the initial dependent variable vector includes the search index of the disease and related meteorological data; and the control variable matrix includes the socioeconomic data related to the search index of the disease.

[0076] Specifically, the correlation coefficient between the disease search index and each candidate variable is calculated using a correlation analysis method for the preprocessed multi-dimensional public data set. The corresponding significance probability value (p-value) is calculated by performing a statistical test on each correlation coefficient; wherein the p-value is used to determine whether the correlation is statistically significant; when the p-value of a variable is less than a preset significance level, the variable is determined to be a variable significantly correlated with the disease search index.

[0077] In this embodiment, the correlation analysis method may use the Pearson correlation coefficient and the Spearman correlation coefficient; the preset significance level is 0.05.

[0078] It should be noted that the statistical test of each correlation coefficient to calculate the corresponding significance probability value (p value) includes the following steps:

[0079] The correlation coefficient r for the jth candidate variable j , use the following formula to get the test statistic t of the j-th candidate variable j :

[0080]

[0081] Among them, m is the total number of samples.

[0082] Since the correlation coefficient is obtained by calculating the mean of the observed values ​​of two variables during correlation analysis, the degrees of freedom are set to m-2 when performing test statistics.

[0083] The test statistic t based on the j-th candidate variable j and the degrees of freedom, using statistical software to calculate the p-value of the j-th candidate variable; illustratively, the statistical software can be the pt() function in R software or the function in the scipy.stats module of Python.

[0084] When the p-value of the candidate variable is less than the preset significance level, the null hypothesis is rejected, and it is considered that there is a significant correlation between the two variables; if the p-value of the candidate variable is greater than or equal to the preset significance level, the null hypothesis cannot be rejected, and it is considered that there is insufficient evidence to show that there is a correlation between the two variables.

[0085] More specifically, the initial dependent variable vector constructed in this embodiment is a set of main variables that need to be predicted or analyzed in the model. Since meteorological factors have a direct and significant correlation with disease transmission and can directly affect the occurrence and epidemic trend of the disease, when the preliminary model is constructed, the disease search index and meteorological data are analyzed together as dependent variables.

[0086] The control variable matrix constructed in this embodiment is a set of variables used to control the changes of other variables in the model. Although socioeconomic data also have an indirect impact on the disease, due to the complex mechanism of action and the existence of a lag effect, the socioeconomic data described in this embodiment plays the role of a covariate, which can eliminate the interference of socioeconomic factors in the analysis of the dependent variable relationship while identifying the potential impact of socioeconomic factors on the epidemic trend of the disease.

[0087] Furthermore, the lag order selection module uses an information criterion method to obtain the optimal lag order of the data in the multidimensional public dataset.

[0088] Specifically, the lag order is the number of past values ​​of the variable considered in the model, that is, the lag order determines how many past time points the model uses to predict the data at the current or future time points; illustratively, when the lag order is 0, it represents the original time series without displacement; when the lag order is 1, it represents that the time series data is shifted one place to the left, that is, the predicted value of the current time point considers the value of the previous time point; when the lag order is 2, the predicted value of the current time point considers the values ​​of the previous two time points.

[0089] In this embodiment, the Akaike Information Criterion (AIC) is used to select the optimal lag order, the purpose of which is to strike a balance between the goodness of fit and complexity of the model to help select the best statistical model. Specifically, the Akaike Information Criterion (AIC) measures the degree of fit of the model to the data by calculating the likelihood function value of the model, and avoids overfitting by penalizing the number of parameters in the model, and selects the model with the smallest AIC value as the best model to obtain the optimal lag order n.

[0090] Furthermore, the preliminary correlation analysis module constructs a preliminary analysis model based on the initial dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the preliminary analysis model to obtain the preliminary linear relationship between the disease search index and related data.

[0091] Wherein, the preliminary analysis model is expressed as:

[0092]

[0093] in, represents the initial dependent variable vector at time t; n represents the optimal lag order; represents the coefficient of the i-th lag order; X t represents the control variable matrix at time t; B pre represents the influence coefficient of the initial control variable matrix; represents the error term at time t.

[0094] Specifically, for the preliminary analysis model, the generalized moment estimation method is used to estimate the lag order coefficients of the preliminary analysis model and the influence coefficients of the control variable matrix, and the relationship between the disease search index and the relevant meteorological data is preliminarily analyzed to capture the dynamic relationship between the relevant meteorological data and the disease search index, and to preliminarily determine which factors are significantly related to the disease search index. According to the preliminary analysis model, important variables can be screened out and parameter estimation can be performed, which improves the efficiency and accuracy of subsequent causal analysis.

[0095] It should be noted that the generalized method of moments (GMM) is a semi-parametric estimation method commonly used in statistics and econometrics. By setting moment conditions and selecting appropriate instrumental variables, it solves the endogeneity problem in the model and captures the dynamic relationship between the disease search index and meteorological data. The parameters obtained by the generalized method of moments (GMM) can be used to determine which meteorological factors are significantly correlated with the disease search index, thereby screening out important variables and improving the efficiency and accuracy of subsequent causal analysis.

[0096] Furthermore, the multidimensional dynamic modeling and analysis module is used to perform further causal analysis based on the preliminary linear relationship between the search index of the disease and related data, construct the multidimensional dynamic analysis model and obtain the dynamic relationship between the search index of the disease and each related data.

[0097] like Figure 4 As shown, the multi-dimensional dynamic modeling and analysis module includes a key data selection module and a multi-dimensional dynamic analysis model building and dynamic relationship analysis module.

[0098] Specifically, the multidimensional dynamic modeling and analysis module conducts causal analysis on the relevant data in the preliminary analysis model, deeply explores the causal relationship between the disease search index and the relevant data, and further screens the variables that have a significant causal impact on the disease search index; based on the results of the causal analysis, the preliminary analysis model is optimized and adjusted to construct a more accurate multidimensional dynamic analysis model. The optimized model can better capture the dynamic relationship and causal mechanism between variables, thereby improving the accuracy of disease prediction.

[0099] Furthermore, the key data selection module, based on the preliminary linear relationship between the search index of the disease and related data obtained by the preliminary analysis model, uses the Granger causality test method to perform a causal test on the initial dependent variable vector, removes the meteorological data with low correlation in the initial dependent variable vector, and obtains a screened dependent variable vector.

[0100] Specifically, the Granger causality test is a method for analyzing the causal relationship between variables in time series data. The meteorological data in the preliminary analysis model are Granger causality tested to determine whether these variables have a significant causal effect on the disease search index. If the past value of a meteorological variable can significantly improve the prediction accuracy of the disease search index, the variable is retained; if the lag term of a meteorological variable has no significant effect, it is considered that its correlation with the disease search index is low and needs to be removed. In this embodiment, the grangercausalitytests function in the statsmodels library in Python can be used to perform the Granger causality test.

[0101] Through the above Granger causality test and screening process, the screened dependent variable vector obtained includes the disease search index and the meteorological data with high correlation therewith, which is used to build a more streamlined and effective model. It should be noted that the Granger causality test will also obtain the p-value of each variable. In this embodiment, when the p-value of the variable is less than or equal to the preset threshold, illustratively, the preset threshold is 0.05, it is considered that the variable has a significant Granger causal relationship with the disease search index; when the p-value of the variable is greater than the preset threshold, it is considered that the variable has no significant Granger causal relationship with the disease search index; the variables with significant Granger causal relationship are retained and the disease search index constitutes the screened dependent variable vector.

[0102] Furthermore, the multidimensional dynamic analysis model construction and dynamic relationship analysis module constructs a multidimensional dynamic analysis model based on the screened dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the multidimensional dynamic analysis model to obtain the search index of the disease and the dynamic relationship between each related data.

[0103] The multi-dimensional dynamic analysis model is expressed as:

[0104]

[0105] Among them, Y t represents the dependent variable vector after screening at time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening; ε t represents the error term at time t.

[0106] Specifically, by using Granger causality analysis on the preliminary linear relationship between the disease search index and related data obtained by the preliminary analysis model, meteorological data with high correlation with the disease search index is screened, and the obtained multidimensional dynamic analysis model can focus more on variables that have a significant impact on the disease search index, thereby improving the predictive ability of the model; the generalized moment estimation (GMM) method is used to estimate the lag order coefficients of the multidimensional dynamic analysis model and the influence coefficients of the control variable matrix, and the obtained multidimensional dynamic analysis model can more accurately capture the dynamic relationship between the disease search index and each related data, thereby improving the model's predictive ability for disease epidemic trends.

[0107] Furthermore, the causal analysis module is used to analyze the impact of each relevant data on the search index of the disease based on the dynamic relationship between the search index of the disease and each relevant data obtained by the multidimensional dynamic analysis model.

[0108] like Figure 5 As shown, the cause-effect analysis module includes: an impact response analysis module and an impact rate ranking module.

[0109] The shock response analysis module performs shock response analysis on the search index and related data of the disease in the multidimensional dynamic analysis model to obtain the dynamic impact of the unit standard deviation change of each related data on the disease search index.

[0110] Specifically, the impact response analysis module obtains the dynamic impact on the disease search index by performing a unit standard deviation change on the meteorological data of the screened dependent variable vector; wherein the unit standard deviation change refers to the amount by which the variable increases by one standard deviation, which is used to simulate a typical change amplitude of the variable in reality. By observing the changes in the disease search index at various time points after the impact, the short-term and long-term effects of each related variable on the disease search index are analyzed. Exemplarily, the impulse response analysis can be performed on the multidimensional dynamic analysis model using the impulse response function (IRF) in Stata software.

[0111] More specifically, based on the impact of each relevant variable on the disease search index, an impact response curve is drawn to show the impact intensity and duration of the change in the disease search index after the impact of each relevant variable.

[0112] Furthermore, the influence rate ranking module uses the variance decomposition method for the search index and relevant data of the disease in the multidimensional dynamic analysis model to obtain the influence rate of each relevant data on the fluctuation of the search index of the disease, and ranks the influence rate.

[0113] Specifically, the variance decomposition method is used to analyze the contribution of each independent variable to the fluctuation of the dependent variable in the time series model, and calculate the explanation ratio of each independent variable to the fluctuation of the dependent variable by decomposing the variance of the dependent variable. In this embodiment, the multidimensional dynamic analysis model can be subjected to variance decomposition using the var module in the statsmodels library in Python or the pvarfevd command in the Stata software to obtain the contribution rate of each variable to the total variance.

[0114] For example, the contribution rate ranking results of the variables for a certain disease are shown in Table 1:

[0115] Table 1 Ranking results of contribution rate of variables

[0116] Variable Name Contribution rate (%) Temperature 38.2 humidity 25.6 GDP 20.4 Air Pollution Index 19.8 Precipitation 15.2 unemployment rate 9.3 CPI 5.1

[0117] Furthermore, the disease prediction module is used to predict the development trend of the disease based on the data that affects the search index of the disease and the degree of influence.

[0118] like Figure 6 As shown, the disease prediction module includes: a trend prediction module, a disease early warning module and a report generation module.

[0119] The trend prediction module predicts the changing trend of the disease search index in the next few days based on the coefficients of each lag order and the influence coefficients of each relevant data of the multidimensional dynamic analysis model.

[0120] Specifically, the trend prediction module is based on historical data, as well as the relationship and influence between various related data. The multidimensional dynamic analysis model can predict the changing trend of the disease search index in the next few days.

[0121] Furthermore, based on the coefficients of each lag order of the multidimensional dynamic analysis model, the influence coefficients of each relevant data, the historical data of the screened dependent variable vector and the control variable matrix, the following formula is used iteratively to predict the change trend of the disease search index of the next day until the change trend of the disease search index in the next few days is obtained:

[0122]

[0123] Among them, Y t+1 represents the vector of the dependent variable after screening on the day after time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the predicted control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening.

[0124] Specifically, in the prediction process, the historical data Y of the known filtered dependent variable vector is used t ,Y t-1 ,…,Y t+n-1 , substituting into the above formula, we can get the predicted value of the disease search index on the day after time t. This iterative prediction method uses the latest prediction result as the input for the next prediction after each prediction, which can gradually update the prediction value and accumulate more information. Compared with directly predicting the data of the next few days, it can more accurately capture the future trend of changes.

[0125] It should be noted that when making short-term forecasts, i.e. forecasts for the next few days or weeks, since current socio-economic data can well reflect the economic environment in the next few days, the control variable matrix in the forecast formula uses current socio-economic data.

[0126] When long-term predictions are needed for certain chronic diseases, that is, predictions for several months, the socioeconomic data may change significantly. At this time, the macroeconomic model is used to predict the socioeconomic data of the next month as the prediction control variable matrix, and the historical monthly data of the screened dependent variable vector is used to predict the changes in the disease search index for the next month; among them, the historical monthly data of the screened dependent variable vector is the average of the daily historical data of that month, which is the historical monthly data.

[0127] Specifically, in the forecasting process, based on the known historical monthly data of the filtered dependent variable vector Use the following formula to predict the disease search index change trend for the next month:

[0128]

[0129] in, represents the vector of the dependent variable after screening in the next month of time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; B represents the influence coefficient of the control variable matrix after screening; Represents the forecast control variable matrix for the next month at time t.

[0130] In this embodiment, the macroeconomic model can use the China Macroeconomic Analysis and Forecasting Model (CMAFM) to obtain socioeconomic data for the next several days.

[0131] Furthermore, when the disease search index prediction value obtained by the trend prediction module is greater than a preset threshold, the disease warning module issues a disease warning based on the impact rate ranking result of each relevant data.

[0132] Specifically, the trend prediction module continuously outputs the predicted value of the disease search index for the next several days, and the disease warning module monitors the predicted value of the disease search index in real time. When its value exceeds a preset threshold, it indicates that the disease may break out or there is a risk of aggravation. According to the contribution rate ranking result of the disease search index, specific disease warning information is generated, indicating the main factors that may lead to increased disease risk and the degree of their influence; wherein, the warning information includes the name of the disease, risk factors, preventive measures and key populations.

[0133] Exemplarily, based on the historical data set of the disease search index, the average value of the disease search index and the standard deviation of the disease search index are obtained, and the preset threshold is set to the disease search index average value + 2× the disease search index standard deviation.

[0134] Furthermore, the report generation module generates a disease search index prediction curve, an influence rate ranking result and an impact response curve based on the disease search index prediction value obtained by the trend prediction module, the influence rate ranking result of each relevant data and the impact response analysis result.

[0135] Specifically, based on the disease search index prediction value obtained by the trend prediction module, a trend curve of the disease search index prediction value over time for the next several days is drawn to discover the potential trend of disease development.

[0136] Based on the impact rate ranking results of each relevant data, an impact rate bar chart is drawn to intuitively display the factors that have the greatest impact on the disease search index, so as to guide the priority allocation of public health resources and the formulation of intervention measures.

[0137] Based on the impact response analysis results, an impact response curve is drawn to show the impact intensity and duration of the change in the disease search index after the impact of each relevant variable.

[0138] It should be noted that the report generation module uses intuitive charts and curves to clarify the degree of influence of each influencing factor and accurately predict the development trend of the disease, so as to quickly grasp the key information of the disease development, optimize the allocation of public health resources, take preventive measures in time, and reduce the risk and impact of disease outbreaks, thereby effectively protecting public health and social stability.

[0139] In summary, a dynamic disease prediction system based on real-time search trends and public data according to an embodiment of the present invention has the following beneficial effects:

[0140] 1. The present invention integrates real-time search trends with multi-dimensional public data including meteorological data and socioeconomic data to perform dynamic time series prediction of the disease search index. The constructed multi-dimensional dynamic analysis model belongs to the enhanced panel vector autoregression (EPVAR) model, which can accurately capture the dynamic relationship between disease-related variables, predict the changing trend of the disease search index, and improve the accuracy of disease prediction.

[0141] 2. The present invention utilizes the real-time retrieval data of the search engine to timely reflect the changing trend of the disease, solves the problem of lagging traditional medical data, and makes disease prediction more timely.

[0142] 3. The present invention utilizes the real-time retrieval data of search engines and various types of easily available public data, which is not only suitable for the prediction of infectious diseases, but can also be effectively applied to the management and prediction of chronic diseases, expanding the application scope of real-time search data and facilitating the analysis of influencing factors and long-term management of chronic diseases.

[0143] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A dynamic disease prediction system based on real-time search trends and public data, characterized in that: include: A data collection module, used to obtain and pre-process a multi-dimensional public data set of continuous time in cities; wherein the multi-dimensional public data set includes a disease search index, meteorological data, and socio-economic data; A dynamic disease analysis module, used to construct and apply a multidimensional dynamic analysis model based on the preprocessed multidimensional public data set, to identify and quantify the data and degree of influence on the search index of the disease; wherein the multidimensional dynamic analysis model is used to capture the dynamic relationship between the data in the multidimensional public data set and obtain the time series relationship between the data; The disease prediction module is used to predict the development trend of the disease based on the data that affects the search index of the disease and the degree of influence.

2. The system according to claim 1, characterized in that: The dynamic disease analysis module includes a multi-dimensional public data set preliminary analysis module, a multi-dimensional dynamic modeling and analysis module, and a causal analysis module; The multi-dimensional public data set preliminary analysis module is used to analyze the data related to the search index of the disease based on the pre-processed multi-dimensional public data set to obtain a preliminary linear relationship between the search index of the disease and the related data; The multidimensional dynamic modeling and analysis module is used to perform further causal analysis based on the preliminary linear relationship between the search index of the disease and the relevant data, construct the multidimensional dynamic analysis model and obtain the dynamic relationship between the search index of the disease and each relevant data; The causal analysis module is used to analyze the impact of each relevant data on the search index of the disease based on the dynamic relationship between the search index of the disease and each relevant data.

3. The system according to claim 2, characterized in that: The multi-dimensional public data set preliminary analysis module includes: a correlation data selection module, an optimal lag order selection module and a correlation preliminary analysis module; The correlation data selection module uses a correlation analysis method to obtain meteorological data and socioeconomic data related to the search index of the disease, and constructs an initial dependent variable vector and a control variable matrix; wherein the initial dependent variable vector includes the search index of the disease and the related meteorological data; and the control variable matrix includes the socioeconomic data related to the search index of the disease; The lag order selection module uses the information criterion method to obtain the optimal lag order of the data in the multidimensional public data set; The preliminary correlation analysis module constructs a preliminary analysis model based on the initial dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the preliminary analysis model to obtain the preliminary linear relationship between the disease search index and related data.

4. The system according to claim 3, characterized in that: The multi-dimensional dynamic modeling and analysis module includes a key data selection module, a multi-dimensional dynamic analysis model construction and a dynamic relationship analysis module; The key data selection module, based on the preliminary linear relationship between the search index of the disease and the relevant data obtained by the preliminary analysis model, uses the Granger causality test method to perform a causal test on the initial dependent variable vector, removes the meteorological data with low correlation in the initial dependent variable vector, and obtains a screened dependent variable vector; The multidimensional dynamic analysis model construction and dynamic relationship analysis module constructs a multidimensional dynamic analysis model based on the screened dependent variable vector, the control variable matrix and the optimal lag order, and uses the generalized moment estimation method to estimate the matrix coefficients of the multidimensional dynamic analysis model to obtain the dynamic relationship between the variables in the multidimensional dynamic analysis model.

5. The system according to claim 4, characterized in that: The preliminary analysis model is expressed as: in, represents the initial dependent variable vector at time t; n represents the optimal lag order; represents the coefficient of the i-th lag order; X t represents the control variable matrix at time t; B pre represents the influence coefficient of the initial control variable matrix; represents the error term at time t; The multi-dimensional dynamic analysis model is expressed as: Among them, Y t represents the dependent variable vector after screening at time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening; ε t represents the error term at time t.

6. The system according to any one of claims 4-5, characterized in that: The causal analysis module includes: an impact response analysis module and an impact rate ranking module; The impact response analysis module performs impact response analysis on the search index and related data of the disease in the multidimensional dynamic analysis model to obtain the dynamic impact of the unit standard deviation change of each related data on the disease search index; The influence rate ranking module uses the variance decomposition method to obtain the influence rate of each relevant data on the search index fluctuation of the disease in the multidimensional dynamic analysis model and each relevant data, and ranks the influence rate.

7. The system according to claim 6, characterized in that: The disease prediction module includes a trend prediction module, a disease early warning module and a report generation module; The trend prediction module predicts the changing trend of the disease search index in the next few days based on the coefficients of each lag order and the influence coefficients of each relevant data of the multidimensional dynamic analysis model; The disease warning module, when the disease search index prediction value obtained by the trend prediction module is greater than a preset threshold, performs a disease warning based on the influence rate ranking result of each relevant data; The report generation module generates a disease search index prediction curve, an influence rate ranking result and an impact response curve respectively based on the disease search index prediction value obtained by the trend prediction module, the influence rate ranking result of each relevant data and the impact response analysis result.

8. The system according to claim 7, characterized in that: Based on the coefficients of each lag order of the multidimensional dynamic analysis model, the influence coefficients of each relevant data, the historical data of the screened dependent variable vector and the control variable matrix, the following formula is used iteratively to predict the change trend of the disease search index of the next day until the change trend of the disease search index in the next few days is obtained: Among them, Y t+1 represents the vector of the dependent variable after screening on the day after time t; n represents the optimal lag order; A i represents the coefficient of the lag order after the i-th causal screening; X t represents the predicted control variable matrix at time t; B represents the influence coefficient of the control variable matrix after screening.

9. The system according to claim 1, characterized in that: The multidimensional public data set is preprocessed, including: removing outliers of each data in the multidimensional public data set and aligning them in time series, and normalizing each data.

10. The system according to claim 9, characterized in that: The meteorological data include daily temperature, humidity, wind speed, precipitation and air pollutant concentrations of each city; The socio-economic data include GDP, unemployment rate and CPI of each city every quarter.

Citation Information

Patent Citations

  • Prediction model establishing device and method as well as computer readable storage medium

    CN107688872A

  • Passenger flow prediction and tourism marketing method based on big data

    CN111951037A

  • Concrete dam uplift pressure safety monitoring method and device

    CN117892408A

  • Method for interactions between traffic and transportation systems based on SDG framework

    US20240161040A1