A method for epidemic prediction based on deep embedding clustering meta-learning

Through the method based on deep embedded clustering meta-learning, the epidemic transmission patterns in multiple regions are learned, and unsupervised meta-learning is used to transfer these patterns to new outbreak areas, which solves the problem of difficult to effectively predict the future development of new outbreak areas in the existing technology, and achieves efficient prediction in the absence of a large amount of historical data.

CN115240871BActive Publication Date: 2025-06-06南昌理工学院 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210887157.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-06-06
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

It is difficult for existing technology to effectively utilize historical data to predict future epidemic development in areas with new outbreaks, especially in the absence of a large amount of historical data.

Method used

A method based on deep embedded cluster meta-learning is adopted to learn fine-grained transmission patterns through time series fragments of epidemic transmission in multiple regions, and the unsupervised meta-learning method is used to transfer these patterns to new outbreak areas for prediction.

Benefits of technology

It has achieved predictions on future epidemic development in areas with new outbreaks, requiring only a small amount of historical data and having good interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240871B_ABST
    Figure CN115240871B_ABST
Patent Text Reader

Abstract

The present invention discloses an epidemic prediction method based on deep embedded clustering meta-learning, comprising the following steps: S1, obtaining historical data, dividing the historical data into a plurality of time series segments matching the data length of a target area, each time series segment comprising a historical segment part and a future segment part; S2, for each time series segment, standardizing the historical segment part and the future segment part thereof respectively, and obtaining a feature set of the time series segment; S3, clustering the time series segments based on an unsupervised clustering model to obtain a plurality of classes, sampling p classes to construct a meta-training set, and obtaining meta-knowledge, initializing parameters of a new task model based on the meta-knowledge, and training the initialized new task model through the meta-training set; S4, obtaining a prediction model, initializing parameters, performing adaptive optimization through multi-step gradient descent, and then predicting the development of the epidemic for the new task in the meta-test set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of epidemic prediction, and more specifically to an epidemic prediction method based on deep embedding clustering meta-learning. Background Art

[0002] Currently, machine / deep learning for predicting influenza or other time series data can be divided into two main categories. First, some researchers focus on finding effective "features". For example, search engine query data is used to predict influenza in Google FluTrends1. Twitter data is also used in other research papers. However, these models are often plagued by unreliable sources of information such as Internet searches. For example, Google's algorithm can easily overfit seasonal terms that are not related to influenza, such as "high school basketball". This example also demonstrates the importance of model interpretability. Second, other researchers focus on finding effective "models", such as RF, Gradient Boosting, Multilayer Perceptron (MLP), Long Short-Term Memory (LSTM), Transformer (TFR), etc. Deep learning-based methods such as Transformer have received more attention for their accuracy, while most of them suffer from poor interpretability. In addition, statistical models and dynamic analysis models are considered to be easily accessible tools for simulating influenza infection patterns, such as SI, SIS, SIR models and their variants. However, their parameters change and the approximation of parameters is difficult, such as the basic reproduction number R0, population mobility, etc. DEFSI combines deep neural network methods with causal models to solve high-resolution ILI incidence prediction. However, most of these models rely heavily on external data to improve accuracy, such as longitude and latitude and climate information.

[0003] Therefore, it is an urgent problem for technical personnel in this field to provide an epidemic prediction method based on deep embedded clustering meta-learning, based on historical data, for newly-breakout areas, using a small amount of initial data to predict the future development of the epidemic. Summary of the invention

[0004] In view of this, the present invention provides an epidemic prediction method based on deep embedded clustering meta-learning; time series fragments of epidemic spread in multiple regions are used to learn fine-grained propagation patterns, and the learned propagation patterns can be used for future predictions in regions with new outbreaks and only a small amount of historical data. Only a small amount of domain knowledge is required to construct meta-learning tasks, and it has good interpretability; an unsupervised meta-learning method based on MAML is used to migrate the disease propagation model from a region where epidemic spread is stable to another region where the epidemic is in an early stage.

[0005] In order to achieve the above object, the present invention adopts the following technical solution:

[0006] A method for epidemic prediction based on deep embedding clustering meta-learning, comprising the following steps:

[0007] S1. Obtain historical data and divide the historical data into multiple time series segments that match the data length of the target area, each time series segment includes a historical segment part and a future segment part;

[0008] S2. For each time series segment, standardize its historical segment part and future segment part respectively, and obtain the feature set of the time series segment;

[0009] S3. Cluster the time series segments based on the unsupervised clustering model to obtain multiple classes, sample p classes to construct a meta-training set, and obtain meta-knowledge. Initialize the parameters of the new task model based on the meta-knowledge, and train the initialized new task model through the meta-training set.

[0010] S4. Obtain the prediction model, initialize the parameters, perform adaptive optimization through multi-step gradient descent, and then predict the development of the epidemic for the new tasks in the meta-test set.

[0011] Preferably, the step S1 specifically includes:

[0012] Get the known historical time series information x of length T in the target region i i , the time series information x i Divide into multiple time series segments of length ω+ΔT

[0013]

[0014] Where M is the number of regions, T i is the total length of historical time series data for region i, is the time series segment of region i at time t, For time series segments The ω data before time t, i.e. the historical fragment, are aligned with the known observation data of the target region i. For time series segments The ΔT data after time t, that is, the future segment part, is aligned with the data to be predicted.

[0015] Preferably, the step S2 specifically includes:

[0016] S21. For the historical fragments and the future fragment part To standardize:

[0017]

[0018]

[0019] in, The time series segments The historical fragment part and the future fragment part The mean of The time series segments The historical fragment part and the future fragment part The variance of the time series is normalized to between 0 and 1;

[0020] S22. For time series segments Based on CNN and RNN, local features and temporal features of the sequence are extracted, and time series fragments are The historical fragments in Corresponding to the features of known data, the embedding representation of the time series fragment is learned only from this part of the features, and the time series fragment set Projected into the embedding space Z, generating a feature set of time series segments

[0021]

[0022] Among them, ξ(·) is the feature encoder, which consists of CNN and RNN. It is a CNN feature extraction operation, which is used to extract local features of time series segments. is the RNN feature extraction operation, which is used to extract the temporal features of the time series fragments, θ c ,θ r They are CNN model parameters and RNN model parameters respectively.

[0023] Preferably, the step S3 specifically includes:

[0024] S31. Time series fragments Clustering is performed and their embedding is learned. Based on the deep clustering model IDEC, clustering loss is used to achieve clustering of given inputs:

[0025]

[0026] Among them, q ij represents a time series segment z measured by a Student's t distribution i and cluster center μ j The similarity, p ijis the target distribution of clustering;

[0027] Feature collection by time series segment Perform clustering to obtain a partition of the time series fragment data set Each cluster is a collection of features of multiple time series segments, and the clustering operation is defined as:

[0028]

[0029]

[0030] Among them, l is the total number of all categories, P i is the i-th cluster, |P i | represents the number of elements in the i-th cluster, z is P i The elements in is the center point of l categories, ||·|| is the binary norm;

[0031] S32. Sample p clusters to construct a meta-training task set M train ={D 1 ,D 2 ,…,D p} is represented by p propagation modes, each cluster D i Divided into Query i and Support i Two parts, corresponding to a prediction task Among them, Support i For tasks Learning adaptation, that is, for updating the basic learner, Query i Used to update meta-learner parameters;

[0032] The minimum mean square error is used as the prediction loss:

[0033]

[0034] Among them, y is the number of confirmed cases of the real epidemic, Predict the results for the model.

[0035] In the base learner learning phase, each task Corresponding to a base learner, based on Support i Data, base learner calculates loss Minimize the loss using gradient descent to find the optimal set of parameters that minimizes the loss:

[0036]

[0037] Among them, θ'i is the optimal parameter for task i, θ is the initial parameter of the model, α is the hyperparameter, is the gradient of task i;

[0038] In the meta-learning phase, Query is used i Data, based on the optimal parameters θ' learned by the base learner i , the meta-learner computes the optimal parameters θ' relative to these i The gradient of θ is used to update the randomly initialized parameter θ, i.e., meta-knowledge, so that θ is adjusted to the optimal value. Under this optimal value state, when it is applied to the prediction of the future epidemic development in a certain region, only a small amount of gradient update is needed to obtain a better prediction effect:

[0039]

[0040] Among them, θ is the initial parameter of the model, β is the hyperparameter, It's a task In Query i The obtained relative to the parameter θ' i gradient.

[0041] Preferably, the step S4 specifically includes:

[0042] For new prediction tasks Assign it to the closest time series segment cluster and sample it to obtain Support test , based on the learned meta-knowledge θ, in Support test Perform gradient descent learning to adapt to new tasks Model.

[0043]

[0044] Among them, θ' test is the model parameter of the new task, θ is the initial parameter, i.e., meta-knowledge, f θ For the prediction model.

[0045] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses an epidemic prediction method based on deep embedded clustering meta-learning; it uses time series fragments of epidemic spread in multiple regions to learn fine-grained propagation patterns, and the learned propagation patterns can be used for future predictions in regions with new outbreaks and only a small amount of historical data. It requires only a small amount of domain knowledge to construct meta-learning tasks and has good interpretability; it uses an unsupervised meta-learning method based on MAML to migrate the disease propagation model from a region where the epidemic spread is stable to another region where the epidemic is in an early stage. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0047] Figure 1 The accompanying drawing is a schematic diagram of the flow structure of the prediction method provided by the present invention.

[0048] Figure 2 The accompanying drawing is a schematic diagram of the model framework structure provided by the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The embodiment of the present invention discloses an epidemic prediction method based on deep embedding clustering meta-learning, comprising the following steps:

[0051] S1. Obtain historical data and divide the historical data into multiple time series segments that match the data length of the target area, each time series segment includes a historical segment part and a future segment part;

[0052] S2. For each time series segment, standardize its historical segment part and future segment part respectively, and obtain the feature set of the time series segment;

[0053] S3. Cluster the time series segments based on the unsupervised clustering model to obtain multiple classes, sample p classes to construct a meta-training set, and obtain meta-knowledge. Initialize the parameters of the new task model based on the meta-knowledge, and train the initialized new task model through the meta-training set.

[0054] S4. Obtain the prediction model, initialize the parameters, perform adaptive optimization through multi-step gradient descent, and then predict the development of the epidemic for the new tasks in the meta-test set.

[0055] To further optimize the above technical solution, step S1 specifically includes:

[0056] Get the known historical time series information x of length T in the target region i i , the time series information x iDivide into multiple time series segments of length ω+ΔT

[0057]

[0058] Where M is the number of regions, T i is the total length of historical time series data for region i, is the time series segment of region i at time t, For time series segments The ω data before time t, i.e. the historical fragment, are aligned with the known observation data of the target region i. For time series segments The ΔT data after time t, that is, the future segment part, is aligned with the data to be predicted.

[0059] Preferably, step S2 specifically includes:

[0060] S21. For the historical fragments and the future fragment part To standardize:

[0061]

[0062]

[0063] in, The time series segments The historical fragment part and the future fragment part The mean of The time series segments The historical fragment part and the future fragment part The variance of the time series is normalized to between 0 and 1;

[0064] S22. For time series segments Based on CNN and RNN, local features and temporal features of the sequence are extracted, and time series fragments are The historical fragments in Corresponding to the features of known data, the embedding representation of the time series fragment is learned only from this part of the features, and the time series fragment set Projected into the embedding space Z, generating a feature set of time series segments

[0065]

[0066] Among them, ξ(·) is the feature encoder, which consists of CNN and RNN. It is a CNN feature extraction operation, which is used to extract local features of time series segments. is the RNN feature extraction operation, which is used to extract the temporal features of the time series fragments, θ c ,θ r They are CNN model parameters and RNN model parameters respectively.

[0067] To further optimize the above technical solution, step S3 specifically includes:

[0068] S31. Time series fragments Clustering is performed and their embedding is learned. Based on the deep clustering model IDEC, clustering loss is used to achieve clustering of given inputs:

[0069]

[0070] Among them, q ij represents a time series segment z measured by a Student's t distribution i and cluster center μ j The similarity, p ij is the target distribution of clustering;

[0071] Feature collection by time series segment Perform clustering to obtain a partition of the time series fragment data set Each cluster is a collection of features of multiple time series segments, and the clustering operation is defined as:

[0072]

[0073]

[0074] Among them, l is the total number of all categories, P i is the i-th cluster, |P i | represents the number of elements in the i-th cluster, z is P i The elements in is the center point of l categories, ||·|| is the binary norm;

[0075] S32. Sample p clusters to construct a meta-training task set M train ={D 1 ,D 2 ,…,D p} is represented by p propagation modes, each cluster D i Divided into Query i and Support i Two parts, corresponding to a prediction task Among them, Support i For tasks Learning adaptation, that is, for updating the basic learner, Query i Used to update meta-learner parameters;

[0076] The minimum mean square error is used as the prediction loss:

[0077]

[0078] Among them, y is the number of confirmed cases of the real epidemic, Predict the results for the model.

[0079] In the base learner learning phase, each task Corresponding to a base learner, based on Support i Data, base learner calculates loss Minimize the loss using gradient descent to find the optimal set of parameters that minimizes the loss:

[0080]

[0081] Among them, θ' i is the optimal parameter for task i, θ is the initial parameter of the model, α is the hyperparameter, is the gradient of task i;

[0082] In the meta-learning phase, Query is used i Data, based on the optimal parameters θ' learned by the base learner i , the meta-learner computes the optimal parameters θ' relative to these i The gradient of θ is used to update the randomly initialized parameter θ, i.e., meta-knowledge, so that θ is adjusted to the optimal value. Under this optimal value state, when it is applied to the prediction of the future epidemic development in a certain region, only a small amount of gradient update is needed to obtain a better prediction effect:

[0083]

[0084] Among them, θ is the initial parameter of the model, β is the hyperparameter, It's a task In Query i The obtained relative to the parameter θ' i gradient.

[0085] To further optimize the above technical solution, step S4 specifically includes:

[0086] For new prediction tasks Assign it to the cluster of the closest time series segments and sample it to obtain Support test, based on the learned meta-knowledge θ, in Support test Perform gradient descent learning to adapt to new tasks Model.

[0087]

[0088] Among them, θ' test is the model parameter of the new task, θ is the initial parameter, i.e., meta-knowledge, f θ For the prediction model.

[0089] Evaluation indicator: We use root mean square error and Pearson correlation coefficient As a metric, lower RMSE values ​​are better, while higher PCC values ​​are better.

[0090] Comparison method:

[0091] –AR: Standard autoregressive model

[0092] –LSTM: Recurrent Neural Network (RNN) using LSTM cells

[0093] –TPA-LSTM: Attention-based LSTM model (Shih, SY, Sun, FK, Lee, Hy: Temporal pattern attention for multivariate time series forecasting. Machine Learning (2019))

[0094] –ST-GCN

[20] : Spatiotemporal Graph Neural Network

[0095] –CNNRNN-Res: A deep learning model combining CNN, RNN and residual links for epidemiological prediction (Yu, B., Yin, H., Zhu, Z.: Spatio-temporal graph convolutional networks: A deeplearning framework for traffic forecasting. arXiv preprint arXiv:1709.04875(2017))

[0096] –SAIFlu-Net: Self-attention-based influenza prediction model (Jung, S., Moon, J., Park, S., Hwang, E.: Self-attention-based deep learning network for regional influenza forecasting. IEEE JBHI (2021))

[0097] –Cola-GNN: A deep learning model combining CNN, RNN and GCN for epidemic prediction (Deng, S., Wang, S., Rangwala, H., Wang, L., Ning, Y.: Cola-gnn: Cross-location attention based graph neural networks for long-term ili prediction. In: Proc. of CIKM (2020))

[0098] RMSE and PCC performance of different methods on three datasets, horizon = 3, 5, 10, 15. Bold indicates the best result in each column, and underline indicates the second best. * indicates that the results are reported in the corresponding references

[0099]

[0100] We evaluate each model in both short-term (range < 10) and long-term (range ≥ 10) settings. The influenza dataset is shown in the table. The general trend is that the prediction accuracy decreases as the prediction horizon increases, because the larger the horizon, the harder the problem. The large differences in RMSE between different datasets are due to the scale and variance of the datasets.

[0101] We observe that our method outperforms other models on most tasks. Our method achieves 5.6% lower RMSE than the best baseline in the flu forecasting task. In the flu forecasting task, most deep learning based models perform better than statistical models (HA and AR) because they strive to handle the nonlinear characteristics and complex patterns behind the time series.

[0102] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0103] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An epidemic prediction method based on deep embedding clustering meta-learning, It is characterized in that The following steps are involved: S1. Obtain historical data on the number of confirmed cases of real epidemics, and divide the historical data into multiple time series segments that match the data length of the target area, each time series segment includes a historical segment part and a future segment part; S2 specifically includes: S21. For the historical fragments and the future fragment part To standardize; S22. For time series segments Based on CNN and RNN, local features and temporal features of the sequence are extracted, and time series fragments are The historical fragments in Corresponding to the characteristics of known data, the time series fragments are collected Projected into the embedding space Z, generating a feature set of time series segments Among them, ξ(·) is the feature encoder, which consists of two parts: CNN and RNN. It is a CNN feature extraction operation, which is used to extract local features of time series segments. is the RNN feature extraction operation, which is used to extract the temporal features of the time series fragments, θ c ,θ r They are CNN model parameters and RNN model parameters respectively; Step S3 specifically includes: S31. Time series fragments Clustering is performed and their embedding is learned. Based on the deep clustering model IDEC, clustering loss is used to achieve clustering of given inputs: Among them, q ij represents a time series segment z measured by a Student's t distribution i and cluster center μ j The similarity, p ij is the target distribution of clustering; Feature collection by time series segment Perform clustering to obtain a partition of the time series fragment data set Each cluster is a collection of features of multiple time series segments, and the clustering operation is defined as: Among them, l is the total number of all categories, P i is the i-th cluster, |P i | represents the number of elements in the i-th cluster, z is P i The elements in is the center point of l categories, ||·|| is the binary norm; S32. Sample p clusters to construct a meta-training task set M train ={D 1 ,D 2 ,…,D p } is represented by p propagation modes, each cluster D i Divided into Query i and Support i Two parts, corresponding to a prediction task Among them, Support i For tasks Learning adaptation, that is, for updating the basic learner, Query i Used to update meta-learner parameters; The minimum mean square error is used as the prediction loss: Among them, y is the number of confirmed cases of the real epidemic, Predict results for the model; In the base learner learning phase, each task Corresponding to a base learner, based on Support i Data, base learner calculates loss Minimize the loss using gradient descent to find the optimal set of parameters that minimizes the loss: Among them, θ' i is the optimal parameter for task i, θ is the initial parameter of the model, α is the hyperparameter, is the gradient of task i; In the meta-learning phase, Query is used i Data, based on the optimal parameters θ' learned by the base learner i , the meta-learner computes the optimal parameters θ' relative to these i The gradient of θ is used to update the randomly initialized parameter θ, i.e., meta-knowledge, so that θ is adjusted to the optimal value. Under this optimal value state, when it is applied to the prediction of the future epidemic development in a certain region, only a small amount of gradient update is needed to obtain a better prediction effect: Among them, θ is the initial parameter of the model, β is the hyperparameter, It's a task In Query i The obtained relative to the parameter θ' i The gradient of S4. Obtain the prediction model, initialize the parameters, perform adaptive optimization through multi-step gradient descent, and then predict the development of the epidemic for the new tasks in the meta-test set.

2. According to claim 1, an epidemic prediction method based on deep embedding clustering meta-learning, It is characterized in that The step S1 specifically includes: Get the target area d length T d Known historical time series information x d , the time series information x d Divide into multiple time series segments of length ω+ΔT 1≤d≤M,ω≤t≤T d -ΔT Where M is the number of regions, T d is the total length of historical time series data for region d, is the time series segment of region d at time t, For time series segments The ω data before time t, i.e. the historical fragment, are aligned with the known observation data of the target area d. For time series segments The ΔT data after time t, that is, the future segment part, is aligned with the data to be predicted.

3. According to claim 1, an epidemic prediction method based on deep embedding clustering meta-learning, It is characterized in that The step S4 specifically includes: For new prediction tasks Assign it to the cluster of the closest time series segments and sample it to obtain Support test , based on the learned meta-knowledge θ, in Support test Perform gradient descent learning to adapt to new tasks Model: Among them, θ' test is the model parameter of the new task, θ is the initial parameter, i.e., meta-knowledge, f θ For the prediction model.

Citation Information

Patent Citations

  • Regional epidemic trend prediction and early warning method and system

    CN113744888A