Epidemic situation prediction method based on STL (Standard Template Library) and Transform model
By combining the STL and Transformer models, the problem of overfitting or underfitting in epidemic data processing is solved, and more accurate epidemic prediction is achieved. It has strong adaptability and is particularly suitable for complex epidemic predictions.
Patent Information
- Application Number
- CN202510586584.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-10-03
AI Technical Summary
When processing large-scale epidemic data, existing technologies have difficulty effectively capturing long-term dependencies and are prone to overfitting or underfitting, resulting in inaccurate epidemic predictions.
A method combining STL and Transformer models is used to construct a Transformer model for epidemic prediction through seasonal-trend decomposition and data integration. Seasonality, trend and residual components are used to generate prediction results, and the model performance is optimized through the loss function.
It improves the accuracy and adaptability of epidemic predictions and is able to process time series data of different types and complexities, making it particularly suitable for complex epidemic predictions.
Smart Images

Figure CN120748770A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of signal processing and prediction analysis. Background Art
[0002] To proactively understand and assess the potential trends and impacts of epidemic spread, thereby providing a scientific basis for public health decision-making, optimizing resource allocation, developing effective prevention and control strategies, and minimizing the impact of epidemics on human health, socioeconomics, and daily life, various institutions are actively adopting a variety of technical means to strengthen their monitoring, analysis, and forecasting capabilities. However, capturing long-term dependencies in epidemic data is inefficient, particularly when dealing with large volumes of data or data with sudden bursts, which can lead to overfitting or underfitting. Summary of the Invention
[0003] To solve the above technical problems, the technical solution adopted by the present invention is an epidemic prediction method based on STL and Transformer model, comprising the following steps:
[0004] S1: Collection and preprocessing: Collect case numbers, time, and location data through the epidemic monitoring platform to form a time series; remove outliers, fill in missing values, and ensure data integrity and consistency;
[0005] S2: Seasonality-trend decomposition: decompose time series data into trend, seasonality and residual;
[0006] S3: Forecasting data: constructing a Transformer model; inputting the trend, seasonality, and residual into the Transformer model; the Transformer model generates forecast development data;
[0007] S4: Data integration: Integrate the prediction results of the Transformer model with the residual components, specifically by directly adding the prediction results and the residual components to generate the final epidemic prediction results, or by introducing weighted coefficients for fusion based on the statistical characteristics of the residuals.
[0008] Furthermore, the steps of removing outliers and filling missing values are specifically as follows:
[0009] The data is divided into normal values and abnormal values; the range of the normal value is expressed as: O(t)∈[μ-bσ,μ+bσ]; where O(t), μ, and σ represent an arbitrary sequence value, data mean, and standard deviation of the epidemic-related data, respectively; b is a multiple of the standard deviation, and the recommended value range is 2 to 3. The specific value can be determined through experiments based on the fluctuation characteristics of the epidemic data;
[0010] The calculation of the missing value is expressed as:
[0011]
[0012] Among them, x(t j )、x(t j+1 ) represents known data, t j , t j+1 Indicates the time sequence number of the known data; x interpolated (t lack ) indicates missing data, t lack Indicates the time sequence number of the missing data.
[0013] Furthermore, in step S2, the seasonality-trend decomposition process is as follows:
[0014] y t =T t +S t +R t
[0015] in:
[0016] y t is the original data,
[0017] T t is the trend component,
[0018] S t For seasonal ingredients,
[0019] R t is the residual component,
[0020] Use local regression method to smooth time series data and extract trend component T t The smoothing window size is recommended to be determined according to the time span and fluctuation frequency of the data, for example, a sliding window of 7 to 30 days can be selected, and the selection can be optimized through cross-validation; the trend component is subtracted from the original data to obtain the seasonal component S t ; Residual component R t It is the result of subtracting the trend and seasonal components from the original data.
[0021] Furthermore, the Transformer model is:
[0022]
[0023] in:
[0024] Q is the query matrix,
[0025] K is the bond matrix,
[0026] V is the value matrix,
[0027] d kis the dimension of the key;
[0028] The trend component T t As the query matrix Q, the seasonal component S t As the key matrix K, the original data y t As a value matrix V, the dimension of the key is d k Based on the balance between data complexity, computing resources, and model performance, it is set to an integer power of 2. For example, for daily epidemic data, dk can be set to 64 or 128. If the data volume is large or computing resources are limited, it can be reduced to 32. The specific value can be verified through experiments.
[0029] Furthermore, the Transformer model also includes a loss function, which is determined by the following steps:
[0030] Model training: The trend component T obtained in step S2 is t , seasonal component S t The input is input to the Transformer model, the mean square error (MSE) between the output result and the actual data is calculated, and the mean square error (MSE) is minimized through iterative optimization using the Adam optimization algorithm, which is used as the loss function.
[0031] Furthermore, step S4 also includes a step of displaying the prediction results by using a data visualization tool.
[0032] The epidemic prediction method based on the STL and Transformer models of the present invention has strong adaptability, can process time series data of different types and complexities, and is particularly suitable for complex epidemic prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flow chart of the epidemic prediction method based on STL and Transformer models of the present invention. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and beneficial effects of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0035] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0036] See also Figure 1 The epidemic prediction method based on STL and Transformer model includes the following steps:
[0037] S1: Collection and preprocessing: Collect case numbers, time, and location data through the epidemic monitoring platform to form a time series; remove outliers and fill in missing values to ensure data integrity and consistency; the specific steps of removing outliers and filling in missing values are as follows:
[0038] The data are divided into normal values and abnormal values; the value range of the normal value is expressed as:
[0039] O(t)∈[μ-bσ,μ+bσ]; where O(t), μ, and σ represent an arbitrary sequence value, data mean, and standard deviation of the epidemic-related data, respectively; b is a multiple of the standard deviation, and is generally recommended to be in the range of 2 to 3. The specific value can be determined through experiments based on the fluctuation characteristics of the epidemic data;
[0040] The calculation of the missing value is expressed as:
[0041]
[0042] Among them, x(t j )、x(t j+1 ) represents known data, t j , t j+1 Indicates the time sequence number of the known data; x interpolated (t lack ) indicates missing data, t lack Indicates the serial number of the missing data. Furthermore, for time series data, considering the continuity and temporal characteristics of the data, methods such as moving average or exponential smoothing can be used to identify data points that deviate significantly from the overall trend. Depending on the characteristics of the epidemic data, statistical model-based anomaly detection or time series interpolation methods can be used for further optimization.
[0043] In addition to the number of cases, time and location data already mentioned, some key auxiliary data to support epidemic prediction can also be collected. For example, data related to population characteristics, including population density, mobility and degree of contact between people. Secondly, data on environmental factors, such as meteorological conditions, seasonal changes and other factors that may affect the spread of the virus. In addition, data related to prevention and control measures can also be processed data, including public health intervention data such as vaccination status, implementation of social distance measures, and nucleic acid testing volume. Finally, medical resource data such as the number of hospital beds and medical staff are also important information. For multi-dimensional data, key variables can be extracted through feature selection or dimensionality reduction methods (such as principal component analysis) before step S2, and then merged with the time series of the number of cases for STL decomposition, or used as additional input features of the Transformer model in step S3 to enhance the model's modeling ability for external factors.
[0044] S2: Seasonality-trend decomposition: decompose time series data into trend, seasonality and residual;
[0045] The process of seasonal-trend decomposition is:
[0046] y t =T t +S t +R t
[0047] in:
[0048] y t is the original data,
[0049] T t is the trend component,
[0050] S t For seasonal ingredients,
[0051] R t is the residual component,
[0052] Use local regression method to smooth time series data and extract trend component T t The smoothing window size is recommended to be determined according to the time span and fluctuation frequency of the data, for example, a sliding window of 7 to 30 days can be selected, and the selection can be optimized through cross-validation; the trend component is subtracted from the original data to obtain the seasonal component S t ; Residual component R t It is the result of subtracting the trend and seasonal components from the original data. The smoothing window size of the local regression method is determined based on the time span and volatility characteristics of the data, or the optimal value is selected through cross-validation.
[0053] Based on data characteristics, epidemic development patterns often differ significantly across regions. Each region has its own unique population density, climate conditions, and prevention and control measures, which can lead to different seasonal patterns and development trends in different regions. Therefore, when implementing this method for data decomposition, the data can be categorized by region into multiple "sub-data series," each of which can be subjected to STL decomposition to avoid masking the characteristic content of each region.
[0054] On the other hand, from a methodological perspective, performing a separate STL decomposition on each region better aligns with the fundamental assumptions of time series analysis. For example, performing a separate STL decomposition on the Guangdong Province case count-time series more accurately captures the region's unique trends and seasonal characteristics. This approach ensures that the resulting decomposition components are more meaningful and interpretable.
[0055] In terms of forecasting effectiveness, regional STL decomposition provides more accurate predictions. This is because each region's model is trained based on the characteristics of that region's historical data, better reflecting local characteristics and thus improving forecast accuracy. Although this method is computationally intensive, it produces more valuable analytical results.
[0056] S3: Forecasting data: constructing a Transformer model; inputting the trend, seasonality, and residual into the Transformer model; the Transformer model generates forecast development data;
[0057] The Transformer model is:
[0058]
[0059] in:
[0060] Q is the query matrix,
[0061] K is the bond matrix,
[0062] V is the value matrix,
[0063] d k is the dimension of the key;
[0064] The trend component T t As the query matrix Q, the seasonal component S t As the key matrix K, the original data y t As a value matrix V, the dimension of the key is d k Based on the balance between data complexity, computing resources, and model performance, dk is set to an integer power of 2. For example, for daily epidemic data, dk can be set to 64 or 128. If the data volume is large or computing resources are limited, it can be reduced to 32. The specific value can be verified through experimentation. Furthermore, the input data is normalized or standardized before entering the Transformer model to ensure data scale consistency.
[0065] The trend reflects the main direction of the data and is suitable as the basis for queries. In this embodiment, the trend component (Tt) is used as the query matrix (Q). The seasonal component (St) contains periodic regular information and can help the model establish time dependencies. In this embodiment, it is input as the key matrix (K). The original data (yt) contains complete information and helps generate the final predicted value. In this embodiment, it is input as the value matrix (V). The choice of the key dimension (dk) needs to consider the balance between data complexity, computing resources and model performance. Using a power of 2 value, a larger dimension can capture more complex patterns, but it also requires more computing resources.
[0066] Furthermore, the Transformer model also includes a loss function, which is determined by the following steps:
[0067] Model training: The trend component T obtained in step S2 is t , seasonal component S t Input the data into the Transformer model, calculate the mean squared error (MSE) between the output and the actual data, and iteratively optimize the Adam optimization algorithm to minimize the MSE, which serves as the loss function. Depending on the data characteristics, you can adjust the Adam learning rate or use other loss functions such as smoothed L1 loss to improve model robustness.
[0068] The loss function calculates the mean squared error (MSE) between the model's predicted values and the actual values for each model training batch. This error is used to evaluate the model's predictive accuracy and to update model parameters through backpropagation. The MSE loss also considers the characteristics of time series, assigning different weights to trend and seasonal forecast errors. Finally, combined with the Adam optimizer, this loss function is minimized through iterative optimization, continuously improving the model's predictive performance. The choice and use of the loss function directly impacts the model's training results and predictive accuracy.
[0069] S4: Data integration: Integrate the prediction results of the Transformer model with the residual components to generate the final epidemic prediction results; also includes the step of displaying the prediction results by using data visualization tools.
[0070] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. The epidemic prediction method based on STL and Transformer model is characterized by: The steps include: S1: Collection and preprocessing: Collect case numbers, time, and location data through the epidemic monitoring platform to form a time series; remove outliers, fill in missing values, and ensure data integrity and consistency; S2: Seasonality-trend decomposition: decompose time series data into trend, seasonality and residual; S3: Forecasting data: constructing a Transformer model; inputting the trend, seasonality, and residual into the Transformer model; the Transformer model generates forecast development data; S4: Data integration: Integrate the prediction results of the Transformer model with the residual components to generate the final epidemic prediction results.
2. The epidemic prediction method based on STL and Transformer model according to claim 1 is characterized in that: The steps of removing outliers and filling missing values are specifically as follows: The data are divided into normal values and abnormal values; the value range of the normal value is expressed as: O(t)∈[μ-bσ,μ+bσ]; where O(t), μ, and σ represent an arbitrary sequence value, data mean, and standard deviation of the epidemic-related data, respectively; b is a multiple of the standard deviation; The calculation of the missing value is expressed as: Among them, x(t j )、x(t j+1 ) represents known data, t j , t j+1 Indicates the time sequence number of the known data; x interpolated (t lack ) indicates missing data, t lack Indicates the time sequence number of the missing data.
3. The epidemic prediction method based on STL and Transformer model according to claim 1 is characterized in that: In step S2, the seasonality-trend decomposition process is as follows: y t =T t +S t +R t in: y t is the original data, T t is the trend component, S t For seasonal ingredients, R t is the residual component, Use local regression method to smooth time series data and extract trend component T t ; Subtract the trend component from the original data to obtain the seasonal component S t ; Residual component R t It is the result of subtracting the trend and seasonal components from the original data.
4. The epidemic prediction method based on STL and Transformer model according to claim 3 is characterized in that: The Transformer model is: in: Q is the query matrix, K is the bond matrix, V is the value matrix, d k is the dimension of the key; The trend component T t As the query matrix Q, the seasonal component S t As the key matrix K, the original data y t As a value matrix V, the dimension of the key is d k It is set to an integer power of 2 based on the balance between data complexity, computing resources, and model performance.
5. The epidemic prediction method based on STL and Transformer model according to claim 4 is characterized in that: The Transformer model also includes a loss function, which is determined by the following steps: Model training: The trend component T obtained in step S2 is t , seasonal component S t The input is input to the Transformer model, the mean square error (MSE) between the output result and the actual data is calculated, and the mean square error (MSE) is minimized through iterative optimization using the Adam optimization algorithm, which is used as the loss function.
6. The epidemic prediction method based on STL and Transformer model according to claim 5 is characterized in that: Step S4 also includes a step of displaying the prediction results by using a data visualization tool.