Epidemic situation prediction method based on modal decomposition network and deep learning
By combining the modal decomposition network and the Transformer model, the limitations of the LSTM model when processing complex time series data are solved, high-precision prediction of nonlinear and non-stationary epidemic data is achieved, the stability and adaptability of the model are enhanced, and more accurate prediction tools are provided for the public health field.
Patent Information
- Application Number
- CN202510110602.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-06
AI Technical Summary
Existing LSTM models are difficult to effectively separate multi-frequency and multi-scale features when processing complex time series data, resulting in limited prediction accuracy and are prone to overfitting or underfitting when processing large-scale or bursting data.
Combining the modal decomposition network (MDN) and the Transformer model, complex time series data is decomposed into several modal components through MDN, and the self-attention mechanism of the Transformer model is used to capture long-term dependencies in the time series to achieve high-precision prediction.
It improves the prediction accuracy of complex time series data, enhances the stability and generalization capabilities of the model, adapts to epidemic data in different scenarios, provides more accurate and efficient prediction tools, and provides important technical support for public health decision-making.
Smart Images

Figure CN120108765A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to an epidemic prediction method based on modal decomposition network and deep learning. Background Art
[0002] In order to understand and evaluate the possible development trend and impact of epidemic spread in advance, so as to provide scientific basis for public health decision-making, optimize resource allocation, formulate effective prevention and control strategies, and reduce the impact of epidemic on human health, social economy and daily life, various institutions are currently actively adopting a variety of technical means to strengthen the monitoring, analysis and prediction capabilities of epidemic. The conventional methods of epidemic prediction mainly include: methods based on time series prediction and methods based on epidemiological models. Among them, the time series prediction method uses time series data (such as daily newly confirmed cases, cured people, deaths, etc.) for modeling and prediction. Epidemiological models predict the development of the epidemic by analyzing the changes in the population during the spread of the epidemic. Commonly used epidemiological models include SIR model and SEIR model. The former divides the population into three categories: susceptible (S), infected (I) and removed (R). By establishing differential equations to describe the changes in the number of these three types of people, it is possible to predict when the epidemic will reach the inflection point, peak and end. The latter adds the category of latent people (E) to the SIR model, that is, those who have been infected but have not yet shown symptoms, and ultimately more accurately describes the spread of the epidemic, especially in infectious diseases with a long incubation period. For example, the patent document with publication number CN115662651A provides an EMD-LSTM epidemic prediction method based on the traffic network, which integrates multiple algorithm models of EMD and LSTM, introduces traffic network data, and improves the accuracy of epidemic prediction.
[0003] However, the LSTM model is only applicable to nonlinear and non-stationary time series data, so as to be able to learn long-term dependencies. However, it often shows limitations when facing nonlinear and non-stationary complex time series. It is difficult to effectively separate the multi-frequency and multi-scale features contained in the time series in complex data decomposition, resulting in limited prediction accuracy. On the other hand, it is less efficient in capturing long-term dependencies in epidemic data, especially when processing data with large-scale data volumes or burst characteristics, which is prone to overfitting or underfitting. Summary of the invention
[0004] In response to the problems existing in the prior art, the present invention provides an epidemic prediction method based on mode decomposition network and deep learning, which organically combines the mode decomposition network (MDN) and the Transformer model, decomposes complex time series data into several modal components through MDN, and uses the self-attention mechanism of the Transformer model to capture the long-term dependencies in the time series, thereby achieving high-precision prediction of nonlinear and non-stationary data.
[0005] The technical solution of the present invention is achieved in this way:
[0006] A method for epidemic prediction based on modal decomposition network and deep learning, comprising the following steps:
[0007] S1. Collection and preprocessing: Collect various types of epidemic-related data, clean and normalize the epidemic-related data to obtain raw data; epidemic-related data include the number of new cases, cumulative number of infections, population density, distribution of medical resources, temperature, and humidity. Epidemic-related data are collected from public health monitoring systems or open databases.
[0008] S2, modal decomposition: construct a modal decomposition network; input the original data into the modal decomposition network to generate multiple modal components, and the multiple modal components constitute a modal matrix; one type of the original data corresponds to one modal matrix; each modal component represents a different frequency or time scale feature in the sequence signal. This process aims to reduce the complexity of the original time series and improve the prediction accuracy and computational efficiency of the subsequent model.
[0009] S3, predicting data: constructing a Transformer model; inputting the modal matrix into the Transformer model; the Transformer model generates predicted development data; one modal matrix corresponds to one predicted development data;
[0010] S4, data integration: integrating the plurality of forecast development data to obtain a final result; and outputting the final result;
[0011] The epidemic-related data, the original data, the modal matrix, the predicted development data and the final result are all time series, respectively represented by i 、x i 、M i , and The corresponding sequence values are O i (t j ), x i (t j )、Mi (t j ), and Wherein, i represents the serial number of the type of the corresponding original data; t j represents the time, and j represents the sequence number of the time.
[0012] A time series is a set of data points that are arranged in chronological order and are usually collected at continuous and regular time intervals. Time series data is used to analyze and predict trends, seasonal patterns, cyclical changes, and other relevant time-dependent characteristics of a variable over time.
[0013] The time series interval is in seconds, minutes, days, months, quarters, or years.
[0014] In the above data, various types of epidemic-related data are represented as O 1 , O 2 , O 3 ...O n ; The same type of data is represented as O 1 (t 1 ), O 1 (t 2 ), O 1 (t 3 )……O 1 (t m ). The same is true for other data. n is the number of types of epidemic-related data; m is the number of time series.
[0015] The modal decomposition network (MDN) decomposes complex time series into several modal components through a deep learning network structure. Each modal component corresponds to a different frequency or time scale feature in the signal, thereby significantly reducing the complexity of the data and providing a clearer feature structure for subsequent modeling. Compared with traditional linear decomposition methods such as empirical mode decomposition (EMD), MDN has higher adaptability and stability and can better capture subtle features in the data.
[0016] The Transformer model uses its self-attention mechanism to efficiently capture long-term dependencies within and between modal components. Compared with traditional recurrent neural networks (RNN, LSTM, etc.), the Transformer significantly improves model processing efficiency through parallel computing, while being able to focus on long-distance feature associations in time series. This structure has shown excellent performance in feature modeling of complex time series, and is particularly suitable for the analysis of multivariate long time series in epidemic prediction.
[0017] By combining the MDN and Transformer models, the present invention not only improves the prediction accuracy of complex time series data, but also enhances the stability and generalization ability of the model. This method can adapt to epidemic data in different scenarios, including non-stationary characteristics in emergencies and data distribution in different regions and time scales, providing a more accurate and efficient prediction tool for the public health field, and providing important technical support for making scientific epidemic prevention and control decisions.
[0018] As a further optimization of the above scheme, the data cleaning includes removing outliers and filling missing values;
[0019] The epidemic-related data are divided into normal values and abnormal values; the value range of the normal value is expressed as: O(t)∈[μ-bσ,μ+bσ]; wherein O(t), μ and σ represent an arbitrary sequence value, data mean and standard deviation of the epidemic-related data respectively; b is a multiple of the standard deviation;
[0020] The calculation of the missing value is expressed as:
[0021]
[0022] Among them, x(t j )、x(t j+1 ) represents known data, t j ,t j+1 represents the time sequence number of the known data; x interpolated (t lack ) indicates missing data, t lack Indicates the time sequence number of the missing data.
[0023] As a further optimization of the above scheme, the normalization is to map the data to the interval [0,1]. The normalization process is expressed as:
[0024]
[0025] Among them, O i (t j ) and x i (t j ) are respectively the sequence value of the epidemic-related data and the sequence value of the corresponding original data; i , max and O i , min are the same type of epidemic-related data. i The maximum and minimum sequence values.
[0026] As a further optimization of the above solution, in step S2, the process of modal decomposition is expressed as:
[0027] M i =f i (x i ,θ i );
[0028] Among them, f i () is the nonlinear mapping function of the deep neural network, λ i Trainable parameters are parameters in a deep neural network that can be adjusted through the learning process, such as weights.
[0029] Mapping function f corresponding to different types of raw data i and trainable parameters θ i Each type of data (such as new cases, population density, etc.) has different feature patterns and requires a special mapping function to process their respective features. The parameter θ i Need to be trained and optimized separately for specific types of data.
[0030] Traditional modal decomposition methods such as EMD, EEMD, and CEEMD are all iterative screening algorithms based on signal processing, without trainable parameters, and have poor flexibility when processing nonlinear and non-stationary data. This solution uses a deep neural network for nonlinear mapping, contains a trainable parameter θ, and optimizes through a loss function composed of reconstruction error and regularization terms. It can learn the optimal decomposition strategy in an end-to-end manner and better adapt to specific data characteristics.
[0031] The original data is represented as:
[0032] The modal matrix is expressed as:
[0033] Where n is the number of sequence values of the same original data, k is the number of modal components of the same modal matrix, and n>k.
[0034] As a further optimization of the above solution, the modal decomposition network also includes a first loss function, which is expressed as:
[0035]
[0036] The term on the left side of the plus sign in the equation is the reconstruction error, which is used to ensure that the modal components can reconstruct the original time series; Reg(M i ) is a regularization term used to control the complexity of each modal component; λ is a preset regularization weight parameter; L MDN is the first loss value.
[0037] The smaller the first loss value, the better the model is fitting the data and the closer the predicted value is to the real signal. By continuously reducing the first loss value, the decomposition model is optimized.
[0038] As a further optimization of the above scheme, the regularization process is expressed as: That is M i The squared norm of the gradient vector .
[0039] As a further optimization of the above scheme, the Transformer model includes a self-attention mechanism, and the result value of the self-attention mechanism is calculated as follows:
[0040]
[0041] Where Q, K and V are query, key and value matrices respectively; W q , W k and W v are weight matrices respectively; d k is the dimension of the key, used to ensure numerical stability).
[0042] The Transformer model consists of an encoder layer and a decoder layer. The encoder layer contains a multi-head self-attention mechanism and a feed-forward fully connected layer to capture the contextual information of the input data and the internal dependencies of the sequence; the decoder layer also contains a multi-head self-attention mechanism and a feed-forward fully connected layer, but adds an encoder-decoder attention mechanism to generate accurate output sequences. Each layer uses residual connections and layer normalization to stabilize the training process and accelerate convergence. The final output layer converts the output of the decoder into a prediction result.
[0043] The result value of the self-attention mechanism is used to measure the correlation between sequence values at different times in the sequence.
[0044] As a further optimization of the above solution, the Transformer model also includes a second loss function, expressed as:
[0045]
[0046] in, and i Represent the predicted value and the true value respectively; N is the number of the predicted values; λ is the preset regularization weight parameter; L Transformer is the second loss value.
[0047] The second loss value is used to measure the difference between the model prediction output and the actual target output, guiding the model to optimize model performance and improve prediction accuracy.
[0048] By minimizing this loss function, the model can continuously adjust its parameters to reduce the gap between the predicted value and the actual value. The smaller the loss value, the closer the model's predicted output is to the actual target output, thereby achieving the purpose of optimizing model performance and improving prediction accuracy. This is an iterative optimization process, in which the model parameters are continuously updated through the back-propagation algorithm.
[0049] As a further optimization of the above scheme, the final result is calculated as:
[0050] As a further optimization of the above solution, the softmax() function is:
[0051]
[0052] Among them, x i and x j are the i-th and j-th elements of the vector respectively, and exp(·) is the exponential function.
[0053] The softmax function converts the input vector x into a probability distribution where each element has a value between 0 and 1 and the sum of all elements is 1.
[0054] Furthermore, the present invention also provides visualization results, including modal decomposition diagrams, prediction trend diagrams, and residual analysis diagrams. The modal decomposition diagram is used to display the original data (original time series) and its decomposed multimodal components; the prediction trend diagram is used to display the final results, that is, to predict the future epidemic trend and confidence interval; the residual analysis diagram is used to display the distribution characteristics of the evaluation prediction error.
[0055] Compared with the prior art, the present invention achieves the following beneficial effects:
[0056] The present invention provides an epidemic prediction method based on modal decomposition network and deep learning. By combining MDN and Transformer models, high-precision prediction of nonlinear and non-stationary epidemic data is achieved. MDN effectively reduces data complexity, and Transformer captures the long-term dependence of time series. The two work together to improve the accuracy and adaptability of prediction, which is particularly suitable for decision support in complex epidemic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flow chart of an epidemic prediction method based on modal decomposition network and deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solution and advantages of the present invention more clear, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] like Figure 1 As shown, this embodiment provides an epidemic prediction method based on modal decomposition network and deep learning, which involves the operation of multiple data, including epidemic-related data, original data, modal matrix, predicted development data and final results are all time series, respectively represented by O i 、x i 、M i , and The corresponding sequence values are O i (t j ), x i (t j )、M i (t j ), and Where i represents the serial number of the type of the corresponding original data; t j represents the time, and j represents the sequence number of the time.
[0060] The following steps are involved:
[0061] S1. Collection and preprocessing: Collect various types of epidemic-related data, clean and normalize the epidemic-related data, and obtain the original data; epidemic-related data include the number of new cases, cumulative number of infections, population density, distribution of medical resources, temperature, and humidity. Epidemic-related data are collected from public health monitoring systems or open databases.
[0062] In this embodiment, data cleaning includes removing outliers and filling missing values;
[0063] Epidemic-related data are divided into normal values and abnormal values; the range of normal values is expressed as:
[0064] O(t)∈[μ-bσ,μ+bσ]; wherein O(t), μ and σ represent an arbitrary sequence value, data mean and standard deviation of epidemic-related data, respectively; b is a multiple of the standard deviation; in this embodiment, the value of b is 3.
[0065] The calculation of missing values is expressed as:
[0066]
[0067] Among them, x(t j )、x(t j+1 ) represents known data, t j ,t j+1 Indicates the time sequence number of known data;
[0068] x interpolated (t lack ) indicates missing data, t lack Indicates the time sequence number of missing data.
[0069] In this embodiment, normalization is to map the data to the interval [0, 1]. The normalization process is expressed as:
[0070]
[0071] Among them, O i (t j ) and x i (t j ) are the sequence values of epidemic-related data and the sequence values of the corresponding original data respectively; i,max and O i,min They are the same type of epidemic-related data O i The maximum and minimum sequence values.
[0072] S2. Modal decomposition: construct a modal decomposition network; input the original data into the modal decomposition network to generate multiple modal components, and the multiple modal components constitute a modal matrix;
[0073] In this embodiment, the process of modal decomposition is expressed as:
[0074] M i =f i (x i ,θ i );
[0075] Among them, f i () is the nonlinear mapping function of the deep neural network, θ i Trainable parameters are parameters in a deep neural network that can be adjusted through the learning process, such as weights.
[0076] Mapping function f corresponding to different types of raw data i and trainable parameters θ i Each type of data (such as new cases, population density, etc.) has different feature patterns and requires a special mapping function to process their respective features. The parameter θ iNeed to be trained and optimized separately for specific types of data.
[0077] Traditional modal decomposition methods such as EMD, EEMD, and CEEMD are all iterative screening algorithms based on signal processing, without trainable parameters, and have poor flexibility when processing nonlinear and non-stationary data. This solution uses a deep neural network for nonlinear mapping, contains a trainable parameter θ, and optimizes through a loss function composed of reconstruction error and regularization terms. It can learn the optimal decomposition strategy in an end-to-end manner and better adapt to specific data characteristics.
[0078] The original data is represented as:
[0079] The modal matrix is expressed as:
[0080] Where n is the number of sequence values of the same original data, k is the number of modal components of the same modal matrix, and n>k.
[0081] One type of raw data corresponds to one modal matrix; each modal component represents a different frequency or time scale feature in the sequence signal. This process aims to reduce the complexity of the original time series and improve the prediction accuracy and computational efficiency of the subsequent model.
[0082] In this embodiment, the optimization of the modal decomposition network is also involved. The modal decomposition network includes a first loss function, that is, the first loss is worth calculating, which is expressed as:
[0083]
[0084] The term on the left side of the plus sign in the equation is the reconstruction error, which is used to ensure that the modal components can reconstruct the original time series; Reg(M i ) is a regularization term used to suppress high-frequency noise and control the complexity of each modal component; λ is a preset regularization weight parameter; L MDN is the first loss value. In this embodiment, the regularization process is expressed as: That is M i The squared norm of the gradient vector .
[0085] The smaller the first loss value, the better the model is fitting the data and the closer the predicted value is to the real signal. By continuously reducing the first loss value, the decomposition model is optimized.
[0086] S3, prediction data: construct a Transformer model; input the modal matrix into the Transformer model; the Transformer model generates prediction development data; one modal matrix corresponds to one prediction development data;
[0087] In this embodiment, the Transformer model includes a self-attention mechanism, and the result value of the self-attention mechanism is calculated as follows:
[0088]
[0089] Where Q, K and V are query, key and value matrices respectively; W q , W k and W v are weight matrices respectively; d k is the dimension of the key, used to ensure numerical stability). In this embodiment, the softmax() function is:
[0090]
[0091] Among them, x i and x j are the i-th and j-th elements of the vector respectively, and exp(·) is the exponential function.
[0092] The softmax function converts the input vector x into a probability distribution where each element has a value between 0 and 1 and the sum of all elements is 1.
[0093] The Transformer model consists of an encoder layer and a decoder layer. The encoder layer contains a multi-head self-attention mechanism and a feed-forward fully connected layer to capture the contextual information of the input data and the internal dependencies of the sequence; the decoder layer also contains a multi-head self-attention mechanism and a feed-forward fully connected layer, but adds an encoder-decoder attention mechanism to generate accurate output sequences. Each layer uses residual connections and layer normalization to stabilize the training process and accelerate convergence. The final output layer converts the output of the decoder into a prediction result.
[0094] The result value of the self-attention mechanism is used to measure the correlation between sequence values at different times in the sequence.
[0095] In this embodiment, the optimization of the Transformer model is also involved. The Transformer model also includes the calculation of the second loss function, which is expressed as:
[0096]
[0097] in, and i Represent the predicted value and the true value respectively; N is the number of predicted values; λ is the preset regularization weight parameter; L Transformer is the second loss value.
[0098] The second loss value is used to measure the difference between the model prediction output and the actual target output, guiding the model to optimize model performance and improve prediction accuracy.
[0099] By minimizing this loss function, the model can continuously adjust its parameters to reduce the gap between the predicted value and the actual value. The smaller the loss value, the closer the model's predicted output is to the actual target output, thereby achieving the purpose of optimizing model performance and improving prediction accuracy. This is an iterative optimization process, in which the model parameters are continuously updated through the back-propagation algorithm.
[0100] S4, data integration: integrating multiple forecast development data to obtain the final result; outputting the final result. In this embodiment, the calculation of the final result is:
[0101] A time series is a set of data points that are arranged in chronological order and are usually collected at continuous and regular time intervals. Time series data is used to analyze and predict trends, seasonal patterns, cyclical changes, and other relevant time-dependent characteristics of a variable over time.
[0102] The time series interval is in seconds, minutes, days, months, quarters, or years.
[0103] In the above data, various types of epidemic-related data are represented as O 1 , O 2 , O 3 ...O n ; The same type of data is represented as O 1 (t 1 ), O 1 (t 2 ), O 1 (t 3 )……O 1 (t m ). The same is true for other data. n is the number of types of epidemic-related data; m is the number of time series.
[0104] The modal decomposition network (MDN) decomposes complex time series into several modal components through a deep learning network structure. Each modal component corresponds to a different frequency or time scale feature in the signal, thereby significantly reducing the complexity of the data and providing a clearer feature structure for subsequent modeling. Compared with traditional linear decomposition methods such as empirical mode decomposition (EMD), MDN has higher adaptability and stability and can better capture subtle features in the data.
[0105] The Transformer model uses its self-attention mechanism to efficiently capture long-term dependencies within and between modal components. Compared with traditional recurrent neural networks (RNN, LSTM, etc.), the Transformer significantly improves model processing efficiency through parallel computing, while being able to focus on long-distance feature associations in time series. This structure has shown excellent performance in feature modeling of complex time series, and is particularly suitable for the analysis of multivariate long time series in epidemic prediction.
[0106] By combining the MDN and Transformer models, the present invention not only improves the prediction accuracy of complex time series data, but also enhances the stability and generalization ability of the model. This method can adapt to epidemic data in different scenarios, including non-stationary characteristics in emergencies and data distribution in different regions and time scales, providing a more accurate and efficient prediction tool for the public health field, and providing important technical support for making scientific epidemic prevention and control decisions.
[0107] According to the disclosure and teaching of the above description, those skilled in the art to which the present invention belongs may also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the scope of protection of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for the convenience of description and do not constitute any limitation to the present invention.
Claims
1. An epidemic prediction method based on modal decomposition network and deep learning, characterized in that: The following steps are involved: S1. Collection and preprocessing: Collect various types of epidemic-related data, clean and normalize the epidemic-related data, and obtain original data; S2, modal decomposition: construct a modal decomposition network; Inputting the original data into the modal decomposition network to generate multiple modal components, wherein the multiple modal components constitute a modal matrix; one type of the original data corresponds to one modal matrix; S3, predicting data: constructing a Transformer model; inputting the modal matrix into the Transformer model; the Transformer model generates predictive development data; One said modal matrix corresponds to one said predicted development data; S4, data integration: integrating the plurality of forecast development data to obtain a final result; and outputting the final result; The epidemic-related data, the original data, the modal matrix, the predicted development data and the final result are all time series, respectively represented by i 、x i 、M i , and The corresponding sequence values are O i (t j ), x i (t j ), M i (t j ), and Wherein, i represents the serial number of the type of the corresponding original data; t j represents the time, and j represents the sequence number of the time.
2. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The data cleaning includes removing outliers and filling missing values; The epidemic-related data are divided into normal values and abnormal values; the value range of the normal value is expressed as: O(t)∈[μ-bσ,μ+bσ]; wherein, (t), μ and σ represent an arbitrary sequence value, data mean and standard deviation of the epidemic-related data respectively; b is a multiple of the standard deviation; The calculation of the missing value is expressed as: Among them, x(t j )、x(t j+1 ) represents known data, t j ,t j+1 Indicates the time sequence number of the known data; x interpolated (t lack ) indicates missing data, t lack Indicates the time sequence number of the missing data.
3. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The normalization is to map the data to the interval [0,1]. The normalization process is expressed as: Among them, O i (t j ) and x i (t j ) are respectively the sequence value of the epidemic-related data and the sequence value of the corresponding original data; i , max and O i , min are the same type of epidemic-related data. i The maximum and minimum sequence values.
4. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: In step S2, the process of modal decomposition is expressed as: M i =f i (x i,i ); Among them, f i () is the nonlinear mapping function of the deep neural network, θ i is the trainable parameter of the deep neural network; The original data is represented as: The modal matrix is expressed as: Where n is the number of sequence values of the same original data, k is the number of modal components of the same modal matrix, and n>k.
5. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The modal decomposition network also includes a first loss function, expressed as: The term on the left side of the plus sign in the equation is the reconstruction error, which is used to ensure that the modal components can reconstruct the original time series; Reg(M i ) is a regularization term used to control the complexity of each modal component; λ is a preset regularization weight parameter; L MDN is the first loss value.
6. The epidemic prediction method based on modal decomposition network and deep learning according to claim 5 is characterized in that: The regularization process is expressed as: That is M i The squared norm of the gradient vector .
7. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The Transformer model includes a self-attention mechanism, and the result value of the self-attention mechanism is calculated as follows: Where Q, K and V are query, key and value matrices respectively; W q , W k and W v are weight matrices respectively; d k is the dimension of the key, used to ensure numerical stability).
8. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The Transformer model also includes a second loss function, expressed as: in, and i Represent the predicted value and the true value respectively; N is the number of the predicted values; λ is the preset regularization weight parameter; L Transformer is the second loss value.
9. The epidemic prediction method based on modal decomposition network and deep learning according to claim 1 is characterized in that: The final result is calculated as:
10. The epidemic prediction method based on modal decomposition network and deep learning according to claim 7 is characterized in that: The softmax() function is: Among them, x i and x j are the i-th and j-th elements of the vector respectively, and exp(·) is the exponential function.
Citation Information
Patent Citations
EMD-LSTM epidemic situation prediction method based on traffic network
CN115662651A