Solar wind speed prediction method based on fusion prediction mode and collaborative attention mechanism
Through deep learning models based on fusion prediction mode and synergistic attention mechanism, the problem of insufficient information utilization and uncaptured feature relationships in solar wind speed prediction is solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510345307.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing solar wind speed prediction methods fail to make full use of historical input information and fail to effectively capture the implicit relationships between different feature data, resulting in insufficient prediction accuracy.
Deep learning models based on fusion prediction mode and synergistic attention mechanism are adopted, and the relationship between pattern features and variables in historical data are deeply studied through prediction mode learning and fusion module, feature extraction module and flat layer, and the relationship between variables is dynamically evaluated using the synergistic attention mechanism.
Improve the accuracy and reliability of solar wind speed prediction, and improve prediction accuracy by better mining and analyzing the complex dependencies between input and output.
Smart Images

Figure CN120277610A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of solar wind speed prediction, and particularly relates to a solar wind speed prediction method based on a fusion prediction mode and a collaborative attention mechanism. Background Art
[0002] Solar wind speed prediction refers to predicting the solar wind speed after a certain period of time based on historical solar-related monitoring data and other relevant information, which is an important task. With the preparation of human space missions to the moon and Mars, the impact of space weather on the Earth and the solar system has become increasingly important. More and more high-tech systems are exposed to the space environment, and the impact of space weather on humans is also growing. In the solar system, space weather is mainly affected by the solar wind and the interplanetary counterparts of coronal mass ejections (CMEs) - interplanetary coronal mass ejections (ICMEs). Therefore, establishing a reliable solar wind prediction model can reduce the harm caused by space weather to human society and is also of great significance for establishing an effective space weather forecasting system.
[0003] Over the years, many studies have also been conducted on predicting the characteristics of the solar wind. Generally speaking, three main methods are used: manual monitoring, physics-based models, and empirical and semi-empirical models. Initially, some experts judged the state of the near-Earth solar wind by observing the activities of the sun in real time, using relevant domain knowledge and combining the solar wind speed measured by the WIND and ACE satellites located at Lagrangian point 1 (L1). Subsequently, scholars developed a series of physics-based models, mainly using magnetohydrodynamics (MHD) models and related improvements for prediction. Taking the bottom of the corona as the inner boundary, mathematical descriptions of physical processes such as the acceleration and heating of the solar wind are added to the equations to generate a more realistic solar wind background. Another type is empirical and semi-empirical models, aiming to establish the relationship between interplanetary space parameters and solar observation parameters. The early classic one is the Wang-Sheeley-Arge (WSA) model, which can predict the solar wind speed and interplanetary magnetic field polarity at L1 point 3 to 4 days in advance using the observed solar magnetic map as input. Subsequently, the development of machine learning has provided new ideas for this type of model. It can well learn the embedding relationship between input and output, and a series of models based on artificial neural networks (ANN), convolutional neural networks (CNN), long short-term memory recurrent neural networks (LSTM), etc. have emerged to predict the solar wind speed.
[0004] Although certain achievements have been made in the practice of the solar wind speed prediction task, there are still some problems and challenges. Currently, existing methods only consider learning prediction information from historical input information, and such utilization of information may not be sufficient; in addition, most models consider that each feature is equally important for the final prediction result. In fact, there are implicit relationships between the solar wind speed and data of different features at different times. For example, coronal holes and CMEs will last for a period of time after reaching the L1 point and affect the solar wind speed, which makes it difficult for most models to accurately predict the solar wind speed during this period. Summary of the Invention
[0005] In view of the above-mentioned prior art, the present invention proposes a solar wind speed prediction method based on a fusion prediction mode and a collaborative attention mechanism. This method deeply learns the core mode features of prediction information in historical data, and comprehensively excavates and analyzes the complex dependence relationship between input and output; through the collaborative attention mechanism, it dynamically and adaptively learns to evaluate the closeness and strength of the relationship between different variables, improving the accuracy and reliability of prediction.
[0006] To solve the above technical problems, a solar wind speed prediction method based on a fusion prediction mode and a collaborative attention mechanism proposed by the present invention uses a deep learning model including a prediction mode learning and fusion module, a feature extraction module, and a flattening layer; the input samples of the deep learning model during training and validation are historical data and prediction data, and the input sample during testing is historical data; the prediction mode learning and fusion module learns multiple modes of prediction information from the prediction data in the input of the deep learning model, and historical data guides the fusion of multiple modes of prediction information to obtain a prediction mode, and the prediction mode is updated during the backpropagation process in the training process of the deep learning model; the feature extraction module includes a first attention mechanism module, a second attention mechanism module, a downsampling module, and a collaborative attention mechanism module; the structures of the first attention mechanism module and the second attention mechanism module are the same; the first attention mechanism module receives the output of the prediction mode learning and fusion module and performs temporal feature learning on the output; the downsampling module performs a downsampling operation on the temporal feature representation output by the first attention mechanism module; the second attention mechanism module aggregates the global information of the temporal features obtained by the downsampling operation; the collaborative attention mechanism module learns the relationship between the output of the second attention mechanism module in the variable dimension to obtain a variable relationship feature representation; the flattening layer maps the variable relationship feature representation obtained by the feature extraction module to a solar wind speed value.
[0007] The solar wind speed prediction method of the present invention includes the following steps:
[0008] Step 1: Data preparation and processing:
[0009] Select five physical quantity attributes from the space weather dataset, namely the solar wind speed (Bulk speed), proton density (Proton density), proton temperature (Proton temperature), flow pressure (Flow pressure), and the vector length (Sigma-b) composed of the standard deviation of the average value of each component;
[0010] Binarize the ICME list as an attribute, set the time points with coronal mass ejection events to 1, and otherwise set to 0; denoted as ICME;
[0011] Calculate the sum of the pixel values of coronal holes in the Atmospheric Imaging Assembly (AIA) image as an attribute, which characterizes the influence of coronal holes, denoted as area;
[0012] Use the multivariate time series composed of the above 7 attribute data, including data for at least 3 consecutive years, sampled hourly; use this multivariate time series as the dataset input to the deep learning model, and divide this dataset into a training set, a validation set, and a test set in a ratio of 5:1:1 and in chronological order. The deep learning model makes predictions for 24 hours, 48 hours, 72 hours, and 96 hours respectively; the target outputs of the deep learning model are the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the observation sequence; the learning rate of the deep learning model is 0.0001 - 0.1, and the batch size is 64 - 512;
[0013] Perform standard deviation normalization (standardScale) processing on the data of each of the above attributes using Equation (1). After processing, the mean of the data is 0 and the standard deviation is 1.
[0014]
[0015] In Equation (1), x* is the processed data, x is the data before processing, μ is the mean of all sample data, and σ is the standard deviation of all sample data;
[0016] Step 2: Prediction mode learning and fusion:
[0017] The processed data first passes through the prediction mode learning and fusion module to learn and fuse the prediction mode, and the operations are as follows:
[0018] 2-1) Divide the standardized historical data into overlapping or non-overlapping historical patches X through the patching segmentation method p ; Let the historical data be X=(x1,…,x L ), x1 to x Lis a vector composed of the above 7 attributes, where L is the window of historical data input, L = 24 to 192, the patch length is P = 8 to 48, and the step size is S = (0.5 to 1)P. The historical patch X p has a quantity of N,
[0019] 2 - 2) Randomly initialize a prediction pattern Z = [z1,…,z n , where n is the number of patterns, n = 1 to 20, for subsequent storage of the patterns learned from the prediction data; at the same time, segment the prediction data in the same way as the historical data segmentation to obtain prediction patches y e ;
[0020] During the model training process, update the learning of different prediction patterns by minimizing the distance between a certain pattern representation z δ in the prediction pattern and the prediction patch y e ;
[0021]
[0022] As the loss function, it will guide the update of the prediction pattern in subsequent model training;
[0023] Use equation (3) to guide the fusion of the prediction pattern through the historical patch X p to obtain the fused patch
[0024]
[0025] where W x is the scientific department weight matrix, b x and b z are the bias terms; the historical patch X p and the fused patch are concatenated as the historical prediction feature representation as the output of the prediction pattern learning and fusion module;
[0026] Step 3, Feature Representation Learning:
[0027] 3 - 1) Temporal Feature Learning, perform multi - head attention calculation on the output of the prediction pattern learning and fusion module through the first attention mechanism module to obtain the temporal feature representation. The output of the attention mechanism is:
[0028]
[0029] In equation (4), Q i , K i , V i are generated by non - linear projection of the historical prediction feature representation, is the scaling factor; at this time, the attention mechanism operation is performed independently for each channel, (·) i The relevant variable corresponding to the i-th dimensional attribute;
[0030] 3-2) Relationship learning between variables:
[0031] Use the downsampling module to perform downsampling on the temporal feature representation, and aggregate the global information of the downsampled temporal feature representation through the second attention mechanism module. The specific operations are as follows:
[0032] p i = Downsample(O i ), p i′ = Attention(p i , O i , O i ) (5)
[0033] Downsample(·) reduces the N + 1 patches in O i to M patches using a multi-layer perceptron, where M = 1 to N;
[0034] In the second attention mechanism module, the downsampled p i is used as the query, and the original O i is used as the key and value, and the long-range modeling ability of the attention mechanism is used to dynamically aggregate the global information to obtain the feature representation p i′ ;
[0035] The co-attention mechanism module captures the feature p i′ aggregating the global information obtained by the second attention mechanism module in the variable dimension, representing the global cointegration relationship between variables. The specific operations are shown in equations (6) and (7):
[0036]
[0037] LayerNorm(·) in equations (6) and (7) is a common layer normalization operation in Transformer, which calculates the mean and variance by statistically calculating the values of all dimensions on each sample; p′ D is p i′ , i = 1, …, M; the representation in the variable dimension, is the feature representation of the relationship between variables obtained through the co-attention mechanism and mapped back to the original dimension;
[0038] Step 4, Output of the predicted solar wind speed:
[0039] The feature representation of the relationship between variables obtained by the co-attention mechanism module Extract the features of the patch corresponding to the predicted position Indicates that the predicted result of the solar wind speed is obtained by using a flat layer with a linear head to obtain the model output,
[0040]
[0041] In Equation (8), Flatten(·) flattens the features and then the flattened feature representation is used for prediction output through the linear layer Linear(·);
[0042] Step 5. Optimize the model to obtain the final predicted result of the solar wind speed:
[0043] Use the trained deep learning model to test the data in the test set, and obtain the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the future of the distance observation sequence respectively; if the predicted value of the solar wind speed reaches the expectation, the predicted result of the solar wind speed output by the model obtained in Step 4 is the final predicted result of the solar wind speed, otherwise, return to Step 2 to adjust the model parameters and retrain.
[0044] Furthermore, in the solar wind speed prediction method of the present invention, in Step 5, the adjusted model parameters include: the learning rate and batch size of the deep learning model; the window size L, patch length P, and step size S of the historical data input in Step 2-1); the number of patterns n in Step 2-1); the number of patches M reduced from N + 1 patches in Step 3-2).
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] Since in the prediction method of the present invention, the pattern features of historical prediction information are considered and the co-attention mechanism is used to fully learn the dependence relationship between different variables of the data, while the commonly used prediction methods in the prior art fail to consider the pattern information and the relationship between variables, the present invention is more suitable for the application scenario and improves a certain prediction accuracy. Description of the Drawings
[0047] Figure 1 is the overall framework of the solar wind speed prediction method of the present invention;
[0048] Figure 2 is Figure 1 the flowchart of the solar wind speed prediction method shown Detailed Embodiments
[0049] The basic concept of a solar wind speed prediction method based on a fusion prediction mode and a collaborative attention mechanism proposed by the present invention is to divide the multivariate time series of historical data and the prediction data into historical patches and predictions through the patching segmentation method. Multiple prediction modes are learned from the prediction patches, and the learned prediction modes are fused with the historical patches as weights to obtain the fused prediction mode information; the prediction mode information is concatenated with the historical patches and then enters the attention mechanism module to obtain a preliminary temporal feature representation; the temporal feature representation is downsampled to obtain the downsampled temporal feature representation, and the global information is further aggregated through the attention mechanism to encapsulate richer temporal information. The feature representation aggregated with the global information captures the global cointegration relationship between variables in the variable dimension through the collaborative attention mechanism to obtain the variable relationship feature representation. Only the feature representation corresponding to the prediction position of the variable relationship feature representation is taken and enters the flat layer with a linear head to obtain the final solar wind speed prediction result, as shown in Figure 1 , the deep learning model in the present invention includes a prediction mode learning and fusion module, a feature extraction module, and a flat layer; the input samples of the deep learning model during training and validation are historical data and prediction data, and the input sample during testing is historical data; the prediction mode learning and fusion module learns multiple modes of prediction information from the prediction data in the input of the deep learning model, and the historical data guides the fusion of multiple modes of prediction information to obtain the prediction mode, and the update of the prediction mode is realized during the backpropagation process in the training process of the deep learning model; the feature extraction module includes a first attention mechanism module, a second attention mechanism module, a downsampling module, and a collaborative attention mechanism module; the structures of the first attention mechanism module and the second attention mechanism module are the same; the first attention mechanism module receives the output of the prediction mode learning and fusion module and performs temporal feature learning on the output; the downsampling module performs a downsampling operation on the temporal feature representation output by the first attention mechanism module; the second attention mechanism module aggregates the global information of the temporal features obtained by the downsampling operation; the collaborative attention mechanism module learns the relationship between the outputs of the second attention mechanism module in the variable dimension to obtain the variable relationship feature representation; the flat layer maps the variable relationship feature representation obtained by the feature extraction module to the solar wind speed value.
[0050] The method flow of the present invention is as Figure 2As shown, first, collect multivariate data and perform certain data processing and standardization; conduct pattern learning and fusion on the prediction information, divide the data into patches, guide the learning pattern through the loss function in the prediction information, and utilize historical information to guide the fusion of the prediction information pattern to obtain the fused prediction pattern information; use the attention mechanism to learn temporal dependencies; perform downsampling and attention mechanism aggregation to obtain longer and richer global information, and learn the relationship features between variables in the variable dimension through the co-attention mechanism; set parameters to train the model; output the predicted results of the solar wind speed; test and evaluate the trained model on the test set. If it does not meet the expectations, such as industry standards or the results of certain papers, the parameters need to be readjusted to train the model again.
[0051] To make the objectives, technical methods, and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and specific embodiments. However, the following embodiments are by no means restrictive of the present invention.
[0052] Problem definition in the present invention: Let a set of multivariate time series sample historical input windows L of solar wind speed data be: (x1,..., x L ), where L is the size of the historical input window, and each x at time step t t is a vector with a dimension of D. The objective of the prediction task is to predict the future values after T time steps x L+T , and f(·) is the mapping relationship learned by the network.
[0053]
[0054] Taking the solar data observed within a certain period of time as an example, the method process of the present invention will be described in detail. Refer to Figure 1 and Figure 2 , and this solar wind speed prediction method includes the following steps:
[0055] Step 1: Data preparation and processing, including:
[0056] In this embodiment, five physical quantity attributes are selected from the OMNI dataset (https: / / omniweb.gsfc.nasa.gov / index.html) released by the National Aeronautics and Space Administration of the United States, namely solar wind speed (Bulk speed), proton density (Proton density), proton temperature (Proton temperature), flow pressure (Flowpressure), and the vector length (Sigma-b) composed of the standard deviation of the average value of each component; the ICME list (http: / / www.srl.caltech.edu / ACE / ASC / DATA / level3 / icmetable2.htm#) is binarized into a new attribute, that is, the time point when a coronal mass ejection event occurs is set to 1, otherwise 0, denoted as ICME; by calculating the sum of the pixel values of coronal holes in the Atmospheric Imaging Assembly (AIA) image as an attribute to characterize the influence of coronal holes, denoted as area; the multivariate time series of the above 7 attributes is used as the input of the model, including data from 2011 to 2017, sampled hourly, with data from 2011 to 2015 as the training set, data from 2016 as the validation set, and data from 2017 as the test set, and predictions are made for 24 hours, 48 hours, 72 hours, and 96 hours respectively. The target outputs are the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the observation sequence;
[0057] The data of each of the above attributes is processed using standard deviation standardization (standardScale) so that the processed data conforms to the standard normal distribution, that is, the mean is 0 and the standard deviation is 1.
[0058]
[0059] x* is the processed data, x is the data before processing, μ is the mean of all sample data, and σ is the standard deviation of all sample data;
[0060] Step 2: Prediction information pattern learning and fusion, including:
[0061] The standardized data is divided into overlapping or non-overlapping patches through the patching segmentation method. Let the historical data be X = (x1,…,x L )), x1 to x Lis a vector composed of L D-dimensional attributes. Here, D is 7 according to the dataset adopted, and L is the window size of data input, which is generally set to a multiple of 24, such as 24, 48, 72, …, 192. In this embodiment, 96 is adopted. An input window that is too short will result in less information, while an input window that is too long will cause the network to be unable to fully learn. Let the patch length be P and the stride be S, and the stride S = (0.5 - 1)P. The size of P is 8 - 48. In this embodiment, P = 16 and S = 8, but the size of P cannot exceed the input window length, and different input window lengths and prediction target lengths are applicable to different P sizes. The appropriate P size can be obtained through experiments within the above settings. During the segmentation process, a series of historical patches X will be generated p , and the number of historical patches is N. Since S repeated last values are filled at the end of the original sequence, so
[0062]
[0063] Through the patching segmentation method, the number of inputs can be reduced from L to approximately L / S, and the memory and computational complexity of the subsequent attention mechanism are reduced quadratically; randomly initialize a variable Z = [z1, …, z n , where n is the number of patterns, which is used to store the pattern information learned from the prediction data later. n = 1 - 20. In this embodiment, n = 10, and this parameter setting is also different in different prediction target tasks. The appropriate number of patterns can be obtained through experiments within the above settings. Segment the prediction data into prediction patches y according to the method of segmenting historical data e ;
[0064] In each round of learning, learn different patterns by minimizing the distance between a certain pattern representation z δ in the pattern information and the prediction sequence feature representation y e ;
[0065]
[0066] As the loss function, it will guide the update of pattern information in subsequent model training. Then, through the historical patches X p guide the fusion of prediction patterns to obtain the fused patches
[0067]
[0068] Among them, W x is the scientific department weight matrix, b x and b z are bias terms. Through the historical patches X pThe weight guidance improves the weights of information related to historical information in each mode and reduces the weights of less relevant information; the historical patch X p is concatenated with the fused patch to form a historical prediction feature representation;
[0069] Step 3, feature representation learning, including:
[0070] (1) Temporal feature learning, through the first attention mechanism, multi-head attention calculation is performed on the historical prediction feature representation, establishing connections between different patches. The multi-head mechanism extends the model's ability to focus on different positions, learns temporal features, and the output of the attention module is O i ,
[0071]
[0072] Q i , K i , V i are generated by non-linear projection from the historical prediction feature representation, is the scaling factor, and the obtained output contains all the dependencies between patches within the input, within the fused prediction information mode, and between the input and the fused prediction information mode; at this time, the attention operation is performed independently for each channel, (·) i corresponds to the relevant variable of the i-th dimensional attribute;
[0073] (2) Learning the relationships between variables. First, perform a downsampling operation on the patches to aggregate global information and reduce the number of patches, so as to encapsulate richer long-term features in each patch while reducing the computational complexity of W;
[0074] p i = Downsample(O i ), p i′ = Attention(p i , O i , O i )
[0075] Downsample(·) reduces the N + 1 patches in O i to M patches (M = 1 to N) by using a multi-layer perceptron. In this embodiment, M is set to 2. When N is too large, consider increasing M as appropriate to alleviate the problem of information loss. The specific setting of M needs to refer to the experimental results; by using the downsampled p i as the query (Query) in the attention mechanism, and using the original O i as the key (Key) and value (Value), the long-range modeling ability of the attention mechanism is used to dynamically aggregate global information to obtain the feature representation p i′This enables each patch to encapsulate richer long-term information, making it possible to capture complex cointegration relationships that only emerge over a sufficiently long time range, alleviating the impact of channel independence in temporal feature learning;
[0076] Subsequently, the co-attention mechanism is used to capture the global cointegration relationships between variables in the variable dimension, adaptively evaluating the strength of these relationships: stronger cointegration relationships are reflected by higher attention weights, while weaker connections receive lower weights;
[0077]
[0078] LayerNorm(·) is a common layer normalization operation in Transformer. Simply put, it calculates the mean and variance by statistically analyzing all dimensions for each sample, stabilizing the distribution of each sample; p′ D is p i′ (i = 1, …, M) aggregation in the representation of the variable dimension, is the feature representation of the relationship between variables obtained by the co-attention mechanism and mapped back to the original dimension;
[0079] Step 4, Output of the predicted solar wind speed, including:
[0080] Taking the above-obtained features Taking the predicted position features indicating that the predicted result of the solar wind speed of the model output is obtained by using a flat layer with a linear head,
[0081]
[0082] In the formula, Flatten(·) flattens the feature mapping dimension, and then the flattened feature representation is passed through a linear layer and then through a linear layer Linear(·) for prediction output;
[0083] Step 5, Optimize the model to obtain the final predicted solar wind speed result: Use the trained model to test the data in the test set to obtain the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the observation sequence. It should be noted that the performance of the model is evaluated using the validation set or the test set, and the root mean square error (RMSE), mean absolute error (MAE), and correlation coefficient (CORR) metrics are used to evaluate the accuracy, generalization ability, and robustness of the model. When the metrics do not meet the set expectations, the parameters can be adjusted and retrained.
[0084] The expectation of this invention is the experimental results of the solar wind speed prediction model WSA-Enlil Solar Wind Prediction published by the International Space Environment Service (ISES).
[0085]
[0086] a is the number of samples, y is the true sample data, is the predicted sample data, is the mean of the predicted sample data, is the mean of the true sample data.
[0087] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can make many improvements and variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A solar wind speed prediction method based on a fusion prediction mode and a collaborative attention mechanism, characterized in that The deep learning model adopted includes a prediction pattern learning and fusion module, a feature extraction module, and a flattening layer; The input samples of the deep learning model during training and validation are historical data and prediction data, and the input samples during testing are historical data; The prediction pattern learning and fusion module learns multiple patterns of prediction information from the prediction data in the input of the deep learning model. The historical data guides the fusion of multiple patterns of prediction information to obtain a prediction pattern, and the update of the prediction pattern is realized during the backpropagation process in the training process of the deep learning model; The feature extraction module includes a first attention mechanism module, a second attention mechanism module, a downsampling module, and a collaborative attention mechanism module; the structures of the first attention mechanism module and the second attention mechanism module are the same; the first attention mechanism module receives the output of the prediction pattern learning and fusion module and performs temporal feature learning on the output; The downsampling module performs a downsampling operation on the temporal feature representation output by the first attention mechanism module; the second attention mechanism module aggregates the global information of the temporal features obtained by the downsampling operation; The collaborative attention mechanism module learns the relationship between the outputs of the second attention mechanism module in the variable dimension to obtain a variable relationship feature representation; The flattening layer maps the variable relationship feature representation obtained by the feature extraction module to the solar wind speed value.
2. The solar wind speed prediction method according to claim 1, characterized in that It includes the following steps: Step 1, data preparation and processing: Select five physical quantity attributes from the space weather dataset, namely solar wind speed (Bulk speed), proton density (Proton density), proton temperature (Proton temperature), flow pressure (Flow pressure), and the vector length (Sigma-b) of the standard deviation of the average value of each component; Binarize the ICME list as an attribute, set the time points with coronal mass ejection events to 1, and otherwise set to 0; denoted as ICME; Calculate the sum of the pixel values of the coronal hole in the Atmospheric Imaging Assembly (AIA) image as an attribute, which characterizes the influence of the coronal hole, denoted as area; Use the multivariate time series composed of the above 7 attribute data, including data of at least 3 consecutive years, sampled once per hour; use this multivariate time series as the dataset input to the deep learning model. This dataset is divided into a training set, a validation set, and a test set in a ratio of 5:1:1 and in chronological order. The deep learning model makes predictions for 24 hours, 48 hours, 72 hours, and 96 hours respectively; the target outputs of the deep learning model are the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the observation sequence; the learning rate of the deep learning model is 0.0001 - 0.1, and the batch size is 64 - 512; Perform standard deviation normalization (standardScale) processing on the data of each of the above attributes using Equation (1). After processing, the mean of the data is 0 and the standard deviation is 1. In formula (1), x* is the data after processing, x is the data before processing, μ is the mean of all sample data, and σ is the standard deviation of all sample data; Step 2, Prediction mode learning and fusion: The processed data first passes through the prediction mode learning and fusion module to learn and fuse the prediction mode, and the operation is as follows: 2-1) Divide the standardized historical data into overlapping or non-overlapping historical patches X through the patching segmentation method p ; Let the historical data be X = (x1, …, x L ), x1 to x L are L vectors composed of the above 7 attributes. L is the window size of the historical data input, L = 24 to 192, the patch length is P = 8 to 48, and the step size is S = (0.5 to 1)P. The number of the historical patches X p is N, 2-2) Randomly initialize the prediction pattern Z = [z1, …, z n , where n is the number of patterns, n = 1 to 20, and it is used to store the patterns learned from the prediction data subsequently; meanwhile, segment the prediction data into prediction patches y according to the method of segmenting the historical data e ; During the process of model training, by minimizing the distance between a certain pattern representation z δ in the prediction pattern and the predicted patch y e to update and learn different prediction patterns; As a loss function, it subsequently guides the update of the prediction pattern during model training; Using the historical patch X through Equation (3) p to guide the prediction mode fusion to obtain the fused patch Among them, W x is the weight matrix of the science department, and b x and b z are the bias terms; the historical patch X p and the fused patch are concatenated into a historical prediction feature representation as the output of the prediction pattern learning and fusion module; Step 3, Feature representation learning: 3-1) Temporal feature learning, the output of the prediction mode learning and fusion module is calculated by the first attention mechanism module through multi-head attention to obtain the temporal feature representation, and the output of the attention mechanism is: Q in formula (4) i , K i , V i are generated by non - linear projection of historical prediction features, is a scaling factor; the attention mechanism operation at this time is carried out independently for each channel, (·) i corresponding to the relevant variable of the i - th dimensional attribute; 3-2) Learning the relationship between variables: The downsampling module is used to perform downsampling operations on the temporal feature representation, and the second attention mechanism module aggregates the global information of the downsampled temporal feature representation. The specific operation is as follows: p i = Downsample(O i ), p i′ = Attention(p i , O i , O i ) (5) Downsample(·) reduces the N+1 patches in O to M patches, where M = 1 to N, by using a multi-layer perceptron i ; In the second attention mechanism module, the downsampled p i is used as the query (Query), and the original O i is used as the key (Key) and value (Value), and the long-range modeling ability of the attention mechanism is utilized to dynamically aggregate global information to obtain the feature representation p i′ ; The co-attention mechanism module captures the feature p aggregated with global information obtained by the second attention mechanism module in the variable dimension i′ , indicating the global cointegration relationship between variables. The specific operations are shown in equations (6) and (7) as follows: The LayerNorm(·) in Equation (6) and Equation (7) is a common layer normalization operation in Transformer, which calculates the mean and variance by statistically analyzing all dimensional values for each sample; p′ D is p i′ , where i = 1, …, M; the representation in the variable dimension is the feature representation of the relationship between variables obtained through the collaborative attention mechanism and mapped back to the original dimension; Step 4, Output of the predicted value of the solar wind speed: The feature representation of the relationship between variables obtained by the collaborative attention mechanism module Extract the features of the corresponding predicted position patch Indicates that the predicted result of the solar wind speed of the model output is obtained by using a flat layer with a linear head In formula (8), Flatten(·) flattens the features and then predicts and outputs the flattened feature representation through the linear layer Linear(·); Step 5, Optimize the model to obtain the final predicted result of the solar wind speed: Use the trained deep learning model to test the data in the test set, and obtain the predicted values of the solar wind speed 24 hours, 48 hours, 72 hours, and 96 hours after the distance observation sequence respectively; if the predicted value of the solar wind speed meets the expectation, the predicted result of the solar wind speed output by the model obtained in step 4 is the final predicted result of the solar wind speed, otherwise, return to step 2 to adjust the model parameters and retrain.
3. The solar wind speed prediction method according to claim 2, wherein In step 5, the adjusted model parameters include: the learning rate and batch size of the deep learning model; the window size L, patch length P, and step size S of the historical data input in step 2-1); the number of modes n in step 2-1), and the number of patches M reduced from N+1 patches in step 3-2).
Citation Information
Cited By
System and method for predicting gale along high-speed rail
CN120744393A
High-speed rail along the line of the wind prediction method and prediction system
CN120744393B
Short-term photovoltaic power prediction device and method based on multi-mode collaborative attention
CN120930888A
Prediction method, system and equipment of coronal mass ejection arrival time and medium
CN122089713A
Methods, systems, equipment, and media for predicting the arrival time of coronal mass ejections.
CN122089713B