Solar flare prediction method based on multiple deep learning algorithms
By combining deep learning algorithms such as 1D-CNN, LSTM and TCN, the problem of insufficient accuracy of traditional flare prediction methods was solved, more accurate flare predictions were achieved, and the capability and safety of space weather forecasting were improved.
Patent Information
- Application Number
- CN202510739750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Traditional flare prediction methods have problems with insufficient prediction accuracy and timeliness in capturing the complex nonlinear evolution of the solar magnetic field, making it difficult to achieve accurate flare predictions.
A variety of deep learning algorithms, including 1D-CNN, LSTM and TCN, are used, combined with feature selection and data balancing processing, to achieve flare prediction through multi-model fusion.
It has significantly improved the accuracy and reliability of flare predictions, reduced the false alarm rate, and provided important technical support for space weather forecasting, spacecraft protection, and Earth space environment monitoring.
Smart Images

Figure CN120632452A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of astronomy, space science and machine learning, and relates to a solar flare prediction method, specifically to a flare prediction method based on multiple deep learning algorithms. Background Art
[0002] Solar flares are violent eruptions caused by violent magnetic reconnection in active solar regions. They release large amounts of high-energy particles and radiation, severely impacting Earth's space environment, power systems, navigation and communications, and spacecraft operations. In particular, large X-class flares and their accompanying coronal mass ejections (CMEs) can cause geomagnetic storms, extreme ionospheric disturbances, and even threaten the safety of space missions. Therefore, accurately predicting the timing, intensity, and spatial distribution of solar flares is a core task of space weather forecasting and is crucial for ensuring the normal operation of our high-tech society.
[0003] Traditional flare prediction methods primarily rely on statistical models and numerical simulations based on physical mechanisms, such as empirical prediction methods based on sunspot classifications (such as the McIntosh and Mount Wilson classifications) or solar magnetic field parameters. However, these methods have significant limitations in capturing the complex, nonlinear evolution of the solar magnetic field, and both prediction accuracy and timeliness need to be improved. In recent years, with the accumulation of solar observation data and breakthroughs in deep learning computer technology, flare prediction methods based on deep learning have become a research hotspot. Deep learning methods can automatically learn high-dimensional feature relationships from large amounts of historical data and are particularly well-suited for processing the complex spatiotemporal correlations in solar activity data. With the rich magnetic field and flare data provided by high-resolution observational instruments such as NASA's SDO / HMI and GOES satellites, flare prediction technology based on deep learning is gradually developing towards real-time and high-precision capabilities. Summary of the Invention
[0004] The present invention provides an efficient and accurate flare prediction method based on multiple deep learning algorithms. This method uses multiple deep learning technologies to express sequence data characteristics, and achieves more accurate prediction of flares through comprehensive processing, thereby improving the efficiency and accuracy of flare prediction and providing important guarantees for space exploration and activities.
[0005] The purpose of the present invention is achieved through the following technical solutions: A flare prediction method based on multiple deep learning algorithms includes the following steps: Step S1, obtaining the flare magnetic field data file; Step S2: perform data cleaning, process missing values and outliers in the data, form a data set and divide it into a training set, a validation set and a test set; Step S3: normalize the data set using the Min-Max method, and proceed to step S4; Step S4: Use the XGBoost algorithm to calculate feature importance and select important features; Step S5: Determine whether the processed training set belongs to a balanced data set. If so, proceed to step S8; if not, proceed to step S6; Step S6: Generate positive samples using the SMOTE algorithm; Step S7: using an undersampling algorithm to reduce negative samples; Step S8: Input the obtained balanced training set into the independently trained 1DCNN, TCN, and LSTM models for prediction, and obtain three prediction results; Step S9: Use the weighted average method to fuse the prediction results obtained by the three models to obtain the final prediction result.
[0006] Compared with the prior art, the present invention has the following advantages: The present invention provides a new flare prediction method, which uses three deep learning algorithms: 1D-CNN, LSTM, and TCN to perform efficient feature extraction and time series modeling of the magnetic field parameters of the solar active region. 1D-CNN is used to extract local spatial features, LSTM captures the long-term dependencies of magnetic field evolution, and TCN combines causal convolution and dilated convolution to enhance time series prediction capabilities. Through multi-model fusion and weighted integration strategies, the present invention can more accurately predict the occurrence time and intensity of flares, improving the reliability and lead time of forecasts. The present invention can not only significantly enhance the space weather forecast capability, but also reduce the false alarm rate, providing important technical support for spacecraft protection, power communication security, and Earth space environment monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 This is a flow chart of the flare prediction method based on multiple deep learning algorithms; Figure 2 The AUC comparison results of different models are shown in Figure 2. DETAILED DESCRIPTION
[0008] The technical solution of the present invention is further described below with reference to the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention that does not depart from the spirit and scope of the technical solution of the present invention should be included in the scope of protection of the present invention.
[0009] The present invention provides a flare prediction method based on multiple deep learning algorithms. First, data cleaning is performed, missing values and outliers are processed, and the data set is divided. Then, Min-Max normalization is used to enhance the numerical stability of the model. In view of the large number of magnetic field features, XGBoost is used for feature selection. In order to solve the problem of class imbalance in the data set, minority class samples are first generated by SMOTE, and then the majority class samples are reduced by undersampling to balance the ratio of positive and negative samples. Then, the data are input into the 1DCNN, TCN and LSTM models respectively, and the prediction results of the three are fused using the weighted average method to obtain the final prediction result. Figure 1 The specific steps are as follows: Step S1, obtaining the flare magnetic field data file.
[0010] In this step, 25 officially published magnetic field characteristics and flare eruption records are obtained. The specific magnetic field characteristics are shown in Table 1: Table 1
[0011] When the magnetic field characteristic data of the active region is obtained, it is converted into a CSV format table. It is then searched in the flare records according to the recording time and active region number. If there is no flare record, the data is a negative sample; if a corresponding flare record is found, the data is a positive sample.
[0012] Step S2: perform data cleaning, process missing values and outliers in the data, form a data set and divide it into training set, validation set and test set.
[0013] In this step, the specific method of data cleaning is: a. Missing value processing: Check the integrity of each data record. If the missing value ratio of a single data record is ≥30% or the core parameters (parameters are defined as: TOTUSJH, USFLUX, R_VALUE, MEANPOT, SHRGT45, MEANJZH and TOTBSQ) are missing, mark the record as invalid data. If the missing value ratio is less than 30% and non-core parameters are missing, fill the missing values of non-core parameters (mean imputation) and do not mark the record as invalid data; b. Outlier detection: The complete data records are checked using the improved robust Z-score method, and out-of-range values (defined as modified_z>3.5) are directly marked as erroneous data; c. Data deletion rules: Delete records that meet any of the following conditions: marked as invalid data, marked as erroneous data.
[0014] In this step, a non-overlapping partitioning strategy is used to divide the dataset into training set, validation set, and test set according to the time series. The time window is 24 consecutive hours, the sliding step is 1 hour, the training set contains samples from the first 70% of the time period, the test set is the last 5%, and the validation set is the remaining 15%. A 24-hour isolation zone is set between the three to prevent information leakage.
[0015] Step S3: Use the Min-Max method to normalize the data set, and then proceed to step S4.
[0016] In this step, the Min-Max method is used to normalize different features with very large spans. Note that the normalization parameters ( , ) is calculated only from the training set and applied simultaneously to the validation set and the test set. The specific formula is as follows: (1) in, and Represent the values before and after scaling, and Represents the minimum and maximum values respectively.
[0017] Step S4: Use the XGBoost algorithm to calculate feature importance and select important features.
[0018] In this step, only the training set data is used to perform the following operations: a. Train the XGBoost model (parameters: max_depth=6, learning_rate=0.3, n_estimators=100); b. Calculate the sum of gains of each feature on the training set according to formula (2) : (2) in, Characterized by The total gain generated across all decision tree node splits, is the characteristic variable whose importance is to be evaluated, For all decision trees passing features The set of instances that undergo node splitting, is the index of the split node in the decision tree, For the node The loss function value before splitting, For the node The sum of the weighted losses of its child nodes after the split.
[0019] c. Calculate feature importance according to formula (3) : (3) in, For the data set features (25 magnetic field parameters in total).
[0020] d. Select Features with scores > 0.05 form the final feature set F_selected.
[0021] e. Apply the feature set F_selected to the training set, validation set, and test set simultaneously to ensure that the data dimensions of the input model are consistent.
[0022] Step S5: Determine whether the processed training set belongs to a balanced data set. If so, proceed to step S8; if not, proceed to step S6.
[0023] In this step, it is necessary to determine whether the training set is balanced. If it is unbalanced, how to perform the corresponding balancing process. The specific method is as follows: Calculate the class imbalance index of the training set, the formula is: (4) in, and represent the number of positive samples and negative samples respectively.
[0024] When IR≥5, proceed to step S6 to perform balancing processing (note: this step is only applied to the training set, and the validation set and test set maintain the original data distribution): When IR<5, the training set is a balanced data set, and step S8 is directly executed.
[0025] Step S6: Use the SMOTE algorithm to generate positive samples.
[0026] In this step, the SMOTE algorithm is used to generate synthetic samples, and the positive samples are amplified to 80% of the negative samples (number of neighbors k = 5), and then step S7 is entered.
[0027] Step S7: Use an undersampling algorithm to reduce negative samples.
[0028] In this step, K-Means undersampling is performed on the negative samples (the number of clusters is the number of positive samples × 2), and then step S8 is entered.
[0029] Step S8: Input the obtained training set into the independently trained 1DCNN, TCN, and LSTM models for prediction, and obtain three prediction results.
[0030] In this step, the processed training set is fed into the 1DCNN, TCN, and LSTM models, and the three models are trained independently. The specific implementation is as follows: (1) Input the training set into each model according to the following specifications: 1DCNN: Input is a 3D tensor of shape (training set size, 24, 15).
[0031] TCN: Input shape is (training set size, 15, 24), then transposed to (training set size, 24, 15) to accommodate temporal convolution.
[0032] LSTM: Input is a 3D time series tensor with shape (training_size, 24, 15).
[0033] (2) The configuration for independent training of each model is as follows: 1DCNN model: a. Structure: First convolutional layer: input channels 15, output channels 32, convolution kernel length 3, stride 1, padding 1.
[0034] ReLU activation function.
[0035] Max pooling layer (pooling window length 2).
[0036] Second convolutional layer: input channels 32, output channels 64, convolution kernel length 3, stride 1, padding 1.
[0037] The flatten layer outputs a 1024-dimensional feature vector (64 channels × 16 time steps).
[0038] The fully connected layer maps to a 1-dimensional output.
[0039] b. Training parameters: Optimizer: Adam (β1=0.9, β2=0.999).
[0040] Initial learning rate: 0.001 (cosine annealing adjustment).
[0041] Batch size: 128.
[0042] Early stopping method: terminate when the validation set loss does not decrease for 15 consecutive rounds.
[0043] TCN model: a. Structure: 4 layers of dilated causal convolution, with dilation coefficients of 1, 2, 4, 4 respectively.
[0044] The number of convolution kernels in each layer is 64, and the kernel length is 3.
[0045] The global max pooling layer compresses the time dimension.
[0046] The fully connected layer outputs a 1D prediction result.
[0047] b. Training parameters: Optimizer: NAdam.
[0048] Initial learning rate: 0.002.
[0049] Weight decay: 1e-5.
[0050] Gradient clipping threshold: 1.0.
[0051] LSTM model: a. Structure: Unidirectional LSTM layer: 2 layers stacked, 32 hidden units per layer.
[0052] Fully connected layer: 32-dimensional input is mapped to 1 dimension.
[0053] b. Training parameters: Dropout rate: 0.3.
[0054] Initial learning rate: 0.001 (step decay, ×0.1 every 50 epochs).
[0055] Sequence length: 24 steps.
[0056] (3) Each model satisfies the following independence conditions: a. Parameter independence: Model weight initializations do not affect each other (1DCNN uses He initialization, TCN uses Xavier initialization, and LSTM uses orthogonal initialization).
[0057] b. Gradient isolation: Cross-model gradient sharing or parameter fusion is prohibited during training.
[0058] Step S9: Use the weighted average method to fuse the prediction results obtained by the three models to obtain the final prediction result.
[0059] In this step, the AUC (Area Under the Receiver Operating Characteristic Curve, ROC curve area) of each model is first calculated on the validation set. 、 、 Calculated by AUC normalization:
[0060] (5)
[0061] in, 、 and Represents a single model , AUC values obtained by TCN and LSTM models.
[0062] Then the final integrated prediction result is calculated according to the weighted average formula (6): (6) in, For the final prediction result, 、 and Single model , TCN and LSTM obtained prediction probability.
[0063] Finally, the test set is used for the final performance test. The specific experimental results are shown in Table 2: Table 2 Prediction effect provided by the present invention
[0064] By comparing with the prediction methods proposed by others, we can obtain the results in Table 3.
[0065] Table 3 Comparison of prediction effects between the present invention and existing algorithms
[0066] The obtained dataset was analyzed for feature importance using the XGBoost method, and the top 15 features were selected for model training. Since flares are low-probability events, there is a high probability of class imbalance in the dataset. When it is determined that the dataset is not balanced, the SMOTE method and undersampling method are used to process the class imbalance of the dataset. Figure 2 The red data in the middle is the data that has been processed with class imbalance, while the blue data represents the unprocessed data set. The data sets are used to train the 1DCNN, TCN, LSTM models and the integrated model of the above three (1DCNN-TCN-LSTM), and the following results can be obtained: Figure 2 The data results are shown in . It can be found that the ensemble model obtained by weighted average method is significantly better than the single model.
Claims
1. A flare prediction method based on multiple deep learning algorithms, characterized by The method comprises the following steps: Step S1, obtaining a flare magnetic field data file; Step S2: perform data cleaning, process missing values and outliers in the data, form a data set and divide it into a training set, a validation set and a test set; Step S3: normalize the data set using the Min-Max method, and proceed to step S4; Step S4: Use the XGBoost algorithm to calculate feature importance and select important features; Step S5: Determine whether the processed training set belongs to a balanced data set. If so, proceed to step S8; if not, proceed to step S6; Step S6: Generate positive samples using the SMOTE algorithm; Step S7: using an undersampling algorithm to reduce negative samples; Step S8: Input the obtained balanced training set into the independently trained 1DCNN, TCN, and LSTM models for prediction, and obtain three prediction results; Step S9: Use the weighted average method to fuse the prediction results obtained by the three models to obtain the final prediction result.
2. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S1, when the magnetic field characteristic data of the active area is obtained, it is converted into a CSV format table and searched in the flare burst record according to the recording time and the active area number. If there is no flare record, the data is a negative sample; if the corresponding flare burst record is found, the data is a positive sample.
3. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S2, the specific steps of data cleaning are as follows: a. Missing value processing: Check the integrity of each data record. If the missing value ratio of a single data record is ≥30% or the core parameter is missing, mark the record as invalid data. If the missing value ratio is less than 30% and the non-core parameter is missing, fill in the missing value of the non-core parameter and do not mark the record as invalid data. b. Outlier detection: The improved robust Z-score method is used to check the complete data records, and out-of-range data are directly marked as erroneous data; c. Data deletion rules: Delete records that meet any of the following conditions: marked as invalid data, marked as erroneous data.
4. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S3, the normalization formula is as follows: in, and Represent the values before and after scaling, and Represents the minimum and maximum values respectively.
5. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that The specific steps of step S4 are as follows: a. Train the XGBoost model; b. Calculate the sum of gains of each feature on the training set : in, Characterized by The total gain generated across all decision tree node splits, is the characteristic variable whose importance is to be evaluated, For all decision trees passing features The set of instances for node splitting, is the index of the split node in the decision tree, For the node The loss function value before splitting, For the node The sum of the weighted losses of its child nodes after splitting; c. Calculate feature importance : in, For the data set Features d. Select Features with F_>0.05 form the final feature set F_selected.
6. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that The specific steps of step S5 are as follows: Calculate the class imbalance index of the training set : in, and Represent the number of positive samples and negative samples respectively; When IR≥5, the process proceeds to step S6 to perform balancing processing; when IR<5, the training set is a balanced data set, and step S8 is directly executed.
7. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S6, the SMOTE algorithm is used to generate synthetic samples, and the positive samples are amplified to 80% of the negative samples, and then step S7 is entered.
8. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S7, K-Means undersampling is performed on the negative samples, and then step S8 is entered.
9. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S8, the training set is input into each model according to the following specifications: 1DCNN: Input is a 3D tensor with shape (training set size, 24, 15); TCN: Input shape is (training set size, 15, 24), then transposed to (training set size, 24, 15) to accommodate temporal convolution; LSTM: Input is a 3D time series tensor with shape (training set size, 24, 15); Each model satisfies the following independence conditions: a. Parameter independence: Model weight initialization does not affect each other. 1DCNN uses He initialization, TCN uses Xavier initialization, and LSTM uses orthogonal initialization. b. Gradient isolation: Cross-model gradient sharing or parameter fusion is prohibited during training.
10. The flare prediction method based on multiple deep learning algorithms according to claim 1, characterized in that In step S9, the weighted average formula is: in, For the final prediction result, 、 and Single model , the predicted probability obtained by TCN and LSTM, 、 、 Represents a single model , the weights of the TCN and LSTM models, 、 and Represents a single model , AUC values obtained by TCN and LSTM models.
Citation Information
Patent Citations
Solar flare dichotomy prediction method based on support vector machine
CN113902053A
Solar flare spectral line inversion method
CN119358245A