Loop current prediction method based on CNN and BiLSTM neural network

Through the dual-current input architecture of CNN and BiLSTM neural networks, the limitations of the ring current proton flux prediction model in complex nonlinear dynamic changes are solved, and efficient and accurate ring current proton flux prediction is achieved, which improves the generalization ability and real-time early warning ability of the model.

CN120492794APending Publication Date: 2025-08-15JIANGSU OCEAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510412766.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing ring current proton flux prediction model has limitations when dealing with complex nonlinear dynamic changes, and it is difficult to capture the real-time dynamic changes of ring current, and the calculation complexity is high, which cannot meet the real-time early warning requirements.

Method used

Using a dual-current input architecture based on CNN and BiLSTM neural networks, the spatial distribution characteristics of satellite data are extracted using convolutional neural networks, and the timing characteristics of geomagnetic index are captured through bidirectional long and short-term memory networks, a deep fully connected network is constructed for cross-modal fusion, achieving efficient and accurate prediction of ring current proton flux.

Benefits of technology

It significantly improves prediction accuracy, reduces errors, improves the generalization ability and robustness of the model, and meets the timeliness requirements of real-time monitoring and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492794A_ABST
    Figure CN120492794A_ABST
Patent Text Reader

Abstract

The invention provides a loop current prediction method based on a CNN and a BiLSTM neural network. The loop current prediction method is used for improving the prediction precision of proton flux data in space weather. The method comprises the following implementation steps of: acquiring and preprocessing data; a CNN neural network model and a BiLSTM neural network model are constructed; training the model; and evaluating and visualizing the model. According to the method, the advantages of the CNN in the aspect of spatial feature extraction and the powerful time sequence processing capability of the BiLSTM neural network are utilized, and finally, the outputs of the CNN and the BiLSTM neural network are fused, so that the spatial-temporal dynamic distribution of the ring current proton flux can be more effectively captured. Experimental results show that the method has excellent performance in a loop current prediction task, is suitable for real-time detection and early warning of space weather, and has relatively high prediction precision and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of space weather prediction, and in particular relates to a ring current prediction method based on CNN and BiLSTM neural networks. Background Art

[0002] In the field of space weather prediction technology, accurate prediction of the proton flux of the circulating current is particularly important. The circulating current refers to the flow of charged particles surrounding the Earth. Its intensity and distribution directly affect changes in the Earth's magnetic field, which in turn affects the normal operation of related facilities such as satellite communications and navigation systems. Traditional circulating current models mainly rely on physical models and statistical models. Although they can provide certain predictive capabilities, they have limitations when dealing with complex nonlinear dynamic changes. For example, Daglis established an empirical distribution model of the circulating current ions by analyzing satellite data, but these models are generally static and cannot accurately capture the real-time dynamic changes of the circulating current. Therefore, finding an efficient and accurate method to predict the proton flux of the circulating current is of great practical significance for the detection and early warning of space weather.

[0003] In recent years, with the advancement of machine learning technology, especially the powerful capabilities of deep learning in predicting time series data, artificial neural networks (ANNs) have gradually been applied to space weather prediction. For example, Bortnik successfully constructed an ANN-based plasma density model using the SymH index. Chu and Zhelavskaya then reconstructed a dynamic plasma density model using multiple input streams (including solar wind parameters and geomagnetic index time series), successfully capturing the ionospheric erosion and replenishment process. These models have demonstrated the potential of machine learning in space weather prediction. However, these models focus on the prediction of plasma density and electron flux, with less attention to the prediction of the ring current proton flux. Moreover, the dynamic changes of the ring current proton flux are closely related to the injection, acceleration, and loss of particles during magnetic storms. In particular, during the recovery period of a magnetic storm, the decay of the ring current proton flux is affected by multiple physical processes such as charge exchange, Coulomb collisions, and wave-particle interactions. Therefore, accurately predicting the spatiotemporal distribution of the ring current proton flux is crucial for capturing the dynamic changes of the ring current during magnetic storms.

[0004] Currently, domestic methods for predicting ring currents include the LSTM ring current prediction model and the RAM-SCB regionalized improved model. The LSTM ring current prediction system, with a unidirectional LSTM at its core, uses single-mode solar wind parameters (such as solar wind speed and interplanetary magnetic field (IMF-Bz)) as input. Finally, dynamic prediction of ring currents is achieved through time series feature extraction. While this method achieves low prediction errors under steady-state conditions, it fails to integrate geomagnetic and satellite data, making it difficult to capture magnetosphere-ionosphere coupling effects. Furthermore, the unidirectional LSTM cannot resolve the reverse energy dissipation process during the ring current decay phase, resulting in significant deviations in peak predictions during magnetic storms. It also suffers from slow convergence. The RAM-SCB regionalized improved model relies on numerical simulations of physical parameters such as the magnetopause reconnection efficiency and particle pitch angle distribution. While this model offers strong physical interpretability, it suffers from high computational complexity. Single simulations rely on supercomputing resources, resulting in long simulation times and insufficient real-time warning requirements. Furthermore, its data dependency is weak and it fails to fully assimilate multi-source observational data, resulting in poor adaptability to the timing of sudden magnetic storms.

[0005] To address the shortcomings of existing models in prediction, this paper will make full use of the powerful feature extraction capabilities of CNN and the ability of LSTM neural networks to effectively capture long-term dependencies in time series, and create a neural network based on CNN and BiLSTM to capture the dynamic changes of ring current more accurately and quickly. Summary of the Invention

[0006] Based on the above technical background and the deficiencies of the existing technology, this patent proposes a ring current prediction method based on CNN and BiLSTM neural networks. The method adopts a dual-stream input architecture of satellite observation data and geomagnetic index data. The convolutional neural network (CNN) is used to effectively extract the spatial distribution characteristics of satellite data, and the bidirectional long short-term memory network (BiLSTM) is used to capture the bidirectional time series characteristics of the geomagnetic index during the magnetic storm process. Subsequently, a deep fully connected network is constructed for cross-modal fusion optimization, so that the model can efficiently and accurately predict the dynamic changes of the ring current proton flux, and realize real-time monitoring and early warning of the ring current dynamics, thereby solving the problem that traditional methods are difficult to capture the complex nonlinear changes of the ring current; specifically, the following steps are included:

[0007] S1: Data collection and preprocessing, specifically including the following steps:

[0008] Data collection: Ground-based measurement data are used, including geomagnetic indices SymH, AsyH, AsyD, and SME. Satellite-based data are used, including satellite orbit parameter L, MLT, LAT, and ring current proton flux data with an energy range of 45keV to 598keV. SymH represents the symmetric ring current index; AsyH represents the asymmetric ring current index; AsyD represents the other component of the asymmetric ring current; SME represents the SuperMAG version of the auroral energy collection index; L represents the magnetic shell parameter; MLT represents the magnetic local time; and LAT represents the latitude.

[0009] Data preprocessing: Perform preprocessing operations on the collected data, including data cleaning, data standardization, and data segmentation. At this time, the data consists of a training set, a validation set, and a test set;

[0010] S2: CNN and BiLSTM neural network model construction, including the following steps:

[0011] Construct satellite data branch CNN network structure;

[0012] Construct a BiLSTM network structure for geomagnetic data branches;

[0013] Build a deep fully connected fusion network structure;

[0014] S3: Model training and optimization, specifically using the Adam optimizer, with an initial learning rate of 1e-4 and a loss function of mean square error. The loss calculation formula is as follows:

[0015]

[0016] Among them, n is the total number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample; during the training process, the learning rate decays exponentially with the number of training rounds, with a decay rate of 0.95; an early stopping strategy is adopted, and when the validation loss does not continue to decrease for 50 consecutive rounds, training is stopped and restored to the optimal model weight;

[0017] S4: Model evaluation and results analysis, by calculating the coefficient of determination R 2 Evaluate the model's ability to explain the target variable. The specific calculation formula is:

[0018]

[0019] Where y is the true mean.

[0020] As a preferred technical solution of the present invention, the data preprocessing in step S1 specifically includes the following steps:

[0021] S1-1: Data cleaning: Perform non-finite value detection on the collected data; for single-point abnormal data, use the forward filling method, that is, replace the current abnormal value with the normal value at the previous moment of the abnormal value; for continuous time-period abnormal data, use the data of the most recent normal time period before this abnormal period for filling and replacement.

[0022] S1-2: Data normalization: Perform Min-Max normalization on the cleaned geomagnetic data and scale the metrics into the interval [-1, 1]. The specific expression is:

[0023]

[0024] where Index norm [[ID=...]]represents the geomagnetic index after completing the normalization operation;

[0025] For the cleaned satellite orbit data, first directly filter and remove the data points with L < 2.5 using physical constraints (the proton flux with L < 2.5 is significantly affected by the atmosphere and has unclear physical meaning, interfering with the subsequent model training). Perform normalization on the filtered L-value data (2.5 < L < 6.57) to unify L-values of different magnitudes to the same scale. The specific expression is:

[0026]

[0027] where L norm is the L-value after the normalization operation;

[0028] For the periodic encoded MLT, convert MLT into a sine and cosine pair to solve the problem of numerical discontinuity caused by its periodic characteristics. The specific expression is:

[0029]

[0030] For the latitude LAT, divide LAT by the maximum value of 0.35 radians to make it consistent with other feature scales and limit its value to the interval [-1, 1]. The specific expression is:

[0031] [[ID=3 June]]

[0032] where LAT norm [[ID=...]]is the scaled latitude value.

[0033] S1-3: Data splitting: Divide the standardized data into training set, validation set and test set according to the time range; after removing the data in the test set, the remaining data is further divided in an 8:2 ratio, where 80% of the data is used as the training set for model parameter training and 20% of the data is used as the validation set.

[0034] S1-4: Data reshaping: reshape the processed data into a three-dimensional tensor form;

[0035] S1-5: Data re-segmentation: Satellite data and geomagnetic data are extracted from the reshaped data to form two independent input branches, which are respectively input into the CNN and BiLSTM neural network models for subsequent feature extraction and fusion processing.

[0036] As a technical preferred solution of the present invention, the satellite data branch network model in step S2 includes, from the input end to the output end, a first convolutional layer, a first batch of normalization layers, a first Dropout layer, a second convolutional layer, a second batch of normalization layers, a flattening layer, a first fully connected layer, and a second Dropout layer; the first convolutional layer uses 64 filters, a convolution kernel size of 3, a convolution step of 1, a padding method of same, and an activation function of ReLU; the first batch of normalization layers is used to perform batch normalization on the convolution output to accelerate model convergence and improve training stability; the first Dropout layer ... The ut layer randomly discards 20% of neurons to prevent overfitting; the second convolutional layer uses 128 filters, a convolution kernel size of 3, a convolution step of 1, a padding method of the same, and an activation function of ReLU, which is used to further extract higher-level spatial features; the second batch normalization layer further performs batch normalization on the convolution output features; the flattening layer is used to convert the multi-dimensional features output by the second batch normalization layer into one-dimensional features; the fully connected layer contains 64 neurons, and the activation function is ReLU, which is used to extract high-dimensional abstract features of the flattening layer features; the Dropout layer randomly discards 30% of neurons.

[0037] As a technical preferred solution of the present invention, the geomagnetic data branch BiLSTM network structure in step S2 adopts a bidirectional long short-term memory network to capture the bidirectional long-term dependency of geomagnetic data in the time dimension, and its network structure includes a first BiLSTM layer, a first batch of normalization layers, a second BiLSTM layer, a fully connected layer and a Dropout layer from the input end to the output end; the first layer of bidirectional BiLSTM contains 128 neurons, which is used to simultaneously capture the feature information of the positive and negative directions of the time series; the first batch of normalization layers is used to batch normalize the output features of the BiLSTM layer, accelerate model convergence and improve training stability; the second layer of BiLSTM uses 64 neurons to further extract high-level features of geomagnetic data and reduce feature dimensions; the fully connected layer contains 64 neurons, and the activation function is ReLU, which is used to nonlinearly combine the time series features extracted by the BiLSTM layer; the Dropout layer randomly abandons 30% of the neurons.

[0038] As a technical preferred solution of the present invention, the deep fully connected fusion network structure described in step S2 is used to fuse the feature information output by the satellite data branch and the geomagnetic data branch, and further extract and optimize the fused features. The specific network structure includes, from the input end to the output end: a first fully connected layer, comprising 128 neurons, with an activation function of ReLU, for fusing the feature information of the satellite data branch and the geomagnetic data branch; a first normalization layer, for batch normalization of the fused features, accelerating the model training process and improving stability; a first Dropout layer, randomly discarding 40% of the neurons, for suppressing overfitting; a second Dense layer, comprising 64 neurons, with an activation function of ReLU, for further extracting and optimizing the fused abstract features; a second batch normalization layer, further performing batch normalization operations on the features to improve model convergence stability; a second Dropout layer, randomly discarding 20% of the neurons, further reducing the risk of model overfitting; an output layer, comprising one neuron, using a linear activation function, for outputting the final prediction result of the ring current proton flux.

[0039] As a technical preferred solution of the present invention, the result analysis in step S4 specifically includes visualization of the loss curve and visualization of the prediction results, specifically:

[0040] Loss curve visualization: plot the training loss and validation loss against the number of epochs, where the horizontal axis represents the number of training epochs and the vertical axis represents the loss value;

[0041] Visualization of prediction results: A scatter plot is used to compare the actual proton flux value with the proton flux value predicted by the model. The scatter density is presented in a color gradient, from purple to yellow, indicating a low to high density. A standard diagonal line is also drawn in the figure; the mean square error (MSE) and the coefficient of determination (R) are marked in the upper left corner of the figure. 2 and other evaluation indicators.

[0042] Compared with the related prior art, the beneficial effects of the present invention are:

[0043] The prediction accuracy is significantly improved: by constructing a dual-stream network model that integrates CNN and BiLSTM, the spatial characteristics of satellite data and the bidirectional time series characteristics of geomagnetic data are jointly optimized. The error of the ring current proton flux predicted by the model (RMSE≈7.5nT) is reduced by about 30% compared with the traditional LSTM model, and the loss value (Loss≈0) is significantly lower than the existing RAM-SCB regionalized improved model (Loss≈0.3-0.5), making the prediction results more accurate and reliable.

[0044] The model has strong generalization ability: it adopts a multimodal feature fusion architecture of satellite observation data and geomagnetic index data, and effectively captures the global evolution of the magnetosphere through cross-modal feature cascading, so that the model can converge quickly within 10-20 epochs. Compared with the traditional LSTM model, the convergence speed is increased by more than 5 times, and the training and verification loss curves are highly consistent, indicating that the model's generalization ability and stability have been significantly enhanced.

[0045] High model robustness: The Huber loss function is used to optimize the training process, making it more robust to abnormal data and more evenly distributed in prediction errors. This effectively reduces the impact of extreme situations or abnormal data on model performance, allowing the model to maintain high prediction accuracy even in complex magnetic storm environments.

[0046] Strong real-time prediction capability: The CNN-BiLSTM network is combined with GPU acceleration technology to achieve efficient model inference, with a single prediction time of less than 10ms. This is a significant improvement in performance compared to the RAM-SCB regionalized improved model (single simulation exceeds 10 minutes), and can meet the timeliness requirements of real-time monitoring and early warning of space weather, with good practical application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of a ring current prediction method based on CNN and BiLSTM neural network of the present invention;

[0048] Figure 2 2 is a graph showing changes in training and validation loss according to an embodiment of the present invention;

[0049] Figure 3 This is a comparison chart of the model prediction results and the true values provided by the embodiment of the present invention;

[0050] Figure 4 The geomagnetic index time series diagram and ion flux heat map of the embodiment provided by the present invention are shown;

[0051] Figure 5 This is a polar coordinate diagram of the spatial distribution of proton flux in an embodiment provided by the present invention. DETAILED DESCRIPTION

[0052] The present invention is further described below with reference to the accompanying drawings and examples. However, the present invention can be implemented in many different ways and should not be construed as limited to the illustrated embodiments; rather, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.

[0053] Example 1: According to Figure 1 As shown, the present invention provides a ring current proton flux prediction method based on the fusion of CNN and BiLSTM neural network. The specific implementation steps are as follows:

[0054] S1: Data collection and preprocessing, which specifically include the following steps: Data collection: Use ground measurement data, which includes geomagnetic indices SymH, AsyH, AsyD, and SME; Use satellite data collection, which includes satellite orbit parameters L value, MLT, LAT, and ring current proton flux data with an energy range of 45 kev to 598 kev. Among them, SymH represents the symmetric ring current index; AsyH represents the asymmetric ring current index; AsyD represents another component of the asymmetric ring current; SME represents the SuperMAG version of the auroral electrojet index; L value represents the magnetic shell parameter; MLT represents magnetic local time; LAT represents latitude. After the data collection is completed, the original data is systematically preprocessed according to the following method:

[0055] First, use the np.isfinite function to screen out non-finite outliers in the original data, including NaN (not a number) and infinite Inf values. For the detected abnormal data, use the forward filling method, that is, replace the abnormal value with the normal value at the previous moment or time period, so as to ensure the continuity and effectiveness of the data set.

[0056] After the data cleaning is completed, perform Min-Max normalization on the cleaned geomagnetic data and scale the indicators to the [-1, 1] interval. The specific expression is:

[0057] [[ID=IO]]

[0058] where Index norm represents the geomagnetic index after the normalization operation;

[0059] For the cleaned satellite orbit data, first directly use physical constraints to filter out data points with L < 2.5 (the proton flux with L < 2.5 is significantly affected by the atmosphere and its physical meaning is not clear, interfering with the training of subsequent models). Perform normalization on the filtered L value (2.5 < L < 6.57) data to unify L values of different magnitudes to the same scale. The specific expression is:

[0060]

[0061] where, L norm is the L value after the normalization operation;[[ID=2i]]

[0062] For the periodic encoding of MLT, convert MLT into a sine and cosine pair to solve the problem of numerical discontinuity caused by its periodic characteristics. The specific expression is:

[0063]

[0064] For latitude LAT, divide LAT by the maximum value of 0.35 radians to make it consistent with other characteristic scales and limit its value to the interval [-1,1]. The specific expression is:

[0065]

[0066] Among them, LAT norm is the scaled latitude value.

[0067] The data was then segmented by date. Data from January 1, 2017, to January 1, 2018, was set as a separate test set for the final model performance evaluation. The remaining data was randomly divided into training and validation sets. The training set accounted for 80% of the data and was used for network training. The validation set accounted for 20% of the data and was used for early stopping and hyperparameter adjustment and optimization. After data segmentation, the data was reshaped into a three-dimensional tensor (number of samples, 1, number of features) to meet the input requirements of the CNN and BiLSTM models.

[0068] Finally, to achieve multimodal feature fusion, the preprocessed data will extract satellite data and geomagnetic data respectively, forming two independent input branches to facilitate the next step of dual-stream network model construction.

[0069] S2: CNN and BiLSTM neural network model construction adopts a dual-stream input architecture of satellite data and geomagnetic data. Each input branch has an independent network structure:

[0070] Satellite Data Branch (CNN Network Structure): The satellite data branch network structure consists of two convolutional layers. The first convolutional layer uses 64 filters with a kernel size of 3, a stride of 1, a ReLU activation function, and the same padding to capture the spatial distribution characteristics of satellite orbit data. This layer is followed by a batch normalization layer to normalize the data to accelerate training. A dropout layer is then added to randomly drop 20% of neurons to reduce the risk of overfitting. The second convolutional layer uses 128 filters with a kernel size of 3, a stride of 1, the same padding, and the ReLU activation function to further extract high-level spatial features. A subsequent batch normalization layer is also added to improve training stability and efficiency. After the convolution operation is completed, the output feature tensor is transformed into a one-dimensional feature vector through the flattening layer (Flatten), and then input into the fully connected layer (Dense). This layer has 64 neurons and the activation function is ReLU. The spatial abstract features are obtained in a nonlinear combination manner. Another layer of Dropout is added to randomly abandon 30% of the neurons to prevent overfitting.

[0071] Geomagnetic data branch (BiLSTM network structure): The geomagnetic data branch uses a bidirectional long short-term memory (BiLSTM) structure to capture the bidirectional long-term dependencies of geomagnetic data: the first layer of BiLSTM contains 128 neurons, and its bidirectional structure effectively captures the long-term dependencies in the forward and reverse directions during the evolution of the magnetic storm, and is then connected to a batch normalization layer for normalization; the second layer of BiLSTM is set to 64 neurons to further extract temporal features and reduce dimensionality; the BiLSTM output features pass through a fully connected layer (64 neurons, ReLU activation function) for high-dimensional abstraction, and finally connected to a Dropout layer (discarding 30% of neurons) to improve model robustness.

[0072] Feature fusion deep fully connected network structure: The output features of the satellite data branch and the geomagnetic data branch are fused in a cascade manner to construct a deep fully connected network structure. Specifically, the fusion network consists of two fully connected layers: the first fully connected layer contains 128 neurons and uses the ReLU activation function to perform nonlinear combination of the fused features. It is then connected to a batch normalization layer and a dropout layer (40% of the neurons are randomly discarded). The second fully connected layer has 64 neurons and further optimizes the fused abstract features. It is also connected to a batch normalization layer and a dropout layer (20% of the neurons are randomly discarded). Finally, a single linear neuron outputs the final proton flux prediction value.

[0073] S3: Model training and optimization, specifically using the Adam optimizer, with the initial learning rate set to 1e-4, and the mean square error selected as the loss function. The CNN-BiLSTM fusion network in the present invention is trained using the Adam optimizer, with an initial learning rate of 1e-4. During the training process, after each epoch, the learning rate decays exponentially by 0.95 times to ensure rapid convergence of the model in the early stage and fine tuning in the later stage. The model training loss function uses the mean square error (MSE). By real-time monitoring of the loss changes of the validation set, when there is no significant decrease in the validation loss for 50 consecutive epochs, the early stopping strategy is automatically triggered to stop training and restore the optimal model weights to ensure the stability and generalization of the model performance. The loss calculation formula is as follows:

[0074]

[0075] Among them, n is the total number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample; during the training process, the learning rate decays exponentially with the number of training rounds, with a decay rate of 0.95; an early stopping strategy is adopted, and when the validation loss does not continue to decrease for 50 consecutive rounds, training is stopped and restored to the optimal model weight;

[0076] S4: Model evaluation and results analysis, by calculating the coefficient of determination R 2 Evaluate the model's ability to explain the target variable. The specific calculation formula is:

[0077]

[0078] in, is the true mean value.

[0079] Example 2: In this example, the CNN-BiLSTM fusion neural network model trained in Example 1 is used, and the prediction performance of the model is comprehensively analyzed through corresponding evaluation methods and visualization tools.

[0080] The specific evaluation results of this embodiment are as follows Figure 2 shown. Figure 2 This is a graph showing the change in the loss value of the model on the training set and the validation set as the number of Epoch training rounds changes. The horizontal axis represents the number of training rounds (Epoch), and the vertical axis is the loss value MSE. The blue curve represents the loss change trend of the training set, and the orange curve represents the loss change trend of the validation set. Figure 2 It can be clearly observed that as the number of training rounds increases, both training loss and validation loss decrease rapidly, with their trends being highly consistent and nearly overlapping. This demonstrates that the model exhibits excellent learning efficiency and excellent generalization performance during training. Convergence (loss approaching 0) is achieved quickly within the first 10-20 epochs of training, far exceeding the 50-100 epochs required by traditional unimodal LSTM models, demonstrating the significant advantages of the multimodal data fusion architecture of this invention.

[0081] Furthermore, in order to visually demonstrate the accuracy of the model in predicting the proton flux of the ring current, the Figure 3 The scatter plot shown in the figure shows that the horizontal axis represents the actual value of the observed ring current proton flux, and the vertical axis represents the proton flux value predicted by the model. Each scatter point represents the true value and the model predicted value of a data sample in the test set. In order to facilitate the judgment of the prediction accuracy, a red diagonal line is drawn in the figure, which is the standard line (y=x) when the predicted value is completely consistent with the true value. The color of the scatter points in the figure gradually changes from purple to yellow, indicating the trend of the scatter point density from low to high. Figure 3 It can be clearly seen that a large number of scattered points are densely distributed near the standard diagonal line, with small deviations, and the true value is highly consistent with the predicted value. At the same time, the important indicators for model performance evaluation are prominently marked in the upper left corner of the graph: mean square error (MSE: 0.070) and coefficient of determination (R 2 :0.891), these two indicators reflect the accuracy of the model and its ability to effectively explain the changes in the proton flux of the ring current, which are significantly better than the existing technology.

[0082] In order to understand the model's ability to capture the actual loop current dynamic process in more detail, Figure 4 The dynamic process of geomagnetic indices (such as SME, SymH, AsyH, etc.) changing with time and the corresponding evolution of ion fluxes are shown in a specific time period (January 1 to July 1, 2017). Figure 4 The upper middle section uses black curves to illustrate the temporal trends of different geomagnetic indices, while the lower section features a heat map visualization of the ion flux. The horizontal axis of the heat map represents the observation date, while the vertical axis represents the proton flux levels of different energy channels. The color gradient from dark to light corresponds to the trend from low to high flux intensity, clearly demonstrating the dynamic changes in proton flux over time. This time series heat map provides more detailed dynamic information for analyzing model predictions, making it easy to understand the changing trends of ring current characteristics during magnetic storms.

[0083] In addition, this embodiment further demonstrates and analyzes the prediction results of the model for the spatial distribution of proton flux at a specific time (e.g., 16:42 on March 1, 2017). Figure 5 As shown in the figure, the spatial distribution polar coordinate diagram intuitively shows the spatial distribution characteristics of proton flux at different magnetic locations. The X-axis and Y-axis represent the spatial coordinate position, and the proton flux value is represented by a color gradient. The color gradually transitions from black in the center to orange and yellow, reflecting the distribution trend of proton flux from high to low. It intuitively shows the details of the spatial distribution of proton flux and provides an effective tool for the refined analysis of space weather events.

[0084] Through the above detailed model performance evaluation and visualization analysis, it can be seen that the CNN-BiLSTM fusion neural network model of the present invention has significant advantages in predicting the ring current proton flux, which is reflected in the improvement of prediction accuracy, the efficiency of multimodal data fusion, the enhancement of model generalization, and the great improvement of real-time prediction performance. Compared with the traditional single-modal LSTM model or the RAM-SCB regionalized model, the present invention shows obvious advantages in prediction error, convergence speed, and computational time. It can provide a strong supporting technology for practical application scenarios such as dynamic monitoring of space weather and magnetic storm warning in the field of space weather, reflecting the important significance and application value of the present invention in practical engineering applications.

[0085] In summary, this embodiment clearly demonstrates the technical effects and technical implementation details of the model of the present invention, verifies the advancement, effectiveness and reliability of the technical solution of the present invention, and lays a good technical foundation for further promotion and application.

[0086] The above embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.

Claims

1. A ring current prediction method based on CNN and BiLSTM neural network, characterized by: The method includes the following steps: S1: Data collection and preprocessing, which specifically includes the following steps: Data collection: Use ground measurement data, which includes geomagnetic indices SymH, AsyH, AsyD, and SME; use satellite to collect data, which includes satellite orbit parameters L value, MLT, LAT, and ring current proton flux data with an energy range of 45 kev to 598 kev. Among them, SymH represents the symmetric ring current index; AsyH represents the asymmetric ring current index; AsyD represents another component of the asymmetric ring current; SME represents the SuperMAG version of the auroral electrojet index; L value represents the magnetic shell parameter; MLT represents the magnetic local time; LAT represents the latitude. Data preprocessing: Perform preprocessing operations on the collected data, including data cleaning, data normalization, and data segmentation. At this time, the data consists of a training set, a validation set, and a test set. S2: Construction of CNN and BiLSTM neural network models, which specifically includes the following steps: Construct a satellite data branch CNN network structure; Construct a geomagnetic data branch BiLSTM network structure; Construct a deep fully connected fusion network structure; S3: Training and optimization of the model. Specifically, use the Adam optimizer, set the initial learning rate to 1e-4, and select the mean squared error as the loss function. The loss calculation formula is as follows: Among them, n is the total number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample; during the training process, the learning rate decays exponentially with the number of training rounds, and the decay rate is 0.95; an early stopping strategy is adopted. When the validation loss does not continue to decrease for 50 consecutive rounds, the training is stopped and the optimal model weight is restored; S4: Model evaluation and results analysis, by calculating the coefficient of determination R 2 Evaluate the model's ability to explain the target variable. The specific calculation formula is: in, is the true mean value.

2. The method for predicting circulating current based on CNN and BiLSTM neural network according to claim 1, characterized in that: The data preprocessing in step S1 specifically includes the following steps: S1-1: Data cleaning: Perform non-finite value detection on the collected data; for single-point abnormal data, use the forward filling method, that is, replace the current abnormal value with the normal value at the previous moment of the abnormal value; for continuous time-period abnormal data, use the data of the nearest normal time period before this abnormal period for filling and replacement. S1-2: Data normalization: Perform Min-Max normalization on the cleaned geomagnetic data and scale the index to the interval [-1, 1]. The specific expression is: Among them, Index norm Indicates the geomagnetic index after normalization operation; For the cleaned satellite orbit data, first directly filter out the data points with L < 2.5 using physical constraints, and perform normalization processing on the filtered L value (2.5 < L < 6.57) data to unify L values of different magnitudes to the same scale. The specific expression is: Among them, L norm is the L value after normalization; For the periodic encoding of MLT, convert MLT into a sine and cosine pair. The specific expression: For the latitude LAT, divide LAT by the maximum value of 0.35 radians to limit its value to the interval [-1, 1]. The specific expression is: Among them, LAT norm is the scaled latitude value; S1-3: Data segmentation: Divide the normalized data into a training set, a validation set, and a test set according to the time range; after removing the data in the test set, the remaining data is divided according to an 8:2 ratio, where 80% of the data is used as the training set for model parameter training, and 20% of the data is used as the validation set. S1-4: Data reshaping: Reshape the processed data into a three-dimensional tensor form; S1-5: Data re-segmentation: Extract satellite data and geomagnetic data from the reshaped data respectively to form two independent input branches, and input them into the CNN and BiLSTM neural network models for subsequent feature extraction and fusion processing.

3. The method for predicting circulating current based on CNN and BiLSTM neural network according to claim 1, characterized in that: The satellite data branch network model in step S2 includes the first convolution layer, the first normalization layer, the first Dropout layer, the second convolution layer, the second normalization layer, the flattening layer, the first fully connected layer and the second Dropout layer from the input end to the output end; the first convolution layer uses 64 filters, the convolution kernel size is 3, the convolution step is 1, the padding method is same, and the activation function is ReLU; the first normalization layer is used to perform batch normalization on the convolution output to accelerate model convergence and improve training stability; the first Dr The opout layer randomly discards 20% of the neurons to prevent overfitting. The second convolutional layer uses 128 filters, a convolution kernel size of 3, a convolution stride of 1, a padding method of same, and an activation function of ReLU to further extract higher-level spatial features. The second batch normalization layer further batch normalizes the convolution output features. The flattening layer is used to convert the multi-dimensional features output by the second batch normalization layer into one-dimensional features. The fully connected layer contains 64 neurons and the activation function is ReLU to extract high-dimensional abstract features of the flattened layer features. The Dropout layer randomly drops 30% of the neurons.

4. The method for predicting circulating current based on CNN and BiLSTM neural network according to claim 1, characterized in that: The geomagnetic data branch BiLSTM network structure described in step S2 uses a bidirectional long short-term memory network to capture the bidirectional long-term dependency of geomagnetic data in the time dimension. Its network structure includes, from the input end to the output end, a first BiLSTM layer, a first batch of normalization layers, a second BiLSTM layer, a fully connected layer, and a Dropout layer. The first bidirectional BiLSTM layer contains 128 neurons and is used to simultaneously capture feature information in both the positive and negative directions of the time series. The first batch of normalization layers is used to batch normalize the output features of the BiLSTM layer to accelerate model convergence and improve training stability. The second BiLSTM layer uses 64 neurons to further extract high-level features of geomagnetic data and reduce feature dimensions; the fully connected layer contains 64 neurons, and the activation function is ReLU, which is used to nonlinearly combine the temporal features extracted by the BiLSTM layer; the Dropout layer randomly discards 30% of the neurons.

5. The method for predicting circulating current based on CNN and BiLSTM neural network according to claim 1, characterized in that: The deep fully connected fusion network structure described in step S2 is used to fuse the feature information output by the satellite data branch and the geomagnetic data branch, and further extract and optimize the fused features. The specific network structure includes, from the input end to the output end: a first fully connected layer, comprising 128 neurons, with an activation function of ReLU, for fusing the feature information of the satellite data branch and the geomagnetic data branch; a first normalization layer, for performing batch normalization on the fused features to accelerate the model training process and improve stability; a first Dropout layer, randomly discarding 40% of the neurons to suppress overfitting; a second Dense layer, comprising 64 neurons, with an activation function of ReLU, for further extracting and optimizing the fused abstract features; a second batch normalization layer, further performing batch normalization on the features to improve the convergence stability of the model; a second Dropout layer, randomly discarding 20% of the neurons to further reduce the risk of model overfitting; an output layer, comprising one neuron, using a linear activation function, for outputting the final prediction result of the ring current proton flux.

6. The method for predicting circulating current based on CNN and BiLSTM neural network according to claim 1, characterized in that: The result analysis in step S4 specifically includes visualization of the loss curve and visualization of the prediction results, specifically: Loss curve visualization: plot the training loss and validation loss against the number of epochs, where the horizontal axis represents the number of training epochs and the vertical axis represents the loss value; Visualization of prediction results: A scatter plot is used to compare the actual proton flux value with the proton flux value predicted by the model. The scatter density is presented in a color gradient, from purple to yellow, indicating a low to high density. A standard diagonal line is also drawn in the figure; the mean square error (MSE) and the coefficient of determination (R) are marked in the upper left corner of the figure. 2 and other evaluation indicators.

Citation Information

Cited By

  • Dynamic current prediction method and device of power device

    CN120893480A

  • Method and apparatus for dynamic current prediction of power devices

    CN120893480B

  • CNN-BiLSTM-based power measurement data dynamic filling method

    CN120950844A