Global ionosphere TEC prediction method based on coding and decoding space-time Transform model

Through the method based on encoding and decoding the spatiotemporal Transformer model, the limitations of the existing TEC prediction model when dealing with long-range space and time dependencies are solved, and the global spatiotemporal features are efficiently extracted, which significantly improves the global TEC prediction accuracy and robustness.

CN120144967AActive Publication Date: 2025-06-13XIANGTAN UNIV

Patent Information

Application Number
CN202510592084.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-13
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing TEC spatiotemporal sequence prediction model has limitations when dealing with long-range spatial and temporal dependencies, and it is difficult to fully extract global spatial and temporal features, and faces the problems of gradient vanishing and gradient explosion.

Method used

The global ionosphere TEC prediction method based on the encoding and decoding spatiotemporal Transformer model is adopted to extract local and global spatiotemporal features through multi-branch convolutional layer and spatiotemporal Transformer layer, and the TEC prediction results are generated using the convolutional decoding layer.

Benefits of technology

It significantly improves the global TEC prediction accuracy, especially in harsh spatial weather conditions, maintains high prediction performance and robustness, demonstrates strong stability and generalization capabilities, and can accurately predict global ionosphere TEC maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144967A_ABST
    Figure CN120144967A_ABST
Patent Text Reader

Abstract

The invention discloses a global ionosphere TEC prediction method based on a coding and decoding space-time Transform model, and belongs to the technical field of ionosphere total electron content prediction, and the method comprises the steps: S1, obtaining and preprocessing TEC, Dst, Kp, SSN and F10.7 data, making a sample, and dividing the data into a training set, a verification set and a test set; s2, a space-time Transform model based on encoding and decoding is constructed, an encoder is composed of a multi-branch convolution layer and a space-time Transform layer, and a decoder is composed of a convolution decoding layer; and S3, training the model by using a training set, optimizing parameters through a loss function, adjusting and optimizing by using a verification set, and evaluating performance and generalization ability by using a test set. According to the method, the global TEC prediction precision is remarkably improved, relatively high prediction performance and robustness can be kept under a severe space weather condition, and efficient extraction of space-time local features and global features is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of total electron content prediction in the ionosphere, and specifically relates to a global ionospheric TEC prediction method based on an encoding-decoding spatio-temporal Transformer model. Background Art

[0002] The total electron content (TEC) of the ionosphere is a key parameter for measuring the state of the ionosphere, which is defined as the total number of electrons integrated over a unit cross-sectional area along the radio wave propagation path. The fluctuations of TEC are significantly affected by solar activity and geomagnetic activity, such as solar flares and geomagnetic storms. These activities can cause large fluctuations in TEC values, which in turn affect the propagation of radio signals on the ground and in the air. Especially during the transmission of Global Positioning System (GPS) signals, these fluctuations can cause signal delays and positioning errors, thereby posing potential threats to the safety of fields such as aviation, national defense, and space exploration. Therefore, accurately predicting TEC is of crucial significance for improving the stability of satellite communication and navigation systems and ensuring the safety of space missions.

[0003] With the continuous increase in data volume and the improvement of computing power, deep learning methods have gradually become a research hotspot for ionospheric TEC prediction. Compared with traditional methods, deep learning can better capture the non-linear features and complex spatio-temporal dependence relationships in the changes of ionospheric TEC, thus showing stronger prediction ability. Currently, there are mainly two construction methods for common TEC spatio-temporal sequence prediction models: one is to combine a convolutional neural network (CNN) with feature extraction units such as LSTM, BiLSTM, and ConvLSTM; the other is to sequentially stack spatio-temporal feature extraction units such as ConvLSTM and BiConvLSTM. However, these models still have two limitations: in terms of spatial feature extraction, CNN can efficiently extract local features of images through convolutional kernels and capture local information. However, due to the limitation of the receptive field, CNN has certain limitations in dealing with long-range spatial dependence relationships and is difficult to fully extract global spatial features; in terms of time feature extraction, although LSTM and its variants can effectively model long-term dependence relationships, due to the nature of its recursive structure, LSTM still faces problems of gradient disappearance and gradient explosion when dealing with long sequences, and it is difficult to capture global time dependence relationships. Summary of the Invention

[0004] In view of the above problems existing in the prior art, the purpose of the present invention is to provide a global ionospheric TEC prediction method based on an encoder-decoder spatio-temporal Transformer model, which significantly improves the global TEC prediction accuracy. Especially under severe space weather conditions, it can still maintain high prediction performance and robustness; it demonstrates strong stability and generalization ability, and can accurately predict the global ionospheric TEC map; it realizes the efficient extraction of spatio-temporal local features and global features.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows: A global ionospheric TEC prediction method based on an encoder-decoder spatio-temporal Transformer model, comprising the following steps: S1: Obtain and preprocess TEC, Dst, Kp, SSN, and F10.7 data, make samples and divide the data into a training set, a validation set, and a test set; where TEC is the total electron content, Dst is the disturbance storm time index, Kp is the planetary K index, SSN is the sunspot number, and F10.7 is the solar F10.7 radiation index. S2: Construct an encoder-decoder spatio-temporal Transformer model, where the encoder consists of a multi-branch convolutional layer and a spatio-temporal Transformer layer, which is responsible for extracting local spatio-temporal features and modeling global dependencies, and the decoder consists of a convolutional decoding layer, which is responsible for decoding the features extracted by the encoder and generating the TEC prediction result. S3: Use the training set to train the model, optimize the parameters through the loss function, tune the parameters using the validation set, and finally use the test set to evaluate the performance and generalization ability.

[0006] As a further improvement of the above technical solution: Step S1 further includes the following steps: S101: Obtain TEC, Dst, Kp, SSN, and F10.7 data for twenty-one consecutive years. S102: Use cubic spline interpolation method to interpolate and fill the missing data in the F10.7 data. S103: Unify the time resolution of TEC, Dst, Kp, SSN, and F10.7, and unify the spatial resolution of TEC, Dst, Kp, SSN, and F10.7. S104: Merge the TEC, Dst, Kp, F10.7, and SSN data at the same moment in the channel dimension to generate multi-channel data with the same dimension at the corresponding moment. S105: Scale the numerical range of each channel to the same range. S106: Use the sliding window method to generate samples. S107: Divide the data into a training set, a validation set, and a test set.

[0007] The step S2 includes the following steps: S201: Construct the input with multi-channel grid data of the past 24 hours; S202: Extract spatio-temporal local features by inputting the input data into a multi-branch convolutional layer; S203: Input the data output by the multi-branch convolutional layer into the spatio-temporal Transformer layer. After the data is processed by dimension reshaping, convolutional embedding, and spatial position encoding, it is passed to the spatial Transformer layer. This layer extracts global spatial features through the self-attention mechanism. After the output of the spatial Transformer layer is reshaped, transposed, and time position encoded, it is input into the temporal Transformer layer. This layer captures time dependencies through the self-attention mechanism and extracts global temporal features. Finally, the output of the temporal Transformer layer is reshaped, transposed, and then input into the convolutional decoding layer; S204: The convolutional decoding layer decodes the features extracted by the encoder to generate the final TEC prediction result.

[0008] In step S202, the multi-branch convolutional layer includes two parallel sub-branches: the first sub-branch is a layer of Conv3D + BN + ReLU units, and the second sub-branch is three stacked layers of Conv3D + BN + ReLU units. Finally, the outputs of the two parallel sub-branches are concatenated in the channel dimension and then passed through a layer of Conv3D + BN + ReLU units. Here, Conv3D represents three-dimensional convolution, BN represents batch normalization, ReLU represents the rectified linear unit, and "+" represents the sequential execution of multiple operations.

[0009] In step S203, both the spatial Transformer layer and the temporal Transformer layer are composed of 6 layers of Transformer encoders. Each layer of the encoder is stacked by a multi-head self-attention module and a multi-layer perceptron module. Layer normalization is applied before each module, and a residual connection is added after each module.

[0010] In step S203, during convolutional embedding, through two-dimensional convolution, the data with dimension ( ) is convolved to obtain an output with a shape of . The spatial dimension is flattened into a one-dimensional sequence to obtain embedding data with a shape of ( ), where , , , the convolutional kernel size is , and the stride is the same as the window size.

[0011] In step S203, the spatial position encoding is as follows: Define a trainable position encoding matrix S, and add it to the embedded data that has undergone convolution to introduce spatial position information. The formula is as follows: ; where is the embedded data, is a trainable position encoding matrix initialized with a normal distribution, as the input to the spatial Transformer.

[0012] In step S203, the temporal position encoding is as follows: Define a trainable position encoding matrix T, and add it to the data obtained by reshaping and transposing the output of the spatial Transformer layer to introduce temporal position information. The formula is as follows: ; where is the data obtained by reshaping and transposing, initialized with a normal distribution, as the input to the temporal Transformer.

[0013] In step S204, the convolutional decoding layer is composed of two layers of UpSampling+Conv3D+BN+ReLU units, one layer of UpSampling+Conv3D+ReLU units, and one layer of Conv3D+ReLU units connected in sequence. Here, UpSampling represents upsampling, Conv3D represents three-dimensional convolution, BN represents batch normalization, ReLU represents the rectified linear unit, and "+" represents the sequential execution of multiple operations.

[0014] In step S3, the mean squared error MSE is used as the loss function, and the Adam optimizer is used for optimization; the root mean squared error RMSE, mean absolute error MAE, and coefficient of determination R² are used as evaluation metrics for the model, and the performance of the trained model is evaluated based on these metrics and the test set data. The formulas for MSE, RMSE, MAE, and R 2 are as follows: ; ; ; ; where is the total number of data points, is the i-th true value, represents the i-th predicted value, is the mean value representing the true value.

[0015] The beneficial effects of the present invention are: (1) Compared with the existing TEC spatiotemporal prediction model, the present invention has significantly improved the global TEC prediction accuracy, especially under severe space weather conditions, and can still maintain high prediction performance and robustness; it shows strong stability and generalization ability, and can accurately predict the global ionospheric TEC map; the method is helpful to ensure the safety of space missions and improve the stability of satellite navigation systems.

[0016] (2) The method combines the advantages of convolutional neural networks (CNNs) in local feature extraction with the ability of Transformer self-attention mechanism in modeling global dependencies, achieving efficient extraction of spatiotemporal local features and global features.

[0017] (3) The encoder can effectively model local and global spatiotemporal features and effectively avoid the gradient vanishing and gradient exploding problems, thereby ensuring the stability of model training. The decoder part consists of a convolutional decoding layer, which is responsible for decoding the features extracted by the encoder and generating prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of the present invention.

[0019] Figure 2 It is the overall structure diagram of the model of the present invention.

[0020] Figure 3 It is a detailed structural diagram of the model of the present invention. DETAILED DESCRIPTION

[0021] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.

[0022] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "upper", etc. can be used here to describe the spatial positional relationship of a device or feature shown in the figure with other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation described in the figure of the device. For example, if the device in the attached drawing is inverted, the device described as "above other devices or structures" or "over other devices or structures" will then be positioned as "below other devices or structures" or "under other devices or structures". Thus, the exemplary term "above" can include both orientations of "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations are made for the spatial relative descriptions used here.

[0023] A global ionospheric TEC prediction method based on an encoding-decoding spatio-temporal Transformer model, as Figure 1 shown, includes the following steps: S1: Obtain and preprocess TEC, Dst, Kp, SSN, and F10.7 data, make samples and divide them into a training set, a validation set, and a test set. Among them, TEC is the total electron content; Dst is the disturbed storm time index, which is an index to measure the intensity of geomagnetic storms; Kp is the planetary K index, which is an index to describe the intensity of geomagnetic activity on a global scale; SSN is the sunspot number, which is an indicator to measure the level of solar activity; F10.7 is the solar F10.7 radiation index, which is the intensity of solar radiation at a wavelength of 10.7 cm.

[0024] Step S1 specifically includes the following steps: S101: Obtain TEC, Dst, Kp, SSN, and F10.7 data from 2003 to 2023. The global TEC map comes from the International GNSS (Global Navigation Satellite System) Service, i.e., IGS; the Kp, SSN, and F10.7 indices come from the Goddard Space Flight Center; the Dst index comes from the Kyoto World Data Center for Geomagnetism (WDC). The time resolution of the TEC map is 2 hours, and the coverage range is 180°W–180°E (longitude) and 87.5°N–87.5°S (latitude), with resolutions of 5° and 2.5° respectively, and the pixel size of each map is 71×73. Dst, Kp, F10.7, and SSN are global indices with different time resolutions, as shown in Table 1.

[0025] Table 1 Units and time resolutions of geomagnetic activity indices and solar activity indices

[0026] S102: Interpolate and fill in the missing values in the F10.7 index using cubic spline interpolation to restore a complete data sequence.

[0027] S103: To unify with the time resolution of the TEC data (2 hours), downsample the Dst index (original time resolution was 1 hour); upsample the Kp index (original time resolution was 3 hours), F10.7 index (original time resolution was 24 hours), and SSN (original time resolution was 24 hours) to unify the time resolution of the data to 2 hours. To match the spatial resolution of the TEC data (71×73), expand the global indices (Dst, Kp, SSN, and F10.7) into grid data of 71×73. Dst, Kp, SSN, and F10.7 are all global indices, which measure space weather phenomena globally and each has only one value at a certain moment. While the TEC data covers the global range and has a total of 71×73 data points. Therefore, to construct multi-channel data, indices such as Dst, Kp, SSN, and F10.7 need to be expanded into grid data of 71×73, the same as the TEC data.

[0028] S104: Merge the TEC data with the Dst, Kp, F10.7, and SSN data at the same moment in the channel dimension to generate multi-channel data of dimension (71, 73, 5) at the corresponding moment.

[0029] S105: To eliminate the influence of the magnitude of data in different channels, this paper uses the Min-Max normalization method to process the data in each channel, scaling the numerical range of each channel to between [0, 1].

[0030] S106: Use the sliding window method to generate samples. Take 48 consecutive hours (24 sheets) of multi-channel grid data as the sliding window. The data for the first 24 hours (12 sheets) is used as the input, and for the data in the last 24 hours (12 sheets), the geomagnetic activity and solar activity index channels are deleted and after inverse normalization processing, it is used as the output. After generating each sample, the window slides backward by 1 sheet of multi-channel grid data; S107: Divide the data into a training set, a validation set, and a test set. Among them, the training set is used for model training and parameter optimization, the validation set is used to evaluate hyperparameters and perform tuning, and the test set is used to evaluate the generalization ability of the model on unseen data to ensure the reliability and performance of the model.

[0031] The division method is as follows: The accuracy of TEC prediction is closely related to the intensity of solar activity. Therefore, select the data for three years as the test set, with one year having relatively active solar activity, one year having relatively weak solar activity, and the remaining year having solar activity intensity between the two. Randomly divide the data for the remaining years into the training set (accounting for 85%) and the validation set (accounting for 15%).

[0032] In this embodiment, the years 2015 with relatively active solar activity, 2019 with relatively weak solar activity, and 2017, the transitional year between the two, are selected as the test set. The data of the remaining years are used as the training set (accounting for 85%) and the validation set (accounting for 15%) respectively.

[0033] S2: Construct an ED-ST-Transformer model (encoding-decoding spatio-temporal Transformer model) based on the encoder-decoder architecture, and the structure is as Figure 2 shown. The encoder consists of a multi-branch convolutional layer and a spatio-temporal Transformer layer, which is responsible for extracting local spatio-temporal features and modeling global dependencies. The decoder consists of a convolutional decoding layer, which is responsible for decoding the features extracted by the encoder and generating the TEC prediction result. It specifically includes the following steps: S201: Construct a model that takes multi-channel grid data for the past 24 hours (time resolution is 2 hours) as input and predicts the global TEC grid map for the next 24 hours. Figure 2 In it, X represents the multi-channel grid input data, the subscript t represents time, and Xt represents the multi-channel input data corresponding to time t.

[0034] S202: Extract spatio-temporal local features by inputting the input data into the multi-branch convolutional layer.

[0035] As Figure 3As shown, the multi-branch convolutional layer contains two parallel sub-branches: one extracts simple features through a single Conv3D+BN+ReLU, and the other extracts complex features through three stacked Conv3D+BN+ReLU. Among them, Conv3D represents three-dimensional convolution, BN represents batch normalization, ReLU represents rectified linear unit, and the "+" represents the sequential combination and execution of multiple operations. Conv3D + BN + ReLU means that these three operations are applied in sequence, that is, first perform the three-dimensional convolution operation (Conv3D), then apply the batch normalization (BN) operation to the convolution result, and then perform the rectified linear unit operation. The same applies to similar expressions below and will not be repeated. Finally, the outputs of the two are concatenated (concatenation: combining multiple tensors along a specific dimension (here refers to the channel dimension)) and further processed through Conv3D+BN+ReLU to extract multi-level spatio-temporal features. The first sub-branch is a layer of Conv3D+BN+ReLU units with 32 channels, a convolution kernel size of 3×3×3, and a stride of 2×2×2; the second sub-branch is three stacked Conv3D+BN+ReLU units. The first layer of convolution has 16 channels, a convolution kernel size of 3×3×3, and a stride of 2×2×2. The second layer of convolution has 32 channels, a convolution kernel size of 3×3×3, and a stride of 1×1×1, with padding='same'. The third layer of convolution has 64 channels, a convolution kernel size of 3×3×3, and a stride of 1×1×1, with padding='same'. Finally, the outputs of the two parallel branches are concatenated in the channel dimension and then passed through a layer of Conv3D+BN+ReLU units with 64 channels, a convolution kernel size of 3×3×3, and a stride of 1×1×1, with padding='same'. Among them, Conv3D is used to capture local features, batch normalization (BN) accelerates training and improves stability, and ReLU activation enhances non-linear modeling. Among them, padding='same' is a way of padding the input data in a convolutional neural network so that the size of the output is the same as the size of the input.

[0036] S203: Spatio-temporal Transformer layer, the structure is as Figure 3 shown. The input of the spatio-temporal Transformer layer is the output of the multi-branch convolutional layer. The spatio-temporal Transformer layer passes the data with a dimension of into the spatio-temporal Transformer layer. The data goes through dimension reshaping ( After convolution embedding and spatial position encoding processing, it is passed to the Spatial Transformer layer, which extracts global spatial features through the self-attention mechanism. The output of the Spatial Transformer layer is reshaped and transposed in dimension ( ), and after temporal position encoding, it is input into the Temporal Transformer layer, which captures temporal dependencies through the self-attention mechanism and extracts global temporal features; finally, the output of the Temporal Transformer layer is reshaped and transposed in dimension ( ) and then input into the convolutional decoding layer. Among them, both the spatial Transformer layer and the temporal Transformer layer are composed of 6 layers of Transformer encoders, and each layer of encoder is alternately stacked by multi-head self-attention (MSA) and multi-layer perceptron (MLP) modules. Layer normalization (Layer Norm) is applied before each module, and a residual connection is added after each module. Among them, the embedding dimension of each layer is 128, and the hidden layer dimension of the multi-layer perceptron (MLP) is 256. Figure 3 In it, Multi-Head Attention is multi-head attention.

[0037] The following explains relevant concepts and processes: For the self-attention mechanism, the self-attention mechanism adopts scaled dot-product attention. By calculating the correlation weights between elements in the sequence, a weighted representation is generated for each element, thereby capturing global dependencies.

[0038] First, the input is multiplied by three weight matrices , , respectively to generate queries (Q), keys (K), and values (V).

[0039] ; Next, calculate the dot product between Q and K, and divide it by the square root of the dimension of the key to avoid the problem of gradient disappearance caused by too large input values. Through the function, calculate the normalized attention weights, and multiply them by V to finally obtain the weighted attention output. The calculation formula is as follows: ; Among them, , represents the set of real numbers, is the number of image patches or time steps, is the embedding dimension; is a learnable parameter matrix; , is the number of attention heads. indicates that X is an N-row M-column matrix, and each element in the matrix is a real number, that is, N, M are the number of rows and columns of the matrix, corresponding to the sequence length of the input data and the feature dimension of each data point, respectively. indicates multiplying Q by the transpose of K.

[0040] For multi-head attention (Multi-Head Attention), multi-head attention enables the model to concurrently focus on subspace information at different positions. Multiple attention heads compute in parallel, learn different features and dependencies, and finally concatenate the outputs and obtain the result through a linear transformation.

[0041] The calculation formula is as follows: ; where Concat means concatenation.

[0042] ; where , is a learnable parameter matrix.

[0043] For the residual connection, a residual connection mechanism based on pre-layer normalization is adopted. First, normalization is performed and then the residual connection is made. The calculation formula is as follows: ; where represents the input of the sublayer, represents the sublayer module (such as the multi-head attention mechanism or the feed-forward network), is the layer normalization operation.

[0044] For the multi-layer perceptron (MLP), the multi-layer perceptron (MLP) consists of two fully connected layers and the activation function ReLU, which is used to enhance the expression ability of the model. The calculation formula is as follows: ; where represents the input of the sublayer, the hidden layer dimension of the first fully connected layer is the hidden layer dimension of the second fully connected layer is . is the output of the multi-layer perceptron (MLP). It is the application of the ReLU (Rectified Linear Unit) activation function: for each input value, if the value is greater than 0, the value is output; if the value is less than or equal to 0, 0 is output. W 1 is a weight matrix, b 1 is the first bias term. : represents the input data and the weight matrix W 1 for the linear transformation (i.e., matrix multiplication), and add the bias term b1. W 2 is another weight matrix, b 2 is the bias term.

[0045] For convolutional embedding, the convolutional embedding is through two-dimensional convolution (the size of the convolutional kernel is , representing the height and width of the convolutional kernel respectively. The stride is the same as the window size), convolve the data of dimension ( ) to obtain an output of shape ; flatten the spatial dimension into a one-dimensional sequence to obtain an embedding representation of shape ( ) . Among them , , . The window usually refers to the size of the convolutional kernel used in the convolution operation. In this step, the size of the convolutional kernel is patch_size_h × patch_size_w, which means that the convolutional kernel slides on the input image and "scans" a region of size patch_size_h × patch_size_w each time. And the "window" refers to this region, that is, the image region covered by the convolutional kernel each time it operates. The stride is the distance that the convolutional kernel moves on the image, which can be in the vertical direction (up and down) and the horizontal direction (left and right) respectively.

[0046] For spatial position encoding, the spatial position encoding is, define a trainable position encoding matrix S, and add it to the data after convolutional embedding to introduce spatial position information. The formula is as follows: ; where, is the input patch embedding representation, is the trainable position encoding matrix initialized to a normal distribution. As the input of the spatial Transformer.

[0047] For the temporal position encoding, the temporal position encoding is defined as follows: a trainable position encoding matrix T is defined and added to the data obtained by reshaping and transposing the output of the Spatial Transformer layer to introduce temporal position information. The formula is as follows: ; ; where is the data obtained by reshaping and transposing, initialized with a normal distribution. is used as the input to the temporal Transformer.

[0048] S204, the convolutional decoding layer, has a structure as Figure 3 shown. This layer decodes the features extracted by the encoder and generates the final TEC prediction result. It is composed of two layers of UpSampling+Conv3D+BN+ReLU units, one layer of UpSampling+Conv3D+ReLU unit, and one layer of Conv3D+ReLU unit connected in sequence. Here, UpSampling represents upsampling. In the two layers of UpSampling+Conv3D+BN+ReLU units, the size of the first UpSampling is 3×3×3, the number of convolutional channels is 128, the size of the convolutional kernel is 3×3×5, and the stride is 2×1×1; the size of the second UpSampling is 3×3×3, the number of convolutional channels is 64, the size of the convolutional kernel is 3×3×5, and the stride is 2×1×1. In the one layer of UpSampling+Conv3D+ReLU unit, the size of UpSampling is 3×2×2, the number of convolutional channels is 32, the size of the convolutional kernel is 3×3×3, and the stride is 2×1×1. In the one layer of Conv3D+ReLU unit, the number of convolutional channels is 1, the size of the convolutional kernel is 3×2×2, and the stride is 1×1×1. Upsampling and expansion are performed through two layers of UpSampling+Conv3D+BN+ReLU; then, spatial information is enhanced using UpSampling+Conv3D+ReLU; finally, final optimization is performed using Conv3D+ReLU to generate the prediction result.

[0049] S3: Train the model using the training set, optimize the parameters through the loss function, tune using the validation set, and finally evaluate the performance and generalization ability using the test set. Specifically: S301: During the model training process, Adam (Adaptive Moment Estimation) is selected as the optimizer and Mean Squared Error (MSE) as the loss function. The Adam optimizer has the advantages of adaptive learning rate and efficient computation, accelerating convergence and improving generalization ability. Mean Squared Error (MSE) as the loss function has the advantages of simple calculation, easy optimization, and good performance for most regression problems. The formula for MSE is as follows: ; where, is the total number of data points, is the i-th true value, represents the i-th predicted value.

[0050] S302: In each training step, the data of the training set is input into the model for forward propagation, the loss between the predicted value and the actual label is calculated, and the model parameters are updated through the backpropagation algorithm. Subsequently, the current model is evaluated using the validation set. The validation set data is only used to evaluate the model performance and adjust hyperparameters, and does not participate in the model training process. Through the evaluation results on the validation set, hyperparameter tuning is performed to optimize the model performance and ensure its generalization ability on unseen data.

[0051] S302: After the model training is completed and adjusted through the validation set, the test set is used to evaluate the final model. Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R²) are used as the evaluation metrics for the model. The formulas for RMSE, MAE, and R 2 are as follows: ; ; ; where, is the total number of data points, is the i-th true value, represents the i-th predicted value, is the average value representing the true values.

[0052] S303: After the model performs well on the test set and reaches the expected performance metrics, the trained model weights and structure are saved. By archiving the model, it can be directly loaded in future applications, avoiding the computational overhead of retraining. The saved model can be used for actual deployment for real-time prediction.

[0053] Finally, it is necessary to clarify here that the above embodiments are only used to further illustrate the technical solutions of the present invention and should not be construed as limiting the protection scope of the present invention. Some non-essential improvements and adjustments made by those skilled in the art based on the above content of the present invention all fall within the protection scope of the present invention.

Claims

1. A global ionospheric TEC prediction method based on the encoding and decoding spatiotemporal Transformer model, characterized in that: The following steps are involved: S1: Obtain and preprocess TEC, Dst, Kp, SSN and F10.7 data, make samples and divide the data into training set, validation set and test set; among them, TEC is the total electron content, Dst is the disturbance storm time index, Kp is the planetary K index, SSN is the sunspot number, F10.7 is the solar F10.7 radiation index, S2: Construct a spatiotemporal Transformer model based on encoding and decoding, in which the encoder consists of a multi-branch convolutional layer and a spatiotemporal Transformer layer. The spatiotemporal Transformer layer consists of a spatial Transformer layer and a temporal Transformer layer, which is responsible for extracting local spatiotemporal features and modeling global dependencies. The input data is input into the multi-branch convolutional layer to extract spatiotemporal local features. The data output by the multi-branch convolutional layer is passed into the spatiotemporal Transformer layer. After dimension reshaping, convolution embedding and spatial position encoding, the data is passed to the spatial Transformer layer, which extracts global spatial features through a self-attention mechanism. The output of the spatial Transformer layer is dimensionally reshaped and transposed, and temporally encoded, and then input into the temporal Transformer layer, which captures temporal dependencies through a self-attention mechanism and extracts global temporal features. The decoder consists of a convolutional decoding layer, which is responsible for decoding the features extracted by the encoder and generating TEC prediction results. S3: Use the training set to train the model, optimize the parameters through the loss function, tune it using the validation set, and finally use the test set to evaluate the performance and generalization ability.

2. The prediction method according to claim 1, characterized in that: Step S1 also includes the following steps: S101: Obtain TEC, Dst, Kp, SSN and F10.7 data for 21 consecutive years; S102: Using cubic spline interpolation method to interpolate and fill in the missing data in F10.7 data; S103: Unify the temporal resolution of TEC, Dst, Kp, SSN, and F10.7, and the spatial resolution of TEC, Dst, Kp, SSN, and F10.7; S104: merging the TEC, Dst, Kp, F10.7 and SSN data at the same time in the channel dimension to generate multi-channel data with the same dimension at the corresponding time; S105: scaling the value range of each channel to the same range; S106: Generate samples using a sliding window method; S107: Divide the data into a training set, a validation set, and a test set.

3. The prediction method according to claim 1, characterized in that: The step S2 comprises the following steps: S201: construct multi-channel grid data with the past 24 hours as input; S202: extracting spatiotemporal local features by inputting input data into a multi-branch convolutional layer; S203: The data output by the multi-branch convolutional layer is passed to the spatiotemporal Transformer layer. After dimension reshaping, convolution embedding and spatial position encoding, the data is passed to the spatial Transformer layer. This layer extracts global spatial features through the self-attention mechanism. The output of the spatial Transformer layer is dimensionally reshaped and transposed, and temporally encoded, and then input to the temporal Transformer layer. This layer captures temporal dependencies through the self-attention mechanism and extracts global temporal features. Finally, the output of the temporal Transformer layer is reshaped and transposed before being input into the convolutional decoding layer; S204: The convolutional decoding layer decodes the features extracted by the encoder to generate the final TEC prediction result.

4. The prediction method according to claim 3, characterized in that: In step S202, the multi-branch convolutional layer includes two parallel sub-branches: the first sub-branch is a layer of Conv3D+BN+ReLU units, and the second sub-branch is three layers of stacked Conv3D+BN+ReLU units. Finally, the outputs of the two parallel sub-branches are concatenated in the channel dimension and then passed through a layer of Conv3D+BN+ReLU units, where Conv3D represents three-dimensional convolution, BN represents batch normalization, ReLU represents rectified linear unit, and "+" represents the sequential execution of multiple operations.

5. The prediction method according to claim 3, characterized in that: In step S203, both the spatial Transformer layer and the temporal Transformer layer are composed of 6 layers of Transformer encoders. Each layer of encoder is composed of a stack of multi-head self-attention modules and multi-layer perceptron modules. Layer normalization is applied before each module, and a residual connection is added after each module.

6. The prediction method according to claim 3, characterized in that: In step S203, when convolution embedding is performed, a two-dimensional convolution is performed on the dimension ( ) is convolved to obtain a shape of Output; Flatten the spatial dimension into a one-dimensional sequence, and get a shape of ( ), where , , , the convolution kernel size is , the step size is consistent with the window size.

7. The prediction method according to claim 6, characterized in that: In step S203, the spatial position encoding is: defining a trainable position encoding matrix S, and adding it to the convolved embedded data to introduce spatial position information, the formula is as follows: ; in, is the embedded data, is a trainable position encoding matrix initialized to a normal distribution, As input to the spatial Transformer.

8. The prediction method according to claim 6 or 7, characterized in that: In step S203, the temporal position encoding is to define a trainable position encoding matrix T and combine it with the output of the spatial Transformer layer after dimension reshaping and transposition. Add together to introduce time position information, the formula is as follows: ; in is the data after dimension reshaping and transposition. Initialized to a normal distribution, As input to the temporal Transformer.

9. The prediction method according to claim 3, characterized in that: In step S204, the convolution decoding layer is composed of two layers of UpSampling+Conv3D+BN+ReLU units, one layer of UpSampling+Conv3D+ReLU units and one layer of Conv3D+ReLU units connected in sequence, where UpSampling represents upsampling, Conv3D represents three-dimensional convolution, BN represents batch normalization, ReLU represents rectified linear unit, and "+" represents the sequential execution of multiple operations.

10. The prediction method according to claim 1, characterized in that: In step S3, the mean square error (MSE) is used as the loss function and the Adam optimizer is used for optimization. The root mean square error (RMSE), mean absolute error (MAE) and coefficient of determination (R²) are used as evaluation indicators of the model, and the performance of the trained model is evaluated based on these indicators and the test set data. MSE, RMSE, MAE and R² are the most accurate. 2 The formula is as follows: ; ; ; ; in, is the total number of data points, is the ith true value, represents the i-th predicted value, is the mean value representing the true value.

Citation Information

Patent Citations

  • Multivariate time series prediction method based on graph neural network and Transform

    CN117113054A

  • HRRP sequence identification method based on TSF-Transform-LS

    CN117540184A

  • Remote sensing image cloud detection method and system based on convolution and Transform cooperation

    CN117809199A

  • Traffic flow prediction method based on space-time bidirectional attention mechanism

    CN118644976A

  • Photovoltaic power generation capacity time sequence prediction method based on improved Transform model

    CN119557586A

Cited By

  • Ionized layer total electron content prediction method and device, computer equipment and medium

    CN121302305A

  • Ionosphere total electron content prediction method and device, computer equipment and medium

    CN121302305B

  • Ionized layer TEC prediction method and device based on parameterized multi-scale gated convolution

    CN121765296A