Low-altitude Thunderstorm Prediction Method in Spatiotemporal Two Dimensions Based on Incremental Learning

By adopting incremental learning methods in low-altitude thunderstorm prediction, a prediction model with space-time and dual dimensions is established, which solves the problems of low training efficiency and forgetting when facing new categories of data, and realizes high-precision electromagnetic waveform classification and prediction.

CN119669910BActive Publication Date: 2025-05-30HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510186179.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

When traditional deep learning algorithms face new categories of electromagnetic waveform data, they have problems of low training efficiency and catastrophic forgetting, making it difficult to effectively classify and predict low-altitude thunderstorms.

Method used

The low-altitude thunderstorm prediction method based on incremental learning is adopted. By establishing a new incremental module when each incremental batch arrives, and model training is performed based on the electromagnetic wave data set of the current and previous incremental batches, the prediction model is gradually iteratively updated.

Benefits of technology

It realizes that the new category of waveform data is gradually learned without losing the learned knowledge, which improves the classification accuracy and prediction ability, and avoids the defect of fixed model parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669910B_ABST
    Figure CN119669910B_ABST
Patent Text Reader

Abstract

The present invention discloses a spatio-temporal two-dimensional low-altitude thunderstorm prediction method based on incremental learning, which relates to the cross technical field of artificial intelligence and meteorology, and includes: collecting electromagnetic wave image data to form T electromagnetic wave data sets for low-altitude thunderstorm prediction. When each incremental batch arrives, a new incremental module and a current incremental batch prediction model are established. The Transformer and CNN feature extractors are used to perform spatio-temporal two-dimensional feature extraction on the electromagnetic wave image data. At the same time, the incremental learning method is used to iteratively train the prediction model through multi-batch electromagnetic wave image data. On the premise that the model does not lose the learned knowledge, new category electromagnetic waveform data is gradually learned, thereby ensuring the recognition and prediction capabilities of the finally obtained trained prediction model in the present invention, avoiding the defect of fixed model parameters, and improving the classification accuracy of the method of the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of artificial intelligence and meteorology, and in particular to a low - altitude thunderstorm prediction method based on incremental learning in spatio - temporal two - dimensions. Background Art

[0002] Low - altitude thunderstorms are typical severe convective weather phenomena. When they occur, they are accompanied by severe meteorological conditions such as strong winds, lightning, and short - term heavy precipitation, posing potential threats to many fields such as aviation safety and public transportation. In today's highly urbanized and industrialized society, the frequent occurrence and suddenness of low - altitude thunderstorms make their prediction an important task for ensuring life and property safety and maintaining the stable operation of critical infrastructure. During the generation of low - altitude thunderstorms, charges inside and outside the thunderstorm clouds continuously accumulate, transfer, and release, and the electric field intensity will fluctuate significantly. This complex electric - field change will generate a series of electromagnetic waveforms with specific characteristics, such as IC (intra - cloud flash), NBE (narrow bipolar event), PB (precursor breakdown), and RS (return stroke), etc. The appearance of these waveforms often indicates the formation of low - altitude thunderstorms. Therefore, the accurate identification of electromagnetic waveforms can provide valuable early warning or alarm information for thunderstorm prediction.

[0003] Although traditional deep - learning algorithms can also accurately classify different types of waveform data, due to the disadvantages of fixed models and parameters, traditional deep - learning algorithms have problems of low training efficiency or catastrophic forgetting when the classification model encounters new categories. Taking the classification of lightning waveforms as an example, assume that the basic classification model is trained on IC and NBE data. At this time, the basic model can effectively classify these two types of data. If later the detection station captures a batch of PB and RS type data that need to be classified, traditional deep - learning has two technical routes. The first is to discard the already trained basic classification model and put the old - category and new - category data together to retrain a model. The model trained in this way has high accuracy, but it does not make full use of the basic model, and when the data volume is large and the training time is long, retraining the model will consume a large amount of computing resources and time, so the training efficiency is low. The second is to fine - tune the basic model with new - category data. This reduces the training cost to a certain extent, but during the fine - tuning process, the model often tends to the new - category data, resulting in a decline in the classification ability of the original categories, that is, the model is prone to catastrophic forgetting of the old - category knowledge and low classification accuracy. Summary of the Invention

[0004] In order to overcome the defects of low classification accuracy for newly added electromagnetic - wave image categories and fixed models and parameters of traditional deep - learning algorithms in the above - mentioned prior art, the present invention proposes a low - altitude thunderstorm prediction method based on incremental learning in spatio - temporal two - dimensions.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A spatio-temporal two-dimensional low-altitude thunderstorm prediction method based on incremental learning, comprising:

[0006] S1: Collect electromagnetic wave images to form T electromagnetic wave data sets {D 0 , D 1 ,..., D t ,..., D T-1} for low-altitude thunderstorm prediction; where D t represents the new category electromagnetic wave data set in the -th incremental batch, T - 1 is the total number of incremental batches, and when t = 0, it represents the initial electromagnetic wave data set;

[0007] S2: When the t-th incremental batch arrives, establish a new incremental module , and train it based on the total electromagnetic wave data set D t model after the t-th incremental batch, i.e., the current incremental batch, and the prediction model F t-1 (x) of the previous incremental batch to obtain the prediction model F t (x) of the current incremental batch; where D t model = {D 0 ∪ D 1 ... D t}, and x represents the electromagnetic wave image;

[0008] According to the method of step S2, after T - 1 incremental batches, the final prediction model F T-1 (x) is obtained;

[0009] S3: Input the electromagnetic wave image to be measured into the final prediction model F T-1 (x) for prediction to obtain the predicted category.

[0010] Preferably, the prediction model F t-1 (x) of the previous incremental batch consists of a feature extractor and a linear classifier, and is expressed as follows:

[0011] ;

[0012] Where Ф t-1 (·) is the feature extractor of the previous incremental batch, and , D is the original feature dimension size of the electromagnetic wave image x, d is the feature dimension size after being extracted by the feature extractor, W t-1 is the parameter of the linear classifier of the previous incremental batch, and , represents the total number of categories in the (t - 1)-th incremental batch stage, and the superscript T represents transpose.

[0013] Preferably, the current incremental batch prediction model F t (x) is composed of the previous incremental batch prediction model F t-1 (x) and the new incremental module and is expressed as follows:

[0014] ;

[0015] where x is the electromagnetic wave image, y is the true value of the category of the electromagnetic wave image, represents training the new incremental module using the cross-entropy loss function, D is the training set obtained from the total electromagnetic wave dataset D t train after the current incremental batch, argmin(.) represents the function to find the optimal parameters, and E is the expectation function. t model

[0016] Preferably, the new incremental module can be expressed as:

[0017] ;

[0018] where φ t (·) is the feature extractor of the new incremental module, which is composed of the CNN spatial feature extractor φ S t (·) and the transformer temporal feature extractor φ t TE (·), and ;

[0019] ; x s and x TE are respectively the normalized matrix and the waveform time matrix after the electromagnetic wave image is preprocessed in the spatial dimension and the temporal dimension; d s and d TE are respectively the feature dimension sizes after being extracted by the CNN spatial feature extractor φ t S (·) and the transformer temporal feature extractor φ t TE (·);

[0020] are the linear classifier parameters of the new incremental module, and , , represents the total number of categories in the t-th incremental batch stage, ​For the linear classifier parameters corresponding to the old category in the new increment module, ;

[0021] For the linear classifier parameters corresponding to the new category in the new increment module, .

[0022] Preferably, the predicted category has the following calculation formula:

[0023] ;

[0024] where 0 is a 0 matrix of the same size as the matrix and S(.) is the softmax, i.e., the normalized exponential function.

[0025] Preferably, the new increment module obtains the optimal parameters through continuous training. The expression of the optimal parameters is:

[0026] ;

[0027] where Dis(.) is Distance(.), representing the function for calculating distance; represents the parameter.

[0028] Preferably, the current increment batch prediction model F t (x) is expressed as:

[0029] ;

[0030] where Ф t (·) is the current increment batch feature extractor; W t is the linear classifier parameter of the current increment batch prediction model F t (x); and are respectively the linear classifier parameters corresponding to the old category and the new category in the new increment module; φ t (·) is the feature extractor of the new increment module; γ 1 , γ 2 are scale factors, and 0 < γ 1 < 1, 0 < γ 2 < 1, and γ is the diagonal matrix composed of γ 1 and γ 2 ; the diagonal matrix composed of γ 1 , γ 2 can be expressed as:

[0031] ;

[0032] Among them, S old is the number of old-class samples in the total electromagnetic wave dataset after the current incremental batch, and S new is the number of new-class samples in the total electromagnetic wave dataset after the current incremental batch.

[0033] Preferably, the loss function of the current incremental batch prediction model F t (x) is calculated as follows:

[0034] ;

[0035] Among them, α 1 , α 2 , α 3 are hyperparameters, and α 1 + α 2 + α 3 = 1, is the first loss function, is the second loss function, is the distillation loss function, y is the true value of the electromagnetic wave category, is the cross-entropy loss function, S(.) is the softmax, i.e., the normalized exponential function, and W t fe is a parameter of a new linear classifier.

[0036] Preferably, the acquisition of the current incremental batch prediction model includes:

[0037] d1: Set the current incremental batch prediction model F t (x) as the teacher model, and at the same time set the student model at the t-th incremental batch stage; among them, the feature extractor of the student model is also jointly composed of a CNN spatial feature extractor and a transformer temporal feature extractor;

[0038] d2: Use the second distillation loss function to train the student model ;

[0039] d3: Use the trained student model as the final current incremental batch prediction model;

[0040] The student model can be expressed as:

[0041] ;

[0042] Among them, is the parameter of the linear classifier of the student model at the t-th incremental batch stage,​ , is the feature extractor of the student model in the current incremental batch stage . They are respectively the CNN spatial feature extractor and the transformer temporal feature extractor of the student model in the current incremental batch stage . , , and d = d s + d TE , x s , x TE are respectively the normalized matrix and the waveform time matrix after the electromagnetic wave image is pre - processed in the spatial dimension and the temporal dimension; d s , d TE are respectively the feature dimension sizes after being extracted by the CNN spatial feature extractor and the transformer temporal feature extractor ; represents the total number of categories in the t - th incremental batch stage.

[0043] Preferably, the processing process of the current incremental batch prediction model includes:

[0044] A1: Pre - process the electromagnetic wave image in the spatial dimension and the temporal dimension respectively to obtain an image pre - processing matrix, and the image pre - processing matrix includes a normalized matrix and a waveform time matrix;

[0045] A2: Input the image pre - processing matrix into the spatio - temporal feature extractor in the current incremental batch prediction model to obtain a temporal feature matrix and a spatial feature matrix; wherein, the spatio - temporal feature extractor includes a CNN spatial feature extractor and a Transformer temporal feature extractor;

[0046] A3: Input the temporal feature matrix and the spatial feature matrix into a feature splicing module for feature splicing to obtain a spatio - temporal feature matrix;

[0047] A4: Input the spatio - temporal feature matrix into a linear classifier to obtain a predicted category.

[0048] The advantages of the present invention are:

[0049] (1) The present invention uses a Transformer and a CNN feature extractor to perform spatio-temporal two-dimensional feature extraction on electromagnetic wave images. At the same time, an incremental learning method is used to iteratively train the prediction model for multi-batch electromagnetic wave image data through different category waveform data. Without losing the learned knowledge of the model, new category waveform data is gradually learned, thus ensuring the recognition and prediction capabilities of the current incremental batch prediction model in the present invention, avoiding the defect of fixed model parameters, and improving the classification accuracy of the method of the present invention.

[0050] (2) The present invention first performs data preprocessing on electromagnetic wave images in the spatial dimension and the time dimension to obtain a normalized matrix and a waveform time matrix. Then, spatio-temporal two-dimensional feature extraction is performed on the electromagnetic wave images, reducing unnecessary channel information and computational complexity, and improving the prediction speed of the method of the present invention.

[0051] (3) The present invention uses feature enhancement technology to extract key features of electromagnetic wave images in different incremental batches, avoiding strong classification biases in the model caused by the imbalance between new and old category waveform data, and improving the accuracy of the model in the electromagnetic waveform classification task.

[0052] (4) The present invention uses feature compression technology to reduce redundant information in the features of electromagnetic wave images, thereby improving the computational efficiency and prediction accuracy of the prediction model in the method of the present invention.

[0053] (5) The present invention uses a first loss function, a second loss function, and a distillation loss function, and adds hyperparameters among the three to form a loss function , and uses the loss function to train the current incremental batch prediction model, gradually optimizing the model parameters, avoiding overfitting of the model to the old category waveform data with a small amount of data, and improving the accuracy of the model in the electromagnetic waveform classification task.

[0054] (6) The present invention sets the current incremental batch prediction model as the teacher model, and at the same time sets a student model in the t-th incremental batch stage. Based on the teacher model, the student model is trained using the second distillation loss function; the trained student model is used as the final current incremental batch prediction model, avoiding the rapid growth of the model's parameter quantity and feature dimension when gradually adding new incremental modules, making it suitable for the incremental learning task of multi-incremental batch data in low-altitude thunderstorm prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is the model architecture diagram of the present invention;

[0056] Figure 2 is the flowchart of the prediction method of the present invention;

[0057] Figure 3 This is the flowchart of the model architecture method of the present invention;

[0058] Figure 4 These are the test accuracies of the method of the present invention when the incremental batches are 0, 1, and 2;

[0059] Figure 5 This is the comparison chart of the test accuracies of the method of the present invention and the finetune method when the incremental batch is 2. Detailed implementation manners

[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0061] Embodiment 1

[0062] As Figure 1-2 shown, the present invention proposes a spatio-temporal two-dimensional low-altitude thunderstorm prediction method based on incremental learning, including:

[0063] S1: Collect the electromagnetic wave images generated by the change of the electric field intensity, perform manual selection and labeling to form T electromagnetic wave data sets {D 0 , D 1 ,..., D t ,..., D T-1} for low-altitude thunderstorm prediction. Among them, D t represents the new category electromagnetic wave data set in the th incremental batch. T - 1 is the total number of incremental batches. When t = 0, it represents the initial electromagnetic wave data set;

[0064] Each electromagnetic wave data set contains electromagnetic wave images and corresponding categories, and the categories in each electromagnetic wave data set do not overlap;

[0065] S2: When each incremental batch arrives, establish the current incremental batch prediction model F t (x), and use the electromagnetic wave data set D t model after the current incremental batch for training to obtain the trained current incremental batch prediction model F t (x);

[0066] Among them, the current incremental batch refers to the t modelIs a mixture of old category data and new category data, i.e., D t model ={D 0 ∪D 1 ...D t}; The current incremental batch prediction model F t (x) is trained based on the previous incremental batch prediction model F t-1 (x), combined with the new category electromagnetic wave data set in the t-th incremental batch. The previous incremental batch is the (t - 1)-th incremental batch; Divide D t model into a training set D t train and a test set D t test at a ratio of 8:2 for model training and testing respectively;

[0067] Finally, the trained final prediction model F T-1 (x) is obtained;

[0068] S3: Input the electromagnetic wave image to be measured into the final prediction model F T-1 (x) for prediction to obtain the predicted category.

[0069] The current incremental batch prediction model F t (x) is trained in the t-th incremental batch stage. In the t-th incremental batch stage, the prediction model F of the (t - 1)-th incremental batch (the previous incremental batch) is saved t-1 (x). The (t - 1)-th incremental batch prediction model F t-1 (x) consists of a feature extractor and a linear classifier, and is expressed as follows:

[0070] ;

[0071] Among them, Ф t-1 (·) is the feature extractor of the previous incremental batch, and , D is the original feature dimension size of the electromagnetic wave image data sample x, W t-1 is the parameter of the linear classifier of the previous incremental batch, and , d represents the feature dimension size after being extracted by the feature extractor, represents the total number of categories in the (t - 1)-th incremental batch stage, and the superscript T represents transpose.

[0072] In the process of low-altitude thunderstorm prediction, when a new type of electromagnetic wave image arrives, fine-tuning or simply freezing F t-1 (·) will reduce the model's processing ability for old category data or lose plasticity for new category data. Inspired by model accumulation, the present invention trains a new incremental module in the t-th incremental batch stage Perform incremental prediction, expressed as:

[0073] ;

[0074] where φ t (·) is the feature extractor of the new increment module, and ; is the parameter of the linear classifier of the new increment module, and ; represents the parameter of the linear classifier of the new increment module, whose size is the concatenation of the size of the parameter of the linear classifier of the old category and the size of the parameter of the linear classifier of the new type of data after the arrival of the new increment batch, that is ; is the parameter corresponding to the size of the parameter of the linear classifier of the old category, used to learn to re - mine the features of the old category after the arrival of the new increment batch; is the parameter of the linear classifier for classifying the new type of data after the arrival of the new increment batch; represents the total number of categories in the t - th increment batch stage.

[0075] Because the electromagnetic wave image data has strong spatial and temporal characteristics, the present invention adopts a spatio - temporal two - dimensional classification technology. The pre - processed electromagnetic wave image data is input into the CNN spatial feature extractor and the transformer temporal feature extractor, respectively obtaining the spatial feature matrix and the temporal feature matrix of the electromagnetic wave image. The two are concatenated to form the spatio - temporal feature matrix of the electromagnetic wave image data. Therefore, it can be further known that the feature extractor φ in the new increment module t (x) is composed of the CNN spatial feature extractor φ t S (x s ) and the transformer temporal feature extractor φ t TE (x TE ). Therefore, the new increment module can be expressed as:

[0076] ;

[0077] where , , x s , x TE are respectively the normalized matrix and the waveform time matrix of the electromagnetic wave image after data pre - processing in the spatial dimension and the temporal dimension; d s , d TE are respectively the results after passing through the CNN spatial feature extractor φ t S(·) and the transformer time feature extractor φ t TE The dimensionality size of the features after (·) extraction.

[0078] Based on the prediction model F of the previous incremental batch (the (t - 1)th incremental batch) t-1 (x) and the new incremental module , obtain the prediction model F of the current incremental batch t (x), which is expressed as follows:

[0079] ;

[0080] Among them, x is the electromagnetic wave image, y is the true value of the electromagnetic wave category, Indicates using the cross - entropy loss function to train the new incremental module , argmin(.) represents the function to find the optimal parameters, and E is the expectation function.

[0081] The final predicted category of the present invention The calculation formula is:

[0082] = ;

[0083] Among them, 0 is a 0 matrix of the same size as , used to unify the dimensionality of the linear classifier parameters, S(.) is the softmax, that is, the normalized exponential, W T t-1 and Ф t-1 (·) are both the frozen parameters of the prediction model F of the previous incremental batch t-1 (x) (will not change during the training process in the t - th incremental batch stage), and the trainable modules are , , .

[0084] The optimization objective of the new incremental module is to minimize the gap between the true value and the category prediction value of the prediction model F of the current incremental batch t (x). The new incremental module obtains the optimal parameters through continuous training. The expression of the optimal parameters is:

[0085] ;

[0086] Among them, Dis(.) is the same as Distance(.), representing the function to calculate the distance; represents the parameter.

[0087] Therefore, during training, the optimization process of the new incremental module can be described as a process of finding an initial loss function that continuously decreases The initial loss function has the following calculation formula:

[0088] .

[0089] Example 2

[0090] Based on Example 1, as the second implementation manner of the present invention, the current incremental batch prediction model F t (x) can be expressed as:

[0091] .

[0092] In the present invention, because the scale of the electromagnetic wave dataset is large, in order to save memory capacity, only part of the old-class electromagnetic wave image data is used when mixing new and old-class electromagnetic wave image data. At the incremental batch stage, in the category space , set the memory data capacity to M, and set C t as the number of waveform categories in t = M / C t , to form the old dataset , where is the pruned old electromagnetic wave dataset.

[0093] Therefore, at the incremental batch stage, as the number of new-class electromagnetic wave images increases, the model will obtain a dataset with an unbalanced ratio of new and old classes , and divide D t model into a training set D t train and a test set D t test at a ratio of 8:2 for model training and testing respectively. Similarly, the old-class electromagnetic wave images in D t train are also much fewer than the new-class electromagnetic wave images. Since the imbalance of D t train will cause a strong classification bias in the model of the present invention, that is, the model will tend to ignore the old-class electromagnetic wave image data with a smaller quantity. Therefore, to strengthen the learning of the old-class sample data and alleviate the classification bias, the present invention adds in the current incremental batch prediction model of A scaling factor is added to the output to obtain the new output. Therefore, the current incremental batch prediction model F t (x) can be expressed as:

[0094] ;

[0095] where 0 < γ 1 < 1, 0 < γ 2 < 1, γ is a diagonal matrix composed of γ 1 and γ 2 . γ 1 and γ 2 are obtained according to the number of samples included in the old and new classes of the current electromagnetic wave dataset. The diagonal matrix composed of γ 1 and γ 2 can be expressed as:

[0096] ;

[0097] where S old is the number of old-class samples in the total electromagnetic wave dataset after the current incremental batch, i.e., S old = M, S new is the number of new-class samples in the total electromagnetic wave dataset after the current incremental batch. Since the new-class data in the electromagnetic wave dataset is much larger than the old-class data, there is γ 1< γ 2 . Through this scaling strategy, the absolute value of the logit value of the old waveform class decreases, and the absolute value of the logit value of the new waveform class increases, thereby forcing the current incremental batch prediction model F t (x) to generate larger logit values for the old class and smaller logit values for the new class to alleviate the classification bias. Therefore, after adding a scaling factor to the basic loss function , the first loss function is obtained:

[0098] .

[0099] Simply making the new incremental module fit the residuals of F t and F t-1 is sometimes insufficient. For example, in extreme cases, the residuals of F t and F t-1 are 0. Therefore, the new incremental module should be further prompted to learn the new-class battery wave image data, and a new linear classifier W t fe is initialized to learn new features. The second loss function is:

[0100] 。

[0101] Training simply using class true values in an imbalanced dataset may lead to overfitting of the old class waveform data with a smaller amount of data. The mitigation method is to encourage F using knowledge distillation t to maintain a similar output to F on the old waveform class, and the distillation loss function t-1 is as follows: :

[0102] 。

[0103] In summary, the total loss function is obtained as follows:

[0104] ;

[0105] where α 1 , α 2 , α 3 are hyperparameters, and α 1 +α 2 +α 3 = 1. The larger the value, the more important the corresponding loss function is. In the incremental batch stage, the entire training process is continuously repeated in multiple rounds. Through multiple forward propagations, calculations of the total loss function , backpropagations, and parameter update loop operations, the parameters of the current incremental batch prediction model are gradually optimized, making the total loss function gradually decrease, thereby gradually improving the accuracy of the model in the waveform classification task. After the iteration ends, the prediction model F t (x) of the t-th incremental batch stage is obtained.

[0106] In this embodiment, α 1 = α 2 = 0.3, α 3 = 0.4, indicating that more importance is attached to the output similarity of the old and new models on the old class waveform data during the optimization process. This setting can effectively mitigate catastrophic forgetting during the incremental process.

[0107] Embodiment 3

[0108] As the third implementation manner of the present invention, the acquisition of the current incremental batch prediction model includes:

[0109] d1: Set the current incremental batch prediction model F t (x) in the second implementation manner as the teacher model, and at the same time set the student model in the t-th incremental batch stage ; wherein, the feature extractor of the student model is also jointly composed of a CNN spatial feature extractor and a transformer temporal feature extractor;

[0110] d2: Based on the teacher model, use the second distillation loss function to train the student model ;

[0111] d3: Use the trained student model as the final current incremental batch prediction model.

[0112] Through Embodiment 3, low-altitude thunderstorm prediction is carried out, avoiding the rapid growth of the number of model parameters and feature dimensions when the model gradually adds new incremental modules, making it suitable for the incremental learning task when applying to multi-incremental batch data of low-altitude thunderstorm prediction.

[0113] The student model can be expressed as:

[0114] ;

[0115] wherein, is the linear classifier parameter of the student model in the t-th incremental batch stage , is the feature extractor of the student model in the t-th incremental batch stage , , are respectively the CNN spatial feature extractor and the transformer temporal feature extractor of the student model in the t-th incremental batch stage , d = d s + d TE , x s , x TE are respectively the normalized matrix and the waveform time matrix after preprocessing the electromagnetic wave image in the spatial dimension and the temporal dimension;

[0116] The calculation formula of the second distillation loss function is:

[0117] .

[0118] Embodiment 4

[0119] Based on the above Embodiment 1 or 2 or 3, the current incremental prediction model processing process includes:

[0120] A1: Preprocess the electromagnetic wave images in the electromagnetic wave dataset in the spatial dimension and the temporal dimension respectively to obtain an image preprocessing matrix, and the image preprocessing matrix includes a normalized matrix and a waveform time matrix;

[0121] The preprocessing in the spatial dimension includes:

[0122] To reduce unnecessary channel information and computational complexity, the electromagnetic wave image is grayscaled and converted into a grayscale waveform picture; all grayscale waveform pictures are adjusted to the same size to ensure that the model receives the same input size; then the grayscale waveform pictures are converted into tensor form to meet the input requirements of the deep learning framework; the mean and standard deviation of the grayscale waveform pictures are calculated for normalization processing to obtain a normalized matrix containing waveform features.

[0123] In the embodiment of the present invention, the unified size is set to 32×32. Therefore, the number of pixel blocks p = 32×32, and the batch size is set to 128. So after the preprocessing in the spatial dimension, the input of the spatio-temporal feature extractor in the incremental prediction model is of size (128, 1, 32, 32).

[0124] Batch size: The advantage of deep learning is that it is good at processing large-scale data and can process a batch of data at one time. Taking image data as an example, in one classification, if the model can process or classify 128 pictures at one time, then its batch size is 128.

[0125] The preprocessing in the time dimension includes:

[0126] On the basis of the completion of the preprocessing in the spatial dimension, the normalized matrix is sliced along the time series (the horizontal axis direction of the image) and divided into S segments of the same size. Assuming the picture size is ( ), then the size of each segment is , and the pixel information in each segment is regarded as the feature vector of one time step. Subsequently, each segment is flattened and stacked together to obtain a waveform time matrix of S×P / S. In this way, a waveform picture is converted into a waveform feature matrix containing a sequence of waveform time features for the input of the subsequent Transformer time feature extractor.

[0127] In the embodiment of the present invention, S = 16 is set, that is, each waveform picture is cut into 16 feature vectors of 32×2. After flattening, 16 feature vectors of 1×64 are obtained. After stacking these 16 feature vectors, a waveform feature matrix of 16×64 is obtained. Then the input size of the Transformer time feature extractor is (128, 16, 64).

[0128] A2: Input the image preprocessing matrix into the spatio-temporal feature extractor in the incremental prediction model to obtain a time feature matrix and a spatial feature matrix; among them, the spatio-temporal feature extractor includes a CNN spatial feature extractor and a Transformer time feature extractor;

[0129] A3: Input the time feature matrix and the spatial feature matrix into the feature concatenation module for feature concatenation to obtain the spatio-temporal feature matrix;

[0130] A4: Input the spatio-temporal feature matrix into the linear classifier to obtain the predicted category.

[0131] The effectiveness and superiority of the above spatio-temporal two-dimensional low-altitude thunderstorm prediction method based on incremental learning are verified in combination with specific experiments below.

[0132] There are 4 types of electromagnetic wave image categories collected in the present invention: IC, NBE, PB, RS

[0133] During the experiment, these 4 waveform categories are randomly sorted, and the electromagnetic wave image data of the first 2 sorted categories are selected as the initial dataset D 0 , and the electromagnetic wave images of the latter two waveform categories are respectively set as D 1 , D 2 . Set IC: 0, NBE: 1, PB: 2, RS: 3. After random sorting, the list [1, 2, 3, 0] is obtained, that is, the electromagnetic wave images of the NBE and PB categories are used as the initial dataset to train the basic classification prediction model, and RS and IC are the new categories that come later.

[0134] In the experiment, the memory data capacity M = 1200, the learning rate is 0.1, the weight decay is 0.0005, the number of training epochs epoch = 30, and the resnet32 (spatial) and vit-base (time) feature extractors are used to extract spatio-temporal features of the electromagnetic wave images.

[0135] The batch size is set to 128. During the preprocessing process, the adjusted size of the image is 32 * 32, so the input size of the CNN spatial feature extractor is 128 * 1 * 32 * 32 (batch size * number of channels * width * length). The electromagnetic wave images are sliced from left to right, and the number of segments is set to 16, that is, each waveform picture is cut into 16 feature vectors of 32 * 2. After flattening, 16 feature vectors of 1 * 64 are obtained. After stacking these 16 feature vectors, a 16 * 64 feature matrix is obtained. Then the input size of the Transformer time feature extractor is (128, 16, 64), and the hyperparameter α of the loss function 1 , α 2 , α 3 , α 1 = α 2 = 0.3, α 3 = 0.4.

[0136] The experimental results of the incremental test prediction accuracy are as Figure 4As shown in the figure, where incremental Test Accuracy is the incremental test accuracy, Epoch is the time, and incremental_batch is the number of incremental batches. As new class batches arrive, the model prediction accuracy will decrease to a certain extent. When the incremental batches are 0, 1, and 2, the average prediction accuracies are 97.44%, 95.81%, and 89.56% respectively.

[0137] In this experiment, the proposed method of the present invention was also compared with the finetune method. The comparison results when the incremental batch is 2 are as Figure 5 shown. According to the experimental results, the prediction effect of the method in this embodiment is significantly better than that of the finetune method. Therefore, it can be seen from the experimental results that the method of the present invention performs well on the electromagnetic wave data set for low-altitude thunderstorm data prediction.

[0138] Of course, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claimed rights.

[0139] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0140] The technologies, shapes, and structures not detailedly described in the present invention are all well-known technologies.

Claims

1. A low-altitude thunderstorm prediction method based on time and space dual dimensions of incremental learning, characterized in that: include: S1: Collect electromagnetic wave images to form T electromagnetic wave datasets {D0, D1, ..., D t ,...,D T-1 }; where D t Indicates The electromagnetic wave dataset of the new category in the incremental batch, T-1 is the total number of incremental batches, when t=0, it represents the initial electromagnetic wave dataset; S2: When the tth incremental batch arrives, a new incremental module is established , and based on the total electromagnetic wave data set D after the tth incremental batch, i.e. the current incremental batch t model and the previous incremental batch prediction model F t-1 (x) is trained to obtain the current incremental batch prediction model F t (x); where D t model ={D0∪D1...D t }, x represents the electromagnetic wave image; According to step S2, after T-1 incremental batches, the final prediction model F is obtained. T-1 (x); S3: Input the electromagnetic wave image to be measured into the final prediction model F T-1 (x) to make a prediction and obtain the prediction category; The current processing of the incremental batch prediction model includes: A1: Preprocess the electromagnetic wave image in spatial dimension and time dimension respectively to obtain the image preprocessing matrix, which includes a normalization matrix and a waveform time matrix; A2: Input the image preprocessing matrix into the spatiotemporal feature extractor in the current incremental batch prediction model to obtain a temporal feature matrix and a spatial feature matrix; the spatiotemporal feature extractor includes a CNN spatial feature extractor and a Transformer temporal feature extractor; A3: Input the temporal feature matrix and the spatial feature matrix into the feature splicing module for feature splicing to obtain the temporal and spatial feature matrix; A4: Input the spatiotemporal feature matrix into the linear classifier to obtain the predicted category.

2. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 1, characterized in that: The previous incremental batch prediction model F t-1 (x) consists of a feature extractor and a linear classifier, expressed as follows: ; Among them, Ф t-1 (·) is the feature extractor of the previous incremental batch, and , D is the original feature dimension size of the electromagnetic wave image x, d is the feature dimension size after being extracted by the feature extractor, W t-1 is the linear classifier parameter of the previous incremental batch, and , Indicates the total number of categories in the t-1th incremental batch stage, with a superscript T Indicates transpose.

3. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 2, characterized in that: The current incremental batch prediction model F t (x) is the prediction model F of the previous incremental batch t-1 (x) and new incremental modules Combined, it is expressed as follows: ; Among them, x is the electromagnetic wave image, y is the true value of the category of the electromagnetic wave image, Indicates the use of cross entropy loss function for the new incremental module Conduct training, D t train is the total electromagnetic wave dataset D after the current incremental batch t model The training set obtained in , argmin(.) represents the function of finding the optimal parameters, and E is the expected function.

4. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 3, characterized in that: The new incremental module It is expressed as: ; Among them, φ t (·) is the feature extractor of the new incremental module, which is composed of the CNN spatial feature extractor φ S t (·) and the transformer temporal feature extractor φ t TE (·) composition, and ; ;x s 、x TE are the normalized matrix and waveform time matrix of the electromagnetic wave image after spatial dimension and time dimension data preprocessing respectively; d s ,d TE They are respectively passed through the CNN spatial feature extractor φ t S (·) and the transformer temporal feature extractor φ t TE (·) Feature dimension size after extraction; are the linear classifier parameters of the new incremental module, and , , represents the total number of categories in the tth incremental batch stage, are the linear classifier parameters corresponding to the old categories in the new incremental module, ; are the linear classifier parameters corresponding to the new categories in the new incremental module, .

5. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 4, characterized in that: Prediction Category The calculation formula is: ; Among them, 0 is the matrix For zero matrices of the same size, S(.) is softmax, which is a normalized exponential function.

6. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 5, characterized in that: The new incremental module Through continuous training, the optimal parameters are obtained , the optimal parameters The expression is: ; Among them, Dis(.) is Distance(.), which means the function of finding the distance; Indicates a parameter.

7. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 2, characterized in that: Current incremental batch prediction model F t (x) is expressed as: ; Among them, Ф t (·) is the current incremental batch feature extractor; W t Predict model F for the current incremental batch t (x) linear classifier parameters; and are the linear classifier parameters corresponding to the old and new categories in the incremental module respectively; φ t (·) is the feature extractor of the new incremental module; γ1, γ2 are scaling factors, and 0<γ1<1, 0<γ2<1, γ is the diagonal matrix composed of γ1 and γ2; the diagonal matrix composed of γ1 and γ2 is expressed as: ; Among them, S old is the number of old category samples in the total electromagnetic wave dataset after the current incremental batch, S new It is the number of new category samples in the total electromagnetic wave dataset after the current incremental batch.

8. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 7, characterized in that: The current incremental batch prediction model F t The loss function of (x) The calculation formula is: ; Among them, α1, α2, α3 are hyperparameters, and α1+α2+α3=1, is the first loss function, is the second loss function, is the distillation loss function, y is the true value of the electromagnetic wave category, is the cross entropy loss function, S(.) is the softmax, i.e. the normalized exponential function, W t fe is a new linear classifier parameter.

9. A time-space dual-dimensional low-altitude thunderstorm prediction method based on incremental learning as claimed in claim 8, characterized in that: The acquisition of the current incremental batch prediction model includes: d1: The current incremental batch prediction model F t (x) is set as the teacher model, and the student model is set at the tth incremental batch stage ; Wherein, the feature extractor of the student model is also composed of a CNN spatial feature extractor and a transformer temporal feature extractor; d2: Using the second distillation loss function Model for students Conduct training; d3: Use the trained student model as the final current incremental batch prediction model; The student model It is expressed as: ; in, The student model for the tth incremental batch stage Linear classifier parameters, , Student model for the current incremental batch stage The feature extractor of They are respectively the student models of the current incremental batch stage CNN spatial feature extractor and transformer temporal feature extractor, , , And d=d s +d TE , x s 、x TE are the normalized matrix and waveform time matrix of the electromagnetic wave image after spatial dimension and time dimension data preprocessing respectively; d s ,d TE After CNN spatial feature extractor and transformer temporal feature extractor The size of the feature dimension after extraction; Represents the total number of categories in the tth incremental batch stage.

Citation Information

Patent Citations

  • Continuous learning image classification method based on double-branch network

    CN116310484A

  • Unbalanced network intrusion detection method based on CNN-Transform fusion module

    CN118764270A