Hail weather identification and classification method based on multi-channel deep residual shrinkage network

GB2621908BActive Publication Date: 2025-06-11HOHAI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
GB2023005494
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-13
Filing Date
2022-12-09
Publication Date
2025-06-11
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing technologies require high signal-to-noise ratio when identifying and classifying hail weather, and are affected by various uncontrollable factors, making it difficult to achieve high-accuracy identification and classification of hail signals under low signal-to-noise ratio conditions.

Method used

A multi-channel deep residual shrinkage network combined with MSST and BN-SMOTE algorithms is used to perform time-frequency analysis and data set balancing, extract signal features, remove noise, and identify and classify hail signals through multi-scale feature extraction and soft thresholding.

Benefits of technology

Under the condition of low signal-to-noise ratio, the accuracy of identification and classification of hail signals is significantly improved, the impact of noise is reduced, and the anti-interference ability and identification ability of the network are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
Patent Text Reader

Abstract

A hail weather identification and classification method based on a multi-channel deep residual shrinkage network, comprising the following steps: acquiring strength data of microwave signals in hailin
Need to check novelty before this filing date? Find Prior Art

Description

Hail weather recognition and classification method based on multi-channel deep residual shrinkage network Technical Field

[0001] The present invention relates to a hail weather recognition and classification method based on a multi-channel deep residual shrinkage network, and belongs to the technical field of meteorological factor monitoring. Background Art

[0002] Hail is a solid precipitation in convective clouds and a common meteorological disaster. It is characterized by suddenness, great destructive power, and rapid movement. It seriously threatens the development of agriculture, electricity, communications, transportation, and other aspects, as well as the safety of human life and property. Therefore, it is of great practical significance to adopt real-time and effective methods to monitor and classify hail.

[0003] Current research on hail focuses on identifying hail clouds, primarily through monitoring and identification using lightning location systems, weather radar, and satellite remote sensing. Lightning location systems use lightning counters to measure lightning frequency and distinguish between thunderstorm clouds and hail clouds, but their deployment costs are high. Weather radar identifies hail clouds by analyzing the unique echo morphology, motion characteristics, and echo parameters, but is susceptible to interference from various factors and can result in large errors. Satellite remote sensing uses infrared cloud imagery to analyze cloud structure and characteristics, comparing them with actual hailfall data, to identify the cloud area of ​​the hail cloud. However, its ability to distinguish smaller affected areas is uncertain, and therefore its use has certain limitations.

[0004] Microwave communication networks offer advantages such as wide coverage, low monitoring costs, small blind spots, stable and reliable operation, and high spatiotemporal resolution. Therefore, they are theoretically well-suited for identifying severe convective weather events such as hail. However, existing traditional machine learning methods typically require a high signal-to-noise ratio (SNR) in the input model data for accurate classification and identification. However, due to various uncontrollable factors, the signals obtained often contain significant noise, making it difficult for traditional models to accurately identify and classify hail signals under low SNR conditions.

[0005] Summary of the Invention

[0006] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides a hail weather recognition and classification method based on a multi-channel deep residual shrinkage network. It combines the use of MSST, BN-SMOTE and a multi-channel deep residual shrinkage network to perform time-frequency analysis, data set balancing and recognition and classification respectively, and can achieve accurate recognition and classification of hail-related microwave signals under low signal-to-noise ratio conditions.

[0007] Technical solution: To solve the above technical problems, the present invention provides a hail weather recognition and classification method based on a multi-channel deep residual shrinkage network, comprising the following steps:

[0008] S1: Obtain microwave signal strength data under hail and non-hail weather conditions and preprocess the data.

[0009] S2: Perform multiple synchronous compression transform (MSST) on the preprocessed data to extract shallow signal features, convert the signal into a two-dimensional time-frequency image, and resize the obtained time-frequency image.

[0010] S3: Construct training and test sets, and use the BN-SMOTE algorithm to oversample the hail sample data (minority class) in the training set to balance the sample data and expand the data set.

[0011] S4: Input the expanded training set into the multi-channel deep residual shrinkage network, perform multi-scale extraction of deep features, remove noise, and output classification results.

[0012] S5: Train and optimize the model, and test the model performance on the test set.

[0013] S6: After the microwave signal data to be tested is processed by MSST, it is input into the trained model to realize hail weather recognition and grade classification.

[0014] Furthermore, the data preprocessing in step S1 specifically includes:

[0015] Missing data were interpolated, and unreasonable data that obviously exceeded the response threshold were eliminated.

[0016] Furthermore, the specific steps of performing the multiple synchronous compression transform (MSST) on the pre-processed data in step S2 include:

[0017] The expression of the selected signal s(u) is:

[0018]

[0019] Where A(t) is the signal amplitude, is the first-order Taylor series expansion of the phase.

[0020] S2-1: Perform STFT on the signal s(u). The time-frequency distribution can be expressed as:

[0021]

[0022] Where ω is the angular frequency and g(·) is the window function.

[0023] Find the partial derivative of the above formula:

[0024]

[0025] When G(t,ω)≠0, the instantaneous frequency estimate It can be expressed as:

[0026]

[0027] S2-2: Perform synchronous compression processing (SST) to compress the STFT result in the frequency direction. Its mathematical expression is as follows:

[0028]

[0029] Where δ(·) is the impulse function and η is the SST output frequency.

[0030] S2-3: Continue to perform SST n times on the obtained time-frequency distribution, then we have:

[0031]

[0032] Wherein, n≥2 is the number of times the synchronous compression process is performed, and n=2 is taken here.

[0033] Through multiple iterations, the instantaneous frequency estimate approaches the true value of the signal, the energy concentration of the time-frequency distribution is improved, and a high-resolution time-frequency image is obtained.

[0034] The time-frequency image obtained above is resized to obtain an image with a size of 224*224, which meets the requirements of network input.

[0035] Furthermore, the specific steps of oversampling the hail sample data in the training set using the BN-SMOTE algorithm in step S3 include:

[0036] We set four different labels: no hail (label 0), light hail (label 1), moderate hail (label 2), and heavy hail (label 3). We used stratified sampling to divide the training and test sets into an 8:2 ratio, ensuring that both sets contain samples of all label categories.

[0037] Define S min is the minority class sample set, including all samples under the labels of light hail, moderate hail, and heavy hail; S max is the majority class sample set, that is, all samples without hail labels; D is the number of new samples to be generated; k1 is the k-nearest neighbor value used to filter minority class samples; k2 is the number of majority class nearest neighbor samples used to generate the majority class set; k3 is the number of minority class nearest neighbor samples used to generate the minority class set.

[0038] S3-1: For each minority class sample ri ∈S min , calculate its nearest neighbor set NN(r i ), where NN(r i ) contains the same i The k1 samples with the closest Euclidean distance.

[0039] Eliminate minority class samples that do not have other minority classes in their k1 nearest neighbors to form a filtered minority class sample set S minf :

[0040] S minf =S min -{r i ∈S min :NN(r i ) has no minority class}

[0041] S3-2: For each minority class sample r i ∈S minf , calculate the majority class sample set N of its nearest neighbors maj (r i ), the set contains the i The k2 majority class samples with the closest Euclidean distance.

[0042] All N maj (r i ) sets are merged to obtain the majority class sample set in the boundary area:

[0043]

[0044] S3-3: For each majority class sample r′ i ∈S bmaj , calculate the minority class sample set N of its nearest neighbors min (r′ i ), the set contains the same i The k3 minority class samples with the closest Euclidean distance.

[0045] For all obtained N min (r′ i ) The minority class samples are combined to obtain the minority class sample set S which is the most difficult to learn in the boundary area. imin :

[0046]

[0047] S3-4: Initialize the set so that S omin =S min .

[0048] From the minority class sample set S iminSelect a sample m1 from the set, and then randomly select another sample m2 to generate a new sample s: s = m1 + α1 × (m2 - m1), where α1 is a random number in [0, 1], and put s into the set S omin Middle: Let S omin =S omin ∪{s}, repeat the above operation D times, end the loop, and output the minority class sample set S after oversampling omin , and add it to the training set to obtain a new training set after oversampling.

[0049] Furthermore, the specific steps of inputting the multi-channel deep residual shrinkage network for feature extraction and classification in step S4 include:

[0050] S4-1: Construct a multi-channel convolutional structure to achieve multi-scale feature extraction and fusion.

[0051] The convolution module consists of four channels with different structures. Channel 1 includes three convolution layers: the first layer uses a convolution kernel of size 1*1, and the second and third layers both use a convolution kernel of size 3*3 (two 3*3 convolution kernels are equivalent to the effect of a 5*5 convolution kernel); Channel 2 includes two convolution layers: the first layer has a convolution kernel size of 1*1, and the second layer has a convolution kernel size of 3*3; Channel 3 includes one convolution layer: the convolution kernel size is 1*1; Channel 4 includes two layers: the first layer is a maximum pooling layer, and the second layer is a convolution layer with a convolution kernel size of 1*1.

[0052] Among them, the first layer of the first three channels and the second layer of channel 4 all use 1*1 convolution kernels, which can achieve dimensionality reduction and increase network depth; the second and third layers of channel 1 use two 3*3 convolution kernels instead of one 5*5 convolution kernel, which not only greatly reduces the amount of calculation, but also increases the network depth, making it easier to extract deeper features; the equivalent convolution kernel size of the second and third layers of channel 1 is 5*5, and the convolution kernel size of the second layer of channel 2 is 3*3. The use of convolution kernels of different sizes adds more complex linear changes and realizes more representative multi-scale feature extraction.

[0053] The ReLu activation function is used after each convolutional layer to increase the nonlinearity and improve the expressiveness of the neural network. To prevent gradient vanishing and accelerate network convergence, a batch normalization (BN) layer is used after each branch.

[0054] Finally, the features extracted by the three branches are fused through the Concatenate layer to aggregate the features with strong correlation and weaken the irrelevant non-critical features.

[0055] S4-2: Input the residual shrinkage module, denoise it through soft thresholding, and further extract effective features.

[0056] The soft threshold function is expressed as follows:

[0057]

[0058] Among them, x is the input feature, y is the output feature, and τ is the threshold.

[0059] The sub-network embedded in the residual shrinkage module can adaptively generate the threshold and ensure that the threshold is positive and not too large.

[0060] The derivative of the soft thresholding output with respect to the input is as follows:

[0061]

[0062] From the above formula, we can see that the derivative of the output with respect to the input is either 0 or 1, which can effectively prevent the gradient from disappearing and exploding.

[0063] S4-3: The extracted high-dimensional features are reduced in dimension through global average pooling (GAP), which greatly reduces the training parameters and avoids overfitting. Finally, the classification results are output through the fully connected layer.

[0064] Finally, a fully connected layer is connected together with Softmax to convert the output of the previous layer into a probability distribution. The output with the highest probability is the current classification result. The Softmax expression is as follows:

[0065]

[0066] Among them, y′ i is the output of the previous layer, P Softmax is the probability of the corresponding hail type, k h =4 is the total number of hail types.

[0067] Furthermore, the specific steps of training, optimizing and testing the model in step S5 include:

[0068] S5-1: Select the Ranked List Loss and Cross Entropy Loss functions in metric learning to jointly guide the optimization network and adjust the parameters:

[0069] definition is the set of all samples, where N is the total number of all samples, (a i ,b i ) is the i-th sample and its corresponding category label, b i ∈[1,2,…,C], C is the total number of categories; are all samples contained in the cth class, where N c is the total number of samples in category c.

[0070] The joint classification loss function expression is as follows:

[0071]

[0072] Where f is the embedding function, λ is the weight of the cross entropy loss function, and L RLL is the Ranked List Loss in metric learning, L CE is the cross entropy loss function.

[0073] S5-2: Adam is used to minimize the loss function. The calculation process is as follows:

[0074] m t =m t-1 β1+(1-β1)g t

[0075]

[0076]

[0077]

[0078] Among them, g t is the gradient of the loss function, m t and v t are the biased first-order moment estimate and second-order moment estimate updated at the t-th iteration, and They are the biased first-order moment estimate and second-order moment estimate for the t-th iteration update, α is the learning rate, β1 and β2 are 0.9 and 0.999 respectively, ε prevents the divisor from being 0, and θ t The updated network parameters for the tth iteration.

[0079] S5-3: The test set is fed into the network to test model performance, using overall accuracy (OA) and the Kappa coefficient as evaluation metrics. OA is the ratio of the number of correctly predicted samples on the test set to the total number of samples in the test set, directly reflecting the proportion of correct classifications. The Kappa coefficient assesses the bias of the model; stronger bias results in lower Kappa values, further measuring classification effectiveness.

[0080] Beneficial effects: The hail weather recognition and classification method based on the multi-channel deep residual shrinkage network of the present invention has the following advantages:

[0081] 1. The use of multiple synchronous compression transform (MSST) for time-frequency analysis greatly reduces the computational burden and eliminates the problem of cross terms, effectively improving the aggregation of time-frequency spectrum and obtaining high-resolution time-frequency images.

[0082] 2. The BN-SMOTE algorithm was used to identify minority hail samples that are difficult to learn, expand the minority sample set, balance the ratio of positive and negative samples in the training set, and prevent the situation where the trained classifier cannot effectively identify hail weather (minority class) due to class imbalance. The classification accuracy of hail weather (minority class) in unbalanced datasets was significantly improved.

[0083] 3. A multi-channel deep residual shrinkage network was constructed, which greatly improved the accuracy of recognition and classification of noisy microwave signals.

[0084] 4. In the constructed network: the multi-channel convolutional structure realizes multi-scale deep feature extraction; the residual shrinkage module enhances the network's ability to extract useful features from noisy signals and remove noise, reduces the difficulty of network training, and effectively prevents the gradient explosion problem.

[0085] 5. Ranked List Loss and cross entropy loss functions are used to jointly guide network training, which maximizes the retention of intra-class features of samples while focusing on the overall distribution of samples, thereby improving the accuracy of the model in identifying hail weather. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] FIG1 is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0087] The present invention will be further described below with reference to the accompanying drawings.

[0088] As shown in Figure 1, a hail weather recognition and classification method based on a multi-channel deep residual shrinkage network includes the following steps:

[0089] S1: Acquire microwave signal strength data in hail and non-hail weather conditions and pre-process the data, including interpolating missing data and removing unreasonable data that clearly exceeds the response threshold;

[0090] S2: Perform multiple synchronous compression transform (MSST) on the preprocessed data to extract shallow signal features, convert the signal into a two-dimensional time-frequency image, and resize the obtained time-frequency image;

[0091] The specific steps of performing the multiple synchronous compression transform (MSST) on the pre-processed data in step S2 include:

[0092] The expression of the selected signal s(u) is:

[0093]

[0094] Where A(t) is the signal amplitude, is the first-order Taylor series expansion of the phase.

[0095] Perform short-time Fourier transform (STFT) on the signal s(u), and the time-frequency distribution can be expressed as:

[0096]

[0097] Where ω is the angular frequency and g(·) is the window function.

[0098] Find the partial derivative of the above formula:

[0099]

[0100] When G(t,ω)≠0, the instantaneous frequency estimate It can be expressed as:

[0101]

[0102] Then perform synchronous compression processing (SST), the mathematical expression is as follows:

[0103]

[0104] Where δ(·) is the impulse function and η is the SST output frequency.

[0105] By performing SST, the STFT result can be compressed in the frequency direction, thereby improving the energy concentration of the time-frequency spectrum. If SST is continued n times on the obtained time-frequency distribution, we have:

[0106]

[0107] Wherein, n≥2 is the number of times the synchronous compression process is performed, and n=2 is taken here.

[0108] Through multiple iterations, the instantaneous frequency estimate approaches the true value of the signal, the energy concentration of the time-frequency distribution is improved, and a high-resolution time-frequency image is obtained.

[0109] The time-frequency image obtained above is resized to obtain an image with a size of 224*224, which meets the requirements of network input.

[0110] S3: Construct training and test sets, and use the BN-SMOTE algorithm to oversample the hail sample data (minority class) in the training set to balance the sample data and expand the data set.

[0111] The step S3 of constructing the training set and the test set specifically includes:

[0112] We set four different labels: no hail (label 0), light hail (label 1), moderate hail (label 2), and heavy hail (label 3). We used stratified sampling to divide the training and test sets into an 8:2 ratio, ensuring that both sets contain samples of all label categories.

[0113] The specific steps of oversampling the hail sample data in the training set using the BN-SMOTE algorithm in step S3 include:

[0114] Define S min is the minority class sample set, including all samples under the labels of light hail, moderate hail, and heavy hail; S max is the majority class sample set, that is, all samples without hail labels; D is the number of new samples to be generated; k1 is the k-nearest neighbor value used to filter minority class samples; k2 is the number of majority class nearest neighbor samples used to generate the majority class set; k3 is the number of minority class nearest neighbor samples used to generate the minority class set.

[0115] S3-1: For each minority class sample r i ∈S min , calculate its nearest neighbor set NN(r i ), where NN(r i ) contains the same i The k1 samples with the closest Euclidean distance.

[0116] Eliminate minority class samples that do not have other minority classes in their k1 nearest neighbors to form a filtered minority class sample set S minf :

[0117] S minf =S min -{r i ∈S min :NN(r i ) has no minority class}

[0118] S3-2: For each minority class sample r i ∈S minf , calculate the majority class sample set N of its nearest neighbors maj (r i ), the set contains the i The k2 majority class samples with the closest Euclidean distance.

[0119] All N maj (r i ) sets are merged to obtain the majority class sample set in the boundary area.

[0120]

[0121] S3-3: For each majority class sample r′ i ∈S bmaj , calculate the minority class sample set N of its nearest neighbors min (r′ i ), the set contains the same i The k3 minority class samples with the closest Euclidean distance.

[0122] For all obtained N min (r′ i ) The minority class samples are combined to obtain the minority class sample set S which is the most difficult to learn in the boundary area. imin :

[0123]

[0124] S3-4: Initialize the set so that S omin =S min .

[0125] Do for j=1...D:

[0126] Step 1: From the minority class sample set S imin Select a sample m1 from the , and then randomly select another sample m2;

[0127] Step 2: Generate a new sample s: s = m1 + α1 × (m2 - m1), where α1 is a random number in [0, 1];

[0128] Step 3: Put s into set S omin Middle: Let S omin =S omin ∪{s}.

[0129] End the above cycle and output the oversampled minority class sample set S omin , add it to the training set to obtain a new training set after oversampling.

[0130] S4: Input the expanded training set into the multi-channel deep residual shrinkage network, perform multi-scale extraction of deep features, remove noise, and output classification results.

[0131] The specific steps of inputting the multi-channel deep residual shrinkage network for feature extraction and classification in step S4 include:

[0132] S4-1: Construct a multi-channel convolutional structure to achieve multi-scale feature extraction and fusion.

[0133] The convolution module consists of four channels with different structures. Channel 1 includes three convolution layers: the first layer uses a convolution kernel of size 1*1, and the second and third layers both use a convolution kernel of size 3*3 (two 3*3 convolution kernels are equivalent to the effect of a 5*5 convolution kernel); Channel 2 includes two convolution layers: the first layer has a convolution kernel size of 1*1, and the second layer has a convolution kernel size of 3*3; Channel 3 includes one convolution layer: the convolution kernel size is 1*1; Channel 4 includes two layers: the first layer is a maximum pooling layer, and the second layer is a convolution layer with a convolution kernel size of 1*1.

[0134] Among them, the first layer of the first three channels and the second layer of channel 4 all use 1*1 convolution kernels, which can achieve dimensionality reduction and increase network depth; the second and third layers of channel 1 use two 3*3 convolution kernels instead of one 5*5 convolution kernel, which not only greatly reduces the amount of calculation, but also increases the network depth, making it easier to extract deeper features; the equivalent convolution kernel size of the second and third layers of channel 1 is 5*5, and the convolution kernel size of the second layer of channel 2 is 3*3. The use of convolution kernels of different sizes adds more complex linear changes and realizes more representative multi-scale feature extraction.

[0135] The ReLu activation function is used after each convolutional layer to increase the nonlinear factor and improve the expressiveness of the neural network. To avoid gradient vanishing and speed up the network convergence, a batch normalization (BN) layer is used after each branch. The specific steps are as follows:

[0136] The first step is to calculate the mean of the batch data:

[0137]

[0138] The second step is to calculate the variance of the batch data:

[0139]

[0140] The third step is standardization:

[0141]

[0142] Step 4: Pan and zoom processing:

[0143]

[0144] Among them, x i and y i are the mini-batch input and output features of the i-th observation, γ and β are the scaling and translation factors, M is the number of batch samples, and ε prevents division by 0.

[0145] Finally, the features extracted by the three branches are fused through the Concatenate layer to aggregate the features with strong correlation and weaken the irrelevant non-critical features.

[0146] S4-2: Input residual shrinkage module, denoise through soft thresholding, and further extract effective features;

[0147] The soft threshold function is expressed as follows:

[0148]

[0149] Among them, x is the input feature, y is the output feature, and τ is the threshold.

[0150] The derivative of the soft thresholding output with respect to the input is as follows:

[0151]

[0152] From the above formula, we can see that the derivative of the output with respect to the input is either 0 or 1, which can effectively prevent the gradient from disappearing and exploding.

[0153] Among them, the sub-network embedded in the residual shrinkage module can adaptively generate the threshold and ensure that the threshold is positive and not too large, as follows:

[0154] First, the output of the last layer of the residual module is absolute-valued, and a one-dimensional vector with the same number of convolution kernels as the previous layer is obtained through global average pooling (GAP). Then, its parameters are scaled to (0, 1) through two layers of fully connected networks and activation functions. The formula is:

[0155]

[0156] Among them, z l is the feature of the lth neuron in the second layer of the fully connected network, α l is the corresponding scaling parameter, and its threshold is as follows:

[0157]

[0158] Among them, τ l is the threshold of the lth channel of the feature map, w and h are the width and height of the feature map respectively.

[0159] S4-3: The extracted high-dimensional features are reduced in dimension through global average pooling (GAP), which greatly reduces the training parameters and avoids overfitting. Finally, the classification results are output through the fully connected layer.

[0160] Finally, a fully connected layer is connected together with Softmax to convert the output of the previous layer into a probability distribution. The output with the highest probability is the current classification result. The Softmax expression is as follows:

[0161]

[0162] Among them, y′ i is the output of the previous layer, P Softmax is the probability of the corresponding hail type, k h =4 is the total number of hail types.

[0163] S5: Train and optimize the model, and test the model performance on the test set.

[0164] The specific steps of training, optimizing and testing the model in step S5 include:

[0165] S5-1: Select the Ranked List Loss and Cross Entropy Loss functions in metric learning to jointly guide the optimization network and adjust the parameters:

[0166] definition is the set of all samples, where N is the total number of all samples, (a i ,b i ) is the i-th sample and its corresponding category label, b i ∈[1,2,…,C], C is the total number of categories; are all samples contained in the cth class, where N c is the total number of samples in category c.

[0167] The pairwise margin loss is used as the basic pairwise constraint to construct a set-based similarity structure, which is expressed as follows:

[0168] L m (a i ,a j ; f) = (1-b ij )[α2-d ij ] + +b ij [d ij -(α2-m)] +

[0169] Among them, α2 is the distance parameter, f is the embedding function, and m is the distance margin between positive and negative samples. i =b j When b ij =1; otherwise, b ij =0.d ij =||f(a i )-f(a j )||2 represents the Euclidean distance difference between two samples.

[0170] The overall loss function expression is as follows:

[0171]

[0172] in, is the positive sample set, is the negative sample set, λ ij is the weight of the negative sample, and its expression is as follows:

[0173]

[0174] Among them, T′ is a hyperparameter.

[0175] Cross entropy loss function:

[0176]

[0177] Among them, N is the total number of training set samples, b i is the actual label, p i is the predicted label.

[0178] The joint classification loss function is expressed as follows:

[0179]

[0180] Where λ is the weight of the cross entropy loss function, which needs to be fine-tuned.

[0181] S5-2: Adam is used to minimize the loss function. The calculation process is as follows:

[0182] m t =m t-1 β1+(1-β1)g t

[0183]

[0184]

[0185]

[0186] Among them, g t is the gradient of the loss function, m t and v t are the biased first-order moment estimate and second-order moment estimate updated at the t-th iteration, and They are the biased first-order moment estimate and second-order moment estimate for the t-th iteration update, α is the learning rate, β1 and β2 are 0.9 and 0.999 respectively, ε prevents the divisor from being 0, and θ t The updated network parameters for the tth iteration.

[0187] S5-3: The test set is fed into the network to test model performance, using overall accuracy (OA) and the Kappa coefficient as evaluation metrics. OA is the ratio of the number of correctly predicted samples on the test set to the total number of samples in the test set, directly reflecting the proportion of correct classifications. The Kappa coefficient assesses the bias of the model; stronger bias results in lower Kappa values, further measuring classification effectiveness.

[0188] S6: After the microwave signal data to be tested is processed by MSST, it is input into the trained model to realize hail weather recognition and grade classification.

[0189] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A hail weather recognition and classification method based on a multi-channel deep residual shrinkage network, characterized in that: The steps include: S1: Obtain microwave signal strength data in hail and non-hail weather conditions and pre-process the data; S2: Perform multiple synchronous compression transform (MSST) on the preprocessed data to extract shallow signal features, convert the signal into a two-dimensional time-frequency image, and resize the obtained time-frequency image; S3: Construct training and test sets, and use the BN-SMOTE algorithm to oversample the hail sample data in the training set, balance the sample data, and expand the data set; S4: Input the expanded training set into the multi-channel deep residual shrinkage network to extract deep features at multiple scales, remove noise, and output the classification results; S5: Train and optimize the model, and test the model performance on the test set; S6: After the microwave signal data to be tested is processed by MSST, it is input into the trained model to realize hail weather recognition and grade classification.

2. The hail weather recognition and classification method based on a multi-channel deep residual shrinkage network according to claim 1 is characterized in that: The data preprocessing in step S1 specifically includes: interpolating missing data and eliminating unreasonable data that obviously exceeds the response threshold.

3. The hail weather recognition and classification method based on a multi-channel deep residual shrinkage network according to claim 1 is characterized in that: The specific steps of performing multiple synchronous compression transformation on the pre-processed data in step S2 include: The expression of the selected signal s(u) is: Where A(t) is the signal amplitude, is the first-order Taylor series expansion of the phase; S2-1: Perform STFT on the signal s(u). The time-frequency distribution can be expressed as: Where ω is the angular frequency and g(·) is the window function; Find the partial derivative of the above formula: When G(t,ω)≠0, the instantaneous frequency estimate It can be expressed as: S2-2: Perform synchronous compression processing SST to compress the STFT result in the frequency direction. Its mathematical expression is as follows: Where, δ(·) is the impulse function, and η is the SST output frequency; S2-3: Continue to perform SST n times on the obtained time-frequency distribution, then we have: Wherein, n≥2 is the number of times the synchronous compression process is performed. The size of the time-frequency image obtained above is adjusted to obtain an image with a size of 224*224, so as to meet the requirements of network input.

4. The hail weather recognition and classification method based on a multi-channel deep residual shrinkage network according to claim 1 is characterized in that: The specific steps of oversampling the hail sample data in the training set using the BN-SMOTE algorithm in step S3 include: Four different labels were set: no hail (label 0), light hail (label 1), moderate hail (label 2), and heavy hail (label 3). Stratified sampling was used to divide the training and test sets into an 8:2 ratio to ensure that both sets contained samples of all label types. Define S min is the minority class sample set, including all samples under the labels of light hail, moderate hail, and heavy hail; S max is the majority class sample set, that is, all samples without hail labels; D is the number of new samples to be generated; k1 is the k-nearest neighbor value used to filter minority class samples; k2 is the number of majority class nearest neighbor samples used to generate the majority class set; k3 is the number of minority class nearest neighbor samples used to generate the minority class set; S3-1: For each minority class sample r i ∈S min , calculate its nearest neighbor set NN(r i ), where NN(r i ) contains the same i The k1 samples with the closest Euclidean distance; Eliminate minority class samples that do not have other minority classes in their k1 nearest neighbors to form a filtered minority class sample set S min f : S min f =S min -{r i ∈S min :NN(r i ) has no minority class} S3-2: For each minority class sample r i ∈S min f , calculate the majority class sample set N of its nearest neighbors maj (r i ), the set contains the i k2 majority class samples with the closest Euclidean distance; All N maj (r i ) sets are merged to obtain the majority class sample set in the boundary area: S3-3: For each majority class sample r′ i ∈S bmaj , calculate the minority class sample set N of its nearest neighbors min (r′ i ), the set contains the same i k3 minority class samples with the closest Euclidean distance; For all obtained N min (r′ i ) The minority class samples are combined to obtain the minority class sample set S which is the most difficult to learn in the boundary area. i min : S3-4: Initialize the set so that S o min =S min ; From the minority class sample set S i min Select a sample m1 from the set, and then randomly select another sample m2 to generate a new sample s: s = m1 + α1 × (m2 - m1), where α1 is a random number in [0, 1], and put s into the set S o min Middle: Let S o min =S o min ∪{s}, repeat the above operation D times, end the loop, and output the minority class sample set S after oversampling o min , and add it to the training set to obtain a new training set after oversampling.

5. The hail recognition and classification method based on a multi-channel deep residual shrinkage network according to claim 1 is characterized in that: The specific steps of inputting the multi-channel deep residual shrinkage network for feature extraction and classification in step S4 include: S4-1: Construct a multi-channel convolutional structure to achieve multi-scale feature extraction and fusion: The convolution module consists of four channels with different structures. Channel 1 includes three convolution layers: the first layer uses a convolution kernel of size 1*1, and the second and third layers both use a convolution kernel of size 3*3 (two 3*3 convolution kernels are equivalent to the effect of a 5*5 convolution kernel); Channel 2 includes two convolution layers: the first layer has a convolution kernel size of 1*1, and the second layer has a convolution kernel size of 3*3; Channel 3 includes one convolution layer: the convolution kernel size is 1*1; Channel 4 includes two layers: the first layer is a maximum pooling layer, and the second layer is a convolution layer with a convolution kernel size of 1*1; The ReLu activation function is used after each convolution layer, and each branch is processed by a batch normalization BN layer; Finally, the features extracted by the three branches are fused through the Concatenate layer to aggregate the highly correlated features and weaken the irrelevant non-critical features; S4-2: Input residual shrinkage module, denoise through soft thresholding, and further extract effective features: The soft threshold function is expressed as follows: Among them, x is the input feature, y is the output feature, and τ is the threshold; The subnetwork embedded in the residual shrinkage module can adaptively generate thresholds, and the derivative of the soft threshold output with respect to the input is as follows: From the above formula, we can see that the derivative of the output with respect to the input is either 0 or 1; S4-3: Reduce the dimensionality of the extracted high-dimensional features through global average pooling, and finally output the classification results through the fully connected layer; Finally, a fully connected layer is connected together with Softmax to convert the output of the previous layer into a probability distribution. The output with the highest probability is the current classification result. The Softmax expression is as follows: Among them, y′ i is the output of the previous layer, P Soft max is the probability of the corresponding hail type, k h =4 is the total number of hail types.

6. The hail recognition and classification method based on a multi-channel deep residual shrinkage network according to claim 1 is characterized in that: The specific steps of training, optimizing and testing the model in step S5 include: S5-1: Select the Ranked List Loss and Cross Entropy Loss functions in metric learning to jointly guide the optimization network and adjust the parameters: definition is the set of all samples, where N is the total number of all samples, (a i ,b i ) is the i-th sample and its corresponding category label, b i ∈[1,2,…,C], C is the total number of categories, L is an ellipsis; are all samples contained in the cth class, where N c is the total number of samples in category c; The joint classification loss function expression is as follows: Where f is the embedding function, λ is the weight of the cross entropy loss function, and L RLL is the Ranked List Loss in metric learning, L CE is the cross entropy loss function; S5-2: Adam is used to minimize the loss function. The calculation process is as follows: m t =m t-1 β1+(1-β1)g t Among them, g t is the gradient of the loss function, m t and v t are the biased first-order moment estimate and second-order moment estimate updated at the t-th iteration, and They are the biased first-order moment estimate and second-order moment estimate for the t-th iteration update, α is the learning rate, β1 and β2 are 0.9 and 0.999 respectively, ε prevents the divisor from being 0, and θ t The network parameters updated for the tth iteration; S5-3: The test set is input into the network to test the model performance, and the overall accuracy OA and Kappa coefficient are used as evaluation indicators. Among them, OA is the ratio of the number of correct samples predicted by the model on the test set to the total number of samples in the test set; the Kappa coefficient gives a bias evaluation of the model.

Citation Information

Patent Citations

  • Hail detection algorithm on basis of meteorological radar data

    CN108802733A

  • Precipitation pattern identification method based on information of attenuation and polarization of microwave links

    CN109581546A

  • BP neural network-based microwave attenuation precipitation particle type identification method

    CN110543893A

  • Rainfall estimation method based on microwave rain attenuation and rainfall monitoring system

    CN111666656A

  • Hail weather identification and classification method based on multi-channel deep residual shrinkage network

    CN114755745A